From 688741ae4777ce9c92bdb4fe33fb98e24c0f020a Mon Sep 17 00:00:00 2001 From: "Jonathan D.A. Jewell" <6759885+hyperpolymath@users.noreply.github.com> Date: Mon, 3 Aug 2026 17:30:45 +0100 Subject: [PATCH] chore: remove the extracted verisimdb tree MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Closes the second and larger of the two duplications. 1058 -> 346 tracked files. verisimdb/ is now a pointer README. Safe only because the recovery landed first: hyperpolymath/verisimdb#219 is merged, and all eight files it carried are confirmed live on that repo's main before anything was deleted here. == This was a fork, not a duplicate lithoglyph/ was removed after verifying every file byte-identical to its own repo. That check FAILED here: only in this directory 108 only upstream 175 common 608, of which 333 differ (331 substantive) Deleting on the lithoglyph precedent would have destroyed ~4,000 lines including an entire tier of test coverage. The transferable lesson is the check, not the deletion. == Recovered before removal (verisimdb#219) Seven files had no upstream counterpart of any kind — not moved, not renamed, not converted — and none was a stub: concurrency_test.exs 444 security_test.exs 436 e2e_verisimdb_test.exs 410 kraft_property_test.exs 328 groove.rs 890 a2ml.rs 360 ram_promotion.rs 472 Plus WHITEPAPER.pdf. The four Elixir suites run against upstream unmodified (5 properties, 43 tests, 0 failures); the Rust modules are declared in their crates' lib.rs and compile (cargo check exit 0), verified by negative control. storage_regenerator.rs (682 lines) was deliberately NOT ported: upstream already has verisim-normalizer/src/regeneration.rs and verisim-api/src/regenerator.rs. Superseded, not lost. The other 100 were triaged as moved (16), converted (26), or superseded prose, config and retired ReScript. == Consequence for the anti-pattern gate .res files: 44 -> 4 Only typeql-experimental/ remains. The gate has been red on ReScript since before this branch; it is now within reach of green for the first time, on the last legacy directory that still carries any. .ts/.tsx 0, package.json 0 — unchanged, both already clear. == Guards and docs updated placement-guard: verisimdb removed from GRANDFATHER, so a new file under verisimdb/ now FAILS rather than warns — same treatment lithoglyph/ got on 2026-07-27. Warning would let the duplication grow back. CLAUDE.md and REGISTRY.adoc: both extractions recorded as done. Two stale REGISTRY rows corrected while there — Lithoglyph and Glyphbase were still listed "to create" when both have existed for over a week, and Glyphbase was described as the "web UI" rather than Airtable-mode delivery. Both files now carry the warning that verisimdb was a fork, so the next extraction is not assumed to be a clean subset. == Nothing is lost $ git ls-tree -r split-history/verisimdb --name-only | wc -l 713 Branch _split_verisimdb and tag split-history/verisimdb, both on origin, 124 commits. Never prune them, nor their lithoglyph/glyphbase/gnpl equivalents. a2ml and k9 validators pass; all workflows parse. --- .github/workflows/placement-guard.yml | 13 +- CLAUDE.md | 18 +- REGISTRY.adoc | 21 +- verisimdb/.cfignore | 5 - verisimdb/.claude/CLAUDE.md | 440 -- verisimdb/.clusterfuzzlite/project.yaml | 6 - verisimdb/.editorconfig | 68 - verisimdb/.gitattributes | 54 - verisimdb/.github/CODEOWNERS | 10 - verisimdb/.github/FUNDING.yml | 7 - .../.github/ISSUE_TEMPLATE/bug_report.md | 38 - verisimdb/.github/ISSUE_TEMPLATE/custom.md | 10 - .../.github/ISSUE_TEMPLATE/documentation.md | 66 - .../.github/ISSUE_TEMPLATE/feature_request.md | 20 - verisimdb/.github/ISSUE_TEMPLATE/question.md | 55 - verisimdb/.github/SUPPORT.md | 30 - verisimdb/.github/dependabot.yml | 48 - verisimdb/.github/workflows/cflite_batch.yml | 32 - verisimdb/.github/workflows/cflite_pr.yml | 36 - verisimdb/.github/workflows/codeql.yml | 40 - verisimdb/.github/workflows/governance.yml | 26 - verisimdb/.github/workflows/hypatia-scan.yml | 179 - verisimdb/.github/workflows/instant-sync.yml | 33 - .../.github/workflows/jekyll-gh-pages.yml | 52 - verisimdb/.github/workflows/mirror.yml | 144 - .../.github/workflows/scorecard-enforcer.yml | 72 - verisimdb/.github/workflows/scorecard.yml | 32 - .../.github/workflows/secret-scanner.yml | 67 - verisimdb/.github/workflows/security-scan.yml | 19 - verisimdb/.gitignore | 117 - verisimdb/.gitlab-ci.yml | 175 - verisimdb/.machine_readable/6a2/AGENTIC.a2ml | 34 - .../.machine_readable/6a2/ECOSYSTEM.a2ml | 20 - verisimdb/.machine_readable/6a2/META.a2ml | 27 - verisimdb/.machine_readable/6a2/NEUROSYM.a2ml | 21 - verisimdb/.machine_readable/6a2/PLAYBOOK.a2ml | 26 - verisimdb/.machine_readable/6a2/STATE.a2ml | 76 - .../.machine_readable/ENSAID_CONFIG.a2ml | 154 - verisimdb/.nojekyll | 0 verisimdb/.verisimdb/config.toml | 53 - verisimdb/.verisimdb/index.json | 12 - .../octads/commit-0087599f9fbd.json | 39 - .../octads/commit-030287996c2c.json | 39 - .../octads/commit-031ad82afba6.json | 39 - .../octads/commit-04ce420ca876.json | 39 - .../octads/commit-0565bb065c0b.json | 39 - .../octads/commit-05b5a1a2cbaa.json | 39 - .../octads/commit-0955ca7ada21.json | 39 - .../octads/commit-0ba909ba3a0b.json | 39 - .../octads/commit-0ca572654610.json | 39 - .../octads/commit-17f903f9fa44.json | 39 - .../octads/commit-1c5b28a1141e.json | 39 - .../octads/commit-208a70f82714.json | 39 - .../octads/commit-20ea93a1be96.json | 39 - .../octads/commit-228ae4bba134.json | 39 - .../octads/commit-23815332901c.json | 39 - .../octads/commit-242713b9c58c.json | 39 - .../octads/commit-2723614d8844.json | 39 - .../octads/commit-2eb26c6b7499.json | 39 - .../octads/commit-2ee7e397d41c.json | 39 - .../octads/commit-30228e2f8185.json | 39 - .../octads/commit-3327beaaa2b3.json | 39 - .../octads/commit-33c8f1c8d141.json | 39 - .../octads/commit-36907e025d4a.json | 39 - .../octads/commit-37139e0b0ba9.json | 39 - .../octads/commit-37dcacf06bf6.json | 39 - .../octads/commit-3beb018c4b99.json | 39 - .../octads/commit-3d7765bff845.json | 39 - .../octads/commit-3da4daef75f9.json | 39 - .../octads/commit-3df322443379.json | 39 - .../octads/commit-3dfad170ecfb.json | 39 - .../octads/commit-45e3a0230e23.json | 39 - .../octads/commit-4a0cbb016b03.json | 39 - .../octads/commit-4ed47fa921d0.json | 39 - .../octads/commit-50ad8033edc3.json | 39 - .../octads/commit-529b85ff223e.json | 39 - .../octads/commit-5427006827c3.json | 39 - .../octads/commit-57aba0bd8ff4.json | 39 - .../octads/commit-591b7006316c.json | 39 - .../octads/commit-5e782e158b79.json | 39 - .../octads/commit-5f6d0db4cc9d.json | 39 - .../octads/commit-63f2c2dec2ce.json | 39 - .../octads/commit-667f86eff63a.json | 39 - .../octads/commit-6d29ea7de601.json | 39 - .../octads/commit-71dc72521bef.json | 39 - .../octads/commit-737a6e822c52.json | 39 - .../octads/commit-74f45a7491e2.json | 39 - .../octads/commit-77d2e9f3f088.json | 39 - .../octads/commit-7b8c073708d5.json | 39 - .../octads/commit-7cec9f8d1b08.json | 39 - .../octads/commit-7de3adf6c8a0.json | 39 - .../octads/commit-8012d86a0882.json | 39 - .../octads/commit-854ea6d13ae7.json | 39 - .../octads/commit-86581c2638b8.json | 39 - .../octads/commit-89ea7188af80.json | 39 - .../octads/commit-8cd1f416878e.json | 39 - .../octads/commit-91a08d99ee55.json | 39 - .../octads/commit-9746771f7457.json | 39 - .../octads/commit-980d6c7a0dc2.json | 39 - .../octads/commit-9cce50aa9df3.json | 39 - .../octads/commit-9cdf85099304.json | 39 - .../octads/commit-9d353c5546f5.json | 39 - .../octads/commit-9e2298460f12.json | 39 - .../octads/commit-9f5c2f37c3f6.json | 39 - .../octads/commit-a0d832b5068b.json | 39 - .../octads/commit-a31e6e33e51f.json | 39 - .../octads/commit-a782219c1f2a.json | 39 - .../octads/commit-a9af6511d111.json | 39 - .../octads/commit-aa63510d99bd.json | 39 - .../octads/commit-b1038bf033e2.json | 39 - .../octads/commit-b62522fca6dc.json | 39 - .../octads/commit-b9d4aae8f5ce.json | 39 - .../octads/commit-ba1f53543438.json | 39 - .../octads/commit-bc2502d10f34.json | 39 - .../octads/commit-bcbfd32da8b9.json | 39 - .../octads/commit-bcd5e7cd743f.json | 39 - .../octads/commit-c0d8094e076f.json | 39 - .../octads/commit-ccb432c96a9d.json | 39 - .../octads/commit-cf2958acf17c.json | 39 - .../octads/commit-d39524c642db.json | 39 - .../octads/commit-d4e6f6be1200.json | 39 - .../octads/commit-d515ccac6a03.json | 39 - .../octads/commit-d6ac4d5f77c0.json | 39 - .../octads/commit-d7a13170c87c.json | 39 - .../octads/commit-d8174107cef6.json | 39 - .../octads/commit-d8593016e679.json | 39 - .../octads/commit-d949b42717bb.json | 39 - .../octads/commit-d961f137140e.json | 39 - .../octads/commit-dd8890c42dde.json | 39 - .../octads/commit-ddc83dd74b3d.json | 39 - .../octads/commit-de2fd383992a.json | 39 - .../octads/commit-dfe015d2ea26.json | 39 - .../octads/commit-e594e11c006e.json | 39 - .../octads/commit-e6dbf191a423.json | 39 - .../octads/commit-e7bd5a33403b.json | 39 - .../octads/commit-e84118929733.json | 39 - .../octads/commit-e9b738078138.json | 39 - .../octads/commit-ea9f52edc373.json | 39 - .../octads/commit-ecceaddc2f85.json | 39 - .../octads/commit-ee34606ba4e1.json | 39 - .../octads/commit-ef3bf59b298f.json | 39 - .../octads/commit-f3821bdec9bb.json | 39 - .../octads/commit-f72914d64062.json | 39 - .../octads/commit-f75fc83b4141.json | 39 - .../octads/commit-fb03b412d792.json | 39 - .../octads/commit-fbf2b307da8e.json | 39 - .../octads/commit-fc71351de8ee.json | 39 - .../octads/commit-fd3b385d4dd8.json | 39 - .../octads/commit-fd7ddf29cec9.json | 39 - verisimdb/.verisimdb/octads/issue-001.json | 37 - verisimdb/.verisimdb/octads/issue-002.json | 37 - verisimdb/.verisimdb/octads/issue-003.json | 37 - verisimdb/.verisimdb/octads/issue-004.json | 37 - verisimdb/.verisimdb/octads/issue-005.json | 37 - verisimdb/.verisimdb/octads/issue-006.json | 37 - verisimdb/.verisimdb/octads/issue-007.json | 37 - verisimdb/.verisimdb/octads/issue-008.json | 37 - verisimdb/.verisimdb/octads/issue-009.json | 37 - verisimdb/.verisimdb/octads/issue-010.json | 37 - verisimdb/.verisimdb/octads/issue-011.json | 37 - verisimdb/.verisimdb/octads/issue-012.json | 37 - verisimdb/.verisimdb/octads/issue-013.json | 37 - verisimdb/.verisimdb/octads/issue-014.json | 37 - verisimdb/.verisimdb/octads/issue-015.json | 37 - verisimdb/.verisimdb/octads/issue-016.json | 37 - verisimdb/.verisimdb/octads/issue-017.json | 37 - verisimdb/.verisimdb/octads/issue-018.json | 37 - verisimdb/.verisimdb/octads/issue-019.json | 37 - verisimdb/.verisimdb/octads/issue-020.json | 37 - verisimdb/.verisimdb/octads/issue-021.json | 37 - verisimdb/.verisimdb/octads/issue-022.json | 37 - verisimdb/.verisimdb/octads/issue-023.json | 37 - verisimdb/.well-known/groove/manifest.json | 39 - verisimdb/.well-known/void.rdf | 16 - verisimdb/.well-known/void.ttl | 75 - verisimdb/0-AI-MANIFEST.a2ml | 146 - verisimdb/ABI-FFI-README.md | 385 -- verisimdb/BEST-IN-CLASS-ROADMAP.md | 354 -- verisimdb/CHANGELOG.adoc | 93 - verisimdb/CODE_OF_CONDUCT.md | 327 -- verisimdb/CONTRIBUTING.md | 187 - verisimdb/Cargo.lock | 5029 ----------------- verisimdb/Cargo.toml | 126 - verisimdb/DEPLOYMENT.adoc | 793 --- verisimdb/EXPLAINME.adoc | 199 - verisimdb/IMPLEMENTATION-ROADMAP.adoc | 509 -- verisimdb/Justfile | 356 -- verisimdb/KNOWN-ISSUES.adoc | 250 - verisimdb/LICENSE | 153 - verisimdb/MAINTAINERS.adoc | 47 - verisimdb/PLANNER-IMPLEMENTATION-STATUS.md | 278 - verisimdb/QUICKSTART-USER.adoc | 185 - verisimdb/README.adoc | 514 +- verisimdb/README.adoc.invariants.md | 2 - verisimdb/RELEASE-NOTES-v0.1.0-alpha.md | 389 -- verisimdb/ROADMAP.adoc | 69 - verisimdb/RSR_OUTLINE.adoc | 218 - verisimdb/SECURITY.md | 375 -- verisimdb/SONNET-TASKS.md | 977 ---- verisimdb/TOPOLOGY.md | 113 - verisimdb/VOID-SETUP.md | 205 - verisimdb/WHITEPAPER.md | 385 -- verisimdb/WHITEPAPER.md.invariants.md | 54 - verisimdb/WHITEPAPER.pdf | Bin 68082 -> 0 bytes verisimdb/admin/deno.json | 24 - verisimdb/admin/gossamer.conf.json | 141 - verisimdb/admin/panels/manifest.json | 97 - verisimdb/admin/public/index.html | 15 - verisimdb/admin/rescript.json | 30 - verisimdb/admin/src/App.res | 904 --- verisimdb/admin/src/Capabilities.res | 123 - verisimdb/admin/src/Model.res | 164 - verisimdb/admin/src/Msg.res | 92 - verisimdb/admin/src/RuntimeBridge.res | 116 - verisimdb/admin/src/VeriSimDbCmd.res | 207 - verisimdb/admin/src/styles.css | 534 -- verisimdb/benches/Cargo.toml | 35 - verisimdb/benches/modality_benchmarks.rs | 624 -- verisimdb/benches/src/lib.rs | 2 - verisimdb/benches/throughput_benchmarks.rs | 369 -- verisimdb/connectors/README.adoc | 315 -- .../connectors/clients/elixir/.formatter.exs | 6 - .../connectors/clients/elixir/.gitignore | 22 - .../clients/elixir/lib/verisim_client.ex | 220 - .../elixir/lib/verisim_client/drift.ex | 90 - .../elixir/lib/verisim_client/error.ex | 103 - .../elixir/lib/verisim_client/federation.ex | 108 - .../elixir/lib/verisim_client/octad.ex | 116 - .../elixir/lib/verisim_client/provenance.ex | 84 - .../elixir/lib/verisim_client/search.ex | 165 - .../elixir/lib/verisim_client/types.ex | 291 - .../clients/elixir/lib/verisim_client/vcl.ex | 60 - verisimdb/connectors/clients/elixir/mix.exs | 52 - .../clients/elixir/test/test_helper.exs | 4 - .../elixir/test/verisim_client_test.exs | 76 - verisimdb/connectors/clients/gleam/gleam.toml | 18 - .../clients/gleam/src/verisimdb_client.gleam | 228 - .../gleam/src/verisimdb_client/codec.gleam | 858 --- .../gleam/src/verisimdb_client/drift.gleam | 92 - .../gleam/src/verisimdb_client/error.gleam | 144 - .../src/verisimdb_client/federation.gleam | 121 - .../gleam/src/verisimdb_client/octad.gleam | 136 - .../src/verisimdb_client/provenance.gleam | 100 - .../gleam/src/verisimdb_client/search.gleam | 226 - .../gleam/src/verisimdb_client/types.gleam | 363 -- .../gleam/src/verisimdb_client/vcl.gleam | 88 - .../gleam/test/verisimdb_client_test.gleam | 173 - .../connectors/clients/julia/Project.toml | 19 - .../clients/julia/src/VeriSimDBClient.jl | 72 - .../connectors/clients/julia/src/client.jl | 187 - .../connectors/clients/julia/src/drift.jl | 70 - .../connectors/clients/julia/src/error.jl | 196 - .../clients/julia/src/federation.jl | 127 - .../connectors/clients/julia/src/octad.jl | 114 - .../clients/julia/src/provenance.jl | 76 - .../connectors/clients/julia/src/search.jl | 163 - .../connectors/clients/julia/src/types.jl | 424 -- verisimdb/connectors/clients/julia/src/vcl.jl | 66 - .../connectors/clients/julia/test/runtests.jl | 126 - .../connectors/clients/rescript/.gitignore | 3 - .../connectors/clients/rescript/deno.json | 29 - .../connectors/clients/rescript/rescript.json | 7 - .../clients/rescript/src/VeriSimClient.res | 162 - .../clients/rescript/src/VeriSimDrift.res | 98 - .../clients/rescript/src/VeriSimError.res | 128 - .../rescript/src/VeriSimFederation.res | 106 - .../clients/rescript/src/VeriSimHexad.res | 147 - .../rescript/src/VeriSimProvenance.res | 104 - .../clients/rescript/src/VeriSimSearch.res | 232 - .../clients/rescript/src/VeriSimTypes.res | 324 -- .../clients/rescript/src/VeriSimVcl.res | 80 - verisimdb/connectors/clients/rust/Cargo.toml | 24 - .../connectors/clients/rust/src/client.rs | 272 - .../connectors/clients/rust/src/drift.rs | 63 - .../connectors/clients/rust/src/error.rs | 54 - .../connectors/clients/rust/src/federation.rs | 97 - verisimdb/connectors/clients/rust/src/lib.rs | 46 - .../connectors/clients/rust/src/octad.rs | 87 - .../connectors/clients/rust/src/provenance.rs | 68 - .../connectors/clients/rust/src/search.rs | 192 - .../connectors/clients/rust/src/types.rs | 392 -- verisimdb/connectors/clients/rust/src/vcl.rs | 72 - .../connectors/clients/vlang/MIGRATION.adoc | 47 - .../shared/json-schema/drift-score.json | 87 - .../connectors/shared/json-schema/error.json | 62 - .../shared/json-schema/federation-result.json | 40 - .../shared/json-schema/modality.json | 28 - .../shared/json-schema/octad-input.json | 215 - .../shared/json-schema/octad-status.json | 47 - .../connectors/shared/json-schema/octad.json | 230 - .../shared/json-schema/provenance-event.json | 59 - .../shared/json-schema/query-params.json | 94 - .../shared/openapi/verisim-api-v1.yaml | 1024 ---- .../shared/proto/verisim_federation.proto | 414 -- .../connectors/test-infra/.gatekeeper.yaml | 68 - .../connectors/test-infra/0-AI-MANIFEST.a2ml | 157 - verisimdb/connectors/test-infra/README.adoc | 344 -- verisimdb/connectors/test-infra/compose.toml | 152 - verisimdb/connectors/test-infra/ct-build.sh | 230 - verisimdb/connectors/test-infra/deploy.k9.ncl | 256 - .../images/Containerfile.clickhouse | 100 - .../test-infra/images/Containerfile.influxdb | 66 - .../test-infra/images/Containerfile.neo4j | 82 - .../images/Containerfile.redis-stack | 116 - .../test-infra/images/Containerfile.surrealdb | 63 - verisimdb/connectors/test-infra/manifest.toml | 136 - .../test-infra/seed/clickhouse-init.sql | 247 - .../test-infra/seed/influxdb-init.sh | 176 - .../connectors/test-infra/seed/minio-init.sh | 193 - .../test-infra/seed/mongodb-init.js | 364 -- .../test-infra/seed/neo4j-init.cypher | 205 - .../connectors/test-infra/seed/redis-init.sh | 286 - .../test-infra/seed/surrealdb-init.surql | 260 - verisimdb/connectors/test-infra/vordr.toml | 174 - verisimdb/container/.gatekeeper.yaml | 100 - verisimdb/container/Containerfile | 125 - verisimdb/container/compose.toml | 94 - verisimdb/container/ct-build.sh | 163 - verisimdb/container/entrypoint.sh | 36 - verisimdb/container/manifest.toml | 66 - verisimdb/contractiles/README.adoc | 19 - verisimdb/contractiles/dust/Dustfile | 1 - verisimdb/contractiles/must/Mustfile | 14 - verisimdb/contractiles/trust/Trustfile | 137 - verisimdb/debugger/.claude/CLAUDE.md | 126 - verisimdb/debugger/Cargo.toml | 54 - verisimdb/debugger/docs/CITATIONS.adoc | 36 - .../debugger/examples/SafeDOMExample.res | 109 - .../debugger/examples/web-project-deno.json | 20 - verisimdb/debugger/fuzz/Cargo.toml | 20 - .../debugger/fuzz/fuzz_targets/fuzz_main.rs | 22 - verisimdb/debugger/src/main.rs | 14 - verisimdb/debugger/tests/integration_test.rs | 7 - verisimdb/demos/drift-detection/run_demo.exs | 536 -- verisimdb/deny.toml | 37 - verisimdb/docs/.well-known/void.rdf | 16 - verisimdb/docs/.well-known/void.ttl | 75 - verisimdb/docs/CITATIONS.adoc | 36 - verisimdb/docs/VCL-SPEC.adoc | 3034 ---------- verisimdb/docs/adoption-strategy.adoc | 405 -- verisimdb/docs/backwards-compatibility.adoc | 669 --- verisimdb/docs/business/business-case.adoc | 588 -- .../business/financials/cost-structure.csv | 14 - .../business/financials/market-sizing.csv | 12 - .../business/financials/revenue-model.csv | 12 - .../business/financials/unit-economics.csv | 13 - .../marketing/feature-comparison.adoc | 446 -- .../docs/business/marketing/one-pager.adoc | 79 - .../marketing/pitch-deck-outline.adoc | 454 -- .../docs/business/marketing/use-cases.adoc | 522 -- verisimdb/docs/business/pr/faq.adoc | 316 -- verisimdb/docs/business/pr/key-messages.adoc | 347 -- verisimdb/docs/business/pr/press-release.adoc | 102 - .../business/strategy/adoption-roadmap.adoc | 546 -- .../business/strategy/economics-analysis.adoc | 441 -- .../business/strategy/functional-units.adoc | 362 -- .../docs/business/strategy/go-to-market.adoc | 447 -- verisimdb/docs/cache-sharing-strategy.adoc | 745 --- verisimdb/docs/caching-strategy.adoc | 594 -- verisimdb/docs/challenges-federated.adoc | 926 --- verisimdb/docs/challenges-hybrid.adoc | 782 --- verisimdb/docs/challenges-standalone.adoc | 439 -- .../consultation-dependent-types-zkp.adoc | 1166 ---- .../consultation-normalization-strategy.adoc | 920 --- verisimdb/docs/deployment-modes.adoc | 260 - .../DESIGN-2026-02-27-level-data-model.md | 138 - ...IGN-2026-02-27-strategic-improvements.adoc | 750 --- .../DESIGN-2026-02-27-vcl-dt-assessment.adoc | 572 -- ...SIGN-2026-02-28-panll-interop-telemetry.md | 178 - verisimdb/docs/drift-handling.adoc | 839 --- verisimdb/docs/error-handling-strategy.adoc | 1259 ----- verisimdb/docs/federation-readiness.adoc | 167 - verisimdb/docs/getting-started.adoc | 583 -- verisimdb/docs/minikanren-integration-v3.adoc | 329 -- verisimdb/docs/normalization-cascade.adoc | 762 --- verisimdb/docs/panll-module-audit.adoc | 325 -- .../verisimdb-federated-consistency.adoc | 840 --- .../papers/verisimdb-idaptik-case-study.adoc | 737 --- .../docs/query-optimization-overview.adoc | 483 -- verisimdb/docs/rescript-registry-types.adoc | 71 - verisimdb/docs/reversibility-design.adoc | 416 -- .../docs/safety-and-fault-tolerance.adoc | 1161 ---- verisimdb/docs/safety-theory-applied.adoc | 1336 ----- verisimdb/docs/security-lessons.lgt | 227 - .../snapshotting-and-truncation-logic.adoc | 77 - ...ical-specification-kraft-metadata-log.adoc | 33 - verisimdb/docs/vcl-architecture.adoc | 707 --- verisimdb/docs/vcl-examples.adoc | 1077 ---- verisimdb/docs/vcl-formal-semantics.adoc | 654 --- verisimdb/docs/vcl-grammar.ebnf | 316 -- verisimdb/docs/vcl-type-system.adoc | 922 --- verisimdb/docs/vcl-vs-sql.adoc | 240 - verisimdb/docs/vcl-vs-vcl-dt.adoc | 271 - .../docs/zkp-and-sanctify-integration.adoc | 65 - verisimdb/elixir-orchestration/.formatter.exs | 5 - .../elixir-orchestration/config/config.exs | 16 - verisimdb/elixir-orchestration/config/dev.exs | 10 - .../elixir-orchestration/config/prod.exs | 11 - .../elixir-orchestration/config/runtime.exs | 9 - .../elixir-orchestration/config/test.exs | 12 - verisimdb/elixir-orchestration/lib/verisim.ex | 67 - .../lib/verisim/api/router.ex | 167 - .../lib/verisim/application.ex | 65 - .../lib/verisim/consensus/kraft_node.ex | 874 --- .../lib/verisim/consensus/kraft_supervisor.ex | 48 - .../lib/verisim/consensus/kraft_transport.ex | 227 - .../lib/verisim/consensus/kraft_wal.ex | 410 -- .../lib/verisim/drift/drift_monitor.ex | 329 -- .../lib/verisim/entity/entity_server.ex | 252 - .../lib/verisim/federation/adapter.ex | 282 - .../verisim/federation/adapters/arangodb.ex | 305 - .../verisim/federation/adapters/clickhouse.ex | 339 -- .../lib/verisim/federation/adapters/duckdb.ex | 361 -- .../federation/adapters/elasticsearch.ex | 322 -- .../verisim/federation/adapters/influxdb.ex | 379 -- .../verisim/federation/adapters/mongodb.ex | 389 -- .../lib/verisim/federation/adapters/neo4j.ex | 427 -- .../federation/adapters/object_storage.ex | 461 -- .../verisim/federation/adapters/postgresql.ex | 431 -- .../lib/verisim/federation/adapters/redis.ex | 380 -- .../lib/verisim/federation/adapters/sqlite.ex | 331 -- .../verisim/federation/adapters/surrealdb.ex | 342 -- .../verisim/federation/adapters/vector_db.ex | 653 --- .../verisim/federation/adapters/verisimdb.ex | 231 - .../lib/verisim/federation/resolver.ex | 466 -- .../lib/verisim/health_checker.ex | 129 - .../lib/verisim/hypatia/dispatch_bridge.ex | 333 -- .../lib/verisim/hypatia/pattern_query.ex | 240 - .../lib/verisim/hypatia/scan_ingester.ex | 332 -- .../lib/verisim/nif_bridge.ex | 110 - .../lib/verisim/query/query_router.ex | 186 - .../lib/verisim/query/vcl_bridge.ex | 744 --- .../lib/verisim/query/vcl_executor.ex | 1812 ------ .../verisim/query/vcl_proof_certificate.ex | 190 - .../lib/verisim/query/vcl_type_checker.ex | 353 -- .../lib/verisim/rust_client.ex | 460 -- .../lib/verisim/schema/schema_registry.ex | 254 - .../lib/verisim/telemetry.ex | 230 - .../lib/verisim/telemetry/collector.ex | 343 -- .../lib/verisim/telemetry/reporter.ex | 350 -- .../lib/verisim/transport.ex | 195 - verisimdb/elixir-orchestration/mix.exs | 87 - .../fixtures/vcl/basic-select.expected.json | 9 - .../test/fixtures/vcl/basic-select.vcl | 1 - .../fixtures/vcl/drift-query.expected.json | 9 - .../test/fixtures/vcl/drift-query.vcl | 1 - .../fixtures/vcl/multi-proof.expected.json | 8 - .../test/fixtures/vcl/multi-proof.vcl | 1 - .../vcl/proof-existence.expected.json | 8 - .../test/fixtures/vcl/proof-existence.vcl | 1 - .../fixtures/vcl/vector-search.expected.json | 8 - .../test/fixtures/vcl/vector-search.vcl | 1 - .../test/integration_test.exs | 352 -- .../test/support/vcl_test_helpers.ex | 317 -- .../elixir-orchestration/test/test_helper.exs | 3 - .../test/verisim/api/router_test.exs | 287 - .../test/verisim/aspect/concurrency_test.exs | 444 -- .../test/verisim/aspect/security_test.exs | 436 -- .../verisim/consensus/kraft_node_test.exs | 302 - .../verisim/consensus/kraft_property_test.exs | 328 -- .../verisim/consensus/kraft_recovery_test.exs | 267 - .../consensus/kraft_transport_test.exs | 192 - .../test/verisim/consensus/kraft_wal_test.exs | 495 -- .../test/verisim/e2e_verisimdb_test.exs | 410 -- .../test/verisim/federation/adapter_test.exs | 428 -- .../federation/adapters/clickhouse_test.exs | 52 - .../federation/adapters/duckdb_test.exs | 58 - .../federation/adapters/influxdb_test.exs | 59 - .../clickhouse_integration_test.exs | 352 -- .../integration/influxdb_integration_test.exs | 374 -- .../integration/mongodb_integration_test.exs | 401 -- .../integration/neo4j_integration_test.exs | 359 -- .../object_storage_integration_test.exs | 381 -- .../integration/redis_integration_test.exs | 358 -- .../surrealdb_integration_test.exs | 373 -- .../federation/adapters/mongodb_test.exs | 93 - .../federation/adapters/neo4j_test.exs | 59 - .../adapters/object_storage_test.exs | 59 - .../federation/adapters/redis_test.exs | 76 - .../federation/adapters/sqlite_test.exs | 58 - .../federation/adapters/surrealdb_test.exs | 52 - .../federation/adapters/vector_db_test.exs | 59 - .../test/verisim/federation/resolver_test.exs | 105 - .../verisim/hypatia/dispatch_bridge_test.exs | 217 - .../verisim/hypatia/pattern_query_test.exs | 159 - .../verisim/hypatia/scan_ingester_test.exs | 258 - .../verisim/query/vcl_crossmodal_test.exs | 280 - .../verisim/query/vcl_dt_integration_test.exs | 422 -- .../test/verisim/query/vcl_dt_test.exs | 263 - .../test/verisim/query/vcl_e2e_test.exs | 613 -- .../verisim/query/vcl_integration_test.exs | 343 -- .../query/vcl_proof_certificate_test.exs | 160 - .../test/verisim/query/vcl_property_test.exs | 117 - .../test/verisim/query/vcl_test.exs | 273 - .../verisim/query/vcl_type_checker_test.exs | 417 -- .../test/verisim/telemetry_test.exs | 265 - .../test/verisim_test.exs | 10 - verisimdb/examples/README.adoc | 59 - verisimdb/examples/SafeDOMExample.res | 109 - verisimdb/examples/load-sample-data.sh | 168 - verisimdb/examples/sample-data/seed.json | 2210 -------- verisimdb/examples/smoke-test.sh | 170 - .../examples/vcl-queries/01-basic-search.vcl | 18 - .../vcl-queries/02-vector-similarity.vcl | 19 - .../vcl-queries/03-cross-modal-join.vcl | 21 - .../vcl-queries/04-drift-detection.vcl | 25 - .../vcl-queries/05-provenance-chain.vcl | 21 - .../vcl-queries/06-temporal-range.vcl | 21 - .../vcl-queries/07-spatial-radius.vcl | 23 - .../vcl-queries/08-proof-existence.vcl | 24 - .../vcl-queries/09-proof-consistency.vcl | 24 - .../vcl-queries/10-multi-modal-pipeline.vcl | 32 - verisimdb/examples/web-project-deno.json | 20 - verisimdb/ffi/zig/build.zig | 94 - verisimdb/ffi/zig/src/main.zig | 274 - verisimdb/ffi/zig/test/integration_test.zig | 182 - verisimdb/fuzz/Cargo.toml | 28 - verisimdb/fuzz/fuzz_targets/fuzz_octad_id.rs | 19 - verisimdb/historiographic-custodian.html | 258 - verisimdb/kraft-comparison.adoc | 69 - verisimdb/lib/verisim/adaptive_learner.ex | 516 -- verisimdb/lib/verisim/circuit_breaker.ex | 301 - verisimdb/lib/verisim/error_recovery.ex | 414 -- verisimdb/lib/verisim/query_cache.ex | 577 -- .../verisim/query_planner_bidirectional.ex | 331 -- verisimdb/lib/verisim/query_planner_config.ex | 274 - verisimdb/lib/verisim/query_router_cached.ex | 334 -- verisimdb/llm-warmup-dev.md | 153 - verisimdb/llm-warmup-user.md | 69 - verisimdb/opsm.toml | 15 - verisimdb/osv-scanner.toml | 4 - verisimdb/playground/build.mjs | 15 - verisimdb/playground/deno.json | 17 - verisimdb/playground/public/index.html | 99 - verisimdb/playground/public/manifest.json | 21 - verisimdb/playground/public/style.css | 370 -- verisimdb/playground/public/sw.js | 31 - verisimdb/playground/rescript.json | 18 - verisimdb/playground/src/ApiClient.res | 278 - verisimdb/playground/src/App.res | 349 -- verisimdb/playground/src/DemoExecutor.res | 98 - verisimdb/playground/src/Examples.res | 107 - verisimdb/playground/src/Formatter.res | 67 - verisimdb/playground/src/Highlighter.res | 83 - verisimdb/playground/src/Linter.res | 140 - verisimdb/playground/src/VclKeywords.res | 38 - verisimdb/proven-coherence.md | 258 - verisimdb/references.bib | 147 - verisimdb/rescript.json | 13 - verisimdb/rust-core/fuzz/Cargo.toml | 22 - .../fuzz/fuzz_targets/fuzz_vcl_parser.rs | 25 - verisimdb/rust-core/verisim-api/Cargo.toml | 60 - verisimdb/rust-core/verisim-api/build.rs | 15 - .../rust-core/verisim-api/proto/verisim.proto | 178 - verisimdb/rust-core/verisim-api/src/a2ml.rs | 360 -- verisimdb/rust-core/verisim-api/src/auth.rs | 778 --- .../rust-core/verisim-api/src/federation.rs | 660 --- .../rust-core/verisim-api/src/graphql.rs | 555 -- verisimdb/rust-core/verisim-api/src/groove.rs | 890 --- verisimdb/rust-core/verisim-api/src/grpc.rs | 444 -- verisimdb/rust-core/verisim-api/src/lib.rs | 2744 --------- verisimdb/rust-core/verisim-api/src/main.rs | 110 - .../verisim-api/src/proof_attempts.rs | 338 -- .../verisim-api/src/proto/verisim.rs | 1445 ----- verisimdb/rust-core/verisim-api/src/rbac.rs | 1084 ---- .../rust-core/verisim-api/src/transaction.rs | 466 -- verisimdb/rust-core/verisim-api/src/vcl.rs | 984 ---- .../rust-core/verisim-document/Cargo.toml | 21 - .../rust-core/verisim-document/src/lib.rs | 332 -- .../tests/property_tests.proptest-regressions | 7 - .../verisim-document/tests/property_tests.rs | 224 - verisimdb/rust-core/verisim-drift/Cargo.toml | 21 - .../rust-core/verisim-drift/src/calculator.rs | 650 --- verisimdb/rust-core/verisim-drift/src/lib.rs | 551 -- verisimdb/rust-core/verisim-graph/Cargo.toml | 34 - verisimdb/rust-core/verisim-graph/src/lib.rs | 427 -- .../verisim-graph/src/oxigraph_backend.rs | 170 - .../verisim-graph/src/redb_backend.rs | 602 -- verisimdb/rust-core/verisim-nif/Cargo.toml | 32 - verisimdb/rust-core/verisim-nif/src/lib.rs | 216 - .../rust-core/verisim-normalizer/Cargo.toml | 32 - .../verisim-normalizer/src/conflict.rs | 1561 ----- .../rust-core/verisim-normalizer/src/lib.rs | 835 --- .../verisim-normalizer/src/regeneration.rs | 1549 ----- .../src/storage_regenerator.rs | 682 --- verisimdb/rust-core/verisim-octad/Cargo.toml | 33 - verisimdb/rust-core/verisim-octad/src/lib.rs | 520 -- .../verisim-octad/src/query_octad.rs | 228 - .../verisim-octad/src/ram_promotion.rs | 472 -- .../rust-core/verisim-octad/src/store.rs | 1616 ------ .../verisim-octad/src/transaction.rs | 1446 ----- .../verisim-octad/tests/atomicity_tests.rs | 351 -- .../tests/crash_recovery_tests.rs | 160 - .../verisim-octad/tests/integration_tests.rs | 362 -- .../verisim-octad/tests/stress_tests.rs | 171 - .../rust-core/verisim-planner/Cargo.toml | 22 - .../rust-core/verisim-planner/src/config.rs | 169 - .../rust-core/verisim-planner/src/cost.rs | 770 --- .../rust-core/verisim-planner/src/error.rs | 23 - .../rust-core/verisim-planner/src/explain.rs | 368 -- .../rust-core/verisim-planner/src/lib.rs | 154 - .../verisim-planner/src/optimizer.rs | 358 -- .../rust-core/verisim-planner/src/plan.rs | 210 - .../rust-core/verisim-planner/src/prepared.rs | 1119 ---- .../rust-core/verisim-planner/src/profiler.rs | 870 --- .../verisim-planner/src/slow_query.rs | 550 -- .../rust-core/verisim-planner/src/stats.rs | 382 -- .../verisim-planner/src/vcl_bridge.rs | 1774 ------ .../rust-core/verisim-provenance/Cargo.toml | 28 - .../rust-core/verisim-provenance/src/lib.rs | 640 --- .../verisim-provenance/src/persistent.rs | 276 - verisimdb/rust-core/verisim-repl/Cargo.toml | 32 - .../rust-core/verisim-repl/src/client.rs | 152 - .../rust-core/verisim-repl/src/completer.rs | 178 - .../rust-core/verisim-repl/src/formatter.rs | 329 -- .../rust-core/verisim-repl/src/highlighter.rs | 238 - .../rust-core/verisim-repl/src/linter.rs | 517 -- verisimdb/rust-core/verisim-repl/src/main.rs | 515 -- .../rust-core/verisim-repl/src/vcl_fmt.rs | 459 -- .../rust-core/verisim-semantic/Cargo.toml | 31 - .../verisim-semantic/src/circuit_compiler.rs | 296 - .../verisim-semantic/src/circuit_registry.rs | 313 - .../rust-core/verisim-semantic/src/lib.rs | 399 -- .../verisim-semantic/src/persistent.rs | 312 - .../verisim-semantic/src/proven_bridge.rs | 283 - .../verisim-semantic/src/sanctify_bridge.rs | 413 -- .../verisim-semantic/src/verification_keys.rs | 246 - .../rust-core/verisim-semantic/src/zkp.rs | 367 -- .../verisim-semantic/src/zkp_bridge.rs | 628 -- .../rust-core/verisim-spatial/Cargo.toml | 26 - .../rust-core/verisim-spatial/src/lib.rs | 610 -- .../verisim-spatial/src/persistent.rs | 292 - .../rust-core/verisim-storage/Cargo.toml | 30 - .../rust-core/verisim-storage/src/backend.rs | 74 - .../rust-core/verisim-storage/src/error.rs | 104 - .../rust-core/verisim-storage/src/lib.rs | 61 - .../rust-core/verisim-storage/src/memory.rs | 282 - .../rust-core/verisim-storage/src/metrics.rs | 380 -- .../verisim-storage/src/redb_backend.rs | 513 -- .../rust-core/verisim-storage/src/typed.rs | 291 - .../rust-core/verisim-temporal/Cargo.toml | 27 - .../rust-core/verisim-temporal/src/diff.rs | 197 - .../rust-core/verisim-temporal/src/lib.rs | 390 -- .../verisim-temporal/src/persistent.rs | 261 - .../verisim-temporal/tests/property_tests.rs | 319 -- verisimdb/rust-core/verisim-tensor/Cargo.toml | 27 - verisimdb/rust-core/verisim-tensor/src/lib.rs | 330 -- .../verisim-tensor/src/persistent.rs | 123 - verisimdb/rust-core/verisim-vector/Cargo.toml | 34 - .../rust-core/verisim-vector/src/hnsw.rs | 669 --- verisimdb/rust-core/verisim-vector/src/lib.rs | 260 - .../verisim-vector/src/persistent.rs | 285 - verisimdb/rust-core/verisim-wal/Cargo.toml | 22 - verisimdb/rust-core/verisim-wal/src/entry.rs | 466 -- verisimdb/rust-core/verisim-wal/src/error.rs | 116 - verisimdb/rust-core/verisim-wal/src/lib.rs | 75 - verisimdb/rust-core/verisim-wal/src/reader.rs | 594 -- .../rust-core/verisim-wal/src/segment.rs | 291 - verisimdb/rust-core/verisim-wal/src/writer.rs | 472 -- verisimdb/scripts/post-commit-hook.sh | 8 - verisimdb/scripts/self-ingest.sh | 401 -- verisimdb/scripts/self-query.sh | 197 - verisimdb/scripts/smoke-test.sh | 149 - verisimdb/scripts/two-node-test.sh | 139 - verisimdb/selur-compose.yml | 262 - verisimdb/site/index.md | 21 - verisimdb/spec/README.adoc | 23 - verisimdb/spec/grammar.ebnf | 318 -- verisimdb/spec/system-specs.md | 179 - verisimdb/src/abi/Foreign.idr | 305 - verisimdb/src/abi/Layout.idr | 220 - verisimdb/src/abi/Types.idr | 348 -- verisimdb/src/registry/KRaftCluster.res | 549 -- verisimdb/src/registry/KRaftSerializer.res | 524 -- verisimdb/src/registry/MetadataLog.res | 375 -- verisimdb/src/registry/Registry.res | 868 --- verisimdb/src/vcl/VCLBidir.res | 852 --- verisimdb/src/vcl/VCLCircuit.res | 71 - verisimdb/src/vcl/VCLContext.res | 247 - verisimdb/src/vcl/VCLError.res | 458 -- verisimdb/src/vcl/VCLExplain.res | 436 -- verisimdb/src/vcl/VCLParser.res | 1195 ---- verisimdb/src/vcl/VCLParser_test.res | 558 -- verisimdb/src/vcl/VCLProofObligation.res | 251 - verisimdb/src/vcl/VCLSubtyping.res | 247 - verisimdb/src/vcl/VCLTypeChecker.res | 354 -- verisimdb/src/vcl/VCLTypes.res | 305 - verisimdb/stapeln.toml | 140 - verisimdb/tests/integration_test.rs | 344 -- verisimdb/vcl-bridge/vcl_parser_port.js | 594 -- verisimdb/verification/PROOF-STATUS.md | 55 - verisimdb/verification/README.adoc | 40 - verisimdb/verification/benchmarks | 1 - verisimdb/verification/fuzzing | 1 - .../proofs/agda/ProvenanceChain.agda | 245 - .../proofs/agda/ProvenanceChain.agdai | Bin 150906 -> 0 bytes .../proofs/idris2/ConnectorSafety.idr | 200 - .../proofs/idris2/DriftMetric.idr | 263 - .../proofs/idris2/FFIOwnership.idr | 113 - .../proofs/idris2/OctadCoherence.idr | 327 -- .../verification/proofs/lean4/RaftSafety.lean | 226 - .../proofs/lean4/VCLSubtyping.lean | 203 - .../proofs/lean4/VCLTypeSoundness.lean | 289 - .../proofs/lean4/WALIntegrity.lean | 202 - .../proofs/lean4/lake-manifest.json | 5 - .../verification/proofs/lean4/lakefile.lean | 9 - .../verification/proofs/tlaplus/.gitignore | 4 - .../proofs/tlaplus/Normalizer.cfg | 25 - .../proofs/tlaplus/Normalizer.tla | 139 - .../proofs/tlaplus/OctadAtomicity.cfg | 24 - .../proofs/tlaplus/OctadAtomicity.tla | 160 - .../verification/proofs/tlaplus/README.adoc | 108 - .../proofs/tlaplus/Serializability.cfg | 20 - .../proofs/tlaplus/Serializability.tla | 202 - verisimdb/verification/tests | 1 - .../verisim-architecture-visualisation.html | 230 - 716 files changed, 87 insertions(+), 165346 deletions(-) delete mode 100644 verisimdb/.cfignore delete mode 100644 verisimdb/.claude/CLAUDE.md delete mode 100644 verisimdb/.clusterfuzzlite/project.yaml delete mode 100644 verisimdb/.editorconfig delete mode 100644 verisimdb/.gitattributes delete mode 100644 verisimdb/.github/CODEOWNERS delete mode 100644 verisimdb/.github/FUNDING.yml delete mode 100644 verisimdb/.github/ISSUE_TEMPLATE/bug_report.md delete mode 100644 verisimdb/.github/ISSUE_TEMPLATE/custom.md delete mode 100644 verisimdb/.github/ISSUE_TEMPLATE/documentation.md delete mode 100644 verisimdb/.github/ISSUE_TEMPLATE/feature_request.md delete mode 100644 verisimdb/.github/ISSUE_TEMPLATE/question.md delete mode 100644 verisimdb/.github/SUPPORT.md delete mode 100644 verisimdb/.github/dependabot.yml delete mode 100644 verisimdb/.github/workflows/cflite_batch.yml delete mode 100644 verisimdb/.github/workflows/cflite_pr.yml delete mode 100644 verisimdb/.github/workflows/codeql.yml delete mode 100644 verisimdb/.github/workflows/governance.yml delete mode 100644 verisimdb/.github/workflows/hypatia-scan.yml delete mode 100644 verisimdb/.github/workflows/instant-sync.yml delete mode 100644 verisimdb/.github/workflows/jekyll-gh-pages.yml delete mode 100644 verisimdb/.github/workflows/mirror.yml delete mode 100644 verisimdb/.github/workflows/scorecard-enforcer.yml delete mode 100644 verisimdb/.github/workflows/scorecard.yml delete mode 100644 verisimdb/.github/workflows/secret-scanner.yml delete mode 100644 verisimdb/.github/workflows/security-scan.yml delete mode 100644 verisimdb/.gitignore delete mode 100644 verisimdb/.gitlab-ci.yml delete mode 100644 verisimdb/.machine_readable/6a2/AGENTIC.a2ml delete mode 100644 verisimdb/.machine_readable/6a2/ECOSYSTEM.a2ml delete mode 100644 verisimdb/.machine_readable/6a2/META.a2ml delete mode 100644 verisimdb/.machine_readable/6a2/NEUROSYM.a2ml delete mode 100644 verisimdb/.machine_readable/6a2/PLAYBOOK.a2ml delete mode 100644 verisimdb/.machine_readable/6a2/STATE.a2ml delete mode 100644 verisimdb/.machine_readable/ENSAID_CONFIG.a2ml delete mode 100644 verisimdb/.nojekyll delete mode 100644 verisimdb/.verisimdb/config.toml delete mode 100644 verisimdb/.verisimdb/index.json delete mode 100644 verisimdb/.verisimdb/octads/commit-0087599f9fbd.json delete mode 100644 verisimdb/.verisimdb/octads/commit-030287996c2c.json delete mode 100644 verisimdb/.verisimdb/octads/commit-031ad82afba6.json delete mode 100644 verisimdb/.verisimdb/octads/commit-04ce420ca876.json delete mode 100644 verisimdb/.verisimdb/octads/commit-0565bb065c0b.json delete mode 100644 verisimdb/.verisimdb/octads/commit-05b5a1a2cbaa.json delete mode 100644 verisimdb/.verisimdb/octads/commit-0955ca7ada21.json delete mode 100644 verisimdb/.verisimdb/octads/commit-0ba909ba3a0b.json delete mode 100644 verisimdb/.verisimdb/octads/commit-0ca572654610.json delete mode 100644 verisimdb/.verisimdb/octads/commit-17f903f9fa44.json delete mode 100644 verisimdb/.verisimdb/octads/commit-1c5b28a1141e.json delete mode 100644 verisimdb/.verisimdb/octads/commit-208a70f82714.json delete mode 100644 verisimdb/.verisimdb/octads/commit-20ea93a1be96.json delete mode 100644 verisimdb/.verisimdb/octads/commit-228ae4bba134.json delete mode 100644 verisimdb/.verisimdb/octads/commit-23815332901c.json delete mode 100644 verisimdb/.verisimdb/octads/commit-242713b9c58c.json delete mode 100644 verisimdb/.verisimdb/octads/commit-2723614d8844.json delete mode 100644 verisimdb/.verisimdb/octads/commit-2eb26c6b7499.json delete mode 100644 verisimdb/.verisimdb/octads/commit-2ee7e397d41c.json delete mode 100644 verisimdb/.verisimdb/octads/commit-30228e2f8185.json delete mode 100644 verisimdb/.verisimdb/octads/commit-3327beaaa2b3.json delete mode 100644 verisimdb/.verisimdb/octads/commit-33c8f1c8d141.json delete mode 100644 verisimdb/.verisimdb/octads/commit-36907e025d4a.json delete mode 100644 verisimdb/.verisimdb/octads/commit-37139e0b0ba9.json delete mode 100644 verisimdb/.verisimdb/octads/commit-37dcacf06bf6.json delete mode 100644 verisimdb/.verisimdb/octads/commit-3beb018c4b99.json delete mode 100644 verisimdb/.verisimdb/octads/commit-3d7765bff845.json delete mode 100644 verisimdb/.verisimdb/octads/commit-3da4daef75f9.json delete mode 100644 verisimdb/.verisimdb/octads/commit-3df322443379.json delete mode 100644 verisimdb/.verisimdb/octads/commit-3dfad170ecfb.json delete mode 100644 verisimdb/.verisimdb/octads/commit-45e3a0230e23.json delete mode 100644 verisimdb/.verisimdb/octads/commit-4a0cbb016b03.json delete mode 100644 verisimdb/.verisimdb/octads/commit-4ed47fa921d0.json delete mode 100644 verisimdb/.verisimdb/octads/commit-50ad8033edc3.json delete mode 100644 verisimdb/.verisimdb/octads/commit-529b85ff223e.json delete mode 100644 verisimdb/.verisimdb/octads/commit-5427006827c3.json delete mode 100644 verisimdb/.verisimdb/octads/commit-57aba0bd8ff4.json delete mode 100644 verisimdb/.verisimdb/octads/commit-591b7006316c.json delete mode 100644 verisimdb/.verisimdb/octads/commit-5e782e158b79.json delete mode 100644 verisimdb/.verisimdb/octads/commit-5f6d0db4cc9d.json delete mode 100644 verisimdb/.verisimdb/octads/commit-63f2c2dec2ce.json delete mode 100644 verisimdb/.verisimdb/octads/commit-667f86eff63a.json delete mode 100644 verisimdb/.verisimdb/octads/commit-6d29ea7de601.json delete mode 100644 verisimdb/.verisimdb/octads/commit-71dc72521bef.json delete mode 100644 verisimdb/.verisimdb/octads/commit-737a6e822c52.json delete mode 100644 verisimdb/.verisimdb/octads/commit-74f45a7491e2.json delete mode 100644 verisimdb/.verisimdb/octads/commit-77d2e9f3f088.json delete mode 100644 verisimdb/.verisimdb/octads/commit-7b8c073708d5.json delete mode 100644 verisimdb/.verisimdb/octads/commit-7cec9f8d1b08.json delete mode 100644 verisimdb/.verisimdb/octads/commit-7de3adf6c8a0.json delete mode 100644 verisimdb/.verisimdb/octads/commit-8012d86a0882.json delete mode 100644 verisimdb/.verisimdb/octads/commit-854ea6d13ae7.json delete mode 100644 verisimdb/.verisimdb/octads/commit-86581c2638b8.json delete mode 100644 verisimdb/.verisimdb/octads/commit-89ea7188af80.json delete mode 100644 verisimdb/.verisimdb/octads/commit-8cd1f416878e.json delete mode 100644 verisimdb/.verisimdb/octads/commit-91a08d99ee55.json delete mode 100644 verisimdb/.verisimdb/octads/commit-9746771f7457.json delete mode 100644 verisimdb/.verisimdb/octads/commit-980d6c7a0dc2.json delete mode 100644 verisimdb/.verisimdb/octads/commit-9cce50aa9df3.json delete mode 100644 verisimdb/.verisimdb/octads/commit-9cdf85099304.json delete mode 100644 verisimdb/.verisimdb/octads/commit-9d353c5546f5.json delete mode 100644 verisimdb/.verisimdb/octads/commit-9e2298460f12.json delete mode 100644 verisimdb/.verisimdb/octads/commit-9f5c2f37c3f6.json delete mode 100644 verisimdb/.verisimdb/octads/commit-a0d832b5068b.json delete mode 100644 verisimdb/.verisimdb/octads/commit-a31e6e33e51f.json delete mode 100644 verisimdb/.verisimdb/octads/commit-a782219c1f2a.json delete mode 100644 verisimdb/.verisimdb/octads/commit-a9af6511d111.json delete mode 100644 verisimdb/.verisimdb/octads/commit-aa63510d99bd.json delete mode 100644 verisimdb/.verisimdb/octads/commit-b1038bf033e2.json delete mode 100644 verisimdb/.verisimdb/octads/commit-b62522fca6dc.json delete mode 100644 verisimdb/.verisimdb/octads/commit-b9d4aae8f5ce.json delete mode 100644 verisimdb/.verisimdb/octads/commit-ba1f53543438.json delete mode 100644 verisimdb/.verisimdb/octads/commit-bc2502d10f34.json delete mode 100644 verisimdb/.verisimdb/octads/commit-bcbfd32da8b9.json delete mode 100644 verisimdb/.verisimdb/octads/commit-bcd5e7cd743f.json delete mode 100644 verisimdb/.verisimdb/octads/commit-c0d8094e076f.json delete mode 100644 verisimdb/.verisimdb/octads/commit-ccb432c96a9d.json delete mode 100644 verisimdb/.verisimdb/octads/commit-cf2958acf17c.json delete mode 100644 verisimdb/.verisimdb/octads/commit-d39524c642db.json delete mode 100644 verisimdb/.verisimdb/octads/commit-d4e6f6be1200.json delete mode 100644 verisimdb/.verisimdb/octads/commit-d515ccac6a03.json delete mode 100644 verisimdb/.verisimdb/octads/commit-d6ac4d5f77c0.json delete mode 100644 verisimdb/.verisimdb/octads/commit-d7a13170c87c.json delete mode 100644 verisimdb/.verisimdb/octads/commit-d8174107cef6.json delete mode 100644 verisimdb/.verisimdb/octads/commit-d8593016e679.json delete mode 100644 verisimdb/.verisimdb/octads/commit-d949b42717bb.json delete mode 100644 verisimdb/.verisimdb/octads/commit-d961f137140e.json delete mode 100644 verisimdb/.verisimdb/octads/commit-dd8890c42dde.json delete mode 100644 verisimdb/.verisimdb/octads/commit-ddc83dd74b3d.json delete mode 100644 verisimdb/.verisimdb/octads/commit-de2fd383992a.json delete mode 100644 verisimdb/.verisimdb/octads/commit-dfe015d2ea26.json delete mode 100644 verisimdb/.verisimdb/octads/commit-e594e11c006e.json delete mode 100644 verisimdb/.verisimdb/octads/commit-e6dbf191a423.json delete mode 100644 verisimdb/.verisimdb/octads/commit-e7bd5a33403b.json delete mode 100644 verisimdb/.verisimdb/octads/commit-e84118929733.json delete mode 100644 verisimdb/.verisimdb/octads/commit-e9b738078138.json delete mode 100644 verisimdb/.verisimdb/octads/commit-ea9f52edc373.json delete mode 100644 verisimdb/.verisimdb/octads/commit-ecceaddc2f85.json delete mode 100644 verisimdb/.verisimdb/octads/commit-ee34606ba4e1.json delete mode 100644 verisimdb/.verisimdb/octads/commit-ef3bf59b298f.json delete mode 100644 verisimdb/.verisimdb/octads/commit-f3821bdec9bb.json delete mode 100644 verisimdb/.verisimdb/octads/commit-f72914d64062.json delete mode 100644 verisimdb/.verisimdb/octads/commit-f75fc83b4141.json delete mode 100644 verisimdb/.verisimdb/octads/commit-fb03b412d792.json delete mode 100644 verisimdb/.verisimdb/octads/commit-fbf2b307da8e.json delete mode 100644 verisimdb/.verisimdb/octads/commit-fc71351de8ee.json delete mode 100644 verisimdb/.verisimdb/octads/commit-fd3b385d4dd8.json delete mode 100644 verisimdb/.verisimdb/octads/commit-fd7ddf29cec9.json delete mode 100644 verisimdb/.verisimdb/octads/issue-001.json delete mode 100644 verisimdb/.verisimdb/octads/issue-002.json delete mode 100644 verisimdb/.verisimdb/octads/issue-003.json delete mode 100644 verisimdb/.verisimdb/octads/issue-004.json delete mode 100644 verisimdb/.verisimdb/octads/issue-005.json delete mode 100644 verisimdb/.verisimdb/octads/issue-006.json delete mode 100644 verisimdb/.verisimdb/octads/issue-007.json delete mode 100644 verisimdb/.verisimdb/octads/issue-008.json delete mode 100644 verisimdb/.verisimdb/octads/issue-009.json delete mode 100644 verisimdb/.verisimdb/octads/issue-010.json delete mode 100644 verisimdb/.verisimdb/octads/issue-011.json delete mode 100644 verisimdb/.verisimdb/octads/issue-012.json delete mode 100644 verisimdb/.verisimdb/octads/issue-013.json delete mode 100644 verisimdb/.verisimdb/octads/issue-014.json delete mode 100644 verisimdb/.verisimdb/octads/issue-015.json delete mode 100644 verisimdb/.verisimdb/octads/issue-016.json delete mode 100644 verisimdb/.verisimdb/octads/issue-017.json delete mode 100644 verisimdb/.verisimdb/octads/issue-018.json delete mode 100644 verisimdb/.verisimdb/octads/issue-019.json delete mode 100644 verisimdb/.verisimdb/octads/issue-020.json delete mode 100644 verisimdb/.verisimdb/octads/issue-021.json delete mode 100644 verisimdb/.verisimdb/octads/issue-022.json delete mode 100644 verisimdb/.verisimdb/octads/issue-023.json delete mode 100644 verisimdb/.well-known/groove/manifest.json delete mode 100644 verisimdb/.well-known/void.rdf delete mode 100644 verisimdb/.well-known/void.ttl delete mode 100644 verisimdb/0-AI-MANIFEST.a2ml delete mode 100644 verisimdb/ABI-FFI-README.md delete mode 100644 verisimdb/BEST-IN-CLASS-ROADMAP.md delete mode 100644 verisimdb/CHANGELOG.adoc delete mode 100644 verisimdb/CODE_OF_CONDUCT.md delete mode 100644 verisimdb/CONTRIBUTING.md delete mode 100644 verisimdb/Cargo.lock delete mode 100644 verisimdb/Cargo.toml delete mode 100644 verisimdb/DEPLOYMENT.adoc delete mode 100644 verisimdb/EXPLAINME.adoc delete mode 100644 verisimdb/IMPLEMENTATION-ROADMAP.adoc delete mode 100644 verisimdb/Justfile delete mode 100644 verisimdb/KNOWN-ISSUES.adoc delete mode 100644 verisimdb/LICENSE delete mode 100644 verisimdb/MAINTAINERS.adoc delete mode 100644 verisimdb/PLANNER-IMPLEMENTATION-STATUS.md delete mode 100644 verisimdb/QUICKSTART-USER.adoc delete mode 100644 verisimdb/README.adoc.invariants.md delete mode 100644 verisimdb/RELEASE-NOTES-v0.1.0-alpha.md delete mode 100644 verisimdb/ROADMAP.adoc delete mode 100644 verisimdb/RSR_OUTLINE.adoc delete mode 100644 verisimdb/SECURITY.md delete mode 100644 verisimdb/SONNET-TASKS.md delete mode 100644 verisimdb/TOPOLOGY.md delete mode 100644 verisimdb/VOID-SETUP.md delete mode 100644 verisimdb/WHITEPAPER.md delete mode 100644 verisimdb/WHITEPAPER.md.invariants.md delete mode 100644 verisimdb/WHITEPAPER.pdf delete mode 100644 verisimdb/admin/deno.json delete mode 100644 verisimdb/admin/gossamer.conf.json delete mode 100644 verisimdb/admin/panels/manifest.json delete mode 100644 verisimdb/admin/public/index.html delete mode 100644 verisimdb/admin/rescript.json delete mode 100644 verisimdb/admin/src/App.res delete mode 100644 verisimdb/admin/src/Capabilities.res delete mode 100644 verisimdb/admin/src/Model.res delete mode 100644 verisimdb/admin/src/Msg.res delete mode 100644 verisimdb/admin/src/RuntimeBridge.res delete mode 100644 verisimdb/admin/src/VeriSimDbCmd.res delete mode 100644 verisimdb/admin/src/styles.css delete mode 100644 verisimdb/benches/Cargo.toml delete mode 100644 verisimdb/benches/modality_benchmarks.rs delete mode 100644 verisimdb/benches/src/lib.rs delete mode 100644 verisimdb/benches/throughput_benchmarks.rs delete mode 100644 verisimdb/connectors/README.adoc delete mode 100644 verisimdb/connectors/clients/elixir/.formatter.exs delete mode 100644 verisimdb/connectors/clients/elixir/.gitignore delete mode 100644 verisimdb/connectors/clients/elixir/lib/verisim_client.ex delete mode 100644 verisimdb/connectors/clients/elixir/lib/verisim_client/drift.ex delete mode 100644 verisimdb/connectors/clients/elixir/lib/verisim_client/error.ex delete mode 100644 verisimdb/connectors/clients/elixir/lib/verisim_client/federation.ex delete mode 100644 verisimdb/connectors/clients/elixir/lib/verisim_client/octad.ex delete mode 100644 verisimdb/connectors/clients/elixir/lib/verisim_client/provenance.ex delete mode 100644 verisimdb/connectors/clients/elixir/lib/verisim_client/search.ex delete mode 100644 verisimdb/connectors/clients/elixir/lib/verisim_client/types.ex delete mode 100644 verisimdb/connectors/clients/elixir/lib/verisim_client/vcl.ex delete mode 100644 verisimdb/connectors/clients/elixir/mix.exs delete mode 100644 verisimdb/connectors/clients/elixir/test/test_helper.exs delete mode 100644 verisimdb/connectors/clients/elixir/test/verisim_client_test.exs delete mode 100644 verisimdb/connectors/clients/gleam/gleam.toml delete mode 100644 verisimdb/connectors/clients/gleam/src/verisimdb_client.gleam delete mode 100644 verisimdb/connectors/clients/gleam/src/verisimdb_client/codec.gleam delete mode 100644 verisimdb/connectors/clients/gleam/src/verisimdb_client/drift.gleam delete mode 100644 verisimdb/connectors/clients/gleam/src/verisimdb_client/error.gleam delete mode 100644 verisimdb/connectors/clients/gleam/src/verisimdb_client/federation.gleam delete mode 100644 verisimdb/connectors/clients/gleam/src/verisimdb_client/octad.gleam delete mode 100644 verisimdb/connectors/clients/gleam/src/verisimdb_client/provenance.gleam delete mode 100644 verisimdb/connectors/clients/gleam/src/verisimdb_client/search.gleam delete mode 100644 verisimdb/connectors/clients/gleam/src/verisimdb_client/types.gleam delete mode 100644 verisimdb/connectors/clients/gleam/src/verisimdb_client/vcl.gleam delete mode 100644 verisimdb/connectors/clients/gleam/test/verisimdb_client_test.gleam delete mode 100644 verisimdb/connectors/clients/julia/Project.toml delete mode 100644 verisimdb/connectors/clients/julia/src/VeriSimDBClient.jl delete mode 100644 verisimdb/connectors/clients/julia/src/client.jl delete mode 100644 verisimdb/connectors/clients/julia/src/drift.jl delete mode 100644 verisimdb/connectors/clients/julia/src/error.jl delete mode 100644 verisimdb/connectors/clients/julia/src/federation.jl delete mode 100644 verisimdb/connectors/clients/julia/src/octad.jl delete mode 100644 verisimdb/connectors/clients/julia/src/provenance.jl delete mode 100644 verisimdb/connectors/clients/julia/src/search.jl delete mode 100644 verisimdb/connectors/clients/julia/src/types.jl delete mode 100644 verisimdb/connectors/clients/julia/src/vcl.jl delete mode 100644 verisimdb/connectors/clients/julia/test/runtests.jl delete mode 100644 verisimdb/connectors/clients/rescript/.gitignore delete mode 100644 verisimdb/connectors/clients/rescript/deno.json delete mode 100644 verisimdb/connectors/clients/rescript/rescript.json delete mode 100644 verisimdb/connectors/clients/rescript/src/VeriSimClient.res delete mode 100644 verisimdb/connectors/clients/rescript/src/VeriSimDrift.res delete mode 100644 verisimdb/connectors/clients/rescript/src/VeriSimError.res delete mode 100644 verisimdb/connectors/clients/rescript/src/VeriSimFederation.res delete mode 100644 verisimdb/connectors/clients/rescript/src/VeriSimHexad.res delete mode 100644 verisimdb/connectors/clients/rescript/src/VeriSimProvenance.res delete mode 100644 verisimdb/connectors/clients/rescript/src/VeriSimSearch.res delete mode 100644 verisimdb/connectors/clients/rescript/src/VeriSimTypes.res delete mode 100644 verisimdb/connectors/clients/rescript/src/VeriSimVcl.res delete mode 100644 verisimdb/connectors/clients/rust/Cargo.toml delete mode 100644 verisimdb/connectors/clients/rust/src/client.rs delete mode 100644 verisimdb/connectors/clients/rust/src/drift.rs delete mode 100644 verisimdb/connectors/clients/rust/src/error.rs delete mode 100644 verisimdb/connectors/clients/rust/src/federation.rs delete mode 100644 verisimdb/connectors/clients/rust/src/lib.rs delete mode 100644 verisimdb/connectors/clients/rust/src/octad.rs delete mode 100644 verisimdb/connectors/clients/rust/src/provenance.rs delete mode 100644 verisimdb/connectors/clients/rust/src/search.rs delete mode 100644 verisimdb/connectors/clients/rust/src/types.rs delete mode 100644 verisimdb/connectors/clients/rust/src/vcl.rs delete mode 100644 verisimdb/connectors/clients/vlang/MIGRATION.adoc delete mode 100644 verisimdb/connectors/shared/json-schema/drift-score.json delete mode 100644 verisimdb/connectors/shared/json-schema/error.json delete mode 100644 verisimdb/connectors/shared/json-schema/federation-result.json delete mode 100644 verisimdb/connectors/shared/json-schema/modality.json delete mode 100644 verisimdb/connectors/shared/json-schema/octad-input.json delete mode 100644 verisimdb/connectors/shared/json-schema/octad-status.json delete mode 100644 verisimdb/connectors/shared/json-schema/octad.json delete mode 100644 verisimdb/connectors/shared/json-schema/provenance-event.json delete mode 100644 verisimdb/connectors/shared/json-schema/query-params.json delete mode 100644 verisimdb/connectors/shared/openapi/verisim-api-v1.yaml delete mode 100644 verisimdb/connectors/shared/proto/verisim_federation.proto delete mode 100644 verisimdb/connectors/test-infra/.gatekeeper.yaml delete mode 100644 verisimdb/connectors/test-infra/0-AI-MANIFEST.a2ml delete mode 100644 verisimdb/connectors/test-infra/README.adoc delete mode 100644 verisimdb/connectors/test-infra/compose.toml delete mode 100755 verisimdb/connectors/test-infra/ct-build.sh delete mode 100644 verisimdb/connectors/test-infra/deploy.k9.ncl delete mode 100644 verisimdb/connectors/test-infra/images/Containerfile.clickhouse delete mode 100644 verisimdb/connectors/test-infra/images/Containerfile.influxdb delete mode 100644 verisimdb/connectors/test-infra/images/Containerfile.neo4j delete mode 100644 verisimdb/connectors/test-infra/images/Containerfile.redis-stack delete mode 100644 verisimdb/connectors/test-infra/images/Containerfile.surrealdb delete mode 100644 verisimdb/connectors/test-infra/manifest.toml delete mode 100644 verisimdb/connectors/test-infra/seed/clickhouse-init.sql delete mode 100755 verisimdb/connectors/test-infra/seed/influxdb-init.sh delete mode 100755 verisimdb/connectors/test-infra/seed/minio-init.sh delete mode 100644 verisimdb/connectors/test-infra/seed/mongodb-init.js delete mode 100644 verisimdb/connectors/test-infra/seed/neo4j-init.cypher delete mode 100755 verisimdb/connectors/test-infra/seed/redis-init.sh delete mode 100644 verisimdb/connectors/test-infra/seed/surrealdb-init.surql delete mode 100644 verisimdb/connectors/test-infra/vordr.toml delete mode 100644 verisimdb/container/.gatekeeper.yaml delete mode 100644 verisimdb/container/Containerfile delete mode 100644 verisimdb/container/compose.toml delete mode 100755 verisimdb/container/ct-build.sh delete mode 100644 verisimdb/container/entrypoint.sh delete mode 100644 verisimdb/container/manifest.toml delete mode 100644 verisimdb/contractiles/README.adoc delete mode 100644 verisimdb/contractiles/dust/Dustfile delete mode 100644 verisimdb/contractiles/must/Mustfile delete mode 100644 verisimdb/contractiles/trust/Trustfile delete mode 100644 verisimdb/debugger/.claude/CLAUDE.md delete mode 100644 verisimdb/debugger/Cargo.toml delete mode 100644 verisimdb/debugger/docs/CITATIONS.adoc delete mode 100644 verisimdb/debugger/examples/SafeDOMExample.res delete mode 100644 verisimdb/debugger/examples/web-project-deno.json delete mode 100644 verisimdb/debugger/fuzz/Cargo.toml delete mode 100644 verisimdb/debugger/fuzz/fuzz_targets/fuzz_main.rs delete mode 100644 verisimdb/debugger/src/main.rs delete mode 100644 verisimdb/debugger/tests/integration_test.rs delete mode 100644 verisimdb/demos/drift-detection/run_demo.exs delete mode 100644 verisimdb/deny.toml delete mode 100644 verisimdb/docs/.well-known/void.rdf delete mode 100644 verisimdb/docs/.well-known/void.ttl delete mode 100644 verisimdb/docs/CITATIONS.adoc delete mode 100644 verisimdb/docs/VCL-SPEC.adoc delete mode 100644 verisimdb/docs/adoption-strategy.adoc delete mode 100644 verisimdb/docs/backwards-compatibility.adoc delete mode 100644 verisimdb/docs/business/business-case.adoc delete mode 100644 verisimdb/docs/business/financials/cost-structure.csv delete mode 100644 verisimdb/docs/business/financials/market-sizing.csv delete mode 100644 verisimdb/docs/business/financials/revenue-model.csv delete mode 100644 verisimdb/docs/business/financials/unit-economics.csv delete mode 100644 verisimdb/docs/business/marketing/feature-comparison.adoc delete mode 100644 verisimdb/docs/business/marketing/one-pager.adoc delete mode 100644 verisimdb/docs/business/marketing/pitch-deck-outline.adoc delete mode 100644 verisimdb/docs/business/marketing/use-cases.adoc delete mode 100644 verisimdb/docs/business/pr/faq.adoc delete mode 100644 verisimdb/docs/business/pr/key-messages.adoc delete mode 100644 verisimdb/docs/business/pr/press-release.adoc delete mode 100644 verisimdb/docs/business/strategy/adoption-roadmap.adoc delete mode 100644 verisimdb/docs/business/strategy/economics-analysis.adoc delete mode 100644 verisimdb/docs/business/strategy/functional-units.adoc delete mode 100644 verisimdb/docs/business/strategy/go-to-market.adoc delete mode 100644 verisimdb/docs/cache-sharing-strategy.adoc delete mode 100644 verisimdb/docs/caching-strategy.adoc delete mode 100644 verisimdb/docs/challenges-federated.adoc delete mode 100644 verisimdb/docs/challenges-hybrid.adoc delete mode 100644 verisimdb/docs/challenges-standalone.adoc delete mode 100644 verisimdb/docs/consultation-dependent-types-zkp.adoc delete mode 100644 verisimdb/docs/consultation-normalization-strategy.adoc delete mode 100644 verisimdb/docs/deployment-modes.adoc delete mode 100644 verisimdb/docs/design/DESIGN-2026-02-27-level-data-model.md delete mode 100644 verisimdb/docs/design/DESIGN-2026-02-27-strategic-improvements.adoc delete mode 100644 verisimdb/docs/design/DESIGN-2026-02-27-vcl-dt-assessment.adoc delete mode 100644 verisimdb/docs/design/DESIGN-2026-02-28-panll-interop-telemetry.md delete mode 100644 verisimdb/docs/drift-handling.adoc delete mode 100644 verisimdb/docs/error-handling-strategy.adoc delete mode 100644 verisimdb/docs/federation-readiness.adoc delete mode 100644 verisimdb/docs/getting-started.adoc delete mode 100644 verisimdb/docs/minikanren-integration-v3.adoc delete mode 100644 verisimdb/docs/normalization-cascade.adoc delete mode 100644 verisimdb/docs/panll-module-audit.adoc delete mode 100644 verisimdb/docs/papers/verisimdb-federated-consistency.adoc delete mode 100644 verisimdb/docs/papers/verisimdb-idaptik-case-study.adoc delete mode 100644 verisimdb/docs/query-optimization-overview.adoc delete mode 100644 verisimdb/docs/rescript-registry-types.adoc delete mode 100644 verisimdb/docs/reversibility-design.adoc delete mode 100644 verisimdb/docs/safety-and-fault-tolerance.adoc delete mode 100644 verisimdb/docs/safety-theory-applied.adoc delete mode 100644 verisimdb/docs/security-lessons.lgt delete mode 100644 verisimdb/docs/snapshotting-and-truncation-logic.adoc delete mode 100644 verisimdb/docs/technical-specification-kraft-metadata-log.adoc delete mode 100644 verisimdb/docs/vcl-architecture.adoc delete mode 100644 verisimdb/docs/vcl-examples.adoc delete mode 100644 verisimdb/docs/vcl-formal-semantics.adoc delete mode 100644 verisimdb/docs/vcl-grammar.ebnf delete mode 100644 verisimdb/docs/vcl-type-system.adoc delete mode 100644 verisimdb/docs/vcl-vs-sql.adoc delete mode 100644 verisimdb/docs/vcl-vs-vcl-dt.adoc delete mode 100644 verisimdb/docs/zkp-and-sanctify-integration.adoc delete mode 100644 verisimdb/elixir-orchestration/.formatter.exs delete mode 100644 verisimdb/elixir-orchestration/config/config.exs delete mode 100644 verisimdb/elixir-orchestration/config/dev.exs delete mode 100644 verisimdb/elixir-orchestration/config/prod.exs delete mode 100644 verisimdb/elixir-orchestration/config/runtime.exs delete mode 100644 verisimdb/elixir-orchestration/config/test.exs delete mode 100644 verisimdb/elixir-orchestration/lib/verisim.ex delete mode 100644 verisimdb/elixir-orchestration/lib/verisim/api/router.ex delete mode 100644 verisimdb/elixir-orchestration/lib/verisim/application.ex delete mode 100644 verisimdb/elixir-orchestration/lib/verisim/consensus/kraft_node.ex delete mode 100644 verisimdb/elixir-orchestration/lib/verisim/consensus/kraft_supervisor.ex delete mode 100644 verisimdb/elixir-orchestration/lib/verisim/consensus/kraft_transport.ex delete mode 100644 verisimdb/elixir-orchestration/lib/verisim/consensus/kraft_wal.ex delete mode 100644 verisimdb/elixir-orchestration/lib/verisim/drift/drift_monitor.ex delete mode 100644 verisimdb/elixir-orchestration/lib/verisim/entity/entity_server.ex delete mode 100644 verisimdb/elixir-orchestration/lib/verisim/federation/adapter.ex delete mode 100644 verisimdb/elixir-orchestration/lib/verisim/federation/adapters/arangodb.ex delete mode 100644 verisimdb/elixir-orchestration/lib/verisim/federation/adapters/clickhouse.ex delete mode 100644 verisimdb/elixir-orchestration/lib/verisim/federation/adapters/duckdb.ex delete mode 100644 verisimdb/elixir-orchestration/lib/verisim/federation/adapters/elasticsearch.ex delete mode 100644 verisimdb/elixir-orchestration/lib/verisim/federation/adapters/influxdb.ex delete mode 100644 verisimdb/elixir-orchestration/lib/verisim/federation/adapters/mongodb.ex delete mode 100644 verisimdb/elixir-orchestration/lib/verisim/federation/adapters/neo4j.ex delete mode 100644 verisimdb/elixir-orchestration/lib/verisim/federation/adapters/object_storage.ex delete mode 100644 verisimdb/elixir-orchestration/lib/verisim/federation/adapters/postgresql.ex delete mode 100644 verisimdb/elixir-orchestration/lib/verisim/federation/adapters/redis.ex delete mode 100644 verisimdb/elixir-orchestration/lib/verisim/federation/adapters/sqlite.ex delete mode 100644 verisimdb/elixir-orchestration/lib/verisim/federation/adapters/surrealdb.ex delete mode 100644 verisimdb/elixir-orchestration/lib/verisim/federation/adapters/vector_db.ex delete mode 100644 verisimdb/elixir-orchestration/lib/verisim/federation/adapters/verisimdb.ex delete mode 100644 verisimdb/elixir-orchestration/lib/verisim/federation/resolver.ex delete mode 100644 verisimdb/elixir-orchestration/lib/verisim/health_checker.ex delete mode 100644 verisimdb/elixir-orchestration/lib/verisim/hypatia/dispatch_bridge.ex delete mode 100644 verisimdb/elixir-orchestration/lib/verisim/hypatia/pattern_query.ex delete mode 100644 verisimdb/elixir-orchestration/lib/verisim/hypatia/scan_ingester.ex delete mode 100644 verisimdb/elixir-orchestration/lib/verisim/nif_bridge.ex delete mode 100644 verisimdb/elixir-orchestration/lib/verisim/query/query_router.ex delete mode 100644 verisimdb/elixir-orchestration/lib/verisim/query/vcl_bridge.ex delete mode 100644 verisimdb/elixir-orchestration/lib/verisim/query/vcl_executor.ex delete mode 100644 verisimdb/elixir-orchestration/lib/verisim/query/vcl_proof_certificate.ex delete mode 100644 verisimdb/elixir-orchestration/lib/verisim/query/vcl_type_checker.ex delete mode 100644 verisimdb/elixir-orchestration/lib/verisim/rust_client.ex delete mode 100644 verisimdb/elixir-orchestration/lib/verisim/schema/schema_registry.ex delete mode 100644 verisimdb/elixir-orchestration/lib/verisim/telemetry.ex delete mode 100644 verisimdb/elixir-orchestration/lib/verisim/telemetry/collector.ex delete mode 100644 verisimdb/elixir-orchestration/lib/verisim/telemetry/reporter.ex delete mode 100644 verisimdb/elixir-orchestration/lib/verisim/transport.ex delete mode 100644 verisimdb/elixir-orchestration/mix.exs delete mode 100644 verisimdb/elixir-orchestration/test/fixtures/vcl/basic-select.expected.json delete mode 100644 verisimdb/elixir-orchestration/test/fixtures/vcl/basic-select.vcl delete mode 100644 verisimdb/elixir-orchestration/test/fixtures/vcl/drift-query.expected.json delete mode 100644 verisimdb/elixir-orchestration/test/fixtures/vcl/drift-query.vcl delete mode 100644 verisimdb/elixir-orchestration/test/fixtures/vcl/multi-proof.expected.json delete mode 100644 verisimdb/elixir-orchestration/test/fixtures/vcl/multi-proof.vcl delete mode 100644 verisimdb/elixir-orchestration/test/fixtures/vcl/proof-existence.expected.json delete mode 100644 verisimdb/elixir-orchestration/test/fixtures/vcl/proof-existence.vcl delete mode 100644 verisimdb/elixir-orchestration/test/fixtures/vcl/vector-search.expected.json delete mode 100644 verisimdb/elixir-orchestration/test/fixtures/vcl/vector-search.vcl delete mode 100644 verisimdb/elixir-orchestration/test/integration_test.exs delete mode 100644 verisimdb/elixir-orchestration/test/support/vcl_test_helpers.ex delete mode 100644 verisimdb/elixir-orchestration/test/test_helper.exs delete mode 100644 verisimdb/elixir-orchestration/test/verisim/api/router_test.exs delete mode 100644 verisimdb/elixir-orchestration/test/verisim/aspect/concurrency_test.exs delete mode 100644 verisimdb/elixir-orchestration/test/verisim/aspect/security_test.exs delete mode 100644 verisimdb/elixir-orchestration/test/verisim/consensus/kraft_node_test.exs delete mode 100644 verisimdb/elixir-orchestration/test/verisim/consensus/kraft_property_test.exs delete mode 100644 verisimdb/elixir-orchestration/test/verisim/consensus/kraft_recovery_test.exs delete mode 100644 verisimdb/elixir-orchestration/test/verisim/consensus/kraft_transport_test.exs delete mode 100644 verisimdb/elixir-orchestration/test/verisim/consensus/kraft_wal_test.exs delete mode 100644 verisimdb/elixir-orchestration/test/verisim/e2e_verisimdb_test.exs delete mode 100644 verisimdb/elixir-orchestration/test/verisim/federation/adapter_test.exs delete mode 100644 verisimdb/elixir-orchestration/test/verisim/federation/adapters/clickhouse_test.exs delete mode 100644 verisimdb/elixir-orchestration/test/verisim/federation/adapters/duckdb_test.exs delete mode 100644 verisimdb/elixir-orchestration/test/verisim/federation/adapters/influxdb_test.exs delete mode 100644 verisimdb/elixir-orchestration/test/verisim/federation/adapters/integration/clickhouse_integration_test.exs delete mode 100644 verisimdb/elixir-orchestration/test/verisim/federation/adapters/integration/influxdb_integration_test.exs delete mode 100644 verisimdb/elixir-orchestration/test/verisim/federation/adapters/integration/mongodb_integration_test.exs delete mode 100644 verisimdb/elixir-orchestration/test/verisim/federation/adapters/integration/neo4j_integration_test.exs delete mode 100644 verisimdb/elixir-orchestration/test/verisim/federation/adapters/integration/object_storage_integration_test.exs delete mode 100644 verisimdb/elixir-orchestration/test/verisim/federation/adapters/integration/redis_integration_test.exs delete mode 100644 verisimdb/elixir-orchestration/test/verisim/federation/adapters/integration/surrealdb_integration_test.exs delete mode 100644 verisimdb/elixir-orchestration/test/verisim/federation/adapters/mongodb_test.exs delete mode 100644 verisimdb/elixir-orchestration/test/verisim/federation/adapters/neo4j_test.exs delete mode 100644 verisimdb/elixir-orchestration/test/verisim/federation/adapters/object_storage_test.exs delete mode 100644 verisimdb/elixir-orchestration/test/verisim/federation/adapters/redis_test.exs delete mode 100644 verisimdb/elixir-orchestration/test/verisim/federation/adapters/sqlite_test.exs delete mode 100644 verisimdb/elixir-orchestration/test/verisim/federation/adapters/surrealdb_test.exs delete mode 100644 verisimdb/elixir-orchestration/test/verisim/federation/adapters/vector_db_test.exs delete mode 100644 verisimdb/elixir-orchestration/test/verisim/federation/resolver_test.exs delete mode 100644 verisimdb/elixir-orchestration/test/verisim/hypatia/dispatch_bridge_test.exs delete mode 100644 verisimdb/elixir-orchestration/test/verisim/hypatia/pattern_query_test.exs delete mode 100644 verisimdb/elixir-orchestration/test/verisim/hypatia/scan_ingester_test.exs delete mode 100644 verisimdb/elixir-orchestration/test/verisim/query/vcl_crossmodal_test.exs delete mode 100644 verisimdb/elixir-orchestration/test/verisim/query/vcl_dt_integration_test.exs delete mode 100644 verisimdb/elixir-orchestration/test/verisim/query/vcl_dt_test.exs delete mode 100644 verisimdb/elixir-orchestration/test/verisim/query/vcl_e2e_test.exs delete mode 100644 verisimdb/elixir-orchestration/test/verisim/query/vcl_integration_test.exs delete mode 100644 verisimdb/elixir-orchestration/test/verisim/query/vcl_proof_certificate_test.exs delete mode 100644 verisimdb/elixir-orchestration/test/verisim/query/vcl_property_test.exs delete mode 100644 verisimdb/elixir-orchestration/test/verisim/query/vcl_test.exs delete mode 100644 verisimdb/elixir-orchestration/test/verisim/query/vcl_type_checker_test.exs delete mode 100644 verisimdb/elixir-orchestration/test/verisim/telemetry_test.exs delete mode 100644 verisimdb/elixir-orchestration/test/verisim_test.exs delete mode 100644 verisimdb/examples/README.adoc delete mode 100644 verisimdb/examples/SafeDOMExample.res delete mode 100755 verisimdb/examples/load-sample-data.sh delete mode 100644 verisimdb/examples/sample-data/seed.json delete mode 100755 verisimdb/examples/smoke-test.sh delete mode 100644 verisimdb/examples/vcl-queries/01-basic-search.vcl delete mode 100644 verisimdb/examples/vcl-queries/02-vector-similarity.vcl delete mode 100644 verisimdb/examples/vcl-queries/03-cross-modal-join.vcl delete mode 100644 verisimdb/examples/vcl-queries/04-drift-detection.vcl delete mode 100644 verisimdb/examples/vcl-queries/05-provenance-chain.vcl delete mode 100644 verisimdb/examples/vcl-queries/06-temporal-range.vcl delete mode 100644 verisimdb/examples/vcl-queries/07-spatial-radius.vcl delete mode 100644 verisimdb/examples/vcl-queries/08-proof-existence.vcl delete mode 100644 verisimdb/examples/vcl-queries/09-proof-consistency.vcl delete mode 100644 verisimdb/examples/vcl-queries/10-multi-modal-pipeline.vcl delete mode 100644 verisimdb/examples/web-project-deno.json delete mode 100644 verisimdb/ffi/zig/build.zig delete mode 100644 verisimdb/ffi/zig/src/main.zig delete mode 100644 verisimdb/ffi/zig/test/integration_test.zig delete mode 100644 verisimdb/fuzz/Cargo.toml delete mode 100644 verisimdb/fuzz/fuzz_targets/fuzz_octad_id.rs delete mode 100644 verisimdb/historiographic-custodian.html delete mode 100644 verisimdb/kraft-comparison.adoc delete mode 100644 verisimdb/lib/verisim/adaptive_learner.ex delete mode 100644 verisimdb/lib/verisim/circuit_breaker.ex delete mode 100644 verisimdb/lib/verisim/error_recovery.ex delete mode 100644 verisimdb/lib/verisim/query_cache.ex delete mode 100644 verisimdb/lib/verisim/query_planner_bidirectional.ex delete mode 100644 verisimdb/lib/verisim/query_planner_config.ex delete mode 100644 verisimdb/lib/verisim/query_router_cached.ex delete mode 100644 verisimdb/llm-warmup-dev.md delete mode 100644 verisimdb/llm-warmup-user.md delete mode 100644 verisimdb/opsm.toml delete mode 100644 verisimdb/osv-scanner.toml delete mode 100644 verisimdb/playground/build.mjs delete mode 100644 verisimdb/playground/deno.json delete mode 100644 verisimdb/playground/public/index.html delete mode 100644 verisimdb/playground/public/manifest.json delete mode 100644 verisimdb/playground/public/style.css delete mode 100644 verisimdb/playground/public/sw.js delete mode 100644 verisimdb/playground/rescript.json delete mode 100644 verisimdb/playground/src/ApiClient.res delete mode 100644 verisimdb/playground/src/App.res delete mode 100644 verisimdb/playground/src/DemoExecutor.res delete mode 100644 verisimdb/playground/src/Examples.res delete mode 100644 verisimdb/playground/src/Formatter.res delete mode 100644 verisimdb/playground/src/Highlighter.res delete mode 100644 verisimdb/playground/src/Linter.res delete mode 100644 verisimdb/playground/src/VclKeywords.res delete mode 100644 verisimdb/proven-coherence.md delete mode 100644 verisimdb/references.bib delete mode 100644 verisimdb/rescript.json delete mode 100644 verisimdb/rust-core/fuzz/Cargo.toml delete mode 100644 verisimdb/rust-core/fuzz/fuzz_targets/fuzz_vcl_parser.rs delete mode 100644 verisimdb/rust-core/verisim-api/Cargo.toml delete mode 100644 verisimdb/rust-core/verisim-api/build.rs delete mode 100644 verisimdb/rust-core/verisim-api/proto/verisim.proto delete mode 100644 verisimdb/rust-core/verisim-api/src/a2ml.rs delete mode 100644 verisimdb/rust-core/verisim-api/src/auth.rs delete mode 100644 verisimdb/rust-core/verisim-api/src/federation.rs delete mode 100644 verisimdb/rust-core/verisim-api/src/graphql.rs delete mode 100644 verisimdb/rust-core/verisim-api/src/groove.rs delete mode 100644 verisimdb/rust-core/verisim-api/src/grpc.rs delete mode 100644 verisimdb/rust-core/verisim-api/src/lib.rs delete mode 100644 verisimdb/rust-core/verisim-api/src/main.rs delete mode 100644 verisimdb/rust-core/verisim-api/src/proof_attempts.rs delete mode 100644 verisimdb/rust-core/verisim-api/src/proto/verisim.rs delete mode 100644 verisimdb/rust-core/verisim-api/src/rbac.rs delete mode 100644 verisimdb/rust-core/verisim-api/src/transaction.rs delete mode 100644 verisimdb/rust-core/verisim-api/src/vcl.rs delete mode 100644 verisimdb/rust-core/verisim-document/Cargo.toml delete mode 100644 verisimdb/rust-core/verisim-document/src/lib.rs delete mode 100644 verisimdb/rust-core/verisim-document/tests/property_tests.proptest-regressions delete mode 100644 verisimdb/rust-core/verisim-document/tests/property_tests.rs delete mode 100644 verisimdb/rust-core/verisim-drift/Cargo.toml delete mode 100644 verisimdb/rust-core/verisim-drift/src/calculator.rs delete mode 100644 verisimdb/rust-core/verisim-drift/src/lib.rs delete mode 100644 verisimdb/rust-core/verisim-graph/Cargo.toml delete mode 100644 verisimdb/rust-core/verisim-graph/src/lib.rs delete mode 100644 verisimdb/rust-core/verisim-graph/src/oxigraph_backend.rs delete mode 100644 verisimdb/rust-core/verisim-graph/src/redb_backend.rs delete mode 100644 verisimdb/rust-core/verisim-nif/Cargo.toml delete mode 100644 verisimdb/rust-core/verisim-nif/src/lib.rs delete mode 100644 verisimdb/rust-core/verisim-normalizer/Cargo.toml delete mode 100644 verisimdb/rust-core/verisim-normalizer/src/conflict.rs delete mode 100644 verisimdb/rust-core/verisim-normalizer/src/lib.rs delete mode 100644 verisimdb/rust-core/verisim-normalizer/src/regeneration.rs delete mode 100644 verisimdb/rust-core/verisim-normalizer/src/storage_regenerator.rs delete mode 100644 verisimdb/rust-core/verisim-octad/Cargo.toml delete mode 100644 verisimdb/rust-core/verisim-octad/src/lib.rs delete mode 100644 verisimdb/rust-core/verisim-octad/src/query_octad.rs delete mode 100644 verisimdb/rust-core/verisim-octad/src/ram_promotion.rs delete mode 100644 verisimdb/rust-core/verisim-octad/src/store.rs delete mode 100644 verisimdb/rust-core/verisim-octad/src/transaction.rs delete mode 100644 verisimdb/rust-core/verisim-octad/tests/atomicity_tests.rs delete mode 100644 verisimdb/rust-core/verisim-octad/tests/crash_recovery_tests.rs delete mode 100644 verisimdb/rust-core/verisim-octad/tests/integration_tests.rs delete mode 100644 verisimdb/rust-core/verisim-octad/tests/stress_tests.rs delete mode 100644 verisimdb/rust-core/verisim-planner/Cargo.toml delete mode 100644 verisimdb/rust-core/verisim-planner/src/config.rs delete mode 100644 verisimdb/rust-core/verisim-planner/src/cost.rs delete mode 100644 verisimdb/rust-core/verisim-planner/src/error.rs delete mode 100644 verisimdb/rust-core/verisim-planner/src/explain.rs delete mode 100644 verisimdb/rust-core/verisim-planner/src/lib.rs delete mode 100644 verisimdb/rust-core/verisim-planner/src/optimizer.rs delete mode 100644 verisimdb/rust-core/verisim-planner/src/plan.rs delete mode 100644 verisimdb/rust-core/verisim-planner/src/prepared.rs delete mode 100644 verisimdb/rust-core/verisim-planner/src/profiler.rs delete mode 100644 verisimdb/rust-core/verisim-planner/src/slow_query.rs delete mode 100644 verisimdb/rust-core/verisim-planner/src/stats.rs delete mode 100644 verisimdb/rust-core/verisim-planner/src/vcl_bridge.rs delete mode 100644 verisimdb/rust-core/verisim-provenance/Cargo.toml delete mode 100644 verisimdb/rust-core/verisim-provenance/src/lib.rs delete mode 100644 verisimdb/rust-core/verisim-provenance/src/persistent.rs delete mode 100644 verisimdb/rust-core/verisim-repl/Cargo.toml delete mode 100644 verisimdb/rust-core/verisim-repl/src/client.rs delete mode 100644 verisimdb/rust-core/verisim-repl/src/completer.rs delete mode 100644 verisimdb/rust-core/verisim-repl/src/formatter.rs delete mode 100644 verisimdb/rust-core/verisim-repl/src/highlighter.rs delete mode 100644 verisimdb/rust-core/verisim-repl/src/linter.rs delete mode 100644 verisimdb/rust-core/verisim-repl/src/main.rs delete mode 100644 verisimdb/rust-core/verisim-repl/src/vcl_fmt.rs delete mode 100644 verisimdb/rust-core/verisim-semantic/Cargo.toml delete mode 100644 verisimdb/rust-core/verisim-semantic/src/circuit_compiler.rs delete mode 100644 verisimdb/rust-core/verisim-semantic/src/circuit_registry.rs delete mode 100644 verisimdb/rust-core/verisim-semantic/src/lib.rs delete mode 100644 verisimdb/rust-core/verisim-semantic/src/persistent.rs delete mode 100644 verisimdb/rust-core/verisim-semantic/src/proven_bridge.rs delete mode 100644 verisimdb/rust-core/verisim-semantic/src/sanctify_bridge.rs delete mode 100644 verisimdb/rust-core/verisim-semantic/src/verification_keys.rs delete mode 100644 verisimdb/rust-core/verisim-semantic/src/zkp.rs delete mode 100644 verisimdb/rust-core/verisim-semantic/src/zkp_bridge.rs delete mode 100644 verisimdb/rust-core/verisim-spatial/Cargo.toml delete mode 100644 verisimdb/rust-core/verisim-spatial/src/lib.rs delete mode 100644 verisimdb/rust-core/verisim-spatial/src/persistent.rs delete mode 100644 verisimdb/rust-core/verisim-storage/Cargo.toml delete mode 100644 verisimdb/rust-core/verisim-storage/src/backend.rs delete mode 100644 verisimdb/rust-core/verisim-storage/src/error.rs delete mode 100644 verisimdb/rust-core/verisim-storage/src/lib.rs delete mode 100644 verisimdb/rust-core/verisim-storage/src/memory.rs delete mode 100644 verisimdb/rust-core/verisim-storage/src/metrics.rs delete mode 100644 verisimdb/rust-core/verisim-storage/src/redb_backend.rs delete mode 100644 verisimdb/rust-core/verisim-storage/src/typed.rs delete mode 100644 verisimdb/rust-core/verisim-temporal/Cargo.toml delete mode 100644 verisimdb/rust-core/verisim-temporal/src/diff.rs delete mode 100644 verisimdb/rust-core/verisim-temporal/src/lib.rs delete mode 100644 verisimdb/rust-core/verisim-temporal/src/persistent.rs delete mode 100644 verisimdb/rust-core/verisim-temporal/tests/property_tests.rs delete mode 100644 verisimdb/rust-core/verisim-tensor/Cargo.toml delete mode 100644 verisimdb/rust-core/verisim-tensor/src/lib.rs delete mode 100644 verisimdb/rust-core/verisim-tensor/src/persistent.rs delete mode 100644 verisimdb/rust-core/verisim-vector/Cargo.toml delete mode 100644 verisimdb/rust-core/verisim-vector/src/hnsw.rs delete mode 100644 verisimdb/rust-core/verisim-vector/src/lib.rs delete mode 100644 verisimdb/rust-core/verisim-vector/src/persistent.rs delete mode 100644 verisimdb/rust-core/verisim-wal/Cargo.toml delete mode 100644 verisimdb/rust-core/verisim-wal/src/entry.rs delete mode 100644 verisimdb/rust-core/verisim-wal/src/error.rs delete mode 100644 verisimdb/rust-core/verisim-wal/src/lib.rs delete mode 100644 verisimdb/rust-core/verisim-wal/src/reader.rs delete mode 100644 verisimdb/rust-core/verisim-wal/src/segment.rs delete mode 100644 verisimdb/rust-core/verisim-wal/src/writer.rs delete mode 100755 verisimdb/scripts/post-commit-hook.sh delete mode 100755 verisimdb/scripts/self-ingest.sh delete mode 100755 verisimdb/scripts/self-query.sh delete mode 100755 verisimdb/scripts/smoke-test.sh delete mode 100755 verisimdb/scripts/two-node-test.sh delete mode 100644 verisimdb/selur-compose.yml delete mode 100644 verisimdb/site/index.md delete mode 100644 verisimdb/spec/README.adoc delete mode 100644 verisimdb/spec/grammar.ebnf delete mode 100644 verisimdb/spec/system-specs.md delete mode 100644 verisimdb/src/abi/Foreign.idr delete mode 100644 verisimdb/src/abi/Layout.idr delete mode 100644 verisimdb/src/abi/Types.idr delete mode 100644 verisimdb/src/registry/KRaftCluster.res delete mode 100644 verisimdb/src/registry/KRaftSerializer.res delete mode 100644 verisimdb/src/registry/MetadataLog.res delete mode 100644 verisimdb/src/registry/Registry.res delete mode 100644 verisimdb/src/vcl/VCLBidir.res delete mode 100644 verisimdb/src/vcl/VCLCircuit.res delete mode 100644 verisimdb/src/vcl/VCLContext.res delete mode 100644 verisimdb/src/vcl/VCLError.res delete mode 100644 verisimdb/src/vcl/VCLExplain.res delete mode 100644 verisimdb/src/vcl/VCLParser.res delete mode 100644 verisimdb/src/vcl/VCLParser_test.res delete mode 100644 verisimdb/src/vcl/VCLProofObligation.res delete mode 100644 verisimdb/src/vcl/VCLSubtyping.res delete mode 100644 verisimdb/src/vcl/VCLTypeChecker.res delete mode 100644 verisimdb/src/vcl/VCLTypes.res delete mode 100644 verisimdb/stapeln.toml delete mode 100644 verisimdb/tests/integration_test.rs delete mode 100644 verisimdb/vcl-bridge/vcl_parser_port.js delete mode 100644 verisimdb/verification/PROOF-STATUS.md delete mode 100644 verisimdb/verification/README.adoc delete mode 120000 verisimdb/verification/benchmarks delete mode 120000 verisimdb/verification/fuzzing delete mode 100644 verisimdb/verification/proofs/agda/ProvenanceChain.agda delete mode 100644 verisimdb/verification/proofs/agda/ProvenanceChain.agdai delete mode 100644 verisimdb/verification/proofs/idris2/ConnectorSafety.idr delete mode 100644 verisimdb/verification/proofs/idris2/DriftMetric.idr delete mode 100644 verisimdb/verification/proofs/idris2/FFIOwnership.idr delete mode 100644 verisimdb/verification/proofs/idris2/OctadCoherence.idr delete mode 100644 verisimdb/verification/proofs/lean4/RaftSafety.lean delete mode 100644 verisimdb/verification/proofs/lean4/VCLSubtyping.lean delete mode 100644 verisimdb/verification/proofs/lean4/VCLTypeSoundness.lean delete mode 100644 verisimdb/verification/proofs/lean4/WALIntegrity.lean delete mode 100644 verisimdb/verification/proofs/lean4/lake-manifest.json delete mode 100644 verisimdb/verification/proofs/lean4/lakefile.lean delete mode 100644 verisimdb/verification/proofs/tlaplus/.gitignore delete mode 100644 verisimdb/verification/proofs/tlaplus/Normalizer.cfg delete mode 100644 verisimdb/verification/proofs/tlaplus/Normalizer.tla delete mode 100644 verisimdb/verification/proofs/tlaplus/OctadAtomicity.cfg delete mode 100644 verisimdb/verification/proofs/tlaplus/OctadAtomicity.tla delete mode 100644 verisimdb/verification/proofs/tlaplus/README.adoc delete mode 100644 verisimdb/verification/proofs/tlaplus/Serializability.cfg delete mode 100644 verisimdb/verification/proofs/tlaplus/Serializability.tla delete mode 120000 verisimdb/verification/tests delete mode 100644 verisimdb/verisim-architecture-visualisation.html diff --git a/.github/workflows/placement-guard.yml b/.github/workflows/placement-guard.yml index ca711149..de99f199 100644 --- a/.github/workflows/placement-guard.yml +++ b/.github/workflows/placement-guard.yml @@ -69,14 +69,15 @@ jobs: # Legacy per-database dirs (grandfathered: warn, do not fail — being extracted). # - # `lithoglyph` was removed from this list on 2026-07-27. Its extraction is - # complete (hyperpolymath/lithoglyph#4) and its 819 files are gone from this - # repo, so a new file under lithoglyph/ is no longer "legacy content not yet - # moved" — it is fresh duplication of a repo that already exists. Warning - # would let exactly the defect the extraction fixed grow back. It now fails. + # `lithoglyph` was removed from this list on 2026-07-27, `verisimdb` on + # 2026-08-03. Both extractions are complete and their files are gone from + # this repo, so a new file under either is no longer "legacy content not + # yet moved" — it is fresh duplication of a repo that already exists. + # Warning would let exactly the defect each extraction fixed grow back. + # Both now fail. # # Move a directory out of this list as each extraction completes. - GRANDFATHER='^(verisimdb|quandledb|nqc|typeql-experimental|verisim-core|verisim-modular-experiment)/' + GRANDFATHER='^(quandledb|nqc|typeql-experimental|verisim-core|verisim-modular-experiment)/' FAIL=0 while IFS= read -r f; do diff --git a/CLAUDE.md b/CLAUDE.md index 073b0bab..f1690df7 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -40,12 +40,18 @@ See **`REGISTRY.adoc`** for the authoritative map. Examples: VeriSimDB → ### Transitional note -`lithoglyph/` is **done**: extraction completed 2026-07-27 and its 819 files were -removed, leaving a pointer README. Everything it held is preserved at its original -path in the `split-history/lithoglyph` tag on `origin` — never prune that tag or the -`_split_lithoglyph` branch. - -The directories `verisimdb/`, `quandledb/`, `nqc/`, +`lithoglyph/` and `verisimdb/` are **done**: extractions completed 2026-07-27 and +2026-08-03, their 819 and 713 files removed, each leaving a pointer README. +Everything they held is preserved at its original path in the +`split-history/{lithoglyph,verisimdb}` tags on `origin` — never prune those tags or +the `_split_*` branches. + +verisimdb was a FORK, not a duplicate (108 files unique here, 331 substantive +diffs), so seven files with no upstream counterpart were ported first — +hyperpolymath/verisimdb#219. Do not assume the next extraction is a clean subset; +check before deleting. + +The directories `quandledb/`, `nqc/`, `typeql-experimental/`, `verisim-core/`, and `verisim-modular-experiment/` are **legacy content being extracted** to their own repos — see `docs/migration/RESITE-DATABASES-TO-OWN-REPOS.adoc`. **Do not grow them.** A CI guard diff --git a/REGISTRY.adoc b/REGISTRY.adoc index 3f5dabbd..2bb74f22 100644 --- a/REGISTRY.adoc +++ b/REGISTRY.adoc @@ -20,8 +20,8 @@ Resite: decisions finalised; execution per |=== | Database | Repository | Status | Query language (repo) | VeriSimDB (cross-modal consistency) | hyperpolymath/verisimdb | exists (canonical) | VCL — nested in verisimdb -| Lithoglyph (narrative-first) | hyperpolymath/lithoglyph | to create | GNPL (narration/projection, over GQLdt) — hyperpolymath/gnpl (exists) -| Glyphbase (Lithoglyph web UI) | hyperpolymath/glyphbase | to create | — +| Lithoglyph (narrative-first) | hyperpolymath/lithoglyph | exists (canonical) | GNPL (narration/projection, over GQLdt) — hyperpolymath/gnpl (exists) +| Glyphbase (Airtable-mode delivery) | hyperpolymath/glyphbase | exists (canonical) | — | QuandleDB (knot-theory) | hyperpolymath/quandledb | exists (canonical) | KRL — hyperpolymath/krl |=== @@ -84,9 +84,8 @@ make a stale sentence true — a submodule would have to be fetched with `submodules: recursive` in every workflow that touches the tree, which is real cost for a coordination repo that only needs to *point* at its members. -`lithoglyph/` now holds a pointer README only; its 819 files were removed once the -extraction completed (2026-07-27). `verisimdb/` is still 714 tracked files awaiting the -same treatment. +`lithoglyph/` and `verisimdb/` now hold pointer READMEs only; their 819 and 713 files +were removed once each extraction completed (2026-07-27 and 2026-08-03). == What belongs in *this* (coordination) repo @@ -99,6 +98,14 @@ same treatment. [NOTE] ==== -The per-database directories currently in this repo are *legacy* and are being extracted -— see `docs/migration/RESITE-DATABASES-TO-OWN-REPOS.adoc`. Do not add new content to them. +The per-database directories still in this repo are *legacy* and are being extracted — +see `docs/migration/RESITE-DATABASES-TO-OWN-REPOS.adoc`. Do not add new content to them. + +`lithoglyph/` and `verisimdb/` are already done and now fail the placement guard rather +than warning. `quandledb/`, `nqc/`, `typeql-experimental/`, `verisim-core/` and +`verisim-modular-experiment/` remain grandfathered. + +**verisimdb was a fork, not a duplicate.** Do not assume the next extraction is a clean +subset of its own repo — verify byte-identity per file first, and port whatever is +unique before deleting. ==== diff --git a/verisimdb/.cfignore b/verisimdb/.cfignore deleted file mode 100644 index bf5c7365..00000000 --- a/verisimdb/.cfignore +++ /dev/null @@ -1,5 +0,0 @@ -target/ -node_modules/ -.git/ -*.rlib -*-lsp diff --git a/verisimdb/.claude/CLAUDE.md b/verisimdb/.claude/CLAUDE.md deleted file mode 100644 index f452eb79..00000000 --- a/verisimdb/.claude/CLAUDE.md +++ /dev/null @@ -1,440 +0,0 @@ -# CLAUDE.md - VeriSimDB AI Assistant Instructions - -## CRITICAL: Instance Policy (Read First) - -**This repository is SOURCE CODE and EXAMPLE DATA. It is NOT a database instance.** - -If you are integrating VeriSimDB into another project (IDApTIK, Burble, etc.): - -1. **NEVER** store application data in this repo's VeriSimDB instance -2. **NEVER** point application code at `localhost:8080` expecting a shared server -3. **ALWAYS** create a dedicated VeriSimDB instance in the consuming project -4. **ALWAYS** use a unique port per project (not 8080) -5. **ALWAYS** copy the client SDK (`connectors/clients//`) into your project - -**Correct pattern:** -``` -your-project/ -├── containers/ -│ └── verisimdb.Containerfile ← YOUR instance -├── src/verisimdb/ -│ └── VeriSimClient.res ← Copied SDK files -└── podman-compose.yml ← YOUR port + volume -``` - -**Port assignments:** -| Project | Port | Volume | -|---------|------|--------| -| VeriSimDB dev/test | 8080 | (in-memory, ephemeral) | -| IDApTIK | 8090 | idaptik-verisimdb-data | -| Burble | 8091 | burble-verisimdb-data | -| Hypatia | 8092 | hypatia-verisimdb-data | - -The `examples/` directory contains example/demo data only. Never treat it as live storage. - -## Project Overview - -VeriSimDB (Veridical Simulacrum Database) is a cross-system entity consistency engine with drift detection, self-normalisation, and formally verified queries. Each entity exists simultaneously across 8 modalities — the octad (Graph, Vector, Tensor, Semantic, Document, Temporal, Provenance, Spatial) — with drift detection and automatic consistency maintenance. Operates as standalone database OR heterogeneous federation coordinator over existing databases. - -## Machine-Readable Artefacts - -The following files in `.machine_readable/` contain structured project metadata: -- `STATE.scm` - Current project state and progress -- `META.scm` - Architecture decisions and development practices -- `ECOSYSTEM.scm` - Position in the ecosystem and related projects - -## Architecture - -``` -┌─────────────────────────────────────────────────────────────┐ -│ Elixir Orchestration Layer │ -│ ├── VeriSim.EntityServer (GenServer per entity) │ -│ ├── VeriSim.DriftMonitor (drift detection coordinator) │ -│ ├── VeriSim.QueryRouter (distributes queries) │ -│ └── VeriSim.SchemaRegistry (type system coordinator) │ -│ ↓ HTTP │ -├─────────────────────────────────────────────────────────────┤ -│ Rust Core (verisim-api) │ -│ ├── verisim-graph (Oxigraph RDF/Property Graph) │ -│ ├── verisim-vector (HNSW similarity search) │ -│ ├── verisim-tensor (ndarray/Burn tensors) │ -│ ├── verisim-semantic (CBOR proof blobs) │ -│ ├── verisim-document (Tantivy full-text) │ -│ ├── verisim-temporal (versioning/time-series) │ -│ ├── verisim-provenance (origin/lineage tracking) │ -│ ├── verisim-spatial (geospatial/R-tree) │ -│ ├── verisim-octad (unified entity → octad evolution) │ -│ ├── verisim-drift (drift detection) │ -│ └── verisim-normalizer (self-normalization) │ -└─────────────────────────────────────────────────────────────┘ -``` - -## Design Philosophy (Marr's Three Levels) - -1. **Computational Level**: What problem are we solving? - - Maintain cross-modal consistency across 8 representations of the same entity (octad) - - Detect and repair drift before it causes data quality issues - - Provide unified querying across all modalities - -2. **Algorithmic Level**: How do we solve it? - - Octad entities: one ID, eight synchronized stores - - Drift detection with configurable thresholds - - Self-normalization triggered by drift events - - OTP supervision for fault tolerance - -3. **Implementational Level**: How is it built? - - Rust for performance-critical modality stores - - Elixir/OTP for distributed coordination - - HTTP API for communication - - Prometheus metrics for observability - -## Language Policy - -### ALLOWED -- **Rust** - Core database engine, modality stores -- **Elixir** - OTP orchestration layer -- **ReScript** - VCL parser, federation registry -- **VCL** - VeriSim Consonance Language (native query interface, NOT SQL) - -### BANNED -- Python - Use Rust instead -- Go - Use Rust instead -- Node.js - Use Elixir instead - -## Build Commands - -### Rust Core -```bash -cd rust-core -cargo build -cargo test -cargo clippy -``` - -### Elixir Orchestration -```bash -cd elixir-orchestration -mix deps.get -mix compile -mix test -``` - -### Full Build -```bash -# Rust first (Elixir depends on it at runtime) -cargo build --release - -# Then Elixir -cd elixir-orchestration && mix compile -``` - -## Container Deployment - -Use Podman (NOT Docker): -```bash -# In-memory (default) -podman build -t verisimdb:latest -f container/Containerfile . -podman run -p 8080:8080 verisimdb:latest - -# Persistent storage (redb graph + file-backed Tantivy + WAL) -podman build -t verisimdb:persistent --build-arg FEATURES=persistent -f container/Containerfile . -podman run -p 8080:8080 -v verisimdb-data:/data verisimdb:persistent -``` - -### Verified Container Deployment (stapeln) - -For supply-chain-verified deployment with the stapeln ecosystem: -```bash -# Build, sign, verify as .ctp bundle -cd container && ./ct-build.sh persistent --push - -# Full stack via selur-compose (rust-core + elixir + svalinn gateway) -selur-compose up --detach -``` - -Key files in `container/`: -- `compose.toml` — selur-compose stack definition (3 services + volumes + networks) -- `.gatekeeper.yaml` — svalinn edge gateway policy (auth, rate limits, trust) -- `manifest.toml` — cerro-torre .ctp bundle manifest (provenance, attestations, security) -- `ct-build.sh` — build/sign/verify pipeline script - -## Test Infrastructure - -Integration test stack in `connectors/test-infra/`: - -```bash -# Start test databases -cd connectors/test-infra && selur-compose up -d -# Or fallback: podman-compose up -d - -# Run integration tests -cd elixir-orchestration && mix test --include integration - -# Stop stack -cd connectors/test-infra && selur-compose down -``` - -Services: MongoDB (27017), Redis Stack (6379), Neo4j (7474/7687), ClickHouse (8123/9000), SurrealDB (8000), InfluxDB (8086), MinIO (9002/9001). - -All images use `cgr.dev/chainguard/wolfi-base:latest`. - -## Key Concepts - -### Octad Entity (formerly Octad) -An Octad is one entity with 8 synchronized representations: -- **Graph**: RDF triples and property graph edges -- **Vector**: Embedding for similarity search -- **Tensor**: Multi-dimensional representation — active research into novel applications (details forthcoming) -- **Semantic**: Type annotations and proof blobs -- **Document**: Full-text searchable content -- **Temporal**: Version history and time-series -- **Provenance**: Origin tracking, transformation chain, actor trail (implemented — hash-chain integrity, actor search) -- **Spatial**: Geospatial coordinates, geometries, proximity queries (implemented — R-tree index, radius/bounds/nearest search) - -### Drift Detection -Drift is measured as divergence between modalities: -- `semantic_vector_drift`: Embedding doesn't match semantic content -- `graph_document_drift`: Graph structure doesn't match document -- `temporal_consistency_drift`: Version history issues -- `tensor_drift`: Tensor representation diverged -- `schema_drift`: Type constraint violations -- `quality_drift`: Overall data quality - -### Self-Normalization -When drift exceeds thresholds, the normalizer: -1. Identifies the most authoritative modality -2. Regenerates drifted modalities from it -3. Validates consistency -4. Updates all modalities atomically - -**StorageRegenerator** (production): Real OctadStore-backed implementation in `rust-core/verisim-normalizer/src/storage_regenerator.rs`. Replaces the dry-run SummaryRegenerator. Implements: -- Document->Vector: FNV-1a trigram hashing to 384-dim embedding -- Document->Semantic: keyword extraction as type annotations -- Document/Semantic/Graph cross-regeneration (6 source->target pairs) -- Weighted merge for Vector/Semantic targets -- Cosine similarity drift measurement (Vector), Jaccard index (Semantic) -- 68 normalizer tests pass (7 StorageRegenerator-specific) - -**NormalizerError variants:** NormalizationFailed, StrategyNotFound, OctadError, ChannelError, MissingModality, StorageError, NoViableSource - -## Code Patterns - -### Creating a Octad (Rust) -```rust -let input = OctadBuilder::new() - .with_document("Title", "Body content") - .with_embedding(vec![0.1, 0.2, ...]) - .with_types(vec!["http://example.org/Document"]) - .with_relationships(vec![("relates_to", "other-entity-id")]) - .build(); - -let octad = store.create(input).await?; -``` - -### Entity Server (Elixir) -```elixir -# Start entity server -{:ok, _pid} = VeriSim.EntityServer.start_link("entity-123") - -# Get state -{:ok, state} = VeriSim.EntityServer.get("entity-123") - -# Update -{:ok, new_state} = VeriSim.EntityServer.update("entity-123", [ - {:modality, :vector, true} -]) -``` - -## Testing - -### Unit Tests -```bash -cargo test # Rust -mix test # Elixir -``` - -### Integration Tests -```bash -cargo test --test integration # Rust integration tests -mix test test/integration # Elixir integration tests -``` - -### Federation Adapter Integration Tests -```bash -# Requires test-infra stack running (see Test Infrastructure section above) -cd elixir-orchestration && mix test --include integration -# 105 tests across 7 adapter test files (MongoDB, Redis, Neo4j, ClickHouse, SurrealDB, InfluxDB, MinIO) -``` - -## GitHub CI Integration (Priority - Sonnet Task) - -VeriSimDB needs to service all ~290 hyperpolymath repos from GitHub Actions CI. - -### Architecture: Git-Backed Flat-File Store - -Instead of running verisimdb as a persistent server in GitHub, use a **git-backed data repo**: - -``` -hyperpolymath/verisimdb-data (new repo) -├── scans/ # panic-attack scan results per repo -│ ├── echidna.json -│ ├── verisimdb.json -│ └── ... -├── hardware/ # hardware-crash-team findings -│ └── latest-scan.json -├── drift/ # drift detection snapshots -│ └── drift-status.json -├── index.json # Master index of all octads -└── .github/workflows/ - └── ingest.yml # Workflow: receive data, update index -``` - -### How It Works - -1. **Each repo's CI** runs panic-attack, produces JSON, pushes to `verisimdb-data` via workflow dispatch -2. **verisimdb-data ingest workflow** receives the JSON, stores it, updates the index -3. **Query** by checking out verisimdb-data and reading JSON (no server needed) -4. **Local dev**: Run `verisim-api` server locally, load from the data repo - -### Implementation Steps (for Sonnet) - -1. Create `verisimdb-data` repo from rsr-template-repo -2. Add `ingest.yml` workflow that accepts repository_dispatch events with scan payloads -3. Add a reusable workflow `scan-and-report.yml` that repos can call: - - Runs `panic-attack assail` on the repo - - Sends results to verisimdb-data via repository_dispatch -4. Add the reusable workflow to 2-3 pilot repos first (echidna, panic-attacker, ambientops) -5. Add a `query.sh` script that clones verisimdb-data and searches the index - -### Future: Persistent Server - -When ready to scale beyond flat files: -- Deploy verisim-api to **Fly.io free tier** (3 shared VMs, 1GB persistent volume) -- Use the Containerfile already in `container/` -- GitHub Actions calls the Fly.io endpoint instead of repository_dispatch -- Keep verisimdb-data as backup/mirror - -## Hypatia Integration Pipeline - -### Data Flow (IMPLEMENTED) - -``` -panic-attack assail → ScanIngester → octad octads → PatternQuery → DispatchBridge → gitbot-fleet - ↑ WORKS ↑ WORKS ↑ WORKS ↑ WORKS ↑ JSONL logged -``` - -### VeriSimDB-Side Modules (elixir-orchestration/lib/verisim/hypatia/) - -1. **ScanIngester** (`scan_ingester.ex`): Ingests panic-attack scan results as octad octad entities - - Builds Document (searchable text), Graph (triples), Temporal (timestamps), Vector (embeddings), - Provenance (scanner origin), Semantic (category tags) modalities - - Falls back to ETS (`:hypatia_scans`) when Rust core unavailable - - API: `ingest_scan/1`, `ingest_file/1`, `ingest_directory/1`, `list_scans/0`, `get_scan/1` - -2. **PatternQuery** (`pattern_query.ex`): Cross-repo pattern analytics over ingested scans - - API: `pipeline_health/0`, `cross_repo_patterns/1`, `severity_distribution/0`, - `category_distribution/0`, `temporal_trends/1`, `repos_by_severity/1`, `weakness_hotspots/0` - -3. **DispatchBridge** (`dispatch_bridge.ex`): Bridge to Hypatia dispatch pipeline - - Reads JSONL dispatch manifests from `verisimdb-data/dispatch/` - - Tracks execution status and feeds outcomes back for drift tracking - - API: `read_pending/1`, `read_dispatch_log/2`, `read_all_dispatch_logs/1`, - `read_outcomes/1`, `summarize/1`, `feedback_to_drift/1`, `ingest_dispatch_summary/1` - -### Remaining: Fleet Dispatch (Live Execution) - -Fleet dispatch is logged to JSONL but not yet executing live GraphQL mutations. -Requires GitHub PAT with `repo` scope — see `verisimdb-data/INTEGRATION.md`. - -## Model Router (Future Tool - Sonnet Task) - -A tool to auto-select Claude model based on task complexity. Architecture: - -``` -User prompt → Haiku classifier → Route to: - ├── Haiku: single-file edits, template creation, simple queries - ├── Sonnet: multi-file implementation, feature work, testing - └── Opus: architecture decisions, debugging, cross-repo design -``` - -### Classification Signals -- File count affected (1 = Haiku, 2-5 = Sonnet, 5+ = Opus) -- Task type (create = Sonnet, debug = Opus, edit = Haiku) -- CLAUDE.md complexity rating per repo -- Presence of "why", "how", "design" in prompt → Opus -- Presence of "add", "create", "implement" → Sonnet -- Presence of "fix typo", "rename", "update version" → Haiku - -### Implementation -- Rust CLI tool or Claude Code hook -- Reads CLAUDE.md to understand repo complexity -- Uses Haiku API call (~0.001 cents) to classify -- Returns recommended model as stdout - -## Known Issues - -See `KNOWN-ISSUES.adoc` at repo root for all honest gaps. All 25 issues resolved. - -Resolved in recent sessions: -- VCL-UT type checker wired end-to-end (Elixir-native + ReScript + Rust ZKP bridge) -- 11 proof types: EXISTENCE, INTEGRITY, CONSISTENCY, PROVENANCE, FRESHNESS, ACCESS, CITATION, CUSTOM, ZKP, PROVEN, SANCTIFY -- Multi-proof parsing: PROOF A(x) AND B(y) splits correctly -- Modality compatibility validation (INTEGRITY needs semantic, PROVENANCE needs provenance, etc.) -- proven library integrated (certificate-based JSON/CBOR bridge) -- verisim-repl builds clean (67 tests pass) -- oxrocksdb-sys C++ dependency eliminated (Oxigraph feature-flagged, redb pure-Rust backend added) -- protoc build dependency eliminated (proto code pre-generated) -- stapeln container ecosystem integrated (compose.toml, .gatekeeper.yaml, manifest.toml, ct-build.sh) -- VCL Playground wired to real backend (ApiClient.res, async execution, demo mode fallback, octad modalities) -- PanLL database module protocol (DatabaseModule.res, DatabaseRegistry.res — VeriSimDB/QuandleDB/LithoGlyph) -- Product telemetry: opt-in collector (ETS), reporter (JSON), 19 telemetry tests, VCL executor + drift monitor wired -- PanLL telemetry dashboard panel with modality heatmap, query patterns, performance metrics - -## Hypatia Integration Status - -**Working (VeriSimDB side — 3 modules, 37 tests):** -- ScanIngester: panic-attack JSON → octad octads (Document, Graph, Temporal, Vector, Provenance, Semantic) -- PatternQuery: cross-repo analytics (pipeline health, severity distribution, temporal trends, hotspots) -- DispatchBridge: reads JSONL dispatch manifests, summarizes outcomes, feeds drift tracking -- Hypatia VCL layer reads verisimdb-data flat files directly -- Built-in Elixir VCL parser (no external Deno/Node needed) -- 954 canonical patterns tracked across 298 repos - -**Needs PAT:** Automated cross-repo dispatch requires a GitHub PAT with `repo` scope. -See `verisimdb-data/INTEGRATION.md` for PAT setup instructions. - -## User Preferences - -- **Container runtime**: Podman > Docker -- **Source hosting**: GitLab > GitHub -- **Package manager**: Cargo (Rust), Mix (Elixir) -- **No Python**: Use Rust for systems, Julia for data processing - -## Repository Structure - -``` -verisimdb/ -├── Cargo.toml # Workspace definition -├── rust-core/ # Rust crates -│ ├── verisim-graph/ -│ ├── verisim-vector/ -│ ├── verisim-tensor/ -│ ├── verisim-semantic/ -│ ├── verisim-document/ -│ ├── verisim-temporal/ -│ ├── verisim-octad/ -│ ├── verisim-drift/ -│ ├── verisim-normalizer/ -│ └── verisim-api/ -├── elixir-orchestration/ # Elixir/OTP layer -│ ├── lib/verisim/ -│ ├── config/ -│ └── mix.exs -├── connectors/ # Federation adapters + client SDKs + test infra -│ ├── clients/ # 6 SDKs: Rust, V, Elixir, ReScript, Julia, Gleam -│ ├── shared/ # JSON Schema, OpenAPI, protobuf -│ └── test-infra/ # selur-compose: 7 databases for integration testing -├── container/ # Containerfiles -├── docs/ # Documentation -└── tests/ # Integration tests -``` diff --git a/verisimdb/.clusterfuzzlite/project.yaml b/verisimdb/.clusterfuzzlite/project.yaml deleted file mode 100644 index 20b88c4a..00000000 --- a/verisimdb/.clusterfuzzlite/project.yaml +++ /dev/null @@ -1,6 +0,0 @@ -# SPDX-License-Identifier: MPL-2.0 -# ClusterFuzzLite configuration for VeriSimDB -# OpenSSF Scorecard Fuzzing compliance - -language: rust -main_repo: 'https://github.com/hyperpolymath/verisimdb' diff --git a/verisimdb/.editorconfig b/verisimdb/.editorconfig deleted file mode 100644 index fc6650ce..00000000 --- a/verisimdb/.editorconfig +++ /dev/null @@ -1,68 +0,0 @@ -# RSR-template-repo - Editor Configuration -# https://editorconfig.org - -root = true - -[*] -charset = utf-8 -end_of_line = lf -indent_size = 2 -indent_style = space -insert_final_newline = true -trim_trailing_whitespace = true - -[*.md] -trim_trailing_whitespace = false - -[*.adoc] -trim_trailing_whitespace = false - -[*.rs] -indent_size = 4 - -[*.ex] -indent_size = 2 - -[*.exs] -indent_size = 2 - -[*.zig] -indent_size = 4 - -[*.ada] -indent_size = 3 - -[*.adb] -indent_size = 3 - -[*.ads] -indent_size = 3 - -[*.hs] -indent_size = 2 - -[*.res] -indent_size = 2 - -[*.resi] -indent_size = 2 - -[*.ncl] -indent_size = 2 - -[*.rkt] -indent_size = 2 - -[*.scm] -indent_size = 2 - -[*.nix] -indent_size = 2 - -[Justfile] -indent_style = space -indent_size = 4 - -[justfile] -indent_style = space -indent_size = 4 diff --git a/verisimdb/.gitattributes b/verisimdb/.gitattributes deleted file mode 100644 index e860a85c..00000000 --- a/verisimdb/.gitattributes +++ /dev/null @@ -1,54 +0,0 @@ -# SPDX-License-Identifier: MPL-2.0 -# RSR-compliant .gitattributes - -* text=auto eol=lf - -# Source -*.rs text eol=lf diff=rust -*.ex text eol=lf diff=elixir -*.exs text eol=lf diff=elixir -*.jl text eol=lf -*.res text eol=lf -*.resi text eol=lf -*.ada text eol=lf diff=ada -*.adb text eol=lf diff=ada -*.ads text eol=lf diff=ada -*.hs text eol=lf -*.chpl text eol=lf -*.scm text eol=lf -*.ncl text eol=lf -*.nix text eol=lf - -# Docs -*.md text eol=lf diff=markdown -*.adoc text eol=lf -*.txt text eol=lf - -# Data -*.json text eol=lf -*.yaml text eol=lf -*.yml text eol=lf -*.toml text eol=lf - -# Config -.gitignore text eol=lf -.gitattributes text eol=lf -justfile text eol=lf -Makefile text eol=lf -Containerfile text eol=lf - -# Scripts -*.sh text eol=lf - -# Binary -*.png binary -*.jpg binary -*.gif binary -*.pdf binary -*.woff2 binary -*.zip binary -*.gz binary - -# Lock files -Cargo.lock text eol=lf -diff -flake.lock text eol=lf -diff diff --git a/verisimdb/.github/CODEOWNERS b/verisimdb/.github/CODEOWNERS deleted file mode 100644 index 890c0e00..00000000 --- a/verisimdb/.github/CODEOWNERS +++ /dev/null @@ -1,10 +0,0 @@ -# SPDX-License-Identifier: MPL-2.0 -# VeriSimDB Code Owners -# See https://docs.github.com/en/repositories/managing-your-repositorys-settings-and-features/customizing-your-repository/about-code-owners - -* @hyperpolymath -/rust-core/ @hyperpolymath -/elixir-orchestration/ @hyperpolymath -/src/vcl/ @hyperpolymath -/container/ @hyperpolymath -/.github/ @hyperpolymath diff --git a/verisimdb/.github/FUNDING.yml b/verisimdb/.github/FUNDING.yml deleted file mode 100644 index 688a442c..00000000 --- a/verisimdb/.github/FUNDING.yml +++ /dev/null @@ -1,7 +0,0 @@ -# SPDX-License-Identifier: MPL-2.0 -# Funding platforms for hyperpolymath projects -# See: https://docs.github.com/en/repositories/managing-your-repositorys-settings-and-features/customizing-your-repository/displaying-a-sponsor-button-in-your-repository - -github: hyperpolymath -ko_fi: hyperpolymath -liberapay: hyperpolymath diff --git a/verisimdb/.github/ISSUE_TEMPLATE/bug_report.md b/verisimdb/.github/ISSUE_TEMPLATE/bug_report.md deleted file mode 100644 index 987aab6b..00000000 --- a/verisimdb/.github/ISSUE_TEMPLATE/bug_report.md +++ /dev/null @@ -1,38 +0,0 @@ ---- -name: Bug report -about: Create a report to help us improve -title: "[Bug]: " -labels: 'bug, priority: unset, triage' -assignees: '' - ---- - -**Describe the bug** -A clear and concise description of what the bug is. - -**To Reproduce** -Steps to reproduce the behavior: -1. Go to '...' -2. Click on '....' -3. Scroll down to '....' -4. See error - -**Expected behavior** -A clear and concise description of what you expected to happen. - -**Screenshots** -If applicable, add screenshots to help explain your problem. - -**Desktop (please complete the following information):** - - OS: [e.g. iOS] - - Browser [e.g. chrome, safari] - - Version [e.g. 22] - -**Smartphone (please complete the following information):** - - Device: [e.g. iPhone6] - - OS: [e.g. iOS8.1] - - Browser [e.g. stock browser, safari] - - Version [e.g. 22] - -**Additional context** -Add any other context about the problem here. diff --git a/verisimdb/.github/ISSUE_TEMPLATE/custom.md b/verisimdb/.github/ISSUE_TEMPLATE/custom.md deleted file mode 100644 index 48d5f81f..00000000 --- a/verisimdb/.github/ISSUE_TEMPLATE/custom.md +++ /dev/null @@ -1,10 +0,0 @@ ---- -name: Custom issue template -about: Describe this issue template's purpose here. -title: '' -labels: '' -assignees: '' - ---- - - diff --git a/verisimdb/.github/ISSUE_TEMPLATE/documentation.md b/verisimdb/.github/ISSUE_TEMPLATE/documentation.md deleted file mode 100644 index 4fcb9f9f..00000000 --- a/verisimdb/.github/ISSUE_TEMPLATE/documentation.md +++ /dev/null @@ -1,66 +0,0 @@ ---- -name: Documentation -about: Report unclear, missing, or incorrect documentation -title: "[DOCS]: " -labels: 'documentation, priority: unset, triage' -assignees: '' - ---- - -name: Documentation -description: Report unclear, missing, or incorrect documentation -title: "[Docs]: " -labels: ["documentation", "triage"] -body: - - type: markdown - attributes: - value: | - Help us improve our documentation by reporting issues or gaps. - - - type: dropdown - id: type - attributes: - label: Documentation issue type - options: - - Missing (documentation doesn't exist) - - Incorrect (information is wrong) - - Unclear (confusing or hard to follow) - - Outdated (no longer accurate) - - Typo or grammar - validations: - required: true - - - type: input - id: location - attributes: - label: Location - description: Where is this documentation? (URL, file path, or section name) - placeholder: README.adoc, section "Installation" - validations: - required: true - - - type: textarea - id: description - attributes: - label: Description - description: What's the problem with the current documentation? - placeholder: Describe what's wrong or missing - validations: - required: true - - - type: textarea - id: suggestion - attributes: - label: Suggested improvement - description: How should it be fixed or improved? - placeholder: The documentation should say... - validations: - required: false - - - type: checkboxes - id: contribution - attributes: - label: Contribution - options: - - label: I would be willing to submit a PR to fix this - required: false diff --git a/verisimdb/.github/ISSUE_TEMPLATE/feature_request.md b/verisimdb/.github/ISSUE_TEMPLATE/feature_request.md deleted file mode 100644 index 3e8fa7e7..00000000 --- a/verisimdb/.github/ISSUE_TEMPLATE/feature_request.md +++ /dev/null @@ -1,20 +0,0 @@ ---- -name: Feature request -about: Suggest an idea for this project -title: '' -labels: 'enhancement, priority: unset, triage' -assignees: '' - ---- - -**Is your feature request related to a problem? Please describe.** -A clear and concise description of what the problem is. Ex. I'm always frustrated when [...] - -**Describe the solution you'd like** -A clear and concise description of what you want to happen. - -**Describe alternatives you've considered** -A clear and concise description of any alternative solutions or features you've considered. - -**Additional context** -Add any other context or screenshots about the feature request here. diff --git a/verisimdb/.github/ISSUE_TEMPLATE/question.md b/verisimdb/.github/ISSUE_TEMPLATE/question.md deleted file mode 100644 index fd0e2a5c..00000000 --- a/verisimdb/.github/ISSUE_TEMPLATE/question.md +++ /dev/null @@ -1,55 +0,0 @@ ---- -name: Question -about: Ask a question about usage or behaviour -title: "[QUESTION]: " -labels: question, triage -assignees: '' - ---- - -name: Question -description: Ask a question about usage or behaviour -title: "[Question]: " -labels: ["question", "triage"] -body: - - type: markdown - attributes: - value: | - Have a question? You can also ask in [Discussions](../discussions) for broader conversations. - - - type: textarea - id: question - attributes: - label: Your question - description: What would you like to know? - placeholder: How do I...? - validations: - required: true - - - type: textarea - id: context - attributes: - label: Context - description: Any relevant context that helps us answer your question - placeholder: I'm trying to achieve X and I've tried Y... - validations: - required: false - - - type: textarea - id: research - attributes: - label: What I've already tried - description: What have you already looked at or attempted? - placeholder: I've read the README and searched issues but... - validations: - required: false - - - type: checkboxes - id: checked - attributes: - label: Pre-submission checklist - options: - - label: I have searched existing issues and discussions - required: true - - label: I have read the documentation - required: true diff --git a/verisimdb/.github/SUPPORT.md b/verisimdb/.github/SUPPORT.md deleted file mode 100644 index 2dfd3c31..00000000 --- a/verisimdb/.github/SUPPORT.md +++ /dev/null @@ -1,30 +0,0 @@ - -# Support - -## How to Get Help - -VeriSimDB is an open-source project maintained by [hyperpolymath](https://github.com/hyperpolymath). - -### For Bug Reports - -Please use the [Bug Report](https://github.com/hyperpolymath/verisimdb/issues/new?template=bug_report.md) issue template. - -### For Feature Requests - -Please use the [Feature Request](https://github.com/hyperpolymath/verisimdb/issues/new?template=feature_request.md) issue template. - -### For Questions - -Please use the [Question](https://github.com/hyperpolymath/verisimdb/issues/new?template=question.md) issue template. - -### For Security Issues - -**Do NOT open a public issue.** Please see [SECURITY.md](../SECURITY.md) for responsible disclosure instructions. - -## Response Times - -This is a solo-maintained project. Response times vary but issues are reviewed regularly. - -## Contributing - -See [CONTRIBUTING.md](../CONTRIBUTING.md) for contribution guidelines. diff --git a/verisimdb/.github/dependabot.yml b/verisimdb/.github/dependabot.yml deleted file mode 100644 index e2dcde95..00000000 --- a/verisimdb/.github/dependabot.yml +++ /dev/null @@ -1,48 +0,0 @@ -# SPDX-License-Identifier: MPL-2.0 -# Dependabot configuration for RSR-compliant repositories -# Covers common ecosystems - remove unused ones for your project - -version: 2 -updates: - # GitHub Actions - always include - - package-ecosystem: "github-actions" - directory: "/" - schedule: - interval: "daily" - groups: - actions: - patterns: - - "*" - - # Rust/Cargo - - package-ecosystem: "cargo" - directory: "/" - schedule: - interval: "daily" - ignore: - - dependency-name: "*" - update-types: ["version-update:semver-patch"] - - # Elixir/Mix - - package-ecosystem: "weeklymix" - directory: "/" - schedule: - interval: "daily" - - # Node.js/npm - - package-ecosystem: "npm" - directory: "/" - schedule: - interval: "daily" - - # Python/pip - - package-ecosystem: "pip" - directory: "/" - schedule: - interval: "daily" - - # Nix flakes - - package-ecosystem: "nix" - directory: "/" - schedule: - interval: "daily" diff --git a/verisimdb/.github/workflows/cflite_batch.yml b/verisimdb/.github/workflows/cflite_batch.yml deleted file mode 100644 index b653b6e6..00000000 --- a/verisimdb/.github/workflows/cflite_batch.yml +++ /dev/null @@ -1,32 +0,0 @@ -# SPDX-License-Identifier: MPL-2.0 -name: ClusterFuzzLite batch fuzzing -on: - schedule: - - cron: '0 0 * * 0' # Weekly on Sunday - workflow_dispatch: - -permissions: read-all - -jobs: - batch-fuzzing: - runs-on: ubuntu-latest - strategy: - fail-fast: false - matrix: - sanitizer: [address, undefined] - steps: - - name: Build Fuzzers (${{ matrix.sanitizer }}) - id: build - uses: google/clusterfuzzlite/actions/build_fuzzers@884713a6c30a92e5e8544c39945cd7cb630abcd1 # v1 - with: - language: rust - sanitizer: ${{ matrix.sanitizer }} - - - name: Run Fuzzers (${{ matrix.sanitizer }}) - id: run - uses: google/clusterfuzzlite/actions/run_fuzzers@884713a6c30a92e5e8544c39945cd7cb630abcd1 # v1 - with: - github-token: ${{ secrets.GITHUB_TOKEN }} - fuzz-seconds: 1800 - mode: 'batch' - sanitizer: ${{ matrix.sanitizer }} diff --git a/verisimdb/.github/workflows/cflite_pr.yml b/verisimdb/.github/workflows/cflite_pr.yml deleted file mode 100644 index 03f06393..00000000 --- a/verisimdb/.github/workflows/cflite_pr.yml +++ /dev/null @@ -1,36 +0,0 @@ -# SPDX-License-Identifier: MPL-2.0 -name: ClusterFuzzLite PR fuzzing -on: - pull_request: - branches: [main] - workflow_dispatch: - -permissions: read-all - -concurrency: - group: ${{ github.workflow }}-${{ github.ref }} - cancel-in-progress: true - -jobs: - pr-fuzzing: - runs-on: ubuntu-latest - strategy: - fail-fast: false - matrix: - sanitizer: [address] - steps: - - name: Build Fuzzers (${{ matrix.sanitizer }}) - id: build - uses: google/clusterfuzzlite/actions/build_fuzzers@884713a6c30a92e5e8544c39945cd7cb630abcd1 # v1 - with: - language: rust - sanitizer: ${{ matrix.sanitizer }} - - - name: Run Fuzzers (${{ matrix.sanitizer }}) - id: run - uses: google/clusterfuzzlite/actions/run_fuzzers@884713a6c30a92e5e8544c39945cd7cb630abcd1 # v1 - with: - github-token: ${{ secrets.GITHUB_TOKEN }} - fuzz-seconds: 300 - mode: 'code-change' - sanitizer: ${{ matrix.sanitizer }} diff --git a/verisimdb/.github/workflows/codeql.yml b/verisimdb/.github/workflows/codeql.yml deleted file mode 100644 index b317db1b..00000000 --- a/verisimdb/.github/workflows/codeql.yml +++ /dev/null @@ -1,40 +0,0 @@ -# SPDX-License-Identifier: MPL-2.0 -name: CodeQL Security Analysis - -on: - push: - branches: [main, master] - pull_request: - branches: [main, master] - schedule: - - cron: '0 6 * * 1' - -permissions: read-all - -jobs: - analyze: - runs-on: ubuntu-latest - permissions: - contents: read - security-events: write - strategy: - fail-fast: false - matrix: - include: - - language: javascript-typescript - build-mode: none - - steps: - - name: Checkout - uses: actions/checkout@8e8c483db84b4bee98b60c0593521ed34d9990e8 # v6.0.1 - - - name: Initialize CodeQL - uses: github/codeql-action/init@cdefb33c0f6224e58673d9004f47f7cb3e328b89 # v3.28.1 - with: - languages: ${{ matrix.language }} - build-mode: ${{ matrix.build-mode }} - - - name: Perform CodeQL Analysis - uses: github/codeql-action/analyze@cdefb33c0f6224e58673d9004f47f7cb3e328b89 # v3.28.1 - with: - category: "/language:${{ matrix.language }}" diff --git a/verisimdb/.github/workflows/governance.yml b/verisimdb/.github/workflows/governance.yml deleted file mode 100644 index b0b1ed6d..00000000 --- a/verisimdb/.github/workflows/governance.yml +++ /dev/null @@ -1,26 +0,0 @@ -# SPDX-License-Identifier: MPL-2.0 -# governance.yml — single wrapper calling the shared estate governance bundle -# in hyperpolymath/standards instead of carrying per-repo copies. -# -# Replaces the per-repo governance scaffolding removed in the same commit: -# quality.yml, guix-nix-policy.yml, npm-bun-blocker.yml, ts-blocker.yml, -# security-policy.yml, rsr-antipattern.yml, wellknown-enforcement.yml, -# workflow-linter.yml -# -# Load-bearing build/security workflows stay standalone in the repo -# (rust-ci, codeql, dependabot, release, scan/mirror/pages plumbing). - -name: Governance - -on: - push: - branches: [main, master] - pull_request: - workflow_dispatch: - -permissions: - contents: read - -jobs: - governance: - uses: hyperpolymath/standards/.github/workflows/governance-reusable.yml@main diff --git a/verisimdb/.github/workflows/hypatia-scan.yml b/verisimdb/.github/workflows/hypatia-scan.yml deleted file mode 100644 index 5b59919d..00000000 --- a/verisimdb/.github/workflows/hypatia-scan.yml +++ /dev/null @@ -1,179 +0,0 @@ -# SPDX-License-Identifier: MPL-2.0 -# Hypatia Neurosymbolic CI/CD Security Scan -name: Hypatia Security Scan - -on: - push: - branches: [ main, master, develop ] - pull_request: - branches: [ main, master ] - schedule: - - cron: '0 0 * * 0' # Weekly on Sunday - workflow_dispatch: - -permissions: read-all - -jobs: - scan: - name: Hypatia Neurosymbolic Analysis - runs-on: ubuntu-latest - - steps: - - name: Checkout repository - uses: actions/checkout@b4ffde65f46336ab88eb53be808477a3936bae11 # v4 - with: - fetch-depth: 0 # Full history for better pattern analysis - - - name: Setup Elixir for Hypatia scanner - uses: erlef/setup-beam@2f0cc07b4b9bea248ae098aba9e1a8a1de5ec24c # v1.18.2 - with: - elixir-version: '1.19.4' - otp-version: '28.3' - - - name: Clone Hypatia - run: | - if [ ! -d "$HOME/hypatia" ]; then - git clone https://github.com/hyperpolymath/hypatia.git "$HOME/hypatia" - fi - - - name: Build Hypatia scanner (if needed) - working-directory: ${{ env.HOME }}/hypatia - run: | - if [ ! -f hypatia-v2 ]; then - echo "Building hypatia-v2 scanner..." - cd scanner - mix deps.get - mix escript.build - mv hypatia ../hypatia-v2 - fi - - - name: Run Hypatia scan - id: scan - run: | - echo "Scanning repository: ${{ github.repository }}" - - # Run scanner - HYPATIA_FORMAT=json "$HOME/hypatia/hypatia-cli.sh" scan . > hypatia-findings.json - - # Count findings - FINDING_COUNT=$(jq '. | length' hypatia-findings.json 2>/dev/null || echo 0) - echo "findings_count=$FINDING_COUNT" >> $GITHUB_OUTPUT - - # Extract severity counts - CRITICAL=$(jq '[.[] | select(.severity == "critical")] | length' hypatia-findings.json) - HIGH=$(jq '[.[] | select(.severity == "high")] | length' hypatia-findings.json) - MEDIUM=$(jq '[.[] | select(.severity == "medium")] | length' hypatia-findings.json) - - echo "critical=$CRITICAL" >> $GITHUB_OUTPUT - echo "high=$HIGH" >> $GITHUB_OUTPUT - echo "medium=$MEDIUM" >> $GITHUB_OUTPUT - - echo "## Hypatia Scan Results" >> $GITHUB_STEP_SUMMARY - echo "- Total findings: $FINDING_COUNT" >> $GITHUB_STEP_SUMMARY - echo "- Critical: $CRITICAL" >> $GITHUB_STEP_SUMMARY - echo "- High: $HIGH" >> $GITHUB_STEP_SUMMARY - echo "- Medium: $MEDIUM" >> $GITHUB_STEP_SUMMARY - - - name: Upload findings artifact - uses: actions/upload-artifact@ea165f8d65b6e75b540449e92b4886f43607fa02 # v4 - with: - name: hypatia-findings - path: hypatia-findings.json - retention-days: 90 - - - name: Submit findings to gitbot-fleet (Phase 2) - if: steps.scan.outputs.findings_count > 0 - env: - GITHUB_TOKEN: ${{ secrets.GITHUB_TOKEN }} - GITHUB_REPOSITORY: ${{ github.repository }} - GITHUB_SHA: ${{ github.sha }} - run: | - echo "📤 Submitting ${{ steps.scan.outputs.findings_count }} findings to gitbot-fleet..." - - # Clone gitbot-fleet to temp directory - FLEET_DIR="/tmp/gitbot-fleet-$$" - git clone https://github.com/hyperpolymath/gitbot-fleet.git "$FLEET_DIR" - - # Run submission script - bash "$FLEET_DIR/scripts/submit-finding.sh" hypatia-findings.json - - # Cleanup - rm -rf "$FLEET_DIR" - - echo "✅ Finding submission complete" - - - name: Check for critical issues - if: steps.scan.outputs.critical > 0 - run: | - echo "⚠️ Critical security issues found!" - echo "Review hypatia-findings.json for details" - # Don't fail the build yet - just warn - # exit 1 - - - name: Generate scan report - run: | - cat << EOF > hypatia-report.md - # Hypatia Security Scan Report - - **Repository:** ${{ github.repository }} - **Scan Date:** $(date -u +"%Y-%m-%d %H:%M:%S UTC") - **Commit:** ${{ github.sha }} - - ## Summary - - | Severity | Count | - |----------|-------| - | Critical | ${{ steps.scan.outputs.critical }} | - | High | ${{ steps.scan.outputs.high }} | - | Medium | ${{ steps.scan.outputs.medium }} | - | **Total**| ${{ steps.scan.outputs.findings_count }} | - - ## Next Steps - - 1. Review findings in the artifact: hypatia-findings.json - 2. Auto-fixable issues will be addressed by robot-repo-automaton (Phase 3) - 3. Manual review required for complex issues - - ## Learning - - These findings feed Hypatia's learning engine to improve future rules. - - --- - *Powered by [Hypatia](https://github.com/hyperpolymath/hypatia) - Neurosymbolic CI/CD Intelligence* - EOF - - cat hypatia-report.md >> $GITHUB_STEP_SUMMARY - - - name: Comment on PR with findings - if: github.event_name == 'pull_request' && steps.scan.outputs.findings_count > 0 - uses: actions/github-script@60a0d83039c74a4aee543508d2ffcb1c3799cdea # v7 - with: - script: | - const fs = require('fs'); - const findings = JSON.parse(fs.readFileSync('hypatia-findings.json', 'utf8')); - - const critical = findings.filter(f => f.severity === 'critical').length; - const high = findings.filter(f => f.severity === 'high').length; - - let comment = `## 🔍 Hypatia Security Scan\n\n`; - comment += `**Findings:** ${findings.length} issues detected\n\n`; - comment += `| Severity | Count |\n|----------|-------|\n`; - comment += `| 🔴 Critical | ${critical} |\n`; - comment += `| 🟠 High | ${high} |\n`; - comment += `| 🟡 Medium | ${findings.length - critical - high} |\n\n`; - - if (critical > 0) { - comment += `⚠️ **Action Required:** Critical security issues found!\n\n`; - } - - comment += `
View findings\n\n`; - comment += `\`\`\`json\n${JSON.stringify(findings.slice(0, 10), null, 2)}\n\`\`\`\n`; - comment += `
\n\n`; - comment += `*Powered by Hypatia Neurosymbolic CI/CD Intelligence*`; - - github.rest.issues.createComment({ - owner: context.repo.owner, - repo: context.repo.repo, - issue_number: context.issue.number, - body: comment - }); diff --git a/verisimdb/.github/workflows/instant-sync.yml b/verisimdb/.github/workflows/instant-sync.yml deleted file mode 100644 index 228dc438..00000000 --- a/verisimdb/.github/workflows/instant-sync.yml +++ /dev/null @@ -1,33 +0,0 @@ -# SPDX-License-Identifier: MPL-2.0 -# Instant Forge Sync - Triggers propagation to all forges on push/release -name: Instant Sync - -on: - push: - branches: [main, master] - release: - types: [published] - -permissions: - contents: read - -jobs: - dispatch: - runs-on: ubuntu-latest - steps: - - name: Trigger Propagation - uses: peter-evans/repository-dispatch@28959ce8df70de7be546dd1250a005dd32156697 # v3 - with: - token: ${{ secrets.FARM_DISPATCH_TOKEN }} - repository: hyperpolymath/.git-private-farm - event-type: propagate - client-payload: |- - { - "repo": "${{ github.event.repository.name }}", - "ref": "${{ github.ref }}", - "sha": "${{ github.sha }}", - "forges": "" - } - - - name: Confirm - run: echo "::notice::Propagation triggered for ${{ github.event.repository.name }}" diff --git a/verisimdb/.github/workflows/jekyll-gh-pages.yml b/verisimdb/.github/workflows/jekyll-gh-pages.yml deleted file mode 100644 index a4882d63..00000000 --- a/verisimdb/.github/workflows/jekyll-gh-pages.yml +++ /dev/null @@ -1,52 +0,0 @@ -# SPDX-License-Identifier: MPL-2.0 -# Sample workflow for building and deploying a Jekyll site to GitHub Pages -name: Deploy Jekyll with GitHub Pages dependencies preinstalled - -on: - # Runs on pushes targeting the default branch - push: - branches: ["main"] - - # Allows you to run this workflow manually from the Actions tab - workflow_dispatch: - -# Sets permissions of the GITHUB_TOKEN to allow deployment to GitHub Pages -permissions: - contents: read - pages: write - id-token: write - -# Allow only one concurrent deployment, skipping runs queued between the run in-progress and latest queued. -# However, do NOT cancel in-progress runs as we want to allow these production deployments to complete. -concurrency: - group: "pages" - cancel-in-progress: false - -jobs: - # Build job - build: - runs-on: ubuntu-latest - steps: - - name: Checkout - uses: actions/checkout@b4ffde65f46336ab88eb53be808477a3936bae11 # v6.0.1 - - name: Setup Pages - uses: actions/configure-pages@983d7736d9b0ae728b81ab479565c72886d7745b # v5 - - name: Build with Jekyll - uses: actions/jekyll-build-pages@44a6e6beabd48582f863aeeb6cb2151cc1716697 # v1 - with: - source: ./ - destination: ./_site - - name: Upload artifact - uses: actions/upload-pages-artifact@56afc609e74202658d3ffba0e8f6dda462b719fa # v4 - - # Deployment job - deploy: - environment: - name: github-pages - url: ${{ steps.deployment.outputs.page_url }} - runs-on: ubuntu-latest - needs: build - steps: - - name: Deploy to GitHub Pages - id: deployment - uses: actions/deploy-pages@d6db90164ac5ed86f2b6aed7e0febac5b3c0c03e # v4 diff --git a/verisimdb/.github/workflows/mirror.yml b/verisimdb/.github/workflows/mirror.yml deleted file mode 100644 index 8fb0c752..00000000 --- a/verisimdb/.github/workflows/mirror.yml +++ /dev/null @@ -1,144 +0,0 @@ -# SPDX-License-Identifier: MPL-2.0 -# SPDX-FileCopyrightText: 2025 Jonathan D.A. Jewell -name: Mirror to Git Forges - -on: - push: - branches: [main] - workflow_dispatch: - -permissions: read-all - -jobs: - mirror-gitlab: - runs-on: ubuntu-latest - if: vars.GITLAB_MIRROR_ENABLED == 'true' - steps: - - uses: actions/checkout@de0fac2e4500dabe0009e67214ff5f5447ce83dd # v6.0.2 - with: - fetch-depth: 0 - - - uses: webfactory/ssh-agent@a6f90b1f127823b31d4d4a8d96047790581349bd # v0.9.1 - with: - ssh-private-key: ${{ secrets.GITLAB_SSH_KEY }} - - - name: Mirror to GitLab - run: | - ssh-keyscan -t ed25519 gitlab.com >> ~/.ssh/known_hosts - git remote add gitlab git@gitlab.com:${{ vars.GITLAB_ORG || vars.MIRROR_ORG || github.repository_owner }}/${{ github.event.repository.name }}.git || true - git push --force gitlab main - - mirror-bitbucket: - runs-on: ubuntu-latest - if: vars.BITBUCKET_MIRROR_ENABLED == 'true' - steps: - - uses: actions/checkout@de0fac2e4500dabe0009e67214ff5f5447ce83dd # v6.0.2 - with: - fetch-depth: 0 - - - uses: webfactory/ssh-agent@a6f90b1f127823b31d4d4a8d96047790581349bd # v0.9.1 - with: - ssh-private-key: ${{ secrets.BITBUCKET_SSH_KEY }} - - - name: Mirror to Bitbucket - run: | - ssh-keyscan -t ed25519 bitbucket.org >> ~/.ssh/known_hosts - git remote add bitbucket git@bitbucket.org:${{ vars.BITBUCKET_ORG || vars.MIRROR_ORG || github.repository_owner }}/${{ github.event.repository.name }}.git || true - git push --force bitbucket main - - mirror-codeberg: - runs-on: ubuntu-latest - if: vars.CODEBERG_MIRROR_ENABLED == 'true' - steps: - - uses: actions/checkout@de0fac2e4500dabe0009e67214ff5f5447ce83dd # v6.0.2 - with: - fetch-depth: 0 - - - uses: webfactory/ssh-agent@a6f90b1f127823b31d4d4a8d96047790581349bd # v0.9.1 - with: - ssh-private-key: ${{ secrets.CODEBERG_SSH_KEY }} - - - name: Mirror to Codeberg - run: | - ssh-keyscan -t ed25519 codeberg.org >> ~/.ssh/known_hosts - git remote add codeberg git@codeberg.org:${{ vars.CODEBERG_ORG || vars.MIRROR_ORG || github.repository_owner }}/${{ github.event.repository.name }}.git || true - git push --force codeberg main - - mirror-sourcehut: - runs-on: ubuntu-latest - if: vars.SOURCEHUT_MIRROR_ENABLED == 'true' - steps: - - uses: actions/checkout@de0fac2e4500dabe0009e67214ff5f5447ce83dd # v6.0.2 - with: - fetch-depth: 0 - - - uses: webfactory/ssh-agent@a6f90b1f127823b31d4d4a8d96047790581349bd # v0.9.1 - with: - ssh-private-key: ${{ secrets.SOURCEHUT_SSH_KEY }} - - - name: Mirror to SourceHut - run: | - ssh-keyscan -t ed25519 git.sr.ht >> ~/.ssh/known_hosts - git remote add sourcehut git@git.sr.ht:~${{ vars.SOURCEHUT_ORG || vars.MIRROR_ORG || github.repository_owner }}/${{ github.event.repository.name }} || true - git push --force sourcehut main - - mirror-disroot: - runs-on: ubuntu-latest - if: vars.DISROOT_MIRROR_ENABLED == 'true' - steps: - - uses: actions/checkout@de0fac2e4500dabe0009e67214ff5f5447ce83dd # v6.0.2 - with: - fetch-depth: 0 - - - uses: webfactory/ssh-agent@a6f90b1f127823b31d4d4a8d96047790581349bd # v0.9.1 - with: - ssh-private-key: ${{ secrets.DISROOT_SSH_KEY }} - - - name: Mirror to Disroot - run: | - ssh-keyscan -t ed25519 git.disroot.org >> ~/.ssh/known_hosts - git remote add disroot git@git.disroot.org:${{ vars.DISROOT_ORG || vars.MIRROR_ORG || github.repository_owner }}/${{ github.event.repository.name }}.git || true - git push --force disroot main - - mirror-gitea: - runs-on: ubuntu-latest - if: vars.GITEA_MIRROR_ENABLED == 'true' - steps: - - uses: actions/checkout@de0fac2e4500dabe0009e67214ff5f5447ce83dd # v6.0.2 - with: - fetch-depth: 0 - - - uses: webfactory/ssh-agent@a6f90b1f127823b31d4d4a8d96047790581349bd # v0.9.1 - with: - ssh-private-key: ${{ secrets.GITEA_SSH_KEY }} - - - name: Mirror to Gitea - run: | - ssh-keyscan -t ed25519 ${{ vars.GITEA_HOST }} >> ~/.ssh/known_hosts - git remote add gitea git@${{ vars.GITEA_HOST }}:${{ vars.GITEA_ORG || vars.MIRROR_ORG || github.repository_owner }}/${{ github.event.repository.name }}.git || true - git push --force gitea main - - mirror-radicle: - runs-on: ubuntu-latest - if: vars.RADICLE_MIRROR_ENABLED == 'true' - steps: - - uses: actions/checkout@de0fac2e4500dabe0009e67214ff5f5447ce83dd # v6.0.2 - with: - fetch-depth: 0 - - - name: Setup Rust - uses: dtolnay/rust-toolchain@efa25f7f19611383d5b0ccf2d1c8914531636bf9 # stable - with: - toolchain: stable - - - name: Install Radicle - run: | - # Install via cargo (safer than curl|sh) - cargo install radicle-cli --locked - echo "$HOME/.cargo/bin" >> $GITHUB_PATH - - - name: Mirror to Radicle - run: | - echo "${{ secrets.RADICLE_KEY }}" > ~/.radicle/keys/radicle - chmod 600 ~/.radicle/keys/radicle - rad sync --announce || echo "Radicle sync attempted" diff --git a/verisimdb/.github/workflows/scorecard-enforcer.yml b/verisimdb/.github/workflows/scorecard-enforcer.yml deleted file mode 100644 index e7c897d5..00000000 --- a/verisimdb/.github/workflows/scorecard-enforcer.yml +++ /dev/null @@ -1,72 +0,0 @@ -# SPDX-License-Identifier: MPL-2.0 -# Prevention workflow - runs OpenSSF Scorecard and fails on low scores -name: OpenSSF Scorecard Enforcer - -on: - push: - branches: [main] - schedule: - - cron: '0 6 * * 1' # Weekly on Monday - workflow_dispatch: - -permissions: read-all - -jobs: - scorecard: - runs-on: ubuntu-latest - permissions: - security-events: write - id-token: write # For OIDC - steps: - - uses: actions/checkout@8e8c483db84b4bee98b60c0593521ed34d9990e8 # v4 - with: - persist-credentials: false - - - name: Run Scorecard - uses: ossf/scorecard-action@4eaacf0543bb3f2c246792bd56e8cdeffafb205a # v2.4.3 - with: - results_file: results.sarif - results_format: sarif - publish_results: true - - - name: Upload SARIF - uses: github/codeql-action/upload-sarif@cdefb33c0f6224e58673d9004f47f7cb3e328b89 # v3 - with: - sarif_file: results.sarif - - - name: Check minimum score - run: | - # Parse score from results - SCORE=$(jq -r '.runs[0].tool.driver.properties.score // 0' results.sarif 2>/dev/null || echo "0") - - echo "OpenSSF Scorecard Score: $SCORE" - - # Minimum acceptable score (0-10 scale) - MIN_SCORE=5 - - if [ "$(echo "$SCORE < $MIN_SCORE" | bc -l)" = "1" ]; then - echo "::error::Scorecard score $SCORE is below minimum $MIN_SCORE" - exit 1 - fi - - # Check specific high-priority items - check-critical: - runs-on: ubuntu-latest - steps: - - uses: actions/checkout@8e8c483db84b4bee98b60c0593521ed34d9990e8 # v4 - - - name: Check SECURITY.md exists - run: | - if [ ! -f "SECURITY.md" ]; then - echo "::error::SECURITY.md is required" - exit 1 - fi - - - name: Check for pinned dependencies - run: | - # Check workflows for unpinned actions - unpinned=$(grep -r "uses:.*@v[0-9]" .github/workflows/*.yml 2>/dev/null | grep -v "#" | head -5 || true) - if [ -n "$unpinned" ]; then - echo "::warning::Found unpinned actions:" - echo "$unpinned" - fi diff --git a/verisimdb/.github/workflows/scorecard.yml b/verisimdb/.github/workflows/scorecard.yml deleted file mode 100644 index d50c271a..00000000 --- a/verisimdb/.github/workflows/scorecard.yml +++ /dev/null @@ -1,32 +0,0 @@ -# SPDX-License-Identifier: MPL-2.0 -name: OSSF Scorecard -on: - push: - branches: [main, master] - schedule: - - cron: '0 4 * * *' - workflow_dispatch: - -permissions: read-all - -jobs: - analysis: - runs-on: ubuntu-latest - permissions: - security-events: write - id-token: write - steps: - - uses: actions/checkout@8e8c483db84b4bee98b60c0593521ed34d9990e8 # v6.0.1 - with: - persist-credentials: false - - - name: Run Scorecard - uses: ossf/scorecard-action@4eaacf0543bb3f2c246792bd56e8cdeffafb205a # v2.3.1 - with: - results_file: results.sarif - results_format: sarif - - - name: Upload results - uses: github/codeql-action/upload-sarif@cdefb33c0f6224e58673d9004f47f7cb3e328b89 # v3.31.8 - with: - sarif_file: results.sarif diff --git a/verisimdb/.github/workflows/secret-scanner.yml b/verisimdb/.github/workflows/secret-scanner.yml deleted file mode 100644 index 83840b33..00000000 --- a/verisimdb/.github/workflows/secret-scanner.yml +++ /dev/null @@ -1,67 +0,0 @@ -# SPDX-License-Identifier: MPL-2.0 -# Prevention workflow - scans for hardcoded secrets before they reach main -name: Secret Scanner - -on: - pull_request: - push: - branches: [main] - -permissions: read-all - -jobs: - trufflehog: - runs-on: ubuntu-latest - steps: - - uses: actions/checkout@8e8c483db84b4bee98b60c0593521ed34d9990e8 # v4 - with: - fetch-depth: 0 # Full history for scanning - - - name: TruffleHog Secret Scan - uses: trufflesecurity/trufflehog@116e7171542d2f1dad8810f00dcfacbe0b809183 # v3 - with: - extra_args: --only-verified --fail - - gitleaks: - runs-on: ubuntu-latest - steps: - - uses: actions/checkout@8e8c483db84b4bee98b60c0593521ed34d9990e8 # v4 - with: - fetch-depth: 0 - - - name: Gitleaks Secret Scan - uses: gitleaks/gitleaks-action@ff98106e4c7b2bc287b24eaf42907196329070c7 # v2 - env: - GITHUB_TOKEN: ${{ secrets.GITHUB_TOKEN }} - - # Rust-specific: Check for hardcoded crypto values - rust-secrets: - runs-on: ubuntu-latest - if: hashFiles('**/Cargo.toml') != '' - steps: - - uses: actions/checkout@8e8c483db84b4bee98b60c0593521ed34d9990e8 # v4 - - - name: Check for hardcoded secrets in Rust - run: | - # Patterns that suggest hardcoded secrets - PATTERNS=( - 'const.*SECRET.*=.*"' - 'const.*KEY.*=.*"[a-zA-Z0-9]{16,}"' - 'const.*TOKEN.*=.*"' - 'let.*api_key.*=.*"' - 'HMAC.*"[a-fA-F0-9]{32,}"' - 'password.*=.*"[^"]+"' - ) - - found=0 - for pattern in "${PATTERNS[@]}"; do - if grep -rn --include="*.rs" -E "$pattern" src/; then - echo "WARNING: Potential hardcoded secret found matching: $pattern" - found=1 - fi - done - - if [ $found -eq 1 ]; then - echo "::error::Potential hardcoded secrets detected. Use environment variables instead." - exit 1 - fi diff --git a/verisimdb/.github/workflows/security-scan.yml b/verisimdb/.github/workflows/security-scan.yml deleted file mode 100644 index d9fff11e..00000000 --- a/verisimdb/.github/workflows/security-scan.yml +++ /dev/null @@ -1,19 +0,0 @@ -# SPDX-License-Identifier: MPL-2.0 - -name: Security Scan - -on: - push: - branches: [main] - schedule: - - cron: '0 0 * * 0' # Weekly on Sunday at midnight - workflow_dispatch: - -permissions: - contents: read - -jobs: - scan: - uses: hyperpolymath/panic-attacker/.github/workflows/scan-and-report.yml@21fc3f00a088c954912936f4a68970621b82c2e6 # main - secrets: - VERISIMDB_PAT: ${{ secrets.VERISIMDB_PAT }} diff --git a/verisimdb/.gitignore b/verisimdb/.gitignore deleted file mode 100644 index 5ba4d5c5..00000000 --- a/verisimdb/.gitignore +++ /dev/null @@ -1,117 +0,0 @@ -# SPDX-License-Identifier: MPL-2.0 -# RSR-compliant .gitignore - merged local+remote - -# OS & Editor -.DS_Store -Thumbs.db -*.swp -*.swo -*~ -.idea/ -.vscode/ - -# Build -/target/ -/_build/ -/build/ -/dist/ -/out/ - -# Dependencies -/node_modules/ -/vendor/ -/deps/ -/.elixir_ls/ - -# Rust -**/*.rs.bk -# Cargo.lock # Keep for binaries - -# Elixir -/cover/ -/doc/ -/elixir-orchestration/_build/ -/elixir-orchestration/deps/ -/elixir-orchestration/cover/ -/elixir-orchestration/doc/ -/elixir-orchestration/*.ez -*.ez -*.beam -erl_crash.dump - -# Julia -*.jl.cov -*.jl.mem -/Manifest.toml - -# ReScript -/lib/bs/ -/.bsb.lock -*.res.mjs - -# Playground build artifacts -/playground/node_modules/ -/playground/lib/ -/playground/public/app.js -/playground/public/app.js.map -/playground/deno.lock - -# Python (SaltStack only) -__pycache__/ -*.py[cod] -.venv/ - -# Ada/SPARK -*.ali -/obj/ -/bin/ - -# Haskell -/.stack-work/ -/dist-newstyle/ - -# Chapel -*.chpl.tmp.* - -# Secrets & Environment -.env -.env.* -.env.local -.env.*.local -*.pem -*.key -secrets/ - -# Test/Coverage -/coverage/ -htmlcov/ - -# Logs -*.log -/logs/ -logs/ - -# Temp -/tmp/ -tmp/ -temp/ -*.tmp -*.bak - -# Data directories (for local dev) -/data/ -/storage/ - -# Container build artifacts -*.tar - -# Crash recovery artifacts -ai-cli-crash-capture/ -target/ -node_modules/ -_build/ -deps/ -.elixir_ls/ -.cache/ -build/ -dist/ diff --git a/verisimdb/.gitlab-ci.yml b/verisimdb/.gitlab-ci.yml deleted file mode 100644 index 7309fa90..00000000 --- a/verisimdb/.gitlab-ci.yml +++ /dev/null @@ -1,175 +0,0 @@ -# SPDX-License-Identifier: MPL-2.0 -# Primary CI/CD - GitLab is the source of truth - -stages: - - security - - lint - - test - - build - -variables: - CARGO_HOME: ${CI_PROJECT_DIR}/.cargo - -cache: - key: ${CI_COMMIT_REF_SLUG} - paths: - - .cargo/ - - target/ - -# ================== -# Security Scanning -# ================== - -trivy: - stage: security - image: aquasec/trivy:latest - script: - - trivy fs --exit-code 0 --severity HIGH,CRITICAL --format table . - - trivy fs --exit-code 1 --severity CRITICAL . - allow_failure: false - -gitleaks: - stage: security - image: zricethezav/gitleaks:latest - script: - - gitleaks detect --source . --verbose --redact - allow_failure: false - -semgrep: - stage: security - image: returntocorp/semgrep - script: - - semgrep --config auto --error . - allow_failure: true - -cargo-audit: - stage: security - image: rust:latest - script: - - cargo install cargo-audit - - cargo audit - rules: - - exists: - - Cargo.toml - -cargo-deny: - stage: security - image: rust:latest - script: - - cargo install cargo-deny - - cargo deny check - rules: - - exists: - - Cargo.toml - allow_failure: true - -mix-audit: - stage: security - image: elixir:latest - script: - - mix local.hex --force - - mix archive.install hex mix_audit --force - - mix deps.get - - mix deps.audit - rules: - - exists: - - mix.exs - allow_failure: true - -# ================== -# Linting -# ================== - -rustfmt: - stage: lint - image: rust:latest - script: - - rustup component add rustfmt - - cargo fmt -- --check - rules: - - exists: - - Cargo.toml - -clippy: - stage: lint - image: rust:latest - script: - - rustup component add clippy - - cargo clippy -- -D warnings - rules: - - exists: - - Cargo.toml - allow_failure: true - -mix-format: - stage: lint - image: elixir:latest - script: - - mix format --check-formatted - rules: - - exists: - - mix.exs - -credo: - stage: lint - image: elixir:latest - script: - - mix local.hex --force - - mix deps.get - - mix credo --strict - rules: - - exists: - - mix.exs - allow_failure: true - -# ================== -# Testing -# ================== - -cargo-test: - stage: test - image: rust:latest - script: - - cargo test --all-features - rules: - - exists: - - Cargo.toml - -mix-test: - stage: test - image: elixir:latest - script: - - mix local.hex --force - - mix deps.get - - mix test - rules: - - exists: - - mix.exs - -# ================== -# Build -# ================== - -cargo-build: - stage: build - image: rust:latest - script: - - cargo build --release - artifacts: - paths: - - target/release/ - expire_in: 1 week - rules: - - exists: - - Cargo.toml - -mix-build: - stage: build - image: elixir:latest - script: - - mix local.hex --force - - mix deps.get - - MIX_ENV=prod mix compile - rules: - - exists: - - mix.exs diff --git a/verisimdb/.machine_readable/6a2/AGENTIC.a2ml b/verisimdb/.machine_readable/6a2/AGENTIC.a2ml deleted file mode 100644 index 1699fe4a..00000000 --- a/verisimdb/.machine_readable/6a2/AGENTIC.a2ml +++ /dev/null @@ -1,34 +0,0 @@ -# SPDX-License-Identifier: MPL-2.0 -# Copyright (c) 2026 Jonathan D.A. Jewell (hyperpolymath) -# -# AGENTIC.a2ml — AI agent constraints and capabilities -[metadata] -version = "0.1.0" -last-updated = "2026-04-11" - -[agent-permissions] -can-edit-source = true -can-edit-tests = true -can-edit-docs = true -can-edit-config = true -can-create-files = true - -[agent-constraints] -# What AI agents must NOT do: -# - Never use banned language patterns (believe_me, unsafeCoerce, etc.) -# - Never commit secrets or credentials -# - Never use banned languages (TypeScript, Python, Go, etc.) -# - Never place state files in repository root (must be in .machine_readable/) -# - Never use AGPL license (use PMPL-1.0-or-later) - -[maintenance-integrity] -fail-closed = true -require-evidence-per-step = true -allow-silent-skip = false -require-rerun-after-fix = true -release-claim-requires-hard-pass = true - -[automation-hooks] -# on-enter: Read 0-AI-MANIFEST.a2ml, then STATE.a2ml -# on-exit: Update STATE.a2ml with session outcomes -# on-commit: Run just validate-rsr diff --git a/verisimdb/.machine_readable/6a2/ECOSYSTEM.a2ml b/verisimdb/.machine_readable/6a2/ECOSYSTEM.a2ml deleted file mode 100644 index 5b198a4d..00000000 --- a/verisimdb/.machine_readable/6a2/ECOSYSTEM.a2ml +++ /dev/null @@ -1,20 +0,0 @@ -# SPDX-License-Identifier: MPL-2.0 -# Copyright (c) 2026 Jonathan D.A. Jewell (hyperpolymath) -# -# ECOSYSTEM.a2ml — Verisimdb ecosystem position -[metadata] -version = "1.0" -last-updated = "2026-04-11" - -[project] -name = "Verisimdb" -purpose = "" -role = "" - -[position-in-ecosystem] -category = "" - -[related-projects] -projects = [ - # No related projects recorded -] diff --git a/verisimdb/.machine_readable/6a2/META.a2ml b/verisimdb/.machine_readable/6a2/META.a2ml deleted file mode 100644 index e2c09eb4..00000000 --- a/verisimdb/.machine_readable/6a2/META.a2ml +++ /dev/null @@ -1,27 +0,0 @@ -# SPDX-License-Identifier: MPL-2.0 -# Copyright (c) 2026 Jonathan D.A. Jewell (hyperpolymath) -# -# META.a2ml — Verisimdb meta-level information -[metadata] -version = "0.1.0" -last-updated = "2026-04-11" - -[project-info] -license = "PMPL-1.0-or-later" -author = "Jonathan D.A. Jewell (hyperpolymath)" - -[architecture-decisions] -decisions = [ - # No ADRs recorded -] - -[development-practices] -versioning = "SemVer" -documentation = "AsciiDoc" -build-tool = "just" - -[maintenance-axes] -scoping-first = true -axis-1 = "must > intend > like" -axis-2 = "corrective > adaptive > perfective" -axis-3 = "systems > compliance > effects" diff --git a/verisimdb/.machine_readable/6a2/NEUROSYM.a2ml b/verisimdb/.machine_readable/6a2/NEUROSYM.a2ml deleted file mode 100644 index e1d34c09..00000000 --- a/verisimdb/.machine_readable/6a2/NEUROSYM.a2ml +++ /dev/null @@ -1,21 +0,0 @@ -# SPDX-License-Identifier: MPL-2.0 -# Copyright (c) 2026 Jonathan D.A. Jewell (hyperpolymath) -# -# NEUROSYM.a2ml — Neurosymbolic integration metadata -[metadata] -version = "0.1.0" -last-updated = "2026-04-11" - -[hypatia-config] -scan-enabled = true -scan-depth = "standard" # quick | standard | deep -report-format = "logtalk" - -[symbolic-rules] -# Custom symbolic rules for this project -# - { name = "no-unsafe-ffi", pattern = "believe_me|unsafeCoerce", severity = "critical" } - -[neural-config] -# Neural pattern detection settings -# confidence-threshold = 0.85 -# model = "hypatia-v2" diff --git a/verisimdb/.machine_readable/6a2/PLAYBOOK.a2ml b/verisimdb/.machine_readable/6a2/PLAYBOOK.a2ml deleted file mode 100644 index 5003fd08..00000000 --- a/verisimdb/.machine_readable/6a2/PLAYBOOK.a2ml +++ /dev/null @@ -1,26 +0,0 @@ -# SPDX-License-Identifier: MPL-2.0 -# Copyright (c) 2026 Jonathan D.A. Jewell (hyperpolymath) -# -# PLAYBOOK.a2ml — Operational playbook -[metadata] -version = "0.1.0" -last-updated = "2026-04-11" - -[deployment] -# method = "gitops" # gitops | manual | ci-triggered -# target = "container" # container | binary | library | wasm - -[incident-response] -# 1. Check .machine_readable/STATE.a2ml for current status -# 2. Review recent commits and CI results -# 3. Run `just validate` to check compliance -# 4. Run `just security` to audit for vulnerabilities - -[release-process] -# 1. Update version in STATE.a2ml, META.a2ml -# 2. Run `just release-preflight` (validate + quality + security + maint-hard-pass) -# 3. Tag and push - -[maintenance-operations] -# Baseline audit: just maint-audit -# Hard release gate: just maint-hard-pass diff --git a/verisimdb/.machine_readable/6a2/STATE.a2ml b/verisimdb/.machine_readable/6a2/STATE.a2ml deleted file mode 100644 index 6cc83b09..00000000 --- a/verisimdb/.machine_readable/6a2/STATE.a2ml +++ /dev/null @@ -1,76 +0,0 @@ -# SPDX-License-Identifier: MPL-2.0 -# Copyright (c) 2026 Jonathan D.A. Jewell (hyperpolymath) -# -# STATE.a2ml — VeriSimDB project state - -[metadata] -project = "verisimdb" -version = "0.1.0" -last-updated = "2026-04-11" -status = "active" -session = "proof_attempts REST API + tantivy 0.26 fix — 2026-04-11" - -[project-context] -name = "VeriSimDB" -purpose = """ -Cross-system entity consistency engine with drift detection, -self-normalisation, and formally verified queries. Octad entities -(8 synchronized modalities). Also hosts the proof_attempts pipeline: -ClickHouse-backed proof attempt log consumed by Hypatia and echidnabot. -""" -completion-percentage = 68 - -[position] -phase = "implementation" # design | implementation | testing | maintenance | archived -maturity = "alpha" # experimental | alpha | beta | production | lts - -[route-to-mvp] -milestones = [ - { id = "M1", title = "8-modality octad store (Rust core)", status = "done" }, - { id = "M2", title = "Elixir OTP orchestration layer", status = "done" }, - { id = "M3", title = "VCL/VCL query parser + executor", status = "done" }, - { id = "M4", title = "Drift detection + self-normalisation", status = "done" }, - { id = "M5", title = "Federation adapters (7 backends)", status = "done" }, - { id = "M6", title = "proof_attempts REST API (GET/POST/strategy/certificates)", status = "done" }, - { id = "M7", title = "Persistent deployment (Fly.io or Stapeln)", status = "planned" }, - { id = "M8", title = "VCL federation executor (multi-store)", status = "planned" }, -] - -[blockers-and-issues] -issues = [ - { id = "B1", description = "VERISIM_CLICKHOUSE_URL not documented in README/DEPLOYMENT.adoc", severity = "low" }, - { id = "B2", description = "proof_attempts GET endpoint returns full rows — large responses for limit>5000", severity = "low" }, - { id = "B3", description = "No authentication on /api/v1/proof_attempts POST — internal-only assumption", severity = "medium" }, -] - -[critical-next-actions] -actions = [ - "Document VERISIM_CLICKHOUSE_URL in DEPLOYMENT.adoc", - "Add pagination (offset param) to GET /api/v1/proof_attempts", - "Consider rate-limiting the proof_attempts POST endpoint", -] - -[session-history] -# 2026-04-17: Echidna ProverKind drift audit (commit 8f573f1 — 28 new -# *TypeChecker variants). VeriSimDB's ProverKind (in -# rust-core/verisim-semantic/src/proven_bridge.rs) is NOT a -# mirror of echidna's: it's a proven-library-domain enum with -# 6 variants (Z3, Lean, Coq, Agda, Idris2, Custom(String)) -# documenting the prover that produced a ProvenCertificate — -# strictly independent of echidna's dispatcher backend list. -# verisim-semantic has no path dep on echidna. No action required. -# Build + 53/53 semantic tests pass post-audit. -# 2026-04-11: Added three proof_attempts routes under /api/v1/: -# GET ?limit=N — list rows for Julia retraining (ClickHouse SELECT) -# POST — insert attempt row (ClickHouse INSERT JSONEachRow) -# GET /strategy?class= — recommendations from mv_proven_certificates -# GET /certificates?class= — PROVEN/pending cert status -# Fixed tantivy 0.26 API break in verisim-document: TopDocs::with_limit(n) -# now requires .order_by_score() — Collector no longer impl'd on TopDocs directly. -# Verified: strategy endpoint returns correct data from 1233 ClickHouse rows. -# 2026-04-05: VCL→VCL rename estate-wide. Tropical bridge wired. -# 2026-04-03: Proof certificate pipeline (PROVEN/SANCTIFY). 11 VCL proof types. - -[maintenance-status] -last-run-utc = "2026-04-11T17:06:00Z" -last-result = "pass" diff --git a/verisimdb/.machine_readable/ENSAID_CONFIG.a2ml b/verisimdb/.machine_readable/ENSAID_CONFIG.a2ml deleted file mode 100644 index 6ddc74ff..00000000 --- a/verisimdb/.machine_readable/ENSAID_CONFIG.a2ml +++ /dev/null @@ -1,154 +0,0 @@ -; SPDX-License-Identifier: MPL-2.0 -; Copyright (c) 2026 Jonathan D.A. Jewell (hyperpolymath) -; -; ENSAID_CONFIG.a2ml — PanLL configuration for VeriSimDB. -; VeriSimDB is a formally verified database with Idris2 ABI, Zig FFI, -; V-lang API gateway, VCL query language, and a registry subsystem. -; This is a key dogfooding project for the full ABI→FFI→Adapter stack. - -; ───────────────────────────────────────────────────────────────────── -; [ensaid] — eNSAID environment identity -; ───────────────────────────────────────────────────────────────────── -[ensaid] -version = "1.0" -name = "verisimdb-dev" -description = "Formally verified database development environment" -humidity = "dry" -; Humidity: "dry" = symbolic-heavy. VeriSimDB is a formally verified system — -; proofs and type checking dominate. Neural assists with query optimisation -; and test generation, but never overrides the formal guarantees. - -; ───────────────────────────────────────────────────────────────────── -; [workspace] — file watching, language detection, build integration -; ───────────────────────────────────────────────────────────────────── -[workspace] -root = "." -watch_extensions = [".idr", ".zig", ".v", ".vcl", ".gleam", ".toml"] -; .idr = ABI definitions, .zig = FFI bridge, .v = API gateway, -; .vcl = VeriSimDB Query Language, .gleam = REPL/tooling -ignore_patterns = ["build/", ".git/", "zig-cache/", "zig-out/", "_build/"] -editor = "vscodium" -build_command = "just build" -test_command = "just test" -language_id = "idris2" -; Primary specification language. Implementation spans Zig + V + Gleam. - -; ───────────────────────────────────────────────────────────────────── -; [preferences] — developer preferences and UI settings -; ───────────────────────────────────────────────────────────────────── -[preferences] -theme = "dark" -font_size = 14 -show_line_numbers = true -auto_save = true -auto_format_on_save = true -tab_size = 2 -panel_layout = "three-column" - -; ───────────────────────────────────────────────────────────────────── -; [panels] — which PanLL panels are enabled for this project -; ───────────────────────────────────────────────────────────────────── -[panels] -enabled = [ - "valence-shell", ; Terminal + command execution - "editor-bridge", ; Editor integration - "build-dashboard", ; Multi-language build status (Idris2 + Zig + V) - "boj", ; BoJ dashboard — VeriSimDB is a BoJ consumer - "database", ; Database inspection — dogfooding VeriSimDB itself - "vm-inspector", ; Query engine / VCL inspection - "security", ; Security scanning (Hypatia) - "reposystem", ; Repository management - "panic-attack", ; Error analysis - "tsdm", ; Type System Design Mode — ABI type exploration - "migration", ; Schema migration tracking -] - -disabled = [ - "network", ; Network handled by V-lang adapter - "container", ; Container builds handled externally - "fleet", ; Single-repo work -] - -; ───────────────────────────────────────────────────────────────────── -; [workflows] — project-specific workflow definitions -; ───────────────────────────────────────────────────────────────────── -[workflows] - -[workflows.zig-on-abi-change] -trigger = "file_save" -match = "src/abi/*.idr" -action = "run_command" -command = "cd ffi/zig && zig build test" -description = "Rebuild and test Zig FFI when Idris2 ABI definitions change" - -[workflows.adapter-on-ffi-change] -trigger = "file_save" -match = "ffi/zig/src/*.zig" -action = "run_command" -command = "cd v-api-gateway && v build ." -description = "Rebuild V-lang API gateway when Zig FFI changes" - -[workflows.vcl-test-on-save] -trigger = "file_save" -match = "src/vcl/*.idr" -action = "run_command" -command = "idris2 --check src/vcl/Parser.idr" -description = "Verify VCL parser when query language definitions change" - -[workflows.registry-test] -trigger = "file_save" -match = "src/registry/*.idr" -action = "run_command" -command = "idris2 --check src/registry/Registry.idr" -description = "Verify registry types when registry definitions change" - -[workflows.debugger-rebuild] -trigger = "file_save" -match = "debugger/src/**" -action = "run_command" -command = "cd debugger && just build" -description = "Rebuild debugger when debugger source changes" - -; ───────────────────────────────────────────────────────────────────── -; [clades] — panel clade taxonomy for this project -; ───────────────────────────────────────────────────────────────────── -[clades] -active = ["builder", "database", "terminal", "scanner", "viewer", "bridge"] -; Database development uses builder (multi-lang compilation), database -; (dogfooding VeriSimDB), terminal (VCL REPL), scanner (security), -; viewer (query plans), bridge (editor integration). - -; ───────────────────────────────────────────────────────────────────── -; [portfolios] — named panel/workflow presets -; ───────────────────────────────────────────────────────────────────── -[portfolios] - -[portfolios.abi-focus] -description = "Formal ABI development — Idris2 type definitions" -panels = ["editor-bridge", "tsdm", "build-dashboard"] -layout = "three-column" -focus = "tsdm" - -[portfolios.ffi-testing] -description = "FFI bridge testing — Zig implementation" -panels = ["valence-shell", "build-dashboard", "panic-attack"] -layout = "three-column" -focus = "build-dashboard" - -[portfolios.query-dev] -description = "VCL query language development" -panels = ["editor-bridge", "vm-inspector", "database"] -layout = "three-column" -focus = "vm-inspector" - -[portfolios.full-stack] -description = "Full VeriSimDB development view" -panels = ["valence-shell", "editor-bridge", "build-dashboard", "boj", "database", "tsdm"] -layout = "three-column" -focus = "build-dashboard" - -[portfolios.integration] -description = "End-to-end integration testing" -panels = ["valence-shell", "database", "panic-attack", "security"] -layout = "two-column" -focus = "valence-shell" diff --git a/verisimdb/.nojekyll b/verisimdb/.nojekyll deleted file mode 100644 index e69de29b..00000000 diff --git a/verisimdb/.verisimdb/config.toml b/verisimdb/.verisimdb/config.toml deleted file mode 100644 index bc5b0a0a..00000000 --- a/verisimdb/.verisimdb/config.toml +++ /dev/null @@ -1,53 +0,0 @@ -# SPDX-License-Identifier: MPL-2.0 -# VeriSimDB Self-Hosted Instance Configuration -# -# This instance stores metadata ABOUT the verisimdb repository itself. -# Dogfooding: the database describes its own development history. - -[instance] -name = "verisimdb-self" -description = "Self-hosted VeriSimDB instance for repository metadata" -version = "0.1.0-alpha" -base_iri = "https://verisim.db/self" - -[storage] -# Flat-file JSON storage (committed to git) -backend = "flat-json" -octad_dir = ".verisimdb/octads" -index_file = ".verisimdb/index.json" - -[ingest] -# Sources to ingest automatically -sources = ["git-log", "known-issues", "panic-attack-scans", "benchmarks"] - -[ingest.git-log] -# Map git commits to octads -# Document: commit message + diff stats -# Graph: file dependencies (which files changed together) -# Vector: embedding of commit message (placeholder until model integrated) -# Tensor: [insertions, deletions, files_changed] per commit -# Semantic: commit type (feat/fix/docs/chore) as typed annotation -# Temporal: commit timestamp + author -enabled = true - -[ingest.known-issues] -# Parse KNOWN-ISSUES.adoc entries into octads -enabled = true - -[ingest.panic-attack] -# Ingest panic-attack scan results -scan_dir = "/tmp" -enabled = false - -[ingest.benchmarks] -# Ingest Criterion benchmark results -bench_dir = "target/criterion" -enabled = false - -[modalities] -vector_dimension = 64 # Small dimension for commit embeddings (hash-based) - -[display] -# How to present octads in CLI/TUI -date_format = "%Y-%m-%d %H:%M" -max_title_length = 80 diff --git a/verisimdb/.verisimdb/index.json b/verisimdb/.verisimdb/index.json deleted file mode 100644 index c844a8b7..00000000 --- a/verisimdb/.verisimdb/index.json +++ /dev/null @@ -1,12 +0,0 @@ -{ - "instance": "verisimdb-self", - "version": "0.1.0-alpha", - "base_iri": "https://verisim.db/self", - "generated_at": "2026-02-13T16:20:02Z", - "stats": { - "total_octads": 131, - "commits": 108, - "known_issues": 23 - }, - "octad_ids": ["commit-0087599f9fbd","commit-030287996c2c","commit-031ad82afba6","commit-04ce420ca876","commit-0565bb065c0b","commit-05b5a1a2cbaa","commit-0955ca7ada21","commit-0ba909ba3a0b","commit-0ca572654610","commit-17f903f9fa44","commit-1c5b28a1141e","commit-208a70f82714","commit-20ea93a1be96","commit-228ae4bba134","commit-23815332901c","commit-242713b9c58c","commit-2723614d8844","commit-2eb26c6b7499","commit-2ee7e397d41c","commit-30228e2f8185","commit-3327beaaa2b3","commit-33c8f1c8d141","commit-36907e025d4a","commit-37139e0b0ba9","commit-37dcacf06bf6","commit-3beb018c4b99","commit-3d7765bff845","commit-3da4daef75f9","commit-3df322443379","commit-3dfad170ecfb","commit-45e3a0230e23","commit-4a0cbb016b03","commit-4ed47fa921d0","commit-50ad8033edc3","commit-529b85ff223e","commit-5427006827c3","commit-57aba0bd8ff4","commit-591b7006316c","commit-5e782e158b79","commit-5f6d0db4cc9d","commit-63f2c2dec2ce","commit-667f86eff63a","commit-6d29ea7de601","commit-71dc72521bef","commit-737a6e822c52","commit-74f45a7491e2","commit-77d2e9f3f088","commit-7b8c073708d5","commit-7cec9f8d1b08","commit-7de3adf6c8a0","commit-8012d86a0882","commit-854ea6d13ae7","commit-86581c2638b8","commit-89ea7188af80","commit-8cd1f416878e","commit-91a08d99ee55","commit-9746771f7457","commit-980d6c7a0dc2","commit-9cce50aa9df3","commit-9cdf85099304","commit-9d353c5546f5","commit-9e2298460f12","commit-9f5c2f37c3f6","commit-a0d832b5068b","commit-a31e6e33e51f","commit-a782219c1f2a","commit-a9af6511d111","commit-aa63510d99bd","commit-b1038bf033e2","commit-b62522fca6dc","commit-b9d4aae8f5ce","commit-ba1f53543438","commit-bc2502d10f34","commit-bcbfd32da8b9","commit-bcd5e7cd743f","commit-c0d8094e076f","commit-ccb432c96a9d","commit-cf2958acf17c","commit-d39524c642db","commit-d4e6f6be1200","commit-d515ccac6a03","commit-d6ac4d5f77c0","commit-d7a13170c87c","commit-d8174107cef6","commit-d8593016e679","commit-d949b42717bb","commit-d961f137140e","commit-dd8890c42dde","commit-ddc83dd74b3d","commit-de2fd383992a","commit-dfe015d2ea26","commit-e594e11c006e","commit-e6dbf191a423","commit-e7bd5a33403b","commit-e84118929733","commit-e9b738078138","commit-ea9f52edc373","commit-ecceaddc2f85","commit-ee34606ba4e1","commit-ef3bf59b298f","commit-f3821bdec9bb","commit-f72914d64062","commit-f75fc83b4141","commit-fb03b412d792","commit-fbf2b307da8e","commit-fc71351de8ee","commit-fd3b385d4dd8","commit-fd7ddf29cec9","issue-001","issue-002","issue-003","issue-004","issue-005","issue-006","issue-007","issue-008","issue-009","issue-010","issue-011","issue-012","issue-013","issue-014","issue-015","issue-016","issue-017","issue-018","issue-019","issue-020","issue-021","issue-022","issue-023"] -} diff --git a/verisimdb/.verisimdb/octads/commit-0087599f9fbd.json b/verisimdb/.verisimdb/octads/commit-0087599f9fbd.json deleted file mode 100644 index 0a9829f8..00000000 --- a/verisimdb/.verisimdb/octads/commit-0087599f9fbd.json +++ /dev/null @@ -1,39 +0,0 @@ -{ - "id": "commit-0087599f9fbd", - "source": "git-log", - "created_at": "2026-01-22T13:07:07Z", - "document": { - "title": "Update STATE.scm with session accomplishments", - "body": "Commit 0087599f by Your Name: Update STATE.scm with session accomplishments", - "fields": { - "type": "commit", - "hash": "0087599f9fbdcd0657fe80dbbb46cd701f3ca8e4", - "author": "Your Name", - "email": "you@example.com", - "commit_type": "other" - } - }, - "graph": { - "relationships": [{"predicate":"modifies","target":"file:.machine_readable/STATE.scm"}] - }, - "vector": { - "embedding": [-0.156250,0.773437,0.031250,-0.945312,-0.164062,0.351562,0.273437,-0.132812,-0.789062,-0.070312,-0.898437,0.593750,0.304687,-0.476562,0.148437,0.898437,0.835937,0.281250,0.687500,0.992187,-0.671875,-0.515625,-0.687500,-0.281250,-0.031250,-0.890625,-0.992187,-1.000000,-0.281250,0.679687,0.500000,0.070312,-0.156250,0.773437,0.031250,-0.945312,-0.164062,0.351562,0.273437,-0.132812,-0.789062,-0.070312,-0.898437,0.593750,0.304687,-0.476562,0.148437,0.898437,0.835937,0.281250,0.687500,0.992187,-0.671875,-0.515625,-0.687500,-0.281250,-0.031250,-0.890625,-0.992187,-1.000000,-0.281250,0.679687,0.500000,0.070312], - "model": "sha256-hash-64d" - }, - "tensor": { - "shape": [1, 3], - "data": [32.0, 1.0, 1.0] - }, - "semantic": { - "types": ["https://verisim.db/self/type/Commit", "https://verisim.db/self/type/other"], - "properties": { - "conventional_commit_type": "other", - "files_changed": "1" - } - }, - "temporal": { - "timestamp": "2026-01-22T13:07:07Z", - "version": 1, - "author": "Your Name" - } -} diff --git a/verisimdb/.verisimdb/octads/commit-030287996c2c.json b/verisimdb/.verisimdb/octads/commit-030287996c2c.json deleted file mode 100644 index f5a02ca6..00000000 --- a/verisimdb/.verisimdb/octads/commit-030287996c2c.json +++ /dev/null @@ -1,39 +0,0 @@ -{ - "id": "commit-030287996c2c", - "source": "git-log", - "created_at": "2026-01-22T09:48:56Z", - "document": { - "title": "feat(security): add ClusterFuzzLite fuzzing for OpenSSF Scorecard", - "body": "Commit 03028799 by Your Name: feat(security): add ClusterFuzzLite fuzzing for OpenSSF Scorecard", - "fields": { - "type": "commit", - "hash": "030287996c2cf730f18795cbe3dfdc4e45052eb7", - "author": "Your Name", - "email": "you@example.com", - "commit_type": "feature" - } - }, - "graph": { - "relationships": [{"predicate":"modifies","target":"file:.clusterfuzzlite/project.yaml"},{"predicate":"modifies","target":"file:.github/workflows/cflite_batch.yml"},{"predicate":"modifies","target":"file:.github/workflows/cflite_pr.yml"},{"predicate":"modifies","target":"file:fuzz/Cargo.toml"},{"predicate":"modifies","target":"file:fuzz/fuzz_targets/fuzz_octad_id.rs"}] - }, - "vector": { - "embedding": [0.468750,0.539062,-0.046875,0.210937,0.890625,-0.085937,-0.585937,0.148437,0.234375,-0.984375,0.218750,0.460937,0.757812,-0.953125,0.007812,0.257812,-0.078125,0.484375,0.468750,0.578125,-0.468750,-0.351562,0.226562,-0.890625,0.164062,0.843750,0.187500,0.125000,-0.429687,0.773437,-0.929687,-0.085937,0.468750,0.539062,-0.046875,0.210937,0.890625,-0.085937,-0.585937,0.148437,0.234375,-0.984375,0.218750,0.460937,0.757812,-0.953125,0.007812,0.257812,-0.078125,0.484375,0.468750,0.578125,-0.468750,-0.351562,0.226562,-0.890625,0.164062,0.843750,0.187500,0.125000,-0.429687,0.773437,-0.929687,-0.085937], - "model": "sha256-hash-64d" - }, - "tensor": { - "shape": [1, 3], - "data": [121.0, 0.0, 5.0] - }, - "semantic": { - "types": ["https://verisim.db/self/type/Commit", "https://verisim.db/self/type/feature"], - "properties": { - "conventional_commit_type": "feature", - "files_changed": "5" - } - }, - "temporal": { - "timestamp": "2026-01-22T09:48:56Z", - "version": 1, - "author": "Your Name" - } -} diff --git a/verisimdb/.verisimdb/octads/commit-031ad82afba6.json b/verisimdb/.verisimdb/octads/commit-031ad82afba6.json deleted file mode 100644 index fe0e3ffb..00000000 --- a/verisimdb/.verisimdb/octads/commit-031ad82afba6.json +++ /dev/null @@ -1,39 +0,0 @@ -{ - "id": "commit-031ad82afba6", - "source": "git-log", - "created_at": "2026-01-22T09:19:58Z", - "document": { - "title": "docs: add PDF version of whitepaper for sharing", - "body": "Commit 031ad82a by Your Name: docs: add PDF version of whitepaper for sharing", - "fields": { - "type": "commit", - "hash": "031ad82afba6203094fe1ab1389e9e150cc17280", - "author": "Your Name", - "email": "you@example.com", - "commit_type": "documentation" - } - }, - "graph": { - "relationships": [{"predicate":"modifies","target":"file:WHITEPAPER.pdf"}] - }, - "vector": { - "embedding": [-0.773437,-0.921875,-0.664062,-0.976562,0.273437,0.843750,0.398437,0.960937,-0.015625,0.382812,-0.414062,0.593750,0.382812,-0.195312,0.640625,-0.976562,0.710937,-0.757812,-0.257812,-0.929687,-0.203125,0.593750,0.773437,-0.757812,-0.640625,0.468750,0.882812,0.312500,0.046875,0.968750,0.429687,0.585937,-0.773437,-0.921875,-0.664062,-0.976562,0.273437,0.843750,0.398437,0.960937,-0.015625,0.382812,-0.414062,0.593750,0.382812,-0.195312,0.640625,-0.976562,0.710937,-0.757812,-0.257812,-0.929687,-0.203125,0.593750,0.773437,-0.757812,-0.640625,0.468750,0.882812,0.312500,0.046875,0.968750,0.429687,0.585937], - "model": "sha256-hash-64d" - }, - "tensor": { - "shape": [1, 3], - "data": [0.0, 0.0, 1.0] - }, - "semantic": { - "types": ["https://verisim.db/self/type/Commit", "https://verisim.db/self/type/documentation"], - "properties": { - "conventional_commit_type": "documentation", - "files_changed": "1" - } - }, - "temporal": { - "timestamp": "2026-01-22T09:19:58Z", - "version": 1, - "author": "Your Name" - } -} diff --git a/verisimdb/.verisimdb/octads/commit-04ce420ca876.json b/verisimdb/.verisimdb/octads/commit-04ce420ca876.json deleted file mode 100644 index 201f2dc1..00000000 --- a/verisimdb/.verisimdb/octads/commit-04ce420ca876.json +++ /dev/null @@ -1,39 +0,0 @@ -{ - "id": "commit-04ce420ca876", - "source": "git-log", - "created_at": "2026-02-13T15:25:03Z", - "document": { - "title": "fix: resolve issue 19 — federation bare query now returns octad data", - "body": "Commit 04ce420c by Jonathan D.A. Jewell: fix: resolve issue 19 — federation bare query now returns octad data", - "fields": { - "type": "commit", - "hash": "04ce420ca876132673cf1edb040cea7f5fb35e17", - "author": "Jonathan D.A. Jewell", - "email": "jonathan.jewell@open.ac.uk", - "commit_type": "bugfix" - } - }, - "graph": { - "relationships": [] - }, - "vector": { - "embedding": [0.328125,-0.851562,-0.515625,0.210937,0.015625,-0.234375,-0.101562,0.890625,-0.367187,-0.929687,-0.687500,-0.968750,0.093750,0.546875,0.976562,0.539062,0.507812,-0.078125,-0.945312,-0.828125,0.070312,0.585937,-0.304687,0.023437,-0.039062,0.812500,0.117187,0.335937,-0.406250,-0.304687,-0.789062,0.085937,0.328125,-0.851562,-0.515625,0.210937,0.015625,-0.234375,-0.101562,0.890625,-0.367187,-0.929687,-0.687500,-0.968750,0.093750,0.546875,0.976562,0.539062,0.507812,-0.078125,-0.945312,-0.828125,0.070312,0.585937,-0.304687,0.023437,-0.039062,0.812500,0.117187,0.335937,-0.406250,-0.304687,-0.789062,0.085937], - "model": "sha256-hash-64d" - }, - "tensor": { - "shape": [1, 3], - "data": [775.0, 151.0, 23.0] - }, - "semantic": { - "types": ["https://verisim.db/self/type/Commit", "https://verisim.db/self/type/bugfix"], - "properties": { - "conventional_commit_type": "bugfix", - "files_changed": "23" - } - }, - "temporal": { - "timestamp": "2026-02-13T15:25:03Z", - "version": 1, - "author": "Jonathan D.A. Jewell" - } -} diff --git a/verisimdb/.verisimdb/octads/commit-0565bb065c0b.json b/verisimdb/.verisimdb/octads/commit-0565bb065c0b.json deleted file mode 100644 index 4fb918ca..00000000 --- a/verisimdb/.verisimdb/octads/commit-0565bb065c0b.json +++ /dev/null @@ -1,39 +0,0 @@ -{ - "id": "commit-0565bb065c0b", - "source": "git-log", - "created_at": "2026-01-22T09:45:32Z", - "document": { - "title": "Update STATE.scm", - "body": "Commit 0565bb06 by Jonathan D.A. Jewell: Update STATE.scm", - "fields": { - "type": "commit", - "hash": "0565bb065c0ba2ec88d7548b1bbc34f62679e1ea", - "author": "Jonathan D.A. Jewell", - "email": "6759885+hyperpolymath@users.noreply.github.com", - "commit_type": "other" - } - }, - "graph": { - "relationships": [{"predicate":"modifies","target":"file:.machine_readable/STATE.scm"}] - }, - "vector": { - "embedding": [-0.304687,-0.210937,-0.226562,0.414062,-0.750000,0.351562,0.015625,-0.664062,0.210937,-0.398437,0.914062,0.093750,0.757812,-0.453125,-0.429687,0.679687,-0.820312,-0.507812,-0.789062,-0.468750,-0.937500,0.562500,-0.554687,-0.882812,0.695312,-0.164062,0.898437,0.695312,0.171875,-0.718750,0.796875,-0.085937,-0.304687,-0.210937,-0.226562,0.414062,-0.750000,0.351562,0.015625,-0.664062,0.210937,-0.398437,0.914062,0.093750,0.757812,-0.453125,-0.429687,0.679687,-0.820312,-0.507812,-0.789062,-0.468750,-0.937500,0.562500,-0.554687,-0.882812,0.695312,-0.164062,0.898437,0.695312,0.171875,-0.718750,0.796875,-0.085937], - "model": "sha256-hash-64d" - }, - "tensor": { - "shape": [1, 3], - "data": [14.0, 102.0, 1.0] - }, - "semantic": { - "types": ["https://verisim.db/self/type/Commit", "https://verisim.db/self/type/other"], - "properties": { - "conventional_commit_type": "other", - "files_changed": "1" - } - }, - "temporal": { - "timestamp": "2026-01-22T09:45:32Z", - "version": 1, - "author": "Jonathan D.A. Jewell" - } -} diff --git a/verisimdb/.verisimdb/octads/commit-05b5a1a2cbaa.json b/verisimdb/.verisimdb/octads/commit-05b5a1a2cbaa.json deleted file mode 100644 index 499ac6bf..00000000 --- a/verisimdb/.verisimdb/octads/commit-05b5a1a2cbaa.json +++ /dev/null @@ -1,39 +0,0 @@ -{ - "id": "commit-05b5a1a2cbaa", - "source": "git-log", - "created_at": "2026-01-17T02:19:22Z", - "document": { - "title": "chore: sync template files and configuration", - "body": "Commit 05b5a1a2 by Jonathan D.A. Jewell: chore: sync template files and configuration", - "fields": { - "type": "commit", - "hash": "05b5a1a2cbaa85ab30d499ebf77d0889922039fa", - "author": "Jonathan D.A. Jewell", - "email": "jonathan.jewell@open.ac.uk", - "commit_type": "chore" - } - }, - "graph": { - "relationships": [{"predicate":"modifies","target":"file:rust-core/verisim-octad/tests/integration_tests.rs"},{"predicate":"modifies","target":"file:rust-core/verisim-vector/src/hnsw.rs"}] - }, - "vector": { - "embedding": [0.101562,0.898437,-0.367187,-0.203125,0.718750,0.656250,0.703125,-0.921875,0.140625,-0.695312,-0.234375,-0.734375,0.640625,0.546875,-0.640625,-0.460937,0.210937,0.781250,0.648437,0.890625,0.820312,-0.648437,-0.742187,0.007812,0.625000,-0.015625,0.687500,-0.445312,0.187500,0.859375,-0.367187,0.101562,0.101562,0.898437,-0.367187,-0.203125,0.718750,0.656250,0.703125,-0.921875,0.140625,-0.695312,-0.234375,-0.734375,0.640625,0.546875,-0.640625,-0.460937,0.210937,0.781250,0.648437,0.890625,0.820312,-0.648437,-0.742187,0.007812,0.625000,-0.015625,0.687500,-0.445312,0.187500,0.859375,-0.367187,0.101562], - "model": "sha256-hash-64d" - }, - "tensor": { - "shape": [1, 3], - "data": [920.0, 0.0, 2.0] - }, - "semantic": { - "types": ["https://verisim.db/self/type/Commit", "https://verisim.db/self/type/chore"], - "properties": { - "conventional_commit_type": "chore", - "files_changed": "2" - } - }, - "temporal": { - "timestamp": "2026-01-17T02:19:22Z", - "version": 1, - "author": "Jonathan D.A. Jewell" - } -} diff --git a/verisimdb/.verisimdb/octads/commit-0955ca7ada21.json b/verisimdb/.verisimdb/octads/commit-0955ca7ada21.json deleted file mode 100644 index b28adf98..00000000 --- a/verisimdb/.verisimdb/octads/commit-0955ca7ada21.json +++ /dev/null @@ -1,39 +0,0 @@ -{ - "id": "commit-0955ca7ada21", - "source": "git-log", - "created_at": "2026-01-16T19:58:41Z", - "document": { - "title": "fix: test failures and Axum 0.8 route syntax", - "body": "Commit 0955ca7a by Jonathan D.A. Jewell: fix: test failures and Axum 0.8 route syntax", - "fields": { - "type": "commit", - "hash": "0955ca7ada210394f7de09376160e25087bdcf55", - "author": "Jonathan D.A. Jewell", - "email": "jonathan.jewell@open.ac.uk", - "commit_type": "bugfix" - } - }, - "graph": { - "relationships": [{"predicate":"modifies","target":"file:rust-core/verisim-api/src/lib.rs"},{"predicate":"modifies","target":"file:rust-core/verisim-drift/src/lib.rs"}] - }, - "vector": { - "embedding": [-0.984375,0.070312,0.054687,-0.835937,-0.851562,-0.421875,0.437500,0.664062,-0.203125,0.023437,-0.023437,-0.539062,-0.304687,0.570312,-0.304687,0.203125,0.687500,0.835937,0.351562,-0.421875,-0.609375,0.976562,0.359375,-0.312500,0.140625,0.382812,-0.992187,-0.929687,0.765625,0.914062,-0.195312,0.328125,-0.984375,0.070312,0.054687,-0.835937,-0.851562,-0.421875,0.437500,0.664062,-0.203125,0.023437,-0.023437,-0.539062,-0.304687,0.570312,-0.304687,0.203125,0.687500,0.835937,0.351562,-0.421875,-0.609375,0.976562,0.359375,-0.312500,0.140625,0.382812,-0.992187,-0.929687,0.765625,0.914062,-0.195312,0.328125], - "model": "sha256-hash-64d" - }, - "tensor": { - "shape": [1, 3], - "data": [8.0, 7.0, 2.0] - }, - "semantic": { - "types": ["https://verisim.db/self/type/Commit", "https://verisim.db/self/type/bugfix"], - "properties": { - "conventional_commit_type": "bugfix", - "files_changed": "2" - } - }, - "temporal": { - "timestamp": "2026-01-16T19:58:41Z", - "version": 1, - "author": "Jonathan D.A. Jewell" - } -} diff --git a/verisimdb/.verisimdb/octads/commit-0ba909ba3a0b.json b/verisimdb/.verisimdb/octads/commit-0ba909ba3a0b.json deleted file mode 100644 index e98edd25..00000000 --- a/verisimdb/.verisimdb/octads/commit-0ba909ba3a0b.json +++ /dev/null @@ -1,39 +0,0 @@ -{ - "id": "commit-0ba909ba3a0b", - "source": "git-log", - "created_at": "2026-01-16T19:10:48Z", - "document": { - "title": "fix: resolve compilation errors in vector, document, and drift crates", - "body": "Commit 0ba909ba by Jonathan D.A. Jewell: fix: resolve compilation errors in vector, document, and drift crates", - "fields": { - "type": "commit", - "hash": "0ba909ba3a0b32ec996b8ae3fcc53dd27140fce4", - "author": "Jonathan D.A. Jewell", - "email": "jonathan.jewell@open.ac.uk", - "commit_type": "bugfix" - } - }, - "graph": { - "relationships": [{"predicate":"modifies","target":"file:rust-core/verisim-document/src/lib.rs"},{"predicate":"modifies","target":"file:rust-core/verisim-drift/src/lib.rs"},{"predicate":"modifies","target":"file:rust-core/verisim-vector/Cargo.toml"},{"predicate":"modifies","target":"file:rust-core/verisim-vector/src/lib.rs"}] - }, - "vector": { - "embedding": [-0.242187,0.257812,0.015625,0.414062,-0.648437,-0.945312,-0.187500,-1.000000,0.250000,-0.929687,-0.757812,-0.296875,-0.273437,0.843750,0.968750,0.171875,-0.679687,-0.148437,-0.515625,-0.796875,0.898437,0.101562,0.968750,0.156250,0.507812,-0.093750,-0.718750,0.781250,0.890625,-0.875000,-0.507812,0.914062,-0.242187,0.257812,0.015625,0.414062,-0.648437,-0.945312,-0.187500,-1.000000,0.250000,-0.929687,-0.757812,-0.296875,-0.273437,0.843750,0.968750,0.171875,-0.679687,-0.148437,-0.515625,-0.796875,0.898437,0.101562,0.968750,0.156250,0.507812,-0.093750,-0.718750,0.781250,0.890625,-0.875000,-0.507812,0.914062], - "model": "sha256-hash-64d" - }, - "tensor": { - "shape": [1, 3], - "data": [52.0, 63.0, 4.0] - }, - "semantic": { - "types": ["https://verisim.db/self/type/Commit", "https://verisim.db/self/type/bugfix"], - "properties": { - "conventional_commit_type": "bugfix", - "files_changed": "4" - } - }, - "temporal": { - "timestamp": "2026-01-16T19:10:48Z", - "version": 1, - "author": "Jonathan D.A. Jewell" - } -} diff --git a/verisimdb/.verisimdb/octads/commit-0ca572654610.json b/verisimdb/.verisimdb/octads/commit-0ca572654610.json deleted file mode 100644 index 89b65512..00000000 --- a/verisimdb/.verisimdb/octads/commit-0ca572654610.json +++ /dev/null @@ -1,39 +0,0 @@ -{ - "id": "commit-0ca572654610", - "source": "git-log", - "created_at": "2026-01-22T11:10:09Z", - "document": { - "title": "Fix VCL grammar: use single quotes per ISO EBNF standard", - "body": "Commit 0ca57265 by Your Name: Fix VCL grammar: use single quotes per ISO EBNF standard", - "fields": { - "type": "commit", - "hash": "0ca572654610b0420c50f4963701c3dca79d7e28", - "author": "Your Name", - "email": "you@example.com", - "commit_type": "other" - } - }, - "graph": { - "relationships": [{"predicate":"modifies","target":"file:docs/vcl-grammar.ebnf"}] - }, - "vector": { - "embedding": [-1.000000,0.898437,0.078125,-0.656250,0.742187,-0.562500,0.210937,0.320312,-0.031250,0.429687,-0.757812,0.234375,-0.375000,-0.859375,-0.757812,-0.281250,-0.421875,-0.031250,0.179687,-0.640625,-0.101562,-0.054687,0.289062,-0.992187,-0.031250,0.125000,-0.742187,0.328125,-0.093750,0.539062,0.375000,0.187500,-1.000000,0.898437,0.078125,-0.656250,0.742187,-0.562500,0.210937,0.320312,-0.031250,0.429687,-0.757812,0.234375,-0.375000,-0.859375,-0.757812,-0.281250,-0.421875,-0.031250,0.179687,-0.640625,-0.101562,-0.054687,0.289062,-0.992187,-0.031250,0.125000,-0.742187,0.328125,-0.093750,0.539062,0.375000,0.187500], - "model": "sha256-hash-64d" - }, - "tensor": { - "shape": [1, 3], - "data": [90.0, 90.0, 1.0] - }, - "semantic": { - "types": ["https://verisim.db/self/type/Commit", "https://verisim.db/self/type/other"], - "properties": { - "conventional_commit_type": "other", - "files_changed": "1" - } - }, - "temporal": { - "timestamp": "2026-01-22T11:10:09Z", - "version": 1, - "author": "Your Name" - } -} diff --git a/verisimdb/.verisimdb/octads/commit-17f903f9fa44.json b/verisimdb/.verisimdb/octads/commit-17f903f9fa44.json deleted file mode 100644 index 6903c626..00000000 --- a/verisimdb/.verisimdb/octads/commit-17f903f9fa44.json +++ /dev/null @@ -1,39 +0,0 @@ -{ - "id": "commit-17f903f9fa44", - "source": "git-log", - "created_at": "2026-01-22T10:41:41Z", - "document": { - "title": "Add comprehensive ASCII diagrams to cache sharing strategy", - "body": "Commit 17f903f9 by Your Name: Add comprehensive ASCII diagrams to cache sharing strategy", - "fields": { - "type": "commit", - "hash": "17f903f9fa445aaa9e60b7509d564810b198224e", - "author": "Your Name", - "email": "you@example.com", - "commit_type": "other" - } - }, - "graph": { - "relationships": [{"predicate":"modifies","target":"file:docs/cache-sharing-strategy.adoc"}] - }, - "vector": { - "embedding": [0.226562,0.906250,-0.070312,-0.617187,-0.828125,-0.007812,0.320312,-0.007812,0.921875,0.132812,0.304687,-0.257812,-0.210937,-0.546875,-0.500000,-0.632812,-0.390625,0.429687,-0.539062,0.507812,0.726562,0.742187,0.335937,-0.437500,0.500000,-0.226562,0.687500,0.882812,-0.132812,0.085937,0.273437,0.359375,0.226562,0.906250,-0.070312,-0.617187,-0.828125,-0.007812,0.320312,-0.007812,0.921875,0.132812,0.304687,-0.257812,-0.210937,-0.546875,-0.500000,-0.632812,-0.390625,0.429687,-0.539062,0.507812,0.726562,0.742187,0.335937,-0.437500,0.500000,-0.226562,0.687500,0.882812,-0.132812,0.085937,0.273437,0.359375], - "model": "sha256-hash-64d" - }, - "tensor": { - "shape": [1, 3], - "data": [745.0, 0.0, 1.0] - }, - "semantic": { - "types": ["https://verisim.db/self/type/Commit", "https://verisim.db/self/type/other"], - "properties": { - "conventional_commit_type": "other", - "files_changed": "1" - } - }, - "temporal": { - "timestamp": "2026-01-22T10:41:41Z", - "version": 1, - "author": "Your Name" - } -} diff --git a/verisimdb/.verisimdb/octads/commit-1c5b28a1141e.json b/verisimdb/.verisimdb/octads/commit-1c5b28a1141e.json deleted file mode 100644 index a7115db5..00000000 --- a/verisimdb/.verisimdb/octads/commit-1c5b28a1141e.json +++ /dev/null @@ -1,39 +0,0 @@ -{ - "id": "commit-1c5b28a1141e", - "source": "git-log", - "created_at": "2026-01-26T02:54:37Z", - "document": { - "title": "Add OPSM integration links", - "body": "Commit 1c5b28a1 by Your Name: Add OPSM integration links", - "fields": { - "type": "commit", - "hash": "1c5b28a1141e34b8f3999d19280f492d215c182d", - "author": "Your Name", - "email": "you@example.com", - "commit_type": "other" - } - }, - "graph": { - "relationships": [{"predicate":"modifies","target":"file:.machine_readable/6scm/ECOSYSTEM.scm"},{"predicate":"modifies","target":"file:.machine_readable/6scm/META.scm"},{"predicate":"modifies","target":"file:README.adoc"},{"predicate":"modifies","target":"file:ROADMAP.adoc"}] - }, - "vector": { - "embedding": [-0.460937,-0.390625,0.960937,0.171875,-0.078125,-0.492187,-0.046875,-0.437500,0.046875,0.164062,-0.921875,0.382812,-0.945312,0.750000,0.039062,-0.921875,0.414062,0.031250,0.875000,-0.468750,0.609375,-0.367187,0.148437,-0.460937,-0.703125,0.796875,0.898437,0.914062,-0.015625,0.757812,-0.382812,0.187500,-0.460937,-0.390625,0.960937,0.171875,-0.078125,-0.492187,-0.046875,-0.437500,0.046875,0.164062,-0.921875,0.382812,-0.945312,0.750000,0.039062,-0.921875,0.414062,0.031250,0.875000,-0.468750,0.609375,-0.367187,0.148437,-0.460937,-0.703125,0.796875,0.898437,0.914062,-0.015625,0.757812,-0.382812,0.187500], - "model": "sha256-hash-64d" - }, - "tensor": { - "shape": [1, 3], - "data": [25.0, 3.0, 4.0] - }, - "semantic": { - "types": ["https://verisim.db/self/type/Commit", "https://verisim.db/self/type/other"], - "properties": { - "conventional_commit_type": "other", - "files_changed": "4" - } - }, - "temporal": { - "timestamp": "2026-01-26T02:54:37Z", - "version": 1, - "author": "Your Name" - } -} diff --git a/verisimdb/.verisimdb/octads/commit-208a70f82714.json b/verisimdb/.verisimdb/octads/commit-208a70f82714.json deleted file mode 100644 index 6d64debd..00000000 --- a/verisimdb/.verisimdb/octads/commit-208a70f82714.json +++ /dev/null @@ -1,39 +0,0 @@ -{ - "id": "commit-208a70f82714", - "source": "git-log", - "created_at": "2026-01-22T09:32:56Z", - "document": { - "title": "chore(security): fix Dependabot and OpenSSF Scorecard issues", - "body": "Commit 208a70f8 by Your Name: chore(security): fix Dependabot and OpenSSF Scorecard issues", - "fields": { - "type": "commit", - "hash": "208a70f82714940aa396d3feb9b3fb7317e5133f", - "author": "Your Name", - "email": "you@example.com", - "commit_type": "chore" - } - }, - "graph": { - "relationships": [{"predicate":"modifies","target":"file:.github/workflows/jekyll-gh-pages.yml"},{"predicate":"modifies","target":"file:.github/workflows/jekyll.yml"},{"predicate":"modifies","target":"file:Cargo.lock"},{"predicate":"modifies","target":"file:Cargo.toml"},{"predicate":"modifies","target":"file:docs/challenges-hybrid.adoc"},{"predicate":"modifies","target":"file:docs/challenges-standalone.adoc"}] - }, - "vector": { - "embedding": [0.335937,-0.507812,0.750000,-0.429687,-0.242187,0.835937,-0.093750,0.734375,-0.804687,-0.664062,-0.578125,-0.703125,-0.781250,0.601562,0.875000,0.539062,0.476562,-0.164062,0.171875,-0.023437,0.648437,0.062500,-0.945312,0.757812,-0.562500,0.250000,0.476562,-0.804687,0.812500,0.117187,-0.859375,0.648437,0.335937,-0.507812,0.750000,-0.429687,-0.242187,0.835937,-0.093750,0.734375,-0.804687,-0.664062,-0.578125,-0.703125,-0.781250,0.601562,0.875000,0.539062,0.476562,-0.164062,0.171875,-0.023437,0.648437,0.062500,-0.945312,0.757812,-0.562500,0.250000,0.476562,-0.804687,0.812500,0.117187,-0.859375,0.648437], - "model": "sha256-hash-64d" - }, - "tensor": { - "shape": [1, 3], - "data": [1234.0, 11.0, 6.0] - }, - "semantic": { - "types": ["https://verisim.db/self/type/Commit", "https://verisim.db/self/type/chore"], - "properties": { - "conventional_commit_type": "chore", - "files_changed": "6" - } - }, - "temporal": { - "timestamp": "2026-01-22T09:32:56Z", - "version": 1, - "author": "Your Name" - } -} diff --git a/verisimdb/.verisimdb/octads/commit-20ea93a1be96.json b/verisimdb/.verisimdb/octads/commit-20ea93a1be96.json deleted file mode 100644 index dabd8191..00000000 --- a/verisimdb/.verisimdb/octads/commit-20ea93a1be96.json +++ /dev/null @@ -1,39 +0,0 @@ -{ - "id": "commit-20ea93a1be96", - "source": "git-log", - "created_at": "2026-02-12T16:34:12Z", - "document": { - "title": "fix: repair broken integration tests and write Round 2 Sonnet tasks", - "body": "Commit 20ea93a1 by Jonathan D.A. Jewell: fix: repair broken integration tests and write Round 2 Sonnet tasks", - "fields": { - "type": "commit", - "hash": "20ea93a1be96ca9c8ba0e04b5f408b1565fc1d99", - "author": "Jonathan D.A. Jewell", - "email": "jonathan.jewell@open.ac.uk", - "commit_type": "bugfix" - } - }, - "graph": { - "relationships": [{"predicate":"modifies","target":"file:SONNET-TASKS.md"},{"predicate":"modifies","target":"file:rust-core/verisim-octad/tests/integration_tests.rs"},{"predicate":"modifies","target":"file:rust-core/verisim-temporal/tests/property_tests.rs"}] - }, - "vector": { - "embedding": [-0.031250,0.195312,0.539062,0.265625,0.617187,0.796875,-0.429687,-0.945312,-0.523437,0.335937,0.335937,-0.523437,-0.171875,-0.703125,-0.460937,0.062500,-0.828125,0.695312,-0.093750,-0.882812,-0.968750,0.039062,-0.148437,-0.718750,-0.140625,0.859375,0.429687,-0.765625,0.359375,0.046875,0.250000,-0.078125,-0.031250,0.195312,0.539062,0.265625,0.617187,0.796875,-0.429687,-0.945312,-0.523437,0.335937,0.335937,-0.523437,-0.171875,-0.703125,-0.460937,0.062500,-0.828125,0.695312,-0.093750,-0.882812,-0.968750,0.039062,-0.148437,-0.718750,-0.140625,0.859375,0.429687,-0.765625,0.359375,0.046875,0.250000,-0.078125], - "model": "sha256-hash-64d" - }, - "tensor": { - "shape": [1, 3], - "data": [731.0, 1097.0, 3.0] - }, - "semantic": { - "types": ["https://verisim.db/self/type/Commit", "https://verisim.db/self/type/bugfix"], - "properties": { - "conventional_commit_type": "bugfix", - "files_changed": "3" - } - }, - "temporal": { - "timestamp": "2026-02-12T16:34:12Z", - "version": 1, - "author": "Jonathan D.A. Jewell" - } -} diff --git a/verisimdb/.verisimdb/octads/commit-228ae4bba134.json b/verisimdb/.verisimdb/octads/commit-228ae4bba134.json deleted file mode 100644 index 0968b72c..00000000 --- a/verisimdb/.verisimdb/octads/commit-228ae4bba134.json +++ /dev/null @@ -1,39 +0,0 @@ -{ - "id": "commit-228ae4bba134", - "source": "git-log", - "created_at": "2026-01-22T19:20:56Z", - "document": { - "title": "chore: commit local changes for sync", - "body": "Commit 228ae4bb by Your Name: chore: commit local changes for sync", - "fields": { - "type": "commit", - "hash": "228ae4bba134c8fc7befb7c5985f25a25f52d46c", - "author": "Your Name", - "email": "you@example.com", - "commit_type": "chore" - } - }, - "graph": { - "relationships": [{"predicate":"modifies","target":"file:ECOSYSTEM.scm"},{"predicate":"modifies","target":"file:META.scm"},{"predicate":"modifies","target":"file:STATE.scm"}] - }, - "vector": { - "embedding": [0.218750,0.921875,-0.023437,-0.476562,0.031250,-0.195312,-0.664062,0.515625,-0.367187,0.843750,-0.500000,0.046875,0.242187,0.976562,-0.476562,-0.906250,-0.007812,0.375000,0.179687,0.398437,-0.281250,0.640625,0.890625,0.726562,-0.320312,-0.156250,-0.226562,0.046875,-0.093750,-0.125000,0.695312,0.140625,0.218750,0.921875,-0.023437,-0.476562,0.031250,-0.195312,-0.664062,0.515625,-0.367187,0.843750,-0.500000,0.046875,0.242187,0.976562,-0.476562,-0.906250,-0.007812,0.375000,0.179687,0.398437,-0.281250,0.640625,0.890625,0.726562,-0.320312,-0.156250,-0.226562,0.046875,-0.093750,-0.125000,0.695312,0.140625], - "model": "sha256-hash-64d" - }, - "tensor": { - "shape": [1, 3], - "data": [0.0, 721.0, 3.0] - }, - "semantic": { - "types": ["https://verisim.db/self/type/Commit", "https://verisim.db/self/type/chore"], - "properties": { - "conventional_commit_type": "chore", - "files_changed": "3" - } - }, - "temporal": { - "timestamp": "2026-01-22T19:20:56Z", - "version": 1, - "author": "Your Name" - } -} diff --git a/verisimdb/.verisimdb/octads/commit-23815332901c.json b/verisimdb/.verisimdb/octads/commit-23815332901c.json deleted file mode 100644 index 63b97552..00000000 --- a/verisimdb/.verisimdb/octads/commit-23815332901c.json +++ /dev/null @@ -1,39 +0,0 @@ -{ - "id": "commit-23815332901c", - "source": "git-log", - "created_at": "2026-01-17T00:23:10Z", - "document": { - "title": "Add STATE.scm and update to PMPL-1.0", - "body": "Commit 23815332 by Jonathan D.A. Jewell: Add STATE.scm and update to PMPL-1.0", - "fields": { - "type": "commit", - "hash": "23815332901cb179418cf749eb477a6e30f9708a", - "author": "Jonathan D.A. Jewell", - "email": "jonathan.jewell@open.ac.uk", - "commit_type": "other" - } - }, - "graph": { - "relationships": [{"predicate":"modifies","target":"file:LICENSE"},{"predicate":"modifies","target":"file:STATE.scm"}] - }, - "vector": { - "embedding": [-0.726562,-0.375000,0.492187,-0.609375,0.281250,-0.492187,0.679687,-0.320312,0.390625,-0.164062,-0.515625,0.789062,-0.546875,-0.210937,-0.093750,0.460937,-0.601562,0.523437,0.570312,0.187500,0.648437,-0.320312,0.710937,-0.781250,0.046875,0.820312,-0.539062,-0.851562,-0.875000,0.789062,-0.992187,-0.578125,-0.726562,-0.375000,0.492187,-0.609375,0.281250,-0.492187,0.679687,-0.320312,0.390625,-0.164062,-0.515625,0.789062,-0.546875,-0.210937,-0.093750,0.460937,-0.601562,0.523437,0.570312,0.187500,0.648437,-0.320312,0.710937,-0.781250,0.046875,0.820312,-0.539062,-0.851562,-0.875000,0.789062,-0.992187,-0.578125], - "model": "sha256-hash-64d" - }, - "tensor": { - "shape": [1, 3], - "data": [178.0, 0.0, 2.0] - }, - "semantic": { - "types": ["https://verisim.db/self/type/Commit", "https://verisim.db/self/type/other"], - "properties": { - "conventional_commit_type": "other", - "files_changed": "2" - } - }, - "temporal": { - "timestamp": "2026-01-17T00:23:10Z", - "version": 1, - "author": "Jonathan D.A. Jewell" - } -} diff --git a/verisimdb/.verisimdb/octads/commit-242713b9c58c.json b/verisimdb/.verisimdb/octads/commit-242713b9c58c.json deleted file mode 100644 index f77cdc31..00000000 --- a/verisimdb/.verisimdb/octads/commit-242713b9c58c.json +++ /dev/null @@ -1,39 +0,0 @@ -{ - "id": "commit-242713b9c58c", - "source": "git-log", - "created_at": "2026-02-08T07:52:40Z", - "document": { - "title": "feat: absorb debugger and practice-mirror into verisimdb monorepo", - "body": "Commit 242713b9 by Jonathan D.A. Jewell: feat: absorb debugger and practice-mirror into verisimdb monorepo", - "fields": { - "type": "commit", - "hash": "242713b9c58cc980aa53f003b66a5eb9035dccdc", - "author": "Jonathan D.A. Jewell", - "email": "jonathan.jewell@open.ac.uk", - "commit_type": "feature" - } - }, - "graph": { - "relationships": [] - }, - "vector": { - "embedding": [-1.000000,0.718750,-0.929687,-0.554687,-0.414062,-0.140625,0.929687,-0.476562,-0.156250,-0.398437,-0.617187,-0.796875,-0.429687,0.023437,0.570312,-0.492187,0.273437,0.601562,0.281250,0.898437,0.367187,-0.343750,0.828125,0.921875,-0.367187,0.085937,-0.281250,-0.843750,0.226562,0.656250,-0.125000,0.578125,-1.000000,0.718750,-0.929687,-0.554687,-0.414062,-0.140625,0.929687,-0.476562,-0.156250,-0.398437,-0.617187,-0.796875,-0.429687,0.023437,0.570312,-0.492187,0.273437,0.601562,0.281250,0.898437,0.367187,-0.343750,0.828125,0.921875,-0.367187,0.085937,-0.281250,-0.843750,0.226562,0.656250,-0.125000,0.578125], - "model": "sha256-hash-64d" - }, - "tensor": { - "shape": [1, 3], - "data": [3777.0, 0.0, 50.0] - }, - "semantic": { - "types": ["https://verisim.db/self/type/Commit", "https://verisim.db/self/type/feature"], - "properties": { - "conventional_commit_type": "feature", - "files_changed": "50" - } - }, - "temporal": { - "timestamp": "2026-02-08T07:52:40Z", - "version": 1, - "author": "Jonathan D.A. Jewell" - } -} diff --git a/verisimdb/.verisimdb/octads/commit-2723614d8844.json b/verisimdb/.verisimdb/octads/commit-2723614d8844.json deleted file mode 100644 index d258ced7..00000000 --- a/verisimdb/.verisimdb/octads/commit-2723614d8844.json +++ /dev/null @@ -1,39 +0,0 @@ -{ - "id": "commit-2723614d8844", - "source": "git-log", - "created_at": "2026-01-22T20:09:49Z", - "document": { - "title": "chore: fix license headers to PMPL-1.0-or-later", - "body": "Commit 2723614d by Your Name: chore: fix license headers to PMPL-1.0-or-later", - "fields": { - "type": "commit", - "hash": "2723614d88447263678dcaddae24b76e917f2212", - "author": "Your Name", - "email": "you@example.com", - "commit_type": "chore" - } - }, - "graph": { - "relationships": [{"predicate":"modifies","target":"file:.github/workflows/cflite_batch.yml"},{"predicate":"modifies","target":"file:.github/workflows/cflite_pr.yml"},{"predicate":"modifies","target":"file:.github/workflows/jekyll-gh-pages.yml"},{"predicate":"modifies","target":"file:.github/workflows/jekyll.yml"}] - }, - "vector": { - "embedding": [-0.820312,0.804687,0.960937,-0.687500,-0.671875,0.460937,0.937500,-0.289062,-0.414062,-0.828125,-0.187500,-0.195312,0.914062,0.484375,0.648437,0.226562,0.937500,0.742187,0.578125,0.539062,-0.164062,-0.546875,0.078125,-0.335937,0.140625,-0.015625,-0.968750,0.085937,0.250000,-0.445312,0.328125,-0.718750,-0.820312,0.804687,0.960937,-0.687500,-0.671875,0.460937,0.937500,-0.289062,-0.414062,-0.828125,-0.187500,-0.195312,0.914062,0.484375,0.648437,0.226562,0.937500,0.742187,0.578125,0.539062,-0.164062,-0.546875,0.078125,-0.335937,0.140625,-0.015625,-0.968750,0.085937,0.250000,-0.445312,0.328125,-0.718750], - "model": "sha256-hash-64d" - }, - "tensor": { - "shape": [1, 3], - "data": [4.0, 4.0, 4.0] - }, - "semantic": { - "types": ["https://verisim.db/self/type/Commit", "https://verisim.db/self/type/chore"], - "properties": { - "conventional_commit_type": "chore", - "files_changed": "4" - } - }, - "temporal": { - "timestamp": "2026-01-22T20:09:49Z", - "version": 1, - "author": "Your Name" - } -} diff --git a/verisimdb/.verisimdb/octads/commit-2eb26c6b7499.json b/verisimdb/.verisimdb/octads/commit-2eb26c6b7499.json deleted file mode 100644 index 70ee3661..00000000 --- a/verisimdb/.verisimdb/octads/commit-2eb26c6b7499.json +++ /dev/null @@ -1,39 +0,0 @@ -{ - "id": "commit-2eb26c6b7499", - "source": "git-log", - "created_at": "2026-01-22T09:20:14Z", - "document": { - "title": "docs: add PDF link to whitepaper in README", - "body": "Commit 2eb26c6b by Your Name: docs: add PDF link to whitepaper in README", - "fields": { - "type": "commit", - "hash": "2eb26c6b7499b6dc4b09b8b9bd734a70bdb5cc37", - "author": "Your Name", - "email": "you@example.com", - "commit_type": "documentation" - } - }, - "graph": { - "relationships": [{"predicate":"modifies","target":"file:README.adoc"}] - }, - "vector": { - "embedding": [0.468750,0.601562,0,-0.648437,0.351562,-0.132812,-0.187500,0.867187,-0.992187,-0.375000,-0.125000,0.820312,0.742187,0.601562,0.875000,-0.101562,-0.468750,0.117187,0.398437,-0.367187,0.757812,-0.601562,-0.273437,0.882812,-0.171875,-0.601562,0.843750,-1.000000,-0.671875,-0.742187,-1.000000,-0.039062,0.468750,0.601562,0,-0.648437,0.351562,-0.132812,-0.187500,0.867187,-0.992187,-0.375000,-0.125000,0.820312,0.742187,0.601562,0.875000,-0.101562,-0.468750,0.117187,0.398437,-0.367187,0.757812,-0.601562,-0.273437,0.882812,-0.171875,-0.601562,0.843750,-1.000000,-0.671875,-0.742187,-1.000000,-0.039062], - "model": "sha256-hash-64d" - }, - "tensor": { - "shape": [1, 3], - "data": [1.0, 1.0, 1.0] - }, - "semantic": { - "types": ["https://verisim.db/self/type/Commit", "https://verisim.db/self/type/documentation"], - "properties": { - "conventional_commit_type": "documentation", - "files_changed": "1" - } - }, - "temporal": { - "timestamp": "2026-01-22T09:20:14Z", - "version": 1, - "author": "Your Name" - } -} diff --git a/verisimdb/.verisimdb/octads/commit-2ee7e397d41c.json b/verisimdb/.verisimdb/octads/commit-2ee7e397d41c.json deleted file mode 100644 index 5dffe987..00000000 --- a/verisimdb/.verisimdb/octads/commit-2ee7e397d41c.json +++ /dev/null @@ -1,39 +0,0 @@ -{ - "id": "commit-2ee7e397d41c", - "source": "git-log", - "created_at": "2026-01-22T01:34:46Z", - "document": { - "title": "Create Technical Specification - KRaft Metadata Log.md", - "body": "Commit 2ee7e397 by Jonathan D.A. Jewell: Create Technical Specification - KRaft Metadata Log.md", - "fields": { - "type": "commit", - "hash": "2ee7e397d41cc0d971e90e29ccdea61b5e3578db", - "author": "Jonathan D.A. Jewell", - "email": "6759885+hyperpolymath@users.noreply.github.com", - "commit_type": "other" - } - }, - "graph": { - "relationships": [{"predicate":"modifies","target":"file:Technical Specification - KRaft Metadata Log.md"}] - }, - "vector": { - "embedding": [-0.343750,-0.640625,-0.992187,0.328125,-0.265625,0.070312,0.078125,-0.031250,-0.906250,0.484375,-0.414062,-0.218750,-0.476562,-0.281250,-0.367187,0.273437,0.562500,0.265625,0.953125,-0.156250,0.867187,0.929687,0.976562,-0.835937,0.851562,-0.398437,-0.726562,-0.234375,0.757812,-0.593750,-0.078125,0.414062,-0.343750,-0.640625,-0.992187,0.328125,-0.265625,0.070312,0.078125,-0.031250,-0.906250,0.484375,-0.414062,-0.218750,-0.476562,-0.281250,-0.367187,0.273437,0.562500,0.265625,0.953125,-0.156250,0.867187,0.929687,0.976562,-0.835937,0.851562,-0.398437,-0.726562,-0.234375,0.757812,-0.593750,-0.078125,0.414062], - "model": "sha256-hash-64d" - }, - "tensor": { - "shape": [1, 3], - "data": [33.0, 0.0, 1.0] - }, - "semantic": { - "types": ["https://verisim.db/self/type/Commit", "https://verisim.db/self/type/other"], - "properties": { - "conventional_commit_type": "other", - "files_changed": "1" - } - }, - "temporal": { - "timestamp": "2026-01-22T01:34:46Z", - "version": 1, - "author": "Jonathan D.A. Jewell" - } -} diff --git a/verisimdb/.verisimdb/octads/commit-30228e2f8185.json b/verisimdb/.verisimdb/octads/commit-30228e2f8185.json deleted file mode 100644 index 3cbf01a8..00000000 --- a/verisimdb/.verisimdb/octads/commit-30228e2f8185.json +++ /dev/null @@ -1,39 +0,0 @@ -{ - "id": "commit-30228e2f8185", - "source": "git-log", - "created_at": "2026-02-13T09:09:37Z", - "document": { - "title": "chore: remove crash capture artifacts, update .gitignore", - "body": "Commit 30228e2f by Jonathan D.A. Jewell: chore: remove crash capture artifacts, update .gitignore", - "fields": { - "type": "commit", - "hash": "30228e2f81850a8af085501eba59c10f1632bcfd", - "author": "Jonathan D.A. Jewell", - "email": "jonathan.jewell@open.ac.uk", - "commit_type": "chore" - } - }, - "graph": { - "relationships": [{"predicate":"modifies","target":"file:.gitignore"},{"predicate":"modifies","target":"file:ai-cli-crash-capture/README.adoc"},{"predicate":"modifies","target":"file:ai-cli-crash-capture/bash/recoverer.sh"},{"predicate":"modifies","target":"file:ai-cli-crash-capture/recoverer-setup.sh"},{"predicate":"modifies","target":"file:ai-cli-crash-capture/systemd/terminal-recoverer-coredump.service"},{"predicate":"modifies","target":"file:ai-cli-crash-capture/systemd/terminal-recoverer.service"}] - }, - "vector": { - "embedding": [0.851562,-0.468750,0.593750,-0.898437,-0.679687,0.265625,-0.554687,-0.585937,-0.992187,-0.765625,0.390625,0.289062,-0.359375,0.789062,0.867187,0.085937,0.593750,-0.671875,0.578125,-0.601562,0.109375,0.257812,-0.726562,0.890625,-0.039062,-0.390625,0.296875,0.265625,-0.125000,0.359375,0.054687,-0.546875,0.851562,-0.468750,0.593750,-0.898437,-0.679687,0.265625,-0.554687,-0.585937,-0.992187,-0.765625,0.390625,0.289062,-0.359375,0.789062,0.867187,0.085937,0.593750,-0.671875,0.578125,-0.601562,0.109375,0.257812,-0.726562,0.890625,-0.039062,-0.390625,0.296875,0.265625,-0.125000,0.359375,0.054687,-0.546875], - "model": "sha256-hash-64d" - }, - "tensor": { - "shape": [1, 3], - "data": [3.0, 222.0, 6.0] - }, - "semantic": { - "types": ["https://verisim.db/self/type/Commit", "https://verisim.db/self/type/chore"], - "properties": { - "conventional_commit_type": "chore", - "files_changed": "6" - } - }, - "temporal": { - "timestamp": "2026-02-13T09:09:37Z", - "version": 1, - "author": "Jonathan D.A. Jewell" - } -} diff --git a/verisimdb/.verisimdb/octads/commit-3327beaaa2b3.json b/verisimdb/.verisimdb/octads/commit-3327beaaa2b3.json deleted file mode 100644 index 20a7a332..00000000 --- a/verisimdb/.verisimdb/octads/commit-3327beaaa2b3.json +++ /dev/null @@ -1,39 +0,0 @@ -{ - "id": "commit-3327beaaa2b3", - "source": "git-log", - "created_at": "2026-01-22T10:42:09Z", - "document": { - "title": "Implement multi-layer caching system for VeriSimDB", - "body": "Commit 3327beaa by Your Name: Implement multi-layer caching system for VeriSimDB", - "fields": { - "type": "commit", - "hash": "3327beaaa2b39538d3fd972d6343291f304d1153", - "author": "Your Name", - "email": "you@example.com", - "commit_type": "other" - } - }, - "graph": { - "relationships": [{"predicate":"modifies","target":"file:docs/caching-strategy.adoc"},{"predicate":"modifies","target":"file:lib/verisim/query_cache.ex"},{"predicate":"modifies","target":"file:lib/verisim/query_router_cached.ex"}] - }, - "vector": { - "embedding": [-0.265625,-0.546875,0.273437,-0.773437,-0.632812,0.328125,0.117187,-0.750000,0.140625,-0.281250,0.359375,-0.289062,0.570312,0.640625,0.750000,0.507812,0.773437,0.453125,-0.179687,-0.312500,0.929687,0.460937,0.062500,-0.882812,-0.640625,0.335937,-0.804687,-0.679687,0.937500,-0.914062,0.914062,0.343750,-0.265625,-0.546875,0.273437,-0.773437,-0.632812,0.328125,0.117187,-0.750000,0.140625,-0.281250,0.359375,-0.289062,0.570312,0.640625,0.750000,0.507812,0.773437,0.453125,-0.179687,-0.312500,0.929687,0.460937,0.062500,-0.882812,-0.640625,0.335937,-0.804687,-0.679687,0.937500,-0.914062,0.914062,0.343750], - "model": "sha256-hash-64d" - }, - "tensor": { - "shape": [1, 3], - "data": [1425.0, 0.0, 3.0] - }, - "semantic": { - "types": ["https://verisim.db/self/type/Commit", "https://verisim.db/self/type/other"], - "properties": { - "conventional_commit_type": "other", - "files_changed": "3" - } - }, - "temporal": { - "timestamp": "2026-01-22T10:42:09Z", - "version": 1, - "author": "Your Name" - } -} diff --git a/verisimdb/.verisimdb/octads/commit-33c8f1c8d141.json b/verisimdb/.verisimdb/octads/commit-33c8f1c8d141.json deleted file mode 100644 index 582a980b..00000000 --- a/verisimdb/.verisimdb/octads/commit-33c8f1c8d141.json +++ /dev/null @@ -1,39 +0,0 @@ -{ - "id": "commit-33c8f1c8d141", - "source": "git-log", - "created_at": "2026-02-08T14:51:52Z", - "document": { - "title": "docs: update STATE.scm with GitHub CI and Hypatia integration status", - "body": "Commit 33c8f1c8 by Jonathan D.A. Jewell: docs: update STATE.scm with GitHub CI and Hypatia integration status", - "fields": { - "type": "commit", - "hash": "33c8f1c8d141ca480cebc88e3a4bdb039f617c49", - "author": "Jonathan D.A. Jewell", - "email": "jonathan.jewell@open.ac.uk", - "commit_type": "documentation" - } - }, - "graph": { - "relationships": [{"predicate":"modifies","target":"file:.machine_readable/STATE.scm"}] - }, - "vector": { - "embedding": [-0.867187,0.476562,-0.695312,0.140625,0.156250,0.820312,0.515625,0.750000,0.867187,0.750000,-0.148437,-0.781250,-0.968750,0.500000,0.460937,-0.750000,0.984375,0.812500,0.093750,-0.414062,-0.625000,0.734375,-0.054687,0.679687,-0.953125,-0.429687,-0.773437,-0.898437,0.539062,-0.421875,-0.882812,0.140625,-0.867187,0.476562,-0.695312,0.140625,0.156250,0.820312,0.515625,0.750000,0.867187,0.750000,-0.148437,-0.781250,-0.968750,0.500000,0.460937,-0.750000,0.984375,0.812500,0.093750,-0.414062,-0.625000,0.734375,-0.054687,0.679687,-0.953125,-0.429687,-0.773437,-0.898437,0.539062,-0.421875,-0.882812,0.140625], - "model": "sha256-hash-64d" - }, - "tensor": { - "shape": [1, 3], - "data": [39.0, 6.0, 1.0] - }, - "semantic": { - "types": ["https://verisim.db/self/type/Commit", "https://verisim.db/self/type/documentation"], - "properties": { - "conventional_commit_type": "documentation", - "files_changed": "1" - } - }, - "temporal": { - "timestamp": "2026-02-08T14:51:52Z", - "version": 1, - "author": "Jonathan D.A. Jewell" - } -} diff --git a/verisimdb/.verisimdb/octads/commit-36907e025d4a.json b/verisimdb/.verisimdb/octads/commit-36907e025d4a.json deleted file mode 100644 index b9c4c813..00000000 --- a/verisimdb/.verisimdb/octads/commit-36907e025d4a.json +++ /dev/null @@ -1,39 +0,0 @@ -{ - "id": "commit-36907e025d4a", - "source": "git-log", - "created_at": "2026-01-16T18:40:40Z", - "document": { - "title": "feat: Initial VeriSimDB repository structure", - "body": "Commit 36907e02 by Jonathan D.A. Jewell: feat: Initial VeriSimDB repository structure", - "fields": { - "type": "commit", - "hash": "36907e025d4add9f863729929ec07c387191b232", - "author": "Jonathan D.A. Jewell", - "email": "jonathan.jewell@open.ac.uk", - "commit_type": "feature" - } - }, - "graph": { - "relationships": [] - }, - "vector": { - "embedding": [-0.226562,0.515625,-0.960937,-0.234375,0.148437,0.437500,0.679687,-0.625000,-0.117187,0.523437,0.890625,0.312500,-0.132812,0.671875,-0.257812,-0.625000,0.210937,0.460937,0.640625,-0.156250,0.101562,0.781250,0.937500,-0.625000,0.570312,0.781250,0.148437,-0.781250,0.085937,0.585937,0.960937,-0.750000,-0.226562,0.515625,-0.960937,-0.234375,0.148437,0.437500,0.679687,-0.625000,-0.117187,0.523437,0.890625,0.312500,-0.132812,0.671875,-0.257812,-0.625000,0.210937,0.460937,0.640625,-0.156250,0.101562,0.781250,0.937500,-0.625000,0.570312,0.781250,0.148437,-0.781250,0.085937,0.585937,0.960937,-0.750000], - "model": "sha256-hash-64d" - }, - "tensor": { - "shape": [1, 3], - "data": [0.0, 0.0, 0.0] - }, - "semantic": { - "types": ["https://verisim.db/self/type/Commit", "https://verisim.db/self/type/feature"], - "properties": { - "conventional_commit_type": "feature", - "files_changed": "0" - } - }, - "temporal": { - "timestamp": "2026-01-16T18:40:40Z", - "version": 1, - "author": "Jonathan D.A. Jewell" - } -} diff --git a/verisimdb/.verisimdb/octads/commit-37139e0b0ba9.json b/verisimdb/.verisimdb/octads/commit-37139e0b0ba9.json deleted file mode 100644 index 0ae4ff14..00000000 --- a/verisimdb/.verisimdb/octads/commit-37139e0b0ba9.json +++ /dev/null @@ -1,39 +0,0 @@ -{ - "id": "commit-37139e0b0ba9", - "source": "git-log", - "created_at": "2026-02-13T15:14:35Z", - "document": { - "title": "feat: add proof costing, adaptive tuning, and post-processing cost model", - "body": "Commit 37139e0b by Jonathan D.A. Jewell: feat: add proof costing, adaptive tuning, and post-processing cost model", - "fields": { - "type": "commit", - "hash": "37139e0b0ba9d5679ea7e2c461c93de3bd364eeb", - "author": "Jonathan D.A. Jewell", - "email": "jonathan.jewell@open.ac.uk", - "commit_type": "feature" - } - }, - "graph": { - "relationships": [{"predicate":"modifies","target":"file:rust-core/verisim-planner/src/cost.rs"},{"predicate":"modifies","target":"file:rust-core/verisim-planner/src/lib.rs"},{"predicate":"modifies","target":"file:rust-core/verisim-planner/src/optimizer.rs"},{"predicate":"modifies","target":"file:rust-core/verisim-planner/src/stats.rs"}] - }, - "vector": { - "embedding": [0.382812,0.820312,-0.023437,0.570312,0.968750,0.625000,0.843750,-0.171875,-0.523437,0.718750,0.078125,-0.312500,0.398437,-0.476562,-0.468750,0.335937,-1.000000,-0.617187,-0.585937,0.570312,-0.281250,0.914062,-0.406250,-0.945312,0.875000,0.085937,-0.359375,0.570312,-0.929687,0.976562,0.085937,-0.039062,0.382812,0.820312,-0.023437,0.570312,0.968750,0.625000,0.843750,-0.171875,-0.523437,0.718750,0.078125,-0.312500,0.398437,-0.476562,-0.468750,0.335937,-1.000000,-0.617187,-0.585937,0.570312,-0.281250,0.914062,-0.406250,-0.945312,0.875000,0.085937,-0.359375,0.570312,-0.929687,0.976562,0.085937,-0.039062], - "model": "sha256-hash-64d" - }, - "tensor": { - "shape": [1, 3], - "data": [626.0, 11.0, 4.0] - }, - "semantic": { - "types": ["https://verisim.db/self/type/Commit", "https://verisim.db/self/type/feature"], - "properties": { - "conventional_commit_type": "feature", - "files_changed": "4" - } - }, - "temporal": { - "timestamp": "2026-02-13T15:14:35Z", - "version": 1, - "author": "Jonathan D.A. Jewell" - } -} diff --git a/verisimdb/.verisimdb/octads/commit-37dcacf06bf6.json b/verisimdb/.verisimdb/octads/commit-37dcacf06bf6.json deleted file mode 100644 index 5bf7f18f..00000000 --- a/verisimdb/.verisimdb/octads/commit-37dcacf06bf6.json +++ /dev/null @@ -1,39 +0,0 @@ -{ - "id": "commit-37dcacf06bf6", - "source": "git-log", - "created_at": "2026-01-22T09:37:47Z", - "document": { - "title": "docs: complete deployment challenges trilogy + security lessons", - "body": "Commit 37dcacf0 by Your Name: docs: complete deployment challenges trilogy + security lessons", - "fields": { - "type": "commit", - "hash": "37dcacf06bf65cfdeb961e7b14febee4d74915c8", - "author": "Your Name", - "email": "you@example.com", - "commit_type": "documentation" - } - }, - "graph": { - "relationships": [{"predicate":"modifies","target":"file:docs/challenges-federated.adoc"},{"predicate":"modifies","target":"file:docs/security-lessons.lgt"}] - }, - "vector": { - "embedding": [-0.226562,-0.375000,-0.273437,-0.867187,0.492187,0.671875,0.812500,0.890625,-0.554687,-0.046875,-0.046875,-0.070312,0.320312,-0.750000,-0.648437,0.468750,0.164062,0.367187,0.812500,-0.429687,0.500000,-0.695312,-0.640625,0.617187,0.375000,0.648437,0.437500,-0.093750,0.023437,0.968750,0.843750,0.937500,-0.226562,-0.375000,-0.273437,-0.867187,0.492187,0.671875,0.812500,0.890625,-0.554687,-0.046875,-0.046875,-0.070312,0.320312,-0.750000,-0.648437,0.468750,0.164062,0.367187,0.812500,-0.429687,0.500000,-0.695312,-0.640625,0.617187,0.375000,0.648437,0.437500,-0.093750,0.023437,0.968750,0.843750,0.937500], - "model": "sha256-hash-64d" - }, - "tensor": { - "shape": [1, 3], - "data": [1153.0, 0.0, 2.0] - }, - "semantic": { - "types": ["https://verisim.db/self/type/Commit", "https://verisim.db/self/type/documentation"], - "properties": { - "conventional_commit_type": "documentation", - "files_changed": "2" - } - }, - "temporal": { - "timestamp": "2026-01-22T09:37:47Z", - "version": 1, - "author": "Your Name" - } -} diff --git a/verisimdb/.verisimdb/octads/commit-3beb018c4b99.json b/verisimdb/.verisimdb/octads/commit-3beb018c4b99.json deleted file mode 100644 index fbcc1086..00000000 --- a/verisimdb/.verisimdb/octads/commit-3beb018c4b99.json +++ /dev/null @@ -1,39 +0,0 @@ -{ - "id": "commit-3beb018c4b99", - "source": "git-log", - "created_at": "2026-01-29T14:12:02Z", - "document": { - "title": "feat(ci): enable Hypatia scanning", - "body": "Commit 3beb018c by Test: feat(ci): enable Hypatia scanning", - "fields": { - "type": "commit", - "hash": "3beb018c4b99731c21bac8c2ab5159520c435f6c", - "author": "Test", - "email": "test@example.com", - "commit_type": "feature" - } - }, - "graph": { - "relationships": [{"predicate":"modifies","target":"file:.github/workflows/hypatia-scan.yml"}] - }, - "vector": { - "embedding": [-0.125000,0.648437,0.984375,0.156250,-0.062500,0.132812,-0.617187,-0.414062,-0.906250,-0.046875,0.789062,-0.664062,0.476562,0.906250,0.664062,0.101562,0.578125,-0.265625,0.390625,-0.171875,-0.718750,-0.265625,0.906250,-0.078125,-0.375000,0.664062,0.523437,0.554687,0.507812,-0.046875,-0.031250,-0.382812,-0.125000,0.648437,0.984375,0.156250,-0.062500,0.132812,-0.617187,-0.414062,-0.906250,-0.046875,0.789062,-0.664062,0.476562,0.906250,0.664062,0.101562,0.578125,-0.265625,0.390625,-0.171875,-0.718750,-0.265625,0.906250,-0.078125,-0.375000,0.664062,0.523437,0.554687,0.507812,-0.046875,-0.031250,-0.382812], - "model": "sha256-hash-64d" - }, - "tensor": { - "shape": [1, 3], - "data": [179.0, 0.0, 1.0] - }, - "semantic": { - "types": ["https://verisim.db/self/type/Commit", "https://verisim.db/self/type/feature"], - "properties": { - "conventional_commit_type": "feature", - "files_changed": "1" - } - }, - "temporal": { - "timestamp": "2026-01-29T14:12:02Z", - "version": 1, - "author": "Test" - } -} diff --git a/verisimdb/.verisimdb/octads/commit-3d7765bff845.json b/verisimdb/.verisimdb/octads/commit-3d7765bff845.json deleted file mode 100644 index aec3e918..00000000 --- a/verisimdb/.verisimdb/octads/commit-3d7765bff845.json +++ /dev/null @@ -1,39 +0,0 @@ -{ - "id": "commit-3d7765bff845", - "source": "git-log", - "created_at": "2026-01-22T01:35:22Z", - "document": { - "title": "Create Rescript Registry Types.md", - "body": "Commit 3d7765bf by Jonathan D.A. Jewell: Create Rescript Registry Types.md", - "fields": { - "type": "commit", - "hash": "3d7765bff8455873c3dd9ddf6fdf9c786c11aad6", - "author": "Jonathan D.A. Jewell", - "email": "6759885+hyperpolymath@users.noreply.github.com", - "commit_type": "other" - } - }, - "graph": { - "relationships": [{"predicate":"modifies","target":"file:Rescript Registry Types.md"}] - }, - "vector": { - "embedding": [-0.015625,0.257812,-0.812500,0.859375,-0.890625,0.046875,-0.781250,0.101562,-0.703125,-0.273437,-0.656250,-0.773437,0.585937,0.601562,-0.648437,0.929687,0.062500,0.507812,-0.140625,-0.210937,-0.492187,-0.429687,-0.109375,0.773437,-0.140625,-0.875000,-0.671875,0.937500,0.976562,0.812500,0.539062,0.132812,-0.015625,0.257812,-0.812500,0.859375,-0.890625,0.046875,-0.781250,0.101562,-0.703125,-0.273437,-0.656250,-0.773437,0.585937,0.601562,-0.648437,0.929687,0.062500,0.507812,-0.140625,-0.210937,-0.492187,-0.429687,-0.109375,0.773437,-0.140625,-0.875000,-0.671875,0.937500,0.976562,0.812500,0.539062,0.132812], - "model": "sha256-hash-64d" - }, - "tensor": { - "shape": [1, 3], - "data": [62.0, 0.0, 1.0] - }, - "semantic": { - "types": ["https://verisim.db/self/type/Commit", "https://verisim.db/self/type/other"], - "properties": { - "conventional_commit_type": "other", - "files_changed": "1" - } - }, - "temporal": { - "timestamp": "2026-01-22T01:35:22Z", - "version": 1, - "author": "Jonathan D.A. Jewell" - } -} diff --git a/verisimdb/.verisimdb/octads/commit-3da4daef75f9.json b/verisimdb/.verisimdb/octads/commit-3da4daef75f9.json deleted file mode 100644 index c08b0d6c..00000000 --- a/verisimdb/.verisimdb/octads/commit-3da4daef75f9.json +++ /dev/null @@ -1,39 +0,0 @@ -{ - "id": "commit-3da4daef75f9", - "source": "git-log", - "created_at": "2026-01-22T11:24:57Z", - "document": { - "title": "Fix ISO EBNF repetition operators throughout grammar", - "body": "Commit 3da4daef by Your Name: Fix ISO EBNF repetition operators throughout grammar", - "fields": { - "type": "commit", - "hash": "3da4daef75f9f778ad71ead2d388dc1444d507e0", - "author": "Your Name", - "email": "you@example.com", - "commit_type": "other" - } - }, - "graph": { - "relationships": [{"predicate":"modifies","target":"file:docs/vcl-grammar.ebnf"}] - }, - "vector": { - "embedding": [-0.242187,-0.656250,-0.242187,0.976562,-0.578125,0.734375,0.843750,-0.843750,-0.281250,-0.234375,-0.718750,0.757812,-0.242187,0.828125,0.500000,-0.882812,-0.070312,0.078125,-0.726562,-0.187500,-0.304687,-0.515625,-0.820312,-0.132812,-0.101562,0.367187,0.703125,0.484375,0.914062,0.789062,-0.718750,-0.953125,-0.242187,-0.656250,-0.242187,0.976562,-0.578125,0.734375,0.843750,-0.843750,-0.281250,-0.234375,-0.718750,0.757812,-0.242187,0.828125,0.500000,-0.882812,-0.070312,0.078125,-0.726562,-0.187500,-0.304687,-0.515625,-0.820312,-0.132812,-0.101562,0.367187,0.703125,0.484375,0.914062,0.789062,-0.718750,-0.953125], - "model": "sha256-hash-64d" - }, - "tensor": { - "shape": [1, 3], - "data": [12.0, 12.0, 1.0] - }, - "semantic": { - "types": ["https://verisim.db/self/type/Commit", "https://verisim.db/self/type/other"], - "properties": { - "conventional_commit_type": "other", - "files_changed": "1" - } - }, - "temporal": { - "timestamp": "2026-01-22T11:24:57Z", - "version": 1, - "author": "Your Name" - } -} diff --git a/verisimdb/.verisimdb/octads/commit-3df322443379.json b/verisimdb/.verisimdb/octads/commit-3df322443379.json deleted file mode 100644 index f2246387..00000000 --- a/verisimdb/.verisimdb/octads/commit-3df322443379.json +++ /dev/null @@ -1,39 +0,0 @@ -{ - "id": "commit-3df322443379", - "source": "git-log", - "created_at": "2026-02-08T14:49:20Z", - "document": { - "title": "feat: add panic-attack security scan workflow (dogfooding)", - "body": "Commit 3df32244 by Jonathan D.A. Jewell: feat: add panic-attack security scan workflow (dogfooding)", - "fields": { - "type": "commit", - "hash": "3df3224433797aa05ad1af68be60771a9db2d17f", - "author": "Jonathan D.A. Jewell", - "email": "jonathan.jewell@open.ac.uk", - "commit_type": "feature" - } - }, - "graph": { - "relationships": [{"predicate":"modifies","target":"file:.github/workflows/security-scan.yml"}] - }, - "vector": { - "embedding": [0.273437,0.328125,0.289062,-0.367187,0.601562,0.875000,-0.445312,-0.171875,-0.312500,0.703125,0.968750,-0.882812,-0.656250,0.734375,-0.335937,-0.773437,0.265625,0.179687,0.304687,0.656250,-0.539062,-0.281250,0.679687,0.476562,-0.062500,0.976562,0.945312,-0.070312,0.445312,-0.968750,0,-0.804687,0.273437,0.328125,0.289062,-0.367187,0.601562,0.875000,-0.445312,-0.171875,-0.312500,0.703125,0.968750,-0.882812,-0.656250,0.734375,-0.335937,-0.773437,0.265625,0.179687,0.304687,0.656250,-0.539062,-0.281250,0.679687,0.476562,-0.062500,0.976562,0.945312,-0.070312,0.445312,-0.968750,0,-0.804687], - "model": "sha256-hash-64d" - }, - "tensor": { - "shape": [1, 3], - "data": [17.0, 0.0, 1.0] - }, - "semantic": { - "types": ["https://verisim.db/self/type/Commit", "https://verisim.db/self/type/feature"], - "properties": { - "conventional_commit_type": "feature", - "files_changed": "1" - } - }, - "temporal": { - "timestamp": "2026-02-08T14:49:20Z", - "version": 1, - "author": "Jonathan D.A. Jewell" - } -} diff --git a/verisimdb/.verisimdb/octads/commit-3dfad170ecfb.json b/verisimdb/.verisimdb/octads/commit-3dfad170ecfb.json deleted file mode 100644 index d049ef6c..00000000 --- a/verisimdb/.verisimdb/octads/commit-3dfad170ecfb.json +++ /dev/null @@ -1,39 +0,0 @@ -{ - "id": "commit-3dfad170ecfb", - "source": "git-log", - "created_at": "2026-02-13T14:43:36Z", - "document": { - "title": "fix: replace executor stubs with real implementations, eliminate believe_me, add caching", - "body": "Commit 3dfad170 by Jonathan D.A. Jewell: fix: replace executor stubs with real implementations, eliminate believe_me, add caching", - "fields": { - "type": "commit", - "hash": "3dfad170ecfbd2c8904d745b4b80c3859e2c16f3", - "author": "Jonathan D.A. Jewell", - "email": "jonathan.jewell@open.ac.uk", - "commit_type": "bugfix" - } - }, - "graph": { - "relationships": [{"predicate":"modifies","target":"file:KNOWN-ISSUES.adoc"},{"predicate":"modifies","target":"file:debugger/src/abi/Foreign.idr"},{"predicate":"modifies","target":"file:elixir-orchestration/lib/verisim/query/vcl_bridge.ex"},{"predicate":"modifies","target":"file:elixir-orchestration/lib/verisim/query/vcl_executor.ex"},{"predicate":"modifies","target":"file:elixir-orchestration/lib/verisim/rust_client.ex"},{"predicate":"modifies","target":"file:practice-mirror/src/abi/Foreign.idr"},{"predicate":"modifies","target":"file:src/abi/Foreign.idr"}] - }, - "vector": { - "embedding": [0.460937,-0.226562,-0.101562,0.375000,0.664062,0.335937,-0.359375,0.796875,0.304687,0,-0.054687,-0.929687,-0.742187,-0.148437,-0.187500,0.375000,0.140625,-0.500000,-0.898437,0.859375,0.390625,-1.000000,0.375000,-0.757812,0.710937,-0.953125,0.875000,-0.492187,-0.250000,0.257812,-0.023437,-0.328125,0.460937,-0.226562,-0.101562,0.375000,0.664062,0.335937,-0.359375,0.796875,0.304687,0,-0.054687,-0.929687,-0.742187,-0.148437,-0.187500,0.375000,0.140625,-0.500000,-0.898437,0.859375,0.390625,-1.000000,0.375000,-0.757812,0.710937,-0.953125,0.875000,-0.492187,-0.250000,0.257812,-0.023437,-0.328125], - "model": "sha256-hash-64d" - }, - "tensor": { - "shape": [1, 3], - "data": [710.0, 96.0, 7.0] - }, - "semantic": { - "types": ["https://verisim.db/self/type/Commit", "https://verisim.db/self/type/bugfix"], - "properties": { - "conventional_commit_type": "bugfix", - "files_changed": "7" - } - }, - "temporal": { - "timestamp": "2026-02-13T14:43:36Z", - "version": 1, - "author": "Jonathan D.A. Jewell" - } -} diff --git a/verisimdb/.verisimdb/octads/commit-45e3a0230e23.json b/verisimdb/.verisimdb/octads/commit-45e3a0230e23.json deleted file mode 100644 index 5bbc63e1..00000000 --- a/verisimdb/.verisimdb/octads/commit-45e3a0230e23.json +++ /dev/null @@ -1,39 +0,0 @@ -{ - "id": "commit-45e3a0230e23", - "source": "git-log", - "created_at": "2026-01-22T00:20:37Z", - "document": { - "title": "Update README.adoc", - "body": "Commit 45e3a023 by Jonathan D.A. Jewell: Update README.adoc", - "fields": { - "type": "commit", - "hash": "45e3a0230e23108c775cffeae48db663996f4eb5", - "author": "Jonathan D.A. Jewell", - "email": "6759885+hyperpolymath@users.noreply.github.com", - "commit_type": "other" - } - }, - "graph": { - "relationships": [{"predicate":"modifies","target":"file:README.adoc"}] - }, - "vector": { - "embedding": [-0.656250,0.289062,0.242187,-0.078125,0.179687,-0.445312,0.343750,-0.132812,-0.171875,-0.273437,0.687500,0.031250,-0.101562,0.882812,0.437500,0.390625,0.859375,-0.906250,-0.617187,-0.492187,0.312500,0.320312,-0.867187,-0.976562,0.531250,0.890625,0.835937,0.226562,0.070312,0.187500,0.976562,-0.078125,-0.656250,0.289062,0.242187,-0.078125,0.179687,-0.445312,0.343750,-0.132812,-0.171875,-0.273437,0.687500,0.031250,-0.101562,0.882812,0.437500,0.390625,0.859375,-0.906250,-0.617187,-0.492187,0.312500,0.320312,-0.867187,-0.976562,0.531250,0.890625,0.835937,0.226562,0.070312,0.187500,0.976562,-0.078125], - "model": "sha256-hash-64d" - }, - "tensor": { - "shape": [1, 3], - "data": [1.0, 1.0, 1.0] - }, - "semantic": { - "types": ["https://verisim.db/self/type/Commit", "https://verisim.db/self/type/other"], - "properties": { - "conventional_commit_type": "other", - "files_changed": "1" - } - }, - "temporal": { - "timestamp": "2026-01-22T00:20:37Z", - "version": 1, - "author": "Jonathan D.A. Jewell" - } -} diff --git a/verisimdb/.verisimdb/octads/commit-4a0cbb016b03.json b/verisimdb/.verisimdb/octads/commit-4a0cbb016b03.json deleted file mode 100644 index 9dc15857..00000000 --- a/verisimdb/.verisimdb/octads/commit-4a0cbb016b03.json +++ /dev/null @@ -1,39 +0,0 @@ -{ - "id": "commit-4a0cbb016b03", - "source": "git-log", - "created_at": "2026-01-22T10:51:08Z", - "document": { - "title": "Implement comprehensive error handling and safety system", - "body": "Commit 4a0cbb01 by Your Name: Implement comprehensive error handling and safety system", - "fields": { - "type": "commit", - "hash": "4a0cbb016b031a2854f0dfd32d5a05e3cd217720", - "author": "Your Name", - "email": "you@example.com", - "commit_type": "other" - } - }, - "graph": { - "relationships": [{"predicate":"modifies","target":"file:README.adoc"},{"predicate":"modifies","target":"file:docs/error-handling-strategy.adoc"},{"predicate":"modifies","target":"file:docs/safety-and-fault-tolerance.adoc"},{"predicate":"modifies","target":"file:lib/verisim/circuit_breaker.ex"},{"predicate":"modifies","target":"file:lib/verisim/error_recovery.ex"},{"predicate":"modifies","target":"file:src/vcl/VCLError.res"}] - }, - "vector": { - "embedding": [-0.078125,-0.617187,-0.625000,-0.687500,0.460937,-0.382812,-0.585937,-0.437500,0.828125,0.414062,0.078125,-0.257812,0.500000,0.906250,0.726562,-0.960937,-0.078125,-0.312500,0.343750,0.828125,0.210937,-0.531250,-0.757812,0.453125,-0.523437,0.328125,-0.320312,-0.414062,0.781250,0.960937,-0.757812,-0.367187,-0.078125,-0.617187,-0.625000,-0.687500,0.460937,-0.382812,-0.585937,-0.437500,0.828125,0.414062,0.078125,-0.257812,0.500000,0.906250,0.726562,-0.960937,-0.078125,-0.312500,0.343750,0.828125,0.210937,-0.531250,-0.757812,0.453125,-0.523437,0.328125,-0.320312,-0.414062,0.781250,0.960937,-0.757812,-0.367187], - "model": "sha256-hash-64d" - }, - "tensor": { - "shape": [1, 3], - "data": [3151.0, 0.0, 6.0] - }, - "semantic": { - "types": ["https://verisim.db/self/type/Commit", "https://verisim.db/self/type/other"], - "properties": { - "conventional_commit_type": "other", - "files_changed": "6" - } - }, - "temporal": { - "timestamp": "2026-01-22T10:51:08Z", - "version": 1, - "author": "Your Name" - } -} diff --git a/verisimdb/.verisimdb/octads/commit-4ed47fa921d0.json b/verisimdb/.verisimdb/octads/commit-4ed47fa921d0.json deleted file mode 100644 index e93b8b4b..00000000 --- a/verisimdb/.verisimdb/octads/commit-4ed47fa921d0.json +++ /dev/null @@ -1,39 +0,0 @@ -{ - "id": "commit-4ed47fa921d0", - "source": "git-log", - "created_at": "2026-02-13T15:17:09Z", - "document": { - "title": "docs: add ZKP scheme decision (PLONK) to Trustfile and best-in-class roadmap", - "body": "Commit 4ed47fa9 by Jonathan D.A. Jewell: docs: add ZKP scheme decision (PLONK) to Trustfile and best-in-class roadmap", - "fields": { - "type": "commit", - "hash": "4ed47fa921d04f9be2ec513593f6cbfbaa5e8907", - "author": "Jonathan D.A. Jewell", - "email": "jonathan.jewell@open.ac.uk", - "commit_type": "documentation" - } - }, - "graph": { - "relationships": [{"predicate":"modifies","target":"file:BEST-IN-CLASS-ROADMAP.md"},{"predicate":"modifies","target":"file:contractiles/trust/Trustfile"}] - }, - "vector": { - "embedding": [-0.820312,0.210937,0.539062,0.921875,0.328125,0.835937,-0.085937,0.539062,0.781250,-0.429687,0.945312,0.125000,-0.257812,-0.546875,-0.320312,-0.210937,-0.554687,-0.820312,-0.875000,0.859375,0.968750,-0.437500,0.726562,-0.859375,0.187500,-0.320312,-0.679687,0.156250,-0.007812,-0.687500,-0.648437,0.203125,-0.820312,0.210937,0.539062,0.921875,0.328125,0.835937,-0.085937,0.539062,0.781250,-0.429687,0.945312,0.125000,-0.257812,-0.546875,-0.320312,-0.210937,-0.554687,-0.820312,-0.875000,0.859375,0.968750,-0.437500,0.726562,-0.859375,0.187500,-0.320312,-0.679687,0.156250,-0.007812,-0.687500,-0.648437,0.203125], - "model": "sha256-hash-64d" - }, - "tensor": { - "shape": [1, 3], - "data": [491.0, 1.0, 2.0] - }, - "semantic": { - "types": ["https://verisim.db/self/type/Commit", "https://verisim.db/self/type/documentation"], - "properties": { - "conventional_commit_type": "documentation", - "files_changed": "2" - } - }, - "temporal": { - "timestamp": "2026-02-13T15:17:09Z", - "version": 1, - "author": "Jonathan D.A. Jewell" - } -} diff --git a/verisimdb/.verisimdb/octads/commit-50ad8033edc3.json b/verisimdb/.verisimdb/octads/commit-50ad8033edc3.json deleted file mode 100644 index 02070aec..00000000 --- a/verisimdb/.verisimdb/octads/commit-50ad8033edc3.json +++ /dev/null @@ -1,39 +0,0 @@ -{ - "id": "commit-50ad8033edc3", - "source": "git-log", - "created_at": "2026-02-13T00:07:30Z", - "document": { - "title": "fix: apply safety triangle fixes (recipe-remove-believe-me,recipe-shell-quote-vars)", - "body": "Commit 50ad8033 by Jonathan D.A. Jewell: fix: apply safety triangle fixes (recipe-remove-believe-me,recipe-shell-quote-vars)", - "fields": { - "type": "commit", - "hash": "50ad8033edc3edf4839a1d55d3f3bf1db208962c", - "author": "Jonathan D.A. Jewell", - "email": "jonathan.jewell@open.ac.uk", - "commit_type": "bugfix" - } - }, - "graph": { - "relationships": [{"predicate":"modifies","target":"file:debugger/src/abi/Foreign.idr"},{"predicate":"modifies","target":"file:practice-mirror/src/abi/Foreign.idr"},{"predicate":"modifies","target":"file:src/abi/Foreign.idr"}] - }, - "vector": { - "embedding": [0.148437,-0.531250,-0.468750,-0.718750,-0.648437,-0.328125,-0.515625,0.429687,-0.031250,-0.593750,0.890625,-0.179687,-0.539062,0.710937,-0.039062,-0.937500,-0.492187,-0.742187,-0.203125,-0.703125,0.210937,0.484375,-0.828125,-0.296875,-0.226562,-0.585937,0.296875,0.992187,-0.289062,0.281250,0.234375,0.343750,0.148437,-0.531250,-0.468750,-0.718750,-0.648437,-0.328125,-0.515625,0.429687,-0.031250,-0.593750,0.890625,-0.179687,-0.539062,0.710937,-0.039062,-0.937500,-0.492187,-0.742187,-0.203125,-0.703125,0.210937,0.484375,-0.828125,-0.296875,-0.226562,-0.585937,0.296875,0.992187,-0.289062,0.281250,0.234375,0.343750], - "model": "sha256-hash-64d" - }, - "tensor": { - "shape": [1, 3], - "data": [3.0, 0.0, 3.0] - }, - "semantic": { - "types": ["https://verisim.db/self/type/Commit", "https://verisim.db/self/type/bugfix"], - "properties": { - "conventional_commit_type": "bugfix", - "files_changed": "3" - } - }, - "temporal": { - "timestamp": "2026-02-13T00:07:30Z", - "version": 1, - "author": "Jonathan D.A. Jewell" - } -} diff --git a/verisimdb/.verisimdb/octads/commit-529b85ff223e.json b/verisimdb/.verisimdb/octads/commit-529b85ff223e.json deleted file mode 100644 index fa348710..00000000 --- a/verisimdb/.verisimdb/octads/commit-529b85ff223e.json +++ /dev/null @@ -1,39 +0,0 @@ -{ - "id": "commit-529b85ff223e", - "source": "git-log", - "created_at": "2026-02-13T14:13:46Z", - "document": { - "title": "docs: add VCL comparison docs, federation readiness, and known issues", - "body": "Commit 529b85ff by Jonathan D.A. Jewell: docs: add VCL comparison docs, federation readiness, and known issues", - "fields": { - "type": "commit", - "hash": "529b85ff223e00dd18e026ca896b689d431a07ee", - "author": "Jonathan D.A. Jewell", - "email": "jonathan.jewell@open.ac.uk", - "commit_type": "documentation" - } - }, - "graph": { - "relationships": [{"predicate":"modifies","target":"file:.claude/CLAUDE.md"},{"predicate":"modifies","target":"file:.machine_readable/STATE.scm"},{"predicate":"modifies","target":"file:KNOWN-ISSUES.adoc"},{"predicate":"modifies","target":"file:debugger/Cargo.toml"},{"predicate":"modifies","target":"file:docs/federation-readiness.adoc"},{"predicate":"modifies","target":"file:docs/vcl-vs-sql.adoc"},{"predicate":"modifies","target":"file:docs/vcl-vs-vcl-dt.adoc"}] - }, - "vector": { - "embedding": [-0.710937,0.031250,-0.835937,-0.117187,0.476562,-0.562500,0.054687,0.851562,-0.835937,-0.539062,0.781250,0.304687,-0.250000,-0.898437,0.468750,0.562500,-0.398437,-0.664062,0.343750,-0.593750,-0.132812,-0.078125,0,-0.718750,0.867187,-0.390625,-0.773437,0.531250,0.804687,-1.000000,-0.507812,-0.171875,-0.710937,0.031250,-0.835937,-0.117187,0.476562,-0.562500,0.054687,0.851562,-0.835937,-0.539062,0.781250,0.304687,-0.250000,-0.898437,0.468750,0.562500,-0.398437,-0.664062,0.343750,-0.593750,-0.132812,-0.078125,0,-0.718750,0.867187,-0.390625,-0.773437,0.531250,0.804687,-1.000000,-0.507812,-0.171875], - "model": "sha256-hash-64d" - }, - "tensor": { - "shape": [1, 3], - "data": [860.0, 9.0, 7.0] - }, - "semantic": { - "types": ["https://verisim.db/self/type/Commit", "https://verisim.db/self/type/documentation"], - "properties": { - "conventional_commit_type": "documentation", - "files_changed": "7" - } - }, - "temporal": { - "timestamp": "2026-02-13T14:13:46Z", - "version": 1, - "author": "Jonathan D.A. Jewell" - } -} diff --git a/verisimdb/.verisimdb/octads/commit-5427006827c3.json b/verisimdb/.verisimdb/octads/commit-5427006827c3.json deleted file mode 100644 index c878bb17..00000000 --- a/verisimdb/.verisimdb/octads/commit-5427006827c3.json +++ /dev/null @@ -1,39 +0,0 @@ -{ - "id": "commit-5427006827c3", - "source": "git-log", - "created_at": "2026-02-08T13:57:17Z", - "document": { - "title": "feat: add verisim-api bin target and update roadmap", - "body": "Commit 54270068 by Jonathan D.A. Jewell: feat: add verisim-api bin target and update roadmap", - "fields": { - "type": "commit", - "hash": "5427006827c361376ea39f9d9e06311c50fe3850", - "author": "Jonathan D.A. Jewell", - "email": "jonathan.jewell@open.ac.uk", - "commit_type": "feature" - } - }, - "graph": { - "relationships": [{"predicate":"modifies","target":"file:ROADMAP.adoc"},{"predicate":"modifies","target":"file:rust-core/verisim-api/src/main.rs"}] - }, - "vector": { - "embedding": [0.906250,-0.101562,-0.695312,-0.718750,0.289062,0.546875,-0.296875,-0.679687,-0.335937,-0.226562,0.859375,-0.648437,0.585937,0.179687,-0.320312,0.632812,-0.218750,-0.484375,0.750000,-0.429687,-0.031250,-0.273437,0.742187,-0.953125,0.859375,-0.031250,0.734375,0.695312,0.742187,0.343750,0.351562,-0.140625,0.906250,-0.101562,-0.695312,-0.718750,0.289062,0.546875,-0.296875,-0.679687,-0.335937,-0.226562,0.859375,-0.648437,0.585937,0.179687,-0.320312,0.632812,-0.218750,-0.484375,0.750000,-0.429687,-0.031250,-0.273437,0.742187,-0.953125,0.859375,-0.031250,0.734375,0.695312,0.742187,0.343750,0.351562,-0.140625], - "model": "sha256-hash-64d" - }, - "tensor": { - "shape": [1, 3], - "data": [65.0, 0.0, 2.0] - }, - "semantic": { - "types": ["https://verisim.db/self/type/Commit", "https://verisim.db/self/type/feature"], - "properties": { - "conventional_commit_type": "feature", - "files_changed": "2" - } - }, - "temporal": { - "timestamp": "2026-02-08T13:57:17Z", - "version": 1, - "author": "Jonathan D.A. Jewell" - } -} diff --git a/verisimdb/.verisimdb/octads/commit-57aba0bd8ff4.json b/verisimdb/.verisimdb/octads/commit-57aba0bd8ff4.json deleted file mode 100644 index 55d72187..00000000 --- a/verisimdb/.verisimdb/octads/commit-57aba0bd8ff4.json +++ /dev/null @@ -1,39 +0,0 @@ -{ - "id": "commit-57aba0bd8ff4", - "source": "git-log", - "created_at": "2026-01-21T02:19:10Z", - "document": { - "title": "Initial commit", - "body": "Commit 57aba0bd by Jonathan D.A. Jewell: Initial commit", - "fields": { - "type": "commit", - "hash": "57aba0bd8ff473dfcdadd2c759a12a63808700e7", - "author": "Jonathan D.A. Jewell", - "email": "6759885+hyperpolymath@users.noreply.github.com", - "commit_type": "other" - } - }, - "graph": { - "relationships": [] - }, - "vector": { - "embedding": [-0.414062,0.867187,-0.406250,0.750000,0,-0.070312,-0.953125,0.054687,0.812500,0.031250,0.835937,0.632812,0.703125,0.117187,-0.125000,-0.789062,0.406250,-0.710937,-0.476562,0.179687,0.031250,-0.328125,0.007812,-0.382812,0.140625,0.242187,-0.273437,-0.039062,0.906250,0.390625,0,-0.062500,-0.414062,0.867187,-0.406250,0.750000,0,-0.070312,-0.953125,0.054687,0.812500,0.031250,0.835937,0.632812,0.703125,0.117187,-0.125000,-0.789062,0.406250,-0.710937,-0.476562,0.179687,0.031250,-0.328125,0.007812,-0.382812,0.140625,0.242187,-0.273437,-0.039062,0.906250,0.390625,0,-0.062500], - "model": "sha256-hash-64d" - }, - "tensor": { - "shape": [1, 3], - "data": [0.0, 0.0, 0.0] - }, - "semantic": { - "types": ["https://verisim.db/self/type/Commit", "https://verisim.db/self/type/other"], - "properties": { - "conventional_commit_type": "other", - "files_changed": "0" - } - }, - "temporal": { - "timestamp": "2026-01-21T02:19:10Z", - "version": 1, - "author": "Jonathan D.A. Jewell" - } -} diff --git a/verisimdb/.verisimdb/octads/commit-591b7006316c.json b/verisimdb/.verisimdb/octads/commit-591b7006316c.json deleted file mode 100644 index a706da67..00000000 --- a/verisimdb/.verisimdb/octads/commit-591b7006316c.json +++ /dev/null @@ -1,39 +0,0 @@ -{ - "id": "commit-591b7006316c", - "source": "git-log", - "created_at": "2026-01-22T09:36:56Z", - "document": { - "title": "Update dependabot.yml", - "body": "Commit 591b7006 by Jonathan D.A. Jewell: Update dependabot.yml", - "fields": { - "type": "commit", - "hash": "591b7006316c7f66326d6796bbd429876d742a9b", - "author": "Jonathan D.A. Jewell", - "email": "6759885+hyperpolymath@users.noreply.github.com", - "commit_type": "other" - } - }, - "graph": { - "relationships": [{"predicate":"modifies","target":"file:.github/dependabot.yml"}] - }, - "vector": { - "embedding": [0.617187,-0.484375,-0.226562,0.968750,-0.171875,0.101562,-0.570312,-0.906250,-0.531250,0.898437,0.750000,0.289062,0.281250,-0.453125,-0.945312,0.750000,-0.179687,0.375000,0.226562,-0.500000,0.593750,0.851562,-0.453125,-0.132812,-0.210937,0.914062,-0.164062,-0.882812,-0.195312,-0.390625,0.546875,0.648437,0.617187,-0.484375,-0.226562,0.968750,-0.171875,0.101562,-0.570312,-0.906250,-0.531250,0.898437,0.750000,0.289062,0.281250,-0.453125,-0.945312,0.750000,-0.179687,0.375000,0.226562,-0.500000,0.593750,0.851562,-0.453125,-0.132812,-0.210937,0.914062,-0.164062,-0.882812,-0.195312,-0.390625,0.546875,0.648437], - "model": "sha256-hash-64d" - }, - "tensor": { - "shape": [1, 3], - "data": [7.0, 7.0, 1.0] - }, - "semantic": { - "types": ["https://verisim.db/self/type/Commit", "https://verisim.db/self/type/other"], - "properties": { - "conventional_commit_type": "other", - "files_changed": "1" - } - }, - "temporal": { - "timestamp": "2026-01-22T09:36:56Z", - "version": 1, - "author": "Jonathan D.A. Jewell" - } -} diff --git a/verisimdb/.verisimdb/octads/commit-5e782e158b79.json b/verisimdb/.verisimdb/octads/commit-5e782e158b79.json deleted file mode 100644 index c20c089d..00000000 --- a/verisimdb/.verisimdb/octads/commit-5e782e158b79.json +++ /dev/null @@ -1,39 +0,0 @@ -{ - "id": "commit-5e782e158b79", - "source": "git-log", - "created_at": "2026-01-30T01:31:39Z", - "document": { - "title": "docs: add checkpoint files for state tracking", - "body": "Commit 5e782e15 by Test: docs: add checkpoint files for state tracking", - "fields": { - "type": "commit", - "hash": "5e782e158b79436607b34289d61aa435e47375f9", - "author": "Test", - "email": "test@example.com", - "commit_type": "documentation" - } - }, - "graph": { - "relationships": [{"predicate":"modifies","target":"file:ECOSYSTEM.scm"},{"predicate":"modifies","target":"file:META.scm"},{"predicate":"modifies","target":"file:STATE.scm"}] - }, - "vector": { - "embedding": [0.179687,0.015625,-0.648437,-0.664062,-0.578125,0.656250,0.632812,0.179687,-0.148437,0.117187,0.328125,0.609375,-0.648437,0.398437,0.343750,-0.867187,0.851562,-0.414062,0.289062,-0.429687,0.656250,0.109375,0.414062,-0.476562,-0.187500,0.257812,-0.312500,-0.843750,-0.578125,0.085937,-0.664062,0.570312,0.179687,0.015625,-0.648437,-0.664062,-0.578125,0.656250,0.632812,0.179687,-0.148437,0.117187,0.328125,0.609375,-0.648437,0.398437,0.343750,-0.867187,0.851562,-0.414062,0.289062,-0.429687,0.656250,0.109375,0.414062,-0.476562,-0.187500,0.257812,-0.312500,-0.843750,-0.578125,0.085937,-0.664062,0.570312], - "model": "sha256-hash-64d" - }, - "tensor": { - "shape": [1, 3], - "data": [138.0, 0.0, 3.0] - }, - "semantic": { - "types": ["https://verisim.db/self/type/Commit", "https://verisim.db/self/type/documentation"], - "properties": { - "conventional_commit_type": "documentation", - "files_changed": "3" - } - }, - "temporal": { - "timestamp": "2026-01-30T01:31:39Z", - "version": 1, - "author": "Test" - } -} diff --git a/verisimdb/.verisimdb/octads/commit-5f6d0db4cc9d.json b/verisimdb/.verisimdb/octads/commit-5f6d0db4cc9d.json deleted file mode 100644 index 507552fc..00000000 --- a/verisimdb/.verisimdb/octads/commit-5f6d0db4cc9d.json +++ /dev/null @@ -1,39 +0,0 @@ -{ - "id": "commit-5f6d0db4cc9d", - "source": "git-log", - "created_at": "2026-01-22T01:32:38Z", - "document": { - "title": "Create historiographic-custodian.html", - "body": "Commit 5f6d0db4 by Jonathan D.A. Jewell: Create historiographic-custodian.html", - "fields": { - "type": "commit", - "hash": "5f6d0db4cc9d3083a784bbb3f5b8adfab2f34d59", - "author": "Jonathan D.A. Jewell", - "email": "6759885+hyperpolymath@users.noreply.github.com", - "commit_type": "other" - } - }, - "graph": { - "relationships": [{"predicate":"modifies","target":"file:historiographic-custodian.html"}] - }, - "vector": { - "embedding": [-0.203125,-0.492187,-0.421875,-0.062500,-0.492187,0.460937,0.570312,-0.843750,0.984375,0.359375,-0.945312,0.281250,-0.421875,-0.117187,0.242187,-0.031250,-0.875000,-0.109375,-0.992187,-0.078125,-0.328125,-0.898437,-0.914062,-0.351562,0.648437,-0.335937,0.187500,-0.687500,0.578125,-0.882812,0.031250,-0.812500,-0.203125,-0.492187,-0.421875,-0.062500,-0.492187,0.460937,0.570312,-0.843750,0.984375,0.359375,-0.945312,0.281250,-0.421875,-0.117187,0.242187,-0.031250,-0.875000,-0.109375,-0.992187,-0.078125,-0.328125,-0.898437,-0.914062,-0.351562,0.648437,-0.335937,0.187500,-0.687500,0.578125,-0.882812,0.031250,-0.812500], - "model": "sha256-hash-64d" - }, - "tensor": { - "shape": [1, 3], - "data": [258.0, 0.0, 1.0] - }, - "semantic": { - "types": ["https://verisim.db/self/type/Commit", "https://verisim.db/self/type/other"], - "properties": { - "conventional_commit_type": "other", - "files_changed": "1" - } - }, - "temporal": { - "timestamp": "2026-01-22T01:32:38Z", - "version": 1, - "author": "Jonathan D.A. Jewell" - } -} diff --git a/verisimdb/.verisimdb/octads/commit-63f2c2dec2ce.json b/verisimdb/.verisimdb/octads/commit-63f2c2dec2ce.json deleted file mode 100644 index afaac589..00000000 --- a/verisimdb/.verisimdb/octads/commit-63f2c2dec2ce.json +++ /dev/null @@ -1,39 +0,0 @@ -{ - "id": "commit-63f2c2dec2ce", - "source": "git-log", - "created_at": "2026-01-22T09:45:56Z", - "document": { - "title": "Delete STATE.scm", - "body": "Commit 63f2c2de by Jonathan D.A. Jewell: Delete STATE.scm", - "fields": { - "type": "commit", - "hash": "63f2c2dec2ce60385991df85ceb898ecd9e1b042", - "author": "Jonathan D.A. Jewell", - "email": "6759885+hyperpolymath@users.noreply.github.com", - "commit_type": "other" - } - }, - "graph": { - "relationships": [{"predicate":"modifies","target":"file:STATE.scm"}] - }, - "vector": { - "embedding": [-0.562500,0.796875,-0.398437,0.546875,0.609375,-0.359375,-0.523437,-0.804687,-0.117187,0.804687,0.625000,-0.968750,0.578125,-0.132812,0.742187,0.453125,-0.070312,-0.937500,0.453125,-0.648437,0.234375,-0.195312,-0.968750,-0.703125,0.164062,0.656250,-0.781250,0.140625,0.148437,-0.492187,-0.179687,-0.695312,-0.562500,0.796875,-0.398437,0.546875,0.609375,-0.359375,-0.523437,-0.804687,-0.117187,0.804687,0.625000,-0.968750,0.578125,-0.132812,0.742187,0.453125,-0.070312,-0.937500,0.453125,-0.648437,0.234375,-0.195312,-0.968750,-0.703125,0.164062,0.656250,-0.781250,0.140625,0.148437,-0.492187,-0.179687,-0.695312], - "model": "sha256-hash-64d" - }, - "tensor": { - "shape": [1, 3], - "data": [0.0, 25.0, 1.0] - }, - "semantic": { - "types": ["https://verisim.db/self/type/Commit", "https://verisim.db/self/type/other"], - "properties": { - "conventional_commit_type": "other", - "files_changed": "1" - } - }, - "temporal": { - "timestamp": "2026-01-22T09:45:56Z", - "version": 1, - "author": "Jonathan D.A. Jewell" - } -} diff --git a/verisimdb/.verisimdb/octads/commit-667f86eff63a.json b/verisimdb/.verisimdb/octads/commit-667f86eff63a.json deleted file mode 100644 index 7ffba9dd..00000000 --- a/verisimdb/.verisimdb/octads/commit-667f86eff63a.json +++ /dev/null @@ -1,39 +0,0 @@ -{ - "id": "commit-667f86eff63a", - "source": "git-log", - "created_at": "2026-02-04T20:13:12Z", - "document": { - "title": "docs: update SCM files with project information", - "body": "Commit 667f86ef by Jonathan D.A. Jewell: docs: update SCM files with project information", - "fields": { - "type": "commit", - "hash": "667f86eff63aa1c601e89cf115b71086e47b8f58", - "author": "Jonathan D.A. Jewell", - "email": "jonathan.jewell@open.ac.uk", - "commit_type": "documentation" - } - }, - "graph": { - "relationships": [{"predicate":"modifies","target":"file:ECOSYSTEM.scm"}] - }, - "vector": { - "embedding": [0.007812,0.101562,-0.679687,0.679687,-0.039062,0.273437,0.281250,-0.562500,-0.281250,-0.976562,0.429687,0.484375,0.289062,0.484375,0.500000,-0.882812,0.265625,0.070312,-0.179687,0.460937,0.234375,0.992187,0.289062,0.742187,0.179687,-0.125000,0.515625,-0.031250,0.375000,-0.734375,-0.609375,-0.718750,0.007812,0.101562,-0.679687,0.679687,-0.039062,0.273437,0.281250,-0.562500,-0.281250,-0.976562,0.429687,0.484375,0.289062,0.484375,0.500000,-0.882812,0.265625,0.070312,-0.179687,0.460937,0.234375,0.992187,0.289062,0.742187,0.179687,-0.125000,0.515625,-0.031250,0.375000,-0.734375,-0.609375,-0.718750], - "model": "sha256-hash-64d" - }, - "tensor": { - "shape": [1, 3], - "data": [1.0, 1.0, 1.0] - }, - "semantic": { - "types": ["https://verisim.db/self/type/Commit", "https://verisim.db/self/type/documentation"], - "properties": { - "conventional_commit_type": "documentation", - "files_changed": "1" - } - }, - "temporal": { - "timestamp": "2026-02-04T20:13:12Z", - "version": 1, - "author": "Jonathan D.A. Jewell" - } -} diff --git a/verisimdb/.verisimdb/octads/commit-6d29ea7de601.json b/verisimdb/.verisimdb/octads/commit-6d29ea7de601.json deleted file mode 100644 index 82e0270e..00000000 --- a/verisimdb/.verisimdb/octads/commit-6d29ea7de601.json +++ /dev/null @@ -1,39 +0,0 @@ -{ - "id": "commit-6d29ea7de601", - "source": "git-log", - "created_at": "2026-02-12T17:07:56Z", - "document": { - "title": "feat: implement 5 Opus architectural tasks — HNSW, VCL bridge, federation, ZKP, KRaft", - "body": "Commit 6d29ea7d by Jonathan D.A. Jewell: feat: implement 5 Opus architectural tasks — HNSW, VCL bridge, federation, ZKP, KRaft", - "fields": { - "type": "commit", - "hash": "6d29ea7de60190e32232bfa5f82ebe6822c2ddf6", - "author": "Jonathan D.A. Jewell", - "email": "jonathan.jewell@open.ac.uk", - "commit_type": "feature" - } - }, - "graph": { - "relationships": [{"predicate":"modifies","target":"file:Cargo.lock"},{"predicate":"modifies","target":"file:Cargo.toml"},{"predicate":"modifies","target":"file:elixir-orchestration/lib/verisim/consensus/kraft_node.ex"},{"predicate":"modifies","target":"file:elixir-orchestration/lib/verisim/consensus/kraft_supervisor.ex"},{"predicate":"modifies","target":"file:elixir-orchestration/lib/verisim/federation/resolver.ex"},{"predicate":"modifies","target":"file:elixir-orchestration/lib/verisim/query/vcl_bridge.ex"},{"predicate":"modifies","target":"file:elixir-orchestration/lib/verisim/query/vcl_executor.ex"},{"predicate":"modifies","target":"file:rust-core/verisim-api/src/federation.rs"},{"predicate":"modifies","target":"file:rust-core/verisim-api/src/lib.rs"},{"predicate":"modifies","target":"file:rust-core/verisim-semantic/Cargo.toml"},{"predicate":"modifies","target":"file:rust-core/verisim-semantic/src/lib.rs"},{"predicate":"modifies","target":"file:rust-core/verisim-semantic/src/zkp.rs"},{"predicate":"modifies","target":"file:rust-core/verisim-vector/Cargo.toml"},{"predicate":"modifies","target":"file:rust-core/verisim-vector/src/hnsw.rs"},{"predicate":"modifies","target":"file:rust-core/verisim-vector/src/lib.rs"},{"predicate":"modifies","target":"file:src/registry/KRaftCluster.res"},{"predicate":"modifies","target":"file:src/registry/KRaftSerializer.res"},{"predicate":"modifies","target":"file:vcl-bridge/vcl_parser_port.js"}] - }, - "vector": { - "embedding": [0.593750,0.109375,0.421875,0.226562,0.382812,0.875000,0.773437,-0.492187,0.265625,-0.007812,-0.648437,0.859375,0.109375,-0.796875,-0.671875,0.093750,-0.929687,0.039062,0.148437,0.921875,-0.382812,-0.070312,-0.742187,-0.398437,0.031250,-0.929687,0.562500,0.492187,0.132812,0.531250,0.617187,-0.312500,0.593750,0.109375,0.421875,0.226562,0.382812,0.875000,0.773437,-0.492187,0.265625,-0.007812,-0.648437,0.859375,0.109375,-0.796875,-0.671875,0.093750,-0.929687,0.039062,0.148437,0.921875,-0.382812,-0.070312,-0.742187,-0.398437,0.031250,-0.929687,0.562500,0.492187,0.132812,0.531250,0.617187,-0.312500], - "model": "sha256-hash-64d" - }, - "tensor": { - "shape": [1, 3], - "data": [3998.0, 367.0, 18.0] - }, - "semantic": { - "types": ["https://verisim.db/self/type/Commit", "https://verisim.db/self/type/feature"], - "properties": { - "conventional_commit_type": "feature", - "files_changed": "18" - } - }, - "temporal": { - "timestamp": "2026-02-12T17:07:56Z", - "version": 1, - "author": "Jonathan D.A. Jewell" - } -} diff --git a/verisimdb/.verisimdb/octads/commit-71dc72521bef.json b/verisimdb/.verisimdb/octads/commit-71dc72521bef.json deleted file mode 100644 index a723d9a0..00000000 --- a/verisimdb/.verisimdb/octads/commit-71dc72521bef.json +++ /dev/null @@ -1,39 +0,0 @@ -{ - "id": "commit-71dc72521bef", - "source": "git-log", - "created_at": "2026-01-22T00:55:50Z", - "document": { - "title": "Create WHITEPAPER.md", - "body": "Commit 71dc7252 by Jonathan D.A. Jewell: Create WHITEPAPER.md", - "fields": { - "type": "commit", - "hash": "71dc72521bef87493483cca7a952b4ba93644083", - "author": "Jonathan D.A. Jewell", - "email": "6759885+hyperpolymath@users.noreply.github.com", - "commit_type": "other" - } - }, - "graph": { - "relationships": [{"predicate":"modifies","target":"file:WHITEPAPER.md"}] - }, - "vector": { - "embedding": [-0.859375,0.273437,-0.015625,-0.375000,0.578125,-0.695312,-0.203125,0.421875,0.125000,0.765625,-0.718750,-0.335937,0.203125,0.242187,0.664062,0.695312,0.039062,0.789062,-0.406250,-0.750000,-0.140625,-0.117187,-0.468750,0.953125,-0.460937,0.992187,-0.242187,0.273437,-0.515625,-0.187500,-0.171875,0.414062,-0.859375,0.273437,-0.015625,-0.375000,0.578125,-0.695312,-0.203125,0.421875,0.125000,0.765625,-0.718750,-0.335937,0.203125,0.242187,0.664062,0.695312,0.039062,0.789062,-0.406250,-0.750000,-0.140625,-0.117187,-0.468750,0.953125,-0.460937,0.992187,-0.242187,0.273437,-0.515625,-0.187500,-0.171875,0.414062], - "model": "sha256-hash-64d" - }, - "tensor": { - "shape": [1, 3], - "data": [385.0, 0.0, 1.0] - }, - "semantic": { - "types": ["https://verisim.db/self/type/Commit", "https://verisim.db/self/type/other"], - "properties": { - "conventional_commit_type": "other", - "files_changed": "1" - } - }, - "temporal": { - "timestamp": "2026-01-22T00:55:50Z", - "version": 1, - "author": "Jonathan D.A. Jewell" - } -} diff --git a/verisimdb/.verisimdb/octads/commit-737a6e822c52.json b/verisimdb/.verisimdb/octads/commit-737a6e822c52.json deleted file mode 100644 index e1b1497b..00000000 --- a/verisimdb/.verisimdb/octads/commit-737a6e822c52.json +++ /dev/null @@ -1,39 +0,0 @@ -{ - "id": "commit-737a6e822c52", - "source": "git-log", - "created_at": "2026-01-22T01:31:02Z", - "document": { - "title": "Create ZKP and Sanctify Integration.md", - "body": "Commit 737a6e82 by Jonathan D.A. Jewell: Create ZKP and Sanctify Integration.md", - "fields": { - "type": "commit", - "hash": "737a6e822c522923a90a0d3f27025983fc64ff70", - "author": "Jonathan D.A. Jewell", - "email": "6759885+hyperpolymath@users.noreply.github.com", - "commit_type": "other" - } - }, - "graph": { - "relationships": [{"predicate":"modifies","target":"file:ZKP and Sanctify Integration.md"}] - }, - "vector": { - "embedding": [0.554687,0.664062,-0.679687,0.148437,-0.601562,-0.328125,0.585937,-0.500000,0.109375,-0.187500,0.023437,-0.656250,-0.648437,-0.625000,-0.109375,-0.164062,0.773437,0.625000,0.281250,-0.601562,0.007812,-0.492187,-0.406250,0.703125,-0.914062,-0.804687,0.218750,-0.898437,-0.804687,0.671875,-0.164062,0.859375,0.554687,0.664062,-0.679687,0.148437,-0.601562,-0.328125,0.585937,-0.500000,0.109375,-0.187500,0.023437,-0.656250,-0.648437,-0.625000,-0.109375,-0.164062,0.773437,0.625000,0.281250,-0.601562,0.007812,-0.492187,-0.406250,0.703125,-0.914062,-0.804687,0.218750,-0.898437,-0.804687,0.671875,-0.164062,0.859375], - "model": "sha256-hash-64d" - }, - "tensor": { - "shape": [1, 3], - "data": [61.0, 0.0, 1.0] - }, - "semantic": { - "types": ["https://verisim.db/self/type/Commit", "https://verisim.db/self/type/other"], - "properties": { - "conventional_commit_type": "other", - "files_changed": "1" - } - }, - "temporal": { - "timestamp": "2026-01-22T01:31:02Z", - "version": 1, - "author": "Jonathan D.A. Jewell" - } -} diff --git a/verisimdb/.verisimdb/octads/commit-74f45a7491e2.json b/verisimdb/.verisimdb/octads/commit-74f45a7491e2.json deleted file mode 100644 index 129a3dfb..00000000 --- a/verisimdb/.verisimdb/octads/commit-74f45a7491e2.json +++ /dev/null @@ -1,39 +0,0 @@ -{ - "id": "commit-74f45a7491e2", - "source": "git-log", - "created_at": "2026-01-22T11:20:11Z", - "document": { - "title": "Comprehensive ISO EBNF compliance fixes for VCL grammar", - "body": "Commit 74f45a74 by Your Name: Comprehensive ISO EBNF compliance fixes for VCL grammar", - "fields": { - "type": "commit", - "hash": "74f45a7491e2b94e48ca2ab07de2ba8330575f12", - "author": "Your Name", - "email": "you@example.com", - "commit_type": "other" - } - }, - "graph": { - "relationships": [{"predicate":"modifies","target":"file:docs/vcl-grammar.ebnf"}] - }, - "vector": { - "embedding": [0.437500,-0.914062,-0.062500,-0.132812,-0.312500,0.921875,-0.476562,-0.335937,-0.054687,0.171875,-0.273437,0.882812,-0.648437,-0.406250,0.390625,-0.046875,0.085937,0.593750,-0.914062,0.906250,-0.578125,-0.437500,-0.312500,-0.437500,0.664062,0.539062,0.671875,-0.085937,0.960937,0.992187,-0.375000,0.359375,0.437500,-0.914062,-0.062500,-0.132812,-0.312500,0.921875,-0.476562,-0.335937,-0.054687,0.171875,-0.273437,0.882812,-0.648437,-0.406250,0.390625,-0.046875,0.085937,0.593750,-0.914062,0.906250,-0.578125,-0.437500,-0.312500,-0.437500,0.664062,0.539062,0.671875,-0.085937,0.960937,0.992187,-0.375000,0.359375], - "model": "sha256-hash-64d" - }, - "tensor": { - "shape": [1, 3], - "data": [55.0, 55.0, 1.0] - }, - "semantic": { - "types": ["https://verisim.db/self/type/Commit", "https://verisim.db/self/type/other"], - "properties": { - "conventional_commit_type": "other", - "files_changed": "1" - } - }, - "temporal": { - "timestamp": "2026-01-22T11:20:11Z", - "version": 1, - "author": "Your Name" - } -} diff --git a/verisimdb/.verisimdb/octads/commit-77d2e9f3f088.json b/verisimdb/.verisimdb/octads/commit-77d2e9f3f088.json deleted file mode 100644 index 7e4467ae..00000000 --- a/verisimdb/.verisimdb/octads/commit-77d2e9f3f088.json +++ /dev/null @@ -1,39 +0,0 @@ -{ - "id": "commit-77d2e9f3f088", - "source": "git-log", - "created_at": "2026-02-13T16:14:44Z", - "document": { - "title": "feat: complete 7-phase security, operations & feature hardening", - "body": "Commit 77d2e9f3 by Jonathan D.A. Jewell: feat: complete 7-phase security, operations & feature hardening", - "fields": { - "type": "commit", - "hash": "77d2e9f3f08841b091440b66ba3664d184620f21", - "author": "Jonathan D.A. Jewell", - "email": "jonathan.jewell@open.ac.uk", - "commit_type": "feature" - } - }, - "graph": { - "relationships": [] - }, - "vector": { - "embedding": [-0.500000,-0.648437,0.539062,-0.890625,0.914062,-0.976562,0.523437,-0.421875,0.218750,0.359375,-0.843750,0.710937,-0.882812,0.976562,0.632812,0.406250,0.562500,-0.429687,0.070312,-0.171875,-0.929687,0.070312,-0.679687,-0.648437,-0.546875,0.046875,0.367187,-0.265625,0.984375,0.867187,-0.328125,0.078125,-0.500000,-0.648437,0.539062,-0.890625,0.914062,-0.976562,0.523437,-0.421875,0.218750,0.359375,-0.843750,0.710937,-0.882812,0.976562,0.632812,0.406250,0.562500,-0.429687,0.070312,-0.171875,-0.929687,0.070312,-0.679687,-0.648437,-0.546875,0.046875,0.367187,-0.265625,0.984375,0.867187,-0.328125,0.078125], - "model": "sha256-hash-64d" - }, - "tensor": { - "shape": [1, 3], - "data": [12560.0, 240.0, 61.0] - }, - "semantic": { - "types": ["https://verisim.db/self/type/Commit", "https://verisim.db/self/type/feature"], - "properties": { - "conventional_commit_type": "feature", - "files_changed": "61" - } - }, - "temporal": { - "timestamp": "2026-02-13T16:14:44Z", - "version": 1, - "author": "Jonathan D.A. Jewell" - } -} diff --git a/verisimdb/.verisimdb/octads/commit-7b8c073708d5.json b/verisimdb/.verisimdb/octads/commit-7b8c073708d5.json deleted file mode 100644 index 9b9ec744..00000000 --- a/verisimdb/.verisimdb/octads/commit-7b8c073708d5.json +++ /dev/null @@ -1,39 +0,0 @@ -{ - "id": "commit-7b8c073708d5", - "source": "git-log", - "created_at": "2026-01-22T12:15:21Z", - "document": { - "title": "Add VCL formal semantics, type system, and enhanced error handling", - "body": "Commit 7b8c0737 by Your Name: Add VCL formal semantics, type system, and enhanced error handling", - "fields": { - "type": "commit", - "hash": "7b8c073708d52387a17083604e8fcafb6e15a145", - "author": "Your Name", - "email": "you@example.com", - "commit_type": "other" - } - }, - "graph": { - "relationships": [{"predicate":"modifies","target":"file:README.adoc"},{"predicate":"modifies","target":"file:docs/error-handling-strategy.adoc"},{"predicate":"modifies","target":"file:docs/vcl-formal-semantics.adoc"},{"predicate":"modifies","target":"file:docs/vcl-type-system.adoc"}] - }, - "vector": { - "embedding": [0.281250,-0.484375,0.898437,0.820312,-0.945312,-0.960937,0.976562,0.031250,0.984375,0.664062,-0.734375,0.226562,-0.882812,0.851562,-0.703125,0.601562,0.406250,0.140625,0.812500,0.703125,-0.101562,-0.671875,-0.070312,0.187500,-0.531250,0.023437,-0.234375,0.039062,-0.882812,0.773437,-0.546875,0.585937,0.281250,-0.484375,0.898437,0.820312,-0.945312,-0.960937,0.976562,0.031250,0.984375,0.664062,-0.734375,0.226562,-0.882812,0.851562,-0.703125,0.601562,0.406250,0.140625,0.812500,0.703125,-0.101562,-0.671875,-0.070312,0.187500,-0.531250,0.023437,-0.234375,0.039062,-0.882812,0.773437,-0.546875,0.585937], - "model": "sha256-hash-64d" - }, - "tensor": { - "shape": [1, 3], - "data": [1957.0, 1.0, 4.0] - }, - "semantic": { - "types": ["https://verisim.db/self/type/Commit", "https://verisim.db/self/type/other"], - "properties": { - "conventional_commit_type": "other", - "files_changed": "4" - } - }, - "temporal": { - "timestamp": "2026-01-22T12:15:21Z", - "version": 1, - "author": "Your Name" - } -} diff --git a/verisimdb/.verisimdb/octads/commit-7cec9f8d1b08.json b/verisimdb/.verisimdb/octads/commit-7cec9f8d1b08.json deleted file mode 100644 index 169d7fae..00000000 --- a/verisimdb/.verisimdb/octads/commit-7cec9f8d1b08.json +++ /dev/null @@ -1,39 +0,0 @@ -{ - "id": "commit-7cec9f8d1b08", - "source": "git-log", - "created_at": "2026-01-22T09:14:29Z", - "document": { - "title": "docs: clarify deployment modes and convert tech docs to AsciiDoc", - "body": "Commit 7cec9f8d by Your Name: docs: clarify deployment modes and convert tech docs to AsciiDoc", - "fields": { - "type": "commit", - "hash": "7cec9f8d1b086b651a001bd606e87070141bb9ad", - "author": "Your Name", - "email": "you@example.com", - "commit_type": "documentation" - } - }, - "graph": { - "relationships": [{"predicate":"modifies","target":"file:Cargo.lock"},{"predicate":"modifies","target":"file:README.adoc"},{"predicate":"modifies","target":"file:docs/deployment-modes.adoc"},{"predicate":"modifies","target":"file:docs/rescript-registry-types.adoc"},{"predicate":"modifies","target":"file:docs/snapshotting-and-truncation-logic.adoc"},{"predicate":"modifies","target":"file:docs/technical-specification-kraft-metadata-log.adoc"},{"predicate":"modifies","target":"file:docs/zkp-and-sanctify-integration.adoc"}] - }, - "vector": { - "embedding": [0.179687,-0.710937,0.367187,0.234375,0.937500,-0.343750,-0.609375,0.875000,-0.117187,0.359375,0.375000,0.437500,0.070312,0.390625,0.578125,-0.828125,0.851562,-0.351562,-0.945312,0.570312,0.789062,0.773437,0.007812,0.742187,0.414062,-0.531250,-0.554687,-0.148437,0.437500,-0.343750,0.570312,0.593750,0.179687,-0.710937,0.367187,0.234375,0.937500,-0.343750,-0.609375,0.875000,-0.117187,0.359375,0.375000,0.437500,0.070312,0.390625,0.578125,-0.828125,0.851562,-0.351562,-0.945312,0.570312,0.789062,0.773437,0.007812,0.742187,0.414062,-0.531250,-0.554687,-0.148437,0.437500,-0.343750,0.570312,0.593750], - "model": "sha256-hash-64d" - }, - "tensor": { - "shape": [1, 3], - "data": [552.0, 39.0, 7.0] - }, - "semantic": { - "types": ["https://verisim.db/self/type/Commit", "https://verisim.db/self/type/documentation"], - "properties": { - "conventional_commit_type": "documentation", - "files_changed": "7" - } - }, - "temporal": { - "timestamp": "2026-01-22T09:14:29Z", - "version": 1, - "author": "Your Name" - } -} diff --git a/verisimdb/.verisimdb/octads/commit-7de3adf6c8a0.json b/verisimdb/.verisimdb/octads/commit-7de3adf6c8a0.json deleted file mode 100644 index 1fc2caf6..00000000 --- a/verisimdb/.verisimdb/octads/commit-7de3adf6c8a0.json +++ /dev/null @@ -1,39 +0,0 @@ -{ - "id": "commit-7de3adf6c8a0", - "source": "git-log", - "created_at": "2026-02-04T19:20:11Z", - "document": { - "title": "docs: add SCM checkpoint files", - "body": "Commit 7de3adf6 by Jonathan D.A. Jewell: docs: add SCM checkpoint files", - "fields": { - "type": "commit", - "hash": "7de3adf6c8a0dd387a1181375235ce044c07a393", - "author": "Jonathan D.A. Jewell", - "email": "jonathan.jewell@open.ac.uk", - "commit_type": "documentation" - } - }, - "graph": { - "relationships": [{"predicate":"modifies","target":"file:ECOSYSTEM.scm"},{"predicate":"modifies","target":"file:META.scm"},{"predicate":"modifies","target":"file:STATE.scm"}] - }, - "vector": { - "embedding": [-0.109375,-0.976562,0.906250,0.218750,0.539062,0.554687,0.648437,-0.265625,-1.000000,-0.085937,-0.921875,-0.562500,-0.679687,0.773437,0.867187,0.617187,-0.421875,0.187500,0.929687,-0.367187,0.421875,0.804687,0.562500,0.726562,0.414062,-0.210937,0.335937,-0.367187,-0.390625,-0.718750,-0.406250,-0.664062,-0.109375,-0.976562,0.906250,0.218750,0.539062,0.554687,0.648437,-0.265625,-1.000000,-0.085937,-0.921875,-0.562500,-0.679687,0.773437,0.867187,0.617187,-0.421875,0.187500,0.929687,-0.367187,0.421875,0.804687,0.562500,0.726562,0.414062,-0.210937,0.335937,-0.367187,-0.390625,-0.718750,-0.406250,-0.664062], - "model": "sha256-hash-64d" - }, - "tensor": { - "shape": [1, 3], - "data": [83.0, 0.0, 3.0] - }, - "semantic": { - "types": ["https://verisim.db/self/type/Commit", "https://verisim.db/self/type/documentation"], - "properties": { - "conventional_commit_type": "documentation", - "files_changed": "3" - } - }, - "temporal": { - "timestamp": "2026-02-04T19:20:11Z", - "version": 1, - "author": "Jonathan D.A. Jewell" - } -} diff --git a/verisimdb/.verisimdb/octads/commit-8012d86a0882.json b/verisimdb/.verisimdb/octads/commit-8012d86a0882.json deleted file mode 100644 index 613e2228..00000000 --- a/verisimdb/.verisimdb/octads/commit-8012d86a0882.json +++ /dev/null @@ -1,39 +0,0 @@ -{ - "id": "commit-8012d86a0882", - "source": "git-log", - "created_at": "2026-01-22T09:47:26Z", - "document": { - "title": "fix(deps): resolve gix-features SHA-1 collision vulnerability", - "body": "Commit 8012d86a by Your Name: fix(deps): resolve gix-features SHA-1 collision vulnerability", - "fields": { - "type": "commit", - "hash": "8012d86a0882bba4e9f8023b3425485edf29db06", - "author": "Your Name", - "email": "you@example.com", - "commit_type": "bugfix" - } - }, - "graph": { - "relationships": [{"predicate":"modifies","target":"file:Cargo.lock"},{"predicate":"modifies","target":"file:Cargo.toml"}] - }, - "vector": { - "embedding": [0.945312,-0.125000,0.218750,0.054687,0.640625,0.585937,-0.210937,-0.007812,0.429687,-0.859375,-0.171875,0.664062,-0.132812,-0.023437,-0.132812,0.695312,-0.390625,0.710937,-0.742187,0.054687,0.359375,-0.515625,0.070312,-0.390625,-0.226562,-0.304687,0.921875,-0.148437,-0.320312,0.953125,-0.421875,0.843750,0.945312,-0.125000,0.218750,0.054687,0.640625,0.585937,-0.210937,-0.007812,0.429687,-0.859375,-0.171875,0.664062,-0.132812,-0.023437,-0.132812,0.695312,-0.390625,0.710937,-0.742187,0.054687,0.359375,-0.515625,0.070312,-0.390625,-0.226562,-0.304687,0.921875,-0.148437,-0.320312,0.953125,-0.421875,0.843750], - "model": "sha256-hash-64d" - }, - "tensor": { - "shape": [1, 3], - "data": [6.0, 2.0, 2.0] - }, - "semantic": { - "types": ["https://verisim.db/self/type/Commit", "https://verisim.db/self/type/bugfix"], - "properties": { - "conventional_commit_type": "bugfix", - "files_changed": "2" - } - }, - "temporal": { - "timestamp": "2026-01-22T09:47:26Z", - "version": 1, - "author": "Your Name" - } -} diff --git a/verisimdb/.verisimdb/octads/commit-854ea6d13ae7.json b/verisimdb/.verisimdb/octads/commit-854ea6d13ae7.json deleted file mode 100644 index b94363dc..00000000 --- a/verisimdb/.verisimdb/octads/commit-854ea6d13ae7.json +++ /dev/null @@ -1,39 +0,0 @@ -{ - "id": "commit-854ea6d13ae7", - "source": "git-log", - "created_at": "2026-01-21T01:59:00Z", - "document": { - "title": "Revert to PMPL-1.0-or-later (ethical/quantum-safe provenance)", - "body": "Commit 854ea6d1 by Your Name: Revert to PMPL-1.0-or-later (ethical/quantum-safe provenance)", - "fields": { - "type": "commit", - "hash": "854ea6d13ae79f6bcb3cda857f537ec331c15f11", - "author": "Your Name", - "email": "you@example.com", - "commit_type": "other" - } - }, - "graph": { - "relationships": [{"predicate":"modifies","target":"file:LICENSE"}] - }, - "vector": { - "embedding": [0.570312,-0.851562,0.976562,-0.414062,-0.625000,-0.718750,0.843750,0.468750,-0.742187,-0.976562,-0.507812,0.015625,0.281250,0.367187,-0.773437,0,-0.953125,-0.132812,0.328125,-1.000000,-0.296875,0.257812,0.585937,0.406250,-0.382812,0.382812,-0.062500,-0.468750,-0.687500,0.148437,-0.554687,-0.382812,0.570312,-0.851562,0.976562,-0.414062,-0.625000,-0.718750,0.843750,0.468750,-0.742187,-0.976562,-0.507812,0.015625,0.281250,0.367187,-0.773437,0,-0.953125,-0.132812,0.328125,-1.000000,-0.296875,0.257812,0.585937,0.406250,-0.382812,0.382812,-0.062500,-0.468750,-0.687500,0.148437,-0.554687,-0.382812], - "model": "sha256-hash-64d" - }, - "tensor": { - "shape": [1, 3], - "data": [12.0, 12.0, 1.0] - }, - "semantic": { - "types": ["https://verisim.db/self/type/Commit", "https://verisim.db/self/type/other"], - "properties": { - "conventional_commit_type": "other", - "files_changed": "1" - } - }, - "temporal": { - "timestamp": "2026-01-21T01:59:00Z", - "version": 1, - "author": "Your Name" - } -} diff --git a/verisimdb/.verisimdb/octads/commit-86581c2638b8.json b/verisimdb/.verisimdb/octads/commit-86581c2638b8.json deleted file mode 100644 index 9ae2e1de..00000000 --- a/verisimdb/.verisimdb/octads/commit-86581c2638b8.json +++ /dev/null @@ -1,39 +0,0 @@ -{ - "id": "commit-86581c2638b8", - "source": "git-log", - "created_at": "2026-02-01T01:49:17Z", - "document": { - "title": "chore: sync", - "body": "Commit 86581c26 by Jonathan D.A. Jewell: chore: sync", - "fields": { - "type": "commit", - "hash": "86581c2638b8b414ceb9fc7d6ecdcf519a7f813f", - "author": "Jonathan D.A. Jewell", - "email": "jonathan.jewell@open.ac.uk", - "commit_type": "chore" - } - }, - "graph": { - "relationships": [] - }, - "vector": { - "embedding": [0.367187,-0.210937,-0.062500,0.625000,-0.882812,0.632812,-0.320312,-0.679687,0.937500,0.929687,-0.125000,0.203125,-0.765625,-0.382812,0.156250,-0.742187,-0.695312,0.101562,0.296875,0.593750,0.554687,0.500000,0.351562,0.640625,-0.812500,-0.500000,-0.453125,0.718750,0,-0.320312,-0.625000,-0.859375,0.367187,-0.210937,-0.062500,0.625000,-0.882812,0.632812,-0.320312,-0.679687,0.937500,0.929687,-0.125000,0.203125,-0.765625,-0.382812,0.156250,-0.742187,-0.695312,0.101562,0.296875,0.593750,0.554687,0.500000,0.351562,0.640625,-0.812500,-0.500000,-0.453125,0.718750,0,-0.320312,-0.625000,-0.859375], - "model": "sha256-hash-64d" - }, - "tensor": { - "shape": [1, 3], - "data": [1600.0, 537.0, 46.0] - }, - "semantic": { - "types": ["https://verisim.db/self/type/Commit", "https://verisim.db/self/type/chore"], - "properties": { - "conventional_commit_type": "chore", - "files_changed": "46" - } - }, - "temporal": { - "timestamp": "2026-02-01T01:49:17Z", - "version": 1, - "author": "Jonathan D.A. Jewell" - } -} diff --git a/verisimdb/.verisimdb/octads/commit-89ea7188af80.json b/verisimdb/.verisimdb/octads/commit-89ea7188af80.json deleted file mode 100644 index 0b1cbcae..00000000 --- a/verisimdb/.verisimdb/octads/commit-89ea7188af80.json +++ /dev/null @@ -1,39 +0,0 @@ -{ - "id": "commit-89ea7188af80", - "source": "git-log", - "created_at": "2026-01-27T01:01:07Z", - "document": { - "title": "chore(deps): bump oneshot in the cargo group across 1 directory", - "body": "Commit 89ea7188 by dependabot[bot]: chore(deps): bump oneshot in the cargo group across 1 directory", - "fields": { - "type": "commit", - "hash": "89ea7188af80186b909a74e7d3baaf4688515288", - "author": "dependabot[bot]", - "email": "49699333+dependabot[bot]@users.noreply.github.com", - "commit_type": "chore" - } - }, - "graph": { - "relationships": [{"predicate":"modifies","target":"file:Cargo.lock"}] - }, - "vector": { - "embedding": [-0.015625,-0.078125,-0.359375,0.984375,0.765625,0.406250,-0.101562,0.898437,0.414062,0.125000,0.843750,-0.875000,-0.671875,0.171875,-0.218750,0.570312,-0.265625,-0.851562,0.765625,0.992187,-0.039062,-0.218750,-0.070312,0.320312,-0.289062,-0.515625,-0.765625,-0.578125,-0.328125,0.468750,0.031250,0.421875,-0.015625,-0.078125,-0.359375,0.984375,0.765625,0.406250,-0.101562,0.898437,0.414062,0.125000,0.843750,-0.875000,-0.671875,0.171875,-0.218750,0.570312,-0.265625,-0.851562,0.765625,0.992187,-0.039062,-0.218750,-0.070312,0.320312,-0.289062,-0.515625,-0.765625,-0.578125,-0.328125,0.468750,0.031250,0.421875], - "model": "sha256-hash-64d" - }, - "tensor": { - "shape": [1, 3], - "data": [0.0, 0.0, 1.0] - }, - "semantic": { - "types": ["https://verisim.db/self/type/Commit", "https://verisim.db/self/type/chore"], - "properties": { - "conventional_commit_type": "chore", - "files_changed": "1" - } - }, - "temporal": { - "timestamp": "2026-01-27T01:01:07Z", - "version": 1, - "author": "dependabot[bot]" - } -} diff --git a/verisimdb/.verisimdb/octads/commit-8cd1f416878e.json b/verisimdb/.verisimdb/octads/commit-8cd1f416878e.json deleted file mode 100644 index e8861318..00000000 --- a/verisimdb/.verisimdb/octads/commit-8cd1f416878e.json +++ /dev/null @@ -1,39 +0,0 @@ -{ - "id": "commit-8cd1f416878e", - "source": "git-log", - "created_at": "2026-01-22T10:18:52Z", - "document": { - "title": "docs(vcl): add VCL grammar, examples, architecture, and parser", - "body": "Commit 8cd1f416 by Your Name: docs(vcl): add VCL grammar, examples, architecture, and parser", - "fields": { - "type": "commit", - "hash": "8cd1f416878e9f1735239b41e14246efa5c6ba12", - "author": "Your Name", - "email": "you@example.com", - "commit_type": "documentation" - } - }, - "graph": { - "relationships": [{"predicate":"modifies","target":"file:README.adoc"},{"predicate":"modifies","target":"file:docs/vcl-architecture.adoc"},{"predicate":"modifies","target":"file:docs/vcl-examples.adoc"},{"predicate":"modifies","target":"file:docs/vcl-grammar.ebnf"},{"predicate":"modifies","target":"file:src/vcl/VCLParser.res"},{"predicate":"modifies","target":"file:src/vcl/VCLParser_test.res"}] - }, - "vector": { - "embedding": [0.046875,0.351562,-0.851562,-0.640625,0.164062,0.132812,0.023437,0.039062,0.601562,-0.210937,-0.210937,0.718750,-0.195312,0.156250,0.546875,0.601562,0,0.328125,-0.046875,-0.367187,-0.117187,-0.781250,0.992187,-0.414062,-0.750000,0.476562,-0.187500,0.820312,0.890625,-0.593750,0.078125,-0.703125,0.046875,0.351562,-0.851562,-0.640625,0.164062,0.132812,0.023437,0.039062,0.601562,-0.210937,-0.210937,0.718750,-0.195312,0.156250,0.546875,0.601562,0,0.328125,-0.046875,-0.367187,-0.117187,-0.781250,0.992187,-0.414062,-0.750000,0.476562,-0.187500,0.820312,0.890625,-0.593750,0.078125,-0.703125], - "model": "sha256-hash-64d" - }, - "tensor": { - "shape": [1, 3], - "data": [2458.0, 0.0, 6.0] - }, - "semantic": { - "types": ["https://verisim.db/self/type/Commit", "https://verisim.db/self/type/documentation"], - "properties": { - "conventional_commit_type": "documentation", - "files_changed": "6" - } - }, - "temporal": { - "timestamp": "2026-01-22T10:18:52Z", - "version": 1, - "author": "Your Name" - } -} diff --git a/verisimdb/.verisimdb/octads/commit-91a08d99ee55.json b/verisimdb/.verisimdb/octads/commit-91a08d99ee55.json deleted file mode 100644 index 139c5e78..00000000 --- a/verisimdb/.verisimdb/octads/commit-91a08d99ee55.json +++ /dev/null @@ -1,39 +0,0 @@ -{ - "id": "commit-91a08d99ee55", - "source": "git-log", - "created_at": "2026-02-04T21:46:11Z", - "document": { - "title": "feat: complete Elixir orchestration layer", - "body": "Commit 91a08d99 by Jonathan D.A. Jewell: feat: complete Elixir orchestration layer", - "fields": { - "type": "commit", - "hash": "91a08d99ee55a480c08e139c5e8bb1f49b62d117", - "author": "Jonathan D.A. Jewell", - "email": "jonathan.jewell@open.ac.uk", - "commit_type": "feature" - } - }, - "graph": { - "relationships": [{"predicate":"modifies","target":"file:elixir-orchestration/lib/verisim/query/vcl_executor.ex"}] - }, - "vector": { - "embedding": [0.390625,-0.710937,0.023437,-0.648437,-0.601562,-0.101562,0.828125,-0.132812,-0.835937,-0.507812,0.242187,0.210937,0.156250,0.937500,-0.406250,0.789062,-0.929687,-0.242187,0.273437,0.671875,-0.101562,0.851562,-1.000000,-0.351562,0.890625,-0.304687,-0.617187,0.171875,0.523437,0.625000,0.054687,0.375000,0.390625,-0.710937,0.023437,-0.648437,-0.601562,-0.101562,0.828125,-0.132812,-0.835937,-0.507812,0.242187,0.210937,0.156250,0.937500,-0.406250,0.789062,-0.929687,-0.242187,0.273437,0.671875,-0.101562,0.851562,-1.000000,-0.351562,0.890625,-0.304687,-0.617187,0.171875,0.523437,0.625000,0.054687,0.375000], - "model": "sha256-hash-64d" - }, - "tensor": { - "shape": [1, 3], - "data": [275.0, 0.0, 1.0] - }, - "semantic": { - "types": ["https://verisim.db/self/type/Commit", "https://verisim.db/self/type/feature"], - "properties": { - "conventional_commit_type": "feature", - "files_changed": "1" - } - }, - "temporal": { - "timestamp": "2026-02-04T21:46:11Z", - "version": 1, - "author": "Jonathan D.A. Jewell" - } -} diff --git a/verisimdb/.verisimdb/octads/commit-9746771f7457.json b/verisimdb/.verisimdb/octads/commit-9746771f7457.json deleted file mode 100644 index 486a4324..00000000 --- a/verisimdb/.verisimdb/octads/commit-9746771f7457.json +++ /dev/null @@ -1,39 +0,0 @@ -{ - "id": "commit-9746771f7457", - "source": "git-log", - "created_at": "2026-01-22T11:57:33Z", - "document": { - "title": "Fix missing comma after repetition block in vector_literal", - "body": "Commit 9746771f by Your Name: Fix missing comma after repetition block in vector_literal", - "fields": { - "type": "commit", - "hash": "9746771f7457a9a219fff46351270ecb0a28d15b", - "author": "Your Name", - "email": "you@example.com", - "commit_type": "other" - } - }, - "graph": { - "relationships": [{"predicate":"modifies","target":"file:docs/vcl-grammar.ebnf"}] - }, - "vector": { - "embedding": [0.484375,-0.632812,0.320312,0.679687,-0.578125,-0.906250,0.429687,-0.093750,0.156250,0.179687,0.585937,-0.578125,-0.195312,0.937500,0.859375,-0.656250,-0.968750,0.968750,-0.078125,0.125000,0.835937,0.750000,-0.835937,-0.515625,-0.164062,0.421875,0.492187,0.070312,0.984375,0.546875,0.609375,0.351562,0.484375,-0.632812,0.320312,0.679687,-0.578125,-0.906250,0.429687,-0.093750,0.156250,0.179687,0.585937,-0.578125,-0.195312,0.937500,0.859375,-0.656250,-0.968750,0.968750,-0.078125,0.125000,0.835937,0.750000,-0.835937,-0.515625,-0.164062,0.421875,0.492187,0.070312,0.984375,0.546875,0.609375,0.351562], - "model": "sha256-hash-64d" - }, - "tensor": { - "shape": [1, 3], - "data": [2.0, 2.0, 1.0] - }, - "semantic": { - "types": ["https://verisim.db/self/type/Commit", "https://verisim.db/self/type/other"], - "properties": { - "conventional_commit_type": "other", - "files_changed": "1" - } - }, - "temporal": { - "timestamp": "2026-01-22T11:57:33Z", - "version": 1, - "author": "Your Name" - } -} diff --git a/verisimdb/.verisimdb/octads/commit-980d6c7a0dc2.json b/verisimdb/.verisimdb/octads/commit-980d6c7a0dc2.json deleted file mode 100644 index 8b402a69..00000000 --- a/verisimdb/.verisimdb/octads/commit-980d6c7a0dc2.json +++ /dev/null @@ -1,39 +0,0 @@ -{ - "id": "commit-980d6c7a0dc2", - "source": "git-log", - "created_at": "2026-01-22T01:39:24Z", - "document": { - "title": "Create Kraft cross-integration to VeriSim.csv", - "body": "Commit 980d6c7a by Jonathan D.A. Jewell: Create Kraft cross-integration to VeriSim.csv", - "fields": { - "type": "commit", - "hash": "980d6c7a0dc2a0f6ef828f907348603247088b48", - "author": "Jonathan D.A. Jewell", - "email": "6759885+hyperpolymath@users.noreply.github.com", - "commit_type": "other" - } - }, - "graph": { - "relationships": [{"predicate":"modifies","target":"file:Kraft cross-integration to VeriSim.csv"}] - }, - "vector": { - "embedding": [0.539062,-0.257812,-0.054687,-0.445312,-0.828125,-0.078125,-0.078125,0.765625,-0.695312,0.843750,0.867187,0.164062,-0.039062,-0.984375,0,0.812500,-0.484375,-0.343750,0.875000,0.304687,-0.132812,0.312500,0.757812,0.093750,0.742187,-0.125000,-0.921875,0.976562,-0.156250,-0.609375,-0.851562,-0.710937,0.539062,-0.257812,-0.054687,-0.445312,-0.828125,-0.078125,-0.078125,0.765625,-0.695312,0.843750,0.867187,0.164062,-0.039062,-0.984375,0,0.812500,-0.484375,-0.343750,0.875000,0.304687,-0.132812,0.312500,0.757812,0.093750,0.742187,-0.125000,-0.921875,0.976562,-0.156250,-0.609375,-0.851562,-0.710937], - "model": "sha256-hash-64d" - }, - "tensor": { - "shape": [1, 3], - "data": [6.0, 0.0, 1.0] - }, - "semantic": { - "types": ["https://verisim.db/self/type/Commit", "https://verisim.db/self/type/other"], - "properties": { - "conventional_commit_type": "other", - "files_changed": "1" - } - }, - "temporal": { - "timestamp": "2026-01-22T01:39:24Z", - "version": 1, - "author": "Jonathan D.A. Jewell" - } -} diff --git a/verisimdb/.verisimdb/octads/commit-9cce50aa9df3.json b/verisimdb/.verisimdb/octads/commit-9cce50aa9df3.json deleted file mode 100644 index eddcce12..00000000 --- a/verisimdb/.verisimdb/octads/commit-9cce50aa9df3.json +++ /dev/null @@ -1,39 +0,0 @@ -{ - "id": "commit-9cce50aa9df3", - "source": "git-log", - "created_at": "2026-02-04T21:54:56Z", - "document": { - "title": "feat: complete ReScript registry, benchmarks, and deployment guide", - "body": "Commit 9cce50aa by Jonathan D.A. Jewell: feat: complete ReScript registry, benchmarks, and deployment guide", - "fields": { - "type": "commit", - "hash": "9cce50aa9df3617f1ad5e2ae55f2580a29ebf855", - "author": "Jonathan D.A. Jewell", - "email": "jonathan.jewell@open.ac.uk", - "commit_type": "feature" - } - }, - "graph": { - "relationships": [{"predicate":"modifies","target":"file:.machine_readable/STATE.scm"},{"predicate":"modifies","target":"file:DEPLOYMENT.adoc"},{"predicate":"modifies","target":"file:benches/modality_benchmarks.rs"},{"predicate":"modifies","target":"file:src/registry/MetadataLog.res"},{"predicate":"modifies","target":"file:src/registry/Registry.res"}] - }, - "vector": { - "embedding": [-0.093750,-0.898437,0.210937,0.664062,-0.226562,-0.859375,-0.546875,0.046875,-0.960937,-0.070312,0.406250,0.359375,0.242187,0.265625,0.445312,-0.390625,-0.992187,-0.835937,-0.601562,-0.203125,-0.921875,-0.117187,-0.039062,-0.500000,-0.132812,0.382812,0.101562,-0.781250,-0.320312,-0.476562,-0.546875,0.804687,-0.093750,-0.898437,0.210937,0.664062,-0.226562,-0.859375,-0.546875,0.046875,-0.960937,-0.070312,0.406250,0.359375,0.242187,0.265625,0.445312,-0.390625,-0.992187,-0.835937,-0.601562,-0.203125,-0.921875,-0.117187,-0.039062,-0.500000,-0.132812,0.382812,0.101562,-0.781250,-0.320312,-0.476562,-0.546875,0.804687], - "model": "sha256-hash-64d" - }, - "tensor": { - "shape": [1, 3], - "data": [1950.0, 13.0, 5.0] - }, - "semantic": { - "types": ["https://verisim.db/self/type/Commit", "https://verisim.db/self/type/feature"], - "properties": { - "conventional_commit_type": "feature", - "files_changed": "5" - } - }, - "temporal": { - "timestamp": "2026-02-04T21:54:56Z", - "version": 1, - "author": "Jonathan D.A. Jewell" - } -} diff --git a/verisimdb/.verisimdb/octads/commit-9cdf85099304.json b/verisimdb/.verisimdb/octads/commit-9cdf85099304.json deleted file mode 100644 index 9784da13..00000000 --- a/verisimdb/.verisimdb/octads/commit-9cdf85099304.json +++ /dev/null @@ -1,39 +0,0 @@ -{ - "id": "commit-9cdf85099304", - "source": "git-log", - "created_at": "2026-01-22T13:06:13Z", - "document": { - "title": "Add consultation papers and complete .machine_readable files", - "body": "Commit 9cdf8509 by Your Name: Add consultation papers and complete .machine_readable files", - "fields": { - "type": "commit", - "hash": "9cdf850993046aaa48e1ecee80165121b03fcf89", - "author": "Your Name", - "email": "you@example.com", - "commit_type": "other" - } - }, - "graph": { - "relationships": [{"predicate":"modifies","target":"file:.machine_readable/AGENTIC.scm"},{"predicate":"modifies","target":"file:.machine_readable/ECOSYSTEM.scm"},{"predicate":"modifies","target":"file:.machine_readable/META.scm"},{"predicate":"modifies","target":"file:.machine_readable/NEUROSYM.scm"},{"predicate":"modifies","target":"file:.machine_readable/PLAYBOOK.scm"},{"predicate":"modifies","target":"file:.machine_readable/STATE.scm"},{"predicate":"modifies","target":"file:AI.djot"},{"predicate":"modifies","target":"file:docs/consultation-dependent-types-zkp.adoc"},{"predicate":"modifies","target":"file:docs/consultation-normalization-strategy.adoc"}] - }, - "vector": { - "embedding": [-0.429687,0.828125,0.570312,0.875000,0.093750,-0.812500,0.242187,0.390625,0.546875,-0.585937,-0.039062,0.609375,-0.710937,-0.414062,-0.359375,-0.578125,0.148437,0.187500,0.914062,-0.796875,0.234375,0.054687,0.656250,-0.679687,0.125000,-0.898437,0.664062,0.460937,0.828125,-0.812500,0.593750,-0.289062,-0.429687,0.828125,0.570312,0.875000,0.093750,-0.812500,0.242187,0.390625,0.546875,-0.585937,-0.039062,0.609375,-0.710937,-0.414062,-0.359375,-0.578125,0.148437,0.187500,0.914062,-0.796875,0.234375,0.054687,0.656250,-0.679687,0.125000,-0.898437,0.664062,0.460937,0.828125,-0.812500,0.593750,-0.289062], - "model": "sha256-hash-64d" - }, - "tensor": { - "shape": [1, 3], - "data": [3919.0, 248.0, 9.0] - }, - "semantic": { - "types": ["https://verisim.db/self/type/Commit", "https://verisim.db/self/type/other"], - "properties": { - "conventional_commit_type": "other", - "files_changed": "9" - } - }, - "temporal": { - "timestamp": "2026-01-22T13:06:13Z", - "version": 1, - "author": "Your Name" - } -} diff --git a/verisimdb/.verisimdb/octads/commit-9d353c5546f5.json b/verisimdb/.verisimdb/octads/commit-9d353c5546f5.json deleted file mode 100644 index eb02d1ed..00000000 --- a/verisimdb/.verisimdb/octads/commit-9d353c5546f5.json +++ /dev/null @@ -1,39 +0,0 @@ -{ - "id": "commit-9d353c5546f5", - "source": "git-log", - "created_at": "2026-01-22T12:08:06Z", - "document": { - "title": "Add VCL examples for VERSION and DRIFT features", - "body": "Commit 9d353c55 by Your Name: Add VCL examples for VERSION and DRIFT features", - "fields": { - "type": "commit", - "hash": "9d353c5546f5fe55859cae8bb4af60dc9b062f67", - "author": "Your Name", - "email": "you@example.com", - "commit_type": "other" - } - }, - "graph": { - "relationships": [{"predicate":"modifies","target":"file:README.adoc"},{"predicate":"modifies","target":"file:docs/vcl-examples.adoc"}] - }, - "vector": { - "embedding": [-0.781250,0.835937,0.734375,0.820312,0.851562,0.531250,0.414062,-0.710937,-0.742187,-0.359375,-0.710937,0.859375,0.992187,-0.210937,0.593750,-0.054687,0.679687,0.164062,-0.898437,-0.796875,-0.546875,-0.140625,0.890625,0.703125,-0.765625,-0.632812,0.953125,0.929687,-0.250000,0.828125,-0.250000,-0.218750,-0.781250,0.835937,0.734375,0.820312,0.851562,0.531250,0.414062,-0.710937,-0.742187,-0.359375,-0.710937,0.859375,0.992187,-0.210937,0.593750,-0.054687,0.679687,0.164062,-0.898437,-0.796875,-0.546875,-0.140625,0.890625,0.703125,-0.765625,-0.632812,0.953125,0.929687,-0.250000,0.828125,-0.250000,-0.218750], - "model": "sha256-hash-64d" - }, - "tensor": { - "shape": [1, 3], - "data": [108.0, 2.0, 2.0] - }, - "semantic": { - "types": ["https://verisim.db/self/type/Commit", "https://verisim.db/self/type/other"], - "properties": { - "conventional_commit_type": "other", - "files_changed": "2" - } - }, - "temporal": { - "timestamp": "2026-01-22T12:08:06Z", - "version": 1, - "author": "Your Name" - } -} diff --git a/verisimdb/.verisimdb/octads/commit-9e2298460f12.json b/verisimdb/.verisimdb/octads/commit-9e2298460f12.json deleted file mode 100644 index 9b293429..00000000 --- a/verisimdb/.verisimdb/octads/commit-9e2298460f12.json +++ /dev/null @@ -1,39 +0,0 @@ -{ - "id": "commit-9e2298460f12", - "source": "git-log", - "created_at": "2026-01-22T00:21:24Z", - "document": { - "title": "Update README.adoc", - "body": "Commit 9e229846 by Jonathan D.A. Jewell: Update README.adoc", - "fields": { - "type": "commit", - "hash": "9e2298460f125d11bb20160108e8f172ddf35ffd", - "author": "Jonathan D.A. Jewell", - "email": "6759885+hyperpolymath@users.noreply.github.com", - "commit_type": "other" - } - }, - "graph": { - "relationships": [{"predicate":"modifies","target":"file:README.adoc"}] - }, - "vector": { - "embedding": [-0.656250,0.289062,0.242187,-0.078125,0.179687,-0.445312,0.343750,-0.132812,-0.171875,-0.273437,0.687500,0.031250,-0.101562,0.882812,0.437500,0.390625,0.859375,-0.906250,-0.617187,-0.492187,0.312500,0.320312,-0.867187,-0.976562,0.531250,0.890625,0.835937,0.226562,0.070312,0.187500,0.976562,-0.078125,-0.656250,0.289062,0.242187,-0.078125,0.179687,-0.445312,0.343750,-0.132812,-0.171875,-0.273437,0.687500,0.031250,-0.101562,0.882812,0.437500,0.390625,0.859375,-0.906250,-0.617187,-0.492187,0.312500,0.320312,-0.867187,-0.976562,0.531250,0.890625,0.835937,0.226562,0.070312,0.187500,0.976562,-0.078125], - "model": "sha256-hash-64d" - }, - "tensor": { - "shape": [1, 3], - "data": [1.0, 0.0, 1.0] - }, - "semantic": { - "types": ["https://verisim.db/self/type/Commit", "https://verisim.db/self/type/other"], - "properties": { - "conventional_commit_type": "other", - "files_changed": "1" - } - }, - "temporal": { - "timestamp": "2026-01-22T00:21:24Z", - "version": 1, - "author": "Jonathan D.A. Jewell" - } -} diff --git a/verisimdb/.verisimdb/octads/commit-9f5c2f37c3f6.json b/verisimdb/.verisimdb/octads/commit-9f5c2f37c3f6.json deleted file mode 100644 index 295262ac..00000000 --- a/verisimdb/.verisimdb/octads/commit-9f5c2f37c3f6.json +++ /dev/null @@ -1,39 +0,0 @@ -{ - "id": "commit-9f5c2f37c3f6", - "source": "git-log", - "created_at": "2026-01-22T11:12:19Z", - "document": { - "title": "Add miniKanren integration as v3 stub", - "body": "Commit 9f5c2f37 by Your Name: Add miniKanren integration as v3 stub", - "fields": { - "type": "commit", - "hash": "9f5c2f37c3f653f2d6621b7dcffba2cd3d685639", - "author": "Your Name", - "email": "you@example.com", - "commit_type": "other" - } - }, - "graph": { - "relationships": [{"predicate":"modifies","target":"file:README.adoc"},{"predicate":"modifies","target":"file:docs/minikanren-integration-v3.adoc"}] - }, - "vector": { - "embedding": [-0.351562,-0.859375,0.734375,0.390625,-0.671875,0.609375,0,-0.750000,-0.585937,-0.429687,-0.898437,-0.656250,0.281250,-0.562500,-0.742187,-0.125000,-0.765625,0.617187,0.843750,0.617187,0.226562,-0.023437,-0.929687,0.476562,-0.132812,0.828125,-0.226562,0.375000,-0.546875,-0.578125,-0.664062,0.953125,-0.351562,-0.859375,0.734375,0.390625,-0.671875,0.609375,0,-0.750000,-0.585937,-0.429687,-0.898437,-0.656250,0.281250,-0.562500,-0.742187,-0.125000,-0.765625,0.617187,0.843750,0.617187,0.226562,-0.023437,-0.929687,0.476562,-0.132812,0.828125,-0.226562,0.375000,-0.546875,-0.578125,-0.664062,0.953125], - "model": "sha256-hash-64d" - }, - "tensor": { - "shape": [1, 3], - "data": [338.0, 4.0, 2.0] - }, - "semantic": { - "types": ["https://verisim.db/self/type/Commit", "https://verisim.db/self/type/other"], - "properties": { - "conventional_commit_type": "other", - "files_changed": "2" - } - }, - "temporal": { - "timestamp": "2026-01-22T11:12:19Z", - "version": 1, - "author": "Your Name" - } -} diff --git a/verisimdb/.verisimdb/octads/commit-a0d832b5068b.json b/verisimdb/.verisimdb/octads/commit-a0d832b5068b.json deleted file mode 100644 index d866797d..00000000 --- a/verisimdb/.verisimdb/octads/commit-a0d832b5068b.json +++ /dev/null @@ -1,39 +0,0 @@ -{ - "id": "commit-a0d832b5068b", - "source": "git-log", - "created_at": "2026-01-22T01:36:17Z", - "document": { - "title": "Create Snapshotting and Truncation Logic.md", - "body": "Commit a0d832b5 by Jonathan D.A. Jewell: Create Snapshotting and Truncation Logic.md", - "fields": { - "type": "commit", - "hash": "a0d832b5068b57ac46a8ffb2f4555c6476be8516", - "author": "Jonathan D.A. Jewell", - "email": "6759885+hyperpolymath@users.noreply.github.com", - "commit_type": "other" - } - }, - "graph": { - "relationships": [{"predicate":"modifies","target":"file:Snapshotting and Truncation Logic.md"}] - }, - "vector": { - "embedding": [-0.265625,0.101562,0.523437,0.218750,-0.531250,-0.007812,0.554687,-0.914062,-0.062500,0.804687,-0.898437,0.375000,0.273437,-0.906250,0.429687,0.632812,-0.789062,-0.390625,0.414062,0.148437,0.140625,0.171875,0.679687,-0.117187,0.703125,0.117187,-0.859375,0.515625,0.546875,-0.703125,0.070312,-0.437500,-0.265625,0.101562,0.523437,0.218750,-0.531250,-0.007812,0.554687,-0.914062,-0.062500,0.804687,-0.898437,0.375000,0.273437,-0.906250,0.429687,0.632812,-0.789062,-0.390625,0.414062,0.148437,0.140625,0.171875,0.679687,-0.117187,0.703125,0.117187,-0.859375,0.515625,0.546875,-0.703125,0.070312,-0.437500], - "model": "sha256-hash-64d" - }, - "tensor": { - "shape": [1, 3], - "data": [72.0, 0.0, 1.0] - }, - "semantic": { - "types": ["https://verisim.db/self/type/Commit", "https://verisim.db/self/type/other"], - "properties": { - "conventional_commit_type": "other", - "files_changed": "1" - } - }, - "temporal": { - "timestamp": "2026-01-22T01:36:17Z", - "version": 1, - "author": "Jonathan D.A. Jewell" - } -} diff --git a/verisimdb/.verisimdb/octads/commit-a31e6e33e51f.json b/verisimdb/.verisimdb/octads/commit-a31e6e33e51f.json deleted file mode 100644 index c1eb7781..00000000 --- a/verisimdb/.verisimdb/octads/commit-a31e6e33e51f.json +++ /dev/null @@ -1,39 +0,0 @@ -{ - "id": "commit-a31e6e33e51f", - "source": "git-log", - "created_at": "2026-01-17T02:19:22Z", - "document": { - "title": "chore: sync template files and configuration", - "body": "Commit a31e6e33 by Jonathan D.A. Jewell: chore: sync template files and configuration", - "fields": { - "type": "commit", - "hash": "a31e6e33e51f166e82890ba85a9f52da4dfb13a0", - "author": "Jonathan D.A. Jewell", - "email": "jonathan.jewell@open.ac.uk", - "commit_type": "chore" - } - }, - "graph": { - "relationships": [{"predicate":"modifies","target":"file:rust-core/verisim-octad/tests/integration_tests.rs"},{"predicate":"modifies","target":"file:rust-core/verisim-vector/src/hnsw.rs"}] - }, - "vector": { - "embedding": [0.101562,0.898437,-0.367187,-0.203125,0.718750,0.656250,0.703125,-0.921875,0.140625,-0.695312,-0.234375,-0.734375,0.640625,0.546875,-0.640625,-0.460937,0.210937,0.781250,0.648437,0.890625,0.820312,-0.648437,-0.742187,0.007812,0.625000,-0.015625,0.687500,-0.445312,0.187500,0.859375,-0.367187,0.101562,0.101562,0.898437,-0.367187,-0.203125,0.718750,0.656250,0.703125,-0.921875,0.140625,-0.695312,-0.234375,-0.734375,0.640625,0.546875,-0.640625,-0.460937,0.210937,0.781250,0.648437,0.890625,0.820312,-0.648437,-0.742187,0.007812,0.625000,-0.015625,0.687500,-0.445312,0.187500,0.859375,-0.367187,0.101562], - "model": "sha256-hash-64d" - }, - "tensor": { - "shape": [1, 3], - "data": [920.0, 0.0, 2.0] - }, - "semantic": { - "types": ["https://verisim.db/self/type/Commit", "https://verisim.db/self/type/chore"], - "properties": { - "conventional_commit_type": "chore", - "files_changed": "2" - } - }, - "temporal": { - "timestamp": "2026-01-17T02:19:22Z", - "version": 1, - "author": "Jonathan D.A. Jewell" - } -} diff --git a/verisimdb/.verisimdb/octads/commit-a782219c1f2a.json b/verisimdb/.verisimdb/octads/commit-a782219c1f2a.json deleted file mode 100644 index ce637172..00000000 --- a/verisimdb/.verisimdb/octads/commit-a782219c1f2a.json +++ /dev/null @@ -1,39 +0,0 @@ -{ - "id": "commit-a782219c1f2a", - "source": "git-log", - "created_at": "2026-02-08T14:41:48Z", - "document": { - "title": "docs: add comprehensive Sonnet task list for GitHub CI integration", - "body": "Commit a782219c by Jonathan D.A. Jewell: docs: add comprehensive Sonnet task list for GitHub CI integration", - "fields": { - "type": "commit", - "hash": "a782219c1f2a1c554b4837ed54fac6524b782b87", - "author": "Jonathan D.A. Jewell", - "email": "jonathan.jewell@open.ac.uk", - "commit_type": "documentation" - } - }, - "graph": { - "relationships": [{"predicate":"modifies","target":"file:SONNET-TASKS.md"}] - }, - "vector": { - "embedding": [0.265625,0.632812,0.515625,-0.492187,0.140625,-0.125000,-0.500000,0.109375,0.625000,0.523437,-0.359375,-0.515625,0.140625,-0.109375,0.289062,0.945312,0.859375,-0.585937,0.585937,0.078125,0.804687,0.437500,0.914062,0.390625,-0.132812,0.851562,-0.078125,-0.546875,-0.914062,-0.828125,-0.789062,-0.484375,0.265625,0.632812,0.515625,-0.492187,0.140625,-0.125000,-0.500000,0.109375,0.625000,0.523437,-0.359375,-0.515625,0.140625,-0.109375,0.289062,0.945312,0.859375,-0.585937,0.585937,0.078125,0.804687,0.437500,0.914062,0.390625,-0.132812,0.851562,-0.078125,-0.546875,-0.914062,-0.828125,-0.789062,-0.484375], - "model": "sha256-hash-64d" - }, - "tensor": { - "shape": [1, 3], - "data": [411.0, 0.0, 1.0] - }, - "semantic": { - "types": ["https://verisim.db/self/type/Commit", "https://verisim.db/self/type/documentation"], - "properties": { - "conventional_commit_type": "documentation", - "files_changed": "1" - } - }, - "temporal": { - "timestamp": "2026-02-08T14:41:48Z", - "version": 1, - "author": "Jonathan D.A. Jewell" - } -} diff --git a/verisimdb/.verisimdb/octads/commit-a9af6511d111.json b/verisimdb/.verisimdb/octads/commit-a9af6511d111.json deleted file mode 100644 index c3555a26..00000000 --- a/verisimdb/.verisimdb/octads/commit-a9af6511d111.json +++ /dev/null @@ -1,39 +0,0 @@ -{ - "id": "commit-a9af6511d111", - "source": "git-log", - "created_at": "2026-01-16T19:11:17Z", - "document": { - "title": "chore: comment out unused hnsw_rs dependency", - "body": "Commit a9af6511 by Jonathan D.A. Jewell: chore: comment out unused hnsw_rs dependency", - "fields": { - "type": "commit", - "hash": "a9af6511d11142d849fb10332cbbd1632ef5f305", - "author": "Jonathan D.A. Jewell", - "email": "jonathan.jewell@open.ac.uk", - "commit_type": "chore" - } - }, - "graph": { - "relationships": [{"predicate":"modifies","target":"file:rust-core/verisim-vector/Cargo.toml"}] - }, - "vector": { - "embedding": [-0.476562,-0.007812,0.039062,-0.875000,-0.101562,0.375000,0.046875,0.710937,0.093750,-0.046875,0.218750,-0.507812,-0.523437,0.429687,-0.531250,0.828125,-0.593750,-0.320312,-0.218750,0.468750,-0.304687,-0.976562,-0.523437,0.710937,-0.804687,0.867187,-0.148437,0.664062,-0.023437,0.757812,0.210937,-0.359375,-0.476562,-0.007812,0.039062,-0.875000,-0.101562,0.375000,0.046875,0.710937,0.093750,-0.046875,0.218750,-0.507812,-0.523437,0.429687,-0.531250,0.828125,-0.593750,-0.320312,-0.218750,0.468750,-0.304687,-0.976562,-0.523437,0.710937,-0.804687,0.867187,-0.148437,0.664062,-0.023437,0.757812,0.210937,-0.359375], - "model": "sha256-hash-64d" - }, - "tensor": { - "shape": [1, 3], - "data": [1.0, 1.0, 1.0] - }, - "semantic": { - "types": ["https://verisim.db/self/type/Commit", "https://verisim.db/self/type/chore"], - "properties": { - "conventional_commit_type": "chore", - "files_changed": "1" - } - }, - "temporal": { - "timestamp": "2026-01-16T19:11:17Z", - "version": 1, - "author": "Jonathan D.A. Jewell" - } -} diff --git a/verisimdb/.verisimdb/octads/commit-aa63510d99bd.json b/verisimdb/.verisimdb/octads/commit-aa63510d99bd.json deleted file mode 100644 index c5ca69af..00000000 --- a/verisimdb/.verisimdb/octads/commit-aa63510d99bd.json +++ /dev/null @@ -1,39 +0,0 @@ -{ - "id": "commit-aa63510d99bd", - "source": "git-log", - "created_at": "2026-01-22T11:28:24Z", - "document": { - "title": "Add commas before optional elements in sequences", - "body": "Commit aa63510d by Your Name: Add commas before optional elements in sequences", - "fields": { - "type": "commit", - "hash": "aa63510d99bd499a37aa83072a24a81d6497fa92", - "author": "Your Name", - "email": "you@example.com", - "commit_type": "other" - } - }, - "graph": { - "relationships": [{"predicate":"modifies","target":"file:docs/vcl-grammar.ebnf"}] - }, - "vector": { - "embedding": [0.007812,-0.132812,0.257812,0.789062,-0.007812,-0.132812,-0.890625,0.164062,0.312500,-0.046875,0.296875,0.328125,0.726562,-0.609375,0.312500,0.054687,-0.218750,0.679687,0.687500,-0.453125,0.523437,-0.859375,-0.875000,-0.500000,-0.382812,0.281250,0.335937,0.468750,0.351562,-0.914062,-0.796875,0.375000,0.007812,-0.132812,0.257812,0.789062,-0.007812,-0.132812,-0.890625,0.164062,0.312500,-0.046875,0.296875,0.328125,0.726562,-0.609375,0.312500,0.054687,-0.218750,0.679687,0.687500,-0.453125,0.523437,-0.859375,-0.875000,-0.500000,-0.382812,0.281250,0.335937,0.468750,0.351562,-0.914062,-0.796875,0.375000], - "model": "sha256-hash-64d" - }, - "tensor": { - "shape": [1, 3], - "data": [6.0, 6.0, 1.0] - }, - "semantic": { - "types": ["https://verisim.db/self/type/Commit", "https://verisim.db/self/type/other"], - "properties": { - "conventional_commit_type": "other", - "files_changed": "1" - } - }, - "temporal": { - "timestamp": "2026-01-22T11:28:24Z", - "version": 1, - "author": "Your Name" - } -} diff --git a/verisimdb/.verisimdb/octads/commit-b1038bf033e2.json b/verisimdb/.verisimdb/octads/commit-b1038bf033e2.json deleted file mode 100644 index ffbaec14..00000000 --- a/verisimdb/.verisimdb/octads/commit-b1038bf033e2.json +++ /dev/null @@ -1,39 +0,0 @@ -{ - "id": "commit-b1038bf033e2", - "source": "git-log", - "created_at": "2026-01-22T01:33:30Z", - "document": { - "title": "Create proven-coherence.md", - "body": "Commit b1038bf0 by Jonathan D.A. Jewell: Create proven-coherence.md", - "fields": { - "type": "commit", - "hash": "b1038bf033e22cb8d4802989869afb1dd563d48a", - "author": "Jonathan D.A. Jewell", - "email": "6759885+hyperpolymath@users.noreply.github.com", - "commit_type": "other" - } - }, - "graph": { - "relationships": [{"predicate":"modifies","target":"file:proven-coherence.md"}] - }, - "vector": { - "embedding": [0.906250,-0.710937,0.820312,0.773437,-0.937500,-0.929687,-0.039062,-0.414062,0.382812,0.632812,0.984375,-0.343750,0.500000,0.953125,-0.656250,-0.453125,-0.171875,0.062500,0.101562,-0.625000,-0.078125,-0.382812,-0.695312,0.484375,0.320312,-0.945312,-0.867187,-0.593750,0.742187,0.875000,-0.265625,-0.625000,0.906250,-0.710937,0.820312,0.773437,-0.937500,-0.929687,-0.039062,-0.414062,0.382812,0.632812,0.984375,-0.343750,0.500000,0.953125,-0.656250,-0.453125,-0.171875,0.062500,0.101562,-0.625000,-0.078125,-0.382812,-0.695312,0.484375,0.320312,-0.945312,-0.867187,-0.593750,0.742187,0.875000,-0.265625,-0.625000], - "model": "sha256-hash-64d" - }, - "tensor": { - "shape": [1, 3], - "data": [258.0, 0.0, 1.0] - }, - "semantic": { - "types": ["https://verisim.db/self/type/Commit", "https://verisim.db/self/type/other"], - "properties": { - "conventional_commit_type": "other", - "files_changed": "1" - } - }, - "temporal": { - "timestamp": "2026-01-22T01:33:30Z", - "version": 1, - "author": "Jonathan D.A. Jewell" - } -} diff --git a/verisimdb/.verisimdb/octads/commit-b62522fca6dc.json b/verisimdb/.verisimdb/octads/commit-b62522fca6dc.json deleted file mode 100644 index b3c584a2..00000000 --- a/verisimdb/.verisimdb/octads/commit-b62522fca6dc.json +++ /dev/null @@ -1,39 +0,0 @@ -{ - "id": "commit-b62522fca6dc", - "source": "git-log", - "created_at": "2026-02-12T22:00:53Z", - "document": { - "title": "docs: add session entry for security-scan.yml workflow automation", - "body": "Commit b62522fc by Jonathan D.A. Jewell: docs: add session entry for security-scan.yml workflow automation", - "fields": { - "type": "commit", - "hash": "b62522fca6dc3a6364d29fa351983466fba59115", - "author": "Jonathan D.A. Jewell", - "email": "jonathan.jewell@open.ac.uk", - "commit_type": "documentation" - } - }, - "graph": { - "relationships": [{"predicate":"modifies","target":"file:.machine_readable/STATE.scm"}] - }, - "vector": { - "embedding": [-0.273437,0.445312,0.414062,0.578125,0.953125,-0.343750,0.773437,0.929687,-0.078125,-0.601562,0.242187,-0.117187,0.914062,-0.523437,-0.023437,0.578125,-0.328125,0.718750,0.843750,0.796875,0.976562,-0.085937,-0.687500,0.023437,-0.429687,0.843750,-0.609375,-0.429687,-0.210937,-0.156250,0.203125,0.703125,-0.273437,0.445312,0.414062,0.578125,0.953125,-0.343750,0.773437,0.929687,-0.078125,-0.601562,0.242187,-0.117187,0.914062,-0.523437,-0.023437,0.578125,-0.328125,0.718750,0.843750,0.796875,0.976562,-0.085937,-0.687500,0.023437,-0.429687,0.843750,-0.609375,-0.429687,-0.210937,-0.156250,0.203125,0.703125], - "model": "sha256-hash-64d" - }, - "tensor": { - "shape": [1, 3], - "data": [12.0, 1.0, 1.0] - }, - "semantic": { - "types": ["https://verisim.db/self/type/Commit", "https://verisim.db/self/type/documentation"], - "properties": { - "conventional_commit_type": "documentation", - "files_changed": "1" - } - }, - "temporal": { - "timestamp": "2026-02-12T22:00:53Z", - "version": 1, - "author": "Jonathan D.A. Jewell" - } -} diff --git a/verisimdb/.verisimdb/octads/commit-b9d4aae8f5ce.json b/verisimdb/.verisimdb/octads/commit-b9d4aae8f5ce.json deleted file mode 100644 index da9ee781..00000000 --- a/verisimdb/.verisimdb/octads/commit-b9d4aae8f5ce.json +++ /dev/null @@ -1,39 +0,0 @@ -{ - "id": "commit-b9d4aae8f5ce", - "source": "git-log", - "created_at": "2026-01-21T02:22:34Z", - "document": { - "title": "Update README.adoc", - "body": "Commit b9d4aae8 by Jonathan D.A. Jewell: Update README.adoc", - "fields": { - "type": "commit", - "hash": "b9d4aae8f5ceffc7e60fa45a78b0fafa34ec7c31", - "author": "Jonathan D.A. Jewell", - "email": "6759885+hyperpolymath@users.noreply.github.com", - "commit_type": "other" - } - }, - "graph": { - "relationships": [{"predicate":"modifies","target":"file:README.adoc"}] - }, - "vector": { - "embedding": [-0.656250,0.289062,0.242187,-0.078125,0.179687,-0.445312,0.343750,-0.132812,-0.171875,-0.273437,0.687500,0.031250,-0.101562,0.882812,0.437500,0.390625,0.859375,-0.906250,-0.617187,-0.492187,0.312500,0.320312,-0.867187,-0.976562,0.531250,0.890625,0.835937,0.226562,0.070312,0.187500,0.976562,-0.078125,-0.656250,0.289062,0.242187,-0.078125,0.179687,-0.445312,0.343750,-0.132812,-0.171875,-0.273437,0.687500,0.031250,-0.101562,0.882812,0.437500,0.390625,0.859375,-0.906250,-0.617187,-0.492187,0.312500,0.320312,-0.867187,-0.976562,0.531250,0.890625,0.835937,0.226562,0.070312,0.187500,0.976562,-0.078125], - "model": "sha256-hash-64d" - }, - "tensor": { - "shape": [1, 3], - "data": [181.0, 9.0, 1.0] - }, - "semantic": { - "types": ["https://verisim.db/self/type/Commit", "https://verisim.db/self/type/other"], - "properties": { - "conventional_commit_type": "other", - "files_changed": "1" - } - }, - "temporal": { - "timestamp": "2026-01-21T02:22:34Z", - "version": 1, - "author": "Jonathan D.A. Jewell" - } -} diff --git a/verisimdb/.verisimdb/octads/commit-ba1f53543438.json b/verisimdb/.verisimdb/octads/commit-ba1f53543438.json deleted file mode 100644 index 3f603d5a..00000000 --- a/verisimdb/.verisimdb/octads/commit-ba1f53543438.json +++ /dev/null @@ -1,39 +0,0 @@ -{ - "id": "commit-ba1f53543438", - "source": "git-log", - "created_at": "2026-02-04T22:03:10Z", - "document": { - "title": "release: VeriSimDB v0.1.0-alpha", - "body": "Commit ba1f5354 by Jonathan D.A. Jewell: release: VeriSimDB v0.1.0-alpha", - "fields": { - "type": "commit", - "hash": "ba1f53543438c8150ad450bc77521dd809b1fc91", - "author": "Jonathan D.A. Jewell", - "email": "jonathan.jewell@open.ac.uk", - "commit_type": "other" - } - }, - "graph": { - "relationships": [{"predicate":"modifies","target":"file:RELEASE-NOTES-v0.1.0-alpha.md"},{"predicate":"modifies","target":"file:benches/Cargo.toml"},{"predicate":"modifies","target":"file:benchmark_results.txt"}] - }, - "vector": { - "embedding": [-0.445312,0.351562,-0.968750,0.046875,0.375000,0.179687,-0.507812,-0.031250,-0.289062,0.320312,-0.046875,0.484375,-0.710937,-0.960937,-0.992187,-0.718750,-0.031250,0.054687,-0.710937,0.421875,-0.843750,0.187500,-0.289062,0.460937,-0.750000,0.375000,0.359375,-0.593750,0.554687,-0.117187,-0.882812,-0.531250,-0.445312,0.351562,-0.968750,0.046875,0.375000,0.179687,-0.507812,-0.031250,-0.289062,0.320312,-0.046875,0.484375,-0.710937,-0.960937,-0.992187,-0.718750,-0.031250,0.054687,-0.710937,0.421875,-0.843750,0.187500,-0.289062,0.460937,-0.750000,0.375000,0.359375,-0.593750,0.554687,-0.117187,-0.882812,-0.531250], - "model": "sha256-hash-64d" - }, - "tensor": { - "shape": [1, 3], - "data": [513.0, 0.0, 3.0] - }, - "semantic": { - "types": ["https://verisim.db/self/type/Commit", "https://verisim.db/self/type/other"], - "properties": { - "conventional_commit_type": "other", - "files_changed": "3" - } - }, - "temporal": { - "timestamp": "2026-02-04T22:03:10Z", - "version": 1, - "author": "Jonathan D.A. Jewell" - } -} diff --git a/verisimdb/.verisimdb/octads/commit-bc2502d10f34.json b/verisimdb/.verisimdb/octads/commit-bc2502d10f34.json deleted file mode 100644 index 62458289..00000000 --- a/verisimdb/.verisimdb/octads/commit-bc2502d10f34.json +++ /dev/null @@ -1,39 +0,0 @@ -{ - "id": "commit-bc2502d10f34", - "source": "git-log", - "created_at": "2026-02-04T10:00:32Z", - "document": { - "title": "general update", - "body": "Commit bc2502d1 by Jonathan D.A. Jewell: general update", - "fields": { - "type": "commit", - "hash": "bc2502d10f3471faa43c9406a02bdb0bfb37a46e", - "author": "Jonathan D.A. Jewell", - "email": "jonathan.jewell@open.ac.uk", - "commit_type": "other" - } - }, - "graph": { - "relationships": [] - }, - "vector": { - "embedding": [-0.578125,0.156250,0.601562,-0.265625,-0.679687,-0.960937,-0.437500,0.679687,0.210937,0.414062,0.953125,-0.781250,-0.085937,-0.085937,0.117187,0.812500,0.281250,0.828125,0.609375,0.320312,-0.343750,0.023437,0.742187,-0.117187,0.445312,0.312500,-0.679687,0.757812,-0.343750,-0.625000,-0.359375,-0.437500,-0.578125,0.156250,0.601562,-0.265625,-0.679687,-0.960937,-0.437500,0.679687,0.210937,0.414062,0.953125,-0.781250,-0.085937,-0.085937,0.117187,0.812500,0.281250,0.828125,0.609375,0.320312,-0.343750,0.023437,0.742187,-0.117187,0.445312,0.312500,-0.679687,0.757812,-0.343750,-0.625000,-0.359375,-0.437500], - "model": "sha256-hash-64d" - }, - "tensor": { - "shape": [1, 3], - "data": [2671.0, 311.0, 25.0] - }, - "semantic": { - "types": ["https://verisim.db/self/type/Commit", "https://verisim.db/self/type/other"], - "properties": { - "conventional_commit_type": "other", - "files_changed": "25" - } - }, - "temporal": { - "timestamp": "2026-02-04T10:00:32Z", - "version": 1, - "author": "Jonathan D.A. Jewell" - } -} diff --git a/verisimdb/.verisimdb/octads/commit-bcbfd32da8b9.json b/verisimdb/.verisimdb/octads/commit-bcbfd32da8b9.json deleted file mode 100644 index ca700b46..00000000 --- a/verisimdb/.verisimdb/octads/commit-bcbfd32da8b9.json +++ /dev/null @@ -1,39 +0,0 @@ -{ - "id": "commit-bcbfd32da8b9", - "source": "git-log", - "created_at": "2026-02-12T16:00:44Z", - "document": { - "title": "fix: resolve 13 audit findings — honest completion, fix stubs, correct license headers", - "body": "Commit bcbfd32d by Jonathan D.A. Jewell: fix: resolve 13 audit findings — honest completion, fix stubs, correct license headers", - "fields": { - "type": "commit", - "hash": "bcbfd32da8b954c5b91e7968009906fc7cb7ffcc", - "author": "Jonathan D.A. Jewell", - "email": "jonathan.jewell@open.ac.uk", - "commit_type": "bugfix" - } - }, - "graph": { - "relationships": [] - }, - "vector": { - "embedding": [0.515625,0.015625,-0.726562,0.023437,-0.242187,0.421875,-0.390625,0.039062,0.914062,0.343750,0.640625,0.492187,0.476562,-0.179687,-0.015625,-0.585937,-0.890625,0.757812,-0.351562,-0.726562,0.304687,-0.445312,0.531250,-0.492187,0.593750,0.156250,0.468750,0.031250,0.429687,-0.015625,0.523437,-0.992187,0.515625,0.015625,-0.726562,0.023437,-0.242187,0.421875,-0.390625,0.039062,0.914062,0.343750,0.640625,0.492187,0.476562,-0.179687,-0.015625,-0.585937,-0.890625,0.757812,-0.351562,-0.726562,0.304687,-0.445312,0.531250,-0.492187,0.593750,0.156250,0.468750,0.031250,0.429687,-0.015625,0.523437,-0.992187], - "model": "sha256-hash-64d" - }, - "tensor": { - "shape": [1, 3], - "data": [1686.0, 472.0, 52.0] - }, - "semantic": { - "types": ["https://verisim.db/self/type/Commit", "https://verisim.db/self/type/bugfix"], - "properties": { - "conventional_commit_type": "bugfix", - "files_changed": "52" - } - }, - "temporal": { - "timestamp": "2026-02-12T16:00:44Z", - "version": 1, - "author": "Jonathan D.A. Jewell" - } -} diff --git a/verisimdb/.verisimdb/octads/commit-bcd5e7cd743f.json b/verisimdb/.verisimdb/octads/commit-bcd5e7cd743f.json deleted file mode 100644 index 5f503f75..00000000 --- a/verisimdb/.verisimdb/octads/commit-bcd5e7cd743f.json +++ /dev/null @@ -1,39 +0,0 @@ -{ - "id": "commit-bcd5e7cd743f", - "source": "git-log", - "created_at": "2026-01-22T12:51:59Z", - "document": { - "title": "Add project checkpoint files (STATE, META, ECOSYSTEM)", - "body": "Commit bcd5e7cd by Your Name: Add project checkpoint files (STATE, META, ECOSYSTEM)", - "fields": { - "type": "commit", - "hash": "bcd5e7cd743fba5e437adf231eec54fe8fd7d9d2", - "author": "Your Name", - "email": "you@example.com", - "commit_type": "other" - } - }, - "graph": { - "relationships": [{"predicate":"modifies","target":"file:ECOSYSTEM.scm"},{"predicate":"modifies","target":"file:META.scm"},{"predicate":"modifies","target":"file:STATE.scm"}] - }, - "vector": { - "embedding": [0.062500,-0.320312,-0.203125,-0.710937,0.640625,0.843750,0.054687,-0.312500,0.656250,-0.476562,-0.046875,-0.945312,0.695312,-0.679687,0.734375,-1.000000,0.023437,-0.906250,-0.226562,-0.429687,-0.226562,0.343750,-0.859375,0.007812,-0.078125,0.882812,-0.867187,0.226562,-0.023437,0.390625,0.414062,-0.273437,0.062500,-0.320312,-0.203125,-0.710937,0.640625,0.843750,0.054687,-0.312500,0.656250,-0.476562,-0.046875,-0.945312,0.695312,-0.679687,0.734375,-1.000000,0.023437,-0.906250,-0.226562,-0.429687,-0.226562,0.343750,-0.859375,0.007812,-0.078125,0.882812,-0.867187,0.226562,-0.023437,0.390625,0.414062,-0.273437], - "model": "sha256-hash-64d" - }, - "tensor": { - "shape": [1, 3], - "data": [721.0, 0.0, 3.0] - }, - "semantic": { - "types": ["https://verisim.db/self/type/Commit", "https://verisim.db/self/type/other"], - "properties": { - "conventional_commit_type": "other", - "files_changed": "3" - } - }, - "temporal": { - "timestamp": "2026-01-22T12:51:59Z", - "version": 1, - "author": "Your Name" - } -} diff --git a/verisimdb/.verisimdb/octads/commit-c0d8094e076f.json b/verisimdb/.verisimdb/octads/commit-c0d8094e076f.json deleted file mode 100644 index facd0b49..00000000 --- a/verisimdb/.verisimdb/octads/commit-c0d8094e076f.json +++ /dev/null @@ -1,39 +0,0 @@ -{ - "id": "commit-c0d8094e076f", - "source": "git-log", - "created_at": "2026-01-31T16:49:57Z", - "document": { - "title": "feat: add VoID (Vocabulary of Interlinked Datasets) metadata", - "body": "Commit c0d8094e by Test: feat: add VoID (Vocabulary of Interlinked Datasets) metadata", - "fields": { - "type": "commit", - "hash": "c0d8094e076fa595e12770cf9c92f76b902777ae", - "author": "Test", - "email": "test@example.com", - "commit_type": "feature" - } - }, - "graph": { - "relationships": [{"predicate":"modifies","target":"file:.well-known/void.rdf"},{"predicate":"modifies","target":"file:.well-known/void.ttl"},{"predicate":"modifies","target":"file:VOID-SETUP.md"}] - }, - "vector": { - "embedding": [0.531250,-0.468750,0.687500,-0.585937,0.421875,-0.453125,-0.273437,-0.390625,-0.828125,0.171875,-0.656250,0.156250,-0.640625,0.945312,0.710937,0.984375,0.968750,0.453125,0.406250,-0.726562,0.250000,0.304687,-0.890625,0.148437,0.054687,-0.859375,-0.414062,0.304687,-0.015625,0.250000,0.695312,-0.843750,0.531250,-0.468750,0.687500,-0.585937,0.421875,-0.453125,-0.273437,-0.390625,-0.828125,0.171875,-0.656250,0.156250,-0.640625,0.945312,0.710937,0.984375,0.968750,0.453125,0.406250,-0.726562,0.250000,0.304687,-0.890625,0.148437,0.054687,-0.859375,-0.414062,0.304687,-0.015625,0.250000,0.695312,-0.843750], - "model": "sha256-hash-64d" - }, - "tensor": { - "shape": [1, 3], - "data": [296.0, 0.0, 3.0] - }, - "semantic": { - "types": ["https://verisim.db/self/type/Commit", "https://verisim.db/self/type/feature"], - "properties": { - "conventional_commit_type": "feature", - "files_changed": "3" - } - }, - "temporal": { - "timestamp": "2026-01-31T16:49:57Z", - "version": 1, - "author": "Test" - } -} diff --git a/verisimdb/.verisimdb/octads/commit-ccb432c96a9d.json b/verisimdb/.verisimdb/octads/commit-ccb432c96a9d.json deleted file mode 100644 index a7f0d019..00000000 --- a/verisimdb/.verisimdb/octads/commit-ccb432c96a9d.json +++ /dev/null @@ -1,39 +0,0 @@ -{ - "id": "commit-ccb432c96a9d", - "source": "git-log", - "created_at": "2026-02-13T15:07:20Z", - "document": { - "title": "feat: add VCL AST to LogicalPlan bridge with 17 tests", - "body": "Commit ccb432c9 by Jonathan D.A. Jewell: feat: add VCL AST to LogicalPlan bridge with 17 tests", - "fields": { - "type": "commit", - "hash": "ccb432c96a9db67f85b80318483d73503afd325f", - "author": "Jonathan D.A. Jewell", - "email": "jonathan.jewell@open.ac.uk", - "commit_type": "feature" - } - }, - "graph": { - "relationships": [{"predicate":"modifies","target":"file:rust-core/verisim-api/src/graphql.rs"},{"predicate":"modifies","target":"file:rust-core/verisim-api/src/lib.rs"},{"predicate":"modifies","target":"file:rust-core/verisim-planner/src/lib.rs"},{"predicate":"modifies","target":"file:rust-core/verisim-planner/src/vcl_bridge.rs"}] - }, - "vector": { - "embedding": [-0.773437,-0.664062,-0.312500,-0.382812,-0.328125,0.367187,-0.460937,-0.773437,0.671875,-0.062500,0.554687,0.515625,-0.757812,0.265625,0.093750,0.125000,-0.781250,0.867187,0.242187,0.898437,-0.765625,0.171875,-0.156250,-0.843750,-0.015625,0.140625,-0.601562,-0.421875,-0.789062,-0.671875,0.148437,-0.125000,-0.773437,-0.664062,-0.312500,-0.382812,-0.328125,0.367187,-0.460937,-0.773437,0.671875,-0.062500,0.554687,0.515625,-0.757812,0.265625,0.093750,0.125000,-0.781250,0.867187,0.242187,0.898437,-0.765625,0.171875,-0.156250,-0.843750,-0.015625,0.140625,-0.601562,-0.421875,-0.789062,-0.671875,0.148437,-0.125000], - "model": "sha256-hash-64d" - }, - "tensor": { - "shape": [1, 3], - "data": [1803.0, 12.0, 4.0] - }, - "semantic": { - "types": ["https://verisim.db/self/type/Commit", "https://verisim.db/self/type/feature"], - "properties": { - "conventional_commit_type": "feature", - "files_changed": "4" - } - }, - "temporal": { - "timestamp": "2026-02-13T15:07:20Z", - "version": 1, - "author": "Jonathan D.A. Jewell" - } -} diff --git a/verisimdb/.verisimdb/octads/commit-cf2958acf17c.json b/verisimdb/.verisimdb/octads/commit-cf2958acf17c.json deleted file mode 100644 index a3522ced..00000000 --- a/verisimdb/.verisimdb/octads/commit-cf2958acf17c.json +++ /dev/null @@ -1,39 +0,0 @@ -{ - "id": "commit-cf2958acf17c", - "source": "git-log", - "created_at": "2026-01-22T10:32:26Z", - "document": { - "title": "feat(query): add query optimization with tuning, bidirectional propagation, EXPLAIN, and reversibility", - "body": "Commit cf2958ac by Your Name: feat(query): add query optimization with tuning, bidirectional propagation, EXPLAIN, and reversibility", - "fields": { - "type": "commit", - "hash": "cf2958acf17c57ff62af7ca67e90f1301c31a15d", - "author": "Your Name", - "email": "you@example.com", - "commit_type": "feature" - } - }, - "graph": { - "relationships": [{"predicate":"modifies","target":"file:README.adoc"},{"predicate":"modifies","target":"file:docs/query-optimization-overview.adoc"},{"predicate":"modifies","target":"file:docs/reversibility-design.adoc"},{"predicate":"modifies","target":"file:lib/verisim/query_planner_bidirectional.ex"},{"predicate":"modifies","target":"file:lib/verisim/query_planner_config.ex"},{"predicate":"modifies","target":"file:src/vcl/VCLExplain.res"}] - }, - "vector": { - "embedding": [-0.085937,0.328125,-0.531250,0.390625,0.148437,0.164062,-0.343750,-0.062500,0.015625,0.359375,-0.804687,-0.585937,-0.429687,-0.554687,0.468750,-0.468750,-0.039062,0.734375,0.617187,0.781250,-0.953125,0.890625,0.476562,0.398437,0.921875,-0.679687,0.148437,0.515625,0.375000,0.296875,0.648437,-0.031250,-0.085937,0.328125,-0.531250,0.390625,0.148437,0.164062,-0.343750,-0.062500,0.015625,0.359375,-0.804687,-0.585937,-0.429687,-0.554687,0.468750,-0.468750,-0.039062,0.734375,0.617187,0.781250,-0.953125,0.890625,0.476562,0.398437,0.921875,-0.679687,0.148437,0.515625,0.375000,0.296875,0.648437,-0.031250], - "model": "sha256-hash-64d" - }, - "tensor": { - "shape": [1, 3], - "data": [1674.0, 0.0, 6.0] - }, - "semantic": { - "types": ["https://verisim.db/self/type/Commit", "https://verisim.db/self/type/feature"], - "properties": { - "conventional_commit_type": "feature", - "files_changed": "6" - } - }, - "temporal": { - "timestamp": "2026-01-22T10:32:26Z", - "version": 1, - "author": "Your Name" - } -} diff --git a/verisimdb/.verisimdb/octads/commit-d39524c642db.json b/verisimdb/.verisimdb/octads/commit-d39524c642db.json deleted file mode 100644 index 5225bc28..00000000 --- a/verisimdb/.verisimdb/octads/commit-d39524c642db.json +++ /dev/null @@ -1,39 +0,0 @@ -{ - "id": "commit-d39524c642db", - "source": "git-log", - "created_at": "2026-02-13T16:01:07Z", - "document": { - "title": "feat: implement VCL REPL, WAL, auth, and normalizer regeneration", - "body": "Commit d39524c6 by Jonathan D.A. Jewell: feat: implement VCL REPL, WAL, auth, and normalizer regeneration", - "fields": { - "type": "commit", - "hash": "d39524c642db54a4463480aa0ef111a4da9f816e", - "author": "Jonathan D.A. Jewell", - "email": "jonathan.jewell@open.ac.uk", - "commit_type": "feature" - } - }, - "graph": { - "relationships": [] - }, - "vector": { - "embedding": [0.835937,0.835937,-0.679687,0.929687,0.328125,0.414062,-0.109375,-0.117187,0.375000,-0.734375,-0.515625,-0.648437,0.367187,0.921875,-0.960937,-0.601562,0.734375,-0.937500,0.687500,0.203125,0.210937,-0.656250,0.039062,-0.585937,-0.304687,-0.421875,-0.539062,0.367187,0.273437,-0.140625,0.429687,0.765625,0.835937,0.835937,-0.679687,0.929687,0.328125,0.414062,-0.109375,-0.117187,0.375000,-0.734375,-0.515625,-0.648437,0.367187,0.921875,-0.960937,-0.601562,0.734375,-0.937500,0.687500,0.203125,0.210937,-0.656250,0.039062,-0.585937,-0.304687,-0.421875,-0.539062,0.367187,0.273437,-0.140625,0.429687,0.765625], - "model": "sha256-hash-64d" - }, - "tensor": { - "shape": [1, 3], - "data": [5927.0, 0.0, 22.0] - }, - "semantic": { - "types": ["https://verisim.db/self/type/Commit", "https://verisim.db/self/type/feature"], - "properties": { - "conventional_commit_type": "feature", - "files_changed": "22" - } - }, - "temporal": { - "timestamp": "2026-02-13T16:01:07Z", - "version": 1, - "author": "Jonathan D.A. Jewell" - } -} diff --git a/verisimdb/.verisimdb/octads/commit-d4e6f6be1200.json b/verisimdb/.verisimdb/octads/commit-d4e6f6be1200.json deleted file mode 100644 index dbf694dc..00000000 --- a/verisimdb/.verisimdb/octads/commit-d4e6f6be1200.json +++ /dev/null @@ -1,39 +0,0 @@ -{ - "id": "commit-d4e6f6be1200", - "source": "git-log", - "created_at": "2026-02-13T14:18:54Z", - "document": { - "title": "feat: add verisim-planner crate with cost-based query planning and triple API scaffolding", - "body": "Commit d4e6f6be by Jonathan D.A. Jewell: feat: add verisim-planner crate with cost-based query planning and triple API scaffolding", - "fields": { - "type": "commit", - "hash": "d4e6f6be1200132c49caf1f285f300fa24e86c67", - "author": "Jonathan D.A. Jewell", - "email": "jonathan.jewell@open.ac.uk", - "commit_type": "feature" - } - }, - "graph": { - "relationships": [{"predicate":"modifies","target":"file:Cargo.lock"},{"predicate":"modifies","target":"file:Cargo.toml"},{"predicate":"modifies","target":"file:PLANNER-IMPLEMENTATION-STATUS.md"},{"predicate":"modifies","target":"file:rust-core/verisim-api/Cargo.toml"},{"predicate":"modifies","target":"file:rust-core/verisim-api/src/graphql.rs"},{"predicate":"modifies","target":"file:rust-core/verisim-api/src/lib.rs"},{"predicate":"modifies","target":"file:rust-core/verisim-planner/Cargo.toml"},{"predicate":"modifies","target":"file:rust-core/verisim-planner/src/config.rs"},{"predicate":"modifies","target":"file:rust-core/verisim-planner/src/cost.rs"},{"predicate":"modifies","target":"file:rust-core/verisim-planner/src/error.rs"},{"predicate":"modifies","target":"file:rust-core/verisim-planner/src/explain.rs"},{"predicate":"modifies","target":"file:rust-core/verisim-planner/src/lib.rs"},{"predicate":"modifies","target":"file:rust-core/verisim-planner/src/optimizer.rs"},{"predicate":"modifies","target":"file:rust-core/verisim-planner/src/plan.rs"},{"predicate":"modifies","target":"file:rust-core/verisim-planner/src/stats.rs"}] - }, - "vector": { - "embedding": [0.234375,-0.742187,0.164062,0.937500,-0.820312,-0.781250,0.289062,-0.031250,0.609375,0.101562,0.312500,0.992187,-0.437500,0.218750,-0.968750,0.625000,0.140625,0.140625,0.890625,0.015625,-0.335937,0.773437,0.062500,-0.812500,-0.132812,-0.164062,-0.609375,-0.335937,-0.296875,0.968750,-0.257812,-0.921875,0.234375,-0.742187,0.164062,0.937500,-0.820312,-0.781250,0.289062,-0.031250,0.609375,0.101562,0.312500,0.992187,-0.437500,0.218750,-0.968750,0.625000,0.140625,0.140625,0.890625,0.015625,-0.335937,0.773437,0.062500,-0.812500,-0.132812,-0.164062,-0.609375,-0.335937,-0.296875,0.968750,-0.257812,-0.921875], - "model": "sha256-hash-64d" - }, - "tensor": { - "shape": [1, 3], - "data": [2758.0, 0.0, 15.0] - }, - "semantic": { - "types": ["https://verisim.db/self/type/Commit", "https://verisim.db/self/type/feature"], - "properties": { - "conventional_commit_type": "feature", - "files_changed": "15" - } - }, - "temporal": { - "timestamp": "2026-02-13T14:18:54Z", - "version": 1, - "author": "Jonathan D.A. Jewell" - } -} diff --git a/verisimdb/.verisimdb/octads/commit-d515ccac6a03.json b/verisimdb/.verisimdb/octads/commit-d515ccac6a03.json deleted file mode 100644 index daabda3f..00000000 --- a/verisimdb/.verisimdb/octads/commit-d515ccac6a03.json +++ /dev/null @@ -1,39 +0,0 @@ -{ - "id": "commit-d515ccac6a03", - "source": "git-log", - "created_at": "2026-01-22T09:38:45Z", - "document": { - "title": "docs: update README with challenge document links", - "body": "Commit d515ccac by Your Name: docs: update README with challenge document links", - "fields": { - "type": "commit", - "hash": "d515ccac6a03c06b9ded971613262b48ca42915e", - "author": "Your Name", - "email": "you@example.com", - "commit_type": "documentation" - } - }, - "graph": { - "relationships": [{"predicate":"modifies","target":"file:README.adoc"}] - }, - "vector": { - "embedding": [-0.054687,-0.625000,0.609375,0.390625,0.367187,0.343750,0.734375,0.406250,0.750000,0.328125,0.421875,-0.640625,0.218750,0.468750,0.546875,0.593750,-0.468750,-1.000000,0.304687,0.375000,0.421875,0.796875,-0.414062,-1.000000,0.039062,-0.960937,0.742187,0.179687,0.101562,0.304687,0.257812,0.101562,-0.054687,-0.625000,0.609375,0.390625,0.367187,0.343750,0.734375,0.406250,0.750000,0.328125,0.421875,-0.640625,0.218750,0.468750,0.546875,0.593750,-0.468750,-1.000000,0.304687,0.375000,0.421875,0.796875,-0.414062,-1.000000,0.039062,-0.960937,0.742187,0.179687,0.101562,0.304687,0.257812,0.101562], - "model": "sha256-hash-64d" - }, - "tensor": { - "shape": [1, 3], - "data": [5.0, 0.0, 1.0] - }, - "semantic": { - "types": ["https://verisim.db/self/type/Commit", "https://verisim.db/self/type/documentation"], - "properties": { - "conventional_commit_type": "documentation", - "files_changed": "1" - } - }, - "temporal": { - "timestamp": "2026-01-22T09:38:45Z", - "version": 1, - "author": "Your Name" - } -} diff --git a/verisimdb/.verisimdb/octads/commit-d6ac4d5f77c0.json b/verisimdb/.verisimdb/octads/commit-d6ac4d5f77c0.json deleted file mode 100644 index 177826b1..00000000 --- a/verisimdb/.verisimdb/octads/commit-d6ac4d5f77c0.json +++ /dev/null @@ -1,39 +0,0 @@ -{ - "id": "commit-d6ac4d5f77c0", - "source": "git-log", - "created_at": "2026-01-25T06:58:34Z", - "document": { - "title": "Remove duplicate CI/CD workflows", - "body": "Commit d6ac4d5f by Your Name: Remove duplicate CI/CD workflows", - "fields": { - "type": "commit", - "hash": "d6ac4d5f77c080ff2ca93a98972c8a6b72e8f404", - "author": "Your Name", - "email": "you@example.com", - "commit_type": "other" - } - }, - "graph": { - "relationships": [{"predicate":"modifies","target":"file:.github/workflows/jekyll.yml"}] - }, - "vector": { - "embedding": [0.234375,-0.937500,-0.078125,-0.289062,0.226562,0.320312,-0.226562,0.242187,-0.914062,-0.773437,0.468750,-0.421875,0.546875,0.140625,0,-0.859375,0.804687,-0.781250,-0.992187,0.328125,-0.234375,0.656250,0.835937,-0.859375,0.500000,-0.335937,0.445312,-0.882812,-0.273437,-0.554687,-0.562500,0.117187,0.234375,-0.937500,-0.078125,-0.289062,0.226562,0.320312,-0.226562,0.242187,-0.914062,-0.773437,0.468750,-0.421875,0.546875,0.140625,0,-0.859375,0.804687,-0.781250,-0.992187,0.328125,-0.234375,0.656250,0.835937,-0.859375,0.500000,-0.335937,0.445312,-0.882812,-0.273437,-0.554687,-0.562500,0.117187], - "model": "sha256-hash-64d" - }, - "tensor": { - "shape": [1, 3], - "data": [0.0, 66.0, 1.0] - }, - "semantic": { - "types": ["https://verisim.db/self/type/Commit", "https://verisim.db/self/type/other"], - "properties": { - "conventional_commit_type": "other", - "files_changed": "1" - } - }, - "temporal": { - "timestamp": "2026-01-25T06:58:34Z", - "version": 1, - "author": "Your Name" - } -} diff --git a/verisimdb/.verisimdb/octads/commit-d7a13170c87c.json b/verisimdb/.verisimdb/octads/commit-d7a13170c87c.json deleted file mode 100644 index 50cb9c4c..00000000 --- a/verisimdb/.verisimdb/octads/commit-d7a13170c87c.json +++ /dev/null @@ -1,39 +0,0 @@ -{ - "id": "commit-d7a13170c87c", - "source": "git-log", - "created_at": "2026-01-22T11:04:30Z", - "document": { - "title": "Implement backwards compatibility and drift handling", - "body": "Commit d7a13170 by Your Name: Implement backwards compatibility and drift handling", - "fields": { - "type": "commit", - "hash": "d7a13170c87c9a2cbbbcee82cdd408ff0cc81242", - "author": "Your Name", - "email": "you@example.com", - "commit_type": "other" - } - }, - "graph": { - "relationships": [{"predicate":"modifies","target":"file:README.adoc"},{"predicate":"modifies","target":"file:docs/backwards-compatibility.adoc"},{"predicate":"modifies","target":"file:docs/drift-handling.adoc"},{"predicate":"modifies","target":"file:docs/safety-theory-applied.adoc"},{"predicate":"modifies","target":"file:docs/vcl-grammar.ebnf"},{"predicate":"modifies","target":"file:lib/verisim/adaptive_learner.ex"}] - }, - "vector": { - "embedding": [-0.578125,0.734375,0.359375,0.554687,0.632812,-0.968750,0.625000,-0.828125,0.859375,0.851562,0.726562,0.734375,-0.507812,-0.218750,-0.398437,-0.460937,-0.593750,-0.070312,-0.921875,0.953125,0.148437,-0.929687,0.281250,-0.632812,0.914062,-0.328125,0.804687,0.390625,0.625000,-0.054687,0.500000,0.609375,-0.578125,0.734375,0.359375,0.554687,0.632812,-0.968750,0.625000,-0.828125,0.859375,0.851562,0.726562,0.734375,-0.507812,-0.218750,-0.398437,-0.460937,-0.593750,-0.070312,-0.921875,0.953125,0.148437,-0.929687,0.281250,-0.632812,0.914062,-0.328125,0.804687,0.390625,0.625000,-0.054687,0.500000,0.609375], - "model": "sha256-hash-64d" - }, - "tensor": { - "shape": [1, 3], - "data": [3438.0, 71.0, 6.0] - }, - "semantic": { - "types": ["https://verisim.db/self/type/Commit", "https://verisim.db/self/type/other"], - "properties": { - "conventional_commit_type": "other", - "files_changed": "6" - } - }, - "temporal": { - "timestamp": "2026-01-22T11:04:30Z", - "version": 1, - "author": "Your Name" - } -} diff --git a/verisimdb/.verisimdb/octads/commit-d8174107cef6.json b/verisimdb/.verisimdb/octads/commit-d8174107cef6.json deleted file mode 100644 index ce6e5711..00000000 --- a/verisimdb/.verisimdb/octads/commit-d8174107cef6.json +++ /dev/null @@ -1,39 +0,0 @@ -{ - "id": "commit-d8174107cef6", - "source": "git-log", - "created_at": "2026-01-22T11:05:33Z", - "document": { - "title": "Convert KRaft comparison from CSV to AsciiDoc", - "body": "Commit d8174107 by Your Name: Convert KRaft comparison from CSV to AsciiDoc", - "fields": { - "type": "commit", - "hash": "d8174107cef68b16f6ecc5d0d4077ffe3e321897", - "author": "Your Name", - "email": "you@example.com", - "commit_type": "other" - } - }, - "graph": { - "relationships": [{"predicate":"modifies","target":"file:Kraft cross-integration to VeriSim.csv"},{"predicate":"modifies","target":"file:README.adoc"},{"predicate":"modifies","target":"file:kraft-comparison.adoc"}] - }, - "vector": { - "embedding": [0.015625,0.132812,0.242187,-0.015625,-0.226562,0.171875,0.421875,0.765625,0.664062,-0.609375,0.328125,-0.273437,-0.453125,0.093750,0.679687,-0.960937,0.328125,0.906250,-0.382812,-0.992187,0.007812,-0.093750,-0.039062,-0.585937,-0.531250,0.007812,0.148437,0.976562,0.757812,-0.296875,0.890625,0.226562,0.015625,0.132812,0.242187,-0.015625,-0.226562,0.171875,0.421875,0.765625,0.664062,-0.609375,0.328125,-0.273437,-0.453125,0.093750,0.679687,-0.960937,0.328125,0.906250,-0.382812,-0.992187,0.007812,-0.093750,-0.039062,-0.585937,-0.531250,0.007812,0.148437,0.976562,0.757812,-0.296875,0.890625,0.226562], - "model": "sha256-hash-64d" - }, - "tensor": { - "shape": [1, 3], - "data": [70.0, 7.0, 3.0] - }, - "semantic": { - "types": ["https://verisim.db/self/type/Commit", "https://verisim.db/self/type/other"], - "properties": { - "conventional_commit_type": "other", - "files_changed": "3" - } - }, - "temporal": { - "timestamp": "2026-01-22T11:05:33Z", - "version": 1, - "author": "Your Name" - } -} diff --git a/verisimdb/.verisimdb/octads/commit-d8593016e679.json b/verisimdb/.verisimdb/octads/commit-d8593016e679.json deleted file mode 100644 index 08f64a99..00000000 --- a/verisimdb/.verisimdb/octads/commit-d8593016e679.json +++ /dev/null @@ -1,39 +0,0 @@ -{ - "id": "commit-d8593016e679", - "source": "git-log", - "created_at": "2026-01-22T12:04:45Z", - "document": { - "title": "Fix optional group syntax: (...)? → [...]", - "body": "Commit d8593016 by Your Name: Fix optional group syntax: (...)? → [...]", - "fields": { - "type": "commit", - "hash": "d8593016e67954a0e1a429bbcde30e8ab252dc86", - "author": "Your Name", - "email": "you@example.com", - "commit_type": "other" - } - }, - "graph": { - "relationships": [{"predicate":"modifies","target":"file:docs/vcl-grammar.ebnf"}] - }, - "vector": { - "embedding": [0.031250,0.429687,0.984375,0.632812,0.593750,-0.445312,0.007812,0.453125,-0.968750,0.109375,0.726562,0.820312,0.484375,0.742187,-0.539062,-0.359375,-0.265625,-0.554687,-0.039062,0.656250,0.937500,0.187500,-0.601562,-0.187500,-0.812500,0.554687,-0.953125,0.421875,-0.085937,-0.031250,-0.304687,0.125000,0.031250,0.429687,0.984375,0.632812,0.593750,-0.445312,0.007812,0.453125,-0.968750,0.109375,0.726562,0.820312,0.484375,0.742187,-0.539062,-0.359375,-0.265625,-0.554687,-0.039062,0.656250,0.937500,0.187500,-0.601562,-0.187500,-0.812500,0.554687,-0.953125,0.421875,-0.085937,-0.031250,-0.304687,0.125000], - "model": "sha256-hash-64d" - }, - "tensor": { - "shape": [1, 3], - "data": [1.0, 1.0, 1.0] - }, - "semantic": { - "types": ["https://verisim.db/self/type/Commit", "https://verisim.db/self/type/other"], - "properties": { - "conventional_commit_type": "other", - "files_changed": "1" - } - }, - "temporal": { - "timestamp": "2026-01-22T12:04:45Z", - "version": 1, - "author": "Your Name" - } -} diff --git a/verisimdb/.verisimdb/octads/commit-d949b42717bb.json b/verisimdb/.verisimdb/octads/commit-d949b42717bb.json deleted file mode 100644 index 31e1aae0..00000000 --- a/verisimdb/.verisimdb/octads/commit-d949b42717bb.json +++ /dev/null @@ -1,39 +0,0 @@ -{ - "id": "commit-d949b42717bb", - "source": "git-log", - "created_at": "2026-02-08T14:15:10Z", - "document": { - "title": "docs: add GitHub CI integration, hypatia pipeline, and model router plans", - "body": "Commit d949b427 by Jonathan D.A. Jewell: docs: add GitHub CI integration, hypatia pipeline, and model router plans", - "fields": { - "type": "commit", - "hash": "d949b42717bbea54904eea9e7b62da50428c6b37", - "author": "Jonathan D.A. Jewell", - "email": "jonathan.jewell@open.ac.uk", - "commit_type": "documentation" - } - }, - "graph": { - "relationships": [{"predicate":"modifies","target":"file:.claude/CLAUDE.md"}] - }, - "vector": { - "embedding": [-0.914062,-0.421875,-0.148437,0.242187,0.687500,-0.156250,0.250000,-0.796875,-0.953125,0.062500,-0.531250,0.914062,-0.578125,-0.820312,-0.507812,0.031250,-0.218750,0.851562,-0.476562,0.617187,-0.468750,0.578125,-0.492187,-0.882812,0.335937,-0.468750,0.781250,0.132812,0.562500,-0.992187,-0.234375,0.031250,-0.914062,-0.421875,-0.148437,0.242187,0.687500,-0.156250,0.250000,-0.796875,-0.953125,0.062500,-0.531250,0.914062,-0.578125,-0.820312,-0.507812,0.031250,-0.218750,0.851562,-0.476562,0.617187,-0.468750,0.578125,-0.492187,-0.882812,0.335937,-0.468750,0.781250,0.132812,0.562500,-0.992187,-0.234375,0.031250], - "model": "sha256-hash-64d" - }, - "tensor": { - "shape": [1, 3], - "data": [102.0, 0.0, 1.0] - }, - "semantic": { - "types": ["https://verisim.db/self/type/Commit", "https://verisim.db/self/type/documentation"], - "properties": { - "conventional_commit_type": "documentation", - "files_changed": "1" - } - }, - "temporal": { - "timestamp": "2026-02-08T14:15:10Z", - "version": 1, - "author": "Jonathan D.A. Jewell" - } -} diff --git a/verisimdb/.verisimdb/octads/commit-d961f137140e.json b/verisimdb/.verisimdb/octads/commit-d961f137140e.json deleted file mode 100644 index 1f2ad0e5..00000000 --- a/verisimdb/.verisimdb/octads/commit-d961f137140e.json +++ /dev/null @@ -1,39 +0,0 @@ -{ - "id": "commit-d961f137140e", - "source": "git-log", - "created_at": "2026-02-13T14:18:41Z", - "document": { - "title": "feat: add VCL bidirectional type system, cross-modal conditions, and mutations", - "body": "Commit d961f137 by Jonathan D.A. Jewell: feat: add VCL bidirectional type system, cross-modal conditions, and mutations", - "fields": { - "type": "commit", - "hash": "d961f137140e7d3ab44503b8f2549298e01aefde", - "author": "Jonathan D.A. Jewell", - "email": "jonathan.jewell@open.ac.uk", - "commit_type": "feature" - } - }, - "graph": { - "relationships": [] - }, - "vector": { - "embedding": [0.289062,-0.593750,0.695312,0.507812,0.460937,-0.945312,0.609375,0.976562,-0.796875,-0.437500,0.093750,0.398437,0.343750,0.890625,0.507812,-0.296875,-0.320312,0.304687,-0.796875,-0.898437,0.289062,0.718750,0.242187,0.460937,-0.796875,-0.578125,0.820312,-0.500000,0.164062,0.429687,0.140625,0.851562,0.289062,-0.593750,0.695312,0.507812,0.460937,-0.945312,0.609375,0.976562,-0.796875,-0.437500,0.093750,0.398437,0.343750,0.890625,0.507812,-0.296875,-0.320312,0.304687,-0.796875,-0.898437,0.289062,0.718750,0.242187,0.460937,-0.796875,-0.578125,0.820312,-0.500000,0.164062,0.429687,0.140625,0.851562], - "model": "sha256-hash-64d" - }, - "tensor": { - "shape": [1, 3], - "data": [5057.0, 484.0, 24.0] - }, - "semantic": { - "types": ["https://verisim.db/self/type/Commit", "https://verisim.db/self/type/feature"], - "properties": { - "conventional_commit_type": "feature", - "files_changed": "24" - } - }, - "temporal": { - "timestamp": "2026-02-13T14:18:41Z", - "version": 1, - "author": "Jonathan D.A. Jewell" - } -} diff --git a/verisimdb/.verisimdb/octads/commit-dd8890c42dde.json b/verisimdb/.verisimdb/octads/commit-dd8890c42dde.json deleted file mode 100644 index ff8b2e6c..00000000 --- a/verisimdb/.verisimdb/octads/commit-dd8890c42dde.json +++ /dev/null @@ -1,39 +0,0 @@ -{ - "id": "commit-dd8890c42dde", - "source": "git-log", - "created_at": "2026-02-04T21:44:55Z", - "document": { - "title": "feat: complete VCL parser and fix license headers", - "body": "Commit dd8890c4 by Jonathan D.A. Jewell: feat: complete VCL parser and fix license headers", - "fields": { - "type": "commit", - "hash": "dd8890c42dde0d7d754603358053d7073c6d19c3", - "author": "Jonathan D.A. Jewell", - "email": "jonathan.jewell@open.ac.uk", - "commit_type": "feature" - } - }, - "graph": { - "relationships": [] - }, - "vector": { - "embedding": [0.781250,0.695312,0.468750,-0.203125,0.671875,0.804687,0.656250,-0.539062,0.421875,-0.265625,0.351562,-0.242187,0,0.281250,-0.914062,-0.750000,-0.914062,0.328125,-0.609375,-0.468750,-0.523437,0.062500,0.242187,0.859375,0.093750,-0.453125,-0.539062,0.085937,0.585937,-0.906250,0.500000,-0.539062,0.781250,0.695312,0.468750,-0.203125,0.671875,0.804687,0.656250,-0.539062,0.421875,-0.265625,0.351562,-0.242187,0,0.281250,-0.914062,-0.750000,-0.914062,0.328125,-0.609375,-0.468750,-0.523437,0.062500,0.242187,0.859375,0.093750,-0.453125,-0.539062,0.085937,0.585937,-0.906250,0.500000,-0.539062], - "model": "sha256-hash-64d" - }, - "tensor": { - "shape": [1, 3], - "data": [385.0, 104.0, 25.0] - }, - "semantic": { - "types": ["https://verisim.db/self/type/Commit", "https://verisim.db/self/type/feature"], - "properties": { - "conventional_commit_type": "feature", - "files_changed": "25" - } - }, - "temporal": { - "timestamp": "2026-02-04T21:44:55Z", - "version": 1, - "author": "Jonathan D.A. Jewell" - } -} diff --git a/verisimdb/.verisimdb/octads/commit-ddc83dd74b3d.json b/verisimdb/.verisimdb/octads/commit-ddc83dd74b3d.json deleted file mode 100644 index b6615846..00000000 --- a/verisimdb/.verisimdb/octads/commit-ddc83dd74b3d.json +++ /dev/null @@ -1,39 +0,0 @@ -{ - "id": "commit-ddc83dd74b3d", - "source": "git-log", - "created_at": "2026-01-16T19:44:16Z", - "document": { - "title": "fix: resolve remaining compilation errors in graph, normalizer, and api crates", - "body": "Commit ddc83dd7 by Jonathan D.A. Jewell: fix: resolve remaining compilation errors in graph, normalizer, and api crates", - "fields": { - "type": "commit", - "hash": "ddc83dd74b3dd8aaebb4e30f91f3b6ca1100f665", - "author": "Jonathan D.A. Jewell", - "email": "jonathan.jewell@open.ac.uk", - "commit_type": "bugfix" - } - }, - "graph": { - "relationships": [{"predicate":"modifies","target":"file:rust-core/verisim-api/Cargo.toml"},{"predicate":"modifies","target":"file:rust-core/verisim-api/src/lib.rs"},{"predicate":"modifies","target":"file:rust-core/verisim-graph/src/lib.rs"},{"predicate":"modifies","target":"file:rust-core/verisim-normalizer/Cargo.toml"}] - }, - "vector": { - "embedding": [0.445312,-0.703125,-0.335937,-0.632812,-0.164062,-0.671875,0.906250,-0.820312,0.304687,0.867187,-0.671875,0.007812,0.742187,0.687500,0.914062,0.664062,-0.445312,0.179687,0.164062,-0.593750,-0.078125,0.523437,0.953125,0.484375,0.976562,0.609375,-0.875000,0.343750,-0.296875,0.953125,0.851562,0.781250,0.445312,-0.703125,-0.335937,-0.632812,-0.164062,-0.671875,0.906250,-0.820312,0.304687,0.867187,-0.671875,0.007812,0.742187,0.687500,0.914062,0.664062,-0.445312,0.179687,0.164062,-0.593750,-0.078125,0.523437,0.953125,0.484375,0.976562,0.609375,-0.875000,0.343750,-0.296875,0.953125,0.851562,0.781250], - "model": "sha256-hash-64d" - }, - "tensor": { - "shape": [1, 3], - "data": [6.0, 4.0, 4.0] - }, - "semantic": { - "types": ["https://verisim.db/self/type/Commit", "https://verisim.db/self/type/bugfix"], - "properties": { - "conventional_commit_type": "bugfix", - "files_changed": "4" - } - }, - "temporal": { - "timestamp": "2026-01-16T19:44:16Z", - "version": 1, - "author": "Jonathan D.A. Jewell" - } -} diff --git a/verisimdb/.verisimdb/octads/commit-de2fd383992a.json b/verisimdb/.verisimdb/octads/commit-de2fd383992a.json deleted file mode 100644 index 37b8d3f6..00000000 --- a/verisimdb/.verisimdb/octads/commit-de2fd383992a.json +++ /dev/null @@ -1,39 +0,0 @@ -{ - "id": "commit-de2fd383992a", - "source": "git-log", - "created_at": "2026-01-22T11:48:09Z", - "document": { - "title": "Fix VCL grammar for ISO/IEC 14977 EBNF compliance", - "body": "Commit de2fd383 by Your Name: Fix VCL grammar for ISO/IEC 14977 EBNF compliance", - "fields": { - "type": "commit", - "hash": "de2fd383992ac9dd114b3810e984952bdbc411e4", - "author": "Your Name", - "email": "you@example.com", - "commit_type": "other" - } - }, - "graph": { - "relationships": [{"predicate":"modifies","target":"file:docs/vcl-grammar.ebnf"}] - }, - "vector": { - "embedding": [-0.429687,0.445312,0.171875,0.179687,0.367187,0.859375,-0.531250,-0.890625,-0.828125,0.531250,0.179687,0.296875,0.859375,-0.328125,0.906250,0.718750,-0.289062,-0.328125,0.710937,0.320312,0.085937,0.414062,0.734375,-0.953125,0.679687,0.320312,0.953125,0.960937,0.828125,-0.968750,-0.945312,-0.968750,-0.429687,0.445312,0.171875,0.179687,0.367187,0.859375,-0.531250,-0.890625,-0.828125,0.531250,0.179687,0.296875,0.859375,-0.328125,0.906250,0.718750,-0.289062,-0.328125,0.710937,0.320312,0.085937,0.414062,0.734375,-0.953125,0.679687,0.320312,0.953125,0.960937,0.828125,-0.968750,-0.945312,-0.968750], - "model": "sha256-hash-64d" - }, - "tensor": { - "shape": [1, 3], - "data": [24.0, 92.0, 1.0] - }, - "semantic": { - "types": ["https://verisim.db/self/type/Commit", "https://verisim.db/self/type/other"], - "properties": { - "conventional_commit_type": "other", - "files_changed": "1" - } - }, - "temporal": { - "timestamp": "2026-01-22T11:48:09Z", - "version": 1, - "author": "Your Name" - } -} diff --git a/verisimdb/.verisimdb/octads/commit-dfe015d2ea26.json b/verisimdb/.verisimdb/octads/commit-dfe015d2ea26.json deleted file mode 100644 index 0d874b0c..00000000 --- a/verisimdb/.verisimdb/octads/commit-dfe015d2ea26.json +++ /dev/null @@ -1,39 +0,0 @@ -{ - "id": "commit-dfe015d2ea26", - "source": "git-log", - "created_at": "2026-01-22T09:03:32Z", - "document": { - "title": "chore: update license identifier to PMPL-1.0-or-later in README", - "body": "Commit dfe015d2 by Your Name: chore: update license identifier to PMPL-1.0-or-later in README", - "fields": { - "type": "commit", - "hash": "dfe015d2ea267bbd8c94148db453c0f8674e847e", - "author": "Your Name", - "email": "you@example.com", - "commit_type": "chore" - } - }, - "graph": { - "relationships": [{"predicate":"modifies","target":"file:README.adoc"}] - }, - "vector": { - "embedding": [-0.554687,0.101562,-0.570312,-0.156250,0.312500,0.148437,0.968750,0.968750,0.664062,-0.265625,-0.906250,0.718750,0.023437,-0.992187,-0.078125,0.328125,0.804687,0.054687,0.132812,0.039062,0.921875,0.421875,0.398437,-0.421875,0.218750,0.062500,0.203125,-0.046875,0.101562,0.554687,0.867187,-0.187500,-0.554687,0.101562,-0.570312,-0.156250,0.312500,0.148437,0.968750,0.968750,0.664062,-0.265625,-0.906250,0.718750,0.023437,-0.992187,-0.078125,0.328125,0.804687,0.054687,0.132812,0.039062,0.921875,0.421875,0.398437,-0.421875,0.218750,0.062500,0.203125,-0.046875,0.101562,0.554687,0.867187,-0.187500], - "model": "sha256-hash-64d" - }, - "tensor": { - "shape": [1, 3], - "data": [2.0, 2.0, 1.0] - }, - "semantic": { - "types": ["https://verisim.db/self/type/Commit", "https://verisim.db/self/type/chore"], - "properties": { - "conventional_commit_type": "chore", - "files_changed": "1" - } - }, - "temporal": { - "timestamp": "2026-01-22T09:03:32Z", - "version": 1, - "author": "Your Name" - } -} diff --git a/verisimdb/.verisimdb/octads/commit-e594e11c006e.json b/verisimdb/.verisimdb/octads/commit-e594e11c006e.json deleted file mode 100644 index bb1a7294..00000000 --- a/verisimdb/.verisimdb/octads/commit-e594e11c006e.json +++ /dev/null @@ -1,39 +0,0 @@ -{ - "id": "commit-e594e11c006e", - "source": "git-log", - "created_at": "2026-02-12T16:11:04Z", - "document": { - "title": "chore: fix clippy warnings across workspace", - "body": "Commit e594e11c by Jonathan D.A. Jewell: chore: fix clippy warnings across workspace", - "fields": { - "type": "commit", - "hash": "e594e11c006e8031c7e4a92c607a49500d41c32f", - "author": "Jonathan D.A. Jewell", - "email": "jonathan.jewell@open.ac.uk", - "commit_type": "chore" - } - }, - "graph": { - "relationships": [{"predicate":"modifies","target":"file:rust-core/verisim-api/src/lib.rs"},{"predicate":"modifies","target":"file:rust-core/verisim-document/src/lib.rs"},{"predicate":"modifies","target":"file:rust-core/verisim-octad/src/lib.rs"},{"predicate":"modifies","target":"file:rust-core/verisim-normalizer/src/lib.rs"},{"predicate":"modifies","target":"file:rust-core/verisim-temporal/src/lib.rs"},{"predicate":"modifies","target":"file:rust-core/verisim-tensor/src/lib.rs"}] - }, - "vector": { - "embedding": [-0.562500,0.515625,-0.101562,-0.515625,-0.164062,0.968750,-0.265625,0.601562,-0.062500,-0.921875,-0.093750,-0.914062,-0.414062,-0.468750,0.445312,0.703125,0.312500,-0.429687,0.781250,0.757812,0.210937,-0.843750,-0.601562,-0.640625,-0.046875,0.281250,0.546875,0.148437,0.054687,0.179687,0.257812,0.765625,-0.562500,0.515625,-0.101562,-0.515625,-0.164062,0.968750,-0.265625,0.601562,-0.062500,-0.921875,-0.093750,-0.914062,-0.414062,-0.468750,0.445312,0.703125,0.312500,-0.429687,0.781250,0.757812,0.210937,-0.843750,-0.601562,-0.640625,-0.046875,0.281250,0.546875,0.148437,0.054687,0.179687,0.257812,0.765625], - "model": "sha256-hash-64d" - }, - "tensor": { - "shape": [1, 3], - "data": [11.0, 17.0, 6.0] - }, - "semantic": { - "types": ["https://verisim.db/self/type/Commit", "https://verisim.db/self/type/chore"], - "properties": { - "conventional_commit_type": "chore", - "files_changed": "6" - } - }, - "temporal": { - "timestamp": "2026-02-12T16:11:04Z", - "version": 1, - "author": "Jonathan D.A. Jewell" - } -} diff --git a/verisimdb/.verisimdb/octads/commit-e6dbf191a423.json b/verisimdb/.verisimdb/octads/commit-e6dbf191a423.json deleted file mode 100644 index 8621bb6a..00000000 --- a/verisimdb/.verisimdb/octads/commit-e6dbf191a423.json +++ /dev/null @@ -1,39 +0,0 @@ -{ - "id": "commit-e6dbf191a423", - "source": "git-log", - "created_at": "2026-01-22T12:00:43Z", - "document": { - "title": "Replace all regex-style character classes with ISO EBNF equivalents", - "body": "Commit e6dbf191 by Your Name: Replace all regex-style character classes with ISO EBNF equivalents", - "fields": { - "type": "commit", - "hash": "e6dbf191a4236c45100de816ba6d7d551df7c10e", - "author": "Your Name", - "email": "you@example.com", - "commit_type": "other" - } - }, - "graph": { - "relationships": [{"predicate":"modifies","target":"file:docs/vcl-grammar.ebnf"}] - }, - "vector": { - "embedding": [-0.195312,-0.765625,-0.445312,-0.937500,0.929687,-0.906250,0.210937,-0.117187,0.046875,-0.828125,0.062500,-0.640625,0.640625,-0.453125,0.312500,-0.453125,-0.429687,-0.835937,0.234375,0.609375,-0.804687,0.421875,0.867187,-0.070312,-0.328125,-0.382812,-0.226562,-0.664062,0.804687,-0.148437,0.414062,-0.906250,-0.195312,-0.765625,-0.445312,-0.937500,0.929687,-0.906250,0.210937,-0.117187,0.046875,-0.828125,0.062500,-0.640625,0.640625,-0.453125,0.312500,-0.453125,-0.429687,-0.835937,0.234375,0.609375,-0.804687,0.421875,0.867187,-0.070312,-0.328125,-0.382812,-0.226562,-0.664062,0.804687,-0.148437,0.414062,-0.906250], - "model": "sha256-hash-64d" - }, - "tensor": { - "shape": [1, 3], - "data": [7.0, 5.0, 1.0] - }, - "semantic": { - "types": ["https://verisim.db/self/type/Commit", "https://verisim.db/self/type/other"], - "properties": { - "conventional_commit_type": "other", - "files_changed": "1" - } - }, - "temporal": { - "timestamp": "2026-01-22T12:00:43Z", - "version": 1, - "author": "Your Name" - } -} diff --git a/verisimdb/.verisimdb/octads/commit-e7bd5a33403b.json b/verisimdb/.verisimdb/octads/commit-e7bd5a33403b.json deleted file mode 100644 index c66d9026..00000000 --- a/verisimdb/.verisimdb/octads/commit-e7bd5a33403b.json +++ /dev/null @@ -1,39 +0,0 @@ -{ - "id": "commit-e7bd5a33403b", - "source": "git-log", - "created_at": "2026-01-22T12:35:57Z", - "document": { - "title": "Make normalization cascade recommendation crystal clear", - "body": "Commit e7bd5a33 by Your Name: Make normalization cascade recommendation crystal clear", - "fields": { - "type": "commit", - "hash": "e7bd5a33403b00a68f475a36fc57a62ea0e911a4", - "author": "Your Name", - "email": "you@example.com", - "commit_type": "other" - } - }, - "graph": { - "relationships": [{"predicate":"modifies","target":"file:docs/normalization-cascade.adoc"}] - }, - "vector": { - "embedding": [-0.179687,0.265625,-0.031250,0.218750,-0.500000,-0.960937,-0.140625,0.078125,0.531250,-0.210937,-1.000000,-0.367187,0.812500,-0.679687,-0.781250,-0.164062,0.406250,0.421875,0.835937,0.046875,-0.671875,0.820312,0.632812,0.406250,0.843750,-0.929687,0.406250,0.015625,-0.062500,-0.476562,0.031250,0.921875,-0.179687,0.265625,-0.031250,0.218750,-0.500000,-0.960937,-0.140625,0.078125,0.531250,-0.210937,-1.000000,-0.367187,0.812500,-0.679687,-0.781250,-0.164062,0.406250,0.421875,0.835937,0.046875,-0.671875,0.820312,0.632812,0.406250,0.843750,-0.929687,0.406250,0.015625,-0.062500,-0.476562,0.031250,0.921875], - "model": "sha256-hash-64d" - }, - "tensor": { - "shape": [1, 3], - "data": [130.0, 0.0, 1.0] - }, - "semantic": { - "types": ["https://verisim.db/self/type/Commit", "https://verisim.db/self/type/other"], - "properties": { - "conventional_commit_type": "other", - "files_changed": "1" - } - }, - "temporal": { - "timestamp": "2026-01-22T12:35:57Z", - "version": 1, - "author": "Your Name" - } -} diff --git a/verisimdb/.verisimdb/octads/commit-e84118929733.json b/verisimdb/.verisimdb/octads/commit-e84118929733.json deleted file mode 100644 index bd8697d8..00000000 --- a/verisimdb/.verisimdb/octads/commit-e84118929733.json +++ /dev/null @@ -1,39 +0,0 @@ -{ - "id": "commit-e84118929733", - "source": "git-log", - "created_at": "2026-01-22T11:08:51Z", - "document": { - "title": "Fix VCL grammar to comply with ISO EBNF standard", - "body": "Commit e8411892 by Your Name: Fix VCL grammar to comply with ISO EBNF standard", - "fields": { - "type": "commit", - "hash": "e84118929733fd5dc0ed83c8ccc36b1c9d632d2f", - "author": "Your Name", - "email": "you@example.com", - "commit_type": "other" - } - }, - "graph": { - "relationships": [{"predicate":"modifies","target":"file:docs/vcl-grammar.ebnf"}] - }, - "vector": { - "embedding": [0.937500,0.406250,0.671875,-0.429687,-0.226562,-0.507812,-0.398437,-0.742187,0.500000,0.226562,-0.015625,0.750000,-0.992187,0.085937,-0.015625,-0.523437,-0.023437,-0.523437,-0.789062,0.507812,0.164062,-0.085937,0.640625,0.437500,-0.421875,-0.843750,-0.921875,0.375000,0.523437,0.640625,0.953125,0.132812,0.937500,0.406250,0.671875,-0.429687,-0.226562,-0.507812,-0.398437,-0.742187,0.500000,0.226562,-0.015625,0.750000,-0.992187,0.085937,-0.015625,-0.523437,-0.023437,-0.523437,-0.789062,0.507812,0.164062,-0.085937,0.640625,0.437500,-0.421875,-0.843750,-0.921875,0.375000,0.523437,0.640625,0.953125,0.132812], - "model": "sha256-hash-64d" - }, - "tensor": { - "shape": [1, 3], - "data": [11.0, 12.0, 1.0] - }, - "semantic": { - "types": ["https://verisim.db/self/type/Commit", "https://verisim.db/self/type/other"], - "properties": { - "conventional_commit_type": "other", - "files_changed": "1" - } - }, - "temporal": { - "timestamp": "2026-01-22T11:08:51Z", - "version": 1, - "author": "Your Name" - } -} diff --git a/verisimdb/.verisimdb/octads/commit-e9b738078138.json b/verisimdb/.verisimdb/octads/commit-e9b738078138.json deleted file mode 100644 index 5853ac9c..00000000 --- a/verisimdb/.verisimdb/octads/commit-e9b738078138.json +++ /dev/null @@ -1,39 +0,0 @@ -{ - "id": "commit-e9b738078138", - "source": "git-log", - "created_at": "2026-01-22T12:20:00Z", - "document": { - "title": "Add normalization cascade architecture analysis", - "body": "Commit e9b73807 by Your Name: Add normalization cascade architecture analysis", - "fields": { - "type": "commit", - "hash": "e9b7380781384bfaf9c17fe7601097b5b09cb88a", - "author": "Your Name", - "email": "you@example.com", - "commit_type": "other" - } - }, - "graph": { - "relationships": [{"predicate":"modifies","target":"file:docs/normalization-cascade.adoc"}] - }, - "vector": { - "embedding": [0.804687,-0.125000,0.117187,-0.234375,-0.312500,0.375000,-0.390625,0.390625,-0.812500,0.132812,0.054687,0.289062,0.171875,0.726562,-0.648437,0.789062,0.101562,-0.804687,0.671875,0.828125,0.359375,-0.820312,0.523437,-0.343750,-0.140625,0.585937,0.523437,-0.617187,0.507812,0.460937,-0.148437,0.632812,0.804687,-0.125000,0.117187,-0.234375,-0.312500,0.375000,-0.390625,0.390625,-0.812500,0.132812,0.054687,0.289062,0.171875,0.726562,-0.648437,0.789062,0.101562,-0.804687,0.671875,0.828125,0.359375,-0.820312,0.523437,-0.343750,-0.140625,0.585937,0.523437,-0.617187,0.507812,0.460937,-0.148437,0.632812], - "model": "sha256-hash-64d" - }, - "tensor": { - "shape": [1, 3], - "data": [632.0, 0.0, 1.0] - }, - "semantic": { - "types": ["https://verisim.db/self/type/Commit", "https://verisim.db/self/type/other"], - "properties": { - "conventional_commit_type": "other", - "files_changed": "1" - } - }, - "temporal": { - "timestamp": "2026-01-22T12:20:00Z", - "version": 1, - "author": "Your Name" - } -} diff --git a/verisimdb/.verisimdb/octads/commit-ea9f52edc373.json b/verisimdb/.verisimdb/octads/commit-ea9f52edc373.json deleted file mode 100644 index a81cda84..00000000 --- a/verisimdb/.verisimdb/octads/commit-ea9f52edc373.json +++ /dev/null @@ -1,39 +0,0 @@ -{ - "id": "commit-ea9f52edc373", - "source": "git-log", - "created_at": "2026-02-05T11:57:34Z", - "document": { - "title": "general update", - "body": "Commit ea9f52ed by Jonathan D.A. Jewell: general update", - "fields": { - "type": "commit", - "hash": "ea9f52edc373eba3e2346374d0197d29d672fc62", - "author": "Jonathan D.A. Jewell", - "email": "jonathan.jewell@open.ac.uk", - "commit_type": "other" - } - }, - "graph": { - "relationships": [{"predicate":"modifies","target":"file:.cfignore"}] - }, - "vector": { - "embedding": [-0.578125,0.156250,0.601562,-0.265625,-0.679687,-0.960937,-0.437500,0.679687,0.210937,0.414062,0.953125,-0.781250,-0.085937,-0.085937,0.117187,0.812500,0.281250,0.828125,0.609375,0.320312,-0.343750,0.023437,0.742187,-0.117187,0.445312,0.312500,-0.679687,0.757812,-0.343750,-0.625000,-0.359375,-0.437500,-0.578125,0.156250,0.601562,-0.265625,-0.679687,-0.960937,-0.437500,0.679687,0.210937,0.414062,0.953125,-0.781250,-0.085937,-0.085937,0.117187,0.812500,0.281250,0.828125,0.609375,0.320312,-0.343750,0.023437,0.742187,-0.117187,0.445312,0.312500,-0.679687,0.757812,-0.343750,-0.625000,-0.359375,-0.437500], - "model": "sha256-hash-64d" - }, - "tensor": { - "shape": [1, 3], - "data": [5.0, 0.0, 1.0] - }, - "semantic": { - "types": ["https://verisim.db/self/type/Commit", "https://verisim.db/self/type/other"], - "properties": { - "conventional_commit_type": "other", - "files_changed": "1" - } - }, - "temporal": { - "timestamp": "2026-02-05T11:57:34Z", - "version": 1, - "author": "Jonathan D.A. Jewell" - } -} diff --git a/verisimdb/.verisimdb/octads/commit-ecceaddc2f85.json b/verisimdb/.verisimdb/octads/commit-ecceaddc2f85.json deleted file mode 100644 index abf4fdea..00000000 --- a/verisimdb/.verisimdb/octads/commit-ecceaddc2f85.json +++ /dev/null @@ -1,39 +0,0 @@ -{ - "id": "commit-ecceaddc2f85", - "source": "git-log", - "created_at": "2026-02-13T15:27:18Z", - "document": { - "title": "fix: resolve issue 7 — complete ReScript registry functions", - "body": "Commit ecceaddc by Jonathan D.A. Jewell: fix: resolve issue 7 — complete ReScript registry functions", - "fields": { - "type": "commit", - "hash": "ecceaddc2f85d49cade63a4b1f9be3335b903d15", - "author": "Jonathan D.A. Jewell", - "email": "jonathan.jewell@open.ac.uk", - "commit_type": "bugfix" - } - }, - "graph": { - "relationships": [{"predicate":"modifies","target":"file:KNOWN-ISSUES.adoc"},{"predicate":"modifies","target":"file:rescript.json"},{"predicate":"modifies","target":"file:src/registry/Registry.res"}] - }, - "vector": { - "embedding": [-0.742187,0.273437,0.195312,0.007812,0.562500,-0.695312,-0.234375,-0.492187,-0.125000,0.390625,-0.562500,-1.000000,-0.757812,0.359375,0.835937,0.453125,0.789062,0.226562,-0.750000,0.531250,0,0.335937,0.328125,0.046875,0.125000,-0.148437,0.953125,0.242187,-0.867187,0.343750,0.781250,0.531250,-0.742187,0.273437,0.195312,0.007812,0.562500,-0.695312,-0.234375,-0.492187,-0.125000,0.390625,-0.562500,-1.000000,-0.757812,0.359375,0.835937,0.453125,0.789062,0.226562,-0.750000,0.531250,0,0.335937,0.328125,0.046875,0.125000,-0.148437,0.953125,0.242187,-0.867187,0.343750,0.781250,0.531250], - "model": "sha256-hash-64d" - }, - "tensor": { - "shape": [1, 3], - "data": [515.0, 30.0, 3.0] - }, - "semantic": { - "types": ["https://verisim.db/self/type/Commit", "https://verisim.db/self/type/bugfix"], - "properties": { - "conventional_commit_type": "bugfix", - "files_changed": "3" - } - }, - "temporal": { - "timestamp": "2026-02-13T15:27:18Z", - "version": 1, - "author": "Jonathan D.A. Jewell" - } -} diff --git a/verisimdb/.verisimdb/octads/commit-ee34606ba4e1.json b/verisimdb/.verisimdb/octads/commit-ee34606ba4e1.json deleted file mode 100644 index 3de10f24..00000000 --- a/verisimdb/.verisimdb/octads/commit-ee34606ba4e1.json +++ /dev/null @@ -1,39 +0,0 @@ -{ - "id": "commit-ee34606ba4e1", - "source": "git-log", - "created_at": "2026-02-13T14:54:42Z", - "document": { - "title": "fix: wire KRaft + Resolver into supervisor, fix single-node consensus bugs", - "body": "Commit ee34606b by Jonathan D.A. Jewell: fix: wire KRaft + Resolver into supervisor, fix single-node consensus bugs", - "fields": { - "type": "commit", - "hash": "ee34606ba4e1d2ea3a623a7eb28b14b05458add8", - "author": "Jonathan D.A. Jewell", - "email": "jonathan.jewell@open.ac.uk", - "commit_type": "bugfix" - } - }, - "graph": { - "relationships": [{"predicate":"modifies","target":"file:elixir-orchestration/lib/verisim/application.ex"},{"predicate":"modifies","target":"file:elixir-orchestration/lib/verisim/consensus/kraft_node.ex"},{"predicate":"modifies","target":"file:elixir-orchestration/lib/verisim/federation/resolver.ex"},{"predicate":"modifies","target":"file:elixir-orchestration/test/verisim/consensus/kraft_node_test.exs"},{"predicate":"modifies","target":"file:elixir-orchestration/test/verisim/federation/resolver_test.exs"}] - }, - "vector": { - "embedding": [0.226562,0.210937,0.421875,-0.445312,-0.273437,-0.882812,0.617187,0.484375,-0.570312,0.281250,0.085937,-0.273437,0.828125,0.429687,-0.507812,-0.023437,-0.734375,0.523437,-0.609375,-0.875000,0.789062,0.281250,0.539062,0.351562,0.335937,-0.453125,-0.625000,-0.570312,-0.007812,-0.468750,0.109375,0.968750,0.226562,0.210937,0.421875,-0.445312,-0.273437,-0.882812,0.617187,0.484375,-0.570312,0.281250,0.085937,-0.273437,0.828125,0.429687,-0.507812,-0.023437,-0.734375,0.523437,-0.609375,-0.875000,0.789062,0.281250,0.539062,0.351562,0.335937,-0.453125,-0.625000,-0.570312,-0.007812,-0.468750,0.109375,0.968750], - "model": "sha256-hash-64d" - }, - "tensor": { - "shape": [1, 3], - "data": [301.0, 4.0, 5.0] - }, - "semantic": { - "types": ["https://verisim.db/self/type/Commit", "https://verisim.db/self/type/bugfix"], - "properties": { - "conventional_commit_type": "bugfix", - "files_changed": "5" - } - }, - "temporal": { - "timestamp": "2026-02-13T14:54:42Z", - "version": 1, - "author": "Jonathan D.A. Jewell" - } -} diff --git a/verisimdb/.verisimdb/octads/commit-ef3bf59b298f.json b/verisimdb/.verisimdb/octads/commit-ef3bf59b298f.json deleted file mode 100644 index 7d4620a8..00000000 --- a/verisimdb/.verisimdb/octads/commit-ef3bf59b298f.json +++ /dev/null @@ -1,39 +0,0 @@ -{ - "id": "commit-ef3bf59b298f", - "source": "git-log", - "created_at": "2026-02-12T21:49:13Z", - "document": { - "title": "feat: pass VERISIMDB_PAT to scan-and-report workflow", - "body": "Commit ef3bf59b by Jonathan D.A. Jewell: feat: pass VERISIMDB_PAT to scan-and-report workflow", - "fields": { - "type": "commit", - "hash": "ef3bf59b298f02fa7c0201b42fd3c0cc8836cc94", - "author": "Jonathan D.A. Jewell", - "email": "jonathan.jewell@open.ac.uk", - "commit_type": "feature" - } - }, - "graph": { - "relationships": [{"predicate":"modifies","target":"file:.github/workflows/security-scan.yml"}] - }, - "vector": { - "embedding": [0.687500,-0.632812,0.968750,0.648437,-0.906250,-0.242187,-0.203125,-0.656250,0.203125,-0.539062,-0.390625,-0.812500,0.773437,0.640625,0.835937,0.617187,0.609375,0.539062,0.265625,0.382812,0.507812,0.632812,-0.382812,-0.742187,0.406250,0.304687,0.179687,0.742187,-0.375000,0.312500,0.203125,0.718750,0.687500,-0.632812,0.968750,0.648437,-0.906250,-0.242187,-0.203125,-0.656250,0.203125,-0.539062,-0.390625,-0.812500,0.773437,0.640625,0.835937,0.617187,0.609375,0.539062,0.265625,0.382812,0.507812,0.632812,-0.382812,-0.742187,0.406250,0.304687,0.179687,0.742187,-0.375000,0.312500,0.203125,0.718750], - "model": "sha256-hash-64d" - }, - "tensor": { - "shape": [1, 3], - "data": [2.0, 0.0, 1.0] - }, - "semantic": { - "types": ["https://verisim.db/self/type/Commit", "https://verisim.db/self/type/feature"], - "properties": { - "conventional_commit_type": "feature", - "files_changed": "1" - } - }, - "temporal": { - "timestamp": "2026-02-12T21:49:13Z", - "version": 1, - "author": "Jonathan D.A. Jewell" - } -} diff --git a/verisimdb/.verisimdb/octads/commit-f3821bdec9bb.json b/verisimdb/.verisimdb/octads/commit-f3821bdec9bb.json deleted file mode 100644 index 00111278..00000000 --- a/verisimdb/.verisimdb/octads/commit-f3821bdec9bb.json +++ /dev/null @@ -1,39 +0,0 @@ -{ - "id": "commit-f3821bdec9bb", - "source": "git-log", - "created_at": "2026-02-04T21:49:29Z", - "document": { - "title": "feat: complete integration tests and update project status", - "body": "Commit f3821bde by Jonathan D.A. Jewell: feat: complete integration tests and update project status", - "fields": { - "type": "commit", - "hash": "f3821bdec9bb3b37883e3a48f1d3786b334f9d5d", - "author": "Jonathan D.A. Jewell", - "email": "jonathan.jewell@open.ac.uk", - "commit_type": "feature" - } - }, - "graph": { - "relationships": [{"predicate":"modifies","target":"file:.machine_readable/STATE.scm"},{"predicate":"modifies","target":"file:elixir-orchestration/test/integration_test.exs"},{"predicate":"modifies","target":"file:tests/integration_test.rs"}] - }, - "vector": { - "embedding": [0.398437,-0.335937,0.960937,0.773437,-0.945312,0.164062,-0.664062,-0.117187,0.914062,-0.718750,0.187500,0.125000,0.539062,0.500000,0.437500,0.507812,-0.742187,-0.156250,0.625000,0.148437,0.453125,-0.304687,-0.562500,0.226562,-0.359375,-0.640625,0.765625,0.171875,-0.484375,-0.734375,0.804687,-0.914062,0.398437,-0.335937,0.960937,0.773437,-0.945312,0.164062,-0.664062,-0.117187,0.914062,-0.718750,0.187500,0.125000,0.539062,0.500000,0.437500,0.507812,-0.742187,-0.156250,0.625000,0.148437,0.453125,-0.304687,-0.562500,0.226562,-0.359375,-0.640625,0.765625,0.171875,-0.484375,-0.734375,0.804687,-0.914062], - "model": "sha256-hash-64d" - }, - "tensor": { - "shape": [1, 3], - "data": [764.0, 63.0, 3.0] - }, - "semantic": { - "types": ["https://verisim.db/self/type/Commit", "https://verisim.db/self/type/feature"], - "properties": { - "conventional_commit_type": "feature", - "files_changed": "3" - } - }, - "temporal": { - "timestamp": "2026-02-04T21:49:29Z", - "version": 1, - "author": "Jonathan D.A. Jewell" - } -} diff --git a/verisimdb/.verisimdb/octads/commit-f72914d64062.json b/verisimdb/.verisimdb/octads/commit-f72914d64062.json deleted file mode 100644 index ababa2f0..00000000 --- a/verisimdb/.verisimdb/octads/commit-f72914d64062.json +++ /dev/null @@ -1,39 +0,0 @@ -{ - "id": "commit-f72914d64062", - "source": "git-log", - "created_at": "2026-01-22T09:14:55Z", - "document": { - "title": "docs: remove converted markdown files and update links", - "body": "Commit f72914d6 by Your Name: docs: remove converted markdown files and update links", - "fields": { - "type": "commit", - "hash": "f72914d640624e22a60397f551ffea351bf891f3", - "author": "Your Name", - "email": "you@example.com", - "commit_type": "documentation" - } - }, - "graph": { - "relationships": [{"predicate":"modifies","target":"file:README.adoc"},{"predicate":"modifies","target":"file:Rescript Registry Types.md"},{"predicate":"modifies","target":"file:Snapshotting and Truncation Logic.md"},{"predicate":"modifies","target":"file:Technical Specification - KRaft Metadata Log.md"},{"predicate":"modifies","target":"file:ZKP and Sanctify Integration.md"}] - }, - "vector": { - "embedding": [-0.171875,-0.007812,0.250000,-0.781250,-0.914062,-0.656250,-0.765625,0.695312,0.218750,-0.578125,-0.398437,0.656250,0.500000,-0.718750,-0.968750,0.398437,0.281250,-0.015625,0.679687,0.554687,0.007812,0.718750,-0.914062,-0.375000,-0.976562,-0.117187,0.007812,0.062500,0.320312,-0.226562,-0.554687,-0.203125,-0.171875,-0.007812,0.250000,-0.781250,-0.914062,-0.656250,-0.765625,0.695312,0.218750,-0.578125,-0.398437,0.656250,0.500000,-0.718750,-0.968750,0.398437,0.281250,-0.015625,0.679687,0.554687,0.007812,0.718750,-0.914062,-0.375000,-0.976562,-0.117187,0.007812,0.062500,0.320312,-0.226562,-0.554687,-0.203125], - "model": "sha256-hash-64d" - }, - "tensor": { - "shape": [1, 3], - "data": [5.0, 230.0, 5.0] - }, - "semantic": { - "types": ["https://verisim.db/self/type/Commit", "https://verisim.db/self/type/documentation"], - "properties": { - "conventional_commit_type": "documentation", - "files_changed": "5" - } - }, - "temporal": { - "timestamp": "2026-01-22T09:14:55Z", - "version": 1, - "author": "Your Name" - } -} diff --git a/verisimdb/.verisimdb/octads/commit-f75fc83b4141.json b/verisimdb/.verisimdb/octads/commit-f75fc83b4141.json deleted file mode 100644 index 4ba10fce..00000000 --- a/verisimdb/.verisimdb/octads/commit-f75fc83b4141.json +++ /dev/null @@ -1,39 +0,0 @@ -{ - "id": "commit-f75fc83b4141", - "source": "git-log", - "created_at": "2026-01-22T01:37:46Z", - "document": { - "title": "Create verisim-architecture-visualisation.html", - "body": "Commit f75fc83b by Jonathan D.A. Jewell: Create verisim-architecture-visualisation.html", - "fields": { - "type": "commit", - "hash": "f75fc83b4141a7565fbd877f018cfb4aa48a9fbe", - "author": "Jonathan D.A. Jewell", - "email": "6759885+hyperpolymath@users.noreply.github.com", - "commit_type": "other" - } - }, - "graph": { - "relationships": [{"predicate":"modifies","target":"file:verisim-architecture-visualisation.html"}] - }, - "vector": { - "embedding": [-0.679687,-0.234375,-0.679687,-0.828125,-0.789062,0.664062,-0.609375,0.414062,0.156250,0.609375,-0.132812,0.320312,-0.718750,0.695312,-0.968750,-0.015625,-0.851562,0.007812,0.148437,-0.085937,0.875000,-0.921875,0.257812,-0.273437,0.859375,0.335937,-0.960937,0.406250,0.718750,0.867187,-0.218750,-0.601562,-0.679687,-0.234375,-0.679687,-0.828125,-0.789062,0.664062,-0.609375,0.414062,0.156250,0.609375,-0.132812,0.320312,-0.718750,0.695312,-0.968750,-0.015625,-0.851562,0.007812,0.148437,-0.085937,0.875000,-0.921875,0.257812,-0.273437,0.859375,0.335937,-0.960937,0.406250,0.718750,0.867187,-0.218750,-0.601562], - "model": "sha256-hash-64d" - }, - "tensor": { - "shape": [1, 3], - "data": [230.0, 0.0, 1.0] - }, - "semantic": { - "types": ["https://verisim.db/self/type/Commit", "https://verisim.db/self/type/other"], - "properties": { - "conventional_commit_type": "other", - "files_changed": "1" - } - }, - "temporal": { - "timestamp": "2026-01-22T01:37:46Z", - "version": 1, - "author": "Jonathan D.A. Jewell" - } -} diff --git a/verisimdb/.verisimdb/octads/commit-fb03b412d792.json b/verisimdb/.verisimdb/octads/commit-fb03b412d792.json deleted file mode 100644 index 748e463e..00000000 --- a/verisimdb/.verisimdb/octads/commit-fb03b412d792.json +++ /dev/null @@ -1,39 +0,0 @@ -{ - "id": "commit-fb03b412d792", - "source": "git-log", - "created_at": "2026-01-16T20:17:50Z", - "document": { - "title": "feat: implement OctadStore and drift detection algorithms", - "body": "Commit fb03b412 by Jonathan D.A. Jewell: feat: implement OctadStore and drift detection algorithms", - "fields": { - "type": "commit", - "hash": "fb03b412d792e3bd6ed9c5b4ff8ccbd77aae042c", - "author": "Jonathan D.A. Jewell", - "email": "jonathan.jewell@open.ac.uk", - "commit_type": "feature" - } - }, - "graph": { - "relationships": [{"predicate":"modifies","target":"file:rust-core/verisim-api/Cargo.toml"},{"predicate":"modifies","target":"file:rust-core/verisim-api/src/lib.rs"},{"predicate":"modifies","target":"file:rust-core/verisim-drift/src/calculator.rs"},{"predicate":"modifies","target":"file:rust-core/verisim-drift/src/lib.rs"},{"predicate":"modifies","target":"file:rust-core/verisim-octad/src/lib.rs"},{"predicate":"modifies","target":"file:rust-core/verisim-octad/src/store.rs"}] - }, - "vector": { - "embedding": [-0.851562,0.531250,-0.273437,0.843750,-0.171875,0.468750,0.265625,0.859375,0.234375,0.859375,0.359375,0.007812,-0.085937,0.460937,0.820312,-0.132812,0.156250,0.851562,0.898437,0.554687,0.343750,-0.921875,0.359375,-0.765625,0.414062,-0.367187,-0.726562,-0.515625,0.148437,0.125000,-0.421875,-0.687500,-0.851562,0.531250,-0.273437,0.843750,-0.171875,0.468750,0.265625,0.859375,0.234375,0.859375,0.359375,0.007812,-0.085937,0.460937,0.820312,-0.132812,0.156250,0.851562,0.898437,0.554687,0.343750,-0.921875,0.359375,-0.765625,0.414062,-0.367187,-0.726562,-0.515625,0.148437,0.125000,-0.421875,-0.687500], - "model": "sha256-hash-64d" - }, - "tensor": { - "shape": [1, 3], - "data": [1653.0, 67.0, 6.0] - }, - "semantic": { - "types": ["https://verisim.db/self/type/Commit", "https://verisim.db/self/type/feature"], - "properties": { - "conventional_commit_type": "feature", - "files_changed": "6" - } - }, - "temporal": { - "timestamp": "2026-01-16T20:17:50Z", - "version": 1, - "author": "Jonathan D.A. Jewell" - } -} diff --git a/verisimdb/.verisimdb/octads/commit-fbf2b307da8e.json b/verisimdb/.verisimdb/octads/commit-fbf2b307da8e.json deleted file mode 100644 index 54d45829..00000000 --- a/verisimdb/.verisimdb/octads/commit-fbf2b307da8e.json +++ /dev/null @@ -1,39 +0,0 @@ -{ - "id": "commit-fbf2b307da8e", - "source": "git-log", - "created_at": "2026-02-13T14:46:12Z", - "document": { - "title": "feat: complete triple API (REST + GraphQL + gRPC) for verisim-planner", - "body": "Commit fbf2b307 by Jonathan D.A. Jewell: feat: complete triple API (REST + GraphQL + gRPC) for verisim-planner", - "fields": { - "type": "commit", - "hash": "fbf2b307da8e28e9f3e70bbd3f9b60fd975b6562", - "author": "Jonathan D.A. Jewell", - "email": "jonathan.jewell@open.ac.uk", - "commit_type": "feature" - } - }, - "graph": { - "relationships": [{"predicate":"modifies","target":"file:Cargo.lock"},{"predicate":"modifies","target":"file:Cargo.toml"},{"predicate":"modifies","target":"file:rust-core/verisim-api/Cargo.toml"},{"predicate":"modifies","target":"file:rust-core/verisim-api/build.rs"},{"predicate":"modifies","target":"file:rust-core/verisim-api/proto/verisim.proto"},{"predicate":"modifies","target":"file:rust-core/verisim-api/src/graphql.rs"},{"predicate":"modifies","target":"file:rust-core/verisim-api/src/grpc.rs"},{"predicate":"modifies","target":"file:rust-core/verisim-api/src/lib.rs"}] - }, - "vector": { - "embedding": [0.085937,-0.335937,0.945312,-0.734375,-0.539062,0.273437,0.742187,-0.375000,-0.750000,-0.937500,-0.539062,0.164062,0.390625,0.109375,0.164062,0.265625,-0.351562,-0.757812,-0.843750,-0.921875,-0.164062,-0.609375,0.414062,-0.500000,-0.851562,-0.031250,-0.640625,-0.343750,0.257812,-0.085937,0.890625,0.234375,0.085937,-0.335937,0.945312,-0.734375,-0.539062,0.273437,0.742187,-0.375000,-0.750000,-0.937500,-0.539062,0.164062,0.390625,0.109375,0.164062,0.265625,-0.351562,-0.757812,-0.843750,-0.921875,-0.164062,-0.609375,0.414062,-0.500000,-0.851562,-0.031250,-0.640625,-0.343750,0.257812,-0.085937,0.890625,0.234375], - "model": "sha256-hash-64d" - }, - "tensor": { - "shape": [1, 3], - "data": [655.0, 9.0, 8.0] - }, - "semantic": { - "types": ["https://verisim.db/self/type/Commit", "https://verisim.db/self/type/feature"], - "properties": { - "conventional_commit_type": "feature", - "files_changed": "8" - } - }, - "temporal": { - "timestamp": "2026-02-13T14:46:12Z", - "version": 1, - "author": "Jonathan D.A. Jewell" - } -} diff --git a/verisimdb/.verisimdb/octads/commit-fc71351de8ee.json b/verisimdb/.verisimdb/octads/commit-fc71351de8ee.json deleted file mode 100644 index 454a2c9d..00000000 --- a/verisimdb/.verisimdb/octads/commit-fc71351de8ee.json +++ /dev/null @@ -1,39 +0,0 @@ -{ - "id": "commit-fc71351de8ee", - "source": "git-log", - "created_at": "2026-01-22T09:09:37Z", - "document": { - "title": "docs: update README to reflect federated architecture", - "body": "Commit fc71351d by Your Name: docs: update README to reflect federated architecture", - "fields": { - "type": "commit", - "hash": "fc71351de8eef2547d520c55cf42b628a5a72098", - "author": "Your Name", - "email": "you@example.com", - "commit_type": "documentation" - } - }, - "graph": { - "relationships": [{"predicate":"modifies","target":"file:README.adoc"}] - }, - "vector": { - "embedding": [0.898437,0.914062,0.992187,0.546875,-0.867187,-0.671875,-0.437500,-0.312500,0.578125,-0.007812,-0.625000,0.406250,0.132812,-0.289062,-0.750000,0.187500,0.304687,-0.679687,-0.078125,-0.281250,-0.187500,0.890625,-1.000000,-0.031250,-0.609375,-0.617187,0.078125,0.039062,0.843750,0.515625,-0.929687,0.335937,0.898437,0.914062,0.992187,0.546875,-0.867187,-0.671875,-0.437500,-0.312500,0.578125,-0.007812,-0.625000,0.406250,0.132812,-0.289062,-0.750000,0.187500,0.304687,-0.679687,-0.078125,-0.281250,-0.187500,0.890625,-1.000000,-0.031250,-0.609375,-0.617187,0.078125,0.039062,0.843750,0.515625,-0.929687,0.335937], - "model": "sha256-hash-64d" - }, - "tensor": { - "shape": [1, 3], - "data": [143.0, 27.0, 1.0] - }, - "semantic": { - "types": ["https://verisim.db/self/type/Commit", "https://verisim.db/self/type/documentation"], - "properties": { - "conventional_commit_type": "documentation", - "files_changed": "1" - } - }, - "temporal": { - "timestamp": "2026-01-22T09:09:37Z", - "version": 1, - "author": "Your Name" - } -} diff --git a/verisimdb/.verisimdb/octads/commit-fd3b385d4dd8.json b/verisimdb/.verisimdb/octads/commit-fd3b385d4dd8.json deleted file mode 100644 index f7148e89..00000000 --- a/verisimdb/.verisimdb/octads/commit-fd3b385d4dd8.json +++ /dev/null @@ -1,39 +0,0 @@ -{ - "id": "commit-fd3b385d4dd8", - "source": "git-log", - "created_at": "2026-01-26T09:13:14Z", - "document": { - "title": "Add opsm.toml", - "body": "Commit fd3b385d by Your Name: Add opsm.toml", - "fields": { - "type": "commit", - "hash": "fd3b385d4dd8a1424fc3d7d9a6a8e9b2a7d50279", - "author": "Your Name", - "email": "you@example.com", - "commit_type": "other" - } - }, - "graph": { - "relationships": [{"predicate":"modifies","target":"file:opsm.toml"}] - }, - "vector": { - "embedding": [0.210937,-0.875000,0.375000,0.593750,-0.460937,-0.398437,0.398437,0.937500,-0.523437,-0.757812,-0.218750,-0.648437,0.257812,-0.296875,0.296875,0.250000,-0.906250,-0.039062,-0.476562,0.226562,0.664062,0.851562,-0.867187,-0.015625,-0.445312,-0.015625,-0.273437,0.007812,0.031250,-0.789062,0.460937,-0.578125,0.210937,-0.875000,0.375000,0.593750,-0.460937,-0.398437,0.398437,0.937500,-0.523437,-0.757812,-0.218750,-0.648437,0.257812,-0.296875,0.296875,0.250000,-0.906250,-0.039062,-0.476562,0.226562,0.664062,0.851562,-0.867187,-0.015625,-0.445312,-0.015625,-0.273437,0.007812,0.031250,-0.789062,0.460937,-0.578125], - "model": "sha256-hash-64d" - }, - "tensor": { - "shape": [1, 3], - "data": [15.0, 0.0, 1.0] - }, - "semantic": { - "types": ["https://verisim.db/self/type/Commit", "https://verisim.db/self/type/other"], - "properties": { - "conventional_commit_type": "other", - "files_changed": "1" - } - }, - "temporal": { - "timestamp": "2026-01-26T09:13:14Z", - "version": 1, - "author": "Your Name" - } -} diff --git a/verisimdb/.verisimdb/octads/commit-fd7ddf29cec9.json b/verisimdb/.verisimdb/octads/commit-fd7ddf29cec9.json deleted file mode 100644 index a0ec6b9f..00000000 --- a/verisimdb/.verisimdb/octads/commit-fd7ddf29cec9.json +++ /dev/null @@ -1,39 +0,0 @@ -{ - "id": "commit-fd7ddf29cec9", - "source": "git-log", - "created_at": "2026-02-13T14:23:32Z", - "document": { - "title": "fix: proof verification short-circuit and vector data parser bracket notation", - "body": "Commit fd7ddf29 by Jonathan D.A. Jewell: fix: proof verification short-circuit and vector data parser bracket notation", - "fields": { - "type": "commit", - "hash": "fd7ddf29cec9002fe1280698fc8f122c3a31cf0d", - "author": "Jonathan D.A. Jewell", - "email": "jonathan.jewell@open.ac.uk", - "commit_type": "bugfix" - } - }, - "graph": { - "relationships": [{"predicate":"modifies","target":"file:elixir-orchestration/lib/verisim/query/vcl_executor.ex"},{"predicate":"modifies","target":"file:src/vcl/VCLParser.res"}] - }, - "vector": { - "embedding": [0.242187,-0.257812,0.023437,-0.484375,0.171875,0.179687,0.312500,0.726562,0.757812,0.562500,0.593750,-0.875000,0.320312,-0.898437,-0.359375,-0.773437,-0.562500,0.843750,-0.843750,0.945312,-0.476562,0.953125,-0.820312,0.257812,-0.328125,0.296875,-0.257812,0.468750,-0.976562,0.937500,-0.710937,0.773437,0.242187,-0.257812,0.023437,-0.484375,0.171875,0.179687,0.312500,0.726562,0.757812,0.562500,0.593750,-0.875000,0.320312,-0.898437,-0.359375,-0.773437,-0.562500,0.843750,-0.843750,0.945312,-0.476562,0.953125,-0.820312,0.257812,-0.328125,0.296875,-0.257812,0.468750,-0.976562,0.937500,-0.710937,0.773437], - "model": "sha256-hash-64d" - }, - "tensor": { - "shape": [1, 3], - "data": [70.0, 61.0, 2.0] - }, - "semantic": { - "types": ["https://verisim.db/self/type/Commit", "https://verisim.db/self/type/bugfix"], - "properties": { - "conventional_commit_type": "bugfix", - "files_changed": "2" - } - }, - "temporal": { - "timestamp": "2026-02-13T14:23:32Z", - "version": 1, - "author": "Jonathan D.A. Jewell" - } -} diff --git a/verisimdb/.verisimdb/octads/issue-001.json b/verisimdb/.verisimdb/octads/issue-001.json deleted file mode 100644 index 39e3fde2..00000000 --- a/verisimdb/.verisimdb/octads/issue-001.json +++ /dev/null @@ -1,37 +0,0 @@ -{ - "id": "issue-001", - "source": "known-issues", - "created_at": "2026-02-13T16:19:57Z", - "document": { - "title": "Normalizer Regeneration Strategies Are Stubs", - "body": "\\n**Location:** `rust-core/verisim-normalizer/`\\n\\n**Resolved:** 2026-02-12. Regeneration strategies now inspect octad data to select the authoritative modality and derive drifted modality content from it, rather than returning hardcoded placeholder strings. Each of the six modality strategies performs actual data transformation.\\n\\n**Original issue:** Every regeneration strategy returned a hardcoded `[regenerated]` placeholder string instead of actually regenerating data.\\n\\n", - "fields": { - "type": "known_issue", - "status": "resolved", - "severity": "resolved" - } - }, - "graph": { - "relationships": [{"predicate": "documented_in", "target": "file:KNOWN-ISSUES.adoc"}] - }, - "vector": { - "embedding": [-0.085937,0.742187,0.414062,-0.437500,0.171875,-0.984375,-0.585937,0.234375,0.054687,-0.765625,0.398437,-0.890625,-0.179687,0.171875,-0.265625,-0.320312,0.156250,0.328125,-0.242187,-0.414062,-0.015625,0.320312,0.078125,-0.687500,0.898437,0.562500,0.984375,-0.062500,-0.328125,0.648437,0.640625,-0.492187,-0.085937,0.742187,0.414062,-0.437500,0.171875,-0.984375,-0.585937,0.234375,0.054687,-0.765625,0.398437,-0.890625,-0.179687,0.171875,-0.265625,-0.320312,0.156250,0.328125,-0.242187,-0.414062,-0.015625,0.320312,0.078125,-0.687500,0.898437,0.562500,0.984375,-0.062500,-0.328125,0.648437,0.640625,-0.492187], - "model": "sha256-hash-64d" - }, - "tensor": { - "shape": [1, 2], - "data": [1.0, 0.0] - }, - "semantic": { - "types": ["https://verisim.db/self/type/KnownIssue", "https://verisim.db/self/type/resolved"], - "properties": { - "status": "resolved", - "severity": "resolved" - } - }, - "temporal": { - "timestamp": "2026-02-13T16:19:57Z", - "version": 1, - "author": "self-ingest" - } -} diff --git a/verisimdb/.verisimdb/octads/issue-002.json b/verisimdb/.verisimdb/octads/issue-002.json deleted file mode 100644 index a9d8cb4e..00000000 --- a/verisimdb/.verisimdb/octads/issue-002.json +++ /dev/null @@ -1,37 +0,0 @@ -{ - "id": "issue-002", - "source": "known-issues", - "created_at": "2026-02-13T16:19:57Z", - "document": { - "title": "Federation Executor Always Returns Empty", - "body": "\\n**Location:** Elixir orchestration layer, `VeriSim.FederationExecutor`\\n\\n**Resolved:** 2026-02-12. Federation executor now performs parallel HTTP fanout to registered peers via reqwest, decomposes queries into per-peer sub-queries, dispatches them, and merges results. Uses `Task.async_stream` for concurrent peer dispatch with configurable timeouts.\\n\\n**Original issue:** The federation executor always returned `{:ok, []}` regardless of query content or number of registered peers.\\n\\n", - "fields": { - "type": "known_issue", - "status": "resolved", - "severity": "resolved" - } - }, - "graph": { - "relationships": [{"predicate": "documented_in", "target": "file:KNOWN-ISSUES.adoc"}] - }, - "vector": { - "embedding": [0.070312,0.437500,-0.531250,-0.726562,0.804687,0.601562,-0.351562,0.703125,0.703125,0.578125,-0.414062,-0.070312,0.804687,-0.812500,-0.101562,0.953125,-0.023437,0.867187,-0.960937,-0.515625,0.398437,0.132812,-0.007812,0.375000,0.250000,0.781250,-0.046875,-0.109375,0.414062,-0.468750,0.789062,0.890625,0.070312,0.437500,-0.531250,-0.726562,0.804687,0.601562,-0.351562,0.703125,0.703125,0.578125,-0.414062,-0.070312,0.804687,-0.812500,-0.101562,0.953125,-0.023437,0.867187,-0.960937,-0.515625,0.398437,0.132812,-0.007812,0.375000,0.250000,0.781250,-0.046875,-0.109375,0.414062,-0.468750,0.789062,0.890625], - "model": "sha256-hash-64d" - }, - "tensor": { - "shape": [1, 2], - "data": [1.0, 0.0] - }, - "semantic": { - "types": ["https://verisim.db/self/type/KnownIssue", "https://verisim.db/self/type/resolved"], - "properties": { - "status": "resolved", - "severity": "resolved" - } - }, - "temporal": { - "timestamp": "2026-02-13T16:19:57Z", - "version": 1, - "author": "self-ingest" - } -} diff --git a/verisimdb/.verisimdb/octads/issue-003.json b/verisimdb/.verisimdb/octads/issue-003.json deleted file mode 100644 index 986feba0..00000000 --- a/verisimdb/.verisimdb/octads/issue-003.json +++ /dev/null @@ -1,37 +0,0 @@ -{ - "id": "issue-003", - "source": "known-issues", - "created_at": "2026-02-13T16:19:57Z", - "document": { - "title": "Federation Resolver Peer Queries Unimplemented", - "body": "\\n**Location:** Elixir orchestration layer, `VeriSim.Federation.Resolver`\\n\\n**Resolved:** 2026-02-12. `query_peer/3` now makes real HTTP requests to peer endpoints via `Req` with configurable timeout, parses JSON responses, and returns structured results with source store attribution and response time tracking. Drift policy filtering (strict/repair/tolerate/latest) is implemented.\\n\\n**Original issue:** The `query_peer/3` function returned `{:error, :not_implemented}`. Peers could be registered but never contacted.\\n\\n", - "fields": { - "type": "known_issue", - "status": "resolved", - "severity": "resolved" - } - }, - "graph": { - "relationships": [{"predicate": "documented_in", "target": "file:KNOWN-ISSUES.adoc"}] - }, - "vector": { - "embedding": [0.328125,-0.820312,0.453125,-0.343750,0.031250,-0.710937,0.656250,-0.320312,-0.148437,0.218750,-0.570312,-0.539062,-0.945312,0.906250,0.640625,-1.000000,-0.765625,0.953125,0.695312,0.703125,0.421875,0.234375,-0.687500,-0.109375,-0.460937,0.031250,-0.085937,0.812500,-0.453125,0.984375,0.648437,0.398437,0.328125,-0.820312,0.453125,-0.343750,0.031250,-0.710937,0.656250,-0.320312,-0.148437,0.218750,-0.570312,-0.539062,-0.945312,0.906250,0.640625,-1.000000,-0.765625,0.953125,0.695312,0.703125,0.421875,0.234375,-0.687500,-0.109375,-0.460937,0.031250,-0.085937,0.812500,-0.453125,0.984375,0.648437,0.398437], - "model": "sha256-hash-64d" - }, - "tensor": { - "shape": [1, 2], - "data": [1.0, 0.0] - }, - "semantic": { - "types": ["https://verisim.db/self/type/KnownIssue", "https://verisim.db/self/type/resolved"], - "properties": { - "status": "resolved", - "severity": "resolved" - } - }, - "temporal": { - "timestamp": "2026-02-13T16:19:58Z", - "version": 1, - "author": "self-ingest" - } -} diff --git a/verisimdb/.verisimdb/octads/issue-004.json b/verisimdb/.verisimdb/octads/issue-004.json deleted file mode 100644 index 4b76b98f..00000000 --- a/verisimdb/.verisimdb/octads/issue-004.json +++ /dev/null @@ -1,37 +0,0 @@ -{ - "id": "issue-004", - "source": "known-issues", - "created_at": "2026-02-13T16:19:58Z", - "document": { - "title": "Drift Auto-Trigger Missing", - "body": "\\n**Location:** Elixir orchestration layer, `VeriSim.DriftMonitor`\\n\\n**Resolved:** 2026-02-12. `DriftMonitor.init/1` now schedules a periodic sweep via `Process.send_after` at a configurable interval (default 60s). The `:sweep` handler performs a full drift sweep including pulling aggregate metrics from the Rust core via `RustClient.drift_status/0`, then reschedules itself. The sweep interval is configurable via the `config` option passed to `start_link`.\\n\\n**Original issue:** Drift detection had to be triggered manually. No scheduled or event-driven trigger existed.\\n\\n", - "fields": { - "type": "known_issue", - "status": "resolved", - "severity": "resolved" - } - }, - "graph": { - "relationships": [{"predicate": "documented_in", "target": "file:KNOWN-ISSUES.adoc"}] - }, - "vector": { - "embedding": [0.078125,-0.171875,-0.351562,-0.593750,-0.046875,-0.242187,0.359375,0.632812,-0.695312,0.765625,0.570312,0.445312,0.070312,0.328125,0.625000,-0.648437,0.343750,-0.976562,0.218750,-0.242187,-0.445312,-0.914062,0.781250,-0.648437,-0.304687,0.695312,-0.578125,0.710937,0.539062,0.718750,-0.078125,0.531250,0.078125,-0.171875,-0.351562,-0.593750,-0.046875,-0.242187,0.359375,0.632812,-0.695312,0.765625,0.570312,0.445312,0.070312,0.328125,0.625000,-0.648437,0.343750,-0.976562,0.218750,-0.242187,-0.445312,-0.914062,0.781250,-0.648437,-0.304687,0.695312,-0.578125,0.710937,0.539062,0.718750,-0.078125,0.531250], - "model": "sha256-hash-64d" - }, - "tensor": { - "shape": [1, 2], - "data": [1.0, 0.0] - }, - "semantic": { - "types": ["https://verisim.db/self/type/KnownIssue", "https://verisim.db/self/type/resolved"], - "properties": { - "status": "resolved", - "severity": "resolved" - } - }, - "temporal": { - "timestamp": "2026-02-13T16:19:58Z", - "version": 1, - "author": "self-ingest" - } -} diff --git a/verisimdb/.verisimdb/octads/issue-005.json b/verisimdb/.verisimdb/octads/issue-005.json deleted file mode 100644 index 679b38b8..00000000 --- a/verisimdb/.verisimdb/octads/issue-005.json +++ /dev/null @@ -1,37 +0,0 @@ -{ - "id": "issue-005", - "source": "known-issues", - "created_at": "2026-02-13T16:19:58Z", - "document": { - "title": "VCL-UT Not Connected to VCL PROOF Runtime", - "body": "\\n**Location:** `vcl-dt/` (Lean type definitions), VCL query router\\n\\n**Current state:** Lean type definitions exist in the `vcl-dt/` directory specifying dependent type contracts for all six PROOF types (EXISTENCE, INTEGRITY, CONSISTENCY, PROVENANCE, FRESHNESS, AUTHORIZATION). The VCL parser correctly parses `PROOF` clauses and the query router routes to the dependent-type execution path. However, the Lean type checker is not invoked at runtime. VCL-UT queries execute but do not generate real proof certificates.\\n\\n**Impact:** The `PROOF` clause is syntactically supported but semantically inert. Queries with `PROOF` clauses return data but the accompanying proof certificates are placeholders, not verifiable. This is the gap between the designed architecture and the current implementation.\\n\\n**Resolution:** Wire the Lean type checker into the VCL-UT execution path. This requires: (a) compiling Lean definitions to an executable checker, (b) generating proof obligations from the VCL AST, (c) invoking the checker and collecting proof witnesses from modality stores, (d) assembling and returning verifiable proof certificates. See link:docs/vcl-vs-vcl-dt.adoc[VCL Slipstream vs VCL-UT] for the phased roadmap.\\n\\n", - "fields": { - "type": "known_issue", - "status": "unknown", - "severity": "medium" - } - }, - "graph": { - "relationships": [{"predicate": "documented_in", "target": "file:KNOWN-ISSUES.adoc"}] - }, - "vector": { - "embedding": [0.882812,0.218750,0.179687,-0.609375,0.796875,-0.476562,-0.171875,-0.328125,0.679687,0.929687,0.976562,0.726562,-0.710937,-0.070312,-0.703125,0.375000,0.734375,-0.156250,0.578125,-0.921875,0.312500,0.164062,-0.445312,-0.757812,0.773437,0.484375,-0.437500,0.148437,-0.687500,0.398437,0.156250,0.546875,0.882812,0.218750,0.179687,-0.609375,0.796875,-0.476562,-0.171875,-0.328125,0.679687,0.929687,0.976562,0.726562,-0.710937,-0.070312,-0.703125,0.375000,0.734375,-0.156250,0.578125,-0.921875,0.312500,0.164062,-0.445312,-0.757812,0.773437,0.484375,-0.437500,0.148437,-0.687500,0.398437,0.156250,0.546875], - "model": "sha256-hash-64d" - }, - "tensor": { - "shape": [1, 2], - "data": [0.0, 1.0] - }, - "semantic": { - "types": ["https://verisim.db/self/type/KnownIssue", "https://verisim.db/self/type/unknown"], - "properties": { - "status": "unknown", - "severity": "medium" - } - }, - "temporal": { - "timestamp": "2026-02-13T16:19:58Z", - "version": 1, - "author": "self-ingest" - } -} diff --git a/verisimdb/.verisimdb/octads/issue-006.json b/verisimdb/.verisimdb/octads/issue-006.json deleted file mode 100644 index 88345e0b..00000000 --- a/verisimdb/.verisimdb/octads/issue-006.json +++ /dev/null @@ -1,37 +0,0 @@ -{ - "id": "issue-006", - "source": "known-issues", - "created_at": "2026-02-13T16:19:58Z", - "document": { - "title": "ZKP / Proven Library Not Integrated", - "body": "\\n**Location:** `docs/zkp-and-sanctify-integration.adoc` (design), `proven-coherence.md` (notes)\\n\\n**Current state:** The integration of zero-knowledge proofs via the sanctify library is documented as a consultation paper and design specification. The integration is not implemented. No ZKP circuits are generated, no proofs are created, and no verification occurs at runtime.\\n\\n**Impact:** Privacy-preserving proofs (where a query result can be verified without revealing the underlying data) are not available. This affects multi-tenant federation and compliance scenarios where data must remain private but provenance must be verifiable.\\n\\n**Resolution:** Implement sanctify integration after VCL-UT (issue 5) is resolved, since ZKP proof generation depends on the dependent type checking pipeline being functional.\\n\\n", - "fields": { - "type": "known_issue", - "status": "unknown", - "severity": "medium" - } - }, - "graph": { - "relationships": [{"predicate": "documented_in", "target": "file:KNOWN-ISSUES.adoc"}] - }, - "vector": { - "embedding": [-0.289062,0.109375,-0.601562,0.453125,-0.351562,-0.398437,-0.320312,0.226562,0.789062,0.875000,0.125000,0.570312,0.359375,0.851562,0.281250,-0.046875,0.523437,-0.234375,0.804687,0.734375,0.250000,-0.695312,-0.687500,0.164062,-0.765625,-0.218750,0.914062,-0.421875,0.406250,0.625000,-0.906250,0.468750,-0.289062,0.109375,-0.601562,0.453125,-0.351562,-0.398437,-0.320312,0.226562,0.789062,0.875000,0.125000,0.570312,0.359375,0.851562,0.281250,-0.046875,0.523437,-0.234375,0.804687,0.734375,0.250000,-0.695312,-0.687500,0.164062,-0.765625,-0.218750,0.914062,-0.421875,0.406250,0.625000,-0.906250,0.468750], - "model": "sha256-hash-64d" - }, - "tensor": { - "shape": [1, 2], - "data": [0.0, 1.0] - }, - "semantic": { - "types": ["https://verisim.db/self/type/KnownIssue", "https://verisim.db/self/type/unknown"], - "properties": { - "status": "unknown", - "severity": "medium" - } - }, - "temporal": { - "timestamp": "2026-02-13T16:19:58Z", - "version": 1, - "author": "self-ingest" - } -} diff --git a/verisimdb/.verisimdb/octads/issue-007.json b/verisimdb/.verisimdb/octads/issue-007.json deleted file mode 100644 index 27e903f9..00000000 --- a/verisimdb/.verisimdb/octads/issue-007.json +++ /dev/null @@ -1,37 +0,0 @@ -{ - "id": "issue-007", - "source": "known-issues", - "created_at": "2026-02-13T16:19:58Z", - "document": { - "title": "ReScript Registry 60% Complete", - "body": "\\n**Location:** `src/registry/Registry.res`, `rescript.json`\\n\\n**Resolved:** 2026-02-13. All 7 previously stubbed functions now have real implementations:\\n\\n- `executeFederatedQuery`: HTTP fan-out to eligible peer stores with trust filtering, response mapping, and result aggregation with limit enforcement.\\n- `achieveConsensus`: Quorum-based consensus with configurable mode (Strong/Quorum/Eventual), concurrent store fetch, agreement calculation.\\n- `replicateOctad`: Fetches octad from source store via HTTP, pushes to all target stores, reports per-target errors.\\n- `checkReplicationStatus`: Examines mapping locations, checks store liveness against maxStoreDowntimeMs, detects trust divergence as data divergence proxy.\\n- `detectByzantineFaults`: Computes median trust across stores for a octad, flags stores deviating >0.3 from median.\\n- `serializeRegistry`/`deserializeRegistry`: Full JSON serialization and deserialization of registry state (stores, mappings, config).\\n\\nAdded `rescript.json` for build configuration.\\n\\n", - "fields": { - "type": "known_issue", - "status": "resolved", - "severity": "resolved" - } - }, - "graph": { - "relationships": [{"predicate": "documented_in", "target": "file:KNOWN-ISSUES.adoc"}] - }, - "vector": { - "embedding": [-0.500000,-0.632812,0.617187,0.476562,-0.687500,-0.773437,-0.523437,0.757812,-0.546875,0.625000,0.523437,-0.773437,-0.101562,0.023437,0.023437,0.562500,0.234375,-0.343750,0.492187,-0.015625,0.179687,-0.812500,-0.320312,0.976562,0.195312,0.390625,0.382812,0.187500,0.226562,0.882812,0.976562,0.398437,-0.500000,-0.632812,0.617187,0.476562,-0.687500,-0.773437,-0.523437,0.757812,-0.546875,0.625000,0.523437,-0.773437,-0.101562,0.023437,0.023437,0.562500,0.234375,-0.343750,0.492187,-0.015625,0.179687,-0.812500,-0.320312,0.976562,0.195312,0.390625,0.382812,0.187500,0.226562,0.882812,0.976562,0.398437], - "model": "sha256-hash-64d" - }, - "tensor": { - "shape": [1, 2], - "data": [1.0, 0.0] - }, - "semantic": { - "types": ["https://verisim.db/self/type/KnownIssue", "https://verisim.db/self/type/resolved"], - "properties": { - "status": "resolved", - "severity": "resolved" - } - }, - "temporal": { - "timestamp": "2026-02-13T16:19:58Z", - "version": 1, - "author": "self-ingest" - } -} diff --git a/verisimdb/.verisimdb/octads/issue-008.json b/verisimdb/.verisimdb/octads/issue-008.json deleted file mode 100644 index 8ee97a48..00000000 --- a/verisimdb/.verisimdb/octads/issue-008.json +++ /dev/null @@ -1,37 +0,0 @@ -{ - "id": "issue-008", - "source": "known-issues", - "created_at": "2026-02-13T16:19:59Z", - "document": { - "title": "`debugger/Cargo.toml` Had Wrong Author", - "body": "\\n**Location:** `debugger/Cargo.toml`\\n\\n**Current state:** Fixed. The `debugger/Cargo.toml` previously listed an incorrect author. It now correctly reads `authors = [\"Jonathan D.A. Jewell \"]`.\\n\\n**Impact:** None (resolved). Documented here for audit trail completeness.\\n\\n", - "fields": { - "type": "known_issue", - "status": "unknown", - "severity": "medium" - } - }, - "graph": { - "relationships": [{"predicate": "documented_in", "target": "file:KNOWN-ISSUES.adoc"}] - }, - "vector": { - "embedding": [-0.343750,-0.703125,-0.125000,-0.085937,0.335937,0.765625,-0.835937,0.632812,0.437500,0.484375,0.703125,-0.968750,0.437500,0.632812,0.851562,-0.007812,0.468750,-0.406250,0.851562,-0.585937,-0.765625,-0.281250,0.992187,-0.265625,-0.156250,0.195312,-0.890625,0.507812,0.921875,-0.523437,-0.867187,-1.000000,-0.343750,-0.703125,-0.125000,-0.085937,0.335937,0.765625,-0.835937,0.632812,0.437500,0.484375,0.703125,-0.968750,0.437500,0.632812,0.851562,-0.007812,0.468750,-0.406250,0.851562,-0.585937,-0.765625,-0.281250,0.992187,-0.265625,-0.156250,0.195312,-0.890625,0.507812,0.921875,-0.523437,-0.867187,-1.000000], - "model": "sha256-hash-64d" - }, - "tensor": { - "shape": [1, 2], - "data": [0.0, 1.0] - }, - "semantic": { - "types": ["https://verisim.db/self/type/KnownIssue", "https://verisim.db/self/type/unknown"], - "properties": { - "status": "unknown", - "severity": "medium" - } - }, - "temporal": { - "timestamp": "2026-02-13T16:19:59Z", - "version": 1, - "author": "self-ingest" - } -} diff --git a/verisimdb/.verisimdb/octads/issue-009.json b/verisimdb/.verisimdb/octads/issue-009.json deleted file mode 100644 index ee956056..00000000 --- a/verisimdb/.verisimdb/octads/issue-009.json +++ /dev/null @@ -1,37 +0,0 @@ -{ - "id": "issue-009", - "source": "known-issues", - "created_at": "2026-02-13T16:19:59Z", - "document": { - "title": "HNSW Implementation Is Real", - "body": "\\n**Location:** `rust-core/verisim-vector/src/hnsw.rs`\\n\\n**Current state:** The HNSW (Hierarchical Navigable Small World) implementation in `verisim-vector` is a genuine, functional implementation at approximately 670 lines of Rust. A previous audit incorrectly claimed this was brute-force search. This is **not** an issue -- it is a correction of a previous mischaracterization.\\n\\n**Impact:** None (this is a positive clarification). The vector modality store uses a real approximate nearest-neighbor algorithm, not a naive linear scan.\\n\\n**Note:** This entry exists to prevent future audits from repeating the same incorrect claim.\\n\\n", - "fields": { - "type": "known_issue", - "status": "unknown", - "severity": "medium" - } - }, - "graph": { - "relationships": [{"predicate": "documented_in", "target": "file:KNOWN-ISSUES.adoc"}] - }, - "vector": { - "embedding": [-0.617187,-0.007812,-0.695312,0.664062,-0.304687,0.484375,-0.617187,0.148437,-0.476562,-0.015625,0.109375,0.421875,0.898437,-0.062500,0.921875,0.570312,0.078125,-0.109375,0.781250,-0.500000,-0.453125,-0.601562,0.328125,-0.937500,-0.601562,0.507812,-0.414062,-0.656250,-0.656250,0.468750,-0.562500,0.578125,-0.617187,-0.007812,-0.695312,0.664062,-0.304687,0.484375,-0.617187,0.148437,-0.476562,-0.015625,0.109375,0.421875,0.898437,-0.062500,0.921875,0.570312,0.078125,-0.109375,0.781250,-0.500000,-0.453125,-0.601562,0.328125,-0.937500,-0.601562,0.507812,-0.414062,-0.656250,-0.656250,0.468750,-0.562500,0.578125], - "model": "sha256-hash-64d" - }, - "tensor": { - "shape": [1, 2], - "data": [0.0, 1.0] - }, - "semantic": { - "types": ["https://verisim.db/self/type/KnownIssue", "https://verisim.db/self/type/unknown"], - "properties": { - "status": "unknown", - "severity": "medium" - } - }, - "temporal": { - "timestamp": "2026-02-13T16:19:59Z", - "version": 1, - "author": "self-ingest" - } -} diff --git a/verisimdb/.verisimdb/octads/issue-010.json b/verisimdb/.verisimdb/octads/issue-010.json deleted file mode 100644 index d46e7e3e..00000000 --- a/verisimdb/.verisimdb/octads/issue-010.json +++ /dev/null @@ -1,37 +0,0 @@ -{ - "id": "issue-010", - "source": "known-issues", - "created_at": "2026-02-13T16:19:59Z", - "document": { - "title": "verisim-api Needs Bin Target Fix", - "body": "\\n**Location:** `rust-core/verisim-api/Cargo.toml`\\n\\n**Resolved:** 2026-02-12. `src/main.rs` now exists with a `main()` function that reads host/port from environment variables (`VERISIM_HOST`/`VERISIM_PORT`) and initializes the HTTP server. `cargo run -p verisim-api` works.\\n\\n**Original issue:** The crate was missing a `src/main.rs` entry point and could not be run as a standalone binary.\\n\\n", - "fields": { - "type": "known_issue", - "status": "resolved", - "severity": "resolved" - } - }, - "graph": { - "relationships": [{"predicate": "documented_in", "target": "file:KNOWN-ISSUES.adoc"}] - }, - "vector": { - "embedding": [-0.914062,-0.187500,-0.648437,0.882812,0.531250,0.953125,-0.210937,0.242187,0.351562,0.179687,0.125000,0.328125,-0.976562,-0.835937,0.500000,0.171875,0.468750,-0.921875,0.382812,0.164062,0.773437,-0.546875,0.031250,0.796875,0.367187,0.343750,-0.296875,-0.382812,-0.156250,0.742187,-0.765625,-0.875000,-0.914062,-0.187500,-0.648437,0.882812,0.531250,0.953125,-0.210937,0.242187,0.351562,0.179687,0.125000,0.328125,-0.976562,-0.835937,0.500000,0.171875,0.468750,-0.921875,0.382812,0.164062,0.773437,-0.546875,0.031250,0.796875,0.367187,0.343750,-0.296875,-0.382812,-0.156250,0.742187,-0.765625,-0.875000], - "model": "sha256-hash-64d" - }, - "tensor": { - "shape": [1, 2], - "data": [1.0, 0.0] - }, - "semantic": { - "types": ["https://verisim.db/self/type/KnownIssue", "https://verisim.db/self/type/resolved"], - "properties": { - "status": "resolved", - "severity": "resolved" - } - }, - "temporal": { - "timestamp": "2026-02-13T16:19:59Z", - "version": 1, - "author": "self-ingest" - } -} diff --git a/verisimdb/.verisimdb/octads/issue-011.json b/verisimdb/.verisimdb/octads/issue-011.json deleted file mode 100644 index 2cae3dcc..00000000 --- a/verisimdb/.verisimdb/octads/issue-011.json +++ /dev/null @@ -1,37 +0,0 @@ -{ - "id": "issue-011", - "source": "known-issues", - "created_at": "2026-02-13T16:19:59Z", - "document": { - "title": "No Performance Baselines Established", - "body": "\\n**Location:** `benches/modality_benchmarks.rs`\\n\\n**Resolved:** 2026-02-13. Criterion benchmarks now cover all 6 modality stores plus cross-modal and drift operations:\\n\\n- **Document:** create_document, search_text (Tantivy full-text, 1000-doc corpus)\\n- **Vector:** insert (128/384/768 dims), search (10k vectors, HNSW)\\n- **Graph:** add_node, add_edge (Oxigraph)\\n- **Tensor:** store_create_64x64, store_get, reduce_sum_axis0 (ndarray)\\n- **Semantic:** register_type, get_type, proof_create_cbor, proof_verify (CBOR + ZKP verification)\\n- **Temporal:** version_create, version_get_by_number, version_get_latest, history_10, history_100\\n- **Octad:** create_octad, get_octad (unified entity)\\n- **Drift:** calculate_drift (semantic vector drift)\\n- **Cross-modal:** vector_similarity_search, fulltext_search (1000 multi-modal octads)\\n\\n**Original issue:** Only document, vector, graph, octad, drift, and cross-modal benchmarks existed. Tensor, semantic, and temporal stores had zero benchmarks.\\n\\n", - "fields": { - "type": "known_issue", - "status": "resolved", - "severity": "resolved" - } - }, - "graph": { - "relationships": [{"predicate": "documented_in", "target": "file:KNOWN-ISSUES.adoc"}] - }, - "vector": { - "embedding": [-0.546875,0.609375,0.289062,-0.554687,-0.335937,0.593750,-0.687500,0.781250,0.484375,0.421875,-0.492187,0.640625,0.164062,0.601562,0.718750,-0.726562,0.320312,-0.984375,-0.218750,-0.093750,-0.585937,-0.179687,0,0.570312,0.843750,0.804687,-0.085937,0.414062,-0.429687,0.351562,0.929687,0.687500,-0.546875,0.609375,0.289062,-0.554687,-0.335937,0.593750,-0.687500,0.781250,0.484375,0.421875,-0.492187,0.640625,0.164062,0.601562,0.718750,-0.726562,0.320312,-0.984375,-0.218750,-0.093750,-0.585937,-0.179687,0,0.570312,0.843750,0.804687,-0.085937,0.414062,-0.429687,0.351562,0.929687,0.687500], - "model": "sha256-hash-64d" - }, - "tensor": { - "shape": [1, 2], - "data": [1.0, 0.0] - }, - "semantic": { - "types": ["https://verisim.db/self/type/KnownIssue", "https://verisim.db/self/type/resolved"], - "properties": { - "status": "resolved", - "severity": "resolved" - } - }, - "temporal": { - "timestamp": "2026-02-13T16:19:59Z", - "version": 1, - "author": "self-ingest" - } -} diff --git a/verisimdb/.verisimdb/octads/issue-012.json b/verisimdb/.verisimdb/octads/issue-012.json deleted file mode 100644 index 226f4221..00000000 --- a/verisimdb/.verisimdb/octads/issue-012.json +++ /dev/null @@ -1,37 +0,0 @@ -{ - "id": "issue-012", - "source": "known-issues", - "created_at": "2026-02-13T16:19:59Z", - "document": { - "title": "Cross-Modal Drift/Consistency Were Stubs", - "body": "\\n**Location:** `elixir-orchestration/lib/verisim/query/vcl_executor.ex`\\n\\n**Resolved:** 2026-02-13. `compute_modality_drift/3` now fetches drift from the Rust drift API when available, falling back to cosine distance between extracted modality embeddings (with content fingerprinting for non-vector modalities). `compute_consistency/4` now computes real scores using the specified metric (COSINE, EUCLIDEAN, DOT_PRODUCT, JACCARD). Both functions previously returned hardcoded constants (0.0 and 0.5).\\n\\n**Original issue:** Cross-modal correlation queries parsed correctly but evaluation returned fake scores, making WHERE DRIFT(...) and CONSISTENT(...) conditions meaningless.\\n\\n", - "fields": { - "type": "known_issue", - "status": "resolved", - "severity": "resolved" - } - }, - "graph": { - "relationships": [{"predicate": "documented_in", "target": "file:KNOWN-ISSUES.adoc"}] - }, - "vector": { - "embedding": [0.015625,0.414062,-0.195312,0.218750,-0.515625,-0.984375,0.664062,0.203125,0.304687,-0.007812,0.773437,-0.601562,-0.046875,0.968750,0.195312,0.570312,0.742187,-0.671875,0.445312,0.757812,0.343750,-0.171875,-0.070312,-0.554687,-0.054687,-0.734375,-0.046875,-0.750000,-0.843750,-0.242187,-0.078125,-0.617187,0.015625,0.414062,-0.195312,0.218750,-0.515625,-0.984375,0.664062,0.203125,0.304687,-0.007812,0.773437,-0.601562,-0.046875,0.968750,0.195312,0.570312,0.742187,-0.671875,0.445312,0.757812,0.343750,-0.171875,-0.070312,-0.554687,-0.054687,-0.734375,-0.046875,-0.750000,-0.843750,-0.242187,-0.078125,-0.617187], - "model": "sha256-hash-64d" - }, - "tensor": { - "shape": [1, 2], - "data": [1.0, 0.0] - }, - "semantic": { - "types": ["https://verisim.db/self/type/KnownIssue", "https://verisim.db/self/type/resolved"], - "properties": { - "status": "resolved", - "severity": "resolved" - } - }, - "temporal": { - "timestamp": "2026-02-13T16:20:00Z", - "version": 1, - "author": "self-ingest" - } -} diff --git a/verisimdb/.verisimdb/octads/issue-013.json b/verisimdb/.verisimdb/octads/issue-013.json deleted file mode 100644 index a2bc56a2..00000000 --- a/verisimdb/.verisimdb/octads/issue-013.json +++ /dev/null @@ -1,37 +0,0 @@ -{ - "id": "issue-013", - "source": "known-issues", - "created_at": "2026-02-13T16:20:00Z", - "document": { - "title": "VCL WHERE Condition Routing Was Broken", - "body": "\\n**Location:** `elixir-orchestration/lib/verisim/query/vcl_executor.ex`\\n\\n**Resolved:** 2026-02-13. `has_fulltext_condition?/1`, `has_vector_condition?/1`, and `has_graph_pattern?/1` now walk the AST recursively to detect actual condition types. `extract_text_query/1`, `extract_vector_query/1`, and `extract_graph_query/1` now parse actual values from the AST instead of returning hardcoded placeholders. All queries were previously routed to `:multi` type regardless of conditions.\\n\\n**Original issue:** All six condition detection and extraction functions were stubs (returning `false` and hardcoded values), causing every query to be routed as a multi-modal query even when a single modality was targeted.\\n\\n", - "fields": { - "type": "known_issue", - "status": "resolved", - "severity": "resolved" - } - }, - "graph": { - "relationships": [{"predicate": "documented_in", "target": "file:KNOWN-ISSUES.adoc"}] - }, - "vector": { - "embedding": [-0.937500,0.109375,0.140625,0.718750,0.031250,-0.164062,0.710937,0.515625,0.992187,0.085937,-0.046875,-0.335937,-0.523437,-0.976562,-0.460937,0.960937,0,0.625000,0.218750,0.554687,-0.625000,0.492187,0.101562,-0.828125,-0.398437,-0.906250,0.148437,0.367187,0.343750,-0.515625,0.171875,-0.875000,-0.937500,0.109375,0.140625,0.718750,0.031250,-0.164062,0.710937,0.515625,0.992187,0.085937,-0.046875,-0.335937,-0.523437,-0.976562,-0.460937,0.960937,0,0.625000,0.218750,0.554687,-0.625000,0.492187,0.101562,-0.828125,-0.398437,-0.906250,0.148437,0.367187,0.343750,-0.515625,0.171875,-0.875000], - "model": "sha256-hash-64d" - }, - "tensor": { - "shape": [1, 2], - "data": [1.0, 0.0] - }, - "semantic": { - "types": ["https://verisim.db/self/type/KnownIssue", "https://verisim.db/self/type/resolved"], - "properties": { - "status": "resolved", - "severity": "resolved" - } - }, - "temporal": { - "timestamp": "2026-02-13T16:20:00Z", - "version": 1, - "author": "self-ingest" - } -} diff --git a/verisimdb/.verisimdb/octads/issue-014.json b/verisimdb/.verisimdb/octads/issue-014.json deleted file mode 100644 index 488d1de7..00000000 --- a/verisimdb/.verisimdb/octads/issue-014.json +++ /dev/null @@ -1,37 +0,0 @@ -{ - "id": "issue-014", - "source": "known-issues", - "created_at": "2026-02-13T16:20:00Z", - "document": { - "title": "Proof Verification Was No-Op", - "body": "\\n**Location:** `elixir-orchestration/lib/verisim/query/vcl_executor.ex`\\n\\n**Resolved:** 2026-02-13. `verify_single_proof/1` now validates proof type, extracts contract names, and checks contract existence against the semantic store. It properly rejects queries with invalid or missing contracts for CITATION, INTEGRITY, and CUSTOM proof types. Previously it always returned `:ok`.\\n\\n**Original issue:** PROOF clauses in VCL queries were parsed but never verified. All proofs silently passed, making the entire proof system decorative.\\n\\n", - "fields": { - "type": "known_issue", - "status": "resolved", - "severity": "resolved" - } - }, - "graph": { - "relationships": [{"predicate": "documented_in", "target": "file:KNOWN-ISSUES.adoc"}] - }, - "vector": { - "embedding": [-0.328125,0.359375,0.335937,0.828125,-0.882812,0.234375,-0.203125,-0.023437,-0.187500,-0.968750,-0.976562,-0.335937,0.460937,0.195312,-0.796875,-0.203125,0.625000,0.085937,0.328125,0.765625,-0.984375,-0.851562,0.515625,-0.539062,0.765625,-0.390625,-0.007812,-0.125000,-0.765625,0.703125,-0.890625,0.945312,-0.328125,0.359375,0.335937,0.828125,-0.882812,0.234375,-0.203125,-0.023437,-0.187500,-0.968750,-0.976562,-0.335937,0.460937,0.195312,-0.796875,-0.203125,0.625000,0.085937,0.328125,0.765625,-0.984375,-0.851562,0.515625,-0.539062,0.765625,-0.390625,-0.007812,-0.125000,-0.765625,0.703125,-0.890625,0.945312], - "model": "sha256-hash-64d" - }, - "tensor": { - "shape": [1, 2], - "data": [1.0, 0.0] - }, - "semantic": { - "types": ["https://verisim.db/self/type/KnownIssue", "https://verisim.db/self/type/resolved"], - "properties": { - "status": "resolved", - "severity": "resolved" - } - }, - "temporal": { - "timestamp": "2026-02-13T16:20:00Z", - "version": 1, - "author": "self-ingest" - } -} diff --git a/verisimdb/.verisimdb/octads/issue-015.json b/verisimdb/.verisimdb/octads/issue-015.json deleted file mode 100644 index c8ba2dd6..00000000 --- a/verisimdb/.verisimdb/octads/issue-015.json +++ /dev/null @@ -1,37 +0,0 @@ -{ - "id": "issue-015", - "source": "known-issues", - "created_at": "2026-02-13T16:20:00Z", - "document": { - "title": "believe_me in Idris2 ABI Files", - "body": "\\n**Location:** `src/abi/Foreign.idr`, `debugger/src/abi/Foreign.idr`, `practice-mirror/src/abi/Foreign.idr`\\n\\n**Resolved:** 2026-02-13. FFI declarations changed from `Bits64 -> AnyPtr -> PrimIO Bits32` to `Bits64 -> (Bits64 -> Bits32 -> Bits32) -> PrimIO Bits32`, making the callback type match the declaration and eliminating the need for `believe_me` casts. Zero `believe_me` calls remain in the codebase.\\n\\n**Original issue:** `registerCallback` used `believe_me` to cast a callback function to `AnyPtr`, a BANNED unsafe pattern that bypasses the type checker.\\n\\n", - "fields": { - "type": "known_issue", - "status": "resolved", - "severity": "resolved" - } - }, - "graph": { - "relationships": [{"predicate": "documented_in", "target": "file:KNOWN-ISSUES.adoc"}] - }, - "vector": { - "embedding": [0.445312,-0.984375,-0.148437,-0.890625,-0.109375,0.695312,0.617187,-0.304687,0.687500,0.210937,-0.234375,-0.937500,0.343750,-0.601562,-0.914062,0.523437,-0.414062,0.960937,-0.625000,0.085937,-0.640625,-0.203125,0.367187,-0.132812,-0.078125,0.773437,0.445312,-1.000000,-0.789062,-0.031250,-0.703125,-0.976562,0.445312,-0.984375,-0.148437,-0.890625,-0.109375,0.695312,0.617187,-0.304687,0.687500,0.210937,-0.234375,-0.937500,0.343750,-0.601562,-0.914062,0.523437,-0.414062,0.960937,-0.625000,0.085937,-0.640625,-0.203125,0.367187,-0.132812,-0.078125,0.773437,0.445312,-1.000000,-0.789062,-0.031250,-0.703125,-0.976562], - "model": "sha256-hash-64d" - }, - "tensor": { - "shape": [1, 2], - "data": [1.0, 0.0] - }, - "semantic": { - "types": ["https://verisim.db/self/type/KnownIssue", "https://verisim.db/self/type/resolved"], - "properties": { - "status": "resolved", - "severity": "resolved" - } - }, - "temporal": { - "timestamp": "2026-02-13T16:20:00Z", - "version": 1, - "author": "self-ingest" - } -} diff --git a/verisimdb/.verisimdb/octads/issue-016.json b/verisimdb/.verisimdb/octads/issue-016.json deleted file mode 100644 index 7eeab405..00000000 --- a/verisimdb/.verisimdb/octads/issue-016.json +++ /dev/null @@ -1,37 +0,0 @@ -{ - "id": "issue-016", - "source": "known-issues", - "created_at": "2026-02-13T16:20:00Z", - "document": { - "title": "Atom Table Exhaustion Risk in VCL Bridge", - "body": "\\n**Location:** `elixir-orchestration/lib/verisim/query/vcl_bridge.ex`\\n\\n**Resolved:** 2026-02-13. All 8 `String.to_atom` calls replaced with `safe_to_atom/1` helper that uses an explicit allowlist map for the known atom values (6 modalities + 5 aggregate functions + `all`), falling back to `String.to_existing_atom/1`.\\n\\n**Original issue:** 8 `String.to_atom` calls could theoretically exhaust the BEAM atom table if fed arbitrary input, since the BEAM atom table is finite and atoms are never garbage collected.\\n\\n", - "fields": { - "type": "known_issue", - "status": "resolved", - "severity": "resolved" - } - }, - "graph": { - "relationships": [{"predicate": "documented_in", "target": "file:KNOWN-ISSUES.adoc"}] - }, - "vector": { - "embedding": [0.117187,0.296875,-0.007812,-0.335937,0.742187,-0.171875,0.320312,0.351562,-0.500000,-0.179687,0.992187,0.257812,-0.718750,-0.781250,-0.273437,-0.773437,0.046875,-0.609375,0.125000,-0.203125,-0.085937,0.078125,-0.906250,0.632812,-0.835937,-0.507812,-0.007812,-0.695312,0.562500,0.945312,0.671875,0.859375,0.117187,0.296875,-0.007812,-0.335937,0.742187,-0.171875,0.320312,0.351562,-0.500000,-0.179687,0.992187,0.257812,-0.718750,-0.781250,-0.273437,-0.773437,0.046875,-0.609375,0.125000,-0.203125,-0.085937,0.078125,-0.906250,0.632812,-0.835937,-0.507812,-0.007812,-0.695312,0.562500,0.945312,0.671875,0.859375], - "model": "sha256-hash-64d" - }, - "tensor": { - "shape": [1, 2], - "data": [1.0, 0.0] - }, - "semantic": { - "types": ["https://verisim.db/self/type/KnownIssue", "https://verisim.db/self/type/resolved"], - "properties": { - "status": "resolved", - "severity": "resolved" - } - }, - "temporal": { - "timestamp": "2026-02-13T16:20:00Z", - "version": 1, - "author": "self-ingest" - } -} diff --git a/verisimdb/.verisimdb/octads/issue-017.json b/verisimdb/.verisimdb/octads/issue-017.json deleted file mode 100644 index 683f8535..00000000 --- a/verisimdb/.verisimdb/octads/issue-017.json +++ /dev/null @@ -1,37 +0,0 @@ -{ - "id": "issue-017", - "source": "known-issues", - "created_at": "2026-02-13T16:20:01Z", - "document": { - "title": "No ETS Caching in RustClient", - "body": "\\n**Location:** `elixir-orchestration/lib/verisim/rust_client.ex`\\n\\n**Resolved:** 2026-02-13. Added ETS-based read-through cache with configurable TTL (30s default for octads, 10s for drift scores). Cache is invalidated on writes (update/delete). Provides `init_cache/0`, `clear_cache/0`, and `invalidate_cache/1` for cache management.\\n\\n**Original issue:** Every RustClient call made a fresh HTTP request to the Rust core, with no caching. Repeated reads of the same octad in a query pipeline (e.g., cross-modal evaluation fetching the same entity multiple times) each incurred full HTTP round-trip latency.\\n\\n", - "fields": { - "type": "known_issue", - "status": "resolved", - "severity": "resolved" - } - }, - "graph": { - "relationships": [{"predicate": "documented_in", "target": "file:KNOWN-ISSUES.adoc"}] - }, - "vector": { - "embedding": [-0.914062,-0.343750,-0.382812,0.078125,-0.304687,-0.640625,-0.117187,0.398437,0.468750,0.429687,0.906250,0.570312,-0.695312,-0.140625,-0.523437,0.804687,-0.234375,-0.960937,0.601562,-0.046875,0.585937,0.007812,0.421875,-0.242187,0.359375,-0.437500,-0.640625,0.875000,0.843750,-0.640625,0.734375,0.296875,-0.914062,-0.343750,-0.382812,0.078125,-0.304687,-0.640625,-0.117187,0.398437,0.468750,0.429687,0.906250,0.570312,-0.695312,-0.140625,-0.523437,0.804687,-0.234375,-0.960937,0.601562,-0.046875,0.585937,0.007812,0.421875,-0.242187,0.359375,-0.437500,-0.640625,0.875000,0.843750,-0.640625,0.734375,0.296875], - "model": "sha256-hash-64d" - }, - "tensor": { - "shape": [1, 2], - "data": [1.0, 0.0] - }, - "semantic": { - "types": ["https://verisim.db/self/type/KnownIssue", "https://verisim.db/self/type/resolved"], - "properties": { - "status": "resolved", - "severity": "resolved" - } - }, - "temporal": { - "timestamp": "2026-02-13T16:20:01Z", - "version": 1, - "author": "self-ingest" - } -} diff --git a/verisimdb/.verisimdb/octads/issue-018.json b/verisimdb/.verisimdb/octads/issue-018.json deleted file mode 100644 index 3d21c3cf..00000000 --- a/verisimdb/.verisimdb/octads/issue-018.json +++ /dev/null @@ -1,37 +0,0 @@ -{ - "id": "issue-018", - "source": "known-issues", - "created_at": "2026-02-13T16:20:01Z", - "document": { - "title": "EXPLAIN Returns Hardcoded Plan", - "body": "\\n**Location:** `elixir-orchestration/lib/verisim/query/vcl_executor.ex`\\n\\n**Resolved:** 2026-02-13. `generate_explain_plan/1` now analyzes the actual query AST to produce cost estimates based on source type, modality count, WHERE clause complexity, cross-modal conditions, GROUP BY presence, and proof obligations. Delegates to the Rust verisim-planner API when available, with local estimation as fallback.\\n\\n**Original issue:** EXPLAIN returned a static plan with hardcoded step names and costs totalling 71ms, regardless of the actual query structure.\\n\\n", - "fields": { - "type": "known_issue", - "status": "resolved", - "severity": "resolved" - } - }, - "graph": { - "relationships": [{"predicate": "documented_in", "target": "file:KNOWN-ISSUES.adoc"}] - }, - "vector": { - "embedding": [-0.234375,0.742187,0.054687,0.171875,-0.062500,-0.468750,0.843750,-0.226562,0.765625,0.882812,-0.765625,0.648437,0.304687,0.210937,0.046875,0.734375,-0.375000,-0.429687,0.484375,0.929687,0.570312,-0.757812,-0.781250,0.781250,0.281250,0.992187,-0.804687,0.750000,0.851562,0.117187,0.164062,-0.750000,-0.234375,0.742187,0.054687,0.171875,-0.062500,-0.468750,0.843750,-0.226562,0.765625,0.882812,-0.765625,0.648437,0.304687,0.210937,0.046875,0.734375,-0.375000,-0.429687,0.484375,0.929687,0.570312,-0.757812,-0.781250,0.781250,0.281250,0.992187,-0.804687,0.750000,0.851562,0.117187,0.164062,-0.750000], - "model": "sha256-hash-64d" - }, - "tensor": { - "shape": [1, 2], - "data": [1.0, 0.0] - }, - "semantic": { - "types": ["https://verisim.db/self/type/KnownIssue", "https://verisim.db/self/type/resolved"], - "properties": { - "status": "resolved", - "severity": "resolved" - } - }, - "temporal": { - "timestamp": "2026-02-13T16:20:01Z", - "version": 1, - "author": "self-ingest" - } -} diff --git a/verisimdb/.verisimdb/octads/issue-019.json b/verisimdb/.verisimdb/octads/issue-019.json deleted file mode 100644 index bbf8d986..00000000 --- a/verisimdb/.verisimdb/octads/issue-019.json +++ /dev/null @@ -1,37 +0,0 @@ -{ - "id": "issue-019", - "source": "known-issues", - "created_at": "2026-02-13T16:20:01Z", - "document": { - "title": "VCL Executor Federation Stub", - "body": "\\n**Location:** `elixir-orchestration/lib/verisim/query/vcl_executor.ex`, `rust-core/verisim-api/src/federation.rs`\\n\\n**Resolved:** 2026-02-13. Added `GET /octads` list endpoint with `?limit=N&offset=M` pagination to the Rust API. The `OctadStore` trait now includes a `list(limit, offset)` method. Federation `query_single_peer()` now falls back to the `/octads` list endpoint when neither text_query nor vector_query is provided, so bare `SELECT * FROM FEDERATION /pattern/*` returns actual octad data instead of empty results.\\n\\n", - "fields": { - "type": "known_issue", - "status": "resolved", - "severity": "resolved" - } - }, - "graph": { - "relationships": [{"predicate": "documented_in", "target": "file:KNOWN-ISSUES.adoc"}] - }, - "vector": { - "embedding": [-0.648437,0.781250,-0.554687,0.281250,-0.757812,-0.359375,0.664062,-0.570312,0.867187,-0.062500,-0.867187,0.851562,-0.617187,0.054687,-0.648437,0.703125,0.554687,-0.984375,-0.507812,-0.078125,-0.250000,0.203125,-0.843750,0.101562,0.453125,-0.023437,0.531250,-0.953125,0.867187,0.296875,0.125000,-0.335937,-0.648437,0.781250,-0.554687,0.281250,-0.757812,-0.359375,0.664062,-0.570312,0.867187,-0.062500,-0.867187,0.851562,-0.617187,0.054687,-0.648437,0.703125,0.554687,-0.984375,-0.507812,-0.078125,-0.250000,0.203125,-0.843750,0.101562,0.453125,-0.023437,0.531250,-0.953125,0.867187,0.296875,0.125000,-0.335937], - "model": "sha256-hash-64d" - }, - "tensor": { - "shape": [1, 2], - "data": [1.0, 0.0] - }, - "semantic": { - "types": ["https://verisim.db/self/type/KnownIssue", "https://verisim.db/self/type/resolved"], - "properties": { - "status": "resolved", - "severity": "resolved" - } - }, - "temporal": { - "timestamp": "2026-02-13T16:20:01Z", - "version": 1, - "author": "self-ingest" - } -} diff --git a/verisimdb/.verisimdb/octads/issue-020.json b/verisimdb/.verisimdb/octads/issue-020.json deleted file mode 100644 index bb3ce31e..00000000 --- a/verisimdb/.verisimdb/octads/issue-020.json +++ /dev/null @@ -1,37 +0,0 @@ -{ - "id": "issue-020", - "source": "known-issues", - "created_at": "2026-02-13T16:20:01Z", - "document": { - "title": "Custom Circuit Hardcoding", - "body": "\\n**Location:** `rust-core/verisim-semantic/src/circuit_registry.rs`, `circuit_compiler.rs`, `verification_keys.rs`, `src/vcl/VCLCircuit.res`\\n\\n**Resolved:** 2026-02-13. Full custom circuit infrastructure implemented:\\n\\n- Circuit Registry: named circuit storage with register/get/verify/list/unregister operations\\n- Circuit Compiler: DSL gates (AND, OR, XOR, NOT, LinearCombination) compiled to R1CS constraints\\n- Verification Key Store: per-circuit keys with rotation support and federation export/import\\n- VCL Circuit DSL: ReScript types for circuit definition with `PROOF CUSTOM \"name\" WITH (param=value)`\\n- 25 semantic tests pass covering circuit operations\\n\\n**Original issue:** The `CUSTOM` proof type used a hardcoded circuit configuration with no DSL, compiler, or registry.\\n\\n", - "fields": { - "type": "known_issue", - "status": "resolved", - "severity": "resolved" - } - }, - "graph": { - "relationships": [{"predicate": "documented_in", "target": "file:KNOWN-ISSUES.adoc"}] - }, - "vector": { - "embedding": [-0.828125,0.257812,0.218750,-0.640625,0.914062,0.265625,0.523437,0.531250,0.328125,0.078125,-0.101562,0.015625,-0.429687,0.898437,-0.242187,0.492187,-0.062500,0.187500,0.984375,0.718750,-0.132812,-1.000000,0.023437,-0.171875,0.234375,0.031250,0.507812,-0.140625,0.640625,0.789062,-0.843750,0.609375,-0.828125,0.257812,0.218750,-0.640625,0.914062,0.265625,0.523437,0.531250,0.328125,0.078125,-0.101562,0.015625,-0.429687,0.898437,-0.242187,0.492187,-0.062500,0.187500,0.984375,0.718750,-0.132812,-1.000000,0.023437,-0.171875,0.234375,0.031250,0.507812,-0.140625,0.640625,0.789062,-0.843750,0.609375], - "model": "sha256-hash-64d" - }, - "tensor": { - "shape": [1, 2], - "data": [1.0, 0.0] - }, - "semantic": { - "types": ["https://verisim.db/self/type/KnownIssue", "https://verisim.db/self/type/resolved"], - "properties": { - "status": "resolved", - "severity": "resolved" - } - }, - "temporal": { - "timestamp": "2026-02-13T16:20:01Z", - "version": 1, - "author": "self-ingest" - } -} diff --git a/verisimdb/.verisimdb/octads/issue-021.json b/verisimdb/.verisimdb/octads/issue-021.json deleted file mode 100644 index 98ea420e..00000000 --- a/verisimdb/.verisimdb/octads/issue-021.json +++ /dev/null @@ -1,37 +0,0 @@ -{ - "id": "issue-021", - "source": "known-issues", - "created_at": "2026-02-13T16:20:02Z", - "document": { - "title": "VCL-UT Not Connected to VCL PROOF Runtime", - "body": "\\n**Location:** `vcl-dt/` (Lean type definitions), VCL query router\\n\\n**Current state:** Lean type definitions exist specifying dependent type contracts for all six PROOF types. The VCL parser parses `PROOF` clauses and the executor routes proofs through `verify_single_proof/1` which validates contract existence. However, the Lean type checker itself is not invoked at runtime. VCL-UT queries execute but proof certificates are not formally verified by the Lean checker.\\n\\n**Impact:** PROOF clauses are semantically checked (contract existence, parameter validation) but not formally verified via the Lean dependent type checker. Custom ZKP circuits work (R1CS verification), but the VCL-UT formal path is not wired.\\n\\n**Resolution:** Wire the Lean type checker into the VCL-UT execution path. See link:docs/vcl-vs-vcl-dt.adoc[VCL Slipstream vs VCL-UT].\\n\\n", - "fields": { - "type": "known_issue", - "status": "open", - "severity": "medium" - } - }, - "graph": { - "relationships": [{"predicate": "documented_in", "target": "file:KNOWN-ISSUES.adoc"}] - }, - "vector": { - "embedding": [0.882812,0.218750,0.179687,-0.609375,0.796875,-0.476562,-0.171875,-0.328125,0.679687,0.929687,0.976562,0.726562,-0.710937,-0.070312,-0.703125,0.375000,0.734375,-0.156250,0.578125,-0.921875,0.312500,0.164062,-0.445312,-0.757812,0.773437,0.484375,-0.437500,0.148437,-0.687500,0.398437,0.156250,0.546875,0.882812,0.218750,0.179687,-0.609375,0.796875,-0.476562,-0.171875,-0.328125,0.679687,0.929687,0.976562,0.726562,-0.710937,-0.070312,-0.703125,0.375000,0.734375,-0.156250,0.578125,-0.921875,0.312500,0.164062,-0.445312,-0.757812,0.773437,0.484375,-0.437500,0.148437,-0.687500,0.398437,0.156250,0.546875], - "model": "sha256-hash-64d" - }, - "tensor": { - "shape": [1, 2], - "data": [0.0, 1.0] - }, - "semantic": { - "types": ["https://verisim.db/self/type/KnownIssue", "https://verisim.db/self/type/open"], - "properties": { - "status": "open", - "severity": "medium" - } - }, - "temporal": { - "timestamp": "2026-02-13T16:20:02Z", - "version": 1, - "author": "self-ingest" - } -} diff --git a/verisimdb/.verisimdb/octads/issue-022.json b/verisimdb/.verisimdb/octads/issue-022.json deleted file mode 100644 index 5f22d6d1..00000000 --- a/verisimdb/.verisimdb/octads/issue-022.json +++ /dev/null @@ -1,37 +0,0 @@ -{ - "id": "issue-022", - "source": "known-issues", - "created_at": "2026-02-13T16:20:02Z", - "document": { - "title": "proven Library Not Integrated", - "body": "\\n**Location:** Design docs only\\n\\n**Current state:** Custom ZKP circuits work via R1CS constraint systems with in-process verification. However, the `proven` library for generating actual zero-knowledge proofs (where a verifier can confirm a property without seeing the data) is not integrated.\\n\\n**Impact:** Privacy-preserving proofs are not available. Current circuit verification requires the witness (private input) to be present, which defeats the zero-knowledge property.\\n\\n**Resolution:** Integrate proven library after VCL-UT (issue 21) is resolved.\\n\\n", - "fields": { - "type": "known_issue", - "status": "open", - "severity": "medium" - } - }, - "graph": { - "relationships": [{"predicate": "documented_in", "target": "file:KNOWN-ISSUES.adoc"}] - }, - "vector": { - "embedding": [0.320312,0,-0.164062,0.171875,-0.179687,-0.679687,-0.937500,-0.484375,-0.703125,-0.820312,0.101562,0.320312,0.265625,-0.500000,0.601562,-0.765625,0.945312,0.804687,-0.125000,-0.687500,-0.742187,-0.335937,0.296875,-0.414062,0.695312,-0.585937,0.585937,0.265625,-0.523437,-0.992187,-0.992187,-0.671875,0.320312,0,-0.164062,0.171875,-0.179687,-0.679687,-0.937500,-0.484375,-0.703125,-0.820312,0.101562,0.320312,0.265625,-0.500000,0.601562,-0.765625,0.945312,0.804687,-0.125000,-0.687500,-0.742187,-0.335937,0.296875,-0.414062,0.695312,-0.585937,0.585937,0.265625,-0.523437,-0.992187,-0.992187,-0.671875], - "model": "sha256-hash-64d" - }, - "tensor": { - "shape": [1, 2], - "data": [0.0, 1.0] - }, - "semantic": { - "types": ["https://verisim.db/self/type/KnownIssue", "https://verisim.db/self/type/open"], - "properties": { - "status": "open", - "severity": "medium" - } - }, - "temporal": { - "timestamp": "2026-02-13T16:20:02Z", - "version": 1, - "author": "self-ingest" - } -} diff --git a/verisimdb/.verisimdb/octads/issue-023.json b/verisimdb/.verisimdb/octads/issue-023.json deleted file mode 100644 index cc134198..00000000 --- a/verisimdb/.verisimdb/octads/issue-023.json +++ /dev/null @@ -1,37 +0,0 @@ -{ - "id": "issue-023", - "source": "known-issues", - "created_at": "2026-02-13T16:20:02Z", - "document": { - "title": "verisim-repl Has Build Issues", - "body": "\\n**Location:** `rust-core/verisim-repl/`\\n\\n**Current state:** The REPL crate was added by a linter with rustyline API incompatibilities (`Completer`, `Highlighter`, `Hinter`, `Validator` traits are private in the installed version). The workspace builds successfully with `--exclude verisim-repl`.\\n\\n**Impact:** Low. The REPL is a developer convenience tool, not a core component.\\n\\n**Resolution:** Update rustyline API calls to match installed version, or pin a compatible version.\\n", - "fields": { - "type": "known_issue", - "status": "open", - "severity": "medium" - } - }, - "graph": { - "relationships": [{"predicate": "documented_in", "target": "file:KNOWN-ISSUES.adoc"}] - }, - "vector": { - "embedding": [0.742187,0.179687,-0.164062,0.296875,0.406250,-0.156250,0.132812,-0.476562,0.429687,0.726562,0.546875,-0.882812,-0.218750,-0.539062,-0.312500,-0.156250,-0.828125,-0.906250,-0.867187,-0.546875,-0.546875,0.820312,-0.921875,-0.960937,-0.812500,0.195312,0.750000,-0.328125,-0.226562,-0.781250,-0.804687,-0.687500,0.742187,0.179687,-0.164062,0.296875,0.406250,-0.156250,0.132812,-0.476562,0.429687,0.726562,0.546875,-0.882812,-0.218750,-0.539062,-0.312500,-0.156250,-0.828125,-0.906250,-0.867187,-0.546875,-0.546875,0.820312,-0.921875,-0.960937,-0.812500,0.195312,0.750000,-0.328125,-0.226562,-0.781250,-0.804687,-0.687500], - "model": "sha256-hash-64d" - }, - "tensor": { - "shape": [1, 2], - "data": [0.0, 1.0] - }, - "semantic": { - "types": ["https://verisim.db/self/type/KnownIssue", "https://verisim.db/self/type/open"], - "properties": { - "status": "open", - "severity": "medium" - } - }, - "temporal": { - "timestamp": "2026-02-13T16:20:02Z", - "version": 1, - "author": "self-ingest" - } -} diff --git a/verisimdb/.well-known/groove/manifest.json b/verisimdb/.well-known/groove/manifest.json deleted file mode 100644 index 812f6832..00000000 --- a/verisimdb/.well-known/groove/manifest.json +++ /dev/null @@ -1,39 +0,0 @@ -{ - "groove_version": "1", - "service_id": "verisimdb", - "service_version": "0.1.0", - "capabilities": { - "octad-storage": { - "type": "octad-storage", - "description": "Cross-modal entity consistency via octad structures", - "protocol": "http", - "endpoint": "/octads", - "requires_auth": false, - "panel_compatible": true - }, - "drift-detection": { - "type": "drift-detection", - "description": "Detect entity drift across temporal versions", - "protocol": "http", - "endpoint": "/api/drift", - "requires_auth": false, - "panel_compatible": true - }, - "temporal-versioning": { - "type": "temporal-versioning", - "description": "Temporal versioning with audit trail", - "protocol": "http", - "endpoint": "/audit", - "requires_auth": false, - "panel_compatible": true - } - }, - "consumes": ["scanning"], - "endpoints": { - "health": "/health", - "groove": "/.well-known/groove", - "grpc": "grpc://localhost:50051" - }, - "health": "/health", - "applicability": ["individual", "team", "massive-open"] -} diff --git a/verisimdb/.well-known/void.rdf b/verisimdb/.well-known/void.rdf deleted file mode 100644 index a2b3dc53..00000000 --- a/verisimdb/.well-known/void.rdf +++ /dev/null @@ -1,16 +0,0 @@ - - - - - verisimdb Dataset - Linked data from verisimdb project - - - - - - diff --git a/verisimdb/.well-known/void.ttl b/verisimdb/.well-known/void.ttl deleted file mode 100644 index 57ce0131..00000000 --- a/verisimdb/.well-known/void.ttl +++ /dev/null @@ -1,75 +0,0 @@ -@prefix void: . -@prefix rdf: . -@prefix rdfs: . -@prefix owl: . -@prefix dcterms: . -@prefix foaf: . -@prefix xsd: . - -# Dataset Description - a void:Dataset ; - dcterms:title "verisimdb Dataset" ; - dcterms:description "Linked data from verisimdb project" ; - dcterms:creator ; - dcterms:publisher ; - dcterms:license ; - dcterms:created "2026-01-31T16:49:57Z"^^xsd:dateTime ; - dcterms:modified "2026-01-31T16:49:57Z"^^xsd:dateTime ; - - # Dataset statistics (update these based on actual data) - void:triples 0 ; - void:entities 0 ; - void:distinctSubjects 0 ; - void:distinctObjects 0 ; - - # Access methods - void:sparqlEndpoint ; - void:dataDump ; - void:dataDump ; - void:dataDump ; - - # Technical details - void:feature ; - void:feature ; - void:feature ; - - # Vocabulary usage (customize based on your data model) - void:vocabulary ; - void:vocabulary ; - void:vocabulary ; - void:vocabulary ; - - # Example linksets (connections to other datasets) - # void:subset ; - # void:subset ; -. - -# Creator information - a foaf:Person ; - foaf:name "Jonathan D.A. Jewell" ; - foaf:mbox ; - foaf:homepage ; - foaf:account ; -. - -# Publisher information - a foaf:Organization ; - foaf:name "hyperpolymath" ; - foaf:homepage ; -. - -# Example linkset to DBpedia (uncomment and customize) -# a void:Linkset ; -# void:linkPredicate owl:sameAs ; -# void:target ; -# void:target ; -# void:triples 0 ; -# . - -# Example linkset to Wikidata (uncomment and customize) -# a void:Linkset ; -# void:linkPredicate owl:sameAs ; -# void:target ; -# void:target ; -# void:triples 0 ; -# . diff --git a/verisimdb/0-AI-MANIFEST.a2ml b/verisimdb/0-AI-MANIFEST.a2ml deleted file mode 100644 index 187964f2..00000000 --- a/verisimdb/0-AI-MANIFEST.a2ml +++ /dev/null @@ -1,146 +0,0 @@ -; SPDX-License-Identifier: MPL-2.0 -; SPDX-FileCopyrightText: 2026 Jonathan D.A. Jewell (hyperpolymath) -; -; 0-AI-MANIFEST.a2ml — Universal AI entry point for VeriSimDB -; Media-Type: application/a2ml - -(manifest - (identity - (name "VeriSimDB") - (full-name "Veridical Simulacrum Database") - (version "0.1.0-alpha") - (repo "https://github.com/hyperpolymath/verisimdb") - (license "PMPL-1.0-or-later") - (author "Jonathan D.A. Jewell ") - (monorepo-parent "nextgen-databases")) - - (purpose - "8-modality (octad) database with self-normalization. Each entity - exists simultaneously across 8 modalities — Graph, Vector, Tensor, - Semantic, Document, Temporal, Provenance, Spatial — with drift - detection and automatic consistency maintenance.") - - (canonical-locations - (ai-instructions ".claude/CLAUDE.md") - (state ".machine_readable/STATE.scm") - (meta ".machine_readable/META.scm") - (ecosystem ".machine_readable/ECOSYSTEM.scm") - (topology "TOPOLOGY.md") - (build "justfile") - (container-build "container/Containerfile") - (container-deploy "container/compose.toml") - (test-infra "connectors/test-infra/compose.toml") - (test-infra-manifest "connectors/test-infra/0-AI-MANIFEST.a2ml") - (integration-tests "elixir-orchestration/test/verisim/federation/adapters/integration/") - (known-issues "KNOWN-ISSUES.adoc") - (changelog "CHANGELOG.adoc") - (deployment "DEPLOYMENT.adoc") - (security ".well-known/security.txt")) - - (tech-stack - (primary "Rust" "Elixir/OTP") - (secondary "ReScript" "Idris2" "Zig") - (query-language "VCL") - (container-runtime "Podman") - (base-image "cgr.dev/chainguard/wolfi-base:latest")) - - (architecture - (rust-core - (description "Performance-critical modality stores") - (location "rust-core/") - (crates "verisim-graph" "verisim-vector" "verisim-tensor" - "verisim-semantic" "verisim-document" "verisim-temporal" - "verisim-octad" "verisim-drift" "verisim-normalizer" - "verisim-api")) - (elixir-otp - (description "Distributed coordination and supervision") - (location "elixir-orchestration/")) - (abi-ffi - (description "Idris2 ABI definitions + Zig FFI bridge") - (abi-location "src/abi/") - (ffi-location "ffi/zig/")) - (data-store - (description "Git-backed flat-file data repo") - (location "verisimdb-data/")) - (test-infra - (description "Containerised test stack for federation adapter integration testing") - (location "connectors/test-infra/") - (services "mongodb" "redis-stack" "neo4j" "clickhouse" "surrealdb" "influxdb" "minio")) - (connectors - (description "Federation adapters (10) and client SDKs (6)") - (location "connectors/") - (adapters "MongoDB" "Redis" "DuckDB" "ClickHouse" "SurrealDB" "SQLite" "Neo4j" "VectorDB" "InfluxDB" "ObjectStorage") - (sdks "Rust" "V" "Elixir" "ReScript" "Julia" "Gleam"))) - - (critical-invariants - (rule "SCM files ONLY in .machine_readable/ — never root") - (rule "All shell scripts must validate untrusted input") - (rule "No hardcoded secrets — use env vars with ${VAR:-} defaults") - (rule "Container images MUST use Chainguard base (cgr.dev)") - (rule "Container runtime is Podman — never Docker") - (rule "VCL is the query language — never raw SQL") - (rule "Octad entities must maintain 8-modality consistency") - (rule "Drift thresholds gate all automatic normalization") - - ;; INSTANCE POLICY — MOST IMPORTANT RULE FOR CONSUMERS - (rule "NEVER store application data in this repository's VeriSimDB instance") - (rule "NEVER point application code at localhost:8080 expecting a shared VeriSimDB") - (rule "Each consuming project MUST run its OWN dedicated VeriSimDB instance") - (rule "Each instance MUST have its own port, data directory, and container volume") - (rule "Copy the client SDK files into your project — do NOT import from this repo") - (rule "This repo contains ONLY source code and EXAMPLE data — not live application data")) - - ;; ═══════════════════════════════════════════════════════════════════ - ;; DEPLOYMENT MODEL — READ THIS BEFORE INTEGRATING - ;; ═══════════════════════════════════════════════════════════════════ - ;; - ;; This repository is the VeriSimDB SOURCE CODE and EXAMPLE DATA. - ;; It is NOT a shared database instance. It is NOT a data store. - ;; - ;; If your project needs VeriSimDB: - ;; 1. Copy the client SDK (connectors/clients//) into your project - ;; 2. Add a VeriSimDB Containerfile to YOUR project's container stack - ;; 3. Choose a UNIQUE port for your instance (not 8080) - ;; 4. Create a DEDICATED data volume for your instance - ;; 5. Configure YOUR client to point at YOUR instance - ;; - ;; Example per-project instance setup: - ;; IDApTIK (game saves) → port 8090, volume idaptik-verisimdb-data - ;; Burble (voice audit) → port 8091, volume burble-verisimdb-data - ;; Hypatia (scan results) → port 8092, volume hypatia-verisimdb-data - ;; - ;; The examples/ directory contains EXAMPLE data for testing and demos. - ;; Do NOT treat it as a database. Do NOT store application data here. - ;; ═══════════════════════════════════════════════════════════════════ - (deployment-model - (type "source-code-and-examples-only") - (not "shared-database-instance") - (not "central-data-store") - (consumer-pattern "copy-sdk-run-own-instance")) - - (container-ecosystem - (selur "container/compose.toml — deployment orchestration") - (stapeln "stapeln.toml — layer-based container builds") - (svalinn "TLS gateway with policy enforcement") - (vordr "Runtime verification and formal proof checking") - (cerro-torre "Image signing with ML-DSA-87 post-quantum crypto") - (rokur "Secret rotation with argon2id")) - - (related-projects - (hypatia "Neurosymbolic CI/CD scanner — consumes verisimdb data") - (gitbot-fleet "Bot orchestration — dispatches fixes from verisimdb findings") - (panic-attacker "Static analysis scanner — produces scan data for verisimdb") - (proven "Formally verified safety library — provides SafeString, SafeJson, etc.") - (lithoglyph "Graph database sibling — shares GQL/GQL-DT patterns") - (quandledb "Knot-theoretic database sibling — KQL query language"))) - -## Taxonomy Index - -- `spec/grammar.ebnf` — @taxonomy: spec/grammar -- `spec/README.adoc` — @taxonomy: spec/index -- `verification/README.adoc` — @taxonomy: verification/index - -### New RSR Standard Directories - -- `spec/` — Canonical specification files -- `verification/` — Unified verification gateway (symlinks to proofs, tests, conformance, benchmarks, fuzzing) diff --git a/verisimdb/ABI-FFI-README.md b/verisimdb/ABI-FFI-README.md deleted file mode 100644 index e6a32bbf..00000000 --- a/verisimdb/ABI-FFI-README.md +++ /dev/null @@ -1,385 +0,0 @@ -{{~ Aditionally delete this line and fill out the template below ~}} - -# {{PROJECT}} ABI/FFI Documentation - -## Overview - -This library follows the **Hyperpolymath RSR Standard** for ABI and FFI design: - -- **ABI (Application Binary Interface)** defined in **Idris2** with formal proofs -- **FFI (Foreign Function Interface)** implemented in **Zig** for C compatibility -- **Generated C headers** bridge Idris2 ABI to Zig FFI -- **Any language** can call through standard C ABI - -## Architecture - -``` -┌─────────────────────────────────────────────┐ -│ ABI Definitions (Idris2) │ -│ src/abi/ │ -│ - Types.idr (Type definitions) │ -│ - Layout.idr (Memory layout proofs) │ -│ - Foreign.idr (FFI declarations) │ -└─────────────────┬───────────────────────────┘ - │ - │ generates (at compile time) - ▼ -┌─────────────────────────────────────────────┐ -│ C Headers (auto-generated) │ -│ generated/abi/{{project}}.h │ -└─────────────────┬───────────────────────────┘ - │ - │ imported by - ▼ -┌─────────────────────────────────────────────┐ -│ FFI Implementation (Zig) │ -│ ffi/zig/src/main.zig │ -│ - Implements C-compatible functions │ -│ - Zero-cost abstractions │ -│ - Memory-safe by default │ -└─────────────────┬───────────────────────────┘ - │ - │ compiled to lib{{project}}.so/.a - ▼ -┌─────────────────────────────────────────────┐ -│ Any Language via C ABI │ -│ - Rust, ReScript, Julia, Python, etc. │ -└─────────────────────────────────────────────┘ -``` - -## Directory Structure - -``` -{{project}}/ -├── src/ -│ ├── abi/ # ABI definitions (Idris2) -│ │ ├── Types.idr # Core type definitions with proofs -│ │ ├── Layout.idr # Memory layout verification -│ │ └── Foreign.idr # FFI function declarations -│ └── lib/ # Core library (any language) -│ -├── ffi/ -│ └── zig/ # FFI implementation (Zig) -│ ├── build.zig # Build configuration -│ ├── build.zig.zon # Dependencies -│ ├── src/ -│ │ └── main.zig # C-compatible FFI implementation -│ ├── test/ -│ │ └── integration_test.zig -│ └── include/ -│ └── {{project}}.h # C header (optional, can be generated) -│ -├── generated/ # Auto-generated files -│ └── abi/ -│ └── {{project}}.h # Generated from Idris2 ABI -│ -└── bindings/ # Language-specific wrappers (optional) - ├── rust/ - ├── rescript/ - └── julia/ -``` - -## Why Idris2 for ABI? - -### 1. **Formal Verification** - -Idris2's dependent types allow proving properties about the ABI at compile-time: - -```idris --- Prove struct size is correct -public export -exampleStructSize : HasSize ExampleStruct 16 - --- Prove field alignment is correct -public export -fieldAligned : Divides 8 (offsetOf ExampleStruct.field) - --- Prove ABI is platform-compatible -public export -abiCompatible : Compatible (ABI 1) (ABI 2) -``` - -### 2. **Type Safety** - -Encode invariants that C/Zig cannot express: - -```idris --- Non-null pointer guaranteed at type level -data Handle : Type where - MkHandle : (ptr : Bits64) -> {auto 0 nonNull : So (ptr /= 0)} -> Handle - --- Array with length proof -data Buffer : (n : Nat) -> Type where - MkBuffer : Vect n Byte -> Buffer n -``` - -### 3. **Platform Abstraction** - -Platform-specific types with compile-time selection: - -```idris -CInt : Platform -> Type -CInt Linux = Bits32 -CInt Windows = Bits32 - -CSize : Platform -> Type -CSize Linux = Bits64 -CSize Windows = Bits64 -``` - -### 4. **Safe Evolution** - -Prove that new ABI versions are backward-compatible: - -```idris --- Compiler enforces compatibility -abiUpgrade : ABI 1 -> ABI 2 -abiUpgrade old = MkABI2 { - -- Must preserve all v1 fields - v1_compat = old, - -- Can add new fields - new_features = defaults -} -``` - -## Why Zig for FFI? - -### 1. **C ABI Compatibility** - -Zig exports C-compatible functions naturally: - -```zig -export fn library_function(param: i32) i32 { - return param * 2; -} -``` - -### 2. **Memory Safety** - -Compile-time safety without runtime overhead: - -```zig -// Null check enforced at compile time -const handle = init() orelse return error.InitFailed; -defer free(handle); -``` - -### 3. **Cross-Compilation** - -Built-in cross-compilation to any platform: - -```bash -zig build -Dtarget=x86_64-linux -zig build -Dtarget=aarch64-macos -zig build -Dtarget=x86_64-windows -``` - -### 4. **Zero Dependencies** - -No runtime, no libc required (unless explicitly needed): - -```zig -// Minimal binary size -pub const lib = @import("std"); -// Only includes what you use -``` - -## Building - -### Build FFI Library - -```bash -cd ffi/zig -zig build # Build debug -zig build -Doptimize=ReleaseFast # Build optimized -zig build test # Run tests -``` - -### Generate C Header from Idris2 ABI - -```bash -cd src/abi -idris2 --cg c-header Types.idr -o ../../generated/abi/{{project}}.h -``` - -### Cross-Compile - -```bash -cd ffi/zig - -# Linux x86_64 -zig build -Dtarget=x86_64-linux - -# macOS ARM64 -zig build -Dtarget=aarch64-macos - -# Windows x86_64 -zig build -Dtarget=x86_64-windows -``` - -## Usage - -### From C - -```c -#include "{{project}}.h" - -int main() { - void* handle = {{project}}_init(); - if (!handle) return 1; - - int result = {{project}}_process(handle, 42); - if (result != 0) { - const char* err = {{project}}_last_error(); - fprintf(stderr, "Error: %s\n", err); - } - - {{project}}_free(handle); - return 0; -} -``` - -Compile with: -```bash -gcc -o example example.c -l{{project}} -L./zig-out/lib -``` - -### From Idris2 - -```idris -import {{PROJECT}}.ABI.Foreign - -main : IO () -main = do - Just handle <- init - | Nothing => putStrLn "Failed to initialize" - - Right result <- process handle 42 - | Left err => putStrLn $ "Error: " ++ errorDescription err - - free handle - putStrLn "Success" -``` - -### From Rust - -```rust -#[link(name = "{{project}}")] -extern "C" { - fn {{project}}_init() -> *mut std::ffi::c_void; - fn {{project}}_free(handle: *mut std::ffi::c_void); - fn {{project}}_process(handle: *mut std::ffi::c_void, input: u32) -> i32; -} - -fn main() { - unsafe { - let handle = {{project}}_init(); - assert!(!handle.is_null()); - - let result = {{project}}_process(handle, 42); - assert_eq!(result, 0); - - {{project}}_free(handle); - } -} -``` - -### From Julia - -```julia -const lib{{project}} = "lib{{project}}" - -function init() - handle = ccall((:{{project}}_init, lib{{project}}), Ptr{Cvoid}, ()) - handle == C_NULL && error("Failed to initialize") - handle -end - -function process(handle, input) - result = ccall((:{{project}}_process, lib{{project}}), Cint, (Ptr{Cvoid}, UInt32), handle, input) - result -end - -function cleanup(handle) - ccall((:{{project}}_free, lib{{project}}), Cvoid, (Ptr{Cvoid},), handle) -end - -# Usage -handle = init() -try - result = process(handle, 42) - println("Result: $result") -finally - cleanup(handle) -end -``` - -## Testing - -### Unit Tests (Zig) - -```bash -cd ffi/zig -zig build test -``` - -### Integration Tests - -```bash -cd ffi/zig -zig build test-integration -``` - -### ABI Verification (Idris2) - -```idris --- Compile-time verification -%runElab verifyABI - --- Runtime checks -main : IO () -main = do - verifyLayoutsCorrect - verifyAlignmentsCorrect - putStrLn "ABI verification passed" -``` - -## Contributing - -When modifying the ABI/FFI: - -1. **Update ABI first** (`src/abi/*.idr`) - - Modify type definitions - - Update proofs - - Ensure backward compatibility - -2. **Generate C header** - ```bash - idris2 --cg c-header src/abi/Types.idr -o generated/abi/{{project}}.h - ``` - -3. **Update FFI implementation** (`ffi/zig/src/main.zig`) - - Implement new functions - - Match ABI types exactly - -4. **Add tests** - - Unit tests in Zig - - Integration tests - - ABI verification tests - -5. **Update documentation** - - Function signatures - - Usage examples - - Migration guide (if breaking changes) - -## License - -PMPL-1.0-or-later - -## See Also - -- [Idris2 Documentation](https://idris2.readthedocs.io) -- [Zig Documentation](https://ziglang.org/documentation/master/) -- [Rhodium Standard Repositories](https://github.com/hyperpolymath/rhodium-standard-repositories) -- [FFI Migration Guide](../ffi-migration-guide.md) -- [ABI Migration Guide](../abi-migration-guide.md) diff --git a/verisimdb/BEST-IN-CLASS-ROADMAP.md b/verisimdb/BEST-IN-CLASS-ROADMAP.md deleted file mode 100644 index 7632b0c7..00000000 --- a/verisimdb/BEST-IN-CLASS-ROADMAP.md +++ /dev/null @@ -1,354 +0,0 @@ -# VeriSimDB Best-in-Class Roadmap - -Criticality-ordered plan. Toolchain first, then infrastructure, then ecosystem. - -**Completed prerequisites** (this session): -- [x] verisim-planner crate (cost-based query planning) -- [x] Triple API (REST + GraphQL + gRPC) -- [x] VCL AST → LogicalPlan bridge -- [x] Proof obligation costing (per-type: existence→ZKP) -- [x] Adaptive tuning (actual vs estimated latency feedback) -- [x] Post-processing + cross-modal cost models -- [x] ZKP scheme decision: PLONK (documented in Trustfile) - ---- - -## Phase 1: VCL Toolchain (HIGHEST PRIORITY) - -### 1.1 VCL REPL — Rust CLI -**Criticality: CRITICAL** | Effort: Medium | Crate: `verisim-repl` - -Every database has an interactive shell. Without one, VeriSimDB is unusable -for exploration and debugging. - -- Rust CLI binary (`vcl` command) -- Readline/rustyline for input editing, history, multiline -- HTTP client to verisim-api (configurable endpoint) -- Commands: `\connect`, `\explain`, `\timing`, `\format json|table|csv` -- Syntax highlighting for VCL keywords -- Tab completion for modalities, proof types -- Output formatters: table (default), JSON, CSV -- `.vclrc` config file support - -### 1.2 VCL REPL — Elixir IEx Extension -**Criticality: HIGH** | Effort: Small | Module: `VeriSim.VCL.IEx` - -- `use VeriSim.VCL.IEx` in IEx sessions -- `vcl("SELECT GRAPH FROM OCTAD ...")` function -- Pretty-printed results with modality indicators -- `vcl_explain/1` for EXPLAIN output -- Direct in-process execution (no HTTP round-trip) - -### 1.3 VCL Language Server (LSP) -**Criticality: HIGH** | Effort: Large | Crate: `verisim-lsp` - -- Diagnostics: parse errors, unknown modalities, type mismatches -- Completions: modality names, field names, proof types, keywords -- Hover: modality documentation, cost estimates -- Go-to-definition: proof contracts → semantic store -- VS Code extension + Neovim plugin -- Uses existing ReScript parser via JSON bridge - -### 1.4 VCL Formatter -**Criticality: MEDIUM** | Effort: Small | Module in `verisim-repl` - -- `vcl fmt` subcommand -- Canonical formatting for VCL queries -- Keyword uppercasing, consistent indentation -- Integrates with LSP `textDocument/formatting` - ---- - -## Phase 2: Storage & Durability (CRITICAL for deployment) - -### 2.1 Write-Ahead Log (WAL) -**Criticality: CRITICAL** | Effort: Large | Crate: `verisim-wal` - -Without WAL, any crash loses all data. Non-negotiable for deployment. - -- Append-only log of all write operations -- Configurable sync mode: `fsync` (safe) vs `async` (fast) -- Recovery: replay WAL on startup to rebuild state -- Log compaction: periodic checkpoint + truncation -- Per-modality WAL segments for parallel recovery - -### 2.2 ACID Transactions -**Criticality: CRITICAL** | Effort: Large | Module in `verisim-octad` - -Cross-modality atomicity. A octad update must either succeed across all -modalities or fail completely. - -- Transaction manager with begin/commit/rollback -- MVCC (Multi-Version Concurrency Control) for isolation -- Undo log for rollback across modalities -- Deadlock detection (modality-level locking) -- Configurable isolation levels: read-committed, serializable - -### 2.3 Persistence Layer Abstraction -**Criticality: HIGH** | Effort: Large | Trait: `StorageBackend` - -Currently in-memory only. Need pluggable backends. - -- `StorageBackend` trait with get/put/delete/scan -- Backends: Memory (current), RocksDB, SQLite, LMDB -- Per-modality backend configuration -- Migration tooling between backends - -### 2.4 Snapshots & Backup -**Criticality: HIGH** | Effort: Medium - -- Point-in-time snapshots (consistent across all 6 modalities) -- Incremental backup (WAL-based) -- Restore from snapshot + WAL replay -- Export/import in portable format - ---- - -## Phase 3: Security (CRITICAL for deployment) - -### 3.1 Authentication -**Criticality: CRITICAL** | Effort: Medium - -- API key authentication (REST, GraphQL, gRPC) -- JWT token support for stateless auth -- mTLS for gRPC -- Rate limiting per client - -### 3.2 Authorization (RBAC) -**Criticality: HIGH** | Effort: Large - -Configurable by admin, overridable by user, defaults to global. - -- Role-based access control -- Per-modality permissions (read/write/admin) -- Per-entity access lists -- Admin: set global + per-modality defaults -- User: override within allowed scope -- Audit log of access decisions - -### 3.3 Encryption at Rest -**Criticality: MEDIUM** | Effort: Medium - -- AES-256-GCM for data files -- Key management via rokur (secrets manager) -- Per-modality encryption keys -- Key rotation without downtime - -### 3.4 ZKP Integration (PLONK) -**Criticality: HIGH** | Effort: Large - -Real zero-knowledge proof verification for VCL-UT PROOF clause. -Scheme: PLONK (see contractiles/trust/Trustfile for rationale). - -- `ark-plonk` integration in verisim-semantic -- Circuit definitions for each proof type -- Universal SRS (Structured Reference String) management -- Proof generation API -- Proof verification in query pipeline -- Proof caching (verified proofs don't need re-verification) - ---- - -## Phase 4: Normalizer (Core Differentiator) - -### 4.1 Real Regeneration Strategies -**Criticality: CRITICAL** | Effort: Large - -The normalizer is VeriSimDB's killer feature — currently stubs. - -- Configurable authority ranking per modality - - Admin sets global defaults - - User can override for their entities - - Default: Document > Semantic > Graph > Vector > Tensor > Temporal -- Regeneration strategies: - - `from_authoritative`: regenerate drifted modality from highest-authority - - `merge`: combine information from multiple modalities - - `user_resolve`: flag for manual resolution -- Regeneration pipelines per modality pair -- Validation after regeneration (verify consistency restored) - -### 4.2 Conflict Resolution Policies -**Criticality: HIGH** | Effort: Medium - -- Last-writer-wins (default for non-critical data) -- Modality-priority (configurable ranking) -- Manual resolution queue -- Conflict history tracking - -### 4.3 Normalization Audit Trail -**Criticality: MEDIUM** | Effort: Medium - -- What was normalized, when, why -- Before/after snapshots -- Drift score history -- Admin dashboard for normalization health - ---- - -## Phase 5: Query Engine Maturity - -### 5.1 Query Profiling (EXPLAIN ANALYZE) -**Criticality: HIGH** | Effort: Medium - -Close the loop between estimated and actual costs. - -- `EXPLAIN ANALYZE` mode: execute + measure actual costs -- Per-step actual_ms vs estimated_ms -- Feed results into AdaptiveTuner automatically -- Profile history for query optimization - -### 5.2 Prepared Statements / Query Caching -**Criticality: MEDIUM** | Effort: Medium - -- Parse-once, execute-many for repeated queries -- Plan caching (skip re-optimization for identical plans) -- Parameterized queries (prevent VCL injection) -- Cache invalidation on schema/config changes - -### 5.3 Result Streaming -**Criticality: MEDIUM** | Effort: Medium - -- Server-sent events (REST) -- GraphQL subscriptions -- gRPC server streaming -- Backpressure handling -- Cursor-based pagination - -### 5.4 Slow Query Log -**Criticality: LOW** | Effort: Small - -- Configurable threshold (default: 100ms) -- Log query text, actual cost, plan chosen -- Integration with tracing/Prometheus - ---- - -## Phase 6: Distributed Systems - -### 6.1 Replication -**Criticality: HIGH** (for production) | Effort: Very Large - -- Raft consensus (KRAFT node stubs exist in Elixir) -- Leader-follower replication -- Automatic failover -- Read replicas for scaling queries - -### 6.2 Sharding -**Criticality: MEDIUM** | Effort: Large - -- Octad ID-based consistent hashing -- Per-modality shard assignment -- Cross-shard query routing -- Rebalancing without downtime - -### 6.3 Real Federation -**Criticality: MEDIUM** | Effort: Large - -- Federation resolver returns real results (currently empty) -- Cross-instance VCL queries -- Drift-aware federation (respect drift policies) -- Federation discovery protocol - ---- - -## Phase 7: Ecosystem & Developer Experience - -### 7.1 Benchmarks -**Criticality: HIGH** | Effort: Medium - -Can't claim best-in-class without numbers. - -- Benchmark suite: insert, query, mixed workload -- Compare against: ArangoDB, SurrealDB, Virtuoso -- Publish results with reproducible methodology -- CI-integrated regression benchmarks (criterion) - -### 7.2 Client Libraries -**Criticality: HIGH** | Effort: Medium - -- Rust SDK (typed, async, with VCL builder) -- Elixir SDK (direct BEAM integration) -- ReScript SDK (VCL builder + type-safe results) -- Each SDK: connection pooling, retry logic, auth - -### 7.3 Documentation Site -**Criticality: MEDIUM** | Effort: Medium - -- API reference (auto-generated from proto + GraphQL schema) -- VCL language guide with examples -- Architecture guide (Marr's three levels) -- Tutorial: "Build a multimodal search in 10 minutes" -- Deployment guide (Podman + Containerfile) - -### 7.4 Containerized Deployment -**Criticality: MEDIUM** | Effort: Small - -- Production Containerfile (multi-stage, chainguard base) -- selur-compose configuration -- Health check endpoints -- Graceful shutdown -- Resource limits and monitoring - -### 7.5 Migration Tooling -**Criticality: LOW** | Effort: Medium - -- Schema versioning -- Data migration scripts -- Import from: Neo4j (graph), Milvus (vector), PostgreSQL (document) -- Export to portable formats - ---- - -## Phase 8: Observability - -### 8.1 Metrics Dashboard -**Criticality: MEDIUM** | Effort: Small - -- Grafana dashboard template -- Per-modality latency, throughput, error rate -- Drift score visualization -- Normalization event timeline -- Query cost distribution - -### 8.2 Health Checks with Degraded States -**Criticality: LOW** | Effort: Small - -- Per-modality health (not just binary healthy/unhealthy) -- Degraded mode: some modalities down, others serving -- Dependency health (Elixir orchestration, store backends) - ---- - -## Summary: Priority Execution Order - -| # | Item | Phase | Criticality | -|---|------|-------|-------------| -| 1 | VCL REPL (Rust CLI) | 1.1 | CRITICAL | -| 2 | Write-Ahead Log | 2.1 | CRITICAL | -| 3 | Real normalizer regeneration | 4.1 | CRITICAL | -| 4 | Authentication | 3.1 | CRITICAL | -| 5 | ACID transactions | 2.2 | CRITICAL | -| 6 | VCL REPL (Elixir IEx) | 1.2 | HIGH | -| 7 | ZKP/PLONK integration | 3.4 | HIGH | -| 8 | Authorization (RBAC) | 3.2 | HIGH | -| 9 | Query profiling (EXPLAIN ANALYZE) | 5.1 | HIGH | -| 10 | VCL LSP | 1.3 | HIGH | -| 11 | Persistence backends | 2.3 | HIGH | -| 12 | Benchmarks vs ArangoDB/SurrealDB/Virtuoso | 7.1 | HIGH | -| 13 | Client libraries | 7.2 | HIGH | -| 14 | Snapshots & backup | 2.4 | HIGH | -| 15 | Conflict resolution | 4.2 | HIGH | -| 16 | Replication | 6.1 | HIGH | -| 17 | VCL formatter | 1.4 | MEDIUM | -| 18 | Encryption at rest | 3.3 | MEDIUM | -| 19 | Result streaming | 5.3 | MEDIUM | -| 20 | Prepared statements | 5.2 | MEDIUM | -| 21 | Normalization audit trail | 4.3 | MEDIUM | -| 22 | Documentation site | 7.3 | MEDIUM | -| 23 | Metrics dashboard | 8.1 | MEDIUM | -| 24 | Containerized deployment | 7.4 | MEDIUM | -| 25 | Sharding | 6.2 | MEDIUM | -| 26 | Real federation | 6.3 | MEDIUM | -| 27 | Slow query log | 5.4 | LOW | -| 28 | Health check degraded states | 8.2 | LOW | -| 29 | Migration tooling | 7.5 | LOW | diff --git a/verisimdb/CHANGELOG.adoc b/verisimdb/CHANGELOG.adoc deleted file mode 100644 index 93d7681b..00000000 --- a/verisimdb/CHANGELOG.adoc +++ /dev/null @@ -1,93 +0,0 @@ -// SPDX-License-Identifier: CC-BY-SA-4.0 -// Copyright (c) 2026 Jonathan D.A. Jewell (hyperpolymath) - -= VeriSimDB Changelog -:toc: left -:toclevels: 2 - -All notable changes to VeriSimDB are documented here. This project uses https://semver.org/[Semantic Versioning]. - -== [Unreleased] - -=== Added - -==== Security Hardening (Phase 1) -- **RwLock poisoning protection**: 35+ locations across 7 Rust files converted from `.expect("poisoned")` panics to graceful `map_err` error propagation. Server no longer crashes on lock contention. -- **API error sanitization**: Internal Rust error strings no longer leak to clients. All `ApiError::Internal` responses now log the real error via tracing and return a generic "Internal server error" to the client. -- **Input validation**: All search/list endpoints cap `limit` at 1000. Vector inputs checked for NaN/Inf. Octad IDs validated (≤128 chars, alphanumeric + dash + underscore). Empty query strings return 400. -- **Federation PSK authentication**: Peer registration requires `X-Federation-PSK` header. Configured via `VERISIM_FEDERATION_KEYS` env var. Federation disabled by default when keys are unset. - -==== Supply Chain & CI (Phase 2) -- **deny.toml**: Cargo dependency auditing with vulnerability deny, license allowlist, and wildcard dependency blocking. -- **CODEOWNERS**: All paths owned by @hyperpolymath. -- **SUPPORT.md**: Standard support policy. -- **quality.yml**: CI workflow for cargo deny, clippy, test, fmt, mix test, TruffleHog, EditorConfig. - -==== Operational Hardening (Phase 3) -- **IPv6-only default**: Server binds to `[::]` by default. `VERISIM_ENABLE_IPV4=true` enables dual-stack binding. -- **TLS support**: Optional HTTPS via `VERISIM_TLS_CERT` + `VERISIM_TLS_KEY` env vars using axum-server + rustls. -- **Container hardening**: Non-root user (`verisim`), OCI labels, IPv6 defaults in Containerfile. -- **Prometheus /metrics**: Drift gauge per type, uptime counter. Text/plain format for Prometheus scraping. -- **Readiness probe**: `/ready` checks octad store accessibility and drift detector health. Returns 503 when degraded. -- **Health check**: `/health` reports drift detector status with degraded reason when drift is critical. -- **Structured JSON logging**: `VERISIM_LOG_FORMAT=json` (default) enables JSON-formatted tracing output. - -==== Elixir Stub Completion (Phase 4) -- **QueryRouter**: Semantic queries via `search_text("type:")`, temporal queries via `/octads/{id}/versions`. -- **Federation Resolver**: Real repair via `Req.post` to peer normalizer. Modality-based routing (vector/graph/text/default). -- **SchemaRegistry**: Unknown constraint types now return `{:error, "Unknown constraint type"}` instead of silent `:ok`. -- **EntityServer**: Snapshots state before normalization. Selective modality normalization based on drift scores. -- **HealthChecker GenServer**: Periodic (30s) checks of Rust core, entity registry, and ETS cache. Emits telemetry events. -- **Application supervisor**: Changed to `:rest_for_one` strategy with `max_restarts: 10`. - -==== Rust Stub Completion (Phase 5) -- **TensorRegenerationStrategy**: Regenerates tensor from vector embedding reshape or document TF-IDF. -- **TemporalRepairStrategy**: Fixes timestamp ordering, version sanity, version_count consistency. -- **QualityReconciliationStrategy**: Cascades all strategies in priority order. -- **Adaptive drift thresholds**: `ThresholdPolicy::Adaptive { base, sensitivity }` with `effective_threshold()` method. Threshold = base + (moving_avg * sensitivity). - -==== Custom Circuits — ZKP Infrastructure (Phase 6) -- **Circuit Registry** (`circuit_registry.rs`): In-memory registry mapping circuit names to compiled R1CS verification functions. -- **Circuit Compiler** (`circuit_compiler.rs`): DSL gates (AND, OR, XOR, NOT, LinearCombination) compiled to R1CS constraints. -- **Verification Key Store** (`verification_keys.rs`): Per-circuit keys with rotation, federation export/import as `KeyExportBundle`. -- **VCL Circuit DSL** (`VCLCircuit.res`): ReScript types for circuit definition with `parseCustomProof` and `serializeCircuitDef`. - -==== Homoiconicity — Queries as Octads (Phase 7) -- **QueryOctadBuilder** (`query_octad.rs`): Stores VCL queries as octads across all 6 modalities (document=query text, graph=parse tree, vector=embedding, tensor=cost vector, semantic=proof obligations, temporal=execution history). -- **API endpoints**: `POST /queries` stores a query as octad. `POST /queries/similar` finds similar past queries by vector similarity. `PUT /queries/{id}/optimize` re-plans and updates cost vector. -- **VCL REFLECT keyword**: `SELECT * FROM REFLECT WHERE FULLTEXT CONTAINS 'drift'` queries the query store itself. Meta-circular: REFLECT queries are themselves stored as octads. -- **Elixir REFLECT executor**: Routes REFLECT source queries to the query store via text search. - -=== Changed -- Default bind address changed from `0.0.0.0` to `[::]` (IPv6). -- Elixir `runtime.exs` now reads `VERISIM_RUST_CORE_URL` from env (default: `http://[::1]:8080/api/v1`). -- Application supervisor strategy changed from `:one_for_one` to `:rest_for_one`. -- `DriftThresholds` now uses adaptive policies when configured. -- `create_default_normalizer()` registers all 5 strategies (was 2). - -=== Fixed -- RwLock poisoning panics in 7 Rust files (35+ locations). -- API error leakage exposing internal Rust error strings. -- Missing input validation on all search endpoints. -- Federation open to unauthenticated peer registration. -- Normalizer returning hardcoded placeholder strings. -- Federation resolver returning `{:error, :not_implemented}`. -- SchemaRegistry silently accepting unknown constraint types. - -== [0.1.0-alpha] — 2026-02-08 - -=== Added -- Initial alpha release with 6 modality stores (Graph, Vector, Tensor, Semantic, Document, Temporal). -- VCL parser (ReScript) with dependent-type and slipstream execution paths. -- Elixir/OTP orchestration layer with GenServers. -- Rust HTTP API server (verisim-api) with REST endpoints. -- Octad entity management (create, read, update, delete, search). -- Drift detection with 6 drift types and configurable thresholds. -- Self-normalization with vector and document regeneration strategies. -- Federation architecture with KRaft-inspired consensus. -- GitHub CI integration via verisimdb-data git-backed repo. -- Hypatia VeriSimDB connector for pattern detection. -- Criterion benchmarks for all modality stores. -- VCL grammar (ISO/IEC 14977 EBNF compliant). -- VCL formal semantics (operational + type system). -- Comprehensive documentation (WHITEPAPER, consultation papers, deployment guide). diff --git a/verisimdb/CODE_OF_CONDUCT.md b/verisimdb/CODE_OF_CONDUCT.md deleted file mode 100644 index ab314f8c..00000000 --- a/verisimdb/CODE_OF_CONDUCT.md +++ /dev/null @@ -1,327 +0,0 @@ -# Code of Conduct - - - -## Our Pledge - -We as members, contributors, and leaders pledge to make participation in Nextgen Databases a harassment-free experience for everyone, regardless of age, body size, visible or invisible disability, ethnicity, sex characteristics, gender identity and expression, level of experience, education, socio-economic status, nationality, personal appearance, race, caste, colour, religion, or sexual identity and orientation. - -We pledge to act and interact in ways that contribute to an open, welcoming, diverse, inclusive, and healthy community. - -We recognise that a thriving open source community requires **psychological safety** — an environment where people can contribute, ask questions, make mistakes, and learn without fear of ridicule or retaliation. - ---- - -## Our Standards - -### Expected Behaviour - -The following behaviours contribute to a positive environment: - -**Communication** -- Using welcoming and inclusive language -- Being respectful of differing viewpoints and experiences -- Giving and gracefully accepting constructive feedback -- Assuming good intent while addressing impact -- Communicating clearly and patiently, especially with newcomers - -**Collaboration** -- Focusing on what is best for the community -- Showing empathy and kindness toward other community members -- Being collaborative rather than competitive -- Mentoring and supporting less experienced contributors -- Celebrating others' contributions and successes - -**Professionalism** -- Accepting responsibility and apologising to those affected by our mistakes -- Learning from the experience and avoiding repetition -- Respecting others' time and attention -- Staying on topic in project spaces -- Following project guidelines and conventions - -**Accessibility** -- Using plain language and avoiding unnecessary jargon -- Providing alt text for images and transcripts for audio/video -- Being patient with those using assistive technologies -- Accommodating different communication styles and needs -- Recognising that not everyone communicates the same way - -### Unacceptable Behaviour - -The following behaviours are considered harassment and are unacceptable: - -**Harassment** -- The use of sexualised language or imagery, and sexual attention or advances of any kind -- Trolling, insulting or derogatory comments, and personal or political attacks -- Public or private harassment -- Deliberate intimidation, stalking, or following (online or in-person) -- Unwelcome physical contact or simulated physical contact (e.g., emoji) -- Sustained disruption of talks, events, or online discussions - -**Discrimination** -- Discriminatory jokes and language -- Posting or threatening to post others' personally identifying information ("doxing") -- Advocating for, or encouraging, any of the above behaviour -- Microaggressions — subtle, often unintentional, discriminatory comments or actions - -**Professional Misconduct** -- Publishing others' private information without explicit permission -- Misrepresenting affiliation or contributions -- Plagiarism or claiming credit for others' work -- Retaliating against anyone who reports a Code of Conduct violation -- Other conduct which could reasonably be considered inappropriate in a professional setting - -### Grey Areas - -Some situations require judgement. When uncertain: - -- **Intent vs Impact**: Good intentions do not excuse harmful impact. Focus on making things right. -- **Power Dynamics**: Those with more power (maintainers, employers, experienced contributors) must be especially mindful of their impact. -- **Cultural Differences**: What's acceptable varies by culture. When in doubt, err on the side of caution and ask. -- **Humour**: Jokes at others' expense are rarely funny to everyone. Punch up, not down. - ---- - -## Scope - -This Code of Conduct applies within all community spaces, including: - -**Online Spaces** -- Repository discussions, issues, and pull/merge requests -- Project chat channels (Matrix, Discord, Slack, IRC) -- Mailing lists and forums -- Social media when representing the project -- Video calls and virtual meetings - -**In-Person Spaces** -- Conferences, meetups, and events -- Workshops and training sessions -- Any gathering where you represent the project - -**Representation** -This Code of Conduct also applies when an individual is officially representing the community in public spaces. Examples include: - -- Using an official project email address -- Posting via an official social media account -- Acting as an appointed representative at an event -- Speaking on behalf of the project - ---- - -## Enforcement - -### Reporting - -If you experience or witness unacceptable behaviour, or have any other concerns, please report it as soon as possible. - -**How to Report** - -| Method | Details | Best For | -|--------|---------|----------| -| **Email** | {{CONDUCT_EMAIL}} | Detailed reports, sensitive matters | -| **Private Message** | Contact any maintainer directly | Quick questions, minor issues | -| **Anonymous Form** | [Link to form if available] | When you need anonymity | - -**What to Include** - -- Your contact information (unless anonymous) -- Names/usernames of those involved -- Description of what happened -- When and where it occurred -- Any witnesses -- Any supporting evidence (screenshots, links) -- How you would like us to respond (if you have a preference) - -**What Happens Next** - -1. You will receive acknowledgment within **{{RESPONSE_TIME}}** -2. The {{CONDUCT_TEAM}} will review the report -3. We may ask for additional information -4. We will determine appropriate action -5. We will inform you of the outcome (respecting others' privacy) - -### Confidentiality - -All reports will be handled with discretion: - -- Reporter identity is protected by default -- Details are shared only with those who need to know -- We will ask before naming you in any communication -- Anonymous reports are accepted and investigated - -### Conflicts of Interest - -If a {{CONDUCT_TEAM}} member is involved in an incident: - -- They will recuse themselves from the process -- Another maintainer or external party will handle the report -- We will disclose any potential conflicts - ---- - -## Enforcement Guidelines - -The {{CONDUCT_TEAM}} will follow these guidelines in determining consequences: - -### 1. Correction - -**Community Impact**: Use of inappropriate language or other behaviour deemed unprofessional or unwelcome. - -**Consequence**: A private, written warning providing clarity around the nature of the violation and an explanation of why the behaviour was inappropriate. A public apology may be requested. - -**Duration**: Immediate - -### 2. Warning - -**Community Impact**: A violation through a single incident or series of actions. - -**Consequence**: A warning with consequences for continued behaviour. No interaction with the people involved, including unsolicited interaction with those enforcing the Code of Conduct, for a specified period. This includes avoiding interactions in community spaces as well as external channels like social media. Violating these terms may lead to a temporary or permanent ban. - -**Duration**: 1-4 weeks - -### 3. Temporary Ban - -**Community Impact**: A serious violation of community standards, including sustained inappropriate behaviour. - -**Consequence**: A temporary ban from any sort of interaction or public communication with the community for a specified period. No public or private interaction with the people involved, including unsolicited interaction with those enforcing the Code of Conduct, is allowed during this period. Violating these terms may lead to a permanent ban. - -**Duration**: 1-6 months - -### 4. Permanent Ban - -**Community Impact**: Demonstrating a pattern of violation of community standards, including sustained inappropriate behaviour, harassment of an individual, or aggression toward or disparagement of classes of individuals. - -**Consequence**: A permanent ban from any sort of public interaction within the community. - -**Duration**: Permanent (with appeal rights after 12 months) - -### Enforcement Across Perimeters - -For contributors with elevated access (Perimeter 2 or 1): - -| Level | Additional Consequence | -|-------|----------------------| -| Correction | Noted in contributor record | -| Warning | Access privileges may be temporarily reduced | -| Temporary Ban | Access reduced to Perimeter 3 for ban duration | -| Permanent Ban | All access revoked | - ---- - -## Appeals - -If you believe an enforcement decision was made in error: - -1. **Wait 7 days** after the decision (cooling-off period) -2. **Email** {{CONDUCT_EMAIL}} with subject line "Appeal: [Original Report ID]" -3. **Explain** why you believe the decision should be reconsidered -4. **Provide** any new information not previously available - -**Appeals Process** - -- Appeals are reviewed by a different {{CONDUCT_TEAM}} member than the original -- You will receive a response within 14 days -- The appeals decision is final -- You may only appeal once per incident - -**Grounds for Appeal** - -- Procedural errors in the original investigation -- New evidence not previously available -- Disproportionate response to the violation -- Misunderstanding of facts - ---- - -## Supporting Those Who Report - -We are committed to supporting those who report violations: - -**We Will** -- Believe and take all reports seriously -- Respect your privacy and confidentiality preferences -- Keep you informed of progress (if you wish) -- Take steps to protect you from retaliation -- Provide resources if you need support - -**We Will Not** -- Require you to confront the person directly -- Dismiss reports without investigation -- Reveal your identity without consent -- Tolerate retaliation against reporters -- Rush you to make decisions - ---- - -## Prevention - -Beyond enforcement, we actively work to prevent issues: - -**Onboarding** -- All contributors are expected to read this Code of Conduct -- Perimeter 2 applicants must confirm they've read and understood it -- Maintainers receive additional training on enforcement - -**Culture** -- We model the behaviour we expect -- We intervene early when we see potential issues -- We thank people for positive contributions -- We create opportunities for diverse voices - -**Review** -- This Code of Conduct is reviewed annually -- Community feedback is welcomed -- Changes are communicated clearly - ---- - -## Acknowledgments - -This Code of Conduct is adapted from: - -- [Contributor Covenant](https://www.contributor-covenant.org/), version 2.1 -- [Django Code of Conduct](https://www.djangoproject.com/conduct/) -- [Rust Code of Conduct](https://www.rust-lang.org/policies/code-of-conduct) -- [Python Community Code of Conduct](https://www.python.org/psf/conduct/) - -We thank these communities for their leadership in creating welcoming spaces. - ---- - -## Questions? - -If you have questions about this Code of Conduct: - -- Open a [Discussion](https://github.com/hyperpolymath/nextgen-databases/discussions) (for general questions) -- Email {{CONDUCT_EMAIL}} (for private questions) -- Contact any maintainer directly - ---- - -## Summary - -**Be kind. Be respectful. Be collaborative.** - -We're all here because we care about this project. Let's make it a place where everyone can do their best work. - ---- - -Last updated: 2026 · Based on Contributor Covenant 2.1 diff --git a/verisimdb/CONTRIBUTING.md b/verisimdb/CONTRIBUTING.md deleted file mode 100644 index 2746ddb8..00000000 --- a/verisimdb/CONTRIBUTING.md +++ /dev/null @@ -1,187 +0,0 @@ -# Contributing to VeriSimDB - - - - -Thank you for your interest in contributing to VeriSimDB. This document explains how to get started, our development workflow, and how to submit changes. - -## Quick Start - -### Prerequisites - -- **Rust** (nightly) — `asdf install rust nightly` or `rustup install nightly` -- **Elixir** 1.17+ with Erlang/OTP 27+ — `asdf install elixir` / `asdf install erlang` -- **Podman** — for container builds (never Docker) - -### Setup - -```bash -# Clone the repository -git clone https://github.com/hyperpolymath/verisimdb.git -cd verisimdb - -# Build Rust core -cargo build -cargo test - -# Build Elixir orchestration -cd elixir-orchestration -mix deps.get -mix compile -mix test -``` - -### Repository Structure - -``` -verisimdb/ -├── rust-core/ # Rust crates (14 workspace members) -│ ├── verisim-api/ # HTTP/gRPC API server -│ ├── verisim-graph/ # Graph modality (RDF/Property Graph) -│ ├── verisim-vector/ # Vector modality (HNSW) -│ ├── verisim-tensor/ # Tensor modality (Burn) -│ ├── verisim-semantic/ # Semantic modality (CBOR proofs) -│ ├── verisim-document/ # Document modality (Tantivy) -│ ├── verisim-temporal/ # Temporal modality (versioning) -│ ├── verisim-octad/ # Unified 6-modal entity -│ ├── verisim-drift/ # Drift detection -│ ├── verisim-normalizer/ # Self-normalization -│ ├── verisim-planner/ # Cost-based query planner -│ ├── verisim-repl/ # Interactive VCL REPL -│ ├── verisim-wal/ # Write-ahead log -│ └── verisim-storage/ # Storage backend abstraction -├── elixir-orchestration/ # Elixir/OTP coordination layer -├── playground/ # VCL Playground PWA (ReScript) -├── container/ # Containerfile for Podman builds -├── docs/ # Architecture and design documents -├── contractiles/ # Trust, security, and policy contracts -├── .machine_readable/ # SCM checkpoint files -└── .github/workflows/ # CI/CD pipelines -``` - ---- - -## How to Contribute - -### Reporting Bugs - -**Before reporting**: -1. Search existing issues on [GitHub](https://github.com/hyperpolymath/verisimdb/issues) or [GitLab](https://gitlab.com/hyperpolymath/verisimdb/-/issues) -2. Check if it's already fixed in `main` - -**When reporting**, include: -- Clear, descriptive title -- Environment details (OS, Rust version, Elixir version) -- Steps to reproduce -- Expected vs actual behaviour -- Logs, error messages, or minimal reproduction - -### Suggesting Features - -**Before suggesting**: -1. Check the [roadmap](ROADMAP.adoc) -2. Search existing issues and discussions - -**When suggesting**, include: -- Problem statement (what pain point does this solve?) -- Proposed solution -- Alternatives considered -- Which modality or component it affects - -### Your First Contribution - -Look for issues labelled: -- [`good first issue`](https://github.com/hyperpolymath/verisimdb/labels/good%20first%20issue) — Simple tasks -- [`help wanted`](https://github.com/hyperpolymath/verisimdb/labels/help%20wanted) — Community help needed -- [`documentation`](https://github.com/hyperpolymath/verisimdb/labels/documentation) — Docs improvements - ---- - -## Development Workflow - -### Branch Naming -``` -docs/short-description # Documentation -test/what-added # Test additions -feat/short-description # New features -fix/issue-number-description # Bug fixes -refactor/what-changed # Code improvements -security/what-fixed # Security fixes -``` - -### Commit Messages - -We follow [Conventional Commits](https://www.conventionalcommits.org/): -``` -(): - -[optional body] - -[optional footer] -``` - -Types: `feat`, `fix`, `docs`, `test`, `refactor`, `chore`, `security` - -### Testing - -```bash -# Rust — all tests -cargo test - -# Rust — specific crate -cargo test -p verisim-semantic - -# Elixir -cd elixir-orchestration && mix test - -# Container build verification -podman build -t verisimdb:latest -f container/Containerfile . -``` - -### Code Quality - -```bash -# Rust linting -cargo clippy -- -D warnings - -# Check formatting -cargo fmt --check -``` - ---- - -## Language Policy - -### Allowed Languages - -| Language | Use Case | -|----------|----------| -| **Rust** | Core database engine, modality stores, CLI tools | -| **Elixir** | OTP orchestration, distributed coordination | -| **ReScript** | VCL parser, playground PWA | -| **VCL** | VeriSim Consonance Language (query interface) | - -### Not Accepted - -- TypeScript (use AffineScript instead) -- Python (use Rust or Julia instead) -- Go (use Rust instead) -- Node.js/npm/bun (use Deno if JS runtime needed) - ---- - -## License - -By contributing, you agree that your contributions will be licensed under the **PMPL-1.0-or-later** (Palimpsest License). All source files must include: - -``` -// SPDX-License-Identifier: CC-BY-SA-4.0 -``` - ---- - -## Contact - -- **Issues**: [GitHub](https://github.com/hyperpolymath/verisimdb/issues) or [GitLab](https://gitlab.com/hyperpolymath/verisimdb/-/issues) -- **Security**: See [SECURITY.md](SECURITY.md) for vulnerability reporting -- **Maintainer**: Jonathan D.A. Jewell diff --git a/verisimdb/Cargo.lock b/verisimdb/Cargo.lock deleted file mode 100644 index 25551c78..00000000 --- a/verisimdb/Cargo.lock +++ /dev/null @@ -1,5029 +0,0 @@ -# This file is automatically @generated by Cargo. -# It is not intended for manual editing. -version = 4 - -[[package]] -name = "Inflector" -version = "0.11.4" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "fe438c63458706e03479442743baae6c88256498e6431708f6dfc520a26515d3" - -[[package]] -name = "aho-corasick" -version = "1.1.4" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "ddd31a130427c27518df266943a5308ed92d4b226cc639f5a8f1002816174301" -dependencies = [ - "memchr", -] - -[[package]] -name = "allocator-api2" -version = "0.2.21" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "683d7910e743518b0e34f1186f92494becacb047c7b6bf616c96772180fef923" - -[[package]] -name = "android_system_properties" -version = "0.1.5" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "819e7219dbd41043ac279b19830f2efc897156490d7fd6ea916720117ee66311" -dependencies = [ - "libc", -] - -[[package]] -name = "anes" -version = "0.1.6" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "4b46cbb362ab8752921c97e041f5e366ee6297bd428a31275b9fcf1e380f7299" - -[[package]] -name = "anstream" -version = "1.0.0" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "824a212faf96e9acacdbd09febd34438f8f711fb84e09a8916013cd7815ca28d" -dependencies = [ - "anstyle", - "anstyle-parse", - "anstyle-query", - "anstyle-wincon", - "colorchoice", - "is_terminal_polyfill", - "utf8parse", -] - -[[package]] -name = "anstyle" -version = "1.0.14" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "940b3a0ca603d1eade50a4846a2afffd5ef57a9feac2c0e2ec2e14f9ead76000" - -[[package]] -name = "anstyle-parse" -version = "1.0.0" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "52ce7f38b242319f7cabaa6813055467063ecdc9d355bbb4ce0c68908cd8130e" -dependencies = [ - "utf8parse", -] - -[[package]] -name = "anstyle-query" -version = "1.1.5" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "40c48f72fd53cd289104fc64099abca73db4166ad86ea0b4341abe65af83dadc" -dependencies = [ - "windows-sys 0.61.2", -] - -[[package]] -name = "anstyle-wincon" -version = "3.0.11" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "291e6a250ff86cd4a820112fb8898808a366d8f9f58ce16d1f538353ad55747d" -dependencies = [ - "anstyle", - "once_cell_polyfill", - "windows-sys 0.61.2", -] - -[[package]] -name = "anyhow" -version = "1.0.102" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "7f202df86484c868dbad7eaa557ef785d5c66295e41b460ef922eca0723b842c" - -[[package]] -name = "arc-swap" -version = "1.9.1" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "6a3a1fd6f75306b68087b831f025c712524bcb19aad54e557b1129cfa0a2b207" -dependencies = [ - "rustversion", -] - -[[package]] -name = "ascii_utils" -version = "0.9.3" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "71938f30533e4d95a6d17aa530939da3842c2ab6f4f84b9dae68447e4129f74a" - -[[package]] -name = "async-graphql" -version = "7.2.1" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "1057a9f7ccf2404d94571dec3451ade1cb524790df6f1ada0d19c2a49f6b0f40" -dependencies = [ - "async-graphql-derive", - "async-graphql-parser", - "async-graphql-value", - "async-io", - "async-trait", - "asynk-strim", - "base64", - "bytes", - "fast_chemail", - "fnv", - "futures-util", - "handlebars", - "http", - "indexmap", - "mime", - "multer", - "num-traits", - "pin-project-lite", - "regex", - "serde", - "serde_json", - "serde_urlencoded", - "static_assertions_next", - "tempfile", - "thiserror 2.0.18", -] - -[[package]] -name = "async-graphql-axum" -version = "7.2.1" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "a1e37c5532e4b686acf45e7162bc93da91fc2c702fb0d465efc2c20c8f973795" -dependencies = [ - "async-graphql", - "axum", - "bytes", - "futures-util", - "serde_json", - "tokio", - "tokio-stream", - "tokio-util", - "tower-service", -] - -[[package]] -name = "async-graphql-derive" -version = "7.2.1" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "2e6cbeadc8515e66450fba0985ce722192e28443697799988265d86304d7cc68" -dependencies = [ - "Inflector", - "async-graphql-parser", - "darling 0.23.0", - "proc-macro-crate", - "proc-macro2", - "quote", - "strum", - "syn", - "thiserror 2.0.18", -] - -[[package]] -name = "async-graphql-parser" -version = "7.2.1" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "e64ef70f77a1c689111e52076da1cd18f91834bcb847de0a9171f83624b07fbf" -dependencies = [ - "async-graphql-value", - "pest", - "serde", - "serde_json", -] - -[[package]] -name = "async-graphql-value" -version = "7.2.1" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "3e3ef112905abea9dea592fc868a6873b10ebd3f983e83308f995d6284e9ba41" -dependencies = [ - "bytes", - "indexmap", - "serde", - "serde_json", -] - -[[package]] -name = "async-io" -version = "2.6.0" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "456b8a8feb6f42d237746d4b3e9a178494627745c3c56c6ea55d92ba50d026fc" -dependencies = [ - "autocfg", - "cfg-if", - "concurrent-queue", - "futures-io", - "futures-lite", - "parking", - "polling", - "rustix", - "slab", - "windows-sys 0.61.2", -] - -[[package]] -name = "async-trait" -version = "0.1.89" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "9035ad2d096bed7955a320ee7e2230574d28fd3c3a0f186cbea1ff3c7eed5dbb" -dependencies = [ - "proc-macro2", - "quote", - "syn", -] - -[[package]] -name = "asynk-strim" -version = "0.1.5" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "52697735bdaac441a29391a9e97102c74c6ef0f9b60a40cf109b1b404e29d2f6" -dependencies = [ - "futures-core", - "pin-project-lite", -] - -[[package]] -name = "atomic-waker" -version = "1.1.2" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "1505bd5d3d116872e7271a6d4e16d81d0c8570876c8de68093a09ac269d8aac0" - -[[package]] -name = "autocfg" -version = "1.5.0" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "c08606f8c3cbf4ce6ec8e28fb0014a2c086708fe954eaa885384a6165172e7e8" - -[[package]] -name = "axum" -version = "0.8.9" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "31b698c5f9a010f6573133b09e0de5408834d0c82f8d7475a89fc1867a71cd90" -dependencies = [ - "axum-core", - "base64", - "bytes", - "form_urlencoded", - "futures-util", - "http", - "http-body", - "http-body-util", - "hyper", - "hyper-util", - "itoa", - "matchit", - "memchr", - "mime", - "percent-encoding", - "pin-project-lite", - "serde_core", - "serde_json", - "serde_path_to_error", - "serde_urlencoded", - "sha1", - "sync_wrapper", - "tokio", - "tokio-tungstenite", - "tower", - "tower-layer", - "tower-service", - "tracing", -] - -[[package]] -name = "axum-core" -version = "0.5.6" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "08c78f31d7b1291f7ee735c1c6780ccde7785daae9a9206026862dab7d8792d1" -dependencies = [ - "bytes", - "futures-core", - "http", - "http-body", - "http-body-util", - "mime", - "pin-project-lite", - "sync_wrapper", - "tower-layer", - "tower-service", - "tracing", -] - -[[package]] -name = "axum-server" -version = "0.7.3" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "c1ab4a3ec9ea8a657c72d99a03a824af695bd0fb5ec639ccbd9cd3543b41a5f9" -dependencies = [ - "arc-swap", - "bytes", - "fs-err", - "http", - "http-body", - "hyper", - "hyper-util", - "pin-project-lite", - "rustls", - "rustls-pemfile", - "rustls-pki-types", - "tokio", - "tokio-rustls", - "tower-service", -] - -[[package]] -name = "base64" -version = "0.22.1" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "72b3254f16251a8381aa12e40e3c4d2f0199f8c6508fbecb9d91f575e0fbb8c6" - -[[package]] -name = "bindgen" -version = "0.71.1" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "5f58bf3d7db68cfbac37cfc485a8d711e87e064c3d0fe0435b92f7a407f9d6b3" -dependencies = [ - "bitflags", - "cexpr", - "clang-sys", - "itertools 0.13.0", - "log", - "prettyplease", - "proc-macro2", - "quote", - "regex", - "rustc-hash", - "shlex", - "syn", -] - -[[package]] -name = "bit-set" -version = "0.8.0" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "08807e080ed7f9d5433fa9b275196cfc35414f66a0c79d864dc51a0d825231a3" -dependencies = [ - "bit-vec", -] - -[[package]] -name = "bit-vec" -version = "0.8.0" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "5e764a1d40d510daf35e07be9eb06e75770908c27d411ee6c92109c9840eaaf7" - -[[package]] -name = "bitflags" -version = "2.11.1" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "c4512299f36f043ab09a583e57bceb5a5aab7a73db1805848e8fef3c9e8c78b3" - -[[package]] -name = "bitpacking" -version = "0.9.3" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "96a7139abd3d9cebf8cd6f920a389cf3dc9576172e32f4563f188cae3c3eb019" -dependencies = [ - "crunchy", -] - -[[package]] -name = "block-buffer" -version = "0.10.4" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "3078c7629b62d3f0439517fa394996acacc5cbc91c5a20d8c658e77abd503a71" -dependencies = [ - "generic-array", -] - -[[package]] -name = "bon" -version = "3.9.1" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "f47dbe92550676ee653353c310dfb9cf6ba17ee70396e1f7cf0a2020ad49b2fe" -dependencies = [ - "bon-macros", - "rustversion", -] - -[[package]] -name = "bon-macros" -version = "3.9.1" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "519bd3116aeeb42d5372c29d982d16d0170d3d4a5ed85fc7dd91642ffff3c67c" -dependencies = [ - "darling 0.23.0", - "ident_case", - "prettyplease", - "proc-macro2", - "quote", - "rustversion", - "syn", -] - -[[package]] -name = "bumpalo" -version = "3.20.2" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "5d20789868f4b01b2f2caec9f5c4e0213b41e3e5702a50157d699ae31ced2fcb" - -[[package]] -name = "byteorder" -version = "1.5.0" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "1fd0f2584146f6f2ef48085050886acf353beff7305ebd1ae69500e27c67f64b" - -[[package]] -name = "bytes" -version = "1.11.1" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "1e748733b7cbc798e1434b6ac524f0c1ff2ab456fe201501e6497c8417a4fc33" -dependencies = [ - "serde", -] - -[[package]] -name = "cast" -version = "0.3.0" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "37b2a672a2cb129a2e41c10b1224bb368f9f37a2b16b612598138befd7b37eb5" - -[[package]] -name = "cc" -version = "1.2.60" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "43c5703da9466b66a946814e1adf53ea2c90f10063b86290cc9eb67ce3478a20" -dependencies = [ - "find-msvc-tools", - "jobserver", - "libc", - "shlex", -] - -[[package]] -name = "census" -version = "0.4.2" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "4f4c707c6a209cbe82d10abd08e1ea8995e9ea937d2550646e02798948992be0" - -[[package]] -name = "cesu8" -version = "1.1.0" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "6d43a04d8753f35258c91f8ec639f792891f748a1edbd759cf1dcea3382ad83c" - -[[package]] -name = "cexpr" -version = "0.6.0" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "6fac387a98bb7c37292057cffc56d62ecb629900026402633ae9160df93a8766" -dependencies = [ - "nom", -] - -[[package]] -name = "cfg-if" -version = "1.0.4" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "9330f8b2ff13f34540b44e946ef35111825727b38d33286ef986142615121801" - -[[package]] -name = "cfg_aliases" -version = "0.2.1" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "613afe47fcd5fac7ccf1db93babcb082c5994d996f20b8b159f2ad1658eb5724" - -[[package]] -name = "chrono" -version = "0.4.44" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "c673075a2e0e5f4a1dde27ce9dee1ea4558c7ffe648f576438a20ca1d2acc4b0" -dependencies = [ - "iana-time-zone", - "js-sys", - "num-traits", - "serde", - "wasm-bindgen", - "windows-link", -] - -[[package]] -name = "ciborium" -version = "0.2.2" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "42e69ffd6f0917f5c029256a24d0161db17cea3997d185db0d35926308770f0e" -dependencies = [ - "ciborium-io", - "ciborium-ll", - "serde", -] - -[[package]] -name = "ciborium-io" -version = "0.2.2" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "05afea1e0a06c9be33d539b876f1ce3692f4afea2cb41f740e7743225ed1c757" - -[[package]] -name = "ciborium-ll" -version = "0.2.2" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "57663b653d948a338bfb3eeba9bb2fd5fcfaecb9e199e87e1eda4d9e8b240fd9" -dependencies = [ - "ciborium-io", - "half", -] - -[[package]] -name = "clang-sys" -version = "1.8.1" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "0b023947811758c97c59bf9d1c188fd619ad4718dcaa767947df1cadb14f39f4" -dependencies = [ - "glob", - "libc", - "libloading", -] - -[[package]] -name = "clap" -version = "4.6.1" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "1ddb117e43bbf7dacf0a4190fef4d345b9bad68dfc649cb349e7d17d28428e51" -dependencies = [ - "clap_builder", - "clap_derive", -] - -[[package]] -name = "clap_builder" -version = "4.6.0" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "714a53001bf66416adb0e2ef5ac857140e7dc3a0c48fb28b2f10762fc4b5069f" -dependencies = [ - "anstream", - "anstyle", - "clap_lex", - "strsim", -] - -[[package]] -name = "clap_derive" -version = "4.6.1" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "f2ce8604710f6733aa641a2b3731eaa1e8b3d9973d5e3565da11800813f997a9" -dependencies = [ - "heck", - "proc-macro2", - "quote", - "syn", -] - -[[package]] -name = "clap_lex" -version = "1.1.0" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "c8d4a3bb8b1e0c1050499d1815f5ab16d04f0959b233085fb31653fbfc9d98f9" - -[[package]] -name = "clipboard-win" -version = "5.4.1" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "bde03770d3df201d4fb868f2c9c59e66a3e4e2bd06692a0fe701e7103c7e84d4" -dependencies = [ - "error-code", -] - -[[package]] -name = "colorchoice" -version = "1.0.5" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "1d07550c9036bf2ae0c684c4297d503f838287c83c53686d05370d0e139ae570" - -[[package]] -name = "colored" -version = "3.1.1" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "faf9468729b8cbcea668e36183cb69d317348c2e08e994829fb56ebfdfbaac34" -dependencies = [ - "windows-sys 0.61.2", -] - -[[package]] -name = "combine" -version = "4.6.7" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "ba5a308b75df32fe02788e748662718f03fde005016435c444eea572398219fd" -dependencies = [ - "bytes", - "memchr", -] - -[[package]] -name = "comfy-table" -version = "7.2.2" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "958c5d6ecf1f214b4c2bbbbf6ab9523a864bd136dcf71a7e8904799acfe1ad47" -dependencies = [ - "crossterm", - "unicode-segmentation", - "unicode-width", -] - -[[package]] -name = "concurrent-queue" -version = "2.5.0" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "4ca0197aee26d1ae37445ee532fefce43251d24cc7c166799f4d46817f1d3973" -dependencies = [ - "crossbeam-utils", -] - -[[package]] -name = "core-foundation" -version = "0.10.1" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "b2a6cd9ae233e7f62ba4e9353e81a88df7fc8a5987b8d445b4d90c879bd156f6" -dependencies = [ - "core-foundation-sys", - "libc", -] - -[[package]] -name = "core-foundation-sys" -version = "0.8.7" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "773648b94d0e5d620f64f280777445740e61fe701025087ec8b57f45c791888b" - -[[package]] -name = "cpufeatures" -version = "0.2.17" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "59ed5838eebb26a2bb2e58f6d5b5316989ae9d08bab10e0e6d103e656d1b0280" -dependencies = [ - "libc", -] - -[[package]] -name = "crc32fast" -version = "1.5.0" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "9481c1c90cbf2ac953f07c8d4a58aa3945c425b7185c9154d67a65e4230da511" -dependencies = [ - "cfg-if", -] - -[[package]] -name = "criterion" -version = "0.5.1" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "f2b12d017a929603d80db1831cd3a24082f8137ce19c69e6447f54f5fc8d692f" -dependencies = [ - "anes", - "cast", - "ciborium", - "clap", - "criterion-plot", - "futures", - "is-terminal", - "itertools 0.10.5", - "num-traits", - "once_cell", - "oorandom", - "plotters", - "rayon", - "regex", - "serde", - "serde_derive", - "serde_json", - "tinytemplate", - "tokio", - "walkdir", -] - -[[package]] -name = "criterion-plot" -version = "0.5.0" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "6b50826342786a51a89e2da3a28f1c32b06e387201bc2d19791f622c673706b1" -dependencies = [ - "cast", - "itertools 0.10.5", -] - -[[package]] -name = "crossbeam-channel" -version = "0.5.15" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "82b8f8f868b36967f9606790d1903570de9ceaf870a7bf9fbbd3016d636a2cb2" -dependencies = [ - "crossbeam-utils", -] - -[[package]] -name = "crossbeam-deque" -version = "0.8.6" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "9dd111b7b7f7d55b72c0a6ae361660ee5853c9af73f70c3c2ef6858b950e2e51" -dependencies = [ - "crossbeam-epoch", - "crossbeam-utils", -] - -[[package]] -name = "crossbeam-epoch" -version = "0.9.18" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "5b82ac4a3c2ca9c3460964f020e1402edd5753411d7737aa39c3714ad1b5420e" -dependencies = [ - "crossbeam-utils", -] - -[[package]] -name = "crossbeam-utils" -version = "0.8.21" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "d0a5c400df2834b80a4c3327b3aad3a4c4cd4de0629063962b03235697506a28" - -[[package]] -name = "crossterm" -version = "0.29.0" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "d8b9f2e4c67f833b660cdb0a3523065869fb35570177239812ed4c905aeff87b" -dependencies = [ - "bitflags", - "crossterm_winapi", - "document-features", - "parking_lot", - "rustix", - "winapi", -] - -[[package]] -name = "crossterm_winapi" -version = "0.9.1" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "acdd7c62a3665c7f6830a51635d9ac9b23ed385797f70a83bb8bafe9c572ab2b" -dependencies = [ - "winapi", -] - -[[package]] -name = "crunchy" -version = "0.2.4" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "460fbee9c2c2f33933d720630a6a0bac33ba7053db5344fac858d4b8952d77d5" - -[[package]] -name = "crypto-common" -version = "0.1.7" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "78c8292055d1c1df0cce5d180393dc8cce0abec0a7102adb6c7b1eef6016d60a" -dependencies = [ - "generic-array", - "typenum", -] - -[[package]] -name = "darling" -version = "0.20.11" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "fc7f46116c46ff9ab3eb1597a45688b6715c6e628b5c133e288e709a29bcb4ee" -dependencies = [ - "darling_core 0.20.11", - "darling_macro 0.20.11", -] - -[[package]] -name = "darling" -version = "0.23.0" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "25ae13da2f202d56bd7f91c25fba009e7717a1e4a1cc98a76d844b65ae912e9d" -dependencies = [ - "darling_core 0.23.0", - "darling_macro 0.23.0", -] - -[[package]] -name = "darling_core" -version = "0.20.11" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "0d00b9596d185e565c2207a0b01f8bd1a135483d02d9b7b0a54b11da8d53412e" -dependencies = [ - "fnv", - "ident_case", - "proc-macro2", - "quote", - "strsim", - "syn", -] - -[[package]] -name = "darling_core" -version = "0.23.0" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "9865a50f7c335f53564bb694ef660825eb8610e0a53d3e11bf1b0d3df31e03b0" -dependencies = [ - "ident_case", - "proc-macro2", - "quote", - "strsim", - "syn", -] - -[[package]] -name = "darling_macro" -version = "0.20.11" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "fc34b93ccb385b40dc71c6fceac4b2ad23662c7eeb248cf10d529b7e055b6ead" -dependencies = [ - "darling_core 0.20.11", - "quote", - "syn", -] - -[[package]] -name = "darling_macro" -version = "0.23.0" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "ac3984ec7bd6cfa798e62b4a642426a5be0e68f9401cfc2a01e3fa9ea2fcdb8d" -dependencies = [ - "darling_core 0.23.0", - "quote", - "syn", -] - -[[package]] -name = "dashmap" -version = "6.1.0" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "5041cc499144891f3790297212f32a74fb938e5136a14943f338ef9e0ae276cf" -dependencies = [ - "cfg-if", - "crossbeam-utils", - "hashbrown 0.14.5", - "lock_api", - "once_cell", - "parking_lot_core", -] - -[[package]] -name = "data-encoding" -version = "2.10.0" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "d7a1e2f27636f116493b8b860f5546edb47c8d8f8ea73e1d2a20be88e28d1fea" - -[[package]] -name = "datasketches" -version = "0.2.0" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "c286de4e81ea2590afc24d754e0f83810c566f50a1388fa75ebd57928c0d9745" - -[[package]] -name = "deranged" -version = "0.5.8" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "7cd812cc2bc1d69d4764bd80df88b4317eaef9e773c75226407d9bc0876b211c" -dependencies = [ - "powerfmt", - "serde_core", -] - -[[package]] -name = "derive_builder" -version = "0.20.2" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "507dfb09ea8b7fa618fcf76e953f4f5e192547945816d5358edffe39f6f94947" -dependencies = [ - "derive_builder_macro", -] - -[[package]] -name = "derive_builder_core" -version = "0.20.2" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "2d5bcf7b024d6835cfb3d473887cd966994907effbe9227e8c8219824d06c4e8" -dependencies = [ - "darling 0.20.11", - "proc-macro2", - "quote", - "syn", -] - -[[package]] -name = "derive_builder_macro" -version = "0.20.2" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "ab63b0e2bf4d5928aff72e83a7dace85d7bba5fe12dcc3c5a572d78caffd3f3c" -dependencies = [ - "derive_builder_core", - "syn", -] - -[[package]] -name = "digest" -version = "0.10.7" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "9ed9a281f7bc9b7576e61468ba615a66a5c8cfdff42420a70aa82701a3b1e292" -dependencies = [ - "block-buffer", - "crypto-common", -] - -[[package]] -name = "dirs" -version = "6.0.0" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "c3e8aa94d75141228480295a7d0e7feb620b1a5ad9f12bc40be62411e38cce4e" -dependencies = [ - "dirs-sys", -] - -[[package]] -name = "dirs-sys" -version = "0.5.0" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "e01a3366d27ee9890022452ee61b2b63a67e6f13f58900b651ff5665f0bb1fab" -dependencies = [ - "libc", - "option-ext", - "redox_users", - "windows-sys 0.61.2", -] - -[[package]] -name = "displaydoc" -version = "0.2.5" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "97369cbbc041bc366949bc74d34658d6cda5621039731c6310521892a3a20ae0" -dependencies = [ - "proc-macro2", - "quote", - "syn", -] - -[[package]] -name = "document-features" -version = "0.2.12" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "d4b8a88685455ed29a21542a33abd9cb6510b6b129abadabdcef0f4c55bc8f61" -dependencies = [ - "litrs", -] - -[[package]] -name = "downcast-rs" -version = "2.0.2" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "117240f60069e65410b3ae1bb213295bd828f707b5bec6596a1afc8793ce0cbc" - -[[package]] -name = "either" -version = "1.15.0" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "48c757948c5ede0e46177b7add2e67155f70e33c07fea8284df6576da70b3719" - -[[package]] -name = "encoding_rs" -version = "0.8.35" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "75030f3c4f45dafd7586dd6780965a8c7e8e285a5ecb86713e63a79c5b2766f3" -dependencies = [ - "cfg-if", -] - -[[package]] -name = "endian-type" -version = "0.1.2" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "c34f04666d835ff5d62e058c3995147c06f42fe86ff053337632bca83e42702d" - -[[package]] -name = "equivalent" -version = "1.0.2" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "877a4ace8713b0bcf2a4e7eec82529c029f1d0619886d18145fea96c3ffe5c0f" - -[[package]] -name = "erased-serde" -version = "0.4.10" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "d2add8a07dd6a8d93ff627029c51de145e12686fbc36ecb298ac22e74cf02dec" -dependencies = [ - "serde", - "serde_core", - "typeid", -] - -[[package]] -name = "errno" -version = "0.3.14" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "39cab71617ae0d63f51a36d69f866391735b51691dbda63cf6f96d042b63efeb" -dependencies = [ - "libc", - "windows-sys 0.61.2", -] - -[[package]] -name = "error-code" -version = "3.3.2" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "dea2df4cf52843e0452895c455a1a2cfbb842a1e7329671acf418fdc53ed4c59" - -[[package]] -name = "fast_chemail" -version = "0.9.6" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "495a39d30d624c2caabe6312bfead73e7717692b44e0b32df168c275a2e8e9e4" -dependencies = [ - "ascii_utils", -] - -[[package]] -name = "fastdivide" -version = "0.4.2" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "9afc2bd4d5a73106dd53d10d73d3401c2f32730ba2c0b93ddb888a8983680471" - -[[package]] -name = "fastrand" -version = "2.4.1" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "9f1f227452a390804cdb637b74a86990f2a7d7ba4b7d5693aac9b4dd6defd8d6" - -[[package]] -name = "fd-lock" -version = "4.0.4" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "0ce92ff622d6dadf7349484f42c93271a0d49b7cc4d466a936405bacbe10aa78" -dependencies = [ - "cfg-if", - "rustix", - "windows-sys 0.59.0", -] - -[[package]] -name = "find-msvc-tools" -version = "0.1.9" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "5baebc0774151f905a1a2cc41989300b1e6fbb29aff0ceffa1064fdd3088d582" - -[[package]] -name = "fnv" -version = "1.0.7" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "3f9eec918d3f24069decb9af1554cad7c880e2da24a9afd88aca000531ab82c1" - -[[package]] -name = "foldhash" -version = "0.1.5" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "d9c4f5dac5e15c24eb999c26181a6ca40b39fe946cbe4c263c7209467bc83af2" - -[[package]] -name = "foldhash" -version = "0.2.0" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "77ce24cb58228fbb8aa041425bb1050850ac19177686ea6e0f41a70416f56fdb" - -[[package]] -name = "form_urlencoded" -version = "1.2.2" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "cb4cb245038516f5f85277875cdaa4f7d2c9a0fa0468de06ed190163b1581fcf" -dependencies = [ - "percent-encoding", -] - -[[package]] -name = "fs-err" -version = "3.3.0" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "73fde052dbfc920003cfd2c8e2c6e6d4cc7c1091538c3a24226cec0665ab08c0" -dependencies = [ - "autocfg", - "tokio", -] - -[[package]] -name = "fs4" -version = "0.13.1" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "8640e34b88f7652208ce9e88b1a37a2ae95227d84abec377ccd3c5cfeb141ed4" -dependencies = [ - "rustix", - "windows-sys 0.59.0", -] - -[[package]] -name = "futures" -version = "0.3.32" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "8b147ee9d1f6d097cef9ce628cd2ee62288d963e16fb287bd9286455b241382d" -dependencies = [ - "futures-channel", - "futures-core", - "futures-executor", - "futures-io", - "futures-sink", - "futures-task", - "futures-util", -] - -[[package]] -name = "futures-channel" -version = "0.3.32" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "07bbe89c50d7a535e539b8c17bc0b49bdb77747034daa8087407d655f3f7cc1d" -dependencies = [ - "futures-core", - "futures-sink", -] - -[[package]] -name = "futures-core" -version = "0.3.32" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "7e3450815272ef58cec6d564423f6e755e25379b217b0bc688e295ba24df6b1d" - -[[package]] -name = "futures-executor" -version = "0.3.32" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "baf29c38818342a3b26b5b923639e7b1f4a61fc5e76102d4b1981c6dc7a7579d" -dependencies = [ - "futures-core", - "futures-task", - "futures-util", -] - -[[package]] -name = "futures-io" -version = "0.3.32" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "cecba35d7ad927e23624b22ad55235f2239cfa44fd10428eecbeba6d6a717718" - -[[package]] -name = "futures-lite" -version = "2.6.1" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "f78e10609fe0e0b3f4157ffab1876319b5b0db102a2c60dc4626306dc46b44ad" -dependencies = [ - "futures-core", - "pin-project-lite", -] - -[[package]] -name = "futures-macro" -version = "0.3.32" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "e835b70203e41293343137df5c0664546da5745f82ec9b84d40be8336958447b" -dependencies = [ - "proc-macro2", - "quote", - "syn", -] - -[[package]] -name = "futures-sink" -version = "0.3.32" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "c39754e157331b013978ec91992bde1ac089843443c49cbc7f46150b0fad0893" - -[[package]] -name = "futures-task" -version = "0.3.32" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "037711b3d59c33004d3856fbdc83b99d4ff37a24768fa1be9ce3538a1cde4393" - -[[package]] -name = "futures-util" -version = "0.3.32" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "389ca41296e6190b48053de0321d02a77f32f8a5d2461dd38762c0593805c6d6" -dependencies = [ - "futures-channel", - "futures-core", - "futures-io", - "futures-macro", - "futures-sink", - "futures-task", - "memchr", - "pin-project-lite", - "slab", -] - -[[package]] -name = "generic-array" -version = "0.14.7" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "85649ca51fd72272d7821adaf274ad91c288277713d9c18820d8499a7ff69e9a" -dependencies = [ - "typenum", - "version_check", -] - -[[package]] -name = "getrandom" -version = "0.2.17" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "ff2abc00be7fca6ebc474524697ae276ad847ad0a6b3faa4bcb027e9a4614ad0" -dependencies = [ - "cfg-if", - "libc", - "wasi", -] - -[[package]] -name = "getrandom" -version = "0.3.4" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "899def5c37c4fd7b2664648c28120ecec138e4d395b459e5ca34f9cce2dd77fd" -dependencies = [ - "cfg-if", - "libc", - "r-efi 5.3.0", - "wasip2", -] - -[[package]] -name = "getrandom" -version = "0.4.2" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "0de51e6874e94e7bf76d726fc5d13ba782deca734ff60d5bb2fb2607c7406555" -dependencies = [ - "cfg-if", - "libc", - "r-efi 6.0.0", - "wasip2", - "wasip3", -] - -[[package]] -name = "glob" -version = "0.3.3" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "0cc23270f6e1808e30a928bdc84dea0b9b4136a8bc82338574f23baf47bbd280" - -[[package]] -name = "h2" -version = "0.4.13" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "2f44da3a8150a6703ed5d34e164b875fd14c2cdab9af1252a9a1020bde2bdc54" -dependencies = [ - "atomic-waker", - "bytes", - "fnv", - "futures-core", - "futures-sink", - "http", - "indexmap", - "slab", - "tokio", - "tokio-util", - "tracing", -] - -[[package]] -name = "half" -version = "2.7.1" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "6ea2d84b969582b4b1864a92dc5d27cd2b77b622a8d79306834f1be5ba20d84b" -dependencies = [ - "cfg-if", - "crunchy", - "zerocopy", -] - -[[package]] -name = "handlebars" -version = "6.4.0" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "9b3f9296c208515b87bd915a2f5d1163d4b3f863ba83337d7713cf478055948e" -dependencies = [ - "derive_builder", - "log", - "num-order", - "pest", - "pest_derive", - "serde", - "serde_json", - "thiserror 2.0.18", -] - -[[package]] -name = "hashbrown" -version = "0.14.5" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "e5274423e17b7c9fc20b6e7e208532f9b19825d82dfd615708b70edd83df41f1" - -[[package]] -name = "hashbrown" -version = "0.15.5" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "9229cfe53dfd69f0609a49f65461bd93001ea1ef889cd5529dd176593f5338a1" -dependencies = [ - "foldhash 0.1.5", -] - -[[package]] -name = "hashbrown" -version = "0.16.1" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "841d1cc9bed7f9236f321df977030373f4a4163ae1a7dbfe1a51a2c1a51d9100" -dependencies = [ - "allocator-api2", - "equivalent", - "foldhash 0.2.0", -] - -[[package]] -name = "hashbrown" -version = "0.17.0" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "4f467dd6dccf739c208452f8014c75c18bb8301b050ad1cfb27153803edb0f51" - -[[package]] -name = "heck" -version = "0.5.0" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "2304e00983f87ffb38b55b444b5e3b60a884b5d30c0fca7d82fe33449bbe55ea" - -[[package]] -name = "hermit-abi" -version = "0.5.2" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "fc0fef456e4baa96da950455cd02c081ca953b141298e41db3fc7e36b1da849c" - -[[package]] -name = "hex" -version = "0.4.3" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "7f24254aa9a54b5c858eaee2f5bccdb46aaf0e486a595ed5fd8f86ba55232a70" - -[[package]] -name = "home" -version = "0.5.12" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "cc627f471c528ff0c4a49e1d5e60450c8f6461dd6d10ba9dcd3a61d3dff7728d" -dependencies = [ - "windows-sys 0.61.2", -] - -[[package]] -name = "htmlescape" -version = "0.3.1" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "e9025058dae765dee5070ec375f591e2ba14638c63feff74f13805a72e523163" - -[[package]] -name = "http" -version = "1.4.0" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "e3ba2a386d7f85a81f119ad7498ebe444d2e22c2af0b86b069416ace48b3311a" -dependencies = [ - "bytes", - "itoa", -] - -[[package]] -name = "http-body" -version = "1.0.1" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "1efedce1fb8e6913f23e0c92de8e62cd5b772a67e7b3946df930a62566c93184" -dependencies = [ - "bytes", - "http", -] - -[[package]] -name = "http-body-util" -version = "0.1.3" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "b021d93e26becf5dc7e1b75b1bed1fd93124b374ceb73f43d4d4eafec896a64a" -dependencies = [ - "bytes", - "futures-core", - "http", - "http-body", - "pin-project-lite", -] - -[[package]] -name = "httparse" -version = "1.10.1" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "6dbf3de79e51f3d586ab4cb9d5c3e2c14aa28ed23d180cf89b4df0454a69cc87" - -[[package]] -name = "httpdate" -version = "1.0.3" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "df3b46402a9d5adb4c86a0cf463f42e19994e3ee891101b1841f30a545cb49a9" - -[[package]] -name = "hyper" -version = "1.9.0" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "6299f016b246a94207e63da54dbe807655bf9e00044f73ded42c3ac5305fbcca" -dependencies = [ - "atomic-waker", - "bytes", - "futures-channel", - "futures-core", - "h2", - "http", - "http-body", - "httparse", - "httpdate", - "itoa", - "pin-project-lite", - "smallvec", - "tokio", - "want", -] - -[[package]] -name = "hyper-rustls" -version = "0.27.9" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "33ca68d021ef39cf6463ab54c1d0f5daf03377b70561305bb89a8f83aab66e0f" -dependencies = [ - "http", - "hyper", - "hyper-util", - "rustls", - "tokio", - "tokio-rustls", - "tower-service", -] - -[[package]] -name = "hyper-timeout" -version = "0.5.2" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "2b90d566bffbce6a75bd8b09a05aa8c2cb1fabb6cb348f8840c9e4c90a0d83b0" -dependencies = [ - "hyper", - "hyper-util", - "pin-project-lite", - "tokio", - "tower-service", -] - -[[package]] -name = "hyper-util" -version = "0.1.20" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "96547c2556ec9d12fb1578c4eaf448b04993e7fb79cbaad930a656880a6bdfa0" -dependencies = [ - "base64", - "bytes", - "futures-channel", - "futures-util", - "http", - "http-body", - "hyper", - "ipnet", - "libc", - "percent-encoding", - "pin-project-lite", - "socket2", - "tokio", - "tower-service", - "tracing", -] - -[[package]] -name = "iana-time-zone" -version = "0.1.65" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "e31bc9ad994ba00e440a8aa5c9ef0ec67d5cb5e5cb0cc7f8b744a35b389cc470" -dependencies = [ - "android_system_properties", - "core-foundation-sys", - "iana-time-zone-haiku", - "js-sys", - "log", - "wasm-bindgen", - "windows-core", -] - -[[package]] -name = "iana-time-zone-haiku" -version = "0.1.2" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "f31827a206f56af32e590ba56d5d2d085f558508192593743f16b2306495269f" -dependencies = [ - "cc", -] - -[[package]] -name = "icu_collections" -version = "2.2.0" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "2984d1cd16c883d7935b9e07e44071dca8d917fd52ecc02c04d5fa0b5a3f191c" -dependencies = [ - "displaydoc", - "potential_utf", - "utf8_iter", - "yoke", - "zerofrom", - "zerovec", -] - -[[package]] -name = "icu_locale_core" -version = "2.2.0" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "92219b62b3e2b4d88ac5119f8904c10f8f61bf7e95b640d25ba3075e6cac2c29" -dependencies = [ - "displaydoc", - "litemap", - "tinystr", - "writeable", - "zerovec", -] - -[[package]] -name = "icu_normalizer" -version = "2.2.0" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "c56e5ee99d6e3d33bd91c5d85458b6005a22140021cc324cea84dd0e72cff3b4" -dependencies = [ - "icu_collections", - "icu_normalizer_data", - "icu_properties", - "icu_provider", - "smallvec", - "zerovec", -] - -[[package]] -name = "icu_normalizer_data" -version = "2.2.0" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "da3be0ae77ea334f4da67c12f149704f19f81d1adf7c51cf482943e84a2bad38" - -[[package]] -name = "icu_properties" -version = "2.2.0" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "bee3b67d0ea5c2cca5003417989af8996f8604e34fb9ddf96208a033901e70de" -dependencies = [ - "icu_collections", - "icu_locale_core", - "icu_properties_data", - "icu_provider", - "zerotrie", - "zerovec", -] - -[[package]] -name = "icu_properties_data" -version = "2.2.0" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "8e2bbb201e0c04f7b4b3e14382af113e17ba4f63e2c9d2ee626b720cbce54a14" - -[[package]] -name = "icu_provider" -version = "2.2.0" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "139c4cf31c8b5f33d7e199446eff9c1e02decfc2f0eec2c8d71f65befa45b421" -dependencies = [ - "displaydoc", - "icu_locale_core", - "writeable", - "yoke", - "zerofrom", - "zerotrie", - "zerovec", -] - -[[package]] -name = "id-arena" -version = "2.3.0" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "3d3067d79b975e8844ca9eb072e16b31c3c1c36928edf9c6789548c524d0d954" - -[[package]] -name = "ident_case" -version = "1.0.1" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "b9e0384b61958566e926dc50660321d12159025e767c18e043daf26b70104c39" - -[[package]] -name = "idna" -version = "1.1.0" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "3b0875f23caa03898994f6ddc501886a45c7d3d62d04d2d90788d47be1b1e4de" -dependencies = [ - "idna_adapter", - "smallvec", - "utf8_iter", -] - -[[package]] -name = "idna_adapter" -version = "1.2.1" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "3acae9609540aa318d1bc588455225fb2085b9ed0c4f6bd0d9d5bcd86f1a0344" -dependencies = [ - "icu_normalizer", - "icu_properties", -] - -[[package]] -name = "indexmap" -version = "2.14.0" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "d466e9454f08e4a911e14806c24e16fba1b4c121d1ea474396f396069cf949d9" -dependencies = [ - "equivalent", - "hashbrown 0.17.0", - "serde", - "serde_core", -] - -[[package]] -name = "inventory" -version = "0.3.24" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "a4f0c30c76f2f4ccee3fe55a2435f691ca00c0e4bd87abe4f4a851b1d4dac39b" -dependencies = [ - "rustversion", -] - -[[package]] -name = "ipnet" -version = "2.12.0" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "d98f6fed1fde3f8c21bc40a1abb88dd75e67924f9cffc3ef95607bad8017f8e2" - -[[package]] -name = "iri-string" -version = "0.7.12" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "25e659a4bb38e810ebc252e53b5814ff908a8c58c2a9ce2fae1bbec24cbf4e20" -dependencies = [ - "memchr", - "serde", -] - -[[package]] -name = "is-terminal" -version = "0.4.17" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "3640c1c38b8e4e43584d8df18be5fc6b0aa314ce6ebf51b53313d4306cca8e46" -dependencies = [ - "hermit-abi", - "libc", - "windows-sys 0.61.2", -] - -[[package]] -name = "is_terminal_polyfill" -version = "1.70.2" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "a6cb138bb79a146c1bd460005623e142ef0181e3d0219cb493e02f7d08a35695" - -[[package]] -name = "itertools" -version = "0.10.5" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "b0fd2260e829bddf4cb6ea802289de2f86d6a7a690192fbe91b3f46e0f2c8473" -dependencies = [ - "either", -] - -[[package]] -name = "itertools" -version = "0.13.0" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "413ee7dfc52ee1a4949ceeb7dbc8a33f2d6c088194d9f922fb8318faf1f01186" -dependencies = [ - "either", -] - -[[package]] -name = "itertools" -version = "0.14.0" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "2b192c782037fadd9cfa75548310488aabdbf3d2da73885b31bd0abd03351285" -dependencies = [ - "either", -] - -[[package]] -name = "itoa" -version = "1.0.18" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "8f42a60cbdf9a97f5d2305f08a87dc4e09308d1276d28c869c684d7777685682" - -[[package]] -name = "jni" -version = "0.21.1" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "1a87aa2bb7d2af34197c04845522473242e1aa17c12f4935d5856491a7fb8c97" -dependencies = [ - "cesu8", - "cfg-if", - "combine", - "jni-sys 0.3.1", - "log", - "thiserror 1.0.69", - "walkdir", - "windows-sys 0.45.0", -] - -[[package]] -name = "jni-sys" -version = "0.3.1" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "41a652e1f9b6e0275df1f15b32661cf0d4b78d4d87ddec5e0c3c20f097433258" -dependencies = [ - "jni-sys 0.4.1", -] - -[[package]] -name = "jni-sys" -version = "0.4.1" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "c6377a88cb3910bee9b0fa88d4f42e1d2da8e79915598f65fb0c7ee14c878af2" -dependencies = [ - "jni-sys-macros", -] - -[[package]] -name = "jni-sys-macros" -version = "0.4.1" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "38c0b942f458fe50cdac086d2f946512305e5631e720728f2a61aabcd47a6264" -dependencies = [ - "quote", - "syn", -] - -[[package]] -name = "jobserver" -version = "0.1.34" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "9afb3de4395d6b3e67a780b6de64b51c978ecf11cb9a462c66be7d4ca9039d33" -dependencies = [ - "getrandom 0.3.4", - "libc", -] - -[[package]] -name = "js-sys" -version = "0.3.95" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "2964e92d1d9dc3364cae4d718d93f227e3abb088e747d92e0395bfdedf1c12ca" -dependencies = [ - "cfg-if", - "futures-util", - "once_cell", - "wasm-bindgen", -] - -[[package]] -name = "json-event-parser" -version = "0.2.3" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "574b0cd5e90ee2ba03a66d0611fc9a09c9a0c28b2ecc2dc8a181dd31a53ca5d7" - -[[package]] -name = "lazy_static" -version = "1.5.0" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "bbd2bcb4c963f2ddae06a2efc7e9f3591312473c50c6685e1f298068316e66fe" - -[[package]] -name = "leb128fmt" -version = "0.1.0" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "09edd9e8b54e49e587e4f6295a7d29c3ea94d469cb40ab8ca70b288248a81db2" - -[[package]] -name = "levenshtein_automata" -version = "0.2.1" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "0c2cdeb66e45e9f36bfad5bbdb4d2384e70936afbee843c6f6543f0c551ebb25" - -[[package]] -name = "libc" -version = "0.2.185" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "52ff2c0fe9bc6cb6b14a0592c2ff4fa9ceb83eea9db979b0487cd054946a2b8f" - -[[package]] -name = "libloading" -version = "0.8.9" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "d7c4b02199fee7c5d21a5ae7d8cfa79a6ef5bb2fc834d6e9058e89c825efdc55" -dependencies = [ - "cfg-if", - "windows-link", -] - -[[package]] -name = "libredox" -version = "0.1.16" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "e02f3bb43d335493c96bf3fd3a321600bf6bd07ed34bc64118e9293bdffea46c" -dependencies = [ - "libc", -] - -[[package]] -name = "linux-raw-sys" -version = "0.12.1" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "32a66949e030da00e8c7d4434b251670a91556f4144941d37452769c25d58a53" - -[[package]] -name = "litemap" -version = "0.8.2" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "92daf443525c4cce67b150400bc2316076100ce0b3686209eb8cf3c31612e6f0" - -[[package]] -name = "litrs" -version = "1.0.0" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "11d3d7f243d5c5a8b9bb5d6dd2b1602c0cb0b9db1621bafc7ed66e35ff9fe092" - -[[package]] -name = "lock_api" -version = "0.4.14" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "224399e74b87b5f3557511d98dff8b14089b3dadafcab6bb93eab67d3aace965" -dependencies = [ - "scopeguard", -] - -[[package]] -name = "log" -version = "0.4.29" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "5e5032e24019045c762d3c0f28f5b6b8bbf38563a65908389bf7978758920897" - -[[package]] -name = "lru" -version = "0.16.4" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "7f66e8d5d03f609abc3a39e6f08e4164ebf1447a732906d39eb9b99b7919ef39" -dependencies = [ - "hashbrown 0.16.1", -] - -[[package]] -name = "lz4_flex" -version = "0.13.0" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "db9a0d582c2874f68138a16ce1867e0ffde6c0bb0a0df85e1f36d04146db488a" - -[[package]] -name = "matchers" -version = "0.2.0" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "d1525a2a28c7f4fa0fc98bb91ae755d1e2d1505079e05539e35bc876b5d65ae9" -dependencies = [ - "regex-automata", -] - -[[package]] -name = "matchit" -version = "0.8.4" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "47e1ffaa40ddd1f3ed91f717a33c8c0ee23fff369e3aa8772b9605cc1d22f4c3" - -[[package]] -name = "matrixmultiply" -version = "0.3.10" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "a06de3016e9fae57a36fd14dba131fccf49f74b40b7fbdb472f96e361ec71a08" -dependencies = [ - "autocfg", - "rawpointer", -] - -[[package]] -name = "md-5" -version = "0.10.6" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "d89e7ee0cfbedfc4da3340218492196241d89eefb6dab27de5df917a6d2e78cf" -dependencies = [ - "cfg-if", - "digest", -] - -[[package]] -name = "measure_time" -version = "0.9.0" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "51c55d61e72fc3ab704396c5fa16f4c184db37978ae4e94ca8959693a235fc0e" -dependencies = [ - "log", -] - -[[package]] -name = "memchr" -version = "2.8.0" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "f8ca58f447f06ed17d5fc4043ce1b10dd205e060fb3ce5b979b8ed8e59ff3f79" - -[[package]] -name = "memmap2" -version = "0.9.10" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "714098028fe011992e1c3962653c96b2d578c4b4bce9036e15ff220319b1e0e3" -dependencies = [ - "libc", -] - -[[package]] -name = "mime" -version = "0.3.17" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "6877bb514081ee2a7ff5ef9de3281f14a4dd4bceac4c09388074a6b5df8a139a" - -[[package]] -name = "minimal-lexical" -version = "0.2.1" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "68354c5c6bd36d73ff3feceb05efa59b6acb7626617f4962be322a825e61f79a" - -[[package]] -name = "mio" -version = "1.2.0" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "50b7e5b27aa02a74bac8c3f23f448f8d87ff11f92d3aac1a6ed369ee08cc56c1" -dependencies = [ - "libc", - "wasi", - "windows-sys 0.61.2", -] - -[[package]] -name = "multer" -version = "3.1.0" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "83e87776546dc87511aa5ee218730c92b666d7264ab6ed41f9d215af9cd5224b" -dependencies = [ - "bytes", - "encoding_rs", - "futures-util", - "http", - "httparse", - "memchr", - "mime", - "spin", - "version_check", -] - -[[package]] -name = "murmurhash32" -version = "0.3.1" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "2195bf6aa996a481483b29d62a7663eed3fe39600c460e323f8ff41e90bdd89b" - -[[package]] -name = "ndarray" -version = "0.16.1" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "882ed72dce9365842bf196bdeedf5055305f11fc8c03dee7bb0194a6cad34841" -dependencies = [ - "matrixmultiply", - "num-complex", - "num-integer", - "num-traits", - "portable-atomic", - "portable-atomic-util", - "rawpointer", -] - -[[package]] -name = "nibble_vec" -version = "0.1.0" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "77a5d83df9f36fe23f0c3648c6bbb8b0298bb5f1939c8f2704431371f4b84d43" -dependencies = [ - "smallvec", -] - -[[package]] -name = "nix" -version = "0.29.0" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "71e2746dc3a24dd78b3cfcb7be93368c6de9963d30f43a6a73998a9cf4b17b46" -dependencies = [ - "bitflags", - "cfg-if", - "cfg_aliases", - "libc", -] - -[[package]] -name = "nom" -version = "7.1.3" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "d273983c5a657a70a3e8f2a01329822f3b8c8172b73826411a55751e404a0a4a" -dependencies = [ - "memchr", - "minimal-lexical", -] - -[[package]] -name = "nu-ansi-term" -version = "0.50.3" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "7957b9740744892f114936ab4a57b3f487491bbeafaf8083688b16841a4240e5" -dependencies = [ - "windows-sys 0.61.2", -] - -[[package]] -name = "num-complex" -version = "0.4.6" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "73f88a1307638156682bada9d7604135552957b7818057dcef22705b4d509495" -dependencies = [ - "num-traits", -] - -[[package]] -name = "num-conv" -version = "0.2.1" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "c6673768db2d862beb9b39a78fdcb1a69439615d5794a1be50caa9bc92c81967" - -[[package]] -name = "num-integer" -version = "0.1.46" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "7969661fd2958a5cb096e56c8e1ad0444ac2bbcd0061bd28660485a44879858f" -dependencies = [ - "num-traits", -] - -[[package]] -name = "num-modular" -version = "0.6.1" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "17bb261bf36fa7d83f4c294f834e91256769097b3cb505d44831e0a179ac647f" - -[[package]] -name = "num-order" -version = "1.2.0" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "537b596b97c40fcf8056d153049eb22f481c17ebce72a513ec9286e4986d1bb6" -dependencies = [ - "num-modular", -] - -[[package]] -name = "num-traits" -version = "0.2.19" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "071dfc062690e90b734c0b2273ce72ad0ffa95f0c74596bc250dcfd960262841" -dependencies = [ - "autocfg", -] - -[[package]] -name = "once_cell" -version = "1.21.4" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "9f7c3e4beb33f85d45ae3e3a1792185706c8e16d043238c593331cc7cd313b50" - -[[package]] -name = "once_cell_polyfill" -version = "1.70.2" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "384b8ab6d37215f3c5301a95a4accb5d64aa607f1fcb26a11b5303878451b4fe" - -[[package]] -name = "oneshot" -version = "0.1.13" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "269bca4c2591a28585d6bf10d9ed0332b7d76900a1b02bec41bdc3a2cdcda107" - -[[package]] -name = "oorandom" -version = "11.1.5" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "d6790f58c7ff633d8771f42965289203411a5e5c68388703c06e14f24770b41e" - -[[package]] -name = "openssl-probe" -version = "0.2.1" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "7c87def4c32ab89d880effc9e097653c8da5d6ef28e6b539d313baaacfbafcbe" - -[[package]] -name = "option-ext" -version = "0.2.0" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "04744f49eae99ab78e0d5c0b603ab218f515ea8cfe5a456d7629ad883a3b6e7d" - -[[package]] -name = "ordered-float" -version = "5.3.0" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "b7d950ca161dc355eaf28f82b11345ed76c6e1f6eb1f4f4479e0323b9e2fbd0e" -dependencies = [ - "num-traits", -] - -[[package]] -name = "ownedbytes" -version = "0.9.0" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "2fbd56f7631767e61784dc43f8580f403f4475bd4aaa4da003e6295e1bab4a7e" -dependencies = [ - "stable_deref_trait", -] - -[[package]] -name = "oxigraph" -version = "0.4.11" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "86b57a5334aab94d88e1d24b238c093c5efb0d309614b16ac920f23ad77ee77d" -dependencies = [ - "dashmap", - "getrandom 0.2.17", - "libc", - "oxiri", - "oxrdf", - "oxrdfio", - "oxrocksdb-sys", - "oxsdatatypes", - "rand 0.8.6", - "rustc-hash", - "siphasher", - "sparesults", - "spareval", - "spargebra", - "thiserror 2.0.18", -] - -[[package]] -name = "oxilangtag" -version = "0.1.5" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "23f3f87617a86af77fa3691e6350483e7154c2ead9f1261b75130e21ca0f8acb" -dependencies = [ - "serde", -] - -[[package]] -name = "oxiri" -version = "0.2.11" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "54b4ed3a7192fa19f5f48f99871f2755047fabefd7f222f12a1df1773796a102" - -[[package]] -name = "oxjsonld" -version = "0.1.0" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "13a1a66dc569350f3f4e5eff8a8e1a72b0c9e6ad395bb5805493cb7a2fda185f" -dependencies = [ - "json-event-parser", - "oxiri", - "oxrdf", - "thiserror 2.0.18", -] - -[[package]] -name = "oxrdf" -version = "0.2.4" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "a04761319ef84de1f59782f189d072cbfc3a9a40c4e8bded8667202fbd35b02a" -dependencies = [ - "oxilangtag", - "oxiri", - "oxsdatatypes", - "rand 0.8.6", - "thiserror 2.0.18", -] - -[[package]] -name = "oxrdfio" -version = "0.1.8" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "14d33dd87769786a0bb7de342865e33bf0c6e9872fa76f1ede23e944fdc77898" -dependencies = [ - "oxjsonld", - "oxrdf", - "oxrdfxml", - "oxttl", - "thiserror 2.0.18", -] - -[[package]] -name = "oxrdfxml" -version = "0.1.7" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "d8d4bf9c5331127f01efbd1245d90fd75b7c546a97cb3e95461121ce1ad5b1c8" -dependencies = [ - "oxilangtag", - "oxiri", - "oxrdf", - "quick-xml", - "thiserror 2.0.18", -] - -[[package]] -name = "oxrocksdb-sys" -version = "0.4.11" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "16430f45934d678cb6f9823e7c1bfdbdce9025f670ad85b642e46ffe5609e6ff" -dependencies = [ - "bindgen", - "cc", - "libc", -] - -[[package]] -name = "oxsdatatypes" -version = "0.2.2" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "06fa874d87eae638daae9b4e3198864fe2cce68589f227c0b2cf5b62b1530516" -dependencies = [ - "thiserror 2.0.18", -] - -[[package]] -name = "oxttl" -version = "0.1.8" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "0d385f1776d7cace455ef6b7c54407838eff902ca897303d06eb12a26f4cf8a0" -dependencies = [ - "memchr", - "oxilangtag", - "oxiri", - "oxrdf", - "thiserror 2.0.18", -] - -[[package]] -name = "parking" -version = "2.2.1" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "f38d5652c16fde515bb1ecef450ab0f6a219d619a7274976324d5e377f7dceba" - -[[package]] -name = "parking_lot" -version = "0.12.5" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "93857453250e3077bd71ff98b6a65ea6621a19bb0f559a85248955ac12c45a1a" -dependencies = [ - "lock_api", - "parking_lot_core", -] - -[[package]] -name = "parking_lot_core" -version = "0.9.12" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "2621685985a2ebf1c516881c026032ac7deafcda1a2c9b7850dc81e3dfcb64c1" -dependencies = [ - "cfg-if", - "libc", - "redox_syscall", - "smallvec", - "windows-link", -] - -[[package]] -name = "peg" -version = "0.8.5" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "9928cfca101b36ec5163e70049ee5368a8a1c3c6efc9ca9c5f9cc2f816152477" -dependencies = [ - "peg-macros", - "peg-runtime", -] - -[[package]] -name = "peg-macros" -version = "0.8.5" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "6298ab04c202fa5b5d52ba03269fb7b74550b150323038878fe6c372d8280f71" -dependencies = [ - "peg-runtime", - "proc-macro2", - "quote", -] - -[[package]] -name = "peg-runtime" -version = "0.8.5" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "132dca9b868d927b35b5dd728167b2dee150eb1ad686008fc71ccb298b776fca" - -[[package]] -name = "percent-encoding" -version = "2.3.2" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "9b4f627cb1b25917193a259e49bdad08f671f8d9708acfd5fe0a8c1455d87220" - -[[package]] -name = "pest" -version = "2.8.6" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "e0848c601009d37dfa3430c4666e147e49cdcf1b92ecd3e63657d8a5f19da662" -dependencies = [ - "memchr", - "ucd-trie", -] - -[[package]] -name = "pest_derive" -version = "2.8.6" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "11f486f1ea21e6c10ed15d5a7c77165d0ee443402f0780849d1768e7d9d6fe77" -dependencies = [ - "pest", - "pest_generator", -] - -[[package]] -name = "pest_generator" -version = "2.8.6" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "8040c4647b13b210a963c1ed407c1ff4fdfa01c31d6d2a098218702e6664f94f" -dependencies = [ - "pest", - "pest_meta", - "proc-macro2", - "quote", - "syn", -] - -[[package]] -name = "pest_meta" -version = "2.8.6" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "89815c69d36021a140146f26659a81d6c2afa33d216d736dd4be5381a7362220" -dependencies = [ - "pest", - "sha2", -] - -[[package]] -name = "pin-project" -version = "1.1.11" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "f1749c7ed4bcaf4c3d0a3efc28538844fb29bcdd7d2b67b2be7e20ba861ff517" -dependencies = [ - "pin-project-internal", -] - -[[package]] -name = "pin-project-internal" -version = "1.1.11" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "d9b20ed30f105399776b9c883e68e536ef602a16ae6f596d2c473591d6ad64c6" -dependencies = [ - "proc-macro2", - "quote", - "syn", -] - -[[package]] -name = "pin-project-lite" -version = "0.2.17" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "a89322df9ebe1c1578d689c92318e070967d1042b512afbe49518723f4e6d5cd" - -[[package]] -name = "plotters" -version = "0.3.7" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "5aeb6f403d7a4911efb1e33402027fc44f29b5bf6def3effcc22d7bb75f2b747" -dependencies = [ - "num-traits", - "plotters-backend", - "plotters-svg", - "wasm-bindgen", - "web-sys", -] - -[[package]] -name = "plotters-backend" -version = "0.3.7" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "df42e13c12958a16b3f7f4386b9ab1f3e7933914ecea48da7139435263a4172a" - -[[package]] -name = "plotters-svg" -version = "0.3.7" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "51bae2ac328883f7acdfea3d66a7c35751187f870bc81f94563733a154d7a670" -dependencies = [ - "plotters-backend", -] - -[[package]] -name = "polling" -version = "3.11.0" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "5d0e4f59085d47d8241c88ead0f274e8a0cb551f3625263c05eb8dd897c34218" -dependencies = [ - "cfg-if", - "concurrent-queue", - "hermit-abi", - "pin-project-lite", - "rustix", - "windows-sys 0.61.2", -] - -[[package]] -name = "portable-atomic" -version = "1.13.1" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "c33a9471896f1c69cecef8d20cbe2f7accd12527ce60845ff44c153bb2a21b49" - -[[package]] -name = "portable-atomic-util" -version = "0.2.7" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "c2a106d1259c23fac8e543272398ae0e3c0b8d33c88ed73d0cc71b0f1d902618" -dependencies = [ - "portable-atomic", -] - -[[package]] -name = "potential_utf" -version = "0.1.5" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "0103b1cef7ec0cf76490e969665504990193874ea05c85ff9bab8b911d0a0564" -dependencies = [ - "zerovec", -] - -[[package]] -name = "powerfmt" -version = "0.2.0" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "439ee305def115ba05938db6eb1644ff94165c5ab5e9420d1c1bcedbba909391" - -[[package]] -name = "ppv-lite86" -version = "0.2.21" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "85eae3c4ed2f50dcfe72643da4befc30deadb458a9b590d720cde2f2b1e97da9" -dependencies = [ - "zerocopy", -] - -[[package]] -name = "prettyplease" -version = "0.2.37" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "479ca8adacdd7ce8f1fb39ce9ecccbfe93a3f1344b3d0d97f20bc0196208f62b" -dependencies = [ - "proc-macro2", - "syn", -] - -[[package]] -name = "proc-macro-crate" -version = "3.5.0" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "e67ba7e9b2b56446f1d419b1d807906278ffa1a658a8a5d8a39dcb1f5a78614f" -dependencies = [ - "toml_edit", -] - -[[package]] -name = "proc-macro2" -version = "1.0.106" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "8fd00f0bb2e90d81d1044c2b32617f68fcb9fa3bb7640c23e9c748e53fb30934" -dependencies = [ - "unicode-ident", -] - -[[package]] -name = "prometheus" -version = "0.14.0" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "3ca5326d8d0b950a9acd87e6a3f94745394f62e4dae1b1ee22b2bc0c394af43a" -dependencies = [ - "cfg-if", - "fnv", - "lazy_static", - "memchr", - "parking_lot", - "protobuf", - "thiserror 2.0.18", -] - -[[package]] -name = "proptest" -version = "1.11.0" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "4b45fcc2344c680f5025fe57779faef368840d0bd1f42f216291f0dc4ace4744" -dependencies = [ - "bit-set", - "bit-vec", - "bitflags", - "num-traits", - "rand 0.9.4", - "rand_chacha 0.9.0", - "rand_xorshift", - "regex-syntax", - "rusty-fork", - "tempfile", - "unarray", -] - -[[package]] -name = "prost" -version = "0.14.3" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "d2ea70524a2f82d518bce41317d0fae74151505651af45faf1ffbd6fd33f0568" -dependencies = [ - "bytes", - "prost-derive", -] - -[[package]] -name = "prost-derive" -version = "0.14.3" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "27c6023962132f4b30eb4c172c91ce92d933da334c59c23cddee82358ddafb0b" -dependencies = [ - "anyhow", - "itertools 0.14.0", - "proc-macro2", - "quote", - "syn", -] - -[[package]] -name = "prost-types" -version = "0.14.3" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "8991c4cbdb8bc5b11f0b074ffe286c30e523de90fee5ba8132f1399f23cb3dd7" -dependencies = [ - "prost", -] - -[[package]] -name = "protobuf" -version = "3.7.2" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "d65a1d4ddae7d8b5de68153b48f6aa3bba8cb002b243dbdbc55a5afbc98f99f4" -dependencies = [ - "once_cell", - "protobuf-support", - "thiserror 1.0.69", -] - -[[package]] -name = "protobuf-support" -version = "3.7.2" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "3e36c2f31e0a47f9280fb347ef5e461ffcd2c52dd520d8e216b52f93b0b0d7d6" -dependencies = [ - "thiserror 1.0.69", -] - -[[package]] -name = "quick-error" -version = "1.2.3" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "a1d01941d82fa2ab50be1e79e6714289dd7cde78eba4c074bc5a4374f650dfe0" - -[[package]] -name = "quick-xml" -version = "0.37.5" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "331e97a1af0bf59823e6eadffe373d7b27f485be8748f71471c662c1f269b7fb" -dependencies = [ - "memchr", -] - -[[package]] -name = "quote" -version = "1.0.45" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "41f2619966050689382d2b44f664f4bc593e129785a36d6ee376ddf37259b924" -dependencies = [ - "proc-macro2", -] - -[[package]] -name = "r-efi" -version = "5.3.0" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "69cdb34c158ceb288df11e18b4bd39de994f6657d83847bdffdbd7f346754b0f" - -[[package]] -name = "r-efi" -version = "6.0.0" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "f8dcc9c7d52a811697d2151c701e0d08956f92b0e24136cf4cf27b57a6a0d9bf" - -[[package]] -name = "radix_trie" -version = "0.2.1" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "c069c179fcdc6a2fe24d8d18305cf085fdbd4f922c041943e203685d6a1c58fd" -dependencies = [ - "endian-type", - "nibble_vec", -] - -[[package]] -name = "rand" -version = "0.8.6" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "5ca0ecfa931c29007047d1bc58e623ab12e5590e8c7cc53200d5202b69266d8a" -dependencies = [ - "libc", - "rand_chacha 0.3.1", - "rand_core 0.6.4", -] - -[[package]] -name = "rand" -version = "0.9.4" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "44c5af06bb1b7d3216d91932aed5265164bf384dc89cd6ba05cf59a35f5f76ea" -dependencies = [ - "rand_chacha 0.9.0", - "rand_core 0.9.5", -] - -[[package]] -name = "rand_chacha" -version = "0.3.1" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "e6c10a63a0fa32252be49d21e7709d4d4baf8d231c2dbce1eaa8141b9b127d88" -dependencies = [ - "ppv-lite86", - "rand_core 0.6.4", -] - -[[package]] -name = "rand_chacha" -version = "0.9.0" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "d3022b5f1df60f26e1ffddd6c66e8aa15de382ae63b3a0c1bfc0e4d3e3f325cb" -dependencies = [ - "ppv-lite86", - "rand_core 0.9.5", -] - -[[package]] -name = "rand_core" -version = "0.6.4" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "ec0be4795e2f6a28069bec0b5ff3e2ac9bafc99e6a9a7dc3547996c5c816922c" -dependencies = [ - "getrandom 0.2.17", -] - -[[package]] -name = "rand_core" -version = "0.9.5" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "76afc826de14238e6e8c374ddcc1fa19e374fd8dd986b0d2af0d02377261d83c" -dependencies = [ - "getrandom 0.3.4", -] - -[[package]] -name = "rand_xorshift" -version = "0.4.0" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "513962919efc330f829edb2535844d1b912b0fbe2ca165d613e4e8788bb05a5a" -dependencies = [ - "rand_core 0.9.5", -] - -[[package]] -name = "rawpointer" -version = "0.2.1" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "60a357793950651c4ed0f3f52338f53b2f809f32d83a07f72909fa13e4c6c1e3" - -[[package]] -name = "rayon" -version = "1.12.0" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "fb39b166781f92d482534ef4b4b1b2568f42613b53e5b6c160e24cfbfa30926d" -dependencies = [ - "either", - "rayon-core", -] - -[[package]] -name = "rayon-core" -version = "1.13.0" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "22e18b0f0062d30d4230b2e85ff77fdfe4326feb054b9783a3460d8435c8ab91" -dependencies = [ - "crossbeam-deque", - "crossbeam-utils", -] - -[[package]] -name = "redb" -version = "3.1.3" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "4ba239c1c1693315d3cc0e601db3b3965543afbf48c41730fdca2f069f510f4a" -dependencies = [ - "libc", -] - -[[package]] -name = "redox_syscall" -version = "0.5.18" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "ed2bf2547551a7053d6fdfafda3f938979645c44812fbfcda098faae3f1a362d" -dependencies = [ - "bitflags", -] - -[[package]] -name = "redox_users" -version = "0.5.2" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "a4e608c6638b9c18977b00b475ac1f28d14e84b27d8d42f70e0bf1e3dec127ac" -dependencies = [ - "getrandom 0.2.17", - "libredox", - "thiserror 2.0.18", -] - -[[package]] -name = "regex" -version = "1.12.3" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "e10754a14b9137dd7b1e3e5b0493cc9171fdd105e0ab477f51b72e7f3ac0e276" -dependencies = [ - "aho-corasick", - "memchr", - "regex-automata", - "regex-syntax", -] - -[[package]] -name = "regex-automata" -version = "0.4.14" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "6e1dd4122fc1595e8162618945476892eefca7b88c52820e74af6262213cae8f" -dependencies = [ - "aho-corasick", - "memchr", - "regex-syntax", -] - -[[package]] -name = "regex-lite" -version = "0.1.9" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "cab834c73d247e67f4fae452806d17d3c7501756d98c8808d7c9c7aa7d18f973" - -[[package]] -name = "regex-syntax" -version = "0.8.10" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "dc897dd8d9e8bd1ed8cdad82b5966c3e0ecae09fb1907d58efaa013543185d0a" - -[[package]] -name = "reqwest" -version = "0.13.2" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "ab3f43e3283ab1488b624b44b0e988d0acea0b3214e694730a055cb6b2efa801" -dependencies = [ - "base64", - "bytes", - "futures-channel", - "futures-core", - "futures-util", - "h2", - "http", - "http-body", - "http-body-util", - "hyper", - "hyper-rustls", - "hyper-util", - "js-sys", - "log", - "percent-encoding", - "pin-project-lite", - "rustls", - "rustls-pki-types", - "rustls-platform-verifier", - "serde", - "serde_json", - "serde_urlencoded", - "sync_wrapper", - "tokio", - "tokio-rustls", - "tower", - "tower-http", - "tower-service", - "url", - "wasm-bindgen", - "wasm-bindgen-futures", - "web-sys", -] - -[[package]] -name = "ring" -version = "0.17.14" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "a4689e6c2294d81e88dc6261c768b63bc4fcdb852be6d1352498b114f61383b7" -dependencies = [ - "cc", - "cfg-if", - "getrandom 0.2.17", - "libc", - "untrusted", - "windows-sys 0.52.0", -] - -[[package]] -name = "rustc-hash" -version = "2.1.2" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "94300abf3f1ae2e2b8ffb7b58043de3d399c73fa6f4b73826402a5c457614dbe" - -[[package]] -name = "rustix" -version = "1.1.4" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "b6fe4565b9518b83ef4f91bb47ce29620ca828bd32cb7e408f0062e9930ba190" -dependencies = [ - "bitflags", - "errno", - "libc", - "linux-raw-sys", - "windows-sys 0.61.2", -] - -[[package]] -name = "rustler" -version = "0.36.2" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "e3fe55230a9c379733dd38ee67d4072fa5c558b2e22b76b0e7f924390456e003" -dependencies = [ - "inventory", - "libloading", - "regex-lite", - "rustler_codegen", -] - -[[package]] -name = "rustler_codegen" -version = "0.36.2" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "eb3b8de901ae61418e2036245d28e41ef58080d04f40b68430471ae36a4e84ed" -dependencies = [ - "heck", - "inventory", - "proc-macro2", - "quote", - "syn", -] - -[[package]] -name = "rustls" -version = "0.23.38" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "69f9466fb2c14ea04357e91413efb882e2a6d4a406e625449bc0a5d360d53a21" -dependencies = [ - "log", - "once_cell", - "ring", - "rustls-pki-types", - "rustls-webpki", - "subtle", - "zeroize", -] - -[[package]] -name = "rustls-native-certs" -version = "0.8.3" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "612460d5f7bea540c490b2b6395d8e34a953e52b491accd6c86c8164c5932a63" -dependencies = [ - "openssl-probe", - "rustls-pki-types", - "schannel", - "security-framework", -] - -[[package]] -name = "rustls-pemfile" -version = "2.2.0" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "dce314e5fee3f39953d46bb63bb8a46d40c2f8fb7cc5a3b6cab2bde9721d6e50" -dependencies = [ - "rustls-pki-types", -] - -[[package]] -name = "rustls-pki-types" -version = "1.14.0" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "be040f8b0a225e40375822a563fa9524378b9d63112f53e19ffff34df5d33fdd" -dependencies = [ - "zeroize", -] - -[[package]] -name = "rustls-platform-verifier" -version = "0.6.2" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "1d99feebc72bae7ab76ba994bb5e121b8d83d910ca40b36e0921f53becc41784" -dependencies = [ - "core-foundation", - "core-foundation-sys", - "jni", - "log", - "once_cell", - "rustls", - "rustls-native-certs", - "rustls-platform-verifier-android", - "rustls-webpki", - "security-framework", - "security-framework-sys", - "webpki-root-certs", - "windows-sys 0.61.2", -] - -[[package]] -name = "rustls-platform-verifier-android" -version = "0.1.1" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "f87165f0995f63a9fbeea62b64d10b4d9d8e78ec6d7d51fb2125fda7bb36788f" - -[[package]] -name = "rustls-webpki" -version = "0.103.13" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "61c429a8649f110dddef65e2a5ad240f747e85f7758a6bccc7e5777bd33f756e" -dependencies = [ - "ring", - "rustls-pki-types", - "untrusted", -] - -[[package]] -name = "rustversion" -version = "1.0.22" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "b39cdef0fa800fc44525c84ccb54a029961a8215f9619753635a9c0d2538d46d" - -[[package]] -name = "rusty-fork" -version = "0.3.1" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "cc6bf79ff24e648f6da1f8d1f011e9cac26491b619e6b9280f2b47f1774e6ee2" -dependencies = [ - "fnv", - "quick-error", - "tempfile", - "wait-timeout", -] - -[[package]] -name = "rustyline" -version = "15.0.0" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "2ee1e066dc922e513bda599c6ccb5f3bb2b0ea5870a579448f2622993f0a9a2f" -dependencies = [ - "bitflags", - "cfg-if", - "clipboard-win", - "fd-lock", - "home", - "libc", - "log", - "memchr", - "nix", - "radix_trie", - "unicode-segmentation", - "unicode-width", - "utf8parse", - "windows-sys 0.59.0", -] - -[[package]] -name = "rustyline-derive" -version = "0.11.1" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "5d66de233f908aebf9cc30ac75ef9103185b4b715c6f2fb7a626aa5e5ede53ab" -dependencies = [ - "proc-macro2", - "quote", - "syn", -] - -[[package]] -name = "ryu" -version = "1.0.23" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "9774ba4a74de5f7b1c1451ed6cd5285a32eddb5cccb8cc655a4e50009e06477f" - -[[package]] -name = "same-file" -version = "1.0.6" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "93fc1dc3aaa9bfed95e02e6eadabb4baf7e3078b0bd1b4d7b6b0b68378900502" -dependencies = [ - "winapi-util", -] - -[[package]] -name = "schannel" -version = "0.1.29" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "91c1b7e4904c873ef0710c1f407dde2e6287de2bebc1bbbf7d430bb7cbffd939" -dependencies = [ - "windows-sys 0.61.2", -] - -[[package]] -name = "scopeguard" -version = "1.2.0" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "94143f37725109f92c262ed2cf5e59bce7498c01bcc1502d7b9afe439a4e9f49" - -[[package]] -name = "security-framework" -version = "3.7.0" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "b7f4bc775c73d9a02cde8bf7b2ec4c9d12743edf609006c7facc23998404cd1d" -dependencies = [ - "bitflags", - "core-foundation", - "core-foundation-sys", - "libc", - "security-framework-sys", -] - -[[package]] -name = "security-framework-sys" -version = "2.17.0" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "6ce2691df843ecc5d231c0b14ece2acc3efb62c0a398c7e1d875f3983ce020e3" -dependencies = [ - "core-foundation-sys", - "libc", -] - -[[package]] -name = "semver" -version = "1.0.28" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "8a7852d02fc848982e0c167ef163aaff9cd91dc640ba85e263cb1ce46fae51cd" - -[[package]] -name = "serde" -version = "1.0.228" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "9a8e94ea7f378bd32cbbd37198a4a91436180c5bb472411e48b5ec2e2124ae9e" -dependencies = [ - "serde_core", - "serde_derive", -] - -[[package]] -name = "serde_core" -version = "1.0.228" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "41d385c7d4ca58e59fc732af25c3983b67ac852c1a25000afe1175de458b67ad" -dependencies = [ - "serde_derive", -] - -[[package]] -name = "serde_derive" -version = "1.0.228" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "d540f220d3187173da220f885ab66608367b6574e925011a9353e4badda91d79" -dependencies = [ - "proc-macro2", - "quote", - "syn", -] - -[[package]] -name = "serde_json" -version = "1.0.149" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "83fc039473c5595ace860d8c4fafa220ff474b3fc6bfdb4293327f1a37e94d86" -dependencies = [ - "itoa", - "memchr", - "serde", - "serde_core", - "zmij", -] - -[[package]] -name = "serde_path_to_error" -version = "0.1.20" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "10a9ff822e371bb5403e391ecd83e182e0e77ba7f6fe0160b795797109d1b457" -dependencies = [ - "itoa", - "serde", - "serde_core", -] - -[[package]] -name = "serde_urlencoded" -version = "0.7.1" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "d3491c14715ca2294c4d6a88f15e84739788c1d030eed8c110436aafdaa2f3fd" -dependencies = [ - "form_urlencoded", - "itoa", - "ryu", - "serde", -] - -[[package]] -name = "sha1" -version = "0.10.6" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "e3bf829a2d51ab4a5ddf1352d8470c140cadc8301b2ae1789db023f01cedd6ba" -dependencies = [ - "cfg-if", - "cpufeatures", - "digest", -] - -[[package]] -name = "sha2" -version = "0.10.9" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "a7507d819769d01a365ab707794a4084392c824f54a7a6a7862f8c3d0892b283" -dependencies = [ - "cfg-if", - "cpufeatures", - "digest", -] - -[[package]] -name = "sharded-slab" -version = "0.1.7" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "f40ca3c46823713e0d4209592e8d6e826aa57e928f09752619fc696c499637f6" -dependencies = [ - "lazy_static", -] - -[[package]] -name = "shlex" -version = "1.3.0" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "0fda2ff0d084019ba4d7c6f371c95d8fd75ce3524c3cb8fb653a3023f6323e64" - -[[package]] -name = "signal-hook-registry" -version = "1.4.8" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "c4db69cba1110affc0e9f7bcd48bbf87b3f4fc7c61fc9155afd4c469eb3d6c1b" -dependencies = [ - "errno", - "libc", -] - -[[package]] -name = "siphasher" -version = "1.0.2" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "b2aa850e253778c88a04c3d7323b043aeda9d3e30d5971937c1855769763678e" - -[[package]] -name = "sketches-ddsketch" -version = "0.4.0" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "05e40b6cf54d988dc1a2223531b969c9a9e30906ad90ef64890c27b4bfbb46ea" -dependencies = [ - "serde", -] - -[[package]] -name = "slab" -version = "0.4.12" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "0c790de23124f9ab44544d7ac05d60440adc586479ce501c1d6d7da3cd8c9cf5" - -[[package]] -name = "smallvec" -version = "1.15.1" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "67b1b7a3b5fe4f1376887184045fcf45c69e92af734b7aaddc05fb777b6fbd03" - -[[package]] -name = "socket2" -version = "0.6.3" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "3a766e1110788c36f4fa1c2b71b387a7815aa65f88ce0229841826633d93723e" -dependencies = [ - "libc", - "windows-sys 0.61.2", -] - -[[package]] -name = "sparesults" -version = "0.2.5" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "f478f5ead16b6136bccee7a52ea43a615f8512086708f515e26ce33e0b184036" -dependencies = [ - "json-event-parser", - "memchr", - "oxrdf", - "quick-xml", - "thiserror 2.0.18", -] - -[[package]] -name = "spareval" -version = "0.1.4" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "f9d8ff5f1159e7416ed99160b962fa780851dddee133ef56e6b08a94023ea2c7" -dependencies = [ - "hex", - "json-event-parser", - "md-5", - "oxiri", - "oxrdf", - "oxsdatatypes", - "rand 0.8.6", - "regex", - "rustc-hash", - "sha1", - "sha2", - "sparesults", - "spargebra", - "sparopt", - "thiserror 2.0.18", -] - -[[package]] -name = "spargebra" -version = "0.3.5" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "8907e262be4b4b363218f4688f5654d423a958aa4b8d7c7a7f898be591fa474e" -dependencies = [ - "oxilangtag", - "oxiri", - "oxrdf", - "peg", - "rand 0.8.6", - "thiserror 2.0.18", -] - -[[package]] -name = "sparopt" -version = "0.2.1" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "1790bbdf13560c2afc245ab0f82a489003b3918e668ebd45c65fe46bfd7a1763" -dependencies = [ - "oxrdf", - "rand 0.8.6", - "spargebra", -] - -[[package]] -name = "spin" -version = "0.9.8" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "6980e8d7511241f8acf4aebddbb1ff938df5eebe98691418c4468d0b72a96a67" - -[[package]] -name = "stable_deref_trait" -version = "1.2.1" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "6ce2be8dc25455e1f91df71bfa12ad37d7af1092ae736f3a6cd0e37bc7810596" - -[[package]] -name = "static_assertions_next" -version = "1.1.2" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "d7beae5182595e9a8b683fa98c4317f956c9a2dec3b9716990d20023cc60c766" - -[[package]] -name = "strsim" -version = "0.11.1" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "7da8b5736845d9f2fcb837ea5d9e2628564b3b043a70948a3f0b778838c5fb4f" - -[[package]] -name = "strum" -version = "0.27.2" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "af23d6f6c1a224baef9d3f61e287d2761385a5b88fdab4eb4c6f11aeb54c4bcf" -dependencies = [ - "strum_macros", -] - -[[package]] -name = "strum_macros" -version = "0.27.2" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "7695ce3845ea4b33927c055a39dc438a45b059f7c1b3d91d38d10355fb8cbca7" -dependencies = [ - "heck", - "proc-macro2", - "quote", - "syn", -] - -[[package]] -name = "subtle" -version = "2.6.1" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "13c2bddecc57b384dee18652358fb23172facb8a2c51ccc10d74c157bdea3292" - -[[package]] -name = "syn" -version = "2.0.117" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "e665b8803e7b1d2a727f4023456bbbbe74da67099c585258af0ad9c5013b9b99" -dependencies = [ - "proc-macro2", - "quote", - "unicode-ident", -] - -[[package]] -name = "sync_wrapper" -version = "1.0.2" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "0bf256ce5efdfa370213c1dabab5935a12e49f2c58d15e9eac2870d3b4f27263" -dependencies = [ - "futures-core", -] - -[[package]] -name = "synstructure" -version = "0.13.2" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "728a70f3dbaf5bab7f0c4b1ac8d7ae5ea60a4b5549c8a5914361c99147a709d2" -dependencies = [ - "proc-macro2", - "quote", - "syn", -] - -[[package]] -name = "tantivy" -version = "0.26.0" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "778da245841522199d512d19511b041425d8cff3a8f262b4e1516fceb050289a" -dependencies = [ - "aho-corasick", - "arc-swap", - "base64", - "bitpacking", - "bon", - "byteorder", - "census", - "crc32fast", - "crossbeam-channel", - "datasketches", - "downcast-rs", - "fastdivide", - "fnv", - "fs4", - "htmlescape", - "itertools 0.14.0", - "levenshtein_automata", - "log", - "lru", - "lz4_flex", - "measure_time", - "memmap2", - "once_cell", - "oneshot", - "rayon", - "regex", - "rustc-hash", - "serde", - "serde_json", - "sketches-ddsketch", - "smallvec", - "tantivy-bitpacker", - "tantivy-columnar", - "tantivy-common", - "tantivy-fst", - "tantivy-query-grammar", - "tantivy-stacker", - "tantivy-tokenizer-api", - "tempfile", - "thiserror 2.0.18", - "time", - "typetag", - "uuid", - "winapi", -] - -[[package]] -name = "tantivy-bitpacker" -version = "0.10.0" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "4fed3d674429bcd2de5d0a6d1aa5495fed8afd9c5ecce993019caf7615f53fa4" -dependencies = [ - "bitpacking", -] - -[[package]] -name = "tantivy-columnar" -version = "0.7.0" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "c57166f5bcfd478f370ab8445afb4678dce44801fa5ce5c451aaf8595583c5dc" -dependencies = [ - "downcast-rs", - "fastdivide", - "itertools 0.14.0", - "serde", - "tantivy-bitpacker", - "tantivy-common", - "tantivy-sstable", - "tantivy-stacker", -] - -[[package]] -name = "tantivy-common" -version = "0.11.0" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "bbf10915aa75da3c3b0d58b58853d2e889efbaf32d4982a4c3715dde6bba23e5" -dependencies = [ - "async-trait", - "byteorder", - "ownedbytes", - "serde", - "time", -] - -[[package]] -name = "tantivy-fst" -version = "0.5.0" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "d60769b80ad7953d8a7b2c70cdfe722bbcdcac6bccc8ac934c40c034d866fc18" -dependencies = [ - "byteorder", - "regex-syntax", - "utf8-ranges", -] - -[[package]] -name = "tantivy-query-grammar" -version = "0.26.0" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "dfadb8526b6da90704feb293b0701a6aae62ea14983143344be2dc5ce30f1d82" -dependencies = [ - "fnv", - "nom", - "ordered-float", - "serde", - "serde_json", -] - -[[package]] -name = "tantivy-sstable" -version = "0.7.0" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "8a2cfc3ac5164cbadc28965ffb145a8f47582a60ae5897859ad8d4316596c606" -dependencies = [ - "futures-util", - "itertools 0.14.0", - "tantivy-bitpacker", - "tantivy-common", - "tantivy-fst", -] - -[[package]] -name = "tantivy-stacker" -version = "0.7.0" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "6cbb051742da9d53ca9e8fff43a9b10e319338b24e2c0e15d0372df19ffeb951" -dependencies = [ - "murmurhash32", - "tantivy-common", -] - -[[package]] -name = "tantivy-tokenizer-api" -version = "0.7.0" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "eac258c2c6390673f2685813afeeafcb8c4e0ee7de8dd3fc46838dcc37263f98" -dependencies = [ - "serde", -] - -[[package]] -name = "tempfile" -version = "3.27.0" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "32497e9a4c7b38532efcdebeef879707aa9f794296a4f0244f6f69e9bc8574bd" -dependencies = [ - "fastrand", - "getrandom 0.4.2", - "once_cell", - "rustix", - "windows-sys 0.61.2", -] - -[[package]] -name = "thiserror" -version = "1.0.69" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "b6aaf5339b578ea85b50e080feb250a3e8ae8cfcdff9a461c9ec2904bc923f52" -dependencies = [ - "thiserror-impl 1.0.69", -] - -[[package]] -name = "thiserror" -version = "2.0.18" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "4288b5bcbc7920c07a1149a35cf9590a2aa808e0bc1eafaade0b80947865fbc4" -dependencies = [ - "thiserror-impl 2.0.18", -] - -[[package]] -name = "thiserror-impl" -version = "1.0.69" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "4fee6c4efc90059e10f81e6d42c60a18f76588c3d74cb83a0b242a2b6c7504c1" -dependencies = [ - "proc-macro2", - "quote", - "syn", -] - -[[package]] -name = "thiserror-impl" -version = "2.0.18" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "ebc4ee7f67670e9b64d05fa4253e753e016c6c95ff35b89b7941d6b856dec1d5" -dependencies = [ - "proc-macro2", - "quote", - "syn", -] - -[[package]] -name = "thread_local" -version = "1.1.9" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "f60246a4944f24f6e018aa17cdeffb7818b76356965d03b07d6a9886e8962185" -dependencies = [ - "cfg-if", -] - -[[package]] -name = "time" -version = "0.3.47" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "743bd48c283afc0388f9b8827b976905fb217ad9e647fae3a379a9283c4def2c" -dependencies = [ - "deranged", - "itoa", - "num-conv", - "powerfmt", - "serde_core", - "time-core", - "time-macros", -] - -[[package]] -name = "time-core" -version = "0.1.8" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "7694e1cfe791f8d31026952abf09c69ca6f6fa4e1a1229e18988f06a04a12dca" - -[[package]] -name = "time-macros" -version = "0.2.27" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "2e70e4c5a0e0a8a4823ad65dfe1a6930e4f4d756dcd9dd7939022b5e8c501215" -dependencies = [ - "num-conv", - "time-core", -] - -[[package]] -name = "tinystr" -version = "0.8.3" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "c8323304221c2a851516f22236c5722a72eaa19749016521d6dff0824447d96d" -dependencies = [ - "displaydoc", - "zerovec", -] - -[[package]] -name = "tinytemplate" -version = "1.2.1" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "be4d6b5f19ff7664e8c98d03e2139cb510db9b0a60b55f8e8709b689d939b6bc" -dependencies = [ - "serde", - "serde_json", -] - -[[package]] -name = "tokio" -version = "1.52.1" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "b67dee974fe86fd92cc45b7a95fdd2f99a36a6d7b0d431a231178d3d670bbcc6" -dependencies = [ - "bytes", - "libc", - "mio", - "parking_lot", - "pin-project-lite", - "signal-hook-registry", - "socket2", - "tokio-macros", - "windows-sys 0.61.2", -] - -[[package]] -name = "tokio-macros" -version = "2.7.0" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "385a6cb71ab9ab790c5fe8d67f1645e6c450a7ce006a33de03daa956cf70a496" -dependencies = [ - "proc-macro2", - "quote", - "syn", -] - -[[package]] -name = "tokio-rustls" -version = "0.26.4" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "1729aa945f29d91ba541258c8df89027d5792d85a8841fb65e8bf0f4ede4ef61" -dependencies = [ - "rustls", - "tokio", -] - -[[package]] -name = "tokio-stream" -version = "0.1.18" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "32da49809aab5c3bc678af03902d4ccddea2a87d028d86392a4b1560c6906c70" -dependencies = [ - "futures-core", - "pin-project-lite", - "tokio", -] - -[[package]] -name = "tokio-test" -version = "0.4.5" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "3f6d24790a10a7af737693a3e8f1d03faef7e6ca0cc99aae5066f533766de545" -dependencies = [ - "futures-core", - "tokio", - "tokio-stream", -] - -[[package]] -name = "tokio-tungstenite" -version = "0.29.0" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "8f72a05e828585856dacd553fba484c242c46e391fb0e58917c942ee9202915c" -dependencies = [ - "futures-util", - "log", - "tokio", - "tungstenite", -] - -[[package]] -name = "tokio-util" -version = "0.7.18" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "9ae9cec805b01e8fc3fd2fe289f89149a9b66dd16786abd8b19cfa7b48cb0098" -dependencies = [ - "bytes", - "futures-core", - "futures-io", - "futures-sink", - "pin-project-lite", - "tokio", -] - -[[package]] -name = "toml_datetime" -version = "1.1.1+spec-1.1.0" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "3165f65f62e28e0115a00b2ebdd37eb6f3b641855f9d636d3cd4103767159ad7" -dependencies = [ - "serde_core", -] - -[[package]] -name = "toml_edit" -version = "0.25.11+spec-1.1.0" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "0b59c4d22ed448339746c59b905d24568fcbb3ab65a500494f7b8c3e97739f2b" -dependencies = [ - "indexmap", - "toml_datetime", - "toml_parser", - "winnow", -] - -[[package]] -name = "toml_parser" -version = "1.1.2+spec-1.1.0" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "a2abe9b86193656635d2411dc43050282ca48aa31c2451210f4202550afb7526" -dependencies = [ - "winnow", -] - -[[package]] -name = "tonic" -version = "0.14.5" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "fec7c61a0695dc1887c1b53952990f3ad2e3a31453e1f49f10e75424943a93ec" -dependencies = [ - "async-trait", - "axum", - "base64", - "bytes", - "h2", - "http", - "http-body", - "http-body-util", - "hyper", - "hyper-timeout", - "hyper-util", - "percent-encoding", - "pin-project", - "socket2", - "sync_wrapper", - "tokio", - "tokio-stream", - "tower", - "tower-layer", - "tower-service", - "tracing", -] - -[[package]] -name = "tonic-prost" -version = "0.14.5" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "a55376a0bbaa4975a3f10d009ad763d8f4108f067c7c2e74f3001fb49778d309" -dependencies = [ - "bytes", - "prost", - "tonic", -] - -[[package]] -name = "tower" -version = "0.5.3" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "ebe5ef63511595f1344e2d5cfa636d973292adc0eec1f0ad45fae9f0851ab1d4" -dependencies = [ - "futures-core", - "futures-util", - "indexmap", - "pin-project-lite", - "slab", - "sync_wrapper", - "tokio", - "tokio-util", - "tower-layer", - "tower-service", - "tracing", -] - -[[package]] -name = "tower-http" -version = "0.6.8" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "d4e6559d53cc268e5031cd8429d05415bc4cb4aefc4aa5d6cc35fbf5b924a1f8" -dependencies = [ - "bitflags", - "bytes", - "futures-util", - "http", - "http-body", - "iri-string", - "pin-project-lite", - "tower", - "tower-layer", - "tower-service", -] - -[[package]] -name = "tower-layer" -version = "0.3.3" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "121c2a6cda46980bb0fcd1647ffaf6cd3fc79a013de288782836f6df9c48780e" - -[[package]] -name = "tower-service" -version = "0.3.3" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "8df9b6e13f2d32c91b9bd719c00d1958837bc7dec474d94952798cc8e69eeec3" - -[[package]] -name = "tracing" -version = "0.1.44" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "63e71662fa4b2a2c3a26f570f037eb95bb1f85397f3cd8076caed2f026a6d100" -dependencies = [ - "log", - "pin-project-lite", - "tracing-attributes", - "tracing-core", -] - -[[package]] -name = "tracing-attributes" -version = "0.1.31" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "7490cfa5ec963746568740651ac6781f701c9c5ea257c58e057f3ba8cf69e8da" -dependencies = [ - "proc-macro2", - "quote", - "syn", -] - -[[package]] -name = "tracing-core" -version = "0.1.36" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "db97caf9d906fbde555dd62fa95ddba9eecfd14cb388e4f491a66d74cd5fb79a" -dependencies = [ - "once_cell", - "valuable", -] - -[[package]] -name = "tracing-log" -version = "0.2.0" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "ee855f1f400bd0e5c02d150ae5de3840039a3f54b025156404e34c23c03f47c3" -dependencies = [ - "log", - "once_cell", - "tracing-core", -] - -[[package]] -name = "tracing-serde" -version = "0.2.0" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "704b1aeb7be0d0a84fc9828cae51dab5970fee5088f83d1dd7ee6f6246fc6ff1" -dependencies = [ - "serde", - "tracing-core", -] - -[[package]] -name = "tracing-subscriber" -version = "0.3.23" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "cb7f578e5945fb242538965c2d0b04418d38ec25c79d160cd279bf0731c8d319" -dependencies = [ - "matchers", - "nu-ansi-term", - "once_cell", - "regex-automata", - "serde", - "serde_json", - "sharded-slab", - "smallvec", - "thread_local", - "tracing", - "tracing-core", - "tracing-log", - "tracing-serde", -] - -[[package]] -name = "try-lock" -version = "0.2.5" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "e421abadd41a4225275504ea4d6566923418b7f05506fbc9c0fe86ba7396114b" - -[[package]] -name = "tungstenite" -version = "0.29.0" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "6c01152af293afb9c7c2a57e4b559c5620b421f6d133261c60dd2d0cdb38e6b8" -dependencies = [ - "bytes", - "data-encoding", - "http", - "httparse", - "log", - "rand 0.9.4", - "sha1", - "thiserror 2.0.18", -] - -[[package]] -name = "typeid" -version = "1.0.3" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "bc7d623258602320d5c55d1bc22793b57daff0ec7efc270ea7d55ce1d5f5471c" - -[[package]] -name = "typenum" -version = "1.19.0" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "562d481066bde0658276a35467c4af00bdc6ee726305698a55b86e61d7ad82bb" - -[[package]] -name = "typetag" -version = "0.2.21" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "be2212c8a9b9bcfca32024de14998494cf9a5dfa59ea1b829de98bac374b86bf" -dependencies = [ - "erased-serde", - "inventory", - "once_cell", - "serde", - "typetag-impl", -] - -[[package]] -name = "typetag-impl" -version = "0.2.21" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "27a7a9b72ba121f6f1f6c3632b85604cac41aedb5ddc70accbebb6cac83de846" -dependencies = [ - "proc-macro2", - "quote", - "syn", -] - -[[package]] -name = "ucd-trie" -version = "0.1.7" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "2896d95c02a80c6d6a5d6e953d479f5ddf2dfdb6a244441010e373ac0fb88971" - -[[package]] -name = "unarray" -version = "0.1.4" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "eaea85b334db583fe3274d12b4cd1880032beab409c0d774be044d4480ab9a94" - -[[package]] -name = "unicode-ident" -version = "1.0.24" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "e6e4313cd5fcd3dad5cafa179702e2b244f760991f45397d14d4ebf38247da75" - -[[package]] -name = "unicode-segmentation" -version = "1.13.2" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "9629274872b2bfaf8d66f5f15725007f635594914870f65218920345aa11aa8c" - -[[package]] -name = "unicode-width" -version = "0.2.2" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "b4ac048d71ede7ee76d585517add45da530660ef4390e49b098733c6e897f254" - -[[package]] -name = "unicode-xid" -version = "0.2.6" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "ebc1c04c71510c7f702b52b7c350734c9ff1295c464a03335b00bb84fc54f853" - -[[package]] -name = "untrusted" -version = "0.9.0" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "8ecb6da28b8a351d773b68d5825ac39017e680750f980f3a1a85cd8dd28a47c1" - -[[package]] -name = "url" -version = "2.5.8" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "ff67a8a4397373c3ef660812acab3268222035010ab8680ec4215f38ba3d0eed" -dependencies = [ - "form_urlencoded", - "idna", - "percent-encoding", - "serde", -] - -[[package]] -name = "utf8-ranges" -version = "1.0.5" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "7fcfc827f90e53a02eaef5e535ee14266c1d569214c6aa70133a624d8a3164ba" - -[[package]] -name = "utf8_iter" -version = "1.0.4" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "b6c140620e7ffbb22c2dee59cafe6084a59b5ffc27a8859a5f0d494b5d52b6be" - -[[package]] -name = "utf8parse" -version = "0.2.2" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "06abde3611657adf66d383f00b093d7faecc7fa57071cce2578660c9f1010821" - -[[package]] -name = "uuid" -version = "1.23.1" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "ddd74a9687298c6858e9b88ec8935ec45d22e8fd5e6394fa1bd4e99a87789c76" -dependencies = [ - "getrandom 0.4.2", - "js-sys", - "serde_core", - "wasm-bindgen", -] - -[[package]] -name = "valuable" -version = "0.1.1" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "ba73ea9cf16a25df0c8caa16c51acb937d5712a8429db78a3ee29d5dcacd3a65" - -[[package]] -name = "verisim-api" -version = "0.1.0" -dependencies = [ - "async-graphql", - "async-graphql-axum", - "axum", - "axum-server", - "chrono", - "hex", - "hyper", - "prometheus", - "proptest", - "prost", - "prost-types", - "reqwest", - "rustls", - "serde", - "serde_json", - "sha2", - "thiserror 2.0.18", - "tokio", - "tonic", - "tonic-prost", - "tower", - "tracing", - "tracing-subscriber", - "verisim-document", - "verisim-drift", - "verisim-graph", - "verisim-normalizer", - "verisim-octad", - "verisim-planner", - "verisim-provenance", - "verisim-semantic", - "verisim-spatial", - "verisim-temporal", - "verisim-tensor", - "verisim-vector", -] - -[[package]] -name = "verisim-document" -version = "0.1.0" -dependencies = [ - "async-trait", - "proptest", - "serde", - "serde_json", - "tantivy", - "thiserror 2.0.18", - "tokio", - "tracing", -] - -[[package]] -name = "verisim-drift" -version = "0.1.0" -dependencies = [ - "async-trait", - "chrono", - "prometheus", - "proptest", - "serde", - "thiserror 2.0.18", - "tokio", - "tracing", -] - -[[package]] -name = "verisim-graph" -version = "0.1.0" -dependencies = [ - "async-trait", - "oxigraph", - "proptest", - "redb", - "serde", - "serde_json", - "tempfile", - "thiserror 2.0.18", - "tokio", -] - -[[package]] -name = "verisim-nif" -version = "0.1.0" -dependencies = [ - "rustler", - "serde", - "serde_json", - "tokio", - "verisim-document", - "verisim-drift", - "verisim-graph", - "verisim-normalizer", - "verisim-octad", - "verisim-vector", -] - -[[package]] -name = "verisim-normalizer" -version = "0.1.0" -dependencies = [ - "async-trait", - "chrono", - "futures", - "prometheus", - "proptest", - "serde", - "serde_json", - "thiserror 2.0.18", - "tokio", - "tracing", - "uuid", - "verisim-document", - "verisim-drift", - "verisim-graph", - "verisim-octad", - "verisim-semantic", - "verisim-tensor", - "verisim-vector", -] - -[[package]] -name = "verisim-octad" -version = "0.1.0" -dependencies = [ - "async-trait", - "chrono", - "proptest", - "serde", - "serde_json", - "tempfile", - "thiserror 2.0.18", - "tokio", - "tracing", - "uuid", - "verisim-document", - "verisim-graph", - "verisim-provenance", - "verisim-semantic", - "verisim-spatial", - "verisim-temporal", - "verisim-tensor", - "verisim-vector", - "verisim-wal", -] - -[[package]] -name = "verisim-planner" -version = "0.1.0" -dependencies = [ - "chrono", - "proptest", - "serde", - "serde_json", - "sha2", - "thiserror 2.0.18", - "tokio", - "tracing", -] - -[[package]] -name = "verisim-provenance" -version = "0.1.0" -dependencies = [ - "async-trait", - "chrono", - "proptest", - "serde", - "serde_json", - "sha2", - "tempfile", - "thiserror 2.0.18", - "tokio", - "tracing", - "verisim-storage", -] - -[[package]] -name = "verisim-repl" -version = "0.1.0" -dependencies = [ - "clap", - "colored", - "comfy-table", - "dirs", - "reqwest", - "rustls", - "rustyline", - "rustyline-derive", - "serde", - "serde_json", -] - -[[package]] -name = "verisim-semantic" -version = "0.1.0" -dependencies = [ - "async-trait", - "chrono", - "ciborium", - "hex", - "proptest", - "regex", - "serde", - "serde_json", - "sha2", - "tempfile", - "thiserror 2.0.18", - "tokio", - "tracing", - "verisim-storage", -] - -[[package]] -name = "verisim-spatial" -version = "0.1.0" -dependencies = [ - "async-trait", - "proptest", - "serde", - "serde_json", - "tempfile", - "thiserror 2.0.18", - "tokio", - "tracing", - "verisim-storage", -] - -[[package]] -name = "verisim-storage" -version = "0.1.0" -dependencies = [ - "async-trait", - "redb", - "serde", - "serde_json", - "tempfile", - "thiserror 2.0.18", - "tokio", - "tokio-test", - "tracing", -] - -[[package]] -name = "verisim-temporal" -version = "0.1.0" -dependencies = [ - "async-trait", - "chrono", - "proptest", - "serde", - "serde_json", - "tempfile", - "thiserror 2.0.18", - "tokio", - "tracing", - "verisim-storage", -] - -[[package]] -name = "verisim-tensor" -version = "0.1.0" -dependencies = [ - "async-trait", - "ndarray", - "proptest", - "serde", - "serde_json", - "tempfile", - "thiserror 2.0.18", - "tokio", - "tracing", - "verisim-storage", -] - -[[package]] -name = "verisim-vector" -version = "0.1.0" -dependencies = [ - "async-trait", - "criterion", - "ndarray", - "proptest", - "serde", - "serde_json", - "tempfile", - "thiserror 2.0.18", - "tokio", - "tracing", - "verisim-storage", -] - -[[package]] -name = "verisim-wal" -version = "0.1.0" -dependencies = [ - "chrono", - "crc32fast", - "proptest", - "serde", - "serde_json", - "tempfile", - "thiserror 2.0.18", - "tracing", - "uuid", -] - -[[package]] -name = "verisimdb-benchmarks" -version = "0.1.0" -dependencies = [ - "criterion", - "futures", - "tokio", - "uuid", - "verisim-document", - "verisim-drift", - "verisim-graph", - "verisim-normalizer", - "verisim-octad", - "verisim-provenance", - "verisim-semantic", - "verisim-spatial", - "verisim-temporal", - "verisim-tensor", - "verisim-vector", -] - -[[package]] -name = "version_check" -version = "0.9.5" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "0b928f33d975fc6ad9f86c8f283853ad26bdd5b10b7f1542aa2fa15e2289105a" - -[[package]] -name = "wait-timeout" -version = "0.2.1" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "09ac3b126d3914f9849036f826e054cbabdc8519970b8998ddaf3b5bd3c65f11" -dependencies = [ - "libc", -] - -[[package]] -name = "walkdir" -version = "2.5.0" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "29790946404f91d9c5d06f9874efddea1dc06c5efe94541a7d6863108e3a5e4b" -dependencies = [ - "same-file", - "winapi-util", -] - -[[package]] -name = "want" -version = "0.3.1" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "bfa7760aed19e106de2c7c0b581b509f2f25d3dacaf737cb82ac61bc6d760b0e" -dependencies = [ - "try-lock", -] - -[[package]] -name = "wasi" -version = "0.11.1+wasi-snapshot-preview1" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "ccf3ec651a847eb01de73ccad15eb7d99f80485de043efb2f370cd654f4ea44b" - -[[package]] -name = "wasip2" -version = "1.0.3+wasi-0.2.9" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "20064672db26d7cdc89c7798c48a0fdfac8213434a1186e5ef29fd560ae223d6" -dependencies = [ - "wit-bindgen 0.57.1", -] - -[[package]] -name = "wasip3" -version = "0.4.0+wasi-0.3.0-rc-2026-01-06" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "5428f8bf88ea5ddc08faddef2ac4a67e390b88186c703ce6dbd955e1c145aca5" -dependencies = [ - "wit-bindgen 0.51.0", -] - -[[package]] -name = "wasm-bindgen" -version = "0.2.118" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "0bf938a0bacb0469e83c1e148908bd7d5a6010354cf4fb73279b7447422e3a89" -dependencies = [ - "cfg-if", - "once_cell", - "rustversion", - "wasm-bindgen-macro", - "wasm-bindgen-shared", -] - -[[package]] -name = "wasm-bindgen-futures" -version = "0.4.68" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "f371d383f2fb139252e0bfac3b81b265689bf45b6874af544ffa4c975ac1ebf8" -dependencies = [ - "js-sys", - "wasm-bindgen", -] - -[[package]] -name = "wasm-bindgen-macro" -version = "0.2.118" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "eeff24f84126c0ec2db7a449f0c2ec963c6a49efe0698c4242929da037ca28ed" -dependencies = [ - "quote", - "wasm-bindgen-macro-support", -] - -[[package]] -name = "wasm-bindgen-macro-support" -version = "0.2.118" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "9d08065faf983b2b80a79fd87d8254c409281cf7de75fc4b773019824196c904" -dependencies = [ - "bumpalo", - "proc-macro2", - "quote", - "syn", - "wasm-bindgen-shared", -] - -[[package]] -name = "wasm-bindgen-shared" -version = "0.2.118" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "5fd04d9e306f1907bd13c6361b5c6bfc7b3b3c095ed3f8a9246390f8dbdee129" -dependencies = [ - "unicode-ident", -] - -[[package]] -name = "wasm-encoder" -version = "0.244.0" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "990065f2fe63003fe337b932cfb5e3b80e0b4d0f5ff650e6985b1048f62c8319" -dependencies = [ - "leb128fmt", - "wasmparser", -] - -[[package]] -name = "wasm-metadata" -version = "0.244.0" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "bb0e353e6a2fbdc176932bbaab493762eb1255a7900fe0fea1a2f96c296cc909" -dependencies = [ - "anyhow", - "indexmap", - "wasm-encoder", - "wasmparser", -] - -[[package]] -name = "wasmparser" -version = "0.244.0" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "47b807c72e1bac69382b3a6fb3dbe8ea4c0ed87ff5629b8685ae6b9a611028fe" -dependencies = [ - "bitflags", - "hashbrown 0.15.5", - "indexmap", - "semver", -] - -[[package]] -name = "web-sys" -version = "0.3.95" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "4f2dfbb17949fa2088e5d39408c48368947b86f7834484e87b73de55bc14d97d" -dependencies = [ - "js-sys", - "wasm-bindgen", -] - -[[package]] -name = "webpki-root-certs" -version = "1.0.7" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "f31141ce3fc3e300ae89b78c0dd67f9708061d1d2eda54b8209346fd6be9a92c" -dependencies = [ - "rustls-pki-types", -] - -[[package]] -name = "winapi" -version = "0.3.9" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "5c839a674fcd7a98952e593242ea400abe93992746761e38641405d28b00f419" -dependencies = [ - "winapi-i686-pc-windows-gnu", - "winapi-x86_64-pc-windows-gnu", -] - -[[package]] -name = "winapi-i686-pc-windows-gnu" -version = "0.4.0" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "ac3b87c63620426dd9b991e5ce0329eff545bccbbb34f3be09ff6fb6ab51b7b6" - -[[package]] -name = "winapi-util" -version = "0.1.11" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "c2a7b1c03c876122aa43f3020e6c3c3ee5c05081c9a00739faf7503aeba10d22" -dependencies = [ - "windows-sys 0.61.2", -] - -[[package]] -name = "winapi-x86_64-pc-windows-gnu" -version = "0.4.0" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "712e227841d057c1ee1cd2fb22fa7e5a5461ae8e48fa2ca79ec42cfc1931183f" - -[[package]] -name = "windows-core" -version = "0.62.2" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "b8e83a14d34d0623b51dce9581199302a221863196a1dde71a7663a4c2be9deb" -dependencies = [ - "windows-implement", - "windows-interface", - "windows-link", - "windows-result", - "windows-strings", -] - -[[package]] -name = "windows-implement" -version = "0.60.2" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "053e2e040ab57b9dc951b72c264860db7eb3b0200ba345b4e4c3b14f67855ddf" -dependencies = [ - "proc-macro2", - "quote", - "syn", -] - -[[package]] -name = "windows-interface" -version = "0.59.3" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "3f316c4a2570ba26bbec722032c4099d8c8bc095efccdc15688708623367e358" -dependencies = [ - "proc-macro2", - "quote", - "syn", -] - -[[package]] -name = "windows-link" -version = "0.2.1" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "f0805222e57f7521d6a62e36fa9163bc891acd422f971defe97d64e70d0a4fe5" - -[[package]] -name = "windows-result" -version = "0.4.1" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "7781fa89eaf60850ac3d2da7af8e5242a5ea78d1a11c49bf2910bb5a73853eb5" -dependencies = [ - "windows-link", -] - -[[package]] -name = "windows-strings" -version = "0.5.1" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "7837d08f69c77cf6b07689544538e017c1bfcf57e34b4c0ff58e6c2cd3b37091" -dependencies = [ - "windows-link", -] - -[[package]] -name = "windows-sys" -version = "0.45.0" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "75283be5efb2831d37ea142365f009c02ec203cd29a3ebecbc093d52315b66d0" -dependencies = [ - "windows-targets 0.42.2", -] - -[[package]] -name = "windows-sys" -version = "0.52.0" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "282be5f36a8ce781fad8c8ae18fa3f9beff57ec1b52cb3de0789201425d9a33d" -dependencies = [ - "windows-targets 0.52.6", -] - -[[package]] -name = "windows-sys" -version = "0.59.0" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "1e38bc4d79ed67fd075bcc251a1c39b32a1776bbe92e5bef1f0bf1f8c531853b" -dependencies = [ - "windows-targets 0.52.6", -] - -[[package]] -name = "windows-sys" -version = "0.61.2" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "ae137229bcbd6cdf0f7b80a31df61766145077ddf49416a728b02cb3921ff3fc" -dependencies = [ - "windows-link", -] - -[[package]] -name = "windows-targets" -version = "0.42.2" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "8e5180c00cd44c9b1c88adb3693291f1cd93605ded80c250a75d472756b4d071" -dependencies = [ - "windows_aarch64_gnullvm 0.42.2", - "windows_aarch64_msvc 0.42.2", - "windows_i686_gnu 0.42.2", - "windows_i686_msvc 0.42.2", - "windows_x86_64_gnu 0.42.2", - "windows_x86_64_gnullvm 0.42.2", - "windows_x86_64_msvc 0.42.2", -] - -[[package]] -name = "windows-targets" -version = "0.52.6" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "9b724f72796e036ab90c1021d4780d4d3d648aca59e491e6b98e725b84e99973" -dependencies = [ - "windows_aarch64_gnullvm 0.52.6", - "windows_aarch64_msvc 0.52.6", - "windows_i686_gnu 0.52.6", - "windows_i686_gnullvm", - "windows_i686_msvc 0.52.6", - "windows_x86_64_gnu 0.52.6", - "windows_x86_64_gnullvm 0.52.6", - "windows_x86_64_msvc 0.52.6", -] - -[[package]] -name = "windows_aarch64_gnullvm" -version = "0.42.2" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "597a5118570b68bc08d8d59125332c54f1ba9d9adeedeef5b99b02ba2b0698f8" - -[[package]] -name = "windows_aarch64_gnullvm" -version = "0.52.6" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "32a4622180e7a0ec044bb555404c800bc9fd9ec262ec147edd5989ccd0c02cd3" - -[[package]] -name = "windows_aarch64_msvc" -version = "0.42.2" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "e08e8864a60f06ef0d0ff4ba04124db8b0fb3be5776a5cd47641e942e58c4d43" - -[[package]] -name = "windows_aarch64_msvc" -version = "0.52.6" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "09ec2a7bb152e2252b53fa7803150007879548bc709c039df7627cabbd05d469" - -[[package]] -name = "windows_i686_gnu" -version = "0.42.2" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "c61d927d8da41da96a81f029489353e68739737d3beca43145c8afec9a31a84f" - -[[package]] -name = "windows_i686_gnu" -version = "0.52.6" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "8e9b5ad5ab802e97eb8e295ac6720e509ee4c243f69d781394014ebfe8bbfa0b" - -[[package]] -name = "windows_i686_gnullvm" -version = "0.52.6" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "0eee52d38c090b3caa76c563b86c3a4bd71ef1a819287c19d586d7334ae8ed66" - -[[package]] -name = "windows_i686_msvc" -version = "0.42.2" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "44d840b6ec649f480a41c8d80f9c65108b92d89345dd94027bfe06ac444d1060" - -[[package]] -name = "windows_i686_msvc" -version = "0.52.6" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "240948bc05c5e7c6dabba28bf89d89ffce3e303022809e73deaefe4f6ec56c66" - -[[package]] -name = "windows_x86_64_gnu" -version = "0.42.2" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "8de912b8b8feb55c064867cf047dda097f92d51efad5b491dfb98f6bbb70cb36" - -[[package]] -name = "windows_x86_64_gnu" -version = "0.52.6" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "147a5c80aabfbf0c7d901cb5895d1de30ef2907eb21fbbab29ca94c5b08b1a78" - -[[package]] -name = "windows_x86_64_gnullvm" -version = "0.42.2" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "26d41b46a36d453748aedef1486d5c7a85db22e56aff34643984ea85514e94a3" - -[[package]] -name = "windows_x86_64_gnullvm" -version = "0.52.6" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "24d5b23dc417412679681396f2b49f3de8c1473deb516bd34410872eff51ed0d" - -[[package]] -name = "windows_x86_64_msvc" -version = "0.42.2" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "9aec5da331524158c6d1a4ac0ab1541149c0b9505fde06423b02f5ef0106b9f0" - -[[package]] -name = "windows_x86_64_msvc" -version = "0.52.6" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "589f6da84c646204747d1270a2a5661ea66ed1cced2631d546fdfb155959f9ec" - -[[package]] -name = "winnow" -version = "1.0.1" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "09dac053f1cd375980747450bfc7250c264eaae0583872e845c0c7cd578872b5" -dependencies = [ - "memchr", -] - -[[package]] -name = "wit-bindgen" -version = "0.51.0" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "d7249219f66ced02969388cf2bb044a09756a083d0fab1e566056b04d9fbcaa5" -dependencies = [ - "wit-bindgen-rust-macro", -] - -[[package]] -name = "wit-bindgen" -version = "0.57.1" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "1ebf944e87a7c253233ad6766e082e3cd714b5d03812acc24c318f549614536e" - -[[package]] -name = "wit-bindgen-core" -version = "0.51.0" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "ea61de684c3ea68cb082b7a88508a8b27fcc8b797d738bfc99a82facf1d752dc" -dependencies = [ - "anyhow", - "heck", - "wit-parser", -] - -[[package]] -name = "wit-bindgen-rust" -version = "0.51.0" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "b7c566e0f4b284dd6561c786d9cb0142da491f46a9fbed79ea69cdad5db17f21" -dependencies = [ - "anyhow", - "heck", - "indexmap", - "prettyplease", - "syn", - "wasm-metadata", - "wit-bindgen-core", - "wit-component", -] - -[[package]] -name = "wit-bindgen-rust-macro" -version = "0.51.0" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "0c0f9bfd77e6a48eccf51359e3ae77140a7f50b1e2ebfe62422d8afdaffab17a" -dependencies = [ - "anyhow", - "prettyplease", - "proc-macro2", - "quote", - "syn", - "wit-bindgen-core", - "wit-bindgen-rust", -] - -[[package]] -name = "wit-component" -version = "0.244.0" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "9d66ea20e9553b30172b5e831994e35fbde2d165325bec84fc43dbf6f4eb9cb2" -dependencies = [ - "anyhow", - "bitflags", - "indexmap", - "log", - "serde", - "serde_derive", - "serde_json", - "wasm-encoder", - "wasm-metadata", - "wasmparser", - "wit-parser", -] - -[[package]] -name = "wit-parser" -version = "0.244.0" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "ecc8ac4bc1dc3381b7f59c34f00b67e18f910c2c0f50015669dde7def656a736" -dependencies = [ - "anyhow", - "id-arena", - "indexmap", - "log", - "semver", - "serde", - "serde_derive", - "serde_json", - "unicode-xid", - "wasmparser", -] - -[[package]] -name = "writeable" -version = "0.6.3" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "1ffae5123b2d3fc086436f8834ae3ab053a283cfac8fe0a0b8eaae044768a4c4" - -[[package]] -name = "yoke" -version = "0.8.2" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "abe8c5fda708d9ca3df187cae8bfb9ceda00dd96231bed36e445a1a48e66f9ca" -dependencies = [ - "stable_deref_trait", - "yoke-derive", - "zerofrom", -] - -[[package]] -name = "yoke-derive" -version = "0.8.2" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "de844c262c8848816172cef550288e7dc6c7b7814b4ee56b3e1553f275f1858e" -dependencies = [ - "proc-macro2", - "quote", - "syn", - "synstructure", -] - -[[package]] -name = "zerocopy" -version = "0.8.48" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "eed437bf9d6692032087e337407a86f04cd8d6a16a37199ed57949d415bd68e9" -dependencies = [ - "zerocopy-derive", -] - -[[package]] -name = "zerocopy-derive" -version = "0.8.48" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "70e3cd084b1788766f53af483dd21f93881ff30d7320490ec3ef7526d203bad4" -dependencies = [ - "proc-macro2", - "quote", - "syn", -] - -[[package]] -name = "zerofrom" -version = "0.1.7" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "69faa1f2a1ea75661980b013019ed6687ed0e83d069bc1114e2cc74c6c04c4df" -dependencies = [ - "zerofrom-derive", -] - -[[package]] -name = "zerofrom-derive" -version = "0.1.7" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "11532158c46691caf0f2593ea8358fed6bbf68a0315e80aae9bd41fbade684a1" -dependencies = [ - "proc-macro2", - "quote", - "syn", - "synstructure", -] - -[[package]] -name = "zeroize" -version = "1.8.2" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "b97154e67e32c85465826e8bcc1c59429aaaf107c1e4a9e53c8d8ccd5eff88d0" - -[[package]] -name = "zerotrie" -version = "0.2.4" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "0f9152d31db0792fa83f70fb2f83148effb5c1f5b8c7686c3459e361d9bc20bf" -dependencies = [ - "displaydoc", - "yoke", - "zerofrom", -] - -[[package]] -name = "zerovec" -version = "0.11.6" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "90f911cbc359ab6af17377d242225f4d75119aec87ea711a880987b18cd7b239" -dependencies = [ - "yoke", - "zerofrom", - "zerovec-derive", -] - -[[package]] -name = "zerovec-derive" -version = "0.11.3" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "625dc425cab0dca6dc3c3319506e6593dcb08a9f387ea3b284dbd52a92c40555" -dependencies = [ - "proc-macro2", - "quote", - "syn", -] - -[[package]] -name = "zmij" -version = "1.0.21" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "b8848ee67ecc8aedbaf3e4122217aff892639231befc6a1b58d29fff4c2cabaa" diff --git a/verisimdb/Cargo.toml b/verisimdb/Cargo.toml deleted file mode 100644 index 72327d2d..00000000 --- a/verisimdb/Cargo.toml +++ /dev/null @@ -1,126 +0,0 @@ -# SPDX-License-Identifier: MPL-2.0 -# VeriSimDB - The Veridical Simulacrum Database -# An 8-core multimodal database (octad) with self-normalization - -[workspace] -resolver = "2" -members = [ - "rust-core/verisim-graph", - "rust-core/verisim-vector", - "rust-core/verisim-tensor", - "rust-core/verisim-semantic", - "rust-core/verisim-document", - "rust-core/verisim-temporal", - "rust-core/verisim-provenance", - "rust-core/verisim-spatial", - "rust-core/verisim-octad", - "rust-core/verisim-normalizer", - "rust-core/verisim-drift", - "rust-core/verisim-planner", - "rust-core/verisim-api", - "rust-core/verisim-repl", - "rust-core/verisim-wal", - "rust-core/verisim-storage", - "rust-core/verisim-nif", - "benches", -] - -[workspace.package] -version = "0.1.0" -edition = "2021" -authors = ["Jonathan D.A. Jewell "] -license = "MPL-2.0" -repository = "https://gitlab.com/hyperpolymath/verisimdb" -homepage = "https://github.com/hyperpolymath/verisimdb" -documentation = "https://github.com/hyperpolymath/verisimdb/tree/main/docs" -keywords = ["database", "multimodal", "drift-detection", "federation", "consistency"] -categories = ["database", "data-structures"] -readme = "README.adoc" - -[workspace.dependencies] -# Graph modality — SimpleGraphStore is default (pure Rust, no C++ linker needed). -# Enable oxigraph-backend feature on verisim-graph for full RDF/SPARQL support. -oxigraph = "0.4" - -# Vector modality (HNSW) -hnsw_rs = "0.3" -ndarray = "0.16" - -# Semantic modality -serde = { version = "1.0", features = ["derive"] } -serde_json = "1.0" -ciborium = "0.2" # CBOR for proof blobs - -# Document modality (LZ4 compression — pure Rust via lz4_flex, no zstd C library) -tantivy = { version = "0.26", default-features = false, features = ["mmap", "lz4-compression"] } - -# Temporal modality -chrono = { version = "0.4", features = ["serde"] } - -# API and networking -axum = "0.8" -tokio = { version = "1", features = ["full"] } -tower = "0.5" -hyper = "1.0" -reqwest = { version = "0.13", default-features = false, features = ["json", "query", "http2", "rustls-no-provider"] } - -# Serialization -# bincode removed — not used in codebase. Use postcard or ciborium for future serialization needs. -# bincode = "2.0.0-rc.3" -postcard = { version = "1.0", features = ["alloc"] } - -# Error handling -thiserror = "2.0" -anyhow = "1.0" - -# Logging and tracing -tracing = "0.1" -tracing-subscriber = { version = "0.3", features = ["env-filter", "json"] } - -# TLS (pure Rust via ring — no OpenSSL, no aws-lc-sys/cmake) -rustls = { version = "0.23", default-features = false, features = ["ring", "std", "tls12", "logging"] } -axum-server = { version = "0.7", default-features = false, features = ["tls-rustls-no-provider"] } - -# Testing -proptest = "1.4" -criterion = "0.5" - -# Async -futures = "0.3" -async-trait = "0.1" - -# Metrics -prometheus = "0.14" - -# UUID -uuid = { version = "1.11", features = ["v4"] } - -# Regex -regex = "1.11" - -# Persistent storage (pure Rust, B-tree, ACID) -redb = "3.1" - -# CRC (WAL integrity) -crc32fast = "1.4" - -# Cryptography (ZKP proofs in semantic store) -sha2 = "0.10" - -# GraphQL -async-graphql = "7.2" -async-graphql-axum = "7.2" - -# gRPC -tonic = "0.14" -tonic-prost = "0.14" -prost = "0.14" -prost-types = "0.14" - -[profile.release] -lto = true -codegen-units = 1 -panic = "abort" - -# NOTE: lru 0.12.5 had RUSTSEC-2026-0002 (IterMut Stacked Borrows, LOW severity) -# Mitigation: tantivy upgraded to 0.26, which depends on lru 0.16.3. diff --git a/verisimdb/DEPLOYMENT.adoc b/verisimdb/DEPLOYMENT.adoc deleted file mode 100644 index 2bb7a517..00000000 --- a/verisimdb/DEPLOYMENT.adoc +++ /dev/null @@ -1,793 +0,0 @@ -// SPDX-License-Identifier: CC-BY-SA-4.0 -= VeriSimDB Production Deployment Guide -:toc: -:toc-placement!: - -[.lead] -**Complete guide for deploying VeriSimDB in production environments** - -toc::[] - -== Overview - -VeriSimDB can be deployed in three modes: - -1. **Standalone** - Single-node database (like PostgreSQL) -2. **Federated** - Coordinator for distributed stores across institutions -3. **Hybrid** - Some modalities local, others federated - -This guide covers all deployment modes with focus on production readiness, security, monitoring, and operational procedures. - -== Prerequisites - -=== Hardware Requirements - -==== Minimum (Development/Testing) - -[cols="2,3"] -|=== -|Component |Specification - -|CPU |4 cores -|RAM |8 GB -|Storage |20 GB SSD -|Network |100 Mbps -|=== - -==== Recommended (Production - Standalone) - -[cols="2,3"] -|=== -|Component |Specification - -|CPU |16 cores (AMD EPYC or Intel Xeon) -|RAM |64 GB ECC -|Storage |500 GB NVMe SSD (RAID 1) -|Network |10 Gbps -|=== - -==== Recommended (Production - Federated Coordinator) - -[cols="2,3"] -|=== -|Component |Specification - -|CPU |8 cores -|RAM |32 GB ECC -|Storage |100 GB NVMe SSD (RAID 1) -|Network |10 Gbps with low latency -|=== - -=== Software Requirements - -* **OS**: Fedora 39+ or RHEL 9+ (other Linux distributions supported) -* **Container Runtime**: Podman 4.0+ (preferred) or Docker 24.0+ -* **Rust**: 1.75+ (for building from source) -* **Elixir**: 1.16+ with Erlang/OTP 26+ -* **Node.js/Deno**: Deno 1.40+ (for ReScript compilation) - -== Deployment Modes - -=== Mode 1: Standalone Database - -Single-node deployment with all 6 modalities local. - -==== Use Cases - -* Small to medium deployments (<10M octads) -* Organizations with single-site requirements -* Development and testing environments -* Air-gapped environments - -==== Architecture - ----- -┌─────────────────────────────────────────┐ -│ Elixir Orchestration (Port 4000) │ -│ HTTP API + WebSocket │ -├─────────────────────────────────────────┤ -│ Rust Core (Port 8080) │ -│ verisim-api HTTP Server │ -├─────────────────────────────────────────┤ -│ Local Modality Stores │ -│ ├── Graph (Oxigraph) │ -│ ├── Vector (HNSW) │ -│ ├── Tensor (ndarray) │ -│ ├── Semantic (CBOR) │ -│ ├── Document (Tantivy) │ -│ └── Temporal (Version tree) │ -└─────────────────────────────────────────┘ ----- - -==== Deployment Steps - -[source,bash] ----- -# 1. Clone repository -git clone https://github.com/hyperpolymath/verisimdb -cd verisimdb - -# 2. Build Rust core -cargo build --release --all-features - -# 3. Build Elixir orchestration -cd elixir-orchestration -mix deps.get -MIX_ENV=prod mix release - -# 4. Create container image -cd .. -podman build -t verisimdb:latest -f container/Containerfile . - -# 5. Create persistent volumes -podman volume create verisimdb-data -podman volume create verisimdb-logs - -# 6. Run container -podman run -d \ - --name verisimdb \ - -p 8080:8080 \ - -p 4000:4000 \ - -v verisimdb-data:/var/lib/verisimdb:Z \ - -v verisimdb-logs:/var/log/verisimdb:Z \ - -e VERISIM_MODE=standalone \ - -e VERISIM_DATA_DIR=/var/lib/verisimdb \ - --restart=unless-stopped \ - verisimdb:latest ----- - -=== Mode 2: Federated Coordinator - -Lightweight coordinator that maps octad IDs to remote store locations. - -==== Use Cases - -* Multi-institutional collaborations -* Distributed knowledge networks -* Organizations with data sovereignty requirements -* Hybrid cloud deployments - -==== Architecture - ----- -┌─────────────────────────────────────────────────────┐ -│ ReScript Registry (Port 3000) │ -│ UUID → Store Mapping + KRaft Metadata Log │ -├─────────────────────────────────────────────────────┤ -│ Elixir Orchestration (Port 4000) │ -│ Federation Query Router │ -├─────────────────────────────────────────────────────┤ -│ Remote Stores (External) │ -│ ├── University A (Graph + Document) │ -│ ├── Research Lab B (Vector + Tensor) │ -│ └── Company C (Semantic + Temporal) │ -└─────────────────────────────────────────────────────┘ ----- - -==== Deployment Steps - -[source,bash] ----- -# 1. Build ReScript registry -cd src/registry -deno bundle Registry.res registry.js - -# 2. Deploy registry -podman run -d \ - --name verisimdb-registry \ - -p 3000:3000 \ - -v verisimdb-registry-data:/var/lib/verisimdb-registry:Z \ - -e VERISIM_MODE=federation \ - -e VERISIM_REGISTRY_PORT=3000 \ - --restart=unless-stopped \ - verisimdb:latest registry - -# 3. Deploy orchestration layer -podman run -d \ - --name verisimdb-coordinator \ - -p 4000:4000 \ - --link verisimdb-registry \ - -e VERISIM_MODE=federation \ - -e VERISIM_REGISTRY_URL=http://verisimdb-registry:3000 \ - --restart=unless-stopped \ - verisimdb:latest coordinator ----- - -=== Mode 3: Hybrid Deployment - -Some modalities local (fast), others federated (shared). - -==== Use Cases - -* Organizations with high-frequency local queries + occasional federated queries -* Caching frequently accessed remote data -* Gradual migration from standalone to federated - -==== Configuration - -[source,toml] ----- -# config/hybrid.toml -[verisimdb] -mode = "hybrid" - -[local_modalities] -document = true -vector = true -temporal = true - -[federated_modalities] -graph = ["https://partner-a.example.org:8080"] -semantic = ["https://partner-b.example.org:8080"] -tensor = ["https://partner-c.example.org:8080"] - -[cache] -enabled = true -ttl_seconds = 3600 -max_size_mb = 10240 ----- - -== Configuration - -=== Environment Variables - -[cols="2,3,2"] -|=== -|Variable |Description |Default - -|`VERISIM_MODE` |Deployment mode: standalone, federation, hybrid |standalone -|`VERISIM_DATA_DIR` |Data directory path |/var/lib/verisimdb -|`VERISIM_CLICKHOUSE_URL` |ClickHouse HTTP endpoint used by `/api/v1/proof_attempts*` |http://localhost:8123 -|`VERISIM_PROOF_ATTEMPTS_TOKEN` |If set, required token for `POST /api/v1/proof_attempts` via `X-Proof-Attempts-Token` or `Authorization: Bearer` |(unset) -|`VERISIM_LOG_LEVEL` |Log level: debug, info, warn, error |info -|`VERISIM_HTTP_PORT` |Rust API HTTP port |8080 -|`VERISIM_ELIXIR_PORT` |Elixir orchestration port |4000 -|`VERISIM_REGISTRY_PORT` |ReScript registry port |3000 -|`VERISIM_ENABLE_METRICS` |Enable Prometheus metrics |true -|`VERISIM_ENABLE_TRACING` |Enable OpenTelemetry tracing |false -|`VERISIM_MAX_CONNECTIONS` |Max concurrent connections |1000 -|`VERISIM_DRIFT_THRESHOLD` |Drift detection threshold (0.0-1.0) |0.7 -|`VERISIM_AUTO_NORMALIZE` |Enable automatic normalization |true -|=== - -=== Configuration File - -[source,toml] ----- -# config/production.toml -[verisimdb] -mode = "standalone" -data_dir = "/var/lib/verisimdb" -log_level = "info" - -[http] -port = 8080 -max_connections = 1000 -request_timeout_ms = 30000 -keep_alive = true - -[orchestration] -port = 4000 -distributed_erlang = true -cluster_cookie = "verisimdb-production-secret" - -[modalities] -[modalities.document] -enabled = true -index_path = "/var/lib/verisimdb/document" -commit_interval_ms = 5000 - -[modalities.vector] -enabled = true -dimension = 384 -distance_metric = "cosine" -hnsw_m = 16 -hnsw_ef_construction = 200 - -[modalities.graph] -enabled = true -storage_path = "/var/lib/verisimdb/graph" - -[drift] -enabled = true -check_interval_ms = 60000 -thresholds = { semantic_vector = 0.7, graph_document = 0.8 } - -[normalization] -enabled = true -max_concurrent = 10 -strategy = "hybrid_push_pull" - -[metrics] -enabled = true -prometheus_port = 9090 - -[tracing] -enabled = false -otlp_endpoint = "http://localhost:4317" ----- - -== Security - -=== Network Security - -==== Firewall Configuration - -[source,bash] ----- -# Allow only necessary ports -firewall-cmd --permanent --add-port=8080/tcp # Rust API -firewall-cmd --permanent --add-port=4000/tcp # Elixir orchestration -firewall-cmd --permanent --add-port=9090/tcp # Prometheus metrics -firewall-cmd --reload - -# Restrict to specific IPs (recommended) -firewall-cmd --permanent --add-rich-rule='rule family="ipv4" source address="10.0.0.0/8" port port="8080" protocol="tcp" accept' ----- - -==== TLS/HTTPS - -[source,bash] ----- -# Generate self-signed certificate (development) -openssl req -x509 -newkey rsa:4096 -nodes \ - -keyout /etc/verisimdb/tls/key.pem \ - -out /etc/verisimdb/tls/cert.pem \ - -days 365 \ - -subj "/CN=verisimdb.example.org" - -# Production: Use Let's Encrypt -certbot certonly --standalone \ - -d verisimdb.example.org \ - --deploy-hook "systemctl reload verisimdb" - -# Configure TLS in config -[http.tls] -enabled = true -cert_path = "/etc/letsencrypt/live/verisimdb.example.org/fullchain.pem" -key_path = "/etc/letsencrypt/live/verisimdb.example.org/privkey.pem" ----- - -=== Authentication & Authorization - -==== API Keys - -[source,bash] ----- -# Generate API key -openssl rand -hex 32 > /etc/verisimdb/api-keys/admin.key - -# Configure in environment -export VERISIM_API_KEY=$(cat /etc/verisimdb/api-keys/admin.key) - -# Use in requests -curl -H "Authorization: Bearer $VERISIM_API_KEY" \ - https://verisimdb.example.org:8080/api/v1/health ----- - -==== Role-Based Access Control (RBAC) - -[source,toml] ----- -# config/rbac.toml -[roles.reader] -permissions = ["octad:read", "search:execute"] - -[roles.writer] -permissions = ["octad:read", "octad:write", "search:execute"] - -[roles.admin] -permissions = ["*"] - -[users] -[users."alice@example.org"] -role = "admin" -api_key_hash = "sha256:..." - -[users."bob@example.org"] -role = "writer" -api_key_hash = "sha256:..." ----- - -=== Data Encryption - -==== Encryption at Rest - -[source,bash] ----- -# Use LUKS for volume encryption -cryptsetup luksFormat /dev/sdb -cryptsetup luksOpen /dev/sdb verisimdb-data -mkfs.ext4 /dev/mapper/verisimdb-data -mount /dev/mapper/verisimdb-data /var/lib/verisimdb - -# Auto-mount on boot -echo "verisimdb-data UUID=$(blkid -s UUID -o value /dev/sdb) none luks" >> /etc/crypttab -echo "/dev/mapper/verisimdb-data /var/lib/verisimdb ext4 defaults 0 2" >> /etc/fstab ----- - -==== Encryption in Transit - -All network communication uses TLS 1.3: -* API endpoints (HTTPS) -* Elixir distributed Erlang (TLS) -* Federated store communication (HTTPS) - -== Monitoring - -=== Prometheus Metrics - -VeriSimDB exposes metrics on port 9090 (configurable): - -[source,bash] ----- -# Scrape configuration -# prometheus.yml -scrape_configs: - - job_name: 'verisimdb' - static_configs: - - targets: ['localhost:9090'] - metrics_path: '/metrics' - scrape_interval: 15s ----- - -==== Key Metrics - -[cols="2,3,2"] -|=== -|Metric |Description |Type - -|`verisim_octads_total` |Total octads in database |Counter -|`verisim_queries_total` |Total queries executed |Counter -|`verisim_query_duration_seconds` |Query latency histogram |Histogram -|`verisim_drift_score` |Current drift score |Gauge -|`verisim_normalizations_total` |Total normalizations performed |Counter -|`verisim_store_health` |Store health status (0-1) |Gauge -|`verisim_memory_usage_bytes` |Memory usage |Gauge -|`verisim_disk_usage_bytes` |Disk usage per modality |Gauge -|=== - -=== Logging - -==== Log Levels - -* `DEBUG` - Detailed trace for development -* `INFO` - Normal operational messages -* `WARN` - Warning conditions -* `ERROR` - Error conditions requiring attention - -==== Log Format - -[source,json] ----- -{ - "timestamp": "2026-02-04T20:00:00Z", - "level": "INFO", - "component": "verisim-api", - "message": "Octad created", - "octad_id": "550e8400-e29b-41d4-a716-446655440000", - "modalities": ["document", "vector"], - "duration_ms": 45 -} ----- - -==== Log Rotation - -[source,bash] ----- -# /etc/logrotate.d/verisimdb -/var/log/verisimdb/*.log { - daily - rotate 30 - compress - delaycompress - notifempty - create 0640 verisimdb verisimdb - sharedscripts - postrotate - podman kill -s HUP verisimdb - endscript -} ----- - -=== Alerting - -==== Prometheus Alerting Rules - -[source,yaml] ----- -# alerts.yml -groups: - - name: verisimdb - interval: 30s - rules: - - alert: HighDriftScore - expr: verisim_drift_score > 0.8 - for: 5m - labels: - severity: warning - annotations: - summary: "High drift detected" - description: "Drift score {{ $value }} exceeds threshold" - - - alert: SlowQueries - expr: histogram_quantile(0.95, verisim_query_duration_seconds) > 1.0 - for: 5m - labels: - severity: warning - annotations: - summary: "Slow queries detected" - description: "95th percentile query latency is {{ $value }}s" - - - alert: StoreUnhealthy - expr: verisim_store_health < 0.5 - for: 2m - labels: - severity: critical - annotations: - summary: "Store unhealthy" - description: "Store {{ $labels.store }} health is {{ $value }}" ----- - -== Backup & Recovery - -=== Backup Strategy - -==== Full Backup - -[source,bash] ----- -#!/bin/bash -# backup-verisimdb.sh - -BACKUP_DIR="/backups/verisimdb" -TIMESTAMP=$(date +%Y%m%d_%H%M%S) -BACKUP_PATH="$BACKUP_DIR/verisimdb_$TIMESTAMP" - -# Stop writes (optional, for consistent backup) -curl -X POST http://localhost:8080/api/v1/admin/read-only - -# Backup data directory -tar -czf "$BACKUP_PATH.tar.gz" /var/lib/verisimdb - -# Backup configuration -tar -czf "$BACKUP_DIR/config_$TIMESTAMP.tar.gz" /etc/verisimdb - -# Resume writes -curl -X POST http://localhost:8080/api/v1/admin/read-write - -# Upload to remote storage -rclone copy "$BACKUP_PATH.tar.gz" remote:verisimdb-backups/ - -# Cleanup old backups (keep last 30 days) -find $BACKUP_DIR -name "verisimdb_*.tar.gz" -mtime +30 -delete - -echo "Backup complete: $BACKUP_PATH.tar.gz" ----- - -==== Incremental Backup - -[source,bash] ----- -# Use rsync for incremental backups -rsync -avz --delete \ - /var/lib/verisimdb/ \ - backup-server:/backups/verisimdb/current/ ----- - -=== Recovery - -==== Restore from Backup - -[source,bash] ----- -# Stop VeriSimDB -podman stop verisimdb - -# Restore data -tar -xzf /backups/verisimdb_20260204_120000.tar.gz -C / - -# Restore configuration -tar -xzf /backups/config_20260204_120000.tar.gz -C / - -# Start VeriSimDB -podman start verisimdb - -# Verify integrity -curl http://localhost:8080/api/v1/health ----- - -==== Point-in-Time Recovery - -VeriSimDB's temporal modality supports point-in-time recovery: - -[source,bash] ----- -# Restore octad to specific timestamp -curl -X POST http://localhost:8080/api/v1/admin/restore \ - -H "Content-Type: application/json" \ - -d '{ - "octad_id": "550e8400-e29b-41d4-a716-446655440000", - "timestamp": "2026-02-04T12:00:00Z" - }' ----- - -== Performance Tuning - -=== Benchmarking - -[source,bash] ----- -# Run benchmarks -cargo bench --bench modality_benchmarks - -# Results location -open target/criterion/report/index.html ----- - -=== Optimization Tips - -==== Vector Store - -* Use smaller dimensions (128-384) for faster similarity search -* Tune HNSW parameters: `M=16`, `ef_construction=200` -* Consider quantization for large datasets - -==== Document Store - -* Increase Tantivy commit interval for write-heavy workloads -* Use smaller index segments for read-heavy workloads -* Enable compression for large document bodies - -==== Graph Store - -* Use SPARQL query optimization -* Index frequently queried predicates -* Partition large graphs by domain - -==== Drift Detection - -* Adjust thresholds based on workload -* Disable for write-heavy applications -* Use async normalization - -== Troubleshooting - -=== Common Issues - -==== High Memory Usage - -[source,bash] ----- -# Check memory stats -podman stats verisimdb - -# Reduce vector dimension -# config.toml -[modalities.vector] -dimension = 128 # Instead of 384 - -# Enable disk-based caching -[cache] -strategy = "disk" -max_memory_mb = 1024 ----- - -==== Slow Queries - -[source,bash] ----- -# Enable query profiling -export VERISIM_LOG_LEVEL=debug -export VERISIM_PROFILE_QUERIES=true - -# Check slow query log -grep "slow_query" /var/log/verisimdb/api.log - -# Use EXPLAIN for query plans -curl -X POST http://localhost:8080/api/v1/query/explain \ - -d '{"query": "SELECT * FROM..."}' ----- - -==== Drift Normalization Failures - -[source,bash] ----- -# Check normalization status -curl http://localhost:8080/api/v1/normalizer/status - -# Manual trigger -curl -X POST http://localhost:8080/api/v1/normalizer/trigger/$HEXAD_ID - -# Check drift scores -curl http://localhost:8080/api/v1/drift/entity/$HEXAD_ID ----- - -== Operational Procedures - -=== Health Checks - -[source,bash] ----- -# Basic health -curl http://localhost:8080/api/v1/health - -# Detailed status -curl http://localhost:8080/api/v1/status ----- - -=== Upgrades - -[source,bash] ----- -# 1. Backup before upgrade -./backup-verisimdb.sh - -# 2. Pull new image -podman pull verisimdb:v0.2.0 - -# 3. Stop current container -podman stop verisimdb - -# 4. Run new version -podman run -d \ - --name verisimdb-new \ - -p 8080:8080 \ - -v verisimdb-data:/var/lib/verisimdb:Z \ - verisimdb:v0.2.0 - -# 5. Verify new version -curl http://localhost:8080/api/v1/health - -# 6. Remove old container -podman rm verisimdb -podman rename verisimdb-new verisimdb ----- - -=== Scaling - -==== Vertical Scaling - -* Increase CPU cores for parallel query processing -* Increase RAM for larger in-memory indexes -* Use faster NVMe storage for modality stores - -==== Horizontal Scaling (Federation) - -* Deploy multiple standalone instances -* Use ReScript registry to coordinate -* Distribute octads across instances by hash - -== Production Checklist - -=== Before Go-Live - -- [ ] Hardware meets minimum requirements -- [ ] TLS certificates configured -- [ ] API authentication enabled -- [ ] Firewall rules configured -- [ ] Monitoring and alerting configured -- [ ] Backup automation configured -- [ ] Recovery procedures tested -- [ ] Load testing completed -- [ ] Security audit completed -- [ ] Documentation updated - -=== Post-Deployment - -- [ ] Monitor metrics for 24 hours -- [ ] Verify backup completion -- [ ] Test recovery procedures -- [ ] Document any issues -- [ ] Schedule regular maintenance -- [ ] Plan capacity expansion -- [ ] Review security logs - -== Support - -For production support: - -* **Documentation**: https://verisimdb.hyperpolymath.org/docs -* **Issues**: https://github.com/hyperpolymath/verisimdb/issues -* **Discussions**: https://github.com/hyperpolymath/verisimdb/discussions -* **Security**: security@hyperpolymath.org diff --git a/verisimdb/EXPLAINME.adoc b/verisimdb/EXPLAINME.adoc deleted file mode 100644 index b484d96f..00000000 --- a/verisimdb/EXPLAINME.adoc +++ /dev/null @@ -1,199 +0,0 @@ -// SPDX-License-Identifier: CC-BY-SA-4.0 - -= VeriSimDB — Show Me The Receipts -:toc: preamble -:icons: font - -The README makes claims. This file backs them up with code paths, honest -caveats, and enough structural detail for an external reviewer to know where -to look when something goes wrong. - -== Claim 1: "100% drift detection and repair rate across 1000 entities" - -[quote, README.adoc §Drift Detection Demo] -____ -Detection rate: 100.0% -Repair rate: 100.0% -Consistency rate: 100.0% -____ - -=== How it works - -`demos/drift-detection/run_demo.exs` creates 1,000 octad entities, corrupts 50 -of them (semantic/vector/graph divergence), then runs the full -`VeriSim.DriftMonitor` → `StorageRegenerator` pipeline. The demo exits cleanly -only when `VeriSim.EntityServer.get/1` returns consistent state for every -entity. - -The detection path: `rust-core/verisim-drift/` measures divergence between -modality pairs using cosine similarity (Vector ↔ Semantic) and Jaccard index -(Document ↔ Semantic). Scores above configurable thresholds trigger a -`DriftEvent`, which `VeriSim.DriftMonitor` routes to the normalizer. - -The repair path: `rust-core/verisim-normalizer/src/storage_regenerator.rs` -(`StorageRegenerator`) identifies the most authoritative modality and -regenerates drifted modalities from it — six cross-regeneration pairs (Document -→ Vector, Document → Semantic, Semantic → Graph, etc.). Regeneration uses -FNV-1a trigram hashing to 384-dim embeddings for Document → Vector, and -keyword extraction for Document → Semantic. - -=== Honest caveat - -The demo runs in-memory on a single node — both detection and repair are -synchronous in this configuration. Production deployments with the Elixir OTP -supervision tree and persistent storage (redb + Tantivy WAL) have higher -latency. The 100% rates reflect the in-memory ephemeral case; network partitions -or storage engine failures are handled by OTP supervision restart, not by the -normalizer. - -=== Code path - -| File | What it does | -|------|-------------| -| `rust-core/verisim-drift/src/lib.rs` | Drift score computation across modality pairs | -| `rust-core/verisim-normalizer/src/storage_regenerator.rs` | Cross-modality regeneration (68 tests) | -| `elixir-orchestration/lib/verisim/drift_monitor.ex` | OTP GenServer: drift event coordinator | -| `elixir-orchestration/lib/verisim/entity_server.ex` | Per-entity GenServer under DynamicSupervisor | -| `demos/drift-detection/run_demo.exs` | Runnable demo: 1000 entities, 50 corrupted | - -== Claim 2: "Eight modalities per entity — Graph, Vector, Tensor, Semantic, Document, Temporal, Provenance, Spatial" - -[quote, README.adoc §The Octad] -____ -Each entity in VeriSimDB can have representations across eight modalities. -Drift detection operates across all of them. -____ - -=== How it works - -`rust-core/verisim-octad/` defines the unified `OctadStore` type. Each octad -entity carries one optional value per modality. The eight Rust storage backends -are independent crates: - -| Modality | Crate | Storage engine | -|----------|-------|----------------| -| Graph | `verisim-graph` | Pure Rust `SimpleGraphStore` (RDF triples + property graph) | -| Vector | `verisim-vector` | HNSW in-memory similarity index | -| Tensor | `verisim-tensor` | ndarray / Burn multi-dimensional arrays | -| Semantic | `verisim-semantic` | CBOR proof blobs (ciborium) | -| Document | `verisim-document` | Tantivy full-text (LZ4 compressed) | -| Temporal | `verisim-temporal` | chrono version history + time-series | -| Provenance | `verisim-provenance` | SHA-256 hash-chain origin tracking | -| Spatial | `verisim-spatial` | R-tree geospatial index (radius/bounds/nearest) | - -`OctadBuilder` in `rust-core/verisim-octad/` provides a fluent API to populate -any subset of modalities before creating an entity. Unset modalities are `None` -— no phantom drift signals. - -=== Honest caveat - -Tensor modality is the most research-stage of the eight. The storage engine -compiles and tests pass, but real-world use cases (beyond numeric arrays) are -still being refined. The README acknowledges "novel applications, details -forthcoming." Do not rely on the Tensor modality for production data without -reviewing the open issues in `docs/`. - -== Claim 3: "VCL query language with dependent types and proof certificates (VCL-UT)" - -[quote, README.adoc §How It Compares] -____ -Query language: VCL (with dependent types) -Formal verification: VCL-UT (proof certificates) -____ - -=== How it works - -VCL (VeriSim Consonance Language) is the native query interface — NOT SQL. -The built-in Elixir VCL parser lives in `elixir-orchestration/lib/verisim/vcl/` -and translates VCL ASTs against flat-file stores via `FileExecutor`. The parser -handles `FETCH`, `FILTER`, `GROUP`, `FEDERATION`, and `PROOF` clauses. - -VCL-UT extends VCL with typed proof certificates. Eleven proof types are -supported: `EXISTENCE`, `INTEGRITY`, `CONSISTENCY`, `PROVENANCE`, `FRESHNESS`, -`ACCESS`, `CITATION`, `CUSTOM`, `ZKP`, `PROVEN`, `SANCTIFY`. Multi-proof -queries (`PROOF A(x) AND B(y)`) parse and split correctly. - -The Idris2 ABI layer (`src/abi/`) provides formal dependent-type specifications -for VCL query shapes; proof certificates are verified against the ABI before -storage. The `proven` library integration bridges to external certificate-based -JSON/CBOR verification. - -=== Honest caveat - -VCL federation (cross-store queries against remote VeriSimDB instances) is -currently local-only. The `FileExecutor` handles `FEDERATION` clauses against -local flat files; the multi-store remote executor is planned but not -implemented. Do not use VCL for cross-instance federation in production. - -The ReScript VCL Playground in `panll/` connects to the real backend API with -a demo-mode fallback — the fallback is active when `verisim-api` is not running -locally. - -== Dogfooded Across The Account - -[cols="1,2,2", options="header"] -|=== -| Technology / Pattern | Role in VeriSimDB | Also Used In - -| *Rust workspace (12 crates)* -| `rust-core/` — each modality is a separate crate (`verisim-graph`, `-vector`, - `-tensor`, `-semantic`, `-document`, `-temporal`, `-provenance`, `-spatial`, - `-octad`, `-drift`, `-normalizer`, `-api`); Cargo workspace coordinates them -| ephapax (17-crate Rust workspace), panic-attack (Rust analysis engine), - hypatia (`src/rust/` adapters + CLI), a2ml-rs, k9-rs - -| *Elixir/OTP orchestration* -| `elixir-orchestration/` — one `EntityServer` GenServer per entity under - `DynamicSupervisor`; `DriftMonitor`, `QueryRouter`, `SchemaRegistry` as OTP - GenServers; fault isolation: one crashed entity does not affect others -| hypatia (8 OTP GenServers), burble (room-per-GenServer, DynamicSupervisor), - gitbot-fleet (bot GenServers) - -| *Idris2 ABI + Zig FFI standard* -| `src/abi/` — Idris2 ABI definitions for VCL proof types and octad entity - contracts; Zig FFI bridge for C-ABI runtime interop -| gossamer (`src/interface/abi/`), burble (`src/Burble/ABI/`), ephapax (`idris2/`), - typed-wasm, tangle — universal ABI/FFI pattern across the estate - -| *Stapeln container ecosystem* -| `container/compose.toml` (selur-compose stack), `container/.gatekeeper.yaml` - (svalinn edge gateway), `container/manifest.toml` (.ctp bundle manifest); - Chainguard base images throughout -| burble (`containers/compose.toml`), hypatia (Containerfile), boj-server, - all containerised repos via stapeln policy - -| *VCL (VeriSim Consonance Language)* -| Native query interface; built-in Elixir parser + FileExecutor; VCL-UT proof - certificates; ReScript Playground panel via PanLL -| hypatia (`lib/vcl/` — VCL client GenServer, file executor, cross-repo analytics); - nextgen-databases monorepo siblings (QuandleDB, LithoGlyph) -|=== - -== File Map - -[cols="1,3", options="header"] -|=== -| Path | What's There - -| `rust-core/verisim-octad/` | Unified `OctadStore` + `OctadBuilder` — the entity model -| `rust-core/verisim-drift/` | Drift score computation (cosine similarity, Jaccard index) -| `rust-core/verisim-normalizer/src/storage_regenerator.rs` | Cross-modality regeneration (68 tests, 6 source→target pairs) -| `rust-core/verisim-api/` | HTTP API — REST endpoint (8080/8090/8091/8092 per project) -| `elixir-orchestration/lib/verisim/entity_server.ex` | Per-entity GenServer -| `elixir-orchestration/lib/verisim/drift_monitor.ex` | Drift event coordinator -| `elixir-orchestration/lib/verisim/vcl/` | VCL parser, FileExecutor, query functions -| `elixir-orchestration/lib/verisim/hypatia/` | Hypatia integration: ScanIngester, PatternQuery, DispatchBridge (37 tests) -| `connectors/clients/` | 5 client SDKs (Rust, Elixir, ReScript, Julia, Gleam) -| `connectors/test-infra/` | selur-compose: 7 databases for federation adapter integration tests -| `container/Containerfile` | Chainguard-based container (in-memory or persistent build arg) -| `container/compose.toml` | Full stack: rust-core + elixir + svalinn gateway -| `demos/drift-detection/run_demo.exs` | Runnable demo: 1000 entities, 50 corrupted, 100% detect+repair -| `src/abi/` | Idris2 ABI definitions — VCL types, proof certificates, octad contracts -|=== - -== Questions? - -Open an issue or reach out at j.d.a.jewell@open.ac.uk. For instance -deployment, each consuming project should run its own VeriSimDB container on a -unique port — do not share the dev instance across projects. See -`CLAUDE.md §CRITICAL: Instance Policy` for the port assignment table. diff --git a/verisimdb/IMPLEMENTATION-ROADMAP.adoc b/verisimdb/IMPLEMENTATION-ROADMAP.adoc deleted file mode 100644 index 00facc57..00000000 --- a/verisimdb/IMPLEMENTATION-ROADMAP.adoc +++ /dev/null @@ -1,509 +0,0 @@ -= VeriSimDB Implementation Roadmap -:toc: -:toc-placement!: - -// SPDX-License-Identifier: CC-BY-SA-4.0 - -[.lead] -**Bringing VeriSimDB from 10% to 70% completion** - -toc::[] - -== Current Status - -* **Completion**: 10% -* **Phase**: Implementation ramp-up -* **What exists**: Architecture docs, VCL grammar, Rust crate scaffolds, Elixir stubs -* **What's missing**: Actual modality store implementations, VCL execution engine - -== Goal - -**Week 16 target**: 70% completion -- All 6 modality stores operational -- VCL execution engine working -- Testing framework complete -- Basic federation capability - ---- - -== Milestone V1: Rust Modality Stores (Weeks 1-6) - -=== Priority Order - -Implement in this sequence: - -1. **verisim-document** (Week 1-2) - Foundation, most familiar -2. **verisim-temporal** (Week 2-3) - Version history, time-travel -3. **verisim-graph** (Week 3-4) - RDF + property graph -4. **verisim-vector** (Week 4-5) - HNSW embeddings -5. **verisim-semantic** (Week 5-6) - Type annotations, CBOR -6. **verisim-tensor** (Week 5-6) - ndarray storage - -=== Week 1-2: verisim-document - -**Goal**: Full-text search with Tantivy - -```rust -// verisim-document/src/lib.rs -use tantivy::schema::*; -use tantivy::{Index, IndexWriter}; -use uuid::Uuid; - -pub struct DocumentStore { - index: Index, - schema: Schema, -} - -impl DocumentStore { - pub fn create_document(&mut self, octad_id: Uuid, title: &str, body: &str) -> Result<()>; - pub fn search(&self, query: &str, limit: usize) -> Result>; - pub fn get(&self, octad_id: Uuid) -> Result>; - pub fn update(&mut self, octad_id: Uuid, title: &str, body: &str) -> Result<()>; - pub fn delete(&mut self, octad_id: Uuid) -> Result<()>; -} -``` - -**Tests**: -- Unit tests for CRUD operations -- Property-based tests (proptest) -- Search relevance tests - -**Deliverable**: `cargo test` passes, search works - ---- - -=== Week 2-3: verisim-temporal - -**Goal**: Version history and time-travel queries - -```rust -// verisim-temporal/src/lib.rs -use chrono::{DateTime, Utc}; -use uuid::Uuid; - -pub struct TemporalStore { - // Version history storage -} - -impl TemporalStore { - pub fn record_version(&mut self, octad_id: Uuid, timestamp: DateTime, data: &[u8]) -> Result<()>; - pub fn get_at_time(&self, octad_id: Uuid, timestamp: DateTime) -> Result>>; - pub fn get_history(&self, octad_id: Uuid, start: DateTime, end: DateTime) -> Result>; - pub fn diff(&self, octad_id: Uuid, t1: DateTime, t2: DateTime) -> Result; -} -``` - -**Tests**: -- Version recording and retrieval -- Time-travel accuracy -- Diff correctness - -**Deliverable**: Time-travel queries work - ---- - -=== Week 3-4: verisim-graph - -**Goal**: RDF + property graph with Oxigraph - -```rust -// verisim-graph/src/lib.rs -use oxigraph::store::Store; -use uuid::Uuid; - -pub struct GraphStore { - store: Store, -} - -impl GraphStore { - pub fn add_triple(&mut self, subject: &str, predicate: &str, object: &str) -> Result<()>; - pub fn add_edge(&mut self, from: Uuid, to: Uuid, edge_type: &str, props: &HashMap) -> Result<()>; - pub fn traverse(&self, from: Uuid, edge_type: &str, depth: usize) -> Result>; - pub fn sparql_query(&self, query: &str) -> Result; -} -``` - -**Tests**: -- Triple insertion and retrieval -- Graph traversal -- SPARQL subset queries - -**Deliverable**: Graph queries work - ---- - -=== Week 4-5: verisim-vector - -**Goal**: HNSW similarity search - -```rust -// verisim-vector/src/lib.rs -use uuid::Uuid; - -pub struct VectorStore { - // HNSW index -} - -impl VectorStore { - pub fn insert(&mut self, octad_id: Uuid, embedding: &[f32]) -> Result<()>; - pub fn search_similar(&self, query: &[f32], k: usize, threshold: f32) -> Result>; - pub fn cosine_similarity(&self, v1: &[f32], v2: &[f32]) -> f32; - pub fn euclidean_distance(&self, v1: &[f32], v2: &[f32]) -> f32; -} -``` - -**Tests**: -- Embedding insertion -- Similarity search correctness -- Distance metric validation - -**Deliverable**: Vector similarity search works - ---- - -=== Week 5-6: verisim-semantic + verisim-tensor - -**verisim-semantic**: Type annotations + CBOR proof blobs - -```rust -// verisim-semantic/src/lib.rs -pub struct SemanticStore { - // Type annotations -} - -impl SemanticStore { - pub fn add_type(&mut self, octad_id: Uuid, type_uri: &str) -> Result<()>; - pub fn get_types(&self, octad_id: Uuid) -> Result>; - pub fn store_proof_blob(&mut self, octad_id: Uuid, contract: &str, blob: &[u8]) -> Result<()>; - pub fn verify_proof(&self, octad_id: Uuid, contract: &str) -> Result; -} -``` - -**verisim-tensor**: Multi-dimensional arrays - -```rust -// verisim-tensor/src/lib.rs -use ndarray::ArrayD; - -pub struct TensorStore { - // Tensor storage -} - -impl TensorStore { - pub fn store(&mut self, octad_id: Uuid, tensor: ArrayD) -> Result<()>; - pub fn get(&self, octad_id: Uuid) -> Result>>; - pub fn slice(&self, octad_id: Uuid, indices: &[usize]) -> Result>; -} -``` - -**Deliverable**: Both stores operational - ---- - -== Milestone V2: Octad Entity Layer (Weeks 7-8) - -**Goal**: Unified entity abstraction across modalities - -```rust -// verisim-octad/src/lib.rs -use uuid::Uuid; - -pub struct Octad { - pub id: Uuid, - pub document: Option, - pub graph: Option, - pub vector: Option, - pub temporal: Option, - pub semantic: Option, - pub tensor: Option, -} - -pub struct OctadStore { - document_store: DocumentStore, - graph_store: GraphStore, - vector_store: VectorStore, - temporal_store: TemporalStore, - semantic_store: SemanticStore, - tensor_store: TensorStore, -} - -impl OctadStore { - pub fn create(&mut self, octad: Octad) -> Result; - pub fn get(&self, id: Uuid) -> Result>; - pub fn update(&mut self, id: Uuid, octad: Octad) -> Result<()>; - pub fn query_cross_modal(&self, predicate: CrossModalPredicate) -> Result>; -} -``` - -**Drift Detection**: - -```rust -// verisim-drift/src/lib.rs -pub struct DriftDetector { - // Drift detection logic -} - -impl DriftDetector { - pub fn detect(&self, octad_id: Uuid, modality1: Modality, modality2: Modality) -> Result; - pub fn repair(&mut self, octad_id: Uuid, strategy: RepairStrategy) -> Result<()>; -} -``` - -**Deliverable**: Cross-modal queries work, drift detection operational - ---- - -== Milestone V3: VCL Execution Engine (Weeks 9-11) - -**Goal**: Execute VCL queries across all modalities - -=== Week 9: Parser Integration - -```rust -// verisim-api/src/vcl_bridge.rs -use rescript_parser::VCLParser; // ReScript → Rust FFI - -pub fn parse_vcl(query: &str) -> Result; -pub fn execute_vcl(ast: VCLAst, octad_store: &OctadStore) -> Result; -``` - -**Integration**: -- ReScript VCL parser → Rust execution engine -- AST serialization (CBOR or JSON) - ---- - -=== Week 10: Query Planner - -```rust -// verisim-api/src/query_planner.rs -pub struct QueryPlan { - pub steps: Vec, - pub estimated_cost: f64, -} - -pub fn plan_query(ast: VCLAst) -> Result; -``` - -**Optimization**: -- Push predicates to modality stores -- Minimize cross-modal joins -- Cache query plans - ---- - -=== Week 11: EXPLAIN Functionality - -```rust -// verisim-api/src/explain.rs -pub struct ExplainOutput { - pub plan: QueryPlan, - pub modalities_queried: Vec, - pub estimated_time: f64, - pub hints: Vec, -} - -pub fn explain(query: &str) -> Result; -``` - -**Deliverable**: VCL queries execute, EXPLAIN works - ---- - -== Milestone V4: Testing & Stability (Weeks 12-14) - -=== Week 12: Property-Based Tests - -```toml -# Cargo.toml -[dev-dependencies] -proptest = "1.4" -``` - -```rust -// tests/property_tests.rs -use proptest::prelude::*; - -proptest! { - #[test] - fn insert_then_retrieve(title in "\\w{1,50}", body in "\\w{10,200}") { - let mut store = OctadStore::new()?; - let id = store.create(octad_with_document(title, body))?; - let retrieved = store.get(id)?; - assert!(retrieved.is_some()); - } -} -``` - ---- - -=== Week 13: Fuzz Testing - -```yaml -# .clusterfuzzlite/project.yaml -language: rust -``` - -```rust -// fuzz/fuzz_targets/fuzz_vcl.rs -#![no_main] -use libfuzzer_sys::fuzz_target; -use verisim_api::parse_vcl; - -fuzz_target!(|data: &[u8]| { - if let Ok(s) = std::str::from_utf8(data) { - let _ = parse_vcl(s); - } -}); -``` - ---- - -=== Week 14: Integration Tests - -```rust -// tests/integration_test.rs -#[test] -fn test_cross_modal_query() { - let store = setup_test_store(); - - // Insert document - let id = store.create_document("Test", "Body")?; - - // Add embedding - store.add_vector(id, &[0.1, 0.2, 0.3])?; - - // Cross-modal query - let query = "SELECT * FROM octads WHERE DOCUMENT MATCHES 'Test' AND VECTOR SIMILAR TO [0.1, 0.2, 0.3]"; - let results = store.query(query)?; - - assert_eq!(results.len(), 1); -} -``` - -**Deliverable**: Test coverage >80% - ---- - -== Milestone V5: Performance & Documentation (Weeks 15-16) - -=== Week 15: Performance - -```rust -// verisim-api/src/cache.rs -use lru::LruCache; - -pub struct QueryCache { - plan_cache: LruCache, - result_cache: LruCache, -} -``` - -**Optimizations**: -- Query plan caching -- Connection pooling -- Batch operations - ---- - -=== Week 16: Documentation - -**Create**: -- `QUICKSTART.adoc` - Get running in 5 minutes -- `API-REFERENCE.adoc` - All Rust APIs documented -- `VCL-TUTORIAL.adoc` - 10 example queries -- `DEPLOYMENT-GUIDE.adoc` - Standalone mode setup - -**Deliverable**: VeriSimDB at 70% completion - ---- - -== Weekly Progress Tracking - -**Every Friday**: - -```scheme -;; Update STATE.scm -(snapshot (date "2026-02-XX") (session "week-N") - (accomplishments - "Completed verisim-document CRUD" - "Property-based tests passing") - (blockers - "HNSW performance needs optimization") - (next-week - "Begin verisim-temporal implementation")) -``` - ---- - -== Testing Strategy - -=== Unit Tests -```bash -cargo test --package verisim-document -cargo test --package verisim-temporal -# ... etc -``` - -=== Integration Tests -```bash -cargo test --test integration_test -``` - -=== Fuzz Tests -```bash -cargo fuzz run fuzz_vcl -- -max_total_time=300 -``` - ---- - -## Checkpoints - -### Week 2 -- [ ] verisim-document complete -- [ ] Property tests passing -- [ ] Completion: 10% → 20% - -### Week 6 -- [ ] All 6 modality stores complete -- [ ] Unit tests passing -- [ ] Completion: 20% → 40% - -### Week 8 -- [ ] Octad layer complete -- [ ] Drift detection working -- [ ] Completion: 40% → 55% - -### Week 11 -- [ ] VCL execution engine complete -- [ ] EXPLAIN working -- [ ] Completion: 55% → 65% - -### Week 16 ✓ -- [ ] **All milestones V1-V5 complete** -- [ ] **Testing coverage >80%** -- [ ] **Completion: 70%** -- [ ] **Production-ready PoC** - ---- - -## Next Steps - -1. **Today**: Review this roadmap -2. **This week**: Begin verisim-document implementation -3. **Friday**: Update STATE.scm with Week 1 progress - ---- - -## No Integration with FormBD - -**Important**: This roadmap is VeriSimDB-only. FormBD is a separate project. - -**The only connection**: VeriSimDB could eventually federate to FormBD as an external data source: - -```sql --- VCL federating to FormBD (future feature) -SELECT * FROM octads@formbd_instance WHERE modality = 'document'; -``` - -But that's just treating FormBD like any other database endpoint (PostgreSQL, MongoDB, etc.). diff --git a/verisimdb/Justfile b/verisimdb/Justfile deleted file mode 100644 index 29c8d292..00000000 --- a/verisimdb/Justfile +++ /dev/null @@ -1,356 +0,0 @@ -# SPDX-License-Identifier: MPL-2.0 -# SPDX-FileCopyrightText: 2026 Jonathan D.A. Jewell (hyperpolymath) -# -# justfile — VeriSimDB -# Run with: just - -set shell := ["bash", "-euo", "pipefail", "-c"] - -# Default recipe: show help -default: - @just --list - -# ── Build ────────────────────────────────────────────────────── - -# Build Rust core (release) -build: - OPENSSL_NO_VENDOR=1 cargo build --release - -# Build Rust core (debug) -build-dev: - OPENSSL_NO_VENDOR=1 cargo build - -# Build Elixir orchestration layer -build-elixir: - cd elixir-orchestration && mix deps.get && mix compile - -# Build everything (Rust + Elixir) -build-all: build build-elixir - -# Compile Idris2 ABI definitions -build-abi: - cd src/abi && idris2 --build hypatia-abi.ipkg - -# Build Zig FFI bridge -build-ffi: - cd ffi/zig && zig build - -# Model-check TLA+ specifications (e.g. V5 Octad transaction atomicity) -verify-tlaplus: - #!/usr/bin/env bash - # Uses host Java if available, otherwise an ephemeral eclipse-temurin:21-jre - # container. Honours $TLA2TOOLS_JAR for a pre-fetched jar location. - set -euo pipefail - SPEC_DIR="verification/proofs/tlaplus" - TLA2TOOLS_JAR="${TLA2TOOLS_JAR:-$HOME/.local/share/tla2tools.jar}" - if [ ! -f "$TLA2TOOLS_JAR" ]; then - mkdir -p "$(dirname "$TLA2TOOLS_JAR")" - echo "Fetching tla2tools.jar -> $TLA2TOOLS_JAR" - curl -sSL -o "$TLA2TOOLS_JAR" \ - https://github.com/tlaplus/tlaplus/releases/latest/download/tla2tools.jar - fi - run_tlc() { - local spec="$1" cfg="$2" - echo "== TLC on $spec ($cfg)" - if command -v java >/dev/null 2>&1; then - (cd "$SPEC_DIR" && java -XX:+UseParallelGC -cp "$TLA2TOOLS_JAR" tlc2.TLC \ - -workers auto -config "$cfg" "$spec") - else - podman run --rm \ - -v "$PWD/$SPEC_DIR:/work:Z" \ - -v "$TLA2TOOLS_JAR:/tla2tools.jar:ro,Z" \ - -w /work docker.io/library/eclipse-temurin:21-jre \ - java -XX:+UseParallelGC -cp /tla2tools.jar tlc2.TLC \ - -workers auto -config "$cfg" "$spec" - fi - } - run_tlc OctadAtomicity.tla OctadAtomicity.cfg - run_tlc Normalizer.tla Normalizer.cfg - run_tlc Serializability.tla Serializability.cfg - -# ── Test ─────────────────────────────────────────────────────── - -# Run Rust tests -test: - OPENSSL_NO_VENDOR=1 cargo test - -# Run Elixir tests -test-elixir: - cd elixir-orchestration && mix test - -# Run Rust integration tests -test-integration: - OPENSSL_NO_VENDOR=1 cargo test --test integration - -# Run all tests (Rust + Elixir) -test-all: test test-elixir - -# ── Lint & Format ────────────────────────────────────────────── - -# Format all Rust code -fmt: - cargo fmt - -# Run clippy lints -lint: - cargo clippy -- -D warnings - -# Format Elixir code -fmt-elixir: - cd elixir-orchestration && mix format - -# ── Run ──────────────────────────────────────────────────────── - -# Run verisimdb API server (dev mode) -serve: - RUST_LOG=debug cargo run -p verisim-api - -# Run Elixir OTP orchestrator -serve-otp: - cd elixir-orchestration && MIX_ENV=dev mix run --no-halt - -# ── Container ────────────────────────────────────────────────── - -# Build container image with Podman -container-build: - podman build -t verisimdb:latest -f container/Containerfile . - -# Run container locally -container-run: - podman run --rm -p 8080:8080 verisimdb:latest - -# Build with stapeln layers -stapeln-build: - @if command -v stapeln &>/dev/null; then \ - stapeln build --config stapeln.toml --target production; \ - else \ - echo "stapeln not found — falling back to podman build"; \ - just container-build; \ - fi - -# Deploy full stack with selur -deploy: - @if command -v selur &>/dev/null; then \ - selur seal && podman-compose -f selur-compose.yml up -d; \ - else \ - echo "selur not found — using podman-compose directly"; \ - podman-compose -f selur-compose.yml up -d; \ - fi - -# Stop deployed stack -deploy-stop: - podman-compose -f selur-compose.yml down - -# Sign container with cerro-torre -container-sign: - @if command -v cerro-torre &>/dev/null; then \ - cerro-torre sign verisimdb:latest --algorithm ML-DSA-87; \ - else \ - echo "cerro-torre not found — skipping image signing"; \ - fi - -# ── Security ─────────────────────────────────────────────────── - -# Run panic-attack static analysis -panic-scan: - @if [ -x "/var$REPOS_DIR/panic-attacker/target/release/panic-attack" ]; then \ - /var$REPOS_DIR/panic-attacker/target/release/panic-attack assail . --verbose; \ - else \ - echo "panic-attack not built — run 'cd /var$REPOS_DIR/panic-attacker && cargo build --release'"; \ - fi - -# Run hypatia neurosymbolic scan -hypatia-scan: - @if command -v hypatia-v2 &>/dev/null; then \ - hypatia-v2 . --severity=critical --severity=high; \ - else \ - echo "hypatia-v2 not found — run via CI workflow instead"; \ - fi - -# Run vordr runtime verification -vordr-verify: - @if command -v vordr &>/dev/null; then \ - vordr verify --target localhost:8080 --policy strict; \ - else \ - echo "vordr not found — skipping runtime verification"; \ - fi - -# Check license compliance -license-check: - @echo "Checking for banned AGPL-3.0 headers..." - @if grep -rl "AGPL-3.0" --include='*.rs' --include='*.ex' --include='*.exs' --include='*.idr' --include='*.zig' --include='*.yml' . 2>/dev/null; then \ - echo "FAIL: Found AGPL-3.0 headers"; \ - exit 1; \ - else \ - echo "PASS: No AGPL-3.0 headers found"; \ - fi - -# Validate SCM files are in .machine_readable/ only -check-scm: - @for f in STATE.scm META.scm ECOSYSTEM.scm; do \ - if [ -f "$$f" ]; then \ - echo "ERROR: $$f found in root"; exit 1; \ - fi; \ - done - @echo "PASS: No SCM files in root" - -# ── Clean ────────────────────────────────────────────────────── - -# Clean all build artifacts -clean: - cargo clean - cd elixir-orchestration && mix clean 2>/dev/null || true - @echo "Cleaned." - -# Run panic-attacker pre-commit scan -assail: - @command -v panic-attack >/dev/null 2>&1 && panic-attack assail . || echo "panic-attack not found — install from https://github.com/hyperpolymath/panic-attacker" - -# ── Onboarding ──────────────────────────────────────────────── - -# Check all required tools are installed -doctor: - #!/usr/bin/env bash - set -euo pipefail - ok=0; fail=0 - check() { - if "$@" >/dev/null 2>&1; then - echo " [ok] $1" - ((ok++)) - else - echo " [MISSING] $1 — $2" - ((fail++)) - fi - } - echo "=== VeriSimDB Doctor ===" - check rustc --version "install via asdf: asdf install rust nightly" - check cargo --version "comes with Rust" - check rustup --version "https://rustup.rs" - check pkg-config --version "sudo dnf install pkg-config" - if pkg-config --exists openssl 2>/dev/null; then - echo " [ok] openssl-devel (pkg-config)" - ((ok++)) - else - echo " [MISSING] openssl-devel — sudo dnf install openssl-devel" - ((fail++)) - fi - check elixir --version "asdf install elixir 1.17.3-otp-27" - check mix --version "comes with Elixir" - check erl -version "asdf install erlang 27.2" - check zig version "asdf install zig 0.14.0" - check just --version "cargo install just" - check podman --version "sudo dnf install podman (optional, for containers)" - if command -v idris2 >/dev/null 2>&1; then - echo " [ok] idris2 (optional — ABI layer)" - ((ok++)) - else - echo " [info] idris2 not found (optional — only for ABI definitions)" - fi - echo "" - echo "Result: $ok passed, $fail failed" - if [ "$fail" -gt 0 ]; then - echo "Fix the MISSING items above, then re-run: just doctor" - exit 1 - else - echo "All prerequisites satisfied." - fi - -# Auto-install missing tools where possible -heal: - #!/usr/bin/env bash - set -euo pipefail - echo "=== VeriSimDB Heal ===" - if ! command -v rustc &>/dev/null; then - echo "Installing Rust via asdf..." - asdf install rust nightly || echo "Try: curl --proto '=https' --tlsv1.2 -sSf https://sh.rustup.rs | sh" - fi - if ! command -v just &>/dev/null; then - echo "Installing just..." - cargo install just - fi - if ! command -v elixir &>/dev/null; then - echo "Installing Elixir via asdf..." - asdf install elixir 1.17.3-otp-27 || echo "Try: asdf plugin add elixir && asdf install elixir 1.17.3-otp-27" - fi - if ! command -v zig &>/dev/null; then - echo "Installing Zig via asdf..." - asdf install zig 0.14.0 || echo "Try: asdf plugin add zig && asdf install zig 0.14.0" - fi - if ! pkg-config --exists openssl 2>/dev/null; then - echo "openssl-devel missing — run: sudo dnf install openssl-devel" - fi - if ! command -v podman &>/dev/null; then - echo "Podman missing — run: sudo dnf install podman podman-compose" - fi - echo "" - echo "Re-run 'just doctor' to verify." - -# Guided tour of the codebase -tour: - #!/usr/bin/env bash - set -euo pipefail - echo "=== VeriSimDB Tour ===" - echo "" - echo "1. ARCHITECTURE" - echo " Rust core (rust-core/) provides 10 modality crates:" - echo " graph, vector, tensor, semantic, document, temporal," - echo " provenance, spatial, octad, drift, normalizer, api" - echo " Elixir OTP (elixir-orchestration/) coordinates them." - echo "" - echo "2. BUILD & RUN" - echo " just build Build Rust release" - echo " just build-elixir Build Elixir layer" - echo " just serve Start API on :8080" - echo "" - echo "3. THE OCTAD" - echo " Every entity is stored across 8 modalities simultaneously." - echo " Drift between them is detected and self-healed." - echo "" - echo "4. QUERY LANGUAGE" - echo " VCL (VeriSim Consonance Language) — NOT SQL." - echo " See docs/ and playground/ for examples." - echo "" - echo "5. FEDERATION" - echo " 10 adapters (MongoDB, Redis, Neo4j, ClickHouse, SurrealDB," - echo " SQLite, DuckDB, VectorDB, InfluxDB, ObjectStorage)." - echo " 6 client SDKs (Rust, V, Elixir, ReScript, Julia, Gleam)." - echo "" - echo "6. CONTAINERS" - echo " just container-build Build with Podman" - echo " just container-run Run on :8080" - echo "" - echo "7. KEY FILES" - echo " Cargo.toml Workspace definition" - echo " elixir-orchestration/ OTP layer" - echo " connectors/ Federation + SDKs" - echo " container/ Containerfile + compose" - echo " .claude/CLAUDE.md Full AI context" - echo "" - echo "Run 'just' to see all available recipes." - -# What to do when things go wrong -help-me: - #!/usr/bin/env bash - echo "=== VeriSimDB Help ===" - echo "" - echo "BUILD FAILS:" - echo " 'openssl' errors -> sudo dnf install openssl-devel" - echo " 'protoc' errors -> Proto code is pre-generated, check Cargo features" - echo " 'oxrocksdb' errors -> Eliminated; if seen, run: cargo clean && just build" - echo " Elixir errors -> cd elixir-orchestration && mix deps.get" - echo "" - echo "RUNTIME ISSUES:" - echo " Port 8080 in use -> Change port: VERISIM_PORT=8081 just serve" - echo " 'connection refused'-> Is the Rust API running? just serve" - echo " Drift not detected -> Check thresholds in config/config.exs" - echo "" - echo "TESTING:" - echo " Integration tests need the test-infra stack running:" - echo " cd connectors/test-infra && podman-compose up -d" - echo " Then: just test-integration" - echo "" - echo "STILL STUCK?" - echo " 1. just doctor (check prerequisites)" - echo " 2. just heal (auto-install what's missing)" - echo " 3. cargo clean && just build (fresh build)" - echo " 4. Read .claude/CLAUDE.md for full context" diff --git a/verisimdb/KNOWN-ISSUES.adoc b/verisimdb/KNOWN-ISSUES.adoc deleted file mode 100644 index bf2f76d1..00000000 --- a/verisimdb/KNOWN-ISSUES.adoc +++ /dev/null @@ -1,250 +0,0 @@ -// SPDX-License-Identifier: CC-BY-SA-4.0 -// Copyright (c) 2026 Jonathan D.A. Jewell (hyperpolymath) - -= Known Issues and Honest Gaps -:toc: left -:toclevels: 2 -:sectnums: - -== Overview - -This document lists known issues, incomplete implementations, and honest gaps in VeriSimDB as of 2026-02-13. Each entry describes the current state, why it matters, and what needs to happen to resolve it. - -A major 7-phase hardening effort was completed on 2026-02-13, resolving most previously-open issues. The remaining open items are documented below. - -VeriSimDB is a working system with real implementations in its core modality stores, but several higher-level features are stubs or partially implemented. This document does not hide behind "future work" euphemisms -- it states plainly what does not work. - -== Issues - -=== 1. Normalizer Regeneration Strategies Are Stubs — ✅ RESOLVED - -**Location:** `rust-core/verisim-normalizer/` - -**Resolved:** 2026-02-12. Regeneration strategies now inspect octad data to select the authoritative modality and derive drifted modality content from it, rather than returning hardcoded placeholder strings. Each of the six modality strategies performs actual data transformation. - -**Original issue:** Every regeneration strategy returned a hardcoded `[regenerated]` placeholder string instead of actually regenerating data. - -=== 2. Federation Executor Always Returns Empty — ✅ RESOLVED - -**Location:** Elixir orchestration layer, `VeriSim.FederationExecutor` - -**Resolved:** 2026-02-12. Federation executor now performs parallel HTTP fanout to registered peers via reqwest, decomposes queries into per-peer sub-queries, dispatches them, and merges results. Uses `Task.async_stream` for concurrent peer dispatch with configurable timeouts. - -**Original issue:** The federation executor always returned `{:ok, []}` regardless of query content or number of registered peers. - -=== 3. Federation Resolver Peer Queries Unimplemented — ✅ RESOLVED - -**Location:** Elixir orchestration layer, `VeriSim.Federation.Resolver` - -**Resolved:** 2026-02-12. `query_peer/3` now makes real HTTP requests to peer endpoints via `Req` with configurable timeout, parses JSON responses, and returns structured results with source store attribution and response time tracking. Drift policy filtering (strict/repair/tolerate/latest) is implemented. - -**Original issue:** The `query_peer/3` function returned `{:error, :not_implemented}`. Peers could be registered but never contacted. - -=== 4. Drift Auto-Trigger Missing — ✅ RESOLVED - -**Location:** Elixir orchestration layer, `VeriSim.DriftMonitor` - -**Resolved:** 2026-02-12. `DriftMonitor.init/1` now schedules a periodic sweep via `Process.send_after` at a configurable interval (default 60s). The `:sweep` handler performs a full drift sweep including pulling aggregate metrics from the Rust core via `RustClient.drift_status/0`, then reschedules itself. The sweep interval is configurable via the `config` option passed to `start_link`. - -**Original issue:** Drift detection had to be triggered manually. No scheduled or event-driven trigger existed. - -=== 5. VCL-UT Not Connected to VCL PROOF Runtime — ✅ RESOLVED - -**Location:** `elixir-orchestration/lib/verisim/query/vcl_type_checker.ex`, `vcl_executor.ex`, `vcl_bridge.ex` - -**Resolved:** 2026-02-28. VCL-UT proof pipeline is now fully wired end-to-end with a three-tier type checking strategy: - -1. **ReScript bidirectional type checker** (VCLBidir.res, 852 lines) — full formal system with subtyping, invoked via VCLBridge.typecheck/2 when Deno subprocess available -2. **Elixir-native type checker** (VCLTypeChecker, 320 lines) — validates proof types, modality compatibility, composition rules; generates structured obligations with witness fields and circuit names. Used when ReScript subprocess unavailable. -3. **Bare AST extraction** — last resort, no longer needed since the native checker handles all cases - -Additionally: - -- Built-in parser now splits multi-proof specs (`PROOF A(x) AND B(y)` → separate proof specs) -- Added CONSISTENCY and FRESHNESS proof types (now 11 total: EXISTENCE, INTEGRITY, CONSISTENCY, PROVENANCE, FRESHNESS, ACCESS, CITATION, CUSTOM, ZKP, PROVEN, SANCTIFY) -- Type checker validates modality compatibility (e.g., INTEGRITY requires semantic modality) -- Proof obligations carry witness fields, circuit names, and time estimates to Rust endpoints -- 34 type checker tests + 109 query tests pass - -**Original issue:** PROOF clauses were parsed but type checking only worked when the ReScript subprocess was running (which it typically wasn't). Queries silently fell back to unvalidated AST extraction. - -=== 6. ZKP / Proven Library Not Integrated - -**Location:** `docs/zkp-and-sanctify-integration.adoc` (design), `proven-coherence.md` (notes) - -**Current state:** The integration of zero-knowledge proofs via the sanctify library is documented as a consultation paper and design specification. The integration is not implemented. No ZKP circuits are generated, no proofs are created, and no verification occurs at runtime. - -**Impact:** Privacy-preserving proofs (where a query result can be verified without revealing the underlying data) are not available. This affects multi-tenant federation and compliance scenarios where data must remain private but provenance must be verifiable. - -**Resolution:** Implement sanctify integration after VCL-UT (issue 5) is resolved, since ZKP proof generation depends on the dependent type checking pipeline being functional. - -=== 7. ReScript Registry 60% Complete — ✅ RESOLVED - -**Location:** `src/registry/Registry.res`, `rescript.json` - -**Resolved:** 2026-02-13. All 7 previously stubbed functions now have real implementations: - -- `executeFederatedQuery`: HTTP fan-out to eligible peer stores with trust filtering, response mapping, and result aggregation with limit enforcement. -- `achieveConsensus`: Quorum-based consensus with configurable mode (Strong/Quorum/Eventual), concurrent store fetch, agreement calculation. -- `replicateOctad`: Fetches octad from source store via HTTP, pushes to all target stores, reports per-target errors. -- `checkReplicationStatus`: Examines mapping locations, checks store liveness against maxStoreDowntimeMs, detects trust divergence as data divergence proxy. -- `detectByzantineFaults`: Computes median trust across stores for a octad, flags stores deviating >0.3 from median. -- `serializeRegistry`/`deserializeRegistry`: Full JSON serialization and deserialization of registry state (stores, mappings, config). - -Added `rescript.json` for build configuration. - -=== 8. `debugger/Cargo.toml` Had Wrong Author - -**Location:** `debugger/Cargo.toml` - -**Current state:** Fixed. The `debugger/Cargo.toml` previously listed an incorrect author. It now correctly reads `authors = ["Jonathan D.A. Jewell "]`. - -**Impact:** None (resolved). Documented here for audit trail completeness. - -=== 9. HNSW Implementation Is Real - -**Location:** `rust-core/verisim-vector/src/hnsw.rs` - -**Current state:** The HNSW (Hierarchical Navigable Small World) implementation in `verisim-vector` is a genuine, functional implementation at approximately 670 lines of Rust. A previous audit incorrectly claimed this was brute-force search. This is **not** an issue -- it is a correction of a previous mischaracterization. - -**Impact:** None (this is a positive clarification). The vector modality store uses a real approximate nearest-neighbor algorithm, not a naive linear scan. - -**Note:** This entry exists to prevent future audits from repeating the same incorrect claim. - -=== 10. verisim-api Needs Bin Target Fix — ✅ RESOLVED - -**Location:** `rust-core/verisim-api/Cargo.toml` - -**Resolved:** 2026-02-12. `src/main.rs` now exists with a `main()` function that reads host/port from environment variables (`VERISIM_HOST`/`VERISIM_PORT`) and initializes the HTTP server. `cargo run -p verisim-api` works. - -**Original issue:** The crate was missing a `src/main.rs` entry point and could not be run as a standalone binary. - -=== 11. No Performance Baselines Established — ✅ RESOLVED - -**Location:** `benches/modality_benchmarks.rs` - -**Resolved:** 2026-02-13. Criterion benchmarks now cover all 6 modality stores plus cross-modal and drift operations: - -- **Document:** create_document, search_text (Tantivy full-text, 1000-doc corpus) -- **Vector:** insert (128/384/768 dims), search (10k vectors, HNSW) -- **Graph:** add_node, add_edge (Oxigraph) -- **Tensor:** store_create_64x64, store_get, reduce_sum_axis0 (ndarray) -- **Semantic:** register_type, get_type, proof_create_cbor, proof_verify (CBOR + ZKP verification) -- **Temporal:** version_create, version_get_by_number, version_get_latest, history_10, history_100 -- **Octad:** create_octad, get_octad (unified entity) -- **Drift:** calculate_drift (semantic vector drift) -- **Cross-modal:** vector_similarity_search, fulltext_search (1000 multi-modal octads) - -**Original issue:** Only document, vector, graph, octad, drift, and cross-modal benchmarks existed. Tensor, semantic, and temporal stores had zero benchmarks. - -=== 12. Cross-Modal Drift/Consistency Were Stubs — ✅ RESOLVED - -**Location:** `elixir-orchestration/lib/verisim/query/vcl_executor.ex` - -**Resolved:** 2026-02-13. `compute_modality_drift/3` now fetches drift from the Rust drift API when available, falling back to cosine distance between extracted modality embeddings (with content fingerprinting for non-vector modalities). `compute_consistency/4` now computes real scores using the specified metric (COSINE, EUCLIDEAN, DOT_PRODUCT, JACCARD). Both functions previously returned hardcoded constants (0.0 and 0.5). - -**Original issue:** Cross-modal correlation queries parsed correctly but evaluation returned fake scores, making WHERE DRIFT(...) and CONSISTENT(...) conditions meaningless. - -=== 13. VCL WHERE Condition Routing Was Broken — ✅ RESOLVED - -**Location:** `elixir-orchestration/lib/verisim/query/vcl_executor.ex` - -**Resolved:** 2026-02-13. `has_fulltext_condition?/1`, `has_vector_condition?/1`, and `has_graph_pattern?/1` now walk the AST recursively to detect actual condition types. `extract_text_query/1`, `extract_vector_query/1`, and `extract_graph_query/1` now parse actual values from the AST instead of returning hardcoded placeholders. All queries were previously routed to `:multi` type regardless of conditions. - -**Original issue:** All six condition detection and extraction functions were stubs (returning `false` and hardcoded values), causing every query to be routed as a multi-modal query even when a single modality was targeted. - -=== 14. Proof Verification Was No-Op — ✅ RESOLVED - -**Location:** `elixir-orchestration/lib/verisim/query/vcl_executor.ex` - -**Resolved:** 2026-02-13. `verify_single_proof/1` now validates proof type, extracts contract names, and checks contract existence against the semantic store. It properly rejects queries with invalid or missing contracts for CITATION, INTEGRITY, and CUSTOM proof types. Previously it always returned `:ok`. - -**Original issue:** PROOF clauses in VCL queries were parsed but never verified. All proofs silently passed, making the entire proof system decorative. - -=== 15. believe_me in Idris2 ABI Files — ✅ RESOLVED - -**Location:** `src/abi/Foreign.idr`, `debugger/src/abi/Foreign.idr`, `practice-mirror/src/abi/Foreign.idr` - -**Resolved:** 2026-02-13. FFI declarations changed from `Bits64 -> AnyPtr -> PrimIO Bits32` to `Bits64 -> (Bits64 -> Bits32 -> Bits32) -> PrimIO Bits32`, making the callback type match the declaration and eliminating the need for `believe_me` casts. Zero `believe_me` calls remain in the codebase. - -**Original issue:** `registerCallback` used `believe_me` to cast a callback function to `AnyPtr`, a BANNED unsafe pattern that bypasses the type checker. - -=== 16. Atom Table Exhaustion Risk in VCL Bridge — ✅ RESOLVED - -**Location:** `elixir-orchestration/lib/verisim/query/vcl_bridge.ex` - -**Resolved:** 2026-02-13. All 8 `String.to_atom` calls replaced with `safe_to_atom/1` helper that uses an explicit allowlist map for the known atom values (6 modalities + 5 aggregate functions + `all`), falling back to `String.to_existing_atom/1`. - -**Original issue:** 8 `String.to_atom` calls could theoretically exhaust the BEAM atom table if fed arbitrary input, since the BEAM atom table is finite and atoms are never garbage collected. - -=== 17. No ETS Caching in RustClient — ✅ RESOLVED - -**Location:** `elixir-orchestration/lib/verisim/rust_client.ex` - -**Resolved:** 2026-02-13. Added ETS-based read-through cache with configurable TTL (30s default for octads, 10s for drift scores). Cache is invalidated on writes (update/delete). Provides `init_cache/0`, `clear_cache/0`, and `invalidate_cache/1` for cache management. - -**Original issue:** Every RustClient call made a fresh HTTP request to the Rust core, with no caching. Repeated reads of the same octad in a query pipeline (e.g., cross-modal evaluation fetching the same entity multiple times) each incurred full HTTP round-trip latency. - -=== 18. EXPLAIN Returns Hardcoded Plan — ✅ RESOLVED - -**Location:** `elixir-orchestration/lib/verisim/query/vcl_executor.ex` - -**Resolved:** 2026-02-13. `generate_explain_plan/1` now analyzes the actual query AST to produce cost estimates based on source type, modality count, WHERE clause complexity, cross-modal conditions, GROUP BY presence, and proof obligations. Delegates to the Rust verisim-planner API when available, with local estimation as fallback. - -**Original issue:** EXPLAIN returned a static plan with hardcoded step names and costs totalling 71ms, regardless of the actual query structure. - -=== 19. VCL Executor Federation Stub — ✅ RESOLVED - -**Location:** `elixir-orchestration/lib/verisim/query/vcl_executor.ex`, `rust-core/verisim-api/src/federation.rs` - -**Resolved:** 2026-02-13. Added `GET /octads` list endpoint with `?limit=N&offset=M` pagination to the Rust API. The `OctadStore` trait now includes a `list(limit, offset)` method. Federation `query_single_peer()` now falls back to the `/octads` list endpoint when neither text_query nor vector_query is provided, so bare `SELECT * FROM FEDERATION /pattern/*` returns actual octad data instead of empty results. - -=== 20. Custom Circuit Hardcoding — ✅ RESOLVED - -**Location:** `rust-core/verisim-semantic/src/circuit_registry.rs`, `circuit_compiler.rs`, `verification_keys.rs`, `src/vcl/VCLCircuit.res` - -**Resolved:** 2026-02-13. Full custom circuit infrastructure implemented: - -- Circuit Registry: named circuit storage with register/get/verify/list/unregister operations -- Circuit Compiler: DSL gates (AND, OR, XOR, NOT, LinearCombination) compiled to R1CS constraints -- Verification Key Store: per-circuit keys with rotation support and federation export/import -- VCL Circuit DSL: ReScript types for circuit definition with `PROOF CUSTOM "name" WITH (param=value)` -- 25 semantic tests pass covering circuit operations - -**Original issue:** The `CUSTOM` proof type used a hardcoded circuit configuration with no DSL, compiler, or registry. - -=== 21. VCL-UT Not Connected to VCL PROOF Runtime — ✅ RESOLVED - -**Location:** `elixir-orchestration/lib/verisim/query/vcl_type_checker.ex` - -**Resolved:** 2026-02-28. Superseded by issue #5 resolution. The VCL-UT pipeline is now wired with an Elixir-native type checker that validates proof types, modality compatibility, and composition rules without requiring the ReScript subprocess or a Lean checker. The Lean formal verification path remains aspirational (no Lean files were ever created); the practical implementation uses the ReScript bidirectional type checker (852 lines) backed by an Elixir fallback (320 lines). - -**Note:** Formal verification via Lean/Idris2 would be a future enhancement for mathematical proof certificates. The current system provides structural type checking and real cryptographic proofs (SHA-256 commitments, Merkle proofs, R1CS circuit verification) via the Rust ZKP bridge. - -=== 22. proven Library Not Integrated — ✅ RESOLVED - -**Location:** `rust-core/verisim-semantic/src/proven_bridge.rs` - -**Resolved:** 2026-02-13. Created `proven_bridge.rs` (277 LOC) that parses JSON/CBOR proof certificates from the proven library, verifies signatures, and converts certificates to ProofBlob format for storage in the semantic modality. VCL executor routes `PROOF PROVEN(...)` queries through the bridge. Integration is certificate-based (JSON exchange) rather than direct Idris2 FFI. - -=== 23. verisim-repl Has Build Issues — ✅ RESOLVED - -**Location:** `rust-core/verisim-repl/` - -**Resolved:** 2026-02-13. REPL builds and all 67 tests pass. Fixed by a previous session that resolved rustyline API incompatibilities. Also added `rustls` dependency with ring crypto provider installation for TLS-free HTTP client. - -=== 24. oxrocksdb-sys C++ Dependency — ✅ RESOLVED - -**Location:** `rust-core/verisim-graph/` - -**Resolved:** 2026-02-28. Oxigraph was already feature-flagged as optional (`oxigraph-backend`, off by default) in Phase 6. A pure-Rust persistent graph backend using **redb** (B-tree, ACID, single-file) was added as `redb-backend` feature in verisim-graph. The default graph store is `SimpleGraphStore` (in-memory, zero C/C++ deps). The Containerfile no longer references clang-19 or any C++ toolchain. - -**Original issue:** Oxigraph 0.4 unconditionally pulled in `oxrocksdb-sys` (~400K lines of C++), requiring clang-19 and adding ~1GB + ~15 min to container builds. - -=== 25. protoc Binary Required at Build Time — ✅ RESOLVED - -**Location:** `rust-core/verisim-api/src/proto/verisim.rs` - -**Resolved:** 2026-02-28. Protobuf code is now pre-generated and committed at `src/proto/verisim.rs`. The `build.rs` is a no-op. The `prost-build`/`tonic-build` build-dependencies were removed. The Containerfile no longer installs protoc. - -**Original issue:** `prost-build` shelled out to the `protoc` binary to compile `verisim.proto`. To regenerate after changing the proto, install protoc and run: `protoc --prost_out=src/proto proto/verisim.proto`. diff --git a/verisimdb/LICENSE b/verisimdb/LICENSE deleted file mode 100644 index ec540b34..00000000 --- a/verisimdb/LICENSE +++ /dev/null @@ -1,153 +0,0 @@ -SPDX-License-Identifier: MPL-2.0 -SPDX-FileCopyrightText: 2024-2025 Palimpsest Stewardship Council - -================================================================================ -PALIMPSEST-MPL LICENSE VERSION 1.0 -================================================================================ - -File-level copyleft with ethical use and quantum-safe provenance - -Based on Mozilla Public License 2.0 - --------------------------------------------------------------------------------- -PREAMBLE --------------------------------------------------------------------------------- - -This License extends the Mozilla Public License 2.0 (MPL-2.0) with provisions -for ethical use, post-quantum cryptographic provenance, and emotional lineage -protection. The base MPL-2.0 terms apply except where explicitly modified by -the Exhibits below. - -Like a palimpsest manuscript where each layer builds upon what came before, -this license recognizes that creative works carry history, context, and meaning -that transcend mere code or text. - --------------------------------------------------------------------------------- -SECTION 1: BASE LICENSE --------------------------------------------------------------------------------- - -This License incorporates the full text of Mozilla Public License 2.0 by -reference. The complete MPL-2.0 text is available at: -https://www.mozilla.org/en-US/MPL/2.0/ - -All terms, conditions, and definitions from MPL-2.0 apply except where -explicitly modified by the Exhibits in this License. - --------------------------------------------------------------------------------- -SECTION 2: ADDITIONAL DEFINITIONS --------------------------------------------------------------------------------- - -2.1. "Emotional Lineage" - means the narrative, cultural, symbolic, and contextual meaning embedded - in Covered Software, including but not limited to: protest traditions, - cultural heritage, trauma narratives, and community stories. - -2.2. "Provenance Metadata" - means cryptographically signed attribution information attached to or - associated with Covered Software, including author identities, timestamps, - modification history, and lineage references. - -2.3. "Non-Interpretive System" - means any automated system that processes Covered Software without - preserving or considering its Emotional Lineage, including but not - limited to: AI training pipelines, content aggregators, and automated - summarization tools. - -2.4. "Quantum-Safe Signature" - means a cryptographic signature using algorithms resistant to attacks - by quantum computers, as specified in Exhibit B. - --------------------------------------------------------------------------------- -SECTION 3: ETHICAL USE REQUIREMENTS --------------------------------------------------------------------------------- - -In addition to the rights and obligations under MPL-2.0: - -3.1. Emotional Lineage Preservation - You must make reasonable efforts to preserve and communicate the - Emotional Lineage of Covered Software when distributing or creating - derivative works. This includes maintaining narrative context, cultural - attributions, and symbolic meaning where documented. - -3.2. Non-Interpretive System Notice - If You use Covered Software as input to a Non-Interpretive System, You - must: - (a) document such use in a publicly accessible manner; and - (b) not claim that outputs of such systems carry the Emotional Lineage - of the original work without explicit permission from Contributors. - -3.3. Ethical Use Declaration - Commercial use of Covered Software requires acknowledgment that You have - read and understood Exhibit A (Ethical Use Guidelines) and agree to act - in good faith accordance with its principles. - -See Exhibit A for complete Ethical Use Guidelines. - --------------------------------------------------------------------------------- -SECTION 4: PROVENANCE REQUIREMENTS --------------------------------------------------------------------------------- - -4.1. Metadata Preservation - You must not strip, alter, or obscure Provenance Metadata from Covered - Software except where technically necessary and with clear documentation - of any changes. - -4.2. Quantum-Safe Provenance (Optional) - Contributors may sign their Contributions using Quantum-Safe Signatures. - If Quantum-Safe Signatures are present, You must preserve them in all - distributions. - -4.3. Lineage Chain - When creating derivative works, You should extend the provenance chain - to include Your own contributions, maintaining cryptographic linkage to - prior Contributors where feasible. - -See Exhibit B for Quantum-Safe Provenance specifications. - --------------------------------------------------------------------------------- -SECTION 5: GOVERNANCE --------------------------------------------------------------------------------- - -5.1. Stewardship Council - This License is maintained by the Palimpsest Stewardship Council, which - may issue clarifications, interpretive guidance, and future versions. - -5.2. Version Selection - You may use Covered Software under this version of the License or any - later version published by the Palimpsest Stewardship Council. - -5.3. Dispute Resolution - Disputes regarding interpretation of Ethical Use Requirements (Section 3) - should first be submitted to the Palimpsest Stewardship Council for - non-binding guidance before pursuing legal remedies. - --------------------------------------------------------------------------------- -SECTION 6: COMPATIBILITY --------------------------------------------------------------------------------- - -6.1. MPL-2.0 Compatibility - Covered Software under this License may be combined with software under - MPL-2.0. The combined work must comply with both licenses. - -6.2. Secondary Licenses - The Secondary License provisions of MPL-2.0 Section 3.3 apply to this - License. - --------------------------------------------------------------------------------- -EXHIBITS --------------------------------------------------------------------------------- - -Exhibit A - Ethical Use Guidelines -Exhibit B - Quantum-Safe Provenance Specification - -See separate files: -- EXHIBIT-A-ETHICAL-USE.txt -- EXHIBIT-B-QUANTUM-SAFE.txt - --------------------------------------------------------------------------------- -END OF PALIMPSEST-MPL LICENSE VERSION 1.0 --------------------------------------------------------------------------------- - -For questions about this License: -- Repository: https://github.com/hyperpolymath/palimpsest-license -- Council: contact via repository Issues diff --git a/verisimdb/MAINTAINERS.adoc b/verisimdb/MAINTAINERS.adoc deleted file mode 100644 index 48d97817..00000000 --- a/verisimdb/MAINTAINERS.adoc +++ /dev/null @@ -1,47 +0,0 @@ -// SPDX-License-Identifier: CC-BY-SA-4.0 -= Maintainers -:toc: preamble - -This document lists the maintainers of this project and their responsibilities. - -== Current Maintainers - -[cols="2,3,2",options="header"] -|=== -| Name | Role | Contact - -| Jonathan D.A. Jewell -| Lead Maintainer -| https://github.com/hyperpolymath[@hyperpolymath] -|=== - -== Responsibilities - -Maintainers are responsible for: - -* Reviewing and merging pull requests -* Triaging issues and feature requests -* Ensuring code quality and security standards -* Managing releases and versioning -* Upholding the project's code of conduct - -== Becoming a Maintainer - -Contributors who demonstrate: - -* Consistent, high-quality contributions -* Understanding of the project's goals and standards -* Constructive participation in discussions -* Commitment to the project's long-term health - -May be invited to become maintainers at the discretion of existing maintainers. - -== Decision Making - -* Routine decisions (bug fixes, minor improvements) can be made by any maintainer -* Significant changes require discussion and consensus among maintainers -* Breaking changes or major features should be discussed in issues before implementation - -== Contact - -For questions about project governance, open an issue or contact the maintainers listed above. diff --git a/verisimdb/PLANNER-IMPLEMENTATION-STATUS.md b/verisimdb/PLANNER-IMPLEMENTATION-STATUS.md deleted file mode 100644 index 863e9af0..00000000 --- a/verisimdb/PLANNER-IMPLEMENTATION-STATUS.md +++ /dev/null @@ -1,278 +0,0 @@ -# verisim-planner Implementation Status - -**Last updated:** 2026-02-13 -**Author:** Claude Opus 4.6 session -**Recovery doc:** If this machine crashes, this file tells you exactly where things stand. - -## Overview - -Adding a cost-based query planner (`verisim-planner`) to VeriSimDB's Rust core, plus -triple API (REST + GraphQL + gRPC) and VCL integration. - ---- - -## Phase 1: verisim-planner crate — COMPLETE - -**Status:** Done. Builds clean, 39/39 tests pass, zero clippy warnings. - -### Files Created - -| File | Lines | Status | -|------|-------|--------| -| `rust-core/verisim-planner/Cargo.toml` | 16 | Done | -| `rust-core/verisim-planner/src/lib.rs` | ~120 | Done — Modality enum, re-exports | -| `rust-core/verisim-planner/src/error.rs` | ~20 | Done — PlannerError | -| `rust-core/verisim-planner/src/plan.rs` | ~130 | Done — LogicalPlan, PhysicalPlan, PlanNode | -| `rust-core/verisim-planner/src/cost.rs` | ~220 | Done — CostModel, CostEstimate, BaseCost | -| `rust-core/verisim-planner/src/stats.rs` | ~110 | Done — StatisticsCollector | -| `rust-core/verisim-planner/src/config.rs` | ~100 | Done — PlannerConfig, OptimizationMode | -| `rust-core/verisim-planner/src/optimizer.rs` | ~200 | Done — Planner with optimize() + explain() | -| `rust-core/verisim-planner/src/explain.rs` | ~210 | Done — ExplainOutput text/JSON rendering | - -### Files Modified - -| File | Change | Status | -|------|--------|--------| -| `Cargo.toml` (workspace root) | Added `verisim-planner` to members | Done | -| `rust-core/verisim-api/Cargo.toml` | Added `verisim-planner` dependency | Done | -| `rust-core/verisim-api/src/lib.rs` | Added Planner to AppState, 5 REST endpoints | Done | - -### REST Endpoints Added - -| Endpoint | Method | Status | -|----------|--------|--------| -| `/query/plan` | POST | Done | -| `/query/explain` | POST | Done | -| `/planner/config` | GET | Done | -| `/planner/config` | PUT | Done | -| `/planner/stats` | GET | Done | - -### Verification - -```bash -cargo build -p verisim-planner # Clean build, 0 warnings -cargo test -p verisim-planner # 39/39 pass -cargo build -p verisim-api # Clean build -cargo test -p verisim-api # 7/7 pass -cargo clippy -p verisim-planner # 0 warnings -``` - ---- - -## Phase 2: Triple API — IN PROGRESS - -### 2a. REST — COMPLETE (done in Phase 1) - -### 2b. GraphQL — NOT STARTED - -**Plan:** Add `async-graphql` + `async-graphql-axum` to verisim-api. - -**Dependencies needed:** -```toml -async-graphql = "7" # Latest stable (8.x is rc only) -async-graphql-axum = "7" -``` - -**Files to create/modify:** -- `rust-core/verisim-api/src/graphql.rs` — Schema types, Query root, Mutation root -- `rust-core/verisim-api/src/lib.rs` — Add `/graphql` route, add schema to AppState - -**GraphQL Schema (planned):** -```graphql -type Query { - health: Health! - octad(id: ID!): Octad - searchText(query: String!, limit: Int): [SearchResult!]! - driftStatus: [DriftStatus!]! - plannerConfig: PlannerConfig! - plannerStats: PlannerStats! - explainPlan(plan: LogicalPlanInput!): ExplainOutput! -} - -type Mutation { - createOctad(input: OctadInput!): Octad! - updateOctad(id: ID!, input: OctadInput!): Octad! - deleteOctad(id: ID!): Boolean! - optimizePlan(plan: LogicalPlanInput!): PhysicalPlan! - updatePlannerConfig(config: PlannerConfigInput!): PlannerConfig! -} -``` - -**Feasibility:** YES — async-graphql has native axum integration, same AppState pattern. ~200-300 lines. - -### 2c. gRPC — NOT STARTED - -**Plan:** Add `tonic` + `prost` to workspace, create `.proto` definitions. - -**Dependencies needed:** -```toml -# workspace Cargo.toml -tonic = "0.14" -prost = "0.14" -tonic-build = "0.14" # build dependency -``` - -**Files to create/modify:** -- `rust-core/verisim-api/proto/verisim.proto` — Service + message definitions -- `rust-core/verisim-api/build.rs` — tonic-build code generation -- `rust-core/verisim-api/src/grpc.rs` — Service implementation -- `rust-core/verisim-api/src/lib.rs` — Start gRPC server alongside HTTP - -**Proto Schema (planned):** -```protobuf -service VeriSimPlanner { - rpc OptimizePlan(LogicalPlan) returns (PhysicalPlan); - rpc ExplainPlan(LogicalPlan) returns (ExplainOutput); - rpc GetConfig(Empty) returns (PlannerConfig); - rpc SetConfig(PlannerConfig) returns (PlannerConfig); - rpc GetStats(Empty) returns (StatsSnapshot); -} - -service VeriSimOctad { - rpc Create(OctadRequest) returns (OctadResponse); - rpc Get(OctadId) returns (OctadResponse); - rpc Update(UpdateOctadRequest) returns (OctadResponse); - rpc Delete(OctadId) returns (Empty); - rpc SearchText(TextSearchRequest) returns (SearchResults); - rpc SearchVector(VectorSearchRequest) returns (SearchResults); -} -``` - -**Feasibility:** YES — tonic has mature axum co-hosting support. ~300-400 lines + proto. Runs on separate port (50051) or multiplexed. - ---- - -## Phase 3: VCL AST → LogicalPlan Bridge — NOT STARTED - -**Plan:** Add `rust-core/verisim-planner/src/vcl_bridge.rs` that deserializes the ReScript VCL AST JSON format into the planner's `LogicalPlan`. - -**Key mappings:** -- `AST.modality` → `Modality` (direct, except `All` → expand to 6) -- `AST.source` → `QuerySource` (Octad/Federation/Store) -- `AST.simpleCondition` variants → `ConditionKind` variants -- `AST.query.limit` → `PlanNode.early_limit` -- `AST.query.proof` → proof obligation nodes (Phase 4) - -**New endpoint:** `POST /query/vcl` — accepts VCL AST JSON, returns PhysicalPlan. - -**Feasibility:** YES — both sides use JSON serde. ~150-200 lines. - ---- - -## Phase 4: Proof Obligation Costing — NOT STARTED - -**Plan:** Add to verisim-planner: -- `src/proof.rs` — ProofObligation, ProofPlanNode, CompositionStrategy types -- Modify `cost.rs` — add proof cost estimation -- Modify `explain.rs` — include proof section in EXPLAIN output -- Modify `plan.rs` — add proof obligations to PhysicalPlan - -**Cost values (from VCLProofObligation.res):** -| Proof Type | Estimated Cost | -|-----------|---------------| -| Existence | 50ms | -| Citation | 100ms | -| Access | 150ms | -| Integrity | 200ms | -| Provenance | 300ms | -| Custom | 500ms | - -**Feasibility:** YES — pure type additions + cost arithmetic. ~200 lines. - ---- - -## Phase 5: Statistics Feedback + Adaptive Tuning — NOT STARTED - -**Plan:** -- Add `POST /planner/stats/record` endpoint (modality, latency_ms, rows_returned) -- Implement adaptive mode tuning in `config.rs`: - - Track last 50 queries per modality - - Compare estimated vs actual - - Shift mode if average error exceeds ±0.3 - -**Feasibility:** YES — StatisticsCollector already has `record_execution()`. Need endpoint + tuning logic. ~100-150 lines. - ---- - -## Phase 6: Cross-Modal + PostProcessing Costing — NOT STARTED - -**Plan:** -- Add `ConditionKind::CrossModalCompare`, `DriftCheck`, `ConsistencyCheck` -- Add PostProcessing cost estimation to optimizer (GroupBy=20ms, Sort=15ms, Aggregate=10ms) -- Handle `All` modality expansion - -**Feasibility:** YES — extending existing enums + adding cost branches. ~100 lines. - ---- - -## Dependency Graph - -``` -Phase 1 (DONE) ─┬─► Phase 2b (GraphQL) - ├─► Phase 2c (gRPC) - ├─► Phase 3 (VCL Bridge) ──► Phase 4 (Proof Costing) - ├─► Phase 5 (Statistics) - └─► Phase 6 (Cross-Modal) -``` - -All phases are independent except Phase 4 depends on Phase 3 (bridge must exist before proof obligations can flow through it). - ---- - -## Feasibility Summary - -| Phase | Feasible? | Effort | Risk | -|-------|-----------|--------|------| -| 1. Planner crate | YES — DONE | Done | None | -| 2b. GraphQL | YES | ~300 lines | Low — async-graphql is mature | -| 2c. gRPC | YES | ~400 lines + proto | Low — tonic is mature | -| 3. VCL Bridge | YES | ~200 lines | Low — JSON↔JSON mapping | -| 4. Proof Costing | YES | ~200 lines | Low — pure arithmetic | -| 5. Stats Feedback | YES | ~150 lines | Low — extending existing code | -| 6. Cross-Modal | YES | ~100 lines | Low — extending existing enums | - -**Total remaining:** ~1350 lines across 6 phases. All feasible, all low risk. - ---- - -## If Machine Crashes — Recovery Steps - -```bash -cd /var$REPOS_DIR/verisimdb - -# 1. Verify Phase 1 is intact -cargo build -p verisim-planner && cargo test -p verisim-planner -cargo build -p verisim-api && cargo test -p verisim-api - -# 2. Check git status -git status -git diff --stat - -# 3. If uncommitted, commit immediately: -git add rust-core/verisim-planner/ Cargo.toml rust-core/verisim-api/ -git commit -m "feat: add verisim-planner crate with cost-based query planning" - -# 4. Resume from whichever phase is next (check this file) -``` - -## Files That Must Not Be Lost - -These are the new files from this session: -``` -rust-core/verisim-planner/Cargo.toml -rust-core/verisim-planner/src/lib.rs -rust-core/verisim-planner/src/error.rs -rust-core/verisim-planner/src/plan.rs -rust-core/verisim-planner/src/cost.rs -rust-core/verisim-planner/src/stats.rs -rust-core/verisim-planner/src/config.rs -rust-core/verisim-planner/src/optimizer.rs -rust-core/verisim-planner/src/explain.rs -``` - -Modified files: -``` -Cargo.toml (workspace root — added verisim-planner member) -rust-core/verisim-api/Cargo.toml (added verisim-planner dep) -rust-core/verisim-api/src/lib.rs (added planner to AppState + 5 endpoints) -``` diff --git a/verisimdb/QUICKSTART-USER.adoc b/verisimdb/QUICKSTART-USER.adoc deleted file mode 100644 index ebefdc35..00000000 --- a/verisimdb/QUICKSTART-USER.adoc +++ /dev/null @@ -1,185 +0,0 @@ -// SPDX-License-Identifier: CC-BY-SA-4.0 -// Copyright (c) 2026 Jonathan D.A. Jewell (hyperpolymath) -// -// QUICKSTART-USER.adoc — Get VeriSimDB running from zero -= VeriSimDB Quickstart -:toc: macro -:icons: font - -toc::[] - -== What is VeriSimDB? - -VeriSimDB is an 8-modality database engine. Every entity is stored -simultaneously across Graph, Vector, Tensor, Semantic, Document, -Temporal, Provenance, and Spatial representations (the "octad"). -Drift between modalities is detected and self-healed automatically. - -The Rust core provides the modality stores. An Elixir/OTP layer -orchestrates coordination, fault tolerance, and the HTTP API. - -== Prerequisites - -You need these tools installed before building VeriSimDB. - -[cols="1,1,2"] -|=== -| Tool | Version | Install (Fedora) - -| Rust (nightly) -| 1.85+ -| `asdf install rust nightly` or `rustup install nightly` - -| Elixir -| 1.17+ -| `asdf install elixir 1.17.3-otp-27` - -| Erlang/OTP -| 27+ -| `asdf install erlang 27.2` - -| Zig -| 0.14+ -| `asdf install zig 0.14.0` - -| Idris2 _(optional, ABI layer)_ -| 0.7+ -| `asdf install idris2 0.7.0` - -| Podman _(optional, containers)_ -| 4+ -| `sudo dnf install podman podman-compose` - -| just -| 1.0+ -| `cargo install just` - -| openssl-devel -| -- -| `sudo dnf install openssl-devel` - -| pkg-config -| -- -| `sudo dnf install pkg-config` -|=== - -TIP: Run `just doctor` to verify all prerequisites are present. - -== Clone and Build - -[source,bash] ----- -cd ~/Documents/hyperpolymath-repos -# Already cloned if you have the nextgen-databases monorepo: -# cd nextgen-databases/verisimdb -# OR standalone: -git clone https://github.com/hyperpolymath/verisimdb -cd verisimdb ----- - -=== Build Rust Core (release) - -[source,bash] ----- -just build ----- - -This runs `cargo build --release` with `OPENSSL_NO_VENDOR=1`. - -=== Build Elixir Layer - -[source,bash] ----- -just build-elixir ----- - -Fetches Mix dependencies and compiles the OTP supervision tree. - -=== Build Everything - -[source,bash] ----- -just build-all ----- - -== First Run - -=== Option A: Rust API Server (standalone) - -[source,bash] ----- -just serve ----- - -Opens `http://localhost:8080`. The Rust API server provides direct -access to the modality stores. - -=== Option B: Elixir OTP Orchestrator - -[source,bash] ----- -just serve-otp ----- - -Starts the full OTP supervision tree with entity servers, drift -monitoring, and query routing. - -=== Option C: Container (Podman) - -[source,bash] ----- -just container-build -just container-run ----- - -Runs VeriSimDB in a Chainguard-based container on port 8080. - -== Verify It Works - -[source,bash] ----- -# Health check -curl http://localhost:8080/health - -# Create an octad entity (example) -curl -X POST http://localhost:8080/api/v1/entities \ - -H "Content-Type: application/json" \ - -d '{"document": {"title": "Test", "body": "Hello"}, "types": ["test"]}' ----- - -== Run Tests - -[source,bash] ----- -just test # Rust unit tests -just test-elixir # Elixir unit tests -just test-all # Both ----- - -== Key Recipes - -[source,bash] ----- -just # List all recipes -just build # Build Rust core (release) -just build-all # Build Rust + Elixir -just serve # Start API server (dev) -just test-all # Run all tests -just lint # Clippy lints -just fmt # Format Rust code -just doctor # Check prerequisites -just tour # Guided walkthrough -just help-me # What to do when stuck ----- - -== Instance Policy - -This repo is *source code and examples only*. If you are integrating -VeriSimDB into another project, you MUST create your own instance: - -1. Copy the client SDK from `connectors/clients//` -2. Add a Containerfile to YOUR project -3. Choose a unique port (not 8080) -4. Create a dedicated data volume - -See `.claude/CLAUDE.md` for port assignments per project. diff --git a/verisimdb/README.adoc b/verisimdb/README.adoc index f9fc9d83..08d9c5b0 100644 --- a/verisimdb/README.adoc +++ b/verisimdb/README.adoc @@ -1,485 +1,79 @@ // SPDX-License-Identifier: CC-BY-SA-4.0 +// SPDX-FileCopyrightText: 2026 Jonathan D.A. Jewell += VeriSimDB — extracted +:icons: font -= VeriSimDB -:toc: -:toc-placement!: +VeriSimDB no longer lives here. It is developed at: -**Cross-system data consistency that catches drift before it causes damage.** +https://github.com/hyperpolymath/verisimdb[hyperpolymath/verisimdb] -toc::[] +The 713 files that used to sit under this directory were removed on 2026-08-03, +after the work that existed *only* here was recovered upstream +(`hyperpolymath/verisimdb#219`, merged). -== The Problem: Silent Data Drift +== This was a fork, not a duplicate — and that mattered -Your data lives in multiple systems — graphs, vector stores, document indexes, time-series databases. When one system's view of an entity silently diverges from the others, you get **data drift**: stale embeddings, broken provenance chains, graph edges referencing deleted documents, spatial coordinates that no longer match textual descriptions. +`lithoglyph/` was removed by verifying every file was byte-identical to its own +repo first. **That check failed here.** Measured before anything was touched: -Traditional tools detect drift _after_ it causes downstream failures. VeriSimDB detects and repairs it _continuously_, before anyone notices. - -== What VeriSimDB Does - -VeriSimDB is a **cross-modal consistency engine**. Each entity exists simultaneously across up to 8 representations (the **octad**) with automatic drift detection and self-normalisation: - ----- -┌─────────────────────────────────────────────────────────────┐ -│ ONE ENTITY, EIGHT VIEWS │ -│ │ -│ Graph ─── Vector ─── Tensor ─── Semantic │ -│ │ │ │ -│ Document ─ Temporal ─ Provenance ─ Spatial │ -│ │ -│ ↕ Continuous drift detection ↕ │ -│ ↕ Automatic self-normalisation ↕ │ -└─────────────────────────────────────────────────────────────┘ ----- - -When VeriSimDB detects that an entity's vector embedding has diverged from its document content, or that a provenance chain's hash integrity is broken, or that spatial coordinates no longer match location references in the text — it identifies the drift type, scores the severity, and triggers repair automatically. - -== Repository Scope - -This repository is the plain upstream core for cloning and deployment. - -Consumer repositories should keep their own lightweight VeriSim runtime profiles (ports, launch policy, integration toggles) rather than vendoring this repository or sharing a single fixed local port. - -Canonical bot policy files for this repository live in `/.machine_readable/bot_directives/`. Root-level `/.bot_directives/` is treated as legacy. - -=== Drift Detection Demo - -[source,bash] ----- -cd elixir-orchestration && mix run ../demos/drift-detection/run_demo.exs ----- - -Sample output: - ----- -╔══════════════════════════════════════════════════════════════════╗ -║ DEMO RESULTS ║ -╠══════════════════════════════════════════════════════════════════╣ -║ Entities created: 1000 ║ -║ Entities corrupted: 50 ║ -║ ║ -║ DETECTION ║ -║ Detection rate: 100.0% ║ -║ ║ -║ REPAIR ║ -║ Repair rate: 100.0% ║ -║ ║ -║ VERIFICATION ║ -║ Consistency rate: 100.0% ║ -║ System health: healthy ║ -║ ║ -║ TIMING ║ -║ Total: ~2500ms ║ -╚══════════════════════════════════════════════════════════════════╝ ----- - -=== How It Compares - -[cols="1,1,1,1"] -|=== -|Capability |Great Expectations |Monte Carlo |VeriSimDB - -|Schema validation -|Yes -|Yes -|Yes - -|Drift detection -|Manual rules -|ML-based -|**Cross-modal, continuous** - -|Self-repair -|No -|No -|**Yes (automatic normalisation)** - -|Multi-representation -|No -|No -|**8 modalities per entity** - -|Provenance tracking -|No -|Partial -|**Hash-chain verified** - -|Query language -|No (Python API) -|No (UI) -|**VCL (with dependent types)** - -|Formal verification -|No -|No -|**VCL-UT (proof certificates)** +[cols="3,1"] |=== +| | Files -== The Octad: Eight Modalities - -Each entity in VeriSimDB can have representations across eight modalities. Drift detection operates across _all_ of them: - -[cols="1,2,1"] +| Only in this directory | **108** +| Only in `hyperpolymath/verisimdb` | 175 +| Common | 608 +| — of those, differing in content | **333** (331 substantive) |=== -|Modality |Purpose |Storage - -|**Graph** -|RDF triples and property graph edges -|Pure Rust (SimpleGraphStore) - -|**Vector** -|Embeddings for similarity search -|HNSW (in-memory) -|**Tensor** -|Multi-dimensional numeric data -|ndarray / Burn +Deleting on the lithoglyph precedent would have destroyed roughly 4,000 lines of +work, including an entire tier of test coverage. The transferable lesson is the +*check*, not the deletion. -|**Semantic** -|Type annotations and CBOR proof blobs -|CBOR (ciborium) +== What was recovered before this removal -|**Document** -|Full-text searchable content -|Tantivy (LZ4 compression) +Seven files had no counterpart upstream of any kind — not moved, not renamed, +not converted — and none was a stub. All are now on `verisimdb` `main`: -|**Temporal** -|Version history and time-series -|In-memory (chrono) - -|**Provenance** -|Origin tracking and transformation chain -|Hash-chain verified (SHA-256) - -|**Spatial** -|Geospatial coordinates and geometries -|R-tree index -|=== - -== Drift Types - -VeriSimDB detects eight categories of cross-modal drift: - -[cols="1,2,1"] +[cols="3,1,4"] |=== -|Drift Type |What It Detects |Default Threshold - -|**Semantic-Vector** -|Embedding diverged from semantic content -|0.3 - -|**Graph-Document** -|Graph structure doesn't match document -|0.4 - -|**Temporal Consistency** -|Version history gaps or conflicts -|0.2 - -|**Tensor** -|Tensor representation diverged -|0.35 - -|**Schema** -|Type constraint violations -|0.1 - -|**Provenance** -|Broken hash-chain integrity -|0.1 - -|**Spatial** -|Coordinates inconsistent with other modalities -|0.3 - -|**Quality** -|Overall weighted degradation -|0.25 +| File | Lines | What + +| `test/verisim/aspect/concurrency_test.exs` | 444 | Concurrency suite +| `test/verisim/aspect/security_test.exs` | 436 | Security suite +| `test/verisim/e2e_verisimdb_test.exs` | 410 | End-to-end suite +| `test/verisim/consensus/kraft_property_test.exs` | 328 | Kraft consensus property suite +| `rust-core/verisim-api/src/groove.rs` | 890 | Groove Protocol lifecycle + `/.well-known/groove` +| `rust-core/verisim-api/src/a2ml.rs` | 360 | A2ML response helpers (the no-JSON-emit rule) +| `rust-core/verisim-octad/src/ram_promotion.rs` | 472 | Octad RAM promotion |=== -== Architecture - -VeriSimDB runs as a two-layer system: Elixir/OTP for distributed coordination, Rust for performance-critical storage. - -=== Standalone Deployment +Plus `WHITEPAPER.pdf`. The four test suites run against upstream unmodified — +*5 properties, 43 tests, 0 failures* — and the Rust modules are wired into their +crates' `lib.rs` and compile (`cargo check` exit 0). ----- -┌─────────────────────────────────────────────────────────────┐ -│ Elixir Orchestration Layer │ -│ ├── VeriSim.EntityServer (GenServer per entity) │ -│ ├── VeriSim.DriftMonitor (drift detection coordinator) │ -│ ├── VeriSim.QueryRouter (distributes queries) │ -│ └── VeriSim.SchemaRegistry (type system coordinator) │ -│ ↓ HTTP │ -├─────────────────────────────────────────────────────────────┤ -│ Rust Modality Stores (All Local, Pure Rust) │ -│ ├── verisim-graph ├── verisim-semantic │ -│ ├── verisim-vector ├── verisim-document │ -│ ├── verisim-tensor ├── verisim-temporal │ -│ ├── verisim-provenance ├── verisim-spatial │ -│ ├── verisim-octad ├── verisim-drift │ -│ └── verisim-normalizer └── verisim-api │ -└─────────────────────────────────────────────────────────────┘ ----- - -=== Federated Deployment - -VeriSimDB can also operate as a **federation coordinator** over existing databases (ArangoDB, PostgreSQL, Elasticsearch, etc.): - ----- -┌────────────────────────────────────────────────────────────────┐ -│ ReScript Registry (Tiny Core) │ -│ ├── UUID → Store Mapping │ -│ ├── KRaft Metadata Log (Raft consensus) │ -│ └── Trust Window Management │ -│ ↓ Federation Protocol │ -├────────────────────────────────────────────────────────────────┤ -│ Federated Stores (Independent Operators) │ -│ ├── University Archive (Graph + Document) │ -│ ├── Research Lab (Vector + Tensor) │ -│ ├── Corporate DB (Semantic + Temporal) │ -│ └── Community Node (All modalities) │ -└────────────────────────────────────────────────────────────────┘ ----- - -== VCL: VeriSim Consonance Language +`storage_regenerator.rs` (682 lines) was deliberately **not** ported: upstream +already has `verisim-normalizer/src/regeneration.rs` and +`verisim-api/src/regenerator.rs`. It was superseded, not lost. -VCL provides unified querying across all eight modalities with two execution paths: +The remaining 100 files were triaged as moved (16), converted (26), or +superseded prose, config and retired ReScript. Full working: +`docs/migration/` and the estate notes named below. -=== Slipstream (Fast, No Proofs) +== Nothing is lost -[source,sql] ----- --- Query across modalities -SELECT GRAPH.*, DOCUMENT.*, VECTOR.* FROM HEXAD 'entity-001' - --- Cross-modal drift detection in WHERE clause -SELECT * FROM HEXAD 'entity-001' - WHERE DRIFT(VECTOR, DOCUMENT) > 0.3 - --- Modality existence checks -SELECT * FROM HEXAD 'entity-001' - WHERE PROVENANCE EXISTS AND TENSOR NOT EXISTS - --- Federation queries -SELECT * FROM FEDERATION /* WITH DRIFT STRICT ----- - -=== VCL-UT (Dependent Types with Proof Certificates) +Every file removed here is preserved at its original path in the annotated tag +`split-history/verisimdb`, which is on `origin` and holds all 713 files across +124 commits: -[source,sql] +[source,console] ---- --- Query with existence proof -SELECT GRAPH.* FROM HEXAD 'entity-001' - PROOF EXISTENCE(entity-001) - --- Multi-proof composition -SELECT * FROM HEXAD 'entity-001' - PROOF EXISTENCE(entity-001) AND PROVENANCE(entity-001) - --- Integrity verification with contract -SELECT GRAPH.* FROM HEXAD 'entity-001' - PROOF INTEGRITY(my_contract) +$ git ls-tree -r split-history/verisimdb --name-only | wc -l +713 +$ git checkout split-history/verisimdb -- ---- -VCL-UT queries return a `ProvedResult` with data AND a proof certificate: - -[source,elixir] ----- -%{ - data: [...], - proof_certificate: %{ - proofs: [%{type: :existence, verified: true, ...}], - obligations: [%{type: :existence, contract: "entity-001"}], - composition: :conjunction, - verified_at: ~U[2026-02-28 12:00:00Z], - query_hash: "sha256:..." - } -} ----- - -== Quick Start - -=== Prerequisites - -* Rust (edition 2021) — **no C++ linker required** (pure Rust build) -* Elixir 1.17+ - -=== Build - -[source,bash] ----- -# Rust core (pure Rust — no C/C++ dependencies) -cargo build - -# Elixir orchestration -cd elixir-orchestration -mix deps.get -mix compile ----- - -=== Run - -[source,bash] ----- -# Start Rust API server -cargo run -p verisim-api - -# Start Elixir orchestration (in another terminal) -cd elixir-orchestration -iex -S mix ----- - -=== Test - -[source,bash] ----- -cargo test # Rust: 510+ tests -cd elixir-orchestration -mix test # Elixir: 160+ tests (VCL, consensus, telemetry, federation, hypatia) ----- - -=== Container - -[source,bash] ----- -podman build -t verisimdb:latest -f container/Containerfile . -podman run -p 8080:8080 verisimdb:latest ----- - -== API Examples - -=== Create a Octad - -[source,bash] ----- -curl -X POST http://localhost:8080/api/v1/octads \ - -H "Content-Type: application/json" \ - -d '{ - "title": "Machine Learning Paper", - "body": "An introduction to neural networks...", - "embedding": [0.1, 0.2, 0.3], - "types": ["http://example.org/Paper"], - "relationships": [["cites", "paper-123"]], - "provenance": { - "event_type": "created", - "actor": "researcher@university.edu", - "source": "https://arxiv.org/abs/2024.12345" - }, - "spatial": { - "latitude": 51.5074, - "longitude": -0.1278, - "geometry_type": "Point" - } - }' ----- - -=== Check Drift Status - -[source,bash] ----- -curl http://localhost:8080/api/v1/drift/status ----- - -=== Trigger Normalisation - -[source,bash] ----- -curl -X POST http://localhost:8080/api/v1/normalizer/trigger/entity-001 ----- - -=== gRPC (port 50051) - -VeriSimDB's gRPC server runs directly on the Rust core (via tonic) on port 50051. The V API gateway does not proxy gRPC traffic — connect directly to the Rust core. - -[source,bash] ----- -# List available gRPC services -grpcurl -plaintext localhost:50051 list - -# Health check -grpcurl -plaintext localhost:50051 verisimdb.VeriSimService/Health - -# Execute a VCL query -grpcurl -plaintext -d '{"query": "SELECT * FROM octads LIMIT 10"}' \ - localhost:50051 verisimdb.VeriSimService/Query - -# Get a specific octad -grpcurl -plaintext -d '{"id": "entity-001"}' \ - localhost:50051 verisimdb.VeriSimService/GetOctad ----- - -== Project Structure - ----- -verisimdb/ -├── rust-core/ # Rust modality stores (pure Rust, no C++) -│ ├── verisim-graph/ # Graph (SimpleGraphStore default, Oxigraph optional) -│ ├── verisim-vector/ # Vector (HNSW similarity search) -│ ├── verisim-tensor/ # Tensor (ndarray/Burn) -│ ├── verisim-semantic/ # Semantic (CBOR proof blobs) -│ ├── verisim-document/ # Document (Tantivy full-text search) -│ ├── verisim-temporal/ # Temporal (version history) -│ ├── verisim-provenance/ # Provenance (hash-chain lineage) -│ ├── verisim-spatial/ # Spatial (R-tree geospatial) -│ ├── verisim-octad/ # Unified octad entity -│ ├── verisim-drift/ # Drift detection -│ ├── verisim-normalizer/ # Self-normalisation -│ ├── verisim-wal/ # Write-ahead log -│ └── verisim-api/ # HTTP/gRPC/GraphQL API -├── elixir-orchestration/ # Elixir/OTP coordination layer -│ ├── lib/verisim/ -│ │ ├── drift/ # DriftMonitor GenServer -│ │ ├── entity/ # EntityServer (per-entity GenServer) -│ │ ├── query/ # VCL parser, executor, bridge -│ │ └── rust_client.ex # HTTP client for Rust core -│ └── test/ -├── demos/ -│ └── drift-detection/ # Drift detection demo script -├── docs/ # Specifications and design documents -│ ├── vcl-grammar.ebnf # VCL formal grammar -│ ├── vcl-formal-semantics.adoc -│ └── vcl-type-system.adoc -├── registry/ # ReScript federation registry -└── .machine_readable/ # STATE.scm, META.scm, ECOSYSTEM.scm ----- - -=== Client SDK Notes - -VeriSimDB ships six client SDKs in `connectors/clients/` (Rust, V, Elixir, ReScript, Julia, Gleam). Consumer projects should copy the relevant SDK into their own codebase rather than depending on this repo directly. - -**Burble** uses the **Elixir client** (`connectors/clients/elixir/`), which runs natively alongside Burble's Elixir control plane. The Gleam client is available for projects targeting the BEAM via Gleam, but Burble does not use it. - -== Use Cases - -=== Data Quality at Scale -Monitor cross-modal consistency across thousands of entities. When a vector embedding drifts from its source document, or a provenance chain breaks, VeriSimDB detects and repairs it automatically — before downstream ML models or analytics pipelines consume stale data. - -=== Open Science Federation -Universities and research institutions federate their archives while maintaining local control. Drift detection ensures retractions propagate correctly across institutional boundaries. - -=== Neurosymbolic AI Pipelines -Vector embeddings and symbolic graphs coexist in the same namespace. AI systems query both modalities without ETL, with formal proof certificates guaranteeing query correctness. - -=== Verifiable Data Exchange -ZKP proofs and dependent types enable tamper-evident knowledge exchange. VCL-UT queries return data with cryptographic proof certificates — you can verify the result without trusting the source. - -== Documentation - -* link:docs/vcl-grammar.ebnf[VCL Grammar] — Formal EBNF specification -* link:docs/vcl-formal-semantics.adoc[VCL Formal Semantics] — Operational semantics -* link:docs/vcl-type-system.adoc[VCL Type System] — Dependent types and bidirectional type checking -* link:docs/drift-handling.adoc[Drift Handling] — Detection, repair, federation drift -* link:docs/safety-and-fault-tolerance.adoc[Safety & Fault Tolerance] — Memory, kernel, platform, supply chain safety -* link:docs/deployment-modes.adoc[Deployment Modes] — Standalone, federated, hybrid -* link:docs/getting-started.adoc[Getting Started] — Step-by-step setup guide -* link:docs/adoption-strategy.adoc[Adoption Strategy] — Target domains and rollout plan -* link:docs/VCL-SPEC.adoc[VCL Specification] — 2785-line normative language spec -* link:docs/papers/[White Papers] — Academic and industry papers - -== License - -PMPL-1.0-or-later - -== Contributing - -See CONTRIBUTING.adoc for guidelines. +**Never prune** `split-history/verisimdb` or the `_split_verisimdb` branch, nor +their `lithoglyph`, `glyphbase` and `gnpl` equivalents. The imports into the +standalone repos were squash merges, so these refs are the only granular history +that exists. diff --git a/verisimdb/README.adoc.invariants.md b/verisimdb/README.adoc.invariants.md deleted file mode 100644 index e9ccff7e..00000000 --- a/verisimdb/README.adoc.invariants.md +++ /dev/null @@ -1,2 +0,0 @@ -# Invariant Path Scan: README.adoc - diff --git a/verisimdb/RELEASE-NOTES-v0.1.0-alpha.md b/verisimdb/RELEASE-NOTES-v0.1.0-alpha.md deleted file mode 100644 index 33cbf0b1..00000000 --- a/verisimdb/RELEASE-NOTES-v0.1.0-alpha.md +++ /dev/null @@ -1,389 +0,0 @@ -# VeriSimDB v0.1.0-alpha Release Notes - -**Release Date:** 2026-02-04 - -**Status:** Alpha Release - Production-Ready - ---- - -## 🎉 Overview - -VeriSimDB v0.1.0-alpha is the first public release of the Veridical Simulacrum Database - a multimodal database with self-normalization capabilities. This release marks the completion of all core functionality and readiness for production deployment. - -VeriSimDB operates as both a standalone database (like PostgreSQL) and a federated coordinator for distributed knowledge networks. - ---- - -## ✨ Key Features - -### 🔍 VCL Query Language (100% Complete) -- **VCLParser.res** - Full parser with combinator library (633 lines) -- **VCLError.res** - Comprehensive error types for all failure modes (14K+ lines) -- **VCLExplain.res** - Query execution plan visualization (7K+ lines) -- **VCLTypeChecker.res** - Dependent-type verification with ZKP integration -- **VCLExecutor** - Bridges ReScript parser to Elixir orchestration -- ISO/IEC 14977 EBNF compliant grammar -- Two execution paths: Slipstream (fast) and Dependent-type (verified) - -### 🗄️ Six Modalities (80% Complete) -Every octad (entity) exists simultaneously across six synchronized representations: - -1. **Graph** (244 LOC) - RDF triples and property graphs via Oxigraph -2. **Vector** (248+ LOC) - HNSW similarity search for embeddings -3. **Tensor** (278 LOC) - Multi-dimensional numeric data with ndarray/Burn -4. **Semantic** (345 LOC) - Type annotations and CBOR proof blobs -5. **Document** (Complete) - Full-text search powered by Tantivy -6. **Temporal** (377+ LOC) - Version history and time-travel queries - -**Additional Components:** -- **Octad Store** (400+ LOC) - Unified entity management -- **Drift Detection** (484+ LOC) - Cross-modal consistency monitoring -- **Normalizer** (406 LOC) - Self-normalization when drift exceeds thresholds -- **HTTP API** (782 LOC) - RESTful API server with Axum - -### 🔄 Elixir/OTP Orchestration (100% Complete) -- **QueryRouter** - Distributes queries across modality stores -- **EntityServer** - GenServer-per-entity model for fault tolerance -- **DriftMonitor** - Coordinates drift detection and normalization -- **SchemaRegistry** - Type system management with constraint validation -- **RustClient** - HTTP client for Rust core communication -- Full supervision tree with OTP fault tolerance - -### 🌐 Federation Registry (100% Complete) -The "tiny core" (<5K LOC) enabling federated deployments: - -- **Registry.res** (400 LOC) - UUID → store location mapping - - Store health tracking with trust scores - - Pattern-based federation queries (`/universities/*`) - - Byzantine fault detection - - Replication management -- **MetadataLog.res** (500 LOC) - KRaft-inspired Raft consensus - - Leader election - - Log replication - - Commit index management - - Term-based consistency - -### 🧪 Integration Tests (100% Complete) -Comprehensive test coverage: - -**Rust Tests (12 tests):** -- Octad CRUD operations -- Cross-modal consistency -- Drift detection -- Vector similarity search -- Fulltext search -- Temporal versioning -- Graph relationships -- Normalization -- Multi-modal queries -- Concurrent operations - -**Elixir Tests (Full stack):** -- RustClient integration -- QueryRouter -- DriftMonitor -- SchemaRegistry -- EntityServer -- VCLExecutor -- End-to-end integration - -### ⚡ Performance Benchmarks (100% Complete) -Criterion-based benchmarks for: -- Document store: create, search (1K docs) -- Vector store: insert, similarity search across 128/384/768 dimensions (10K vectors) -- Graph store: node/edge operations -- Octad operations: create, retrieve -- Drift detection: calculation performance -- Cross-modal queries: combined vector + fulltext - -Run with: `cd benches && cargo bench` - -### 📚 Production Deployment Guide (100% Complete) -Comprehensive 100+ section guide covering: -- Three deployment modes: Standalone, Federated, Hybrid -- Hardware/software requirements -- Complete deployment steps with Podman -- Security: TLS, authentication, RBAC, encryption -- Monitoring: Prometheus metrics, logging, alerting -- Backup & recovery procedures -- Performance tuning -- Troubleshooting -- Operational procedures -- Production checklists - ---- - -## 🏗️ Architecture - -### Standalone Mode -``` -┌─────────────────────────────────────────┐ -│ Elixir Orchestration (Port 4000) │ -│ HTTP API + WebSocket │ -├─────────────────────────────────────────┤ -│ Rust Core (Port 8080) │ -│ verisim-api HTTP Server │ -├─────────────────────────────────────────┤ -│ Local Modality Stores │ -│ ├── Graph (Oxigraph) │ -│ ├── Vector (HNSW) │ -│ ├── Tensor (ndarray) │ -│ ├── Semantic (CBOR) │ -│ ├── Document (Tantivy) │ -│ └── Temporal (Version tree) │ -└─────────────────────────────────────────┘ -``` - -### Federated Mode -``` -┌─────────────────────────────────────────┐ -│ ReScript Registry (Port 3000) │ -│ UUID → Store Mapping + Raft │ -├─────────────────────────────────────────┤ -│ Elixir Orchestration (Port 4000) │ -│ Federation Query Router │ -├─────────────────────────────────────────┤ -│ Remote Stores (Distributed) │ -│ ├── University A (Graph + Document) │ -│ ├── Research Lab B (Vector + Tensor) │ -│ └── Company C (Semantic + Temporal) │ -└─────────────────────────────────────────┘ -``` - ---- - -## 📊 Metrics - -### Code Statistics -- **Total Lines Written:** ~3,839 lines in final session -- **VCL Implementation:** ~22K lines -- **Rust Core:** ~3,860 lines -- **Elixir Orchestration:** ~2,000+ lines -- **ReScript Registry:** ~900 lines -- **Tests:** ~764 lines -- **Benchmarks:** ~500 lines -- **Documentation:** ~1,000+ lines - -### Project Completion -- Overall: **100%** ✅ -- VCL: **100%** ✅ -- Elixir: **100%** ✅ -- Rust Stores: **80%** 🟡 -- Registry: **100%** ✅ -- Tests: **100%** ✅ -- Benchmarks: **100%** ✅ -- Docs: **100%** ✅ - ---- - -## 🚀 Getting Started - -### Installation - -**Prerequisites:** -- Rust 1.75+ -- Elixir 1.16+ with Erlang/OTP 26+ -- Podman 4.0+ or Docker 24.0+ - -**Quick Start (Standalone):** - -```bash -# Clone repository -git clone https://github.com/hyperpolymath/verisimdb -cd verisimdb - -# Build Rust core -cargo build --release --all-features - -# Build Elixir orchestration -cd elixir-orchestration -mix deps.get -MIX_ENV=prod mix release - -# Run with Podman -cd .. -podman build -t verisimdb:v0.1.0-alpha -f container/Containerfile . -podman run -d \ - --name verisimdb \ - -p 8080:8080 \ - -p 4000:4000 \ - -v verisimdb-data:/var/lib/verisimdb:Z \ - verisimdb:v0.1.0-alpha -``` - -**Health Check:** -```bash -curl http://localhost:8080/api/v1/health -``` - -### Example Usage - -**Create a Octad:** -```bash -curl -X POST http://localhost:8080/api/v1/octads \ - -H "Content-Type: application/json" \ - -d '{ - "title": "Research Paper", - "body": "Introduction to multimodal databases...", - "embedding": [0.1, 0.2, ...], - "types": ["http://example.org/Document"] - }' -``` - -**Query with VCL:** -```bash -curl -X POST http://localhost:4000/api/v1/query \ - -H "Content-Type: application/json" \ - -d '{ - "query": "SELECT GRAPH, VECTOR FROM OCTAD abc-123 WHERE FULLTEXT CONTAINS \"machine learning\" LIMIT 10" - }' -``` - ---- - -## 🔧 Configuration - -### Environment Variables -- `VERISIM_MODE` - Deployment mode: standalone, federation, hybrid (default: standalone) -- `VERISIM_DATA_DIR` - Data directory (default: /var/lib/verisimdb) -- `VERISIM_LOG_LEVEL` - Log level: debug, info, warn, error (default: info) -- `VERISIM_HTTP_PORT` - Rust API port (default: 8080) -- `VERISIM_ELIXIR_PORT` - Elixir port (default: 4000) -- `VERISIM_DRIFT_THRESHOLD` - Drift threshold 0.0-1.0 (default: 0.7) -- `VERISIM_AUTO_NORMALIZE` - Enable auto-normalization (default: true) - -See `DEPLOYMENT.adoc` for comprehensive configuration options. - ---- - -## 📈 Performance - -Expected performance characteristics (preliminary benchmarks): - -| Operation | Throughput | Latency (p95) | -|-----------|------------|---------------| -| Document Index | ~1K docs/sec | <50ms | -| Vector Search (10K vectors) | ~500 queries/sec | <100ms | -| Octad Create | ~200 ops/sec | <200ms | -| Fulltext Search | ~1K queries/sec | <50ms | -| Cross-modal Query | ~100 queries/sec | <500ms | - -*Note: Performance varies based on hardware, configuration, and data size* - -Run full benchmarks: `cd benches && cargo bench` - ---- - -## 🔒 Security - -### Features -- TLS 1.3 for all network communication -- API key authentication -- Role-based access control (RBAC) -- Encryption at rest (LUKS) -- Audit logging -- Byzantine fault tolerance in federation mode - -### Known Limitations -- Default API keys must be changed in production -- Self-signed certificates for development only -- ZKP proof verification is stubbed (full implementation pending) - ---- - -## 🐛 Known Issues - -### Minor Issues -1. **RUSTSEC-2026-0002** - lru 0.12.5 (transitive dependency, LOW severity) - - IterMut Stacked Borrows issue - - Waiting for upstream fix in tantivy/ratatui - - Does not affect VeriSimDB functionality - -### Limitations -1. **Rust Modality Stores** - 80% complete - - Core functionality implemented - - Some advanced features pending (compression, partitioning) -2. **ZKP Integration** - Proof generation/verification stubbed - - Architecture and types complete - - Full cryptographic implementation pending proven library integration -3. **Federation** - Registry complete, needs production testing - - KRaft consensus implemented - - Requires multi-node deployment testing - ---- - -## 🗺️ Roadmap - -### v0.2.0 (Q2 2026) -- Complete Rust store implementations (→ 100%) -- Full ZKP proof generation/verification -- Performance optimizations -- Federation production testing -- Additional VCL features (aggregations, joins) - -### v0.3.0 (Q3 2026) -- Horizontal scaling enhancements -- Advanced drift detection strategies -- Query optimizer -- Compression and partitioning -- Plugin system - -### v1.0.0 (Q4 2026) -- Production-hardened -- Full feature set -- Comprehensive benchmarks -- Migration tools -- Enterprise support - ---- - -## 📄 License - -VeriSimDB is licensed under **PMPL-1.0-or-later** (Palimpsest License). - -Third-party components retain their original licenses (MIT, Apache, BSD, etc.). - -See `LICENSE` for full terms. - ---- - -## 🙏 Acknowledgments - -**Development:** -- Jonathan D.A. Jewell (j.d.a.jewell@open.ac.uk) - -**AI Assistance:** -- Claude Sonnet 4.5 (Anthropic) - -**Open Source Libraries:** -- Oxigraph (RDF/SPARQL) -- Tantivy (Full-text search) -- Axum (HTTP server) -- Elixir/OTP (Orchestration) -- And many others (see dependencies) - ---- - -## 📞 Support & Community - -- **Documentation:** https://verisimdb.hyperpolymath.org/docs -- **Repository:** https://github.com/hyperpolymath/verisimdb -- **Issues:** https://github.com/hyperpolymath/verisimdb/issues -- **Discussions:** https://github.com/hyperpolymath/verisimdb/discussions -- **Security:** security@hyperpolymath.org - ---- - -## 🎯 Next Steps - -1. **Deploy:** Follow `DEPLOYMENT.adoc` for your environment -2. **Test:** Run integration tests and benchmarks -3. **Experiment:** Create octads and run VCL queries -4. **Feedback:** Report issues and contribute -5. **Community:** Join discussions and share use cases - ---- - -**VeriSimDB v0.1.0-alpha** - Bridging the Map and the Territory - -*Released with ❤️ by the hyperpolymath project* diff --git a/verisimdb/ROADMAP.adoc b/verisimdb/ROADMAP.adoc deleted file mode 100644 index b4bf21f1..00000000 --- a/verisimdb/ROADMAP.adoc +++ /dev/null @@ -1,69 +0,0 @@ -// SPDX-License-Identifier: CC-BY-SA-4.0 -= Verisimdb Roadmap - -== Current Status - -Initial development phase. - -== Milestones - -=== v0.1.0 - Foundation -* [ ] Core functionality -* [ ] Basic documentation -* [ ] CI/CD pipeline - -=== v1.0.0 - Stable Release -* [ ] Full feature set -* [ ] Comprehensive tests -* [ ] Production ready - -== Future Directions - -_To be determined based on community feedback._ - -== Enhancements Identified (2026-02-08) - -=== Immediate Fixes - -* Fix tensor reduction operations (max/min/prod use sum_axis - TODO in code) -* Fix unused variable warnings (verisim-normalizer) -* Implement document search highlighting (TODO in code) -* Add binary target to verisim-api for `cargo run` (currently fails with "a bin target must be available") - -=== Integration - -* panic-attack result schema: define a octad schema for storing static analysis results -** Document modality: full JSON report as searchable text -** Graph modality: file -> weakness -> attack-recommendation triples -** Temporal modality: track results over time -** Vector modality: embed weakness descriptions for similarity search -* MCP server mode: expose verisimdb as an MCP tool - -=== Storage - -* Default to persistent storage backends (RocksDB for graph, file-based for document) -* Add configuration file for storage paths and options -* Implement proper data directory management - -=== Query Language (VCL) - -* Complete the ReScript parser integration with the Rust API -* Add VCL endpoint to HTTP API -* VCL REPL for interactive querying - -=== Observability - -* Prometheus metrics endpoint -* Structured logging with tracing -* Health dashboard - -== OPSM Integration - -[source] ----- -OPSM Core - | - v -verisimdb (storage backend candidate for OPSM metadata and audits) - ----- diff --git a/verisimdb/RSR_OUTLINE.adoc b/verisimdb/RSR_OUTLINE.adoc deleted file mode 100644 index 9f323746..00000000 --- a/verisimdb/RSR_OUTLINE.adoc +++ /dev/null @@ -1,218 +0,0 @@ -= RSR Template Repository - -image:[Palimpsest-MPL-1.0,link="https://github.com/hyperpolymath/palimpsest-license"] image:[Palimpsest,link="https://github.com/hyperpolymath/palimpsest-license"] -:toc: -:sectnums: - -// Badges -image:https://img.shields.io/badge/RSR-Infrastructure-cd7f32[RSR Infrastructure] -image:https://img.shields.io/badge/Phase-Maintenance-brightgreen[Phase] -image:https://img.shields.io/badge/Guix-Primary-purple?logo=gnu[Guix] - -== Overview - -**The canonical template for RSR (Rhodium Standard Repository) projects.** - -This repository provides the standardized structure, configuration, and tooling for all 139 repos in the hyperpolymath ecosystem. Use it to: - -* Bootstrap new projects with RSR compliance -* Reference the standard directory structure -* Copy configuration templates (Justfile, STATE.scm, etc.) - -== Quick Start - -[source,bash] ----- -# Clone the template -git clone https://github.com/hyperpolymath/RSR-template-repo my-project -cd my-project - -# Remove template git history -rm -rf .git -git init - -# Customize -sed -i 's/RSR-template-repo/my-project/g' Justfile guix.scm README.adoc - -# Enter development environment -guix shell -D -f guix.scm - -# Validate compliance -just validate-rsr ----- - -== What's Included - -[cols="1,3"] -|=== -|File/Directory |Purpose - -|`.editorconfig` -|Editor configuration (indent, charset) - -|`.gitignore` -|Standard ignore patterns - -|`.guix-channel` -|Guix channel definition - -|`.well-known/` -|RFC-compliant metadata (security.txt, ai.txt, humans.txt) - -|`docs/` -|Documentation directory - -|`guix.scm` -|Guix package definition - -|`justfile` -|Task runner with 50+ recipes - -|`LICENSE.txt` -|PMPL-1.0-or-later (Palimpsest License) - -|`README.adoc` -|This file - -|`RSR_COMPLIANCE.adoc` -|Compliance tracking - -|`STATE.scm` -|Project state checkpoint -|=== - -== Justfile Features - -The template Justfile provides: - -* **~10 billion recipe combinations** via matrix recipes -* **Cookbook generation**: `just cookbook` → `docs/just-cookbook.adoc` -* **Man page generation**: `just man` → `docs/man/project.1` -* **RSR validation**: `just validate-rsr` -* **STATE.scm management**: `just state-touch`, `just state-phase` -* **Container support**: `just container-build`, `just container-push` -* **CI matrix**: `just ci-matrix [stage] [depth]` - -=== Key Recipes - -[source,bash] ----- -just # Show all recipes -just help # Detailed help -just info # Project info -just combinations # Show matrix options - -just build # Build (debug) -just test # Run tests -just quality # Format + lint + test -just ci # Full CI pipeline - -just validate # RSR + STATE validation -just docs # Generate all docs -just cookbook # Generate Justfile docs - -just guix-shell # Guix dev environment -just container-build # Build container ----- - -== Directory Structure - -[source] ----- -project/ -├── .editorconfig # Editor settings -├── .gitignore # Git ignore -├── .guix-channel # Guix channel -├── .well-known/ # RFC metadata -│ ├── ai.txt -│ ├── humans.txt -│ └── security.txt -├── config/ # Nickel configs (optional) -├── docs/ # Documentation -│ ├── generated/ -│ ├── man/ -│ └── just-cookbook.adoc -├── guix.scm # Guix package -├── Justfile # Task runner -├── LICENSE.txt # Dual license -├── README.adoc # Overview -├── RSR_COMPLIANCE.adoc # Compliance -├── src/ # Source code -├── STATE.scm # State checkpoint -└── tests/ # Tests ----- - -== RSR Compliance - -=== Language Tiers - -* **Tier 1** (Gold): Rust, Elixir, Zig, Ada, Haskell, ReScript -* **Tier 2** (Silver): Nickel, Racket, Guile Scheme, Nix -* **Infrastructure**: Guix channels, derivations - -=== Required Files - -* `.editorconfig` -* `.gitignore` -* `justfile` -* `README.adoc` -* `RSR_COMPLIANCE.adoc` -* `LICENSE.txt` (PMPL-1.0-or-later) -* `.well-known/security.txt` -* `.well-known/ai.txt` -* `.well-known/humans.txt` -* `guix.scm` OR `flake.nix` - -=== Prohibited - -* Python outside `salt/` directory -* TypeScript/JavaScript (use ReScript) -* CUE (use Guile/Nickel) -* `Dockerfile` (use `Containerfile`) - -== STATE.scm - -The STATE.scm file tracks project state: - -[source,scheme] ----- -(define state - `((metadata - (project . "my-project") - (updated . "2025-12-10")) - (position - (phase . implementation) ; design|implementation|testing|maintenance|archived - (maturity . beta)) ; experimental|alpha|beta|production|lts - (ecosystem - (part-of . ("RSR Framework")) - (depends-on . ())))) ----- - -== Badge Schema - -Generate badges from STATE.scm: - -[source,bash] ----- -just badges standard ----- - -See `docs/BADGE_SCHEMA.adoc` for the full badge taxonomy. - -== Ecosystem Integration - -This template is part of: - -* **STATE.scm Ecosystem**: Conversation checkpoints -* **RSR Framework**: Repository standards -* **Consent-Aware-HTTP**: .well-known compliance - -== License - -SPDX-License-Identifier: CC-BY-SA-4.0 - -== Links - -* https://github.com/hyperpolymath/elegant-STATE[elegant-STATE] - STATE.scm tooling -* https://github.com/hyperpolymath/conative-gating[conative-gating] - Policy enforcement -* https://rhodium.sh[Rhodium Standard] - RSR documentation diff --git a/verisimdb/SECURITY.md b/verisimdb/SECURITY.md deleted file mode 100644 index fe81dbac..00000000 --- a/verisimdb/SECURITY.md +++ /dev/null @@ -1,375 +0,0 @@ -# Security Policy - - -We take security seriously. We appreciate your efforts to responsibly disclose vulnerabilities and will make every effort to acknowledge your contributions. - -## Table of Contents - -- [Reporting a Vulnerability](#reporting-a-vulnerability) -- [What to Include](#what-to-include) -- [Response Timeline](#response-timeline) -- [Disclosure Policy](#disclosure-policy) -- [Scope](#scope) -- [Safe Harbour](#safe-harbour) -- [Recognition](#recognition) -- [Security Updates](#security-updates) -- [Security Best Practices](#security-best-practices) - ---- - -## Reporting a Vulnerability - -### Preferred Method: GitHub Security Advisories - -The preferred method for reporting security vulnerabilities is through GitHub's Security Advisory feature: - -1. Navigate to [Report a Vulnerability](https://github.com/hyperpolymath/verisimdb/security/advisories/new) -2. Click **"Report a vulnerability"** -3. Complete the form with as much detail as possible -4. Submit — we'll receive a private notification - -This method ensures: - -- End-to-end encryption of your report -- Private discussion space for collaboration -- Coordinated disclosure tooling -- Automatic credit when the advisory is published - -### Alternative: Email - -If you cannot use GitHub Security Advisories, you may email us directly: - -| | | -|---|---| -| **Email** | j.d.a.jewell@open.ac.uk | - -> **⚠️ Important:** Do not report security vulnerabilities through public GitHub issues, pull requests, discussions, or social media. - ---- - -## What to Include - -A good vulnerability report helps us understand and reproduce the issue quickly. - -### Required Information - -- **Description**: Clear explanation of the vulnerability -- **Impact**: What an attacker could achieve (confidentiality, integrity, availability) -- **Affected versions**: Which versions/commits are affected -- **Reproduction steps**: Detailed steps to reproduce the issue - -### Helpful Additional Information - -- **Proof of concept**: Code, scripts, or screenshots demonstrating the vulnerability -- **Attack scenario**: Realistic attack scenario showing exploitability -- **CVSS score**: Your assessment of severity (use [CVSS 3.1 Calculator](https://www.first.org/cvss/calculator/3.1)) -- **CWE ID**: Common Weakness Enumeration identifier if known -- **Suggested fix**: If you have ideas for remediation -- **References**: Links to related vulnerabilities, research, or advisories - -### Example Report Structure - -```markdown -## Summary -[One-sentence description of the vulnerability] - -## Vulnerability Type -[e.g., SQL Injection, XSS, SSRF, Path Traversal, etc.] - -## Affected Component -[File path, function name, API endpoint, etc.] - -## Affected Versions -[Version range or specific commits] - -## Severity Assessment -- CVSS 3.1 Score: [X.X] -- CVSS Vector: [CVSS:3.1/AV:X/AC:X/PR:X/UI:X/S:X/C:X/I:X/A:X] - -## Description -[Detailed technical description] - -## Steps to Reproduce -1. [First step] -2. [Second step] -3. [...] - -## Proof of Concept -[Code, curl commands, screenshots, etc.] - -## Impact -[What can an attacker achieve?] - -## Suggested Remediation -[Optional: your ideas for fixing] - -## References -[Links to related issues, CVEs, research] -``` - ---- - -## Response Timeline - -We commit to the following response times: - -| Stage | Timeframe | Description | -|-------|-----------|-------------| -| **Initial Response** | 48 hours | We acknowledge receipt and confirm we're investigating | -| **Triage** | 7 days | We assess severity, confirm the vulnerability, and estimate timeline | -| **Status Update** | Every 7 days | Regular updates on remediation progress | -| **Resolution** | 90 days | Target for fix development and release (complex issues may take longer) | -| **Disclosure** | 90 days | Public disclosure after fix is available (coordinated with you) | - -> **Note:** These are targets, not guarantees. Complex vulnerabilities may require more time. We'll communicate openly about any delays. - ---- - -## Disclosure Policy - -We follow **coordinated disclosure** (also known as responsible disclosure): - -1. **You report** the vulnerability privately -2. **We acknowledge** and begin investigation -3. **We develop** a fix and prepare a release -4. **We coordinate** disclosure timing with you -5. **We publish** security advisory and fix simultaneously -6. **You may publish** your research after disclosure - -### Our Commitments - -- We will not take legal action against researchers who follow this policy -- We will work with you to understand and resolve the issue -- We will credit you in the security advisory (unless you prefer anonymity) -- We will notify you before public disclosure -- We will publish advisories with sufficient detail for users to assess risk - -### Your Commitments - -- Report vulnerabilities promptly after discovery -- Give us reasonable time to address the issue before disclosure -- Do not access, modify, or delete data beyond what's necessary to demonstrate the vulnerability -- Do not degrade service availability (no DoS testing on production) -- Do not share vulnerability details with others until coordinated disclosure - -### Disclosure Timeline - -``` -Day 0 You report vulnerability -Day 1-2 We acknowledge receipt -Day 7 We confirm vulnerability and share initial assessment -Day 7-90 We develop and test fix -Day 90 Coordinated public disclosure - (earlier if fix is ready; later by mutual agreement) -``` - -If we cannot reach agreement on disclosure timing, we default to 90 days from your initial report. - ---- - -## Scope - -### In Scope ✅ - -The following are within scope for security research: - -- This repository (`hyperpolymath/verisimdb`) and all its code -- Official releases and packages published from this repository -- Documentation that could lead to security issues -- Build and deployment configurations in this repository -- Dependencies (report here, we'll coordinate with upstream) - -### Out of Scope ❌ - -The following are **not** in scope: - -- Third-party services we integrate with (report directly to them) -- Social engineering attacks against maintainers -- Physical security -- Denial of service attacks against production infrastructure -- Spam, phishing, or other non-technical attacks -- Issues already reported or publicly known -- Theoretical vulnerabilities without proof of concept - -### Qualifying Vulnerabilities - -We're particularly interested in: - -- Remote code execution -- SQL injection, command injection, code injection -- Authentication/authorisation bypass -- Cross-site scripting (XSS) and cross-site request forgery (CSRF) -- Server-side request forgery (SSRF) -- Path traversal / local file inclusion -- Information disclosure (credentials, PII, secrets) -- Cryptographic weaknesses -- Deserialisation vulnerabilities -- Memory safety issues (buffer overflows, use-after-free, etc.) -- Supply chain vulnerabilities (dependency confusion, etc.) -- Significant logic flaws - -### Non-Qualifying Issues - -The following generally do not qualify as security vulnerabilities: - -- Missing security headers on non-sensitive pages -- Clickjacking on pages without sensitive actions -- Self-XSS (requires victim to paste code) -- Missing rate limiting (unless it enables a specific attack) -- Username/email enumeration (unless high-risk context) -- Missing cookie flags on non-sensitive cookies -- Software version disclosure -- Verbose error messages (unless exposing secrets) -- Best practice deviations without demonstrable impact - ---- - -## Safe Harbour - -We support security research conducted in good faith. - -### Our Promise - -If you conduct security research in accordance with this policy: - -- ✅ We will not initiate legal action against you -- ✅ We will not report your activity to law enforcement -- ✅ We will work with you in good faith to resolve issues -- ✅ We consider your research authorised under the Computer Fraud and Abuse Act (CFAA), UK Computer Misuse Act, and similar laws -- ✅ We waive any potential claim against you for circumvention of security controls - -### Good Faith Requirements - -To qualify for safe harbour, you must: - -- Comply with this security policy -- Report vulnerabilities promptly -- Avoid privacy violations (do not access others' data) -- Avoid service degradation (no destructive testing) -- Not exploit vulnerabilities beyond proof-of-concept -- Not use vulnerabilities for profit (beyond bug bounties where offered) - -> **⚠️ Important:** This safe harbour does not extend to third-party systems. Always check their policies before testing. - ---- - -## Recognition - -We believe in recognising security researchers who help us improve. - -### Hall of Fame - -Researchers who report valid vulnerabilities will be acknowledged in our [Security Acknowledgments](SECURITY-ACKNOWLEDGMENTS.md) (unless they prefer anonymity). - -Recognition includes: - -- Your name (or chosen alias) -- Link to your website/profile (optional) -- Brief description of the vulnerability class -- Date of report - -### What We Offer - -- ✅ Public credit in security advisories -- ✅ Acknowledgment in release notes -- ✅ Entry in our Hall of Fame -- ✅ Reference/recommendation letter upon request (for significant findings) - -### What We Don't Currently Offer - -- ❌ Monetary bug bounties -- ❌ Hardware or swag -- ❌ Paid security research contracts - -> **Note:** We're a community project with limited resources. Your contributions help everyone who uses this software. - ---- - -## Security Updates - -### Receiving Updates - -To stay informed about security updates: - -- **Watch this repository**: Click "Watch" → "Custom" → Select "Security alerts" -- **GitHub Security Advisories**: Published at [Security Advisories](https://github.com/hyperpolymath/verisimdb/security/advisories) -- **Release notes**: Security fixes noted in [CHANGELOG](CHANGELOG.md) - -### Update Policy - -| Severity | Response | -|----------|----------| -| **Critical/High** | Patch release as soon as fix is ready | -| **Medium** | Included in next scheduled release (or earlier) | -| **Low** | Included in next scheduled release | - -### Supported Versions - - - -| Version | Supported | Notes | -|---------|-----------|-------| -| `main` branch | ✅ Yes | Latest development | -| Latest release | ✅ Yes | Current stable | -| Previous minor release | ✅ Yes | Security fixes backported | -| Older versions | ❌ No | Please upgrade | - ---- - -## Security Best Practices - -When using VeriSimDB, we recommend: - -### General - -- Keep dependencies up to date -- Use the latest stable release -- Subscribe to security notifications -- Review configuration against security documentation -- Follow principle of least privilege - -### For Contributors - -- Never commit secrets, credentials, or API keys -- Use signed commits (`git config commit.gpgsign true`) -- Review dependencies before adding them -- Run security linters locally before pushing -- Report any concerns about existing code - ---- - -## Additional Resources - -- [Security Advisories](https://github.com/hyperpolymath/verisimdb/security/advisories) -- [Changelog](CHANGELOG.md) -- [Contributing Guidelines](CONTRIBUTING.md) -- [CVE Database](https://cve.mitre.org/) -- [CVSS Calculator](https://www.first.org/cvss/calculator/3.1) - ---- - -## Contact - -| Purpose | Contact | -|---------|---------| -| **Security issues** | [Report via GitHub](https://github.com/hyperpolymath/verisimdb/security/advisories/new) or j.d.a.jewell@open.ac.uk | -| **General questions** | [GitHub Discussions](https://github.com/hyperpolymath/verisimdb/discussions) | -| **Other enquiries** | See [README](README.md) for contact information | - ---- - -## Policy Changes - -This security policy may be updated from time to time. Significant changes will be: - -- Committed to this repository with a clear commit message -- Noted in the changelog -- Announced via GitHub Discussions (for major changes) - ---- - -*Thank you for helping keep VeriSimDB and its users safe.* 🛡️ - ---- - -Last updated: 2026 · Policy version: 1.0.0 diff --git a/verisimdb/SONNET-TASKS.md b/verisimdb/SONNET-TASKS.md deleted file mode 100644 index ef9e6c43..00000000 --- a/verisimdb/SONNET-TASKS.md +++ /dev/null @@ -1,977 +0,0 @@ -# SONNET-TASKS.md — VeriSimDB (Round 2) - -**Date:** 2026-02-12 -**Repo:** `/var$REPOS_DIR/verisimdb/` -**Written by:** Opus (for Sonnet to execute) -**Previous round:** All 13 tasks from Round 1 completed successfully -**Honest completion before these tasks:** ~78% -**Target completion after these tasks:** ~88% - ---- - -## Ground Rules - -### Languages -- **Rust** — all crates under `rust-core/`. Edition 2021. Workspace root is `/var$REPOS_DIR/verisimdb/Cargo.toml`. -- **Elixir** — files under `elixir-orchestration/` and `lib/`. Mix project is `elixir-orchestration/mix.exs`. -- **ReScript** — files under `src/vcl/`. Do NOT touch the parser (`VCLParser.res`); it works. - -### What NOT to touch (these work — leave them alone) -- `src/vcl/VCLParser.res` — functional VCL parser -- `src/vcl/VCLError.res` — error types -- `src/vcl/VCLTypeChecker.res` — type checker -- `src/vcl/VCLExplain.res` — AST-based explain (fixed in Round 1) -- `rust-core/verisim-graph/` — Oxigraph integration works -- `rust-core/verisim-drift/` — drift detection works (11 tests pass) -- `rust-core/verisim-normalizer/` — normalization strategies work -- `rust-core/verisim-api/src/lib.rs` — HTTP API works (do NOT rewrite) -- `rust-core/verisim-octad/src/store.rs` — InMemoryOctadStore works (7 tests pass) -- `rust-core/verisim-document/src/lib.rs` — Tantivy + snippets work (2 tests pass) -- `lib/verisim/adaptive_learner.ex` — fully implemented, 4 domains -- `lib/verisim/query_cache.ex` — L1/L2/L3 all implemented (Round 1) -- `lib/verisim/query_router_cached.ex` — regex extraction works (Round 1) -- `elixir-orchestration/lib/verisim/drift/drift_monitor.ex` — sweep implemented (Round 1) - -### Testing requirements -- Every Rust change: `cargo test -p ` must pass -- Full workspace: `cargo test --workspace` — all non-ignored tests pass -- Every Elixir change: `mix compile` in `elixir-orchestration/` must succeed -- Run `cargo clippy --workspace` at end — zero warnings -- Run `cargo build --workspace` at end — must compile clean - -### Current test counts (baseline) -- **56 tests pass**, 4 ignored (persistence), 0 failures, 0 clippy warnings - -### Author attribution -- Git commits: `Jonathan D.A. Jewell ` -- Cargo.toml authors field: `["Jonathan D.A. Jewell "]` - ---- - -## Task 1: Implement Store Persistence (save_to_file / load_from_file) - -**Priority:** HIGH — unlocks 4 ignored integration tests - -### Files to modify -- `/var$REPOS_DIR/verisimdb/rust-core/verisim-vector/src/lib.rs` -- `/var$REPOS_DIR/verisimdb/rust-core/verisim-tensor/src/lib.rs` -- `/var$REPOS_DIR/verisimdb/rust-core/verisim-semantic/src/lib.rs` -- `/var$REPOS_DIR/verisimdb/rust-core/verisim-temporal/src/lib.rs` - -### Problem -Four stores lack persistence: `BruteForceVectorStore`, `InMemoryTensorStore`, `InMemorySemanticStore`, `InMemoryVersionStore`. Integration tests for these are `#[ignore]`'d. - -### What to do - -Use `postcard` (already in workspace dependencies) for serialization. Each store needs two methods. - -**Pattern to follow for all four stores:** - -```rust -use std::path::Path; -use std::fs; - -impl MyStore { - /// Save store contents to a file - pub fn save_to_file(&self, path: impl AsRef) -> Result<(), MyError> { - let data = self.internal_data.read().expect("lock poisoned"); - let bytes = postcard::to_allocvec(&*data) - .map_err(|e| MyError::SerializationError(e.to_string()))?; - fs::write(path, bytes) - .map_err(|e| MyError::SerializationError(e.to_string()))?; - Ok(()) - } - - /// Load store contents from a file - pub fn load_from_file(path: impl AsRef) -> Result { - let bytes = fs::read(path) - .map_err(|e| MyError::SerializationError(e.to_string()))?; - let data: InternalDataType = postcard::from_bytes(&bytes) - .map_err(|e| MyError::SerializationError(e.to_string()))?; - // Reconstruct the store from loaded data - Ok(Self { /* ... */ }) - } -} -``` - -**Store-specific details:** - -#### 1a. BruteForceVectorStore - -The internal data to serialize is `HashMap`. Both `String` and `Embedding` derive `Serialize`/`Deserialize`. - -```rust -/// Serializable snapshot of vector store state -#[derive(Serialize, Deserialize)] -struct VectorStoreSnapshot { - dimension: usize, - metric: DistanceMetric, - embeddings: HashMap, -} - -impl BruteForceVectorStore { - pub fn save_to_file(&self, path: impl AsRef) -> Result<(), VectorError> { - let embeddings = self.embeddings.read().expect("embeddings RwLock poisoned"); - let snapshot = VectorStoreSnapshot { - dimension: self.dimension, - metric: self.metric, - embeddings: embeddings.clone(), - }; - let bytes = postcard::to_allocvec(&snapshot) - .map_err(|e| VectorError::SerializationError(e.to_string()))?; - std::fs::write(path, bytes) - .map_err(|e| VectorError::SerializationError(e.to_string()))?; - Ok(()) - } - - pub fn load_from_file(path: impl AsRef) -> Result { - let bytes = std::fs::read(path) - .map_err(|e| VectorError::SerializationError(e.to_string()))?; - let snapshot: VectorStoreSnapshot = postcard::from_bytes(&bytes) - .map_err(|e| VectorError::SerializationError(e.to_string()))?; - Ok(Self { - dimension: snapshot.dimension, - metric: snapshot.metric, - embeddings: Arc::new(RwLock::new(snapshot.embeddings)), - }) - } - - /// Get basic stats about the store - pub fn stats(&self) -> VectorStoreStats { - let embeddings = self.embeddings.read().expect("embeddings RwLock poisoned"); - VectorStoreStats { - total_vectors: embeddings.len(), - dimension: self.dimension, - } - } -} - -#[derive(Debug, Clone)] -pub struct VectorStoreStats { - pub total_vectors: usize, - pub dimension: usize, -} -``` - -Add `postcard.workspace = true` to `rust-core/verisim-vector/Cargo.toml` under `[dependencies]`. - -#### 1b. InMemoryTensorStore - -The internal data is `HashMap`. `Tensor` already derives `Serialize`/`Deserialize`. - -Add `postcard.workspace = true` to `rust-core/verisim-tensor/Cargo.toml` under `[dependencies]`. - -```rust -impl InMemoryTensorStore { - pub fn save_to_file(&self, path: impl AsRef) -> Result<(), TensorError> { - let tensors = self.tensors.read().expect("tensors RwLock poisoned"); - let bytes = postcard::to_allocvec(&*tensors) - .map_err(|e| TensorError::SerializationError(e.to_string()))?; - std::fs::write(path, bytes) - .map_err(|e| TensorError::SerializationError(e.to_string()))?; - Ok(()) - } - - pub fn load_from_file(path: impl AsRef) -> Result { - let bytes = std::fs::read(path) - .map_err(|e| TensorError::SerializationError(e.to_string()))?; - let tensors: HashMap = postcard::from_bytes(&bytes) - .map_err(|e| TensorError::SerializationError(e.to_string()))?; - Ok(Self { - tensors: Arc::new(RwLock::new(tensors)), - }) - } -} -``` - -#### 1c. InMemorySemanticStore - -Internal data: `HashMap` and `HashMap`. Both derive `Serialize`/`Deserialize`. - -Add `postcard.workspace = true` to `rust-core/verisim-semantic/Cargo.toml` under `[dependencies]`. - -```rust -#[derive(Serialize, Deserialize)] -struct SemanticStoreSnapshot { - types: HashMap, - annotations: HashMap, -} - -impl InMemorySemanticStore { - pub fn save_to_file(&self, path: impl AsRef) -> Result<(), SemanticError> { - let types = self.types.read().expect("types RwLock poisoned"); - let annotations = self.annotations.read().expect("annotations RwLock poisoned"); - let snapshot = SemanticStoreSnapshot { - types: types.clone(), - annotations: annotations.clone(), - }; - let bytes = postcard::to_allocvec(&snapshot) - .map_err(|e| SemanticError::SerializationError(e.to_string()))?; - std::fs::write(path, bytes) - .map_err(|e| SemanticError::SerializationError(e.to_string()))?; - Ok(()) - } - - pub fn load_from_file(path: impl AsRef) -> Result { - let bytes = std::fs::read(path) - .map_err(|e| SemanticError::SerializationError(e.to_string()))?; - let snapshot: SemanticStoreSnapshot = postcard::from_bytes(&bytes) - .map_err(|e| SemanticError::SerializationError(e.to_string()))?; - Ok(Self { - types: Arc::new(RwLock::new(snapshot.types)), - annotations: Arc::new(RwLock::new(snapshot.annotations)), - }) - } -} -``` - -#### 1d. InMemoryVersionStore - -Internal data: `HashMap>>` where `T: Serialize + DeserializeOwned`. Add the bound to the impl block. - -Add `postcard.workspace = true` to `rust-core/verisim-temporal/Cargo.toml` under `[dependencies]`. - -```rust -impl InMemoryVersionStore -where - T: Clone + Send + Sync + Serialize + serde::de::DeserializeOwned + 'static, -{ - pub fn save_to_file(&self, path: impl AsRef) -> Result<(), TemporalError> { - let versions = self.versions.read().expect("versions RwLock poisoned"); - let bytes = postcard::to_allocvec(&*versions) - .map_err(|e| TemporalError::SerializationError(e.to_string()))?; - std::fs::write(path, bytes) - .map_err(|e| TemporalError::SerializationError(e.to_string()))?; - Ok(()) - } - - pub fn load_from_file(path: impl AsRef) -> Result { - let bytes = std::fs::read(path) - .map_err(|e| TemporalError::SerializationError(e.to_string()))?; - let versions = postcard::from_bytes(&bytes) - .map_err(|e| TemporalError::SerializationError(e.to_string()))?; - Ok(Self { - versions: Arc::new(RwLock::new(versions)), - }) - } -} -``` - -**IMPORTANT:** Check if `TemporalError` has a `SerializationError` variant. If not, add one: -```rust -#[error("Serialization error: {0}")] -SerializationError(String), -``` - -### After implementing all four stores - -Remove the `#[ignore]` annotations from the 4 persistence tests in `/var$REPOS_DIR/verisimdb/rust-core/verisim-octad/tests/integration_tests.rs` and update them: - -**test_vector_persistence** (line ~218): -```rust -#[tokio::test] -async fn test_vector_persistence() { - use std::fs; - let temp_path = "/tmp/verisim_integration_vector_test.bin"; - - let store = BruteForceVectorStore::new(64, DistanceMetric::Cosine); - - for i in 0..20 { - let mut vec = vec![0.0f32; 64]; - vec[i % 64] = 1.0; - let embedding = verisim_vector::Embedding::new(format!("vec_{}", i), vec); - store.upsert(&embedding).await.unwrap(); - } - - store.save_to_file(temp_path).unwrap(); - let loaded = BruteForceVectorStore::load_from_file(temp_path).unwrap(); - - assert_eq!(loaded.stats().total_vectors, 20); - - let mut query = vec![0.0f32; 64]; - query[0] = 1.0; - let results = loaded.search(&query, 3).await.unwrap(); - assert_eq!(results.len(), 3); - assert_eq!(results[0].id, "vec_0"); - - fs::remove_file(temp_path).ok(); -} -``` - -**test_tensor_persistence** (line ~240): -```rust -#[tokio::test] -async fn test_tensor_persistence() { - use std::fs; - use verisim_tensor::{Tensor, TensorStore as _}; - - let temp_path = "/tmp/verisim_integration_tensor_test.bin"; - let store = InMemoryTensorStore::new(); - - let t1 = Tensor::new("tensor_1", vec![2, 3], vec![1.0, 2.0, 3.0, 4.0, 5.0, 6.0]).unwrap(); - let t2 = Tensor::new("tensor_2", vec![3, 3], vec![1.0; 9]).unwrap(); - - store.put(&t1).await.unwrap(); - store.put(&t2).await.unwrap(); - - store.save_to_file(temp_path).unwrap(); - let loaded = InMemoryTensorStore::load_from_file(temp_path).unwrap(); - - let retrieved = loaded.get("tensor_1").await.unwrap().unwrap(); - assert_eq!(retrieved.shape, vec![2, 3]); - assert_eq!(retrieved.data, vec![1.0, 2.0, 3.0, 4.0, 5.0, 6.0]); - - let list = loaded.list().await.unwrap(); - assert_eq!(list.len(), 2); - - fs::remove_file(temp_path).ok(); -} -``` - -**test_semantic_persistence** (line ~265): -```rust -#[tokio::test] -async fn test_semantic_persistence() { - use std::fs; - use verisim_semantic::{SemanticStore as _, SemanticType, Constraint, ConstraintKind}; - - let temp_path = "/tmp/verisim_integration_semantic_test.bin"; - let store = InMemorySemanticStore::new(); - - let person_type = SemanticType::new("https://example.org/Person", "Person") - .with_supertype("https://example.org/Entity") - .with_constraint(Constraint { - name: "name_required".to_string(), - kind: ConstraintKind::Required("name".to_string()), - message: "Person must have a name".to_string(), - }); - let org_type = SemanticType::new("https://example.org/Organization", "Organization"); - - store.register_type(&person_type).await.unwrap(); - store.register_type(&org_type).await.unwrap(); - - store.save_to_file(temp_path).unwrap(); - let loaded = InMemorySemanticStore::load_from_file(temp_path).unwrap(); - - let retrieved = loaded.get_type("https://example.org/Person").await.unwrap().unwrap(); - assert_eq!(retrieved.label, "Person"); - assert_eq!(retrieved.constraints.len(), 1); - - let org = loaded.get_type("https://example.org/Organization").await.unwrap(); - assert!(org.is_some()); - - fs::remove_file(temp_path).ok(); -} -``` - -**test_temporal_persistence** (line ~300): -```rust -#[tokio::test] -async fn test_temporal_persistence() { - use std::fs; - use verisim_temporal::TemporalStore as _; - - let temp_path = "/tmp/verisim_integration_temporal_test.bin"; - let store: InMemoryVersionStore = InMemoryVersionStore::new(); - - store.append("entity1", "v1 data".to_string(), "alice", Some("first")).await.unwrap(); - store.append("entity1", "v2 data".to_string(), "bob", Some("second")).await.unwrap(); - store.append("entity2", "other data".to_string(), "charlie", None).await.unwrap(); - - store.save_to_file(temp_path).unwrap(); - let loaded: InMemoryVersionStore = InMemoryVersionStore::load_from_file(temp_path).unwrap(); - - let latest = loaded.latest("entity1").await.unwrap().unwrap(); - assert_eq!(latest.version, 2); - assert_eq!(latest.data, "v2 data"); - - let v1 = loaded.at_version("entity1", 1).await.unwrap().unwrap(); - assert_eq!(v1.data, "v1 data"); - - let history = loaded.history("entity1", 10).await.unwrap(); - assert_eq!(history.len(), 2); - - fs::remove_file(temp_path).ok(); -} -``` - -### Verification -```bash -cd /var$REPOS_DIR/verisimdb - -# Each store individually -cargo test -p verisim-vector -cargo test -p verisim-tensor -cargo test -p verisim-semantic -cargo test -p verisim-temporal - -# Integration tests (should now have 0 ignored) -cargo test -p verisim-octad --test integration_tests -# Must see: 11 passed, 0 ignored, 0 failed - -# Full workspace -cargo test --workspace -cargo clippy --workspace -``` - ---- - -## Task 2: Fix Panic in Temporal Diff compare_values - -**Priority:** MEDIUM — panics are never acceptable in library code - -### Files -- `/var$REPOS_DIR/verisimdb/rust-core/verisim-temporal/src/diff.rs` - -### Problem -Line 99: `(None, None) => panic!("Cannot compare two None values")` — calling `compare_values(None, None)` panics instead of returning a meaningful result. - -### What to do - -Replace line 99 with a graceful `Diff::no_change` that returns a default or special variant. Since comparing two `None` values means "nothing changed" (neither had a value), the semantically correct response is `Diff::NoChange`: - -```rust -(None, None) => Diff { - diff_type: DiffType::NoChange, - old_value: None, - new_value: None, -}, -``` - -Check the `Diff` struct definition to see if this construction is valid. If `Diff` requires `old_value` and `new_value` to be `Some`, you may need a new `DiffType::BothAbsent` variant, or simply: - -```rust -(None, None) => Diff { - diff_type: DiffType::NoChange, - old_value: None, - new_value: None, -}, -``` - -Also add a test: -```rust -#[test] -fn test_compare_values_both_none() { - let diff: Diff = compare_values(None, None); - assert!(!diff.has_change()); - assert_eq!(diff.old_value(), None); - assert_eq!(diff.new_value(), None); -} -``` - -### Verification -```bash -cargo test -p verisim-temporal -# Must see: test_compare_values_both_none ... ok -# Must NOT see any panics -``` - ---- - -## Task 3: Add OctadBuilder Convenience Methods - -**Priority:** LOW — improves API ergonomics - -### Files -- `/var$REPOS_DIR/verisimdb/rust-core/verisim-octad/src/lib.rs` - -### Problem -The builder has `with_types(Vec<&str>)` and `with_relationships(Vec<(&str, &str)>)`, but no singular convenience methods. The integration tests originally used `.with_semantic()` and `.with_relationship()` (singular), which is a more natural API for adding one item. - -### What to do - -Add these methods to `OctadBuilder` (after the existing methods, around line 335): - -```rust -/// Add a single relationship -pub fn with_relationship(self, predicate: &str, target: &str) -> Self { - self.with_relationships(vec![(predicate, target)]) -} - -/// Add semantic types (alias for with_types that accepts owned Strings) -pub fn with_semantic(mut self, type_iris: Vec) -> Self { - let refs: Vec<&str> = type_iris.iter().map(|s| s.as_str()).collect(); - self.with_types(refs) -} - -/// Add semantic properties -pub fn with_properties(mut self, properties: std::collections::HashMap) -> Self { - let existing = self.input.semantic.take().unwrap_or(OctadSemanticInput { - types: Vec::new(), - properties: std::collections::HashMap::new(), - }); - self.input.semantic = Some(OctadSemanticInput { - types: existing.types, - properties, - }); - self -} -``` - -### Verification -```bash -cargo test -p verisim-octad -cargo clippy -p verisim-octad -# All tests pass, no warnings -``` - ---- - -## Task 4: Add Drift-Triggered Normalization HTTP Endpoint - -**Priority:** MEDIUM — enables the Elixir drift monitor to query Rust core - -### Files -- `/var$REPOS_DIR/verisimdb/rust-core/verisim-api/src/lib.rs` - -### Problem -The Elixir drift monitor (`drift_monitor.ex:222`) has a TODO: "Query Rust core when get_drift_summary HTTP endpoint is ready." The Rust API has `/api/drift/status` but no `/api/drift/summary` endpoint that returns per-entity drift scores. - -### What to do - -Add a `GET /api/drift/summary` endpoint to the API. Read `verisim-api/src/lib.rs` first to understand the router structure (it uses Axum). - -The endpoint should return a JSON map of entity IDs to their drift scores: - -```json -{ - "entity-123": { - "semantic_vector_drift": 0.15, - "graph_document_drift": 0.03 - }, - "entity-456": { - "temporal_consistency_drift": 0.42 - } -} -``` - -Implementation approach: -1. Find where the Axum router is defined (look for `Router::new()`) -2. Add `.route("/api/drift/summary", get(drift_summary_handler))` -3. Implement the handler: - -```rust -async fn drift_summary_handler( - State(state): State, -) -> impl IntoResponse { - // Get all entity drift from the drift detector - let drift_detector = &state.drift_detector; - let summary = drift_detector.get_all_drift_scores().await; - Json(summary) -} -``` - -If `DriftDetector` doesn't have `get_all_drift_scores()`, add it to `verisim-drift/src/lib.rs`: - -```rust -impl DriftDetector { - /// Get drift scores for all entities that have been checked - pub async fn get_all_drift_scores(&self) -> HashMap> { - // Return the tracked metrics per entity - let metrics = self.metrics.read().await; - metrics.iter().map(|(id, m)| { - let scores: HashMap = m.iter() - .map(|(dt, metric)| (format!("{:?}", dt), metric.current_value())) - .collect(); - (id.clone(), scores) - }).collect() - } -} -``` - -**Check the actual `DriftDetector` API first** — it may already track per-entity metrics. Read `verisim-drift/src/lib.rs` to understand what's available. - -### Verification -```bash -cargo test -p verisim-api -cargo build -p verisim-api -# Start the server and test: -# curl http://localhost:8080/api/drift/summary -``` - ---- - -## Task 5: Make Query Cache Configuration Dynamic - -**Priority:** LOW — currently works with hardcoded defaults - -### Files -- `/var$REPOS_DIR/verisimdb/elixir-orchestration/lib/verisim/query/query_cache.ex` - -### Problem -Line 574: `# TODO: Make this configurable` — the `get_config()` function returns hardcoded `@default_config`. Cache TTL, max size, and eviction policy should be configurable at runtime. - -### What to do - -1. Read the file first to understand the current `get_config/0` and `@default_config`. - -2. Add a `configure/1` function to the GenServer that accepts a config map and stores it in state: - -```elixir -def configure(config) when is_map(config) do - GenServer.call(__MODULE__, {:configure, config}) -end -``` - -3. Handle the call in `handle_call`: -```elixir -def handle_call({:configure, new_config}, _from, state) do - merged = Map.merge(state.config, new_config) - {:reply, :ok, %{state | config: merged}} -end -``` - -4. Update `get_config/0` to read from state instead of returning `@default_config`: -```elixir -def get_config do - GenServer.call(__MODULE__, :get_config) -end -``` - -5. Handle: -```elixir -def handle_call(:get_config, _from, state) do - {:reply, state.config, state} -end -``` - -6. Ensure `init/1` initializes with `@default_config`: -```elixir -initial_state = %{ - config: @default_config, - # ... other state fields -} -``` - -### Verification -```bash -cd /var$REPOS_DIR/verisimdb/elixir-orchestration -mix compile -# No warnings -``` - ---- - -## Task 6: Add Normalizer Repair Strategies for Remaining Drift Types - -**Priority:** MEDIUM — normalizer only handles 2 of 6 drift types - -### Files -- `/var$REPOS_DIR/verisimdb/rust-core/verisim-normalizer/src/lib.rs` - -### Problem -The normalizer has strategies for `SemanticVectorDrift` and `GraphDocumentDrift` only. It has no strategies for: -- `TemporalConsistencyDrift` -- `TensorDrift` -- `SchemaDrift` -- `QualityDrift` - -### What to do - -Add strategy implementations for the remaining 4 drift types. Follow the same pattern as `SemanticVectorStrategy` and `GraphDocumentStrategy`. - -```rust -/// Strategy for temporal consistency drift -pub struct TemporalRepairStrategy; - -#[async_trait] -impl NormalizationStrategy for TemporalRepairStrategy { - fn name(&self) -> &str { - "temporal-consistency-repair" - } - - fn applies_to(&self, drift_type: DriftType) -> bool { - matches!(drift_type, DriftType::TemporalConsistencyDrift) - } - - async fn normalize( - &self, - octad: &Octad, - _drift_event: &DriftEvent, - ) -> Result { - // Repair temporal consistency by re-indexing version history - let changes = vec![NormalizationChange { - modality: "temporal".to_string(), - field: "version_history".to_string(), - old_value: None, - new_value: "[re-indexed from current state]".to_string(), - reason: "Temporal consistency drift detected".to_string(), - }]; - - Ok(NormalizationResult { - entity_id: octad.id.clone(), - normalization_type: NormalizationType::TemporalRepair, - success: true, - changes, - duration_ms: 0, - completed_at: Utc::now(), - }) - } -} - -/// Strategy for tensor drift -pub struct TensorSyncStrategy; - -#[async_trait] -impl NormalizationStrategy for TensorSyncStrategy { - fn name(&self) -> &str { - "tensor-sync" - } - - fn applies_to(&self, drift_type: DriftType) -> bool { - matches!(drift_type, DriftType::TensorDrift) - } - - async fn normalize( - &self, - octad: &Octad, - _drift_event: &DriftEvent, - ) -> Result { - let changes = vec![NormalizationChange { - modality: "tensor".to_string(), - field: "representation".to_string(), - old_value: None, - new_value: "[synchronized from source data]".to_string(), - reason: "Tensor drift detected".to_string(), - }]; - - Ok(NormalizationResult { - entity_id: octad.id.clone(), - normalization_type: NormalizationType::TensorSync, - success: true, - changes, - duration_ms: 0, - completed_at: Utc::now(), - }) - } -} - -/// Strategy for schema drift -pub struct SchemaRepairStrategy; - -#[async_trait] -impl NormalizationStrategy for SchemaRepairStrategy { - fn name(&self) -> &str { - "schema-repair" - } - - fn applies_to(&self, drift_type: DriftType) -> bool { - matches!(drift_type, DriftType::SchemaDrift) - } - - async fn normalize( - &self, - octad: &Octad, - _drift_event: &DriftEvent, - ) -> Result { - let changes = vec![NormalizationChange { - modality: "semantic".to_string(), - field: "schema_constraints".to_string(), - old_value: None, - new_value: "[re-validated against type registry]".to_string(), - reason: "Schema drift detected".to_string(), - }]; - - Ok(NormalizationResult { - entity_id: octad.id.clone(), - normalization_type: NormalizationType::FullReconciliation, - success: true, - changes, - duration_ms: 0, - completed_at: Utc::now(), - }) - } -} - -/// Strategy for general quality drift -pub struct QualityReconciliationStrategy; - -#[async_trait] -impl NormalizationStrategy for QualityReconciliationStrategy { - fn name(&self) -> &str { - "quality-reconciliation" - } - - fn applies_to(&self, drift_type: DriftType) -> bool { - matches!(drift_type, DriftType::QualityDrift) - } - - async fn normalize( - &self, - octad: &Octad, - _drift_event: &DriftEvent, - ) -> Result { - let changes = vec![NormalizationChange { - modality: "all".to_string(), - field: "cross_modal_consistency".to_string(), - old_value: None, - new_value: "[full reconciliation performed]".to_string(), - reason: "Quality drift detected — full reconciliation triggered".to_string(), - }]; - - Ok(NormalizationResult { - entity_id: octad.id.clone(), - normalization_type: NormalizationType::FullReconciliation, - success: true, - changes, - duration_ms: 0, - completed_at: Utc::now(), - }) - } -} -``` - -Register all new strategies in `create_default_normalizer`: - -```rust -pub async fn create_default_normalizer(drift_detector: Arc) -> Normalizer { - let normalizer = Normalizer::with_defaults(drift_detector); - normalizer.register_strategy(Arc::new(SemanticVectorStrategy)).await; - normalizer.register_strategy(Arc::new(GraphDocumentStrategy)).await; - normalizer.register_strategy(Arc::new(TemporalRepairStrategy)).await; - normalizer.register_strategy(Arc::new(TensorSyncStrategy)).await; - normalizer.register_strategy(Arc::new(SchemaRepairStrategy)).await; - normalizer.register_strategy(Arc::new(QualityReconciliationStrategy)).await; - normalizer -} -``` - -Add tests: - -```rust -#[tokio::test] -async fn test_all_drift_types_have_strategies() { - let drift_detector = Arc::new(DriftDetector::new(DriftThresholds::default())); - let normalizer = create_default_normalizer(drift_detector).await; - - let strategies = normalizer.strategies().await; - assert_eq!(strategies.len(), 6); - assert!(strategies.contains(&"semantic-vector-sync".to_string())); - assert!(strategies.contains(&"graph-document-sync".to_string())); - assert!(strategies.contains(&"temporal-consistency-repair".to_string())); - assert!(strategies.contains(&"tensor-sync".to_string())); - assert!(strategies.contains(&"schema-repair".to_string())); - assert!(strategies.contains(&"quality-reconciliation".to_string())); -} - -#[tokio::test] -async fn test_handle_tensor_drift() { - let drift_detector = Arc::new(DriftDetector::new(DriftThresholds::default())); - let normalizer = create_default_normalizer(drift_detector).await; - - let octad = create_test_octad(); - let event = DriftEvent::new(DriftType::TensorDrift, 0.5, "Test tensor drift"); - - let result = normalizer.handle_drift(&octad, &event).await.unwrap(); - assert!(result.is_some()); - assert!(result.unwrap().success); -} -``` - -### Verification -```bash -cargo test -p verisim-normalizer -# Must see: test_all_drift_types_have_strategies ... ok -# Must see: test_handle_tensor_drift ... ok -``` - ---- - -## Task 7: Update STATE.scm After Round 2 Completion - -**Priority:** DO THIS LAST — after all other tasks - -### Files -- `/var$REPOS_DIR/verisimdb/.machine_readable/STATE.scm` - -### What to do - -After completing Tasks 1-6, update: - -1. `overall-completion` from 75 to ~85 -2. Update component percentages: - - `rust-modality-stores` from 85 to 92 (persistence added) - - `integration-tests` from 70 to 90 (persistence tests unignored) - - `elixir-orchestration` from 70 to 75 (cache config dynamic) - -3. Update `blocked-on` — remove items completed, keep remaining - -4. Add session to `session-history`: -```scheme -(session - (date . "2026-02-12") - (phase . "persistence-and-polish") - (accomplishments - "- Implemented save_to_file/load_from_file on all 4 modality stores - - Unignored 4 persistence integration tests - - Fixed panic in temporal diff compare_values - - Added OctadBuilder convenience methods - - Added drift summary HTTP endpoint - - Made query cache configuration dynamic - - Added 4 remaining normalizer strategies (6/6 drift types covered) - - Updated STATE.scm completion percentages") - (key-decisions - "- postcard for serialization (already in workspace, no new deps) - - Persistence via file snapshots (not WAL or append log)")) -``` - -### Verification -```bash -grep "overall-completion" /var$REPOS_DIR/verisimdb/.machine_readable/STATE.scm -# Must show 85 (not 75 or 100) -``` - ---- - -## Final Verification - -After completing ALL 7 tasks: - -```bash -cd /var$REPOS_DIR/verisimdb - -# 1. Full workspace compiles -cargo build --workspace - -# 2. ALL tests pass (including formerly ignored persistence tests) -cargo test --workspace -# Expected: 60+ passed, 0 ignored, 0 failed - -# 3. No Clippy warnings -cargo clippy --workspace -- -D warnings - -# 4. Elixir compiles -cd elixir-orchestration && mix compile && cd .. - -# 5. No panics in diff module -cargo test -p verisim-temporal -- test_compare_values_both_none - -# 6. All 6 normalizer strategies registered -cargo test -p verisim-normalizer -- test_all_drift_types_have_strategies - -# 7. Binary still works -cargo build -p verisim-api -ls target/debug/verisim-api -``` - -If ALL checks pass, commit: - -```bash -git add -A -git commit -m "feat: add store persistence, normalizer strategies, and polish - -- Implement save_to_file/load_from_file on all 4 modality stores (postcard) -- Fix panic in temporal diff compare_values(None, None) -- Add OctadBuilder convenience methods (with_relationship, with_semantic) -- Add GET /api/drift/summary endpoint for Elixir integration -- Make query cache configuration dynamic -- Add 4 remaining normalizer strategies (6/6 drift types covered) -- Unignore 4 persistence integration tests -- Update STATE.scm to ~85% completion" -``` - -Then push: -```bash -git push origin main -git push gitlab main -``` diff --git a/verisimdb/TOPOLOGY.md b/verisimdb/TOPOLOGY.md deleted file mode 100644 index 3969c60a..00000000 --- a/verisimdb/TOPOLOGY.md +++ /dev/null @@ -1,113 +0,0 @@ - - - -# TOPOLOGY.md — VeriSimDB - -## System Architecture - -``` - ┌─────────────────────────────────────┐ - │ svalinn (TLS gateway) │ - │ ML-DSA-87 · policy: strict │ - └────────────────┬────────────────────┘ - │ :8443 - ┌────────────────▼────────────────────┐ - │ Elixir OTP Orchestration │ - │ ┌──────────┐ ┌──────────────────┐ │ - │ │ Entity │ │ Drift │ │ - │ │ Server │ │ Monitor │ │ - │ │(GenServer│ │(threshold-gated) │ │ - │ │ per │ │ │ │ - │ │ octad) │ │ Schema Registry │ │ - │ └──────────┘ └──────────────────┘ │ - │ ┌──────────┐ ┌──────────────────┐ │ - │ │ Query │ │ VCL Parser │ │ - │ │ Router │ │ (ReScript) │ │ - │ └──────────┘ └──────────────────┘ │ - └────────────────┬────────────────────┘ - │ HTTP :8080 - ┌────────────────────────────────▼────────────────────────────────┐ - │ Rust Core (verisim-api) │ - │ │ - │ ┌──────────┐ ┌──────────┐ ┌──────────┐ ┌──────────────────┐ │ - │ │ Graph │ │ Vector │ │ Tensor │ │ Semantic │ │ - │ │ Oxigraph │ │ HNSW │ │ndarray/ │ │ CBOR proofs │ │ - │ │ RDF + │ │ ANN │ │ Burn │ │ ZKP blobs │ │ - │ │ property │ │ search │ │ compute │ │ │ │ - │ └──────────┘ └──────────┘ └──────────┘ └──────────────────┘ │ - │ ┌──────────┐ ┌──────────┐ ┌──────────┐ ┌──────────────────┐ │ - │ │ Document │ │ Temporal │ │ Octad │ │ Normalizer │ │ - │ │ Tantivy │ │ version │ │ unified │ │ self-normalizing │ │ - │ │ full-text│ │ + time │ │ entity │ │ drift repair │ │ - │ │ index │ │ series │ │ 6-modal │ │ │ │ - │ └──────────┘ └──────────┘ └──────────┘ └──────────────────┘ │ - └────────────────────────────────────────────────────────────────┘ - │ - ┌──────────────────────┼──────────────────────┐ - ▼ ▼ ▼ - ┌──────────────┐ ┌──────────────────┐ ┌──────────────────┐ - │ vordr │ │ cerro-torre │ │ rokur │ - │ runtime │ │ image signing │ │ secret │ - │ verification │ │ ML-DSA-87 │ │ rotation │ - │ formal proofs│ │ SBOM + SLSA 3 │ │ argon2id │ - └──────────────┘ └──────────────────┘ └──────────────────┘ - - Data flow: - panic-attack → verisimdb-data (flat-file) → hypatia (rules) → gitbot-fleet (fixes) -``` - -## Completion Dashboard - -| Component | Progress | Status | -|------------------------|------------------------------|--------------| -| verisim-graph | `████████░░` 80% | Active | -| verisim-vector | `████████░░` 80% | Active | -| verisim-tensor | `███████░░░` 70% | Active | -| verisim-semantic | `██████░░░░` 60% | Active | -| verisim-document | `████████░░` 80% | Active | -| verisim-temporal | `███████░░░` 70% | Active | -| verisim-octad | `████████░░` 80% | Active | -| verisim-drift | `███████░░░` 70% | Active | -| verisim-normalizer | `██████░░░░` 60% | Active | -| verisim-api | `████████░░` 80% | Active | -| Elixir OTP layer | `███████░░░` 70% | Active | -| VCL parser | `█████████░` 95% | Active | -| VCL-UT (Lean checker) | `░░░░░░░░░░` 0% | Not started | -| Idris2 ABI | `████░░░░░░` 40% | In progress | -| Zig FFI | `████░░░░░░` 40% | In progress | -| Containerfile | `██████████` 100% | Complete | -| selur-compose | `██████████` 100% | Complete | -| stapeln.toml | `██████████` 100% | Complete | -| verisimdb-data (CI) | `████████░░` 80% | Active | -| Hypatia integration | `████░░░░░░` 40% | In progress | -| proven integration | `░░░░░░░░░░` 0% | Planned | -| **Overall** | `██████░░░░` **65%** | | - -## Key Dependencies - -``` -verisimdb -├── oxigraph (graph store — RDF + property graph) -├── hnsw_rs (vector similarity — ANN search) -├── tantivy (document store — full-text indexing) -├── burn (tensor compute — ML inference) -├── ndarray (tensor operations) -├── chrono (temporal versioning) -├── cbor (semantic proof blobs) -├── axum (HTTP API framework) -├── tokio (async runtime) -├── elixir 1.18 / OTP 27 (orchestration) -│ -├── Container ecosystem: -│ ├── svalinn (TLS gateway + policy enforcement) -│ ├── vordr (runtime verification) -│ ├── cerro-torre (image signing, ML-DSA-87) -│ ├── rokur (secret rotation, argon2id) -│ └── stapeln (layer-based builds) -│ -└── Data pipeline: - ├── panic-attacker → scan results - ├── verisimdb-data → flat-file store - ├── hypatia → rule engine - └── gitbot-fleet → automated fixes -``` diff --git a/verisimdb/VOID-SETUP.md b/verisimdb/VOID-SETUP.md deleted file mode 100644 index 49b969e1..00000000 --- a/verisimdb/VOID-SETUP.md +++ /dev/null @@ -1,205 +0,0 @@ -# VoID (Vocabulary of Interlinked Datasets) Setup - -## What is VoID? - -VoID is an RDF vocabulary for expressing metadata about RDF datasets. It helps: -- Discover datasets -- Understand dataset structure -- Find linked data connections -- Enable SPARQL endpoints -- Interoperate with semantic web tools - -## Why VoID for verisimdb? - -VoID enables verisimdb.dev to: -1. **Publish structured data** as Linked Open Data -2. **Connect to other datasets** (DBpedia, Wikidata, Schema.org) -3. **Enable semantic queries** via SPARQL -4. **Improve discoverability** in semantic web search engines -5. **Support research** and data integration - -## VoID Files - -- `.well-known/void.ttl` - Turtle format (human-readable) -- `.well-known/void.rdf` - RDF/XML format (tool-compatible) - -## Accessing VoID Metadata - -```bash -# Turtle format -curl https://verisimdb.dev/.well-known/void.ttl - -# RDF/XML format -curl https://verisimdb.dev/.well-known/void.rdf -``` - -## Example: Querying with SPARQL - -```sparql -PREFIX void: -PREFIX dcterms: - -SELECT ?dataset ?title ?triples -WHERE { - ?dataset a void:Dataset ; - dcterms:title ?title ; - void:triples ?triples . -} -``` - -## Integration with verisimdb - -VoID is **perfect** for verisimdb because: - -1. **Semantic Database**: verisimdb can expose its data as RDF -2. **Linksets**: Connect verisimdb entities to external datasets -3. **SPARQL Endpoint**: Query verisimdb using SPARQL -4. **Schema Alignment**: Map verisimdb schema to standard ontologies - -### Example verisimdb Integration - -```turtle -# verisimdb dataset with linksets - a void:Dataset ; - dcterms:title "VerisimDB Verified Data" ; - void:triples 1000000 ; - void:entities 50000 ; - - # Link to DBpedia - void:subset ; - - # Link to Wikidata - void:subset ; - - # SPARQL endpoint - void:sparqlEndpoint ; -. - -# Linkset to DBpedia - a void:Linkset ; - void:linkPredicate owl:sameAs ; - void:target ; - void:target ; - void:triples 25000 ; -. -``` - -## Serving VoID via SSG - -For static site generators (SSG), VoID files can be: - -1. **Pre-generated** during build -2. **Served as static files** from .well-known/ -3. **Content-negotiated** (Turtle for browsers, RDF/XML for tools) - -### Example: ReScript SSG Integration - -```rescript -// void-generator.res -let generateVoID = (dataset: Dataset.t) => { - let ttl = ` -@prefix void: . - - a void:Dataset ; - void:triples ${Int.toString(dataset.tripleCount)} ; - void:entities ${Int.toString(dataset.entityCount)} . -` - // Write to .well-known/void.ttl - Node.Fs.writeFileSync(".well-known/void.ttl", ttl) -} -``` - -## Linking to External Datasets - -### DBpedia - -```turtle -void:subset [ - a void:Linkset ; - void:linkPredicate owl:sameAs ; - void:target ; - void:exampleResource ; -] . -``` - -### Wikidata - -```turtle -void:subset [ - a void:Linkset ; - void:linkPredicate owl:sameAs ; - void:target ; - void:exampleResource ; -] . -``` - -### Schema.org - -```turtle -void:vocabulary ; -void:vocabularyPartition [ - void:class schema:Person ; - void:entities 1000 ; -] . -``` - -## SPARQL Endpoint (Future) - -To add a SPARQL endpoint: - -1. **Cloudflare Worker** can proxy SPARQL queries -2. **GitHub Pages** can serve static SPARQL results -3. **Dedicated backend** (for dynamic queries) - -```javascript -// sparql-worker.js (Cloudflare Worker) -addEventListener('fetch', event => { - event.respondWith(handleSPARQL(event.request)) -}) - -async function handleSPARQL(request) { - const query = await request.text() - // Parse SPARQL query - // Execute against RDF store - // Return results as JSON-LD or Turtle -} -``` - -## Validation - -Validate VoID files: - -```bash -# Using rapper (RDF parser) -rapper -i turtle .well-known/void.ttl - -# Using Apache Jena -riot --validate .well-known/void.ttl -``` - -## Discovery - -VoID metadata is discoverable via: -- **SPARQL endpoints**: `https://verisimdb.dev/sparql` -- **.well-known/**: `https://verisimdb.dev/.well-known/void.ttl` -- **HTTP Headers**: `Link: ; rel="meta"` -- **HTML ``**: `` - -## Next Steps for verisimdb - -1. **Export verisimdb data as RDF** (Turtle, N-Triples, RDF/XML) -2. **Update VoID statistics** (triple count, entities, etc.) -3. **Create linksets** to DBpedia, Wikidata, Schema.org -4. **Deploy SPARQL endpoint** (Cloudflare Worker or dedicated server) -5. **Add content negotiation** (serve different formats based on Accept header) - -## Resources - -- VoID Specification: https://www.w3.org/TR/void/ -- VoID Guide: https://semanticweb.org/wiki/VoID -- LOD Cloud: https://lod-cloud.net/ -- Linked Data: https://www.w3.org/DesignIssues/LinkedData.html - -## License - -PMPL-1.0-or-later diff --git a/verisimdb/WHITEPAPER.md b/verisimdb/WHITEPAPER.md deleted file mode 100644 index a8c30fe2..00000000 --- a/verisimdb/WHITEPAPER.md +++ /dev/null @@ -1,385 +0,0 @@ - -# White Paper -## **VeriSimDB: A Tiny Core for Universal Federated Knowledge** -**Author:** Jonathan Jewell – Hyper‑Polymath -**Date:** 2 November 2025 -**Version:** 1.0 - ---- - -### Table of Contents - -1. [Executive Summary](#executive-summary) -2. [1. Introduction](#1-introduction) - - 2.1. The Fragmentation Problem - - 2.2. Emerging Opportunities -3. [2. The Tiny Core Architecture](#2-the-tiny-core-architecture) - - 2.1. ReScript Registry (Memory #1) - - 2.2. Elixir Orchestration & Drift Management (Memory #5) - - 2.3. Rust Modality Crates (Memory #4) - - 2.4. WASM‑based Public Proxy (Memory #2) - - 2.5. Zero‑Trust Signing (sactify‑php, Memory #2) -4. [3. Universal Federated Store Integration](#3-universal-federated-store-integration) -5. [4. Drift‑Tolerant Knowledge Semantics](#4-drift-tolerant-knowledge-semantics) - - 4.1. Detection Strategies - - 4.2. Repair Policies -6. [5. Zero‑Trust Security Design](#5-zero-trust-security-design) -7. [6. Modality‑Agnostic Federation](#6-modality-agnostic-federation) - - 6.1. Graph (Oxigraph) - - 6.2. Vector (HNSW) - - 6.3. Tensor (ndarray/Burn) - - 6.4. Semantic (CBOR) - - 6.5. Document (Tantivy) - - 6.6. Temporal (Version Trees) -8. [7. Real‑World Use Cases](#7-real-world-use-cases) -9. [8. Comparative Landscape](#8-comparative-landscape) -10. [9. Implementation Roadmap](#9-implementation-roadmap) -11. [10. Ethical & Philosophical Implications](#10-ethical--philosophical-implications) -12. [11. Conclusions & Call to Action](#11-conclusions--call-to-action) -13. [References](#references) - ---- - -## Executive Summary - -**VeriSimDB** is a minimalistic, address‑space‑centric core that enables **any** data modality to be federated across heterogeneous, independently‑operated stores while preserving **drift tolerance**, **ethical governance**, and **Zero‑Trust security**. - -- **Core size:** < 5 k LOC of ReScript + Elixir orchestration. -- **Key capabilities:** Global UUID namespace, on‑demand modality loading, drift detection/repair, immutable audit trails, and modular plug‑in support for Graph, Vector, Tensor, Semantic, Document, and Temporal data. -- **Why it matters:** Modern research, open‑science, and AI pipelines demand *interoperable* knowledge that can evolve without forcing a monolithic consistency model. VeriSimDB provides that missing “tiny core” while keeping the heavy lifting (storage, modality logic) in the federated nodes. - -The remainder of this paper details the architecture rationale, design choices, security model, and practical pathways for adoption. - ---- - -## 1. Introduction - -### 1.1. The Fragmentation Problem - -| Symptom | Example | Consequence | -|---------|---------|-------------| -| **Data silos** | University repositories, proprietary archives | Redundant copies, missed collaborations | -| **Inconsistent semantics** | Retraction of a paper, legal reinterpretation of a record | Knowledge drift leads to erroneous aggregations | -| **Operational brittleness** | Legacy mainframes, proprietary APIs | Integration costs explode | - -Traditional federated systems (e.g., IPFS, Solid, Dat) solve *some* of these issues but enforce **either strict consistency** or **pure peer‑to‑peer autonomy**. Neither can reconcile the need for *controlled drift* (where some updates must propagate, others must stay local) nor enforce *fine‑grained ethical constraints* on who can read or write a particular knowledge artifact. - -### 1.2. Emerging Opportunities - -1. **Neurosymbolic AI** – hybridization of embeddings (vector) and symbolic graphs demands a federation capable of simultaneously serving different modalities. -2. **Open Science & FAIR data** – funding agencies now require data provenance and versioning across institutional boundaries. -3. **Decentralized Web (Web3)** – community governance models rely on verifiable, tamper‑evident data exchanges. - -These trends converge on a single requirement: **a universal, address‑able namespace that can host heterogeneous knowledge units (Octads) while allowing controlled drift and auditable trust boundaries.** - ---- - -## 2. The Tiny Core Architecture - -The VeriSimDB **core** is deliberately *tiny*: it only provides *namespace resolution* and *lightweight coordination*. All modality‑specific logic lives in federated crates that can be versioned, replaced, or upgraded independently. - -### 2.1. ReScript Registry – The Global Namespace (Memory #1) - -- **Function**: Maps each **Octad UUID** (128‑bit) to a *store identifier* and a *metadata bundle* (modality list, access policy hash). -- **Implementation**: Pure ReScript, compiled to a tiny JavaScript module that runs inside a **WASM sandbox**. -- **Why ReScript?** - - Strong static typing eliminates runtime reinterpretation bugs. - - Immutable data structures naturally mirror the *address‑space* semantics. - - You retain exclusive ownership of the registry logic (Memory #1). - -```rescript -// registry.res -type storeId = string; -type octadId = string; - -type storeMeta = { - endpoint: string, - modalities: array, - policyHash: string, -}; - -var registry: map = /* empty */; -``` - -The registry is **stateless** – all mutations are persisted as *signed append‑only events* (see §5). - -### 2.2. Elixir Orchestration & Drift Management (Memory #5) - -- **Framework**: Elixir + GenStage pipelines. -- **Responsibilities**: - 1. **Synchronization** – polling or push‑based updates from stores. - 2. **Drift Detection** – statistical & formal checks (see §4). - 3. **Repair Scheduling** – dispatch manual or automated repair jobs. -- **Why Elixir?** - - Fault‑tolerant actor model fits long‑running coordination. - - GenStage enables composable stages (fetch → detect → repair). - - Your production stack already ships with Elixir (Memory #5). - -### 2.3. Rust Modality Crates – Store‑Side Logic (Memory #4) - -| Modality | Crate | Core API | Performance Note | -|----------|-------|----------|-------------------| -| Graph | `verisim-graph-rs` | `load_graph(octad_id) -> Oxigraph` | Zero‑copy `Arc` over MMAP files | -| Vector | `verisim-vector-rs` | `search(embedding) -> nearest` | HNSW built on `hnsw-sys` (sub‑ms latency) | -| Tensor | `verisim-tensor-rs` | `load_tensor(octad_id) -> ndarray::Array` | Integrates with `burn` for on‑the‑fly inference | -| Document | `verisim-doc-rs` | `fulltext_search(query) -> tantivy::Result` | Uses Tantivy’s inverted index, store‑level compression | -| Semantic | `verisim-semantic-rs` | `type_proof(cbor_blob) -> enum` | CBOR schema validation via `cbor-rs` | -| Temporal | `verisim-temporal-rs` | `versions(octad_id) -> tree` | Merkle‑tree snapshots for deterministic replay | - -All crates expose a **C‑ABI** friendly entry point (`#[no_mangle] pub extern "C"`), enabling direct calls from the Elixir orchestrator without marshalling overhead. - -### 2.4. WASM Proxy – Public API (Memory #2) - -- **Purpose**: Serve HTTP/HTTPS endpoints to *any* client (browser, CLI, third‑party AI) while preserving **statelessness** and **sandboxing**. -- **Composition**: - - **Memory #2** – the sandbox itself (the WebAssembly module). - - **Interacts** with the ReScript registry to resolve UUID → store mapping. - - **Enforces** access signatures (see §5). -- **Benefits**: - - Language‑agnostic: any client can issue a simple JSON‑RPC call (`/octad/:id`). - - No direct filesystem access; all I/O is mediated by the core’s policy engine. - -### 2.5. Zero‑Trust Signing – sactify‑php (Memory #2) - -- **Mechanism**: Every request must carry a **cryptographic signature** generated by `sactify-php`. -- **Workflow**: - 1. Client loads its private key (hardware security module optional). - 2. Computes `sign(payload, private_key)`. - 3. POSTs `{payload, signature}` to the WASM proxy. - 4. Proxy verifies using the public key associated with the client’s *identity claim* (stored in the registry). -- **Why sactify‑php?** It is already part of your security toolbox (Memory #2) and provides **tamper‑evident audit logs** that can be appended to the immutable `verisim-temporal` log. - ---- - -## 3. Universal Federated Store Integration - -### 3.1. Registration Flow - -1. **Discover** – A store publishes a *registration manifest* (JSON) containing: - - `store_id` (UUID) - - `endpoints` (graph, vector, …) - - `supported_modalities` (array) - - `policy_hash` (SHA‑256 of its internal access policy) -2. **Commit** – The manifest is signed with the store’s private key and posted to `/registry` (the ReScript core). -3. **Acknowledge** – Elixir orchestrator adds the entry to the registry and dispatches a **registration event** to downstream modules. - -### 3.2. Octad Lifecycle - -| Phase | Actor | Action | Core Interaction | -|-------|-------|--------|------------------| -| **Store creation** | Store | Emits a *Octad* (UUID + payload + modality tags) | Registry entry created; metadata attached | -| **Fetch** | Client | Requests `GET /octad/:id` | WASM proxy resolves UUID → store, forwards request | -| **Sync** | Orchestrator | Pulls updates from all stores that host the Octad | GenStage pipelines perform *diff* ingestion | -| **Drift repair** | Orchestrator/Store | Detects change; decides to *auto‑repair* or *prompt user* | Repair job scheduled; signed by policy holder | - -Stores effectively behave as *virtual memory pages*: a Octad’s vector modality may reside on Store A, while its document modality resides on Store B. The core never copies data; it merely **routes** the request and enforces policy. - ---- - -## 4. Drift‑Tolerant Knowledge Semantics - -Knowledge is *inherently dynamic*. VeriSimDB embraces this through a **drift taxonomy** that separates **statistical drift** (harmless updates) from **formal drift** (semantic or ethical changes). - -### 4.1. Detection Strategies - -| Modality | Statistic | Thresholding Method | -|----------|-----------|---------------------| -| Vector embeddings | Cosine similarity (pairwise) | Empirical quantile from recent batch | -| Tensor fields | Frobenius norm of delta | Adaptive sigma based on runtime statistics | -| Graph edges | Edge‑addition rate | Sliding‑window Poisson test | -| Document content | TF‑IDF cosine (section level) | Pre‑trained classifier for “retraction” vs “Revision” | -| Temporal versions | Version‑tree depth increase | Formal rule: `max_depth ≤ 3` per policy | - -Statistical drift is **sampled** every *N* minutes (configurable). If the sample exceeds a *soft* threshold, the system **marks** the Octad as *potentially drifted* but does **not** automatically repair. - -### 4.2. Repair Policies - -1. **Automatic (non‑critical)** – e.g., a new research paper version with updated references is accepted without human review if the drift score < *critical* level. -2. **Semi‑automatic** – e.g., a change to a *definition* in a taxonomy triggers a **confidence‑scoped alert** visible to domain custodians. -3. **Manual** – e.g., a contested historical narrative alteration requires a signed **Ethics Review** from an authorized committee; the signature is recorded in the immutable temporal log. - -Repair actions are **policy‑driven** (encoded in CBOR Semantic proofs). The core can be extended with new *drift‑resolution rules* without touching the underlying registry. - ---- - -## 5. Zero‑Trust Security Design - -### 5.1. Core Principles - -| Principle | Mechanism | -|-----------|-----------| -| **Never trust by default** | Every request must present a **cryptographic signature** created with the requester’s private key. | -| **Least privilege** | Access policies are encoded as **hashes** in the registry; verification is delegated to the store that actually hosts the data. | -| **Auditability** | All interactions are appended to an **immutable temporal ledger** (`verisim-temporal`). Each log entry contains: timestamp, UUID, signer, policy hash, and a hash chain linking to the prior entry. | -| **Isolation** | Stores run in **rootless containers** (`svalinn/vordr`, Memory #4). No container shares a network namespace unless explicitly allowed. | - -### 5.2. Signature Flow (Figure 1 – omitted) - -1. **Client** creates payload: `{action: "read", octad_id: "0x12AB…"}` -2. **Client** signs payload → `sig = sactify-php sign(payload, priv_key)` -3. **Client** POSTs `{payload, sig}` to `/octad/:id` (WASM proxy) -4. **Proxy** resolves UUID → `store_id` -5. **Proxy** verifies `sig` against the public key associated with the caller’s DID (Decentralized Identifier stored in registry) -6. **Proxy** forwards request to the identified store only if verification succeeds. -7. **Store** returns data; proxy records the transaction hash in `verisim-temporal`. - -All signatures are **timestamped**; replay attacks are impossible because the timestamp is part of the signed blob. - ---- - -## 6. Modality‑Agnostic Federation - -VeriSimDB treats *every* data type as a **first‑class modality**. Below is the canonical mapping that the core knows about, together with the recommended Rust crate. - -| Modality | Symbolic Name | Core Metadata Tag | Rust Crate | Example Use | -|----------|---------------|-------------------|------------|-------------| -| Graph | `graph` | `"graph"` | `verisim-graph-rs` (Oxigraph) | Citation networks, social graphs | -| Vector | `vector` | `"vector"` | `verisim-vector-rs` (HNSW) | Embedding search, similarity | -| Tensor | `tensor` | `"tensor"` | `verisim-tensor-rs` (ndarray/Burn) | Sensor streams, ML model weights | -| Semantic | `semantic` | `"semantic"` | `verisim-semantic-rs` (CBOR) | Type proofs, ontology annotations | -| Document | `document` | `"document"` | `verisim-doc-rs` (Tantivy) | Full‑text articles, legal codes | -| Temporal | `temporal` | `"temporal"` | `verisim-temporal-rs` (Merkle-tree) | Version histories, draft revisions | - -Adding a **new modality** (e.g., *audio* or *geospatial*) requires only: - -1. A Rust crate exposing `load_(octad_id)`. -2. An entry in the *modality registry* (a static map inside the core). -3. Optional *validation* logic (e.g., schema checks). - -Because all modality bundles are **self‑describing** (CBOR carries type metadata), the core never needs to be recompiled to support new data types. - ---- - -## 7. Real‑World Use Cases - -### 7.1. Open Science - -- **Participants**: University data repositories, pre‑print servers (arXiv, bioRxiv), clinical trial registries. -- **Workflow**: - 1. Each repo registers its endpoint. - 2. When a paper is uploaded, a Octad is minted containing *citation graph*, *embedding vector*, and *document* modalities. - 3. The core synchronizes the Octad across all repositories. - 4. Retraction events trigger *drift detection* and optional *manual review*. -- **Benefit**: Researchers can query a *global citation graph* without harvesting each repository individually, while preserving institutional data sovereignty. - -### 7.2. Digital Humanities - -- **Participants**: National archives, museum collections, private libraries. -- **Scenario**: Reconciling divergent historical accounts (e.g., competing narratives of a war). -- **How VeriSimDB Helps**: - - Each archive tags its version with a *semantic proof* (“revision‑v1”, “revision‑v2”). - - Drift detection flags when a contested narrative gains prominence. - - Ethical review committees sign off on additions. -- **Outcome**: A **multivocal timeline** that can be explored without erasing prior editions, facilitating transparent historiography. - -### 7.3. Neurosymbolic AI - -- **Participants**: AI labs, knowledge‑graph platforms, reinforcement‑learning research groups. -- **Integration**: - - Vector modality stores embeddings from transformer models. - - Graph modality holds relational facts. - - Tensor modality persists latent state tensors of live models. -- **Outcome**: A single Octad can be traversed from *semantic proof* → *graph* → *vector* → *tensor* without moving data, enabling **reason‑driven retrieval** and **runtime grounding** of AI predictions. - -### 7.4. Healthcare & Personal Data - -- **Participants**: Hospitals, patient‑generated health apps, public health agencies. -- **Privacy Model**: - - Each patient owns a *personal namespace* of Octads. - - Access requires **patient‑signed policy**; policy hash is stored in the registry. - - Drift detection respects *clinical relevance thresholds* (e.g., only version changes for lab results > 10% shift trigger review). -- **Result**: A **patient‑centric knowledge graph** that can be securely shared across institutions while respecting consent and regulatory constraints. - -### 7.5. Web3 & Decentralized Wikis - -- **Participants**: DAO‑governed knowledge bases, decentralized social platforms. -- **Mechanics**: - - Governance tokens are mapped to *signature authorities*. - - Proposals to alter a Octad must carry a quorum of signatures. - - The immutable temporal log provides a *public audit trail* for governance disputes. -- **Impact**: Knowledge ownership becomes **token‑backed yet ethically governed**, aligning with Web3 principles of transparency and accountability. - ---- - -## 8. Comparative Landscape - -| System | Federation Model | Drift Handling | Multimodality | Zero‑Trust Built‑In | Core Stack Overlap | -|--------|------------------|----------------|---------------|----------------------|--------------------| -| **VeriSimDB** | *Namespace‑centric* (HD‑like) | ✔︎ Statistical & formal | ✔︎ Graph, Vector, Tensor, Semantic, Document, Temporal | ✔︎ sactify‑php signatures, WASM sandbox, immutable logs | ReScript (core), Elixir (orch), Rust (store), WASM (proxy), sactify‑php (signing) | -| **Solid Project** | Pod‑based (Web‑oriented) | ✘ No systematic drift detection | ✘ Primarily JSON‑LD/Document | ✔︎ DID‑based auth (but not signed payloads by default) | Mostly JavaScript/TypeScript | -| **IPFS** | Content‑addressed (block‑level) | ✘ Immutable by design | ✘ Limited to raw bytes | ✘ No built‑in auth beyond TLS | Written in Go, Rust | -| **Dat** | Append‑only log | ✘ Linear log only | ✘ Primarily document‑oriented | ✘ Stateless, no signing by default | Node.js/Rust | -| **AlphaFold DB** | Centralized repository | ✘ No drift, versioned per release | ✘ Domain‑specific (protein structures) | ✘ No external auth | Python/C++ | - -*Key Takeaway*: VeriSimDB uniquely **combines** a **tiny universal namespace**, **drift‑aware semantics**, **first‑class multimodal support**, and **Zero‑Trust enforcement** within a stack you already own. - ---- - -## 9. Implementation Roadmap - -| Milestone | Duration | Deliverable | Owner | -|-----------|----------|-------------|-------| -| **0. Foundations** | 1 wk | Project scaffolding (GitHub, CI) | You | -| **1. ReScript Registry** | 1 wk | `registry.res` compiled to WASM; API spec | You | -| **2. Elixir Orchestrator** | 2 wks | Registration, sync pipelines, drift detection stub | You | -| **3. Rust Modality Crates** | 3 wks | `verisim-graph-rs`, `verisim-vector-rs`, … with test suites | You | -| **4. WASM Proxy** | 1 wk | `/octad/:id` endpoint, signature verification hook | You | -| **5. sactify‑php Integration** | 1 wk | Signature verification service (deployed as side‑car) | You | -| **6. Pilot Stores** | 2 wks | 3 test stores (e.g., GitHub repo, local Oxigraph instance, mock paper DB) | You + collaborators | -| **7. Drift Engine** | 1 wk | Full statistical + proven‑library rule engine | You | -| **8. Security Hardening** | 1 wk | Auditable logs, rootless container sandboxing | You | -| **9. Documentation & Release** | 1 wk | Public repo, README, API docs | You | -| **Total** | **13 weeks** | **Beta‑ready VeriSimDB core** | — | - -*Post‑Beta*: community‑driven extension of modalities, governance plug‑ins, and integration SDKs (Python, Go, R). - ---- - -## 10. Ethical & Philosophical Implications - -1. **Epistemic Humility** – By modeling *drift* as a first‑class concept, VeriSimDB operationalizes the idea that **knowledge is provisional**. This aligns with contemporary philosophy of science (Kuhnian paradigm shifts) and supports *open‑minded* revision without forced consensus. - -2. **Bias Detection** – Statistical drift signals can expose systematic biases (e.g., a dominant narrative gaining disproportionate representation). The system can surface these signals to domain experts, facilitating **bias remediation** rather than concealment. - -3. **Agency & Ownership** – Each Octad is *individually owned* (via its UUID) yet *participates* in a global namespace. This balances **individual sovereignty** with **collective epistemic infrastructure**—a model resonant with contemporary debates on data‑property rights. - -4. **Governance Transparency** – Immutable temporal logs provide an auditable *public ledger of epistemic decisions*. This satisfies the demands of **participatory governance** in open‑science consortia and DAO communities. - -These philosophical underpinnings are not merely academic; they directly inform **policy decisions** (e.g., who may sign a drift‑repair request) and **product design** (e.g., exposing drift metrics in UI dashboards). - ---- - -## 11. Conclusions & Call to Action - -- **VeriSimDB** provides the *missing glue* for universal federated knowledge: a **tiny, addressable core**, **drift‑aware semantics**, **plug‑in modality support**, and **Zero‑Trust security**—all built on the stack you already own (ReScript, Elixir, Rust, WASM, sactify‑php). -- The architecture is deliberately **minimalist**, allowing rapid iteration and independent evolution of each module. -- Early adopters can **pilot** the system with a handful of research repositories, gaining immediate benefits in **knowledge discoverability**, **reproducibility**, and **ethical governance**. - -**Next Steps** - -1. **Clone & explore** the reference implementation (`hyperpolymath/verisimdb`). -2. **Register** a test store (e.g., a local Oxigraph Graph DB). -3. **Contribute** a new modality (e.g., *audio* embeddings) and submit a pull request. -4. **Join** the community Slack/Discord channel for design reviews and governance discussions. - -Together we can **re‑imagine** how disparate knowledge sources collaborate—without forcing a monolithic consistency model, and while honoring the messy, evolving nature of human understanding. - ---- - -## References - -1. Jewell, J. *VeriSimDB: A Tiny Core for Universal Federated Knowledge* (2025). GitHub repository: https://github.com/hyperpolymath/verisimdb -2. **CRDTs and Convergent Replicated Data Types** – Shapiro, M. et al., 2011. -3. **Sactify – Tamper‑Evident Signing for Distributed Systems** – PHP RFC (2023). -4. **Oxigraph – RDF Graph Database** – https://oxigraph.org (2024). -5. **HNSW Library for Approximate Nearest Neighbor Search** – Malkov, Y., Yashunin, D., 2020. -6. **Tantivy – Full‑Text Search Engine** – https://github.com/tewksbury-commercial/tantivy (2023). -7. **Merkle Tree Versioning for Auditable Provenance** – Zhang, L. et al., 2022. -8. **Zero‑Trust Architecture** – NIST SP 800‑207 (2020). -9. **Ethics of Knowledge Drift** – Jewell, J., 2024. *Philosophy of Science Review*, 39(2). - ---- - -*Prepared by Jonathan Jewell – Hyper‑Polymath (Neurosymbolic AI, Distributed Systems, Ethics).* - ---- diff --git a/verisimdb/WHITEPAPER.md.invariants.md b/verisimdb/WHITEPAPER.md.invariants.md deleted file mode 100644 index f5586d9d..00000000 --- a/verisimdb/WHITEPAPER.md.invariants.md +++ /dev/null @@ -1,54 +0,0 @@ -# Invariant Path Scan: WHITEPAPER.md - -## Invariant: ip-c815621f018fa4a1 - -⚠️ **ISSUE DETECTED / 🔍 REVIEW REQUIRED** - -**Source Text:** Neither can reconcile the need for *controlled drift* (where some updates - -**Target Text:** propagate, others must stay local) nor enforce *fine‑grained ethical constraints* on who can read or write a particular knowledge artifact - -**Invariant Type:** normative_bridge - -**Notes:** auto-generated heuristic suggestion; editable - ---- -## Invariant: ip-e410eefcc307b214 - -⚠️ **ISSUE DETECTED / 🔍 REVIEW REQUIRED** - -**Source Text:** - **Mechanism**: Every request - -**Target Text:** carry a **cryptographic signature** generated by `sactify-php` - -**Invariant Type:** normative_bridge - -**Notes:** auto-generated heuristic suggestion; editable - ---- -## Invariant: ip-138b4c8346f0e9a7 - -⚠️ **ISSUE DETECTED / 🔍 REVIEW REQUIRED** - -**Source Text:** | **Never trust by default** | Every request - -**Target Text:** present a **cryptographic signature** created with the requester’s private key - -**Invariant Type:** normative_bridge - -**Notes:** auto-generated heuristic suggestion; editable - ---- -## Invariant: ip-d14f1cbb7d279f6a - -⚠️ **ISSUE DETECTED / 🔍 REVIEW REQUIRED** - -**Source Text:** - Proposals to alter a Octad - -**Target Text:** carry a quorum of signatures - -**Invariant Type:** normative_bridge - -**Notes:** auto-generated heuristic suggestion; editable - ---- diff --git a/verisimdb/WHITEPAPER.pdf b/verisimdb/WHITEPAPER.pdf deleted file mode 100644 index 897aaf3dc2564ff30036a861d0e1470d8f762d47..0000000000000000000000000000000000000000 GIT binary patch literal 0 HcmV?d00001 literal 68082 zcma&MQ*bU^(6$-dwzXp?JGPxXv2EM7ZQHhO+ctOXdB3SSn7?KYzO!|-s=8NKue-0l z$rVJz=$Pm^V94+DQj1~Oh!~0N3@u@Jco@Vites69i5SGJ4V+CxOpNS|O&DZMY|Wg_ ziI_QAIQaNroSYp^3~XTBH)n7+YPVaTeR>UZK~+wBfd+v7pyb=bu-4Wc6$%oZUh1BZ z)6c*B1}MRfIE@QS;q0`=9!#f`CYbF_)X{#ugChPqMY_#;SS`2wI=-KUxTwna)#{a= zyp6f9Im`yLs@-av@h)ZR^=WtCK7HZS{>G-tYKnFFKhN!H|2TNKo#A)=yxajrwvYDx ze!$<7uhxY|$T`_%_u}yJy>@_2OM-Cx2|jU#TGkW{FqW*FqA(V8kChW0O@uEi^t3lR-UWT6ThJfm(#yFgHL`w9e3x8{%`QO^BlmSc3!Xc?)TfP z{2l+#Gxf-`Mf=x38=%|SbkE*~d%*KjSnFfo_S3*|cz<`{U>wR%`EyLyMZs3xTG?wF zVXmW}EGW59_dqu4v`~JqX23s=xdkR5J`1E~>klNBc|+gCKHB=dW&*X9FYF)Q0GgXG z?6&d7Tog2tA~k+rPgS1eNLK`BlyMyz%1HeTcj3(T_%j1%yce3H^Bpo zsa9rei1b~Mkl*j?agJ~@Xq~qBklbcQ9Ocky31&_&qPdTB6Yk!~cc2&UV&@DH-7NBz zRqydyCQs=5F?9yA*%c97d)Rco7e4Oz;S8lkMk;O<42i06$%J>NI=<80h~7ZcLE3ZP zoi_U~=<6VJ!{D?7%n>@-bU?vu6KM|ymyteAUYxi#<1&yqW%^Duo1GVD3$MXaNfMzZ zsK{Q-+$9+=VAXpA0Afs0L|yj#-?MirLA0j#JH}Yky8Z2S9dKin%w&XygUSn|nU`}M zblM*Y352R@1t=~$5Ab@9Wi5jmYXK!S)u6&+%}EAfhUTKwZ@oW3HPMOY33+mb6qph2@eM@- z?uXMw_j)tk)L!E!gU6?X1H-g4MS+G!7&3*926=EEJGW^lBB4WV58pW2K4V(kkED6DD9f31A*uLEbi!#B%LT z&3BrX#mNwa>fpi+o~d?~H!L_DfBCaGc%Pkv=s0q77Fi*&M>ZE;{NNVKBR6b8#H=$o z8gE?yw-=PY^_{-8=)Y@MfRX^WTj-gXGsjACR;)0Ni#jw|^7nemlj^{F;Cex{$M+PVq^eD2W*G%HpfuY_y~m}(4Clm-ZW^hM z%{jt*8xXE610%joiWccGT7)*qjNOGin_VJ6Ayi8t6k>hjSa^oE>#p5i@JF3;OpFjo`UESDtpvePqBP)J{ zjmS6%BtlGmcOuViIC>OR=ImUhz?G{y|AuFX{&I2{W{N;QbaI`=rP^Vlm6DwhKvAnq zSeOE0A@s0@2Bm;`zu0hU`!UV(B2j-S`43tRhEx1U17tvXoGa*ujMpt&ZrY$dSs*_h zv=DJQr7YPHl@m$$`qMTa)3N%`BJ-kHJALpa39CY30$=F0(L;YgWUd|d0=G-3ObtRj zSvU^|rBL|WshoTGM~dDeLuWsUBWedwTXR^docQ#{?H0kPTI0StbCwWXdS(_oQPTNF zVQzEQ>v`i!rwNq|FljJjTkWMh(i_6#y#B0-1)2VVOL4RrN^_{GbEwg$ zq+M#_3AOf&EQF>EEo8MoZ7*fKS{eLyihciceyEBln6?0*!?#Xs{SJM2;&aXvW?MTO z0rSURvGA-<0ryn5L~ofEl1}{bNx1H#3s$#bC!Y4OPikF}D9s?#UVy%YUNvi7MV&zP zc9U=imQ2BgTqYZiWCZo!w$c$Q@3F94qf;kCEBdz!4@1Y`5IWfGM5nv6icy9Cn%4E69%Da6H8P2AJ$Sr3}yjCtp92+9S#z=o9)9&SI z{3XDKv#k+-SfPngu5t$p{`4*j3=<~Ck^`kCXbt60Q^+{kJf`q9IE z77Ws?V{Q(^K!|m{`Byn}hBVV&hum3d$lb*CZ8RLy)^SWUM5UQ3(kTC2NT5oI2`NnW zJ&!K}D*1=3?Q(-#)u3%q{E`F(=bvGKYW~owR zqpX;XylWdv2I3lVG|cU)@uKote#|WaHU=P5(W0?W6@k&LS=tU(iUsApdQPPAoZTvYIZVk(AjD_GVY3_Ne9mNWS1 z53z9Xc&QNxVS;rKii)D49}*@0QvDtkINF2*NP1iBQDYfZJ9zG(16ju`Af8A1YxTwE z7-_3(N(~J*L+;GMaW$@ui9;8l`v(0L%I;P9TeoaJuFcre_4hZDwRTRJ-rWvvoB8}| z0DOb3v{o}*4(4QNr`3p=HWlIbsRJK= z`4NdCWLl1v1I%@3;(SQe^VyjvFJ~rfm53{%{D<&?8iI9+Mi@bZs}OYO6E2RTo_45jEa= zl`pK)8XUu<>~hP{JGilb$4nO5Re)zJvutMU&@PtJ{jnT@0tFn7j=HPIiq>JwCM=O z?AM}f7Sp)M=;M?Q4({IOUNW@VTEuvJdIm--(PDbZ@{}2L(QzXm5FPH;%gs5jqlu$| zMsC<56hE)iXNUU5{Wm@bPS?G=gfHhe-oQQe1=sj{_VCbBoV+?p{Fzuq*z&JTu7OKK z*K?wICwJ|UVM=gfb2$%E*SKd9q0%rnK9h}n7B|x>7p&xR?znUAO{D0g@ce^pk7#~D z)x=p&J|`VH8+ndqgCnl&9<{l?UVh7^} zK}IS%;H6v39X&X zS;4GM0Z$?4oQlf(_{hK5*txw=ndag&*;wBn$V5@xt-k=MgMJFxKeO6Q!)^S%Uf(aD zMX>NZ>x}eVC?JZ}*aRrD-4J6{0;WutVXoxO53PxpWxugsm?@5She+8o{##s=HpJ&= zll?2A-~@|V_F&46=vvFwUl2hs*YK?yIr@QTV;JCG*{>R37w%G9=JR1(V5q{IB0J%=9Q2Z=ris0y^(k|-qQeIYVk z>XL>V$d#9>)54;BsEsg$9c{C;Eqjww{Awc<3&u(A8l-qJg75R$EHO&v7fECsaf=iT z7pZ6K26WzuqU4luSGox`Vv(C*%E64w$B9KGUe7CPaeBSJl*dvLr-Da^c%fRT*^kaE zS{QvAh14K#L??W~OCKuGRvu_kuOaWyx?^hy%pT|I%i1hP8jn|G%hHms?QNC@Bk=6~ zcmf&V{nJvMA5 zEfQqtvE1C(3S!*v+slrd^@XzKN3Gqi!o*ONnvz6zyN3PLq?sv-&Q0>fq;OhYG{)KM z(vDAU93Hc?(eJ+c&k9K8#N?s^$mRL7?%tHZgO^C1zqE15!8>LNCI{B)mtn;n-fc-uoAeM~3?1psgG_O31_6pMBz zQ7O!$%PjjA8=KT4I6meGv1Jm=IcW{#*K+PpBB1-%bLLwJ#ijarwy3`#+C)LUi>xQm38%)g%!Im14nv)2 z4I9SnuEeECbd02{ z+P^+2{xypdbHkb~X$2JW$|ZtjvtcMu1*s%kH)b`f1Zopgq$VA&WoyM{QV3O4HHq95 zP9kN&J`@jy#WFxkZM!EZs0r6OEebTr$`XS_C$;LOrHOoQ2VJU4 z7!ft_Pwq=ul+0N49nr+Tap$k1fm{Y#PvN`jX^&M&%*PURZ~0qKpETykgYk&BGeNGQ z;J0k-M!I$`tGKp9Q{reFT!j*0r11yQJ{5L~7x7PK*WaNRFQdEYm~Jw~K>;Cw`oS`W zYJ1%hO|q&(d+QP2{k6Dr`PsDwy*{`h_mq1q|13DEB?vwq`M4(c@$>yp27Eru+eR_u zwR4VaYj*K&NlYV~pa`U)Xe%by>E%$K&D|3&8N4q;Fjjv)ZfD&5{Q3Y0TTHEVGLVn0 zedo5UD|4=wbwkUahw5Dhxx=9`bT7qS;(SigK^&;Ds^WYjBpeOpjUj~+Qra31@n@!J z7$NZfOYVKv(OS1qu+H|@Gp~G9W9LqRl9rH0O~-i8TwWaWq^9bY-cXltCpceCDte_k z^O$Jc9=fXS~ zv39RfHq~#aFIewGGcJP8ioI|!+IzYXn)xutf#mE9=O81JTl+6~PM0+-rfc41n2KP) zXt)tB>27Y9jh-$qTrLOSobtH2L|Aj@hv&pU8x&T1uXbIF;u~6ll?%464dC-f9VNuG z#~w40ON|x8m0}#00=egre}9%fP`eKx+BWNx%o>7btSF1oP+wI45%*_#nbr+=3Cy7-?4tnB?Y=K!{o#Lu$`F7sxZwrr!ZvR!yxx8%jdK>p zkrI=r1VUa+zrN&DcgS;mm`;y14D5PR$X&036Tz2vI$Uu5YboQO3u|Z^*3HGd^Q5vbTRk7vJD@xLZJ`P&H zv~p$0nHfw;<%K`NtqZ8+ZJQ0-)uN`J^vB2>r*zu)_RoH;4``cWl+Oe_Hh6 zk0^5aazDWY=bfVTkZXP03qyk<@Lcl!fi-inp8Wqx@W1~n3I1O{VEcd5?Hl|p+w0bc z9lZm>^1-cp5HJWOGJz~GbGcY~xrBVgm(5vanC$=PHUhcgA)I8yw#zOKsfGt#Y;%w&U6rr2m?eMYw&y_2aU(c8KGa04B@2B4bJ@2oRp6=J#YgZrN z?!;3t5Gok3znbemDK48k!r%7K%P>8DpWlZwuM0}6n)tVqtsOf(SYqJc6hso;%p0Zi zcj2KggRg&CI=)(b`rlh7?=~7hwLRNBH*Q;Bx9BL^i$bse*!83c186xi&hS=0fS#MZ z-&=tc+&lOdt@lcM#%~ab92G1WJoU_J;!vg@*>q@HA@3Ltc2lJdFq0#vAkOq&9 zp0|AG_9_(hZ5}u>^5pmxDsL)Wscv2$Z~WnHh@apv&9!2vYtB&UH9q(A7}UJ)b_6yt z$@41*$DDi!v$CK(8z^#iI}B)CqYYFf$s;uF*ZaO%3d{R>=um$#{{A?p&MFX13y@G z^+u*(kNC8a(c8#Kd}V2ISmZ}TjAV->!oZTy%C;h-GEj|&<|ncUNW;H> zl-rWgqis`WSyL&V4E~djlx%XhzDkZ6s3v^44@@J8PgmEi7*8pEMt#prkeaPkmnt&j%BZPW$4x^W1yiV}Y<4cT>rfe#mD5PZ7P~H$Zt;+>wWnptGkA#e`#wFp`}tA#oEI$S zp3PR*2@bLTRRj%a-o5|MgW_4w)YI4WWqb(Lv@X|K#!m}FLgDpZvQyraL!pHNVA74Y zbmc3OGl|Qj)<>PGTCtBVdye+--h4@fxW2A4v*grh2L%e#(*Q)>AOefga&aP`PLutU zy`=W1bi^>L7)OEVbf(w#Zw3xhwwU?q-16>zc+X-3(GfOkyZg5Xsk#$B_ip+2=W=r^ zPgvs1aRpd^EHR_z^7wszo{oB!kqeQ0QLh$*2E2$Q6mo~mBUAx*E?mA!CxmK^^FeG~ z?gW+wreWU!rD}M|o%VAY5z5j7GCkfmyK)p7^<7;L7k4cQsj-XYiun0*J1)y~?n{L= z`xX$Z7=v=8Sy#5-Jc@_9V@Okbnm~m(u8tG^ z_fi9daDLt^>_C9f(tM#Ddx&x7@)JZbHH8U1+ibG_c4YNjv9z~5vD)|`C0C(M(*Rpn z?l7%5B`>^4mZ19*dT?;$Y+vfH^)=?HC?x|$ZI`Nb=Q~>&s@M&24j2%{eo`R zbEiUZy4oyik8P|~%D@fqdkq5CkY=^Xb7G#wN_F4mqWQsiv69iP0#`)J6oA@7Ft(Fz z^khd`8^4+&eC*hqtXp~3%Fjp-41!;^!jSCr> zs{~f^vXsQy9V5A^u#X@VF6lS$8e1Jhx2O_XM((Oxe3*JoNRKvZl+1d{Fr%KuS}wD8 z^H-nhUhSL;I#(S|*{j>6X85(}q!@8-XsjruE3kBF8+aI1jB>Xvo($@)ELqN3I?_V$ zuU}rUvj|e>g@~9RGTZF%zI3DJ$L{Z-sjnUG@K1 zy$%Op0qfIuadQ~r##|^!Tg@*I37PGp&1&Jg$!w}x7L$l~FyvcJx=##Sf83^x`Tg8q zlpO+CV%CL1P3iPobl>>v3~y>BOVq<;hc*MaZl#0UC>CZv6x|+(7&)P7HZSdSDv#ZO z8*@1quav<{TxPQK-oBeU%3}zsi(nrIAJ^1v)_?hHKh3dmqS7|?6{jhcn>rz*=bmzz z(x>#4swA==`}7EBA}x({Chq=vuL8=ztUa0K%R^4#cuDaxCbHX5FfJaWuMH#g>J$v zkK~)n6A-w8GjAYcsfOJZV;hc0+x<7ZG8W+0`yc5_3nZ0v%Cg}LjPg3pO3dC#GyMZ+M^glo059)@e-<@3jigVQ3%l6T?1b_ zb9A;~K^@`;zzX`LCE9;6)4QN+sjB*e?cCW?o^!=s;;koi%1g}+ahrGg)c#l5SbxErRUX4vdrZH*(+>vkN(a zH9JpiHnwu@SOA?yEN`|o02M;5qba8Hm$SWd#S}ZJg9tB*oLqIQw6VUw4B&S7)laNp z<9qErYtCAAlFbz_+dYDbC(32DMJ2l^N5Jd&=k%{b4ia!JLoq8O+0Lm(x+*p6P9!<& zJ_x~wdDRE74X=kT^cVC=dSY`wGdu1YvW*hQ)|@M8Dk`Xj`|!B5?AW*5qH>#aFwR8$ zmB5^e{Rqgfyo5q%g3rZ4tQsXX9)@jDPUAI=!eel9;)QrGWpy0V*&J%8L%{+Vf8YaXfD=!cra@5X<1^MsWD3NMN;|IPzf)ZxAs)wewYGGWr zEKn2`W}wU3K^)febT2Ekyi9BVn~E62G~V5>-op38Sy9z^cy}9n?>atRlB0np>ZLR| zXFjI_(R!*M0EWNqK!6{{*;*JT4Lm>=p5uA78B7uZtu}0Dp=XQ1^U;bFIEDq+#$X>U zyl0HwxgWF5Z;l!mX+%Zz#AIv57dsqe15`mG(_6zSM=YZvr*_;B_CHyS)J3X;HY9YE zDR05)S&^A%c~@HTY?>{yVm$W9oxyn#n_upo((K@Fp7HMCUww2l!xvbWHf#`wOIm?z zoAIu$lNH_abY9L(jD7D@@AD2RF<;{YkfOa=%b9AAFYB>GO^hDIX42cWim+~Bb>et> zqhUvb5Bf68_bh>Jp4~2g5^*&U{~R#l zqy&LDAw^?aAc^Tp$36Ct+Eju_;qYpS)?|p;`Knb7tez&&y0T#)khuox`}ydu)Eb)c zD5sD7>_d3hZ~eNky(hSJu}@}`wc@|RMH@wn2zqvNFrO&OSwb;%;#|21XDM;wW`Cnr z$e(n|fO+&;L_FT0Q+8YWm8%a~^Wy1eB*MC`R zmyCSysuERoYsAD4{;LNJyd#ANH2dVmN?~fe5Vh*9Hcm>It(toadV=bK+*JlllUE9B zy78~%c{eOs{}(Dn7!>21^xeY&QK_Eyb!cys`O2M4%`9cTNt6d}^jhn-<;@GY_xEzc%vp`sv$xDki6d z%AxLzHreAQU6<1s!Zx|71hH9Xrc+H|kxLAl;;+Unhtq5JSoDPmqe*x^@l;(f-tB4+ zjQS=G>P{hbOmP2`uDSrwb{Y|}wA3F|EB;}c2G&Rsi5B-|o|Rdj64I*7@4R}FZ$6FCno9e6(|3?)f16lryIY1f!N4eV_Fz#FD^=l6EGMRN(H{+ykFq%=v(9e z@6gZ0^uL_?|Nb9~w;OHgxZ_sj-nqK-rU3JIU!c+dS-dsC4`1axfFU@ZFUWl+!Xusi zg8r!#jR4R}sl0MN5HnrpI!k-6K32sgQv&~f2@d)NCI~|LLUOWmn2gM9(YWRM~g`Ka~eSfq>`P!7c{k%to`SEpgPYGuG zzH`8dLY)mp`2Kvr?$x3i)-!j;kA7d#;3R1gD!mh@5OjJ!+ykGE#Qw7Ei8yJ-)Lm&< z*lbt)9JQsZTMf;>+WJ}f`e}{M?Vsm9&t3U7Stv>TRfnImVr~2P>F9Xhqu$408s^t- zur&#>nUo|n{9J*rVqZ z?ZJP37#D}g%st&wP!35PlY08}_I|K4b{4ny`*c7yv!9K+m_xvT==Z7AAHNIxhV;?% z_IdHJ)k6GdnkZ&jGa#CGrw%oJPl~MO%-)rp4 z0^8}*&8F>~gMM#1t|IMfiNHtOFUj&Xp(CxMVH?{z^}!-_?N6q|`=;(*>v_vp$svX| zH#e1Caf2~-@6X3Wm0jMY38U$)WpyTuEhB~436~nHyp`_c)U#2sQ_dJaKV9vV$#v(Y zMb_NO@R|b$J7)lt_yZ~DzL{|8-ujbJB)Z{UaxtxFnT2ah2xgFSs~4ylT_yy*rkI;Z7Ef)T zG#60OjDCzurJreDrzwQ1A=D-}xtCcfI0{2U?4~L2bJ|rN2wr8gaZ`DE^iv$e$I&?eCwPOzRjQ5IAs!hcJGP5%M*^|IVp430Hn&RMxKMQvfeXe4|Hgd` z%mkNzk#{2rF^1L?K?x7%ZM(cQ7u_q!_~Nn1J+pO_9#7wU!=84@RTn295#eO4`==65 zk1@7VaDnL-c}j5=gV#&%dX_7sJl+SAsnSKY+y)zFJax>moEgnQ-drLMZ0V$=ZE&9c z6E9t`Ypzqt5HMUHwKsMyYUCu~C? zSw?WrZw0uUx7YE`!SJuwzDIzWi$P1q+onJo zNLRw)>vrZzU<6k)Q>@#|&x&r+>1Jfzi-085WR9$m`e|P!vrQ}n2l=_<=zwrObKyz+`slkzE4+!;% zEb=Ef1q^AwI<_=XMZy%K%Q4*waNPbCwAx&wfn=WvenL<;Xr>9t#rO*-OrfC!TGCV%@z`p| zwz?g^6D*M^6T8{YR|od_DjfYNATFqPjWI2 zDn*jQ21puFRd=FcC;QcK0wfK9c%>IuH#^u<47D2)sgLvd+>OQ|0*Y1WuJ`rX z0$QZRT8c2S=Im58C;NSi>+Q4*W79GbR^VV%iQT1$O2O*f4j0+H3GPBkr|IzwJolkW zBrG%{ROCau$re3stV?k^w9|tg?5)LKD#AU)i@#|LO2WY%Dm3=BJ5#&(}to@8rwr`|sK)6Jjb>D3?B?a>aGWIV)Z&PGY|D#;)rQKp8Q?$NY_NUPk zWwW$lkP@D|IFKUDjbuK$J=5`AG=NscWX;O4^x!eh!e0NtKQ`#=DdI3@9jsRH(gBEI z!81UB*d*X=FA7EJoP5~)<@FOig&zpOW=@X>v5sM*w1p$u~9%*EX09idtxAUpVe z3|)q?>wO^-*H3sUxR|-rS2j!u+J56(;~+O-L8VlHtIoRDC8Gc~SOWO`to53~s&k`Q z%2UC;sU3R~sM$h;Q3(XBy&YJ;%x&>ZX{|*~KWO|!3V2erg_A6E=-KkT0ionJK#*2; z;}1`vmw23&RE7+<55X}4TT&)ZSeh8g3M~OuzNl7wo%w~eCN@a4OxGuQ1%=2A0vu%+ zKqrKZ6GaJHKk_3eE5WFi0<@@*GwqO=#!j8AEO??t61pc9n+QKrPQHReBzQ+$yF0716I9XuS79O^c?W4@JA*b2oJREN@Jrl55 z@l|h*h^>b#lpb|I1N%8+6L_t><=dKc8lyR99hpBd6s1SxPCINW@AhE>N2kBWK^n%V zh+)ABvBv(OI=e&`z;_~K!7g-Q913Hxuwjt63fS@$;(EviWOU^cbG<`a$Uhp<0L-7f zVLQhwa!UMvI#Jrnnx#d(GEcq3zJ>hwTOZ?FwMT_AVV1o5c5|N-426KO$ ziBJ$TY_a8NKWFU%i{rM4!MG1YGqb;;e7-FQ%QdcbTzPM=w!1#b*4`zA!x)Ary|D(K zK8Gf#Sx(uNY>Lk&NC!LtJJMbsA%D}}ksV(wGe7bpa z;NePQponi8O0x*{UcU3T`tS?mEJQPSTdZc(2CMX<9lUkq+q?!IL_{+$$}V(VFIKq5beC`rxGuivq5^`l3tAI{S*s%SguGKLNLahZ`vU{g z=vKFi%a%zGcE0+TT!?|qz;7ha+sM#8{R^m%Y!TV^173gqeL^iW4Q(A(KKyiVxK}ZqsEUIQ?N9@OP{?zDNOuPj%3u~=b zE1EoZt1Fg^-esKgDw>yV(SqNE1g*a~y3q4$h*+wA3|3mM*%z^G6!d$*!5WRPinyTX zt!d6jvgmXkiwu#2MoMWUvE5WH)7Vmh;icD8xVQNDK^macgG`ZOJ-Oqo`?OyIArQpd zx{Wi@s`Yhdm%*jg)Rym|dx-tfhwZxfNuLF{Yk{Yp>&tTgylC$as77vdvFTP;0sakm zdv65P&GOhWIlbP>%3O$h5ijvK-^kR8EUOrxmx2ndorejVcQdDasI@^5rAsa_-S^?wH5 zl^W9|K=tL=EX6%BK{a`^rgdYbdR!f!R37@5PsM@X#07B)UJ$n)f37cg1@Xtl(#ec- z$5%);L`_P=0bUHgHNxiQReOIlJ8F{Mo)A z+|oi7s0$fAzuy!>4k|YH9ATWZ?@!pwq*{&)$0TgKP9FzP*ek=c-vmM-?Ns{?y_6RE z>`#v0m=H-A;kD;h%O7{@BQ<*!v(L3>er3%^6({=NeI`Foo{L6@D|5{b;M2PxUVdMm zp7)=$*S((43+wIo2iK2JHK_}z-{!y#YLmmDK-xU_j;yzZ4L)ZR?We~x|3en5IJbFW(qShqj$uie_% zeb%6!u_(F9scwB+Veg-aOH#c)JhF1OqO}BgMKJUklC^iayWjVZTTQOQP=XfubAa)Hnc0LaOSAW zdHRF3H*IdC>+v}7Y0b6lgg>Z;il1}lttZbRQEOa10Je`3H3}q6XX@fiyS#SMCgRuk znKP|1K>_#%v^h7SGAkuX72>DfM!}-itTThLxFn<8d_`pduM0KsS z$kr?WX6ycnXFWepX0=!g;Uk`LeDDRLfyjnVZlAJd-i@(u9(S>*Om;k@x1H}~t3}LWi4O8AUek20Nl=tqfUsO3-(7cdV=6}= z5+^AJ8dG4`5EXSyk>x(&qHIkE}=eST>|GPrBt1TCrU36Zr zZ%Ux7J9dA@ZdaICPMS^_boG5-X+_1Rxm{axN5*WY!sav^dAC~~aqH#pC8Gd&NYX*J z^A^m+JGmU1Ooj4)3Dq~sY!k@{ea`W$oem5s+qG{LA<>Wdh(UWJ%Bw)z7+WH8+X}*U zAm#NoJ(ALfK|wxVFg^zRw***n=cGS(k^XNFRa$)-tW!j~oBbiIru6OlZxwMi(EfV^ zgc@vo9_l(kAekMQ{<1welc9{#e`2gw5yAhr!|4|yvS|pE;A0S;=n;eUOiD=NPa5z- zulwWE)6c~>qM6MYa*9+`=YJ^Ihy`L`QFNvK#}X`WTe=Ej$7XH+-@d+KhMX0wU(@)h zC?Sw~G6HdE_c9CyW*TQ%2Z>dZ7&wqYz0r)LBOG2adI!zl{5*uqwV9SnE@u>d_etg-JzT-bW!fYCP z21>*nRsU^>wc!Y)h_bmyPENzO+SMuv^&{8>4@M|9t)TW=Ig_eM4^d1_X7&y+-CDTc zGR9$`TDoh1R#IgoBNgWcu~*i$V$tK60hNnZLj5_NcAyCSm?gqU(7 zL&YMMR_)=PBcc7suPTQGkF-S4lt0U7kYo0yOU2`z({tMjDNc;u6%fw04X&(T^&LIu zxPa2^!9BG$9Xx4B;rjSE#DFL{i0_Dy^R8bgeE=xCunfj>Lznb^g*XLyXbnW$w9Dk) z)QxP->2nyGiI<$@C*vQPP3|oPD zbuyD7KmE8w{Wet3i!b!%2-2kgR+Dn46W!N=n9&vOKPGUh#Xe|Y^tA^TmEpU%{uMy3 zprb|5Cd|tBZeYvy&ouN~V;^A4FsEH&8gGVTVKQ-MMBNmWSlqeh6FzM3x@kYGoU|TV zsyHoJp&WI1O^>lu^}$bftiuU@Zraf8G=%#Kp0Lw;mBpi|R}$a3TqMA@qJTi~Fj0~> zjDyxgKd`ng>>qO?q%C+OJQ|{xHwvf(Bls$iJp_>hsd9Wa+T4AoFT~v$0=IGEFeo!J zPqO<*JgeT#3ws5~diNN>W`%U7S2q3wC%sLpZFv1X-)b!_66Hyi^kSH`aTfbypDwCf z$xlECG?b7YK78FVpG(6-Gotkdn}J)#)JYhCk8*F#onaetG?_>g!ZdMM5#@Ums=u9K zf|Wk$$_AQ>m?Tg(AXLJBDjn8L3$h4@L$v9Ayl6Y>3u05sZp7A8(~64Ed55Xmt>^zW z9ui^PwnSwuZy`%V&})0lIGB#T_zB2+>c~!*UxK-skdd#yu0Z#a0O>56ct?L~uo68I zoIqI>W46MUG@#c$Nltdt@iFSYQ%it{c+#ix#LD-qsj_h^)*|B(dO(~; zgP$1X)7qZeJ$D-Do!3Q?&ZhQR&**``4Lx?6wICJ@juoup1d%oN7yNdFiXn>>PhA5k zvS7(F^PsO$C+VW$-lL+N3|CRqMTzOty8UpUs=o#V;jc&;qWWXu&~Y@(}_(4JY88z*U&e1-rwbTK* zLx9cs8rB(#-@!uh;l*0zD z&>pR@74v3Ex>*8d;uh!#mZr^oB{^s{20HFpg75oB+5g4ZI|YXlb=$(RZD+@}ZQHhO z+sTe?+qP}n-m&eQ@BIJ6J$28+sp{vh*$r!Q1xbq z?d^o;H_ti@k+F4tx>QJO`W>E-#D+7){zP0ccM6JV_oSkgWk4Q8ZFUOLvBG* z{zSqCamLqH%?lgG#u@8*YTi6}y(VZQbbiI;qYh4n zx2jXz=mhFN$!#G4^;p@mXMh4?;cI?;;+s3@<*ZI?;!gQK z&r`!={#*x)q1|@7mc%JI=kb96N1*z!ZS!A){Z-6F6Vr|(C9Z2XC@O=wMM56wwK;2~ z1@maNFwlfz{GMIt_O=Qkv{-<%Ln}KuWh)cxRI7db!-}m0e)ueb1*eUM$m?)S5l>bu zVdDBdCmY;ohwG~lr#qTH-FYxPQS%lUCV3oZxCwq-6ao(8k=#shwGObf^i)2E5KTI4J=eTJy`}0^h_fzOj z_9q0>34e7)!qOa*nz}GKOI$^hHTU3hn1DurTBQB?yO^zBG5=|I17h*lJ8)&y<5tV( zEa-NJ_JG-OLGT1_WoEM!$87e0b0UG~kk;kxYu!6qmz-_I@E-vW&|CF#-(6-`ynCUg zGFQ;N3S#DSR@GeUxkfz=4+x^zgV>0>!6WLqgf+253{q*lZ`Fe`x@$qXSvG?fpuo=k z>a&&jV;=94V%+|3ux`wk1$@$)&nuNUS< z#u0HC`x9CYv9HaZvYP96w9iREN>vDLvBTG!z)^lhkc_W*mjCMmeW{~${m%73m^NXt zU9DW$XfyiSPgPfs)IaL!kjdD^8-GQOTB!2V`N1s+(dS3VZgRFbW!2S$@a&9E5xw3O zU}Wy;u0dFOo`a>ZqLtd#Ry}<70r4sfX@(HJuG}noyNZv}>nERnUwx`XgzZ+D{fT)3 zcfY{r#W8M4!M0D6l3y5fhafAy|# z`c|;*&2VyYOdQou?&iGGjaIg6)kaWD3B z*07VE4-G>;d+Vv8pLU^l&?WdedN;Tl#^=jEg(rfYd%HpbVO8_``Th1`=jRU3CgqJ4 zIE48O1+!+QxS~}EZJy3a!hB5s&FG7^HRj?ntR)-Q*+d<%HI=bK@x@nd|@SXR^m7)>(0yWk#<<4sF(r2olihC9^$%p0zDSlwAXOs{0 zlM29-Yl5T)zOTeMux~nyef*=5QfpLBx6nk+sde1Bt(&CWR!p&9 zU7Aw28a?E_q`(sN^nEWoAt-TmQp5mW;)4vm++d5Cih}vD3l+?X#k>%*spP>@yWv;= z7JV8|ZLg&ko}KXYPWm}1A#Nj*1)m71=*cF&d&yi;8XYK}=*c!t-n}rIhT{hz+3xTj zJ%H`eBSlq$I}St-kQ1zAn(M0g)B!V@7@6qL9X22Vixy7k%$V#9#dg|gF5u${WasHb z+wkY_f+_7^+a264Z5y)tuSMyv7*<7wz|rff%98a3QO(5c$>*H`PAga>i^51t1i+

y7;Zd(yH0|-gxr3`)mU?bUtpgPFm$abQj87;&78QM?@PV-FWc!Z5ClGf@6_X>$S0I)(B+e6A7no7B3;6XXy4S z>yP4Ri$crXs_LfG0(y?8_Rz-DCR`VBHXg(Fc512jCG~dBS~tt^&H__RI?HAAWX{c< zaQ2X9ekP@CjxyIm-sUz3v0{Rql|4dtSkhM+<%}f+>*&494y60)Hg)Rwr1e<@=U2-~ zS(;`I5Hk8U%IsM<*`8FyI*jc-ARLY4%pivvQpEAYY|=i50$)-)>VKFm3c_V`6DeXA zI9fT+)@32a;Y~z1qxm+`N~xhsa{cCOe_c#3m~@ zVw7E-lV7wjHQ80kQj!u_Ot7~^pt0h^lf*Cvla+>M1PcM$#DS$rBsy&l7c;>biumVU zD!or?k=<}uGr~SpkH~W!s%jV_9=_?yMrcG{COPW|i8NLq0f@9snRPThlFUFXE*d0m zt|Z%%t{=}e1w+A|2YPhN$V4?WcO$A(4nh!M^r9RB2~llBZ2Y1frO<8|BZI3{*SYuX z?K?R!)!gA>qNLM;P^n4K`u!wCaGdced#XB+a(TtsmN@_|{EGfpAxRAoD z6KX(JJ7sO`nswHh-_>!dB>Ou7_lo`8g z10Y9RSm2w=Xu^r@lwj7%tQC`C91opr;m(v?zMO8(st+zbw{mq7dkU_Z7b$0hVA7;an=;AEx7Wn^)ga?>LE0iRPnA zOlxl$z9_>2m!L0j01uRVjHA_&{oe0!7h_7vD3w}hHB+;{7#vQq{7ccaW!bo5FJ#I_ z0KiF-8I+0DkheCU3%R4Nx3=ff3_LX;QAYv5P>H zs+%NFy)Nd%R8(DXHn%gMoxY{l$y_b0na_N@%=K9#vb3>XF6^*C%=k(+)QXgdZ!Tyn z4WW}zYol1q`bkR&)HmE)6G5(on^=n2dY?g|(r=Y`fDo+cGRsJfq>|f9>lWJDC!kF| zBQ0MvC7pClPoK%6s2;6)25FzOw!aP^eoEKkb^F(Vq`y${5By9+1nw&| z|BF&!2$oE2hWhIjFNF`r&YV|?B>0A^kKd89(W<4cl_NYB^r#R_Nqp+r zuN3P>W)ubp2YV{0yF}H;d$f&OePGFv2ZIhID+~1j0re}KWXGpHtkz@3S4Nux)kr2b+Mkv6{mvmrQ9w`jx-zskqIf; z5n@hf=(JG#tyJ&AW{n57j?vc|P9gr+eI&3pqax$FYzilY?cO^zK;qLkWOcMr`JYTB zPEZDBTG0TqZ=SiyuU`DH8|)uc4ZK$SpDuWd=H!4GL=RZ=uIRrw|Mh?%LSv|VJUB1m z$$+zjCHL_9fbh@|yhCAC(&)e2B6xLF7&9!n${Q)trqS-dk)TlaZJZ7f5%qRcg7~Ml-Fq!6?#|6izj{+ys%)YgtiaN`sv_ibSkB1M0GT zOri2sKaQACW`3>*>SHDEj@T-hft)_kvY6fql0!A`DlGy!W2U?=Q58vul9UPHf)58_ z(yaC^pu-IDTVyH%WFrvqTDJzrBvD=bzDUQI5HNn1q(oOBI494qSRVFNN!S2%ElkJ& z^Q8niIoWj;ZUvYWLsE9mvq&mh4AA?I%d#r;KSy*$y`w?M+IZTFccjA5&%&0nL$OOy0{)z!czXQfrai~PT$AI)=AVm|hU4}y+a(hdp2 z>q{}=NLmmvhC+=b$lA?p3-o8&5I0I)ixz8h?+`ez;VF3kSq4&r{L<{e#LgCBf5v(3 zhw(C@N>mWlvJH8h!-Z@I4^g@EivMe^E0J6N)<2jf)CQiMoD!O$R2I!PW7(A>rl~yo zpSrCc6{&`Ji*_oDH7K`~`w`}#!wv6(a&dkJ{Ct` zcE8>{YS+qGN)12m23WL+hG9?W{pas&Rkp0_jTCB~_Mz-8PQMhJi|P(-n3u%%TRG~t!y)na{KJT!TT~! z6i#;wT0W>aDot`Ueh8Q(>>fh}Q9!?)go#3z7T~i&NbKPwBDfYx(!T%}k6gF`4`3|J zzPKP>EOucSgVXrzq5qs;KFw&hA!`)ey{4_AVVF*uWRx=oD+H8gOr&j3T`8P8e{m~; z7^L)_Emr5)@L`~i=k=mGPAV&)hZwiTPOx!ZPBvMm)KJ2;9Vl!D!f6f~GgG=X*=qag z$?XMw)_EE;!7R{3rB1IHuq4}r06VrH3C9;??Nq9(_MY%szFV@y@xuxx=7-zIhNn_h zb-2?O^ZrEDOtP_F95}g^jhgwQvcgX&0hTFx+T)5$HC5@!d11beIM1T0z!c6vPP_0M z6QbS((-7eZLh{^gQ*pRCujFik8#ifbEJ~2b#EXxHa1~k zfV0u28AiTB&(ap1=0i#}FP^nJi>Ax1&u=TwG-GpE=GpdwX@o8^1m6rh<-f~tEe?%9 zMQ>VBT9DoltO4GaBx!5Z@a7d>zus7f8)N{?Xs32n&iJWpnx;mFuYBwXmL*&c-dzo# zQdvJpSg$=#u~-M`Sj?joMOVQuYuUL(&xbQeEdb^MycyS6Rw2UpEyprcH=v$ra-;S& z4Aky=qNA*aI#q(MTsF(^E+a6pb%7+&^>FX#!wl z+ZsGN;+|m~6M6QkBYp)zU-Lh51p0{(kva^0_ z8ed@`lxdeVfHiQAI<{4E$o|(&W8YB$Tsdo19vdc*PNow=5*b}+ZC!nBl_Ek`4bvPb z#=a@o&T-J+8kj3j<*qOlAhbe6vnzk&#xnuwFcIQz-my;9vYeAiiF;A8a_t3Y8Ah&$ zw0rs@XSr)F?J&mPvLroHT1Mk=WuzEw)J$D<;q(0l|BTPA@IN5>|6R5d6ASZyL-fCd^HYr);qHC*-)#oB;RqX4e;06iyOW z_au%s^5+Wkms`tad?=kdR-)VG_72{@%X}q-A+(eA?A^(1ro}l&*E91vINs-Li_v=T z?cz1S{G!T)(*91RMs7n0}1 z!PCR@eN&~V-l2LezVv#U9Xt$k`2#ySC#z*OCcgV~xEO~2qZ?~{#h&0)C&SnCc=tZ@ z_D+@bGmD>VgiFsCnb{U zc#N}1|B#3FjcK9W==&MfAI_$yG?zWTsOkrO9Vf^*Z2G0|M- zHsP1ds()X+#utR=T>r2*smIHBRH2oX3jOQWuJLphMcz2eTMhj;3+8lQ3z^>yi2!+-l+^xVF-?L`=G3=yB z=&bY+vd{ZuevI~1@!fiN$|yjbLV-XU3DmZQK)(Zr5r=m*4miTy6?pwb{?sv&W>m=8 z^PEkLYD}+3UW0nPPIVFNX5nH2Zq4Pg%1f0*U^Y3(skPwJS+YSd>y+DG z6c3oV?IhVHB1IFlS*Ik)tTD~77}0H+f*~|oH?xHFELcORz>7=`bR=%}x>36-yO_Bk zN(}QgbcR7p@2ZxTNr4=6T_;1!)HwbIq-mQlGo1x0wZnOt-Gc# zH2#2Iu?k-8BaX@2jbA=p|911y((iB*Xt-kqS)ORDHZ;3_(m{lHuLgZX`gZ3u@qb6TsfVfL4(DO z!VGvKVur9o3FJ@(J4OtVM!{+aZqx4yDRx!BX<*wDTjuqvOAs6=e|JbZ>?_ZV&(<82 zRTMO7^z?p3p}O8ewT&pn8K6n`#+&hoZxI^TD30LcStu~!wC$#t0=Wo-nxfBdJ}b#m zxvUoTA}Ym3jJl+Ernd&5LD&FEnGg)1$&0$OYA(0!4VLcWQJ9s3f9v0`==TY?pols+ zcBZ@F!IlXBZNoql27`{9+l{9jlOtgZASpS{pnU`~bl{}&Q2-Jw_+zPXrbku}L2vkf zsUaieauBN(%>D^^dpWAOPZJ8CZEs<th&Qt~u$kw8~^xMY?OBn`8sm|DriI2vU0!_KV@>q<;S_S;lE zGD8zYbCyDokbI z!U&2;q;1;Gu}zJa_nkeVZR>xqgOc|;2A(I7M_k?(Aa-Vsc+eD9r2<}q5(Am@HP zhMyM~#2wlMW>vsm8SXq$ucSVPM8EyIT*(Af(X56TdL0{Kt_EnC`gX-(Te*L}c=)&~ zAB$wAC$!>NtCI>w9#+KNL~VrRu&UPlL*E}7Yj6lLTNdFTJI)8p!>QcJtJ4|q@CPR- z`9Nl1fZ6<^uO*L^N*q==^BQx=^t9aDjs*)g^PD+yVB=VlW+@YotdRxPY8ktQbh@j$ zSK4)wBhDmO0e%1y7vVTvK2~UdFz*Eh49;cP`gDkE=o?~7>4sDuxw@iG?{dD4eVXJk z3DtbLMWlp1H~@wrr_=$Eu4*=;KRivm04{K61PytvP3hYDRKP|kp;oB6)@ouHePwau ze`_1WZ_fto4@fR+`kfIMYzd0X^vuTjp>78wq;?psnPng-Z3Dd4D8RiAzX#A*D~N|Y zL($iz0b-elPNx*>^pGN~a0ZTH)_=$Sr*%P1<-bR;cMFs0=;5f&~Zma8Ne=&3B=0+Ghi0~6url3iABYPR>F-lYLMrE zZ9CE&M`K7{an4UQ99(v zB3HSudB~?VBEbBQq+cvtehyLTxi!MNa)dJj@wJQ)783Lb$pKG+A%cK#umFki&JTIrP4u$G8iUv za&ngOaNIm16r7PCzOD2FJw6iET};g&-3yQ~Rh{(9Z;Z9Mkyu5+6#(a`D|BaZr*m~S z`PYtEEw2tfjgRTOa$lsB%*1WeBx*}vC~Z}o@?K?YBS_n5X{Vhw0Oe&Moh3Xx--lz# zgIrX4hic>MYTFWO$9%q*HG-4PRbyGyBe?O1{BEJ`clHT*KP}pESmRl}^Q*jY?^Q;z zR4dM^l8{FMEJZ0K(dj7`w1A;iU+A6bQiJR*DI$TgzhyS%PJTUD4Rwi=8D>#TY39w! zxV)U*R6uL06-32j0Gg*A>U@PWwjR#K9u9Tsy%eTPu(+-%b%=up263~HI^XKQIelHz zbou^Q&(2p!H>P0@)n2?(L$XwC&nkq->rcAtMK_&+r_lXStivfd461U2hsJF$OAPr| z&UvQZn!-@yq9=FPJ*@}=>M2!H?xrR;H7ygcP2%!KDZ*#HYt#>$jAkh(g5;r@aL6gR-cU*| zSz44Ub9RKq14%Hn!ndn=vm`}MviDjY){|n}ublNSC%8ujeaJS0S@?pP8{n9f!Ma`y z(=i?@;lLwc&dDNz+C;X@oD5-Ed$?A1{cbv5I~s7G7!y+Cg^d&Q` z3EiJvkUtatQmD_w%M?Pp%NbIA6w1q_r|iLPQPf>O@h~b2`N3pGvDqY@>%eG`I?wTn zo%p~iFP*GbJwAs5Pz8O1Op;wp9_2l09*qm!%G=E>+hh@vpvRwZP+o@`+~f@EIVwf- zl<=TLN?!G2fp(8BJ$Zyy`A?vH_5QJ~Bo(iUdiphI7R~}PC2ds542=aUu-)=BL37O0 z2$Ui9s z;SJjWfi&%+EB~Di>C9G6u7i7v7X+P@{A~f56v^}E&`fk!mal-zLU4X@#-ns>U?yzK z^t*kGxy1=~=TrKZf^u{a)h-;p67Md>bLt3T$OyTH>t3f&`&lO!MdW%shLRS!MEuAF zCpwr|$9&~^wBdciPMzF&PfmJ-UoI(>`B!SE?5V3#t5i1PKHl%f2vG9J#Ws0wWt$uS z66C=$->f{P^V&(NH*s3Q2r8L_iaVJ=i)UjbqNde3I8+?Vr#Aecx$1?|_>Q9 zT$J2oVg7LXVA7T1V|$Yc(ZWV^AfdUd?QV_O&F9Zt!2IOr*-|x zeIv*)W}#hGrYOE@`MyUAZvM4euuW<`9NS>Z^%r}!Tn~+2FK8XD5TsZqgl!cDwP(`O7v_ zb&9g}487#-Vo;R$B~2`-IVkh7!EG~aHnVp3`|wO z&Wrto_}k6g-ikXRt5?f=HMeh1?p@@)O8nlKpsahJgU=`LPL9|bn{O{19-gj`)#v@` ze@}y^&WB9QnDc;6$#X<9_&L1ZKX==zr!$(H(3#>)rOem2`GgLm79XF>&JJ99xb(Uhc#GQLnAFKGIh9E|)tN>z^Su=lV5#A?;zEJ`&I#EbTu_oy~hLy*Ibc;EjY zU6MUPk}Q#9sDJ6zLbM!H8KDw;BKtq|t%4u(<38zpPQipaBO+VOb9+_jV&e5@*S!Ri z$XB5z6G@VB3^s!%te(wa?P#sRP}oTr#O~)A8dW@g(!XI!WRTR z_itZq5>Mrkqocrvy(vP_jGJw^(k59$Liifrcpa+ua8M)@^u%$YN2}A5QjX5L1>p=d z@6k~#V|a~yj$XKqiQJ6lv;$Dkso23Kkv#fISezhTv@Y^d?4s`s?mJh-97F19a**ns zf6BG3p!!#|kisiDahZ*)jHhjs0{>2wC;O&dM#8A*^f7E7clhKNouTBBoflIN8C~fD zZPI=(-Pf}(D3V)}e&>;=?#iuOm3Ab-7~NHC)*4?qvNGI7wGUGFXDFW-P`&k)OCe<2 zFJy*%xuTFIxMFnfYgMHcc|A9PoT!voSYT@T8jiI1ig;8tW%YE3zP8`=8g?NRASfR6K*J~@Otqo z>&9RHGzmb*6)Wi~ZkXODXh!y9cRaeu=FGuSl-xH^n=5^dq&Rj6LW?ptr)$8PW#Y@>9fv0w3m zu+MV(HWft+9pxV}&Qdnai5=6>uWyIOrx8_G*{>c*M$E`yvv;MbZxQE0c_#-dpzKY; zYdaR9k72YOMqAQRpZU8ki+vNU(aD5LO3zyA+G-2=id-tiH}mR#palv%#C`!>33ytb zv5qBZw!JBpRy~1)R4$Bb3Po=XAVd)S)}kDJ=bS|x9|}n~itrlFboJGQ7q3YQ?CJ(f zxn9QfsFevu`KF$iP;W_2n`wOo39p{A$TRplf%5cY?d%1mq#G;8I_QB4h|L%)r^=(m zIM=hRIZF{#B{a_g&|v&D82?~UAFwrlZ|v`tzbzS1(h$Kr z-Z!>vFL!@@=6;r09S=WO%M&{NSZb{MIx-o%zY<>mK2AUiZ?yQ@tl5AD$Wg+68Xw#+DbuKVj>A9^0&@VozA(1g`v8B5>KQ#}v`GZ85|eQM zNXIgg88%+NtbGP&Apzp5e&MIHu{rITHQseX3qhO=6Y(B(aVjIb2-xJRi_E$$240K~d0pRtx&`gapVSz)@h__`a2ytyLjzKlY_T<#$kb3O-4Su+2QdIi zWRx3fN>buKb@@(|!TpQXn+lpsn5_tu0y^-AX$c{o_I^*80_40OeB~_X9_aU`KV|sK$U=^^Wc)Iivt}6 zS^ep9WSsz|5|>K~x*=v}d#h&jA2F;<{LX>&3yJA9u=U0Mf-oRgaee*2?w5A~BQ}_0 zba%JE-2=S3oO>ehRSj`*IVFG$v6}CwT$;pnoO?iO>+P|)1wayVF!)(J;B^JKjQK8` z!jaJoj1k*cglYR;6Kc3!7OX`l8D&;84RYIpn8;ruBFuR{KK{|ASWBwHazfb+^Ns(U zN1v9BE{dY-I}B7H#waQqlP7tIa2+*?^&9x++#7ujVCL?~3nRv~TbQh%_Dd^d&1N zmn6C8m~$2I>P`0bSr<;g4J?G<=>{2vKQ_5o>8ZG_JRWj5fFcA0V^g0Z7Ne5Xb@a)?R>xnlqR7s0ZLE%OOSH~muNE~ z@+#6;v%gg2KLt!rYiWs#itD5_V1^8*|}p1X+})$xr4YzxdqUrT_%( zU0r$4)$+EsM8(xrFRjj!jdeI{hfJjgky~_9=kD&P1BiMl8`?&ZV~8q*W>(9F``Qp|Lw-Q|P!}%+LHR1ApQX%rV9z z%JJ?z5bBTn*JO$*TQj>17>hbzV8NJWYQR*qt<5Y@;&%5=paR>fJbNop53Ob?5ma6} zTs5)DzDn@XC-NrV)~bOyXt~)vsY)`ww*%HSZ7KEJ(+b9vn)X{96??V2Sg{D_c0*!_ zDrPuHza*E;ZX<)&fd6!Drf$ouepn=bDB$mvz@vRf#~;dRP)<09Pt zzAq&CUbJ^Ii}jtFg{L^d(6;!rI9LmB8SeY;`bMuq%STWuU}a*>?tvPrufzE7@r>Z3 z-|V=%hVJKrAnd6@k1h;nRkh+`*3a|xWiYc5Z|N-chhHrTiXi(x{O0~9?0@l_BfdY+ zh)?6y|G)M1>ueu)Z20*Hc{iB+*ZP_~p@}Yg74{RNzYv^%Ve9v6eN`HfvUi?&KYss_ z)|B-O->^!KGhaT|Y|2{lY)69C@#{}gnVobzPFu}%t&&EvvJpg)YWr+}^djjmf~>4v zEs#)p7|mqV4J>q8pDgSYc_|K8uk-_2%L zw}+4nTU`8w@vTtT{(kQY{bDBg%Y1TjdbQC3M?rVGbEIxenZB# zp1#%Ee*V5hKG0P342Kkh8&Rcx&|GG-Ej@P2`1uUsvaP_}cCqw0VO-nA;M-b}eQrHE znp}so-M+U$2yzzfAEskVS@FXb%y7if`n(-4qU-hAbu#x1dJ%i+29w%&^$F(Ltm7c5 zaj3~5^r(MNh;z%CIiHVZ_V)O_Z13h~zGMkex{33ooApEyo+F&8sfYSu9s4hh(@x~s z9#drRwNA(8jVx!Ii`(!YY@4kOX#!7NW`1aiGr^rbqq@yIVBG^gVMnNEv~Tr}CU$AI ztQgo!1PtPly`4x3P&WVmeynUXYPyqK-Auol>;^G<$0tk05A{0=B-e}iWxq>XAuE%d zr^O_#R3l4s`cap8r*Z9N73o7>AC+SRrCho|Ye_w#;;wUoHOR~}$yphga4p?I-)OKa zqfyss!F7?`;G&-T{9&v6Y}`NjJusIkazB+S|61FEu9_EK4pHA zograE#07v0_k(`-+gapI@B60Gg;OIXubQ<2i^E&B3SduvYzXd2uKm>zG2 zc;iwgTUlOI1=?GISw$DGet1JqKM9yrn?P+I#nrZcM!5JGmGz?nm z6v=_6{{zV-le}O#;6=(>O@>vzLz>=D zI@-wf2>yqANZ|qV?yB_$zow;u8C@0?jdTkZfyl05Z--rB#w|bCPt{n4$AyC(K*~36 zI8_>iX}GX1cJGKOb;JG<7#0!P57Ft_HhLGvj^RH%vncCYb0fQe-HRM}+Orlzo{R)j z_dBTXkKqWKcsreIRXrp-Zr+&3C>o{!;^{W6tU0 z{~?~GqfIWLv~Cj8KU@|wFv0Ae+BMshOGqGiYpG!$w=l{vXI}}r3neC8`1Y`omCpZe$ST&D&N^)_l5{@RSBmP9^RIPvqCok$U&{m6UHm9c_T&2|FmUV-eQWt7S z*1m;xh84jDIAbD@D0-Nu=olHWS$Bg19?A2G8PbhBcOY*H3yspaqozRSDk~NdSOV-4 zM-jklxUT^(8iE0XsIIG0GgjT?f2HZSk6fUrscxv6wn{4jfoPr3G;>5SjN@NP?08AT zt%U-`q7QA)M7+wQTp6|IXDzDSzTc!90zvkWa@Qb*P5nIKvefD%Ld_hb$~^DulZePy zVPn*!8FceFw;;G}jd;ra4M%NnXOy%XuUy%A9ikuT&pKr<4a*Ki#g{5>21I~Iu9sP_ zOXGT=F_*CvP$3!!(O|<5LsLfioKN{#6jd=3xLY1Qs8H36Wqda^kN7aZHe9QRkR(Ho zBf~mOPcFU;SS$v@>E9*~RPZUue4B=H-GbQcV$PheRxy0Hm2j(G%9v_lKCb9UV$iih zfFMBM4?iM|e0FOKn08j(lYazEI*$~`6&cX?1?tPuoM&KYC##T%#7a13N{7OHp z4nD+UN#{X+`QE`do{WeTuwUsG4ybu20`;ybN%u7>xr|SZI0HFJXYF(shBy?CgdEDX_m6Gb>&FG$9W@yjE0sC#=p;)7 z(16}Yblu1_00~2%J{!2qN^{%@76Y&~Kza@jY=tD1_z`{@Nai6B4$|e~nimW}wL-m? zgKpAZ!lH{ITsc{OSy~gNM+}XUrRv(>F=SY(Xca(J&{k^q`f*U*>u}c$x=Rxx8vQvF z#vaeW17Ro1pCLnmo5i?l2gLFi2dnhG!WXIbJ#s7vl#Fw1<*DPvo>C7lJ4=GWR8=h? zHH{3R3{T@N-;xhru1SCERhiK84R1BI=A(rGFZ`y(6|9XU<2d<(lqV^wrLoC?k$QS5 z1Ok@{2jZoRfGc94JKONR%!cBxGiq7LEI7oV8up!`eMzeY0O37|uUJ^g&N!v6rCUE@ zjls~Zc1W&POQ=oGlvvax&Kt3Y_aJ$A{=9HCAt0Wy9yL1)tfo1lM4bi7U>DE`Yf@yn zCM~6Wob7#t?^N>XqcH1+#ji!*&5+Cf0dfwzPq^r*8gvw&&)3ITyB)vQ5Z2V)aO@?&A5LrOO$z?KL9t z*0_w2`m5^cN$;u=_EEoK?|p>2VbY*#etM}$`q&*y1_@TV4LPcYP{Us*?QOkv>@S(s z2M1GqFzgNim=sH9rbb&h@yx?|Q2nDG!+#LGlq!3N{pMR1DC7`#sa2E zGPg+eHlodSKV=dXsw-NkD#~ZX@0K^GX8>E3q=XX&5>vbHP|?5*2Q`&5jSNQ*w!?Lq%yy{ZFaJm$CO1Mb3=wO*pUo*2fJ9YLN z7mleq2RM$~P6Dz}VDb`#RDX+`FsAlZAJ$i1a$IcWEiDxtZwE36a>SFHZ#88T@N8G+ z({0X>{OU@QI6<^tDZV+exEr9Ag$SfDe7ecD+40)U{m^7(U+~6>9v%)7C0b`l_o;AALsqAR>yv=q0E4yC38F&xNjX0 z%f`ODfmOEQ7ELXhD=!<{k5WLuF&}fqZc=O(vQ<5h(~8CI)4m$S$C)Q06C>tKIrRJb ziO!{UoldL7EGTDp(xGy{;N!%dE4GQQ0TDskLB_;YK)w-6hUfHCX>_ zZLnbIPkE62dH+-Oa*|D1BwB&+$Eft)y{L~tYKsz5W-sEaZPmEkO@ES9gEw5j{%qNG zSVhDhd&pGg2Q=Ev2x_(rV#0CLMz-tK&UN&X+Cu9-Alt zZQEz$9YUX-haz-gGizH4GdYZXjm)5RcT+kQq?sCoSw2Fp#1%zP(;vz06N=du6&0K*K822AFbczw&ekG%5OUv5jN6#xoTBGIkCnbf}`vag-Qb6zt zdm4acG5qr-C*8K-;t%|?c~N`}ZKEq~G4)RUDAt*F&&}2p_`mO-lj~dsZ{`t8<-y0c z53f;C*ze7!=w95Pxr}%xQewCXo^k<9g3!_I5JYO=y|*hqH?&2Y6x+gOi`2{Sj~iFN zE87@vlClycH4?j*CT0vGmuI;)%>5C0Y!w7oUY54G5coCVdP7Pc@ z*!_?8?pxbk@AfbESD6B3`}R9u-$BSo&r>={8nmoGAay6CA7iO>@l_hU4DCNg>|I;k z?3dooYX}fo2deh<5RcDan|eFjaN5z~W9*{sZH{z!FFAAvIayxBj=2xeTx>a$tPcq_ z1rE$ajj9!jO((PPv${LW1rnVnS;$OpgJgd3Z9SkZ5P7MS4N<#d99=dcnl%FIXg5mRjHc!+9f?iJjh2mQ=Sk1=XVg;xLn6L+&qP44Jb|xS^lOnLV~eFC z6K_dCr~;#V487r?t*NCU(^;xEp(w6~<8R{wIII{{EHMx(7&>dzp{-DG0W1V-pn{w= z8nS|GP>1?t+eZI~v3HF1H2Ri=`}DDGtB-Blwr$($W81cE+qP}{cdY3*ZzgkZ=AWBM zR-Wg}{#tAARIOEocWdEJT}T?oV+@2d5D*@;3YHnh-0?3-TG^q(0`-ez(p{!HJR@3F zgQ~9SgtZ;;<*^&wAl36*XW5)6iL~Z5EG!-pxGp1Q#C9A@I6A}pqGc=P9;jg*3v@~M z!Zc#P&7TuUmggA8KiE@eX`DKV$Y5Mp@F^)h>2+>!u`>;^k7UXN(QXIV*UROl18Fae zA>ldP04bpq$261t4@eFJJ|t7T;QpF~q{jeq2O^eX9kGEBXdJ>VTRY;>ls?;DQH8@^(BpD!cAfXbnwL=91(W;+a5d(>rh0J!#RQfFYfA?iSc-a(qq?cr_ObIuDmQnn)L3aYs5irvt(|??;I_h>8#%(QB=@Nv_L3r zO(Z%lc#s6jcpw4^c*S+knpsI(v44+9yvsV)`&IV}aA{li3&_v*0hAvVv{P-e>5hhbpB_qnjJ0bv*NMQ5!`CH4JP7shvV7 zoLu6&E9oa-W3;t~WgOnwU<>tD=260b(jXA5e*?8@M2>K3jid=u`D#MXQzlL1DMEJv zwOl@YIwYdAV}+OH0o2d~Wnl5Kx;>Z8=gu?&$%JZ>3#pps)==A11o1U(3SPG;7iX0W zZwQVJfhwyq*jp7viTBqm*-JE?8i?wkaf%UCCkN10L7qPKW+hfP9#J1|Ic^&G(#j>eu>Ol#=qnxM-1K}_ zN8D)3^-TgJrhQ)W2rR#0<1k3xi!bvPWI=pb3Oc_j(Lv>R$G61wT;C9!Mn+O2%j+2O z7^XwnQK{(2{*vCCtS|TPA^lnmxj|XHKIQK!5laKeNh~=|J@2n)qVwTj+?i=;-zV`r zKzi27@gyTAxMt^TK7mb@^S=N)V!+z0%gJf%n73G>M;`}Bcjx+=9tc?FdSa+j3)7o%Y3OdIZ79Kzb)m$yMFs0pqb2w&*k*uf82Vra_AaP0cgcNgjz;!#=3 z_lL^z{h&6TKf=ug?!p3urMFFJAF;zPh{LaYPx!i#Lmwnvc6!>3wubH{iRp2&VW(am z*6Vix_ezsY5QsNO3nvN^{X>qO(yxe{n+9Go76RiCI_3adL3cTC*PgB_N9{#i$KRI% zS3GDrYlv`ffhB*c4xtCaB0VQ)kIFjs&|RymDKh$Y$bCbd6C}`(+KY9h4(@YWl{0;f zU;Pr){b*zefN_1WSi8%d*q`As!f1K+yLc;ZThIIESjB^CxBUtW4HDbd_S3SNT1=9( zM$uIrLc^jE)}hkzLLXYS-FGV%DI z4{QCq{Cj!mfy2o#U%CFm5ayLBE$!bb4J$>Wkdzk~iW}8c{&h54G7lrraUuEZOpq;Z zHWW!{VN5VicwD4;SOGcQZf%Q5(tH*`b0E&yLQa61f+o}YqtWB^Jag`smB}XqH6(Na zQZj4mIHC!T<5Af>vyW5VxZ;5%Frs<+`&?Xlr6bCnI60*wi@sZ$v~Y}`F+iE1`LIq` zR8m3q$`^%#h;XeDCLT&61yHrHAt=;ARBE6O;}?-<5u9()A2#?sc<({6vaSVtVD)YD zbF%YOqr?3jGc>THa=0kytwr8U)3;%7ttb0O!bU8?(7Qt5JM^%MQMZGDPN!}KS}S?~ zqkc_K$kVXS9>NxrY{6sBYx?d~pxr*pT3-3wrifXq^Ti~=F22kw9)Z!Y=T0ek?w>lCbWo z{XLozlCeB1!Bpq8_9D%=4>R$`kWJ3rkvLhc}Z)Zou7LCEag9_?KoaI#qq(0wHF z7fQofkV~?*Iv)wLeSQBd4~tE^LB_~LQ^(%Y1{J5kLlkS7-aN!aKh1oZu2|HjEntSi z9n)c8cAPb~SEevi=R$;X&SLiH!};;GP}d7t>^%ElWzmI}I1V>>xaWkTbh?*qaRHk_ zxP{YqZ~=iE0ZiIUFGZs(xU#jd;9~I>uvd{OMR`KMd$|Le%!W}R)y`o<32ZCX}nc4PZ^LEXZuy@+6=$oc<-S*!fD*PE>YhwBC@sXq!RduDNOKgx zH~z|E!Xg2ei8uzQ>~6IYQ+Zq`1kDE)D0v5hm-r+pdGR@{S62z%1P=Z)IJx5CZ8uJ7 zeQ|(ry5DH^SZZ>us+_BSiD9_ygEYu zlOng+aqzk$*T&m5S@3-bY6QPVteGrLkA3tw2|WUz39^_N+-p9@=T0<34jTDdo?TLui&*EfMn$x5TFedW zn|2;oeg5sV!lJw51I;mXs8{(>mrO3VePHZKUg&iLbf?y-^WsdCMKSH)dsTrq$2}(s zDWP#(?0D5pg~G)@Y=(kb6^n3OV*h3;x9fd1H&?c|(|WnUzaR2E*p4G5YW)XwS~lit zldKrNs@op=f#8X;tDZ3Jd-VS9&x6w$+X0=`%9K`3LO6JQE@jd<+Za6FcPHb70Qg2D z@SF90>}ZyN_Y}T=`4^UU5Vxx!UhERCAk9-B^79OmPlRiqaGpYl7zS6M$f$<})O_e( zSC66|kE^=LQ!^K{iK9jT8C7Yl9^zudLSn*)S@Bzj=g4zdgjCqEAB|2q^bXVUSSr2A zMh@IJ)dU+3`P8m5blfR2cYKXxDWu2PBtg#b<)pP$n!Jd7yex3xm@tJva(I)a3()qgB%Jcalw+L z;bF+z=xB>riAo~=%-ep>$_HKUCM?ktQGA$L{ntV<$Qan>7o3&>kC-dAW>&mj4kmh5 zw*_pLE{C7~Bt#uvhevs7YLD9DJv#{4;&NmU+=Fpa=RcqEIq;@_f`RdC&cGb(BGE= z1l$A|N#qV{7Zz=JaM3fSN0;w z7GVqN=A9MVG%JK06JZ#0kB8EtFKoSjPEzPHHQb>)&VfW{ zK!N(ga-K)-aMgvU=eV9M6@xC8E5Quu&32Yauwr<)5Fc!_>5cSgu=pcKu#I`2e_8Qo6H1dw?2xJ#6LFm*- z*qPZg9K+$9n6)U|;^AX0urnAMVtY6j!`j$sLEW4`SD|gubCLN@3e4i$FkM&qf`q!8 zl7&;M>61VvF&h_jV!hirO!V!oz{x?I)Seb3Zo$?Q#-92oEbnh-^thrh+VmbuF=oBQ z#5X!!E)HAdOl+Q}hM{%Mb5B}l@J89(Lg(wahzbBO*?zeuRd746(js@Q*SJb*#s%!E zv9m@RHE2D}u7aqxVbhxK^bh+FSBEH#E?e1qv71Ydp!wyx4sGj(Z+l&K}T{-F$OgY>#je#-Z)DP@ktD1oEv#0+GPuaCi8T+j{Kg zQ7+hT(%2y+G#j$9b%`(IuVnshNw_ewRmNpxlv1v1b@jjHYiit?fZVUYhYVa;8s`^- zW_CW{=XI&4(!ZDco$m0!3MV-kpde2Xx?4dGbe`G{piC^!7M?PwXRW#NQidf~5bueeVFf%ZE@ z(x^fcwtBAGo|pQs(zMXE&H%5MG~ zgOvI(naUPqTM7@y&usVI1KwE*+--KPX#4s?XRb^?DKtCRhpxDSD(H|<7_8K1ECGI| z5pswzbwCo74J9y-Wze1{+CbNdG<+C&L9b5h7g>;z|M^C;_Js5!m)T?p1sJ7i0-o^H199$`ShVO0^*`1pdvk&^sq= z+B8LfIE@9Sk)1JtOGJC&hgC3>Tn<*rloO&O3^%crGYIm&V9|5Ghn#A=2CPR{+RHO~e6ZqZTjAmJbn?{IX>8rQ8*o-O=8^Fv+a?f%9B)+(=<4g-N;`8`*l_gLG1 zOcWwScAxWdknTB}AIw)FpJVC0VFbz<2sK7z2!_RN$XeTngNnEor}JzUKFZO&44zNv z(85NQ3s=RdDH?cM%5-j`#vE70KNCE^DAtHd9vO)c{yQ9==Fq<0R+g@dmJ61OgWxU& z;00&)8n%lxuou;s1Xd}a96?P$x_y!@)Fm%a^+K-I7!BOmeK;;7w_a8T{332Z3TZYE z;An&MorU{tt4cY-U}h9VPX*`8S-gk!=w#9Jojhk+%WzSMHsd=k!n3$M;B4jBiQU&09#T=`7|-wU7MQV{)?!=le^vuC2%BpU>s1z$7z* zbw zLSD&>hi1JB%`NZorRT#Oa>&|T?v!a<=e`l*4xd842!6lSU?zF{!FkHCwW!uuR~W0H zU={FpF!$!?ZI+rR*SZ86qqvxEmUP*Xm%l2c==~0^ik(m4OVhmH%>#48dRJr9>K#-i z+5p4Q`!QoH`>8k>_1IBpa9gB{&p_Nlb+4(*kEgj?d|hT$v1a=F^M&}~Y>}_8L4AmN zIZoXg2kQXc`m=otMvn$_+5pjecn}s*M}yaF7$7V93dRbQl`mLP`)wfLdD8`rGWdPY zYL+^?otQeUzuRg*XrDE=G&&-HV!QWZcgOH|D|-6@+{g9I@xl%>`;XEEwvMIQh_WBfuuMT9V&M;Vm?a5uO1U0?69l`t_e zk%07RfEA#J1DVdAD&1oXMGQ8fn-ZtG5!|e)WDv`3@SRC;G#@vmTe7r7jm3yzmiV)s zR47&OGF&YUr)_SUlZ!E0F@P$EJfjaGsS*A9Y`yGmHNoj@+;49I%c=h&_Xdx3gURs7o^ zNRH)&>}k3{ewaT&*d)Z707s^uKVkgfPO!qXK*(|J5s83+y^1&^lptt@dDZlgy|{o9 zhDe<0HgBk4t8a`Iew>VXZ?E?K_IVZK*K6%`)b5cqEuTd}I@zzt0IRRWf+6XiRd#W) zqO@FJcz-ODm$pXj0*xzJGVIjXy8wmL{Hao07Xu)pBotuFT_9*_0=DLfZCKgd+E%-u z^>Zt4xzX~Zp{zrY~5x04r#8 zmndN>#4^|~WWdkn_B{MNtAF4)I4p?2Lqoa2SZMvp)A>9)%FX?f#a=1_I4yXm@yIaa zjDUx2Z__;NLBBFtgAUe}a^-=We7_$#pMbhUD!!rC>Lo)pB^t^o=<1*f6HVADREsEGe_<0<4VKW11GsEPB2Y`_`$ zx+bep&a}5Pmd(LAu2v`%>~l+7qoLdk0Gx-(hVWjawX3K(8Xf$m)1^k^i)P}0w|_BW z?~!o}BHx)hnbGq%L(@%pni%Y`F(BHH$yc2l;ZCEZ0F}m&MlZcg=*ymbU=Bx?Cu}Rs zDH*{{5QeDv^SVtGNn^ay$8Xbaf1h{;g1klP=U9vaxVQC2gj&$*n)EnKmH~-3!tsLm zXu!52o8u6KhT(gk=uw1_Dj*}%dG|&7k|Off@;#y4RqGlp#}qwOtciXditq|Gx)_tS ze=Q9wH#N%%D`3u+nT>_Ytzp6nUb+}B;CKz0G5uW4G5lQg zsDD%G{xA~-{Rjf-%D=rDOqFovJH#Q;omiohb*;`Vu*B!vx^Q~+#u-e${^;|>3PVhA zF~bMx91%wAeEZ{|X0fqgtJDQ6ae=V2yX-;g9>TIxzFjlN@)~TzLyeln?)R)ihUbh8 zQW}EK@dFB7RLfF)M6KpMTk~K`$2WAS>5~+S?^}_1`A}`#1@NUB?>Dqu*0t$AQU3r8 zHvmbYrSPh67)p#LgcUayi{eQ3Xv%%AbHpJci=@rEf1D@P%|&|Eg^$||j1*-ZF-{vt32RgrS1=dX^+X_;Yg!RBraFnUWa8F&BK{&)+rBku4q`K0La zTO6+u2<9Q)DlI%JLQB%tu!rBO9r`$_UM&1Rkglrt1&!2Ey@{A=ed*Qa4|BvLtkE;0 z6Q8boJ4ilX<3ryCH#rO>bOMJT{4nprPtR6s>iU#2!xwBaf>eN=otBf6iJayT4Ld^y zX{Dia@$g^vZ40@suInB>Sn2}30VS+O9^ea^#e#ZBMSJ3vVyyed@QMv4V+|r^{N3HH zK8hZM$o9?VN{BDw^G(sItE18u(p{hNn@NJZoCL|JvuEg#$FAx*h3Ov=T^Fh>LYBzE zDzgpVQL`T91p+3Bo?|cDNu|h&&clNeZr}P2jJpjEv&?!}q}K>1-VWJI4%~vsKzP7e zkEY;(RqM?&{#O*iDJuv)atMb`QwoHWXNvuc*i?pSe@LiVs4MtFfJ&u*lLBO?+Z7+1 zHF7YhrRBFypo|ViYt-R>u!9tct@?h|Z_J?TUz?S8ZY1*uwK39tyY^H31rPy$IMM_> zW!`DNUB$WXUP=IDpWQUw(Y?JAeuLwSTIVub`(g2h!G9xlM2gUAiQ{Myd!7g^4@A5X zCSe#;r^rZ;P~n3^6GJOQQ$uqt!DlEDH{26O!|1{ZnCfLHTjGr}_fJ^6SAXJmd+)>J zW_MoqMpMMhE_O*MBn*`!(!ALcc}h45zL!flL%TGrD;`P0qF9zifH%SvgtnwzAXgy5 zMPiEML5mP5=mtD_2hvElPUEB#2$oVZa=hYxF=o6P&F`D1COCU{Vvr^@-xFX_`+-dx z{pw0rpClS4LElCHHn}AKoE{k1sU9lDd}S3f&x~kxw$|fUg|KN9XXSnB^AI0>Gjxs- z;2hPrfoho;eQ96Y4U``6Bz&ImgE%!`T9ojEr1@pVB%>@0-HkuGp?sAagMcSv#9&9D zHm*XDH-ZBdUiVA2zj5w&Zr|+|J8k&pluw_|zS!O4>(i&fH~y#E=PY&Pm55vfYuFBg z4_OPESQvUFTLgUg<}f@7F{u*iBq?kfokXaFP|^7UB z<7;c4<&L@c0NDEbhIAVxTP3r{L4pqfBk*I`3VcAUKzHDBpb^+9%ntAC)Sj16R>%;X zhvirb-;v&b>^DET$GV&SJ|0PNxh+1v{NL{U-?2ZBZ@C!o4?}k=CQ*&;2yQUt6e?Ys zGZC7l!vzUe9;!`}EVsa~2aMO@lQD((a+Bd!yeu`)V56pdKpksWYa=TN8DJd2 z5{Y}isM7PM19X@}lc_!NnlQ8$UR=c)hnN_8n1k32qHJMCV3zYhxt1a5va>1m)0pR^ zcmxZSVHUwaFmu7cEbu+Ny%}~B?8fq{KQ0U=cD}q!r?Z)SLFh8&Dr&do?3ivxj4WpopI`is4DL_EI zE|n!I$Xvu`+&2R<6qtYlWS9l`>DT^A+*CM!>^t~L;(9qPOi*#elexjosEks->T>bC zSw1|619uD6?6zDgo|+GvihS5`6B510I%r*5?5t+E70k8wv+A9%!9R4IInh0T?oMSp z&`)N)Z>N`N&NpyJD=c#^bsMo~EHBXPc8Q7T(4(`m5S=2184}H@Q)zZ43i9>M_@)Fn zu;Rcz^7M;o?Dq%qDTm4#P`-Cx4uBXJ7>d&*xt1CNmYj0ufv%lbwQC(HBdd|6c=T}9w zeERWg(vGr6w3MT-R6yyPB@5-on(30`sAluR>}jAf+6*#(OQg*UAk)`)V{-fz9pl#! z56a`jT?9J#>Fk))q+V{=uTC&gi=lIGde5sB7Dv#?xsNXH^5Q(!n9y7uDYEb*$0r$R{RpVGifd`Z`LBUm};b5|=RxO_%KMM=Lnmq{{HoYRD5ic=Oz>|8S zVA_&O5Xka*^KkCO6yU5OBHs%?9T>U1Z2t1tjh260h*<#8^?Y28BdJE!Y_;{&n_LGi zYvvJPF%kYHGA8#(>xGJA$lfljnlpW2E``;DK2>SsMB9g#s+TW#Obl_#(X^sm!TJ)E z_VSDD!Tk9}U)yW^I6KZjz?$7;xR560O zYZ=F&O^nq@LV1+F%8mKqm6{updPqv}XTR(=;C7C^ZR8A4@%7VYO+C};1L)B*!A2rN z!5OfYl_ThrDK!wBx4L|54~=0LzFsTp>6U}+JWOx0P;T<}&1vo29sr})QQBgy8N`)R zqg;-iZj0_>MRbV!xuS3kAI3+BWJ1FWVC>|}J22d%*vow4dr3<(Ieb3%gSU3I?YfOd zUPYQJdF%HL%ahS!vU5$S?HZt&7py{G!IyEr8T9(wM4zH%o@tu`cO+it?5~H7qPc#b znE@II`o~h4*lD?);~RnKeY2vE$S^Wi9ON)|5D+0US8T{ENkqL!$Ij`w!!c;&!xM+p z_X(k4pzvTt_fNsqr$a9?(WYsr-yXcMdRgscciNT6X4F!Yt;s!2eLTwLYC7%5h>c;~>TuEDzDjXznjZTnXGP0+90M;itF8379pPeJY4jRv ztU-5Q@7bD?R~A>#p^0s6u`eKG2{L`+l@>;}x#Z>hva=_c^IFs%Z}VHYxT|g({5$k` zG$x(-_-o2>W*JS`D>%$461`D zay||cT;%3cAWNgX4(?JNvHeB5e?^IXb4t2w(l~_)d%ryDF2H5QJT`f~3x2@sQW#B{ zUj*$5mfsvoTP3!z8B_8qo(qBkrX>mcMw$OO)?Q&0 zW-ic{$}}u5DkZD|c351&+%EM@O@Q|LzABPGFO_agKH^9)csx-ZoqDi%A&{bwAe1PX z%&+0bE?38Pt$Qzf_QrbW(NZba$cd@-7tQT+gAxARt|w&lDSj1Y-a)hT+6v@l~CS|d8^#7SygUK3C{|qV(!GWi+Qea zH*BJQhKF|}>lT1&&m>56yp0}Jch-5`64Lp^T-a~Fd;{b&4zdYuz|>(ffb#gVIl_3NQg#=dL%4i&=m4hi++=%O9`>*d^%5C7)5SuAWQ(rF%n8S&SF zB(Yr+C$e_>gfmNwSg_g=?CH`Jjn#>K3=CT)Y-UFE&}9~Euf5QEGgmZh*dpdk`&ru2 z(B4o9q45bKfI*ZZSas!JBDR6)-}tW1rGL?lfwp_W>~fu`vr$^)}{Gh0L+}@6s}9@JCSt=Kk^-d!ruYuiMB!0G#@Zk z8;pO1*WBC)hVBiiPH(J1+2{~dRPop5I-Im5WZM_t%+;s4j z-Ij#^_(>^zgJJT}dl|48%%IzqGC~!C`|95<$KvX>{iY zV|o^#)%!Dqzw?uo-*rP9eYKc*AlJ)0fHEtXCfQGX4o>>|NAb_E6`!=dJr1GG9aCGy z*=*cP{dw1S+Uj4K8LfIe4uXd1xgpuI6B~PL>w5q|>xPiGj`!5LbXCjE zZEZ8_%FsCHhPd9`e3@ZPmVbC$p}zVz?ty-7$NF!PDWYgo<3@EdbKJp}JBJKk7Bhn^Mt^6(XC)d;DR0Zi>J6X@r^ z8yHBzr>Zc>8@>+BzwSkE?+KfdTg$Kd3|3feE%#i+q1@$Mta1 z1a#~tV{?z|Wl~y{2Jr~%FN)1qXBB17Uq3q;Pa%$EGLD4oe`+K4M3&1RKO<%<%QMBZ zx?bzJHsKNaS9c=N##QPt%ZwHKuMMV>DT+p5F-MtW3a!{e&^SiX?R0I*wm^InqBfLH zd#$8iso`#L>=fWw+1GqXv98PVLjzleznmxe{azf+J`9EsMsp3&8EU)Hw88em`Mk4PiRkwT0h7vS<_CQF`RSfCAg z@LGC*%@!=kUnUS4oo&YGG+*zLTi)YR2r9z(@-NNUDI`29;$8rWyKSFe;3{uAVlt~* zvRLeX4>oo;AutS6{4kzHsuZu2=5}egh6dnAzY4d6NDj98yqQ7}JnbV26;QHYd4IYb-4!EoozV9ZeOxFv9pnoUKXTE<=oNQ^;FbkALixqXJZ$tbw4H;-bHt|j;^wrP-?37 z`{o$aH$w`tcp7uZ7fobRW*2dzEo!;gm;UU!Z6fou6Ydi~j*@WKXHd~n(C8_)84Nm> zV3#d=gU^s9@&u!}9sQ{n!yOG<{5AU&Wj-HGoPD@pcIoZPh%a6{hY4-J$ z5{4_A8HnOAm0^=oWWW10uvqluP>~M=|kd(u|+s2(gW%w&SLFtT(8sUCuc=a5jcIR%s&X+ zviZL1zd}6*SFQKu7|#AShUS(N;?LE34*i8Y5+JIQ3S2YvL`Pz>u=rhE$!LDwAbA(c zRJSTp=#ZBkA`4D}Wl4Ef>l4NQ`l>=^7J63=?im?38=?*G>7%El&z;Dg|M_9WX26TH zR*MXfiJ&37&LyrD`*R*1xStrQQ|g1}Vg{PsUA*J*a1<`Nu8O6>;pJ+$u9{J2jXCy) zvg_ba*b@&hTWlg_)mn!^1QkiMGvhh7w7)pzL0~TYbQ52+XOj7W^IK|K9ly6z%ygT{dGcKGyB))MkT)dLxwc{f*#(T(|Ag9n4o0R zn`=YhB-vv63gV_QEC3e<35KwC-fWTd(i-i$u=pAs>Y?1u!sAR)%@pG8xds`AF-+;H zac_jq9J&0ovlK=2cis$FceFKl) z3Bx3ufvJr{KihVBKueT0I+bChh^~`|A9aGq7%*lJ&X&{Zeq{G9WW@X}NA2hLq6B>c zinEdO6tGLKv!B>lI;=t_qa3)LVpH?XHc>wDQhceV)}A~2n2Csrbc#dosTiQI zg-|Kl6f<3kVWv{Dvd^nLZe3L;nkNz3mQ<>iOr7R_u6UN_(71$+60D)!JknWBvm=j% zXo0>z76TgwEjgyrqY(!Ws--P06^+2;X1D=G_*ChX3w+JEVZm}dG{iO84-3ZD>BN@B z%;O@Pp?-@XZ=WuzG4iAur1!&UCtKw^7Q4^P&_n{NY1M~E!a`3c(O)af+kuhdknC>q z2j+T?=3e58=G8H9swWHDTqlYpUTSMP2IJk8x2yAX4>JMsi(?daMN@(RN7E7oz2g6+49%Tpgxe*VWpI9+`~kKlBa`y7QTLIe|2waqmi92hjhtAvDH`jMc>F`>Wk=G{v&&hdHsSz`_qvIl+jh5zdsbdyCCzJ0~RkR1DQLYXLox@ax z0@KNiE;l2ioo>wKiDmNjH0)rv86PGB7q2`zCk+o{u_C5+dQ->5xna(Nhv*W<*BT@c zQzJ3@3j5^=br8Nc@#tHaBA}2sU<7`m&kyYrN`^>lCgxr-jB>;Er*FO9oXE?FWiLAU z+Gi1aVL(g(d?fWcs{S~%Fc0?xIioJUU=SX6e#)RrO6Zne9Vgajy#*1o7{xk;LGkwZ z?F?A9q4cCKbG3*(s(8WU%jH{iSSy1arnG}{2~oi4^a8)fBn|16d`B%$@rlei;)mZ| zgl~Ut&2I{DMv&uvvP;)RF?qXz#ySJxr?v zMR|<;J+(HBzLw&MhtV|;eue6=*PN~;K#-$R$5DB!J7bxEdX1S<-Q<Z`q zL4d9>$%5!+knFD4saP~xoB85_chA-?&*_fSEypRXSpF}ZK5p+b*(g^LbX{dW;0p@1 z+9IfN$GQXQ1^klTOc~XQtW0*mcaOC%t&_L$IoHb|V6ffUk&*lL8~ck3@ZqFwKGob) zVz+2Fi@K%Z1sch34t(aGMOU04BqMbF_xcwN9rXCRqrJek@fB34k~G@FezTKrALo=o zsuXH#M-cf7@3{|onI&o{J287OKEyuS?~TgAlH1DkWwDd1+`Ay$b>G2I^X&*b9*CNc zo-3HWYbYCaB{AjR9-F*dbF0jT)dPoH{L?tQNrlsvGt+)zq9=792_zgLktIIshdkHrKJ8~b8 ztlV1lrEWPk`vj^yo1p^S0Bh{)w*7VAsQolEaC4wjg@8$h2<8%oU>G8d{On8kVcG-5H!NHX*3E?Fq1#sEuAiS(7lqD_p| zyN+`$$0MAHC)u2ciAQ5hnn7zew~C;DrO53_l!rv>yiSSaYL{kEX-PV9^ktKhu8)Uh zgf94Vi2?~yxG+(?`2YERBRAi%JwSc}e`51vg9sq}Pqp;pcP!`qOxyDx0Pf>9O}kD){qEjF++RFxxI%n^psH1XTyYb|s$ z&}ArTww?1K3T!(OF9v%1f-8x1AR_ieL4rr3I}qw`v?ZcM&9Y`dpGcTJ~r0Fm7$8xl0CRm$hVZ7%gy0SZyqC^s=}FV=gv%*x<>Or zIPysIWl3x-?$4|ezN?+Vn_!spE_^5SNc3_g)=s6r*uz4o_~&%C8xG)E;GTm?^C5WW zw4O@(V^hv-h2p68&ycy{aor)(cesy)D%NY4(-2QIx6W{)C|8G|%|IC6qYzH@pSr`Q z0eD;*UA1n%P;cmo^;TXHuCcqZrTkrHJMod-I|RWiEHgQevRH98=Xdg_ELkT%w%hg> z=>`osvaka`DS0xhmdl6AK~=F|&Tc=HY^!Hh@4CilU23we)_)*no<8(wc(q(k*uuq{ zqj%__{aqR}+j)HgZ>H7hRriihE=Mfh#V#8;Bg?Q%mhP~5Ygg%~hI6rEVq-^_aVw`S zD_?9xmsD=FR*plX$lhJD@iOd3&Ih&xKX@x@s@E2cR$2E(E3CX;ds?+Ohg_4iHzO&W z+goDM>~Tbt+5f~3!8C{y1}MC9NFNr(BxBMP^l=uP4gPW^4)Fm{DnqfFK7GS9LN!qb zIiz#=tPL`ZF-##FanpoApRISzTDHw^bk*O7u#c3pSMx2u+I(-f8v&%!R&_H_Rh{91 zaTGL_qDVOWb~*Z|h-euz=Gwts>@Y{N)a*GY&uXLq&jq=XkH%z1Dcw&FYKR`9&^O@} zU$&^GFO8y^DUYIvHyQi5650jX|F}B^4Nsm|uIc{n&ge*g&cf4V_fm{|WOfG{)u_eiH}O)qWb<>%cGZd15I#<5?6cX(nX z@r8Z#FlF2%I0)nj^(fQsF$CmcDYHd_QV%w7l8*u*@DK`%HOIgdcB>7Sq%3|WUsGK= zyWQ6r?$4hqyi5qsm3qA&yImOI$z5hlhR&-CNcmwGL%iuk+>d8k) zDHHu3gPq9j-)CrdjQQdKP7R1tV^VgSl!TVLQ<$W{(1X0LH?c32>pmf0TFl?+wZ52I z7xPFClP6aW%&n=Sr2V{~Xq_b((h(8g43VzN-rV%_T*L)Y-@U(M zwyeTr%4y2-6#4+%gbRlT8780$!G$9q0w$b{?VG7P$M7>?cmDL~p_03R_Jq!kew^=F z80fHl? zYH&0mRtOJ&xo-02`B8w{gyLaMev86+f+QDOGLs2|>lR`0?OEdcFC0H>D~D0a!`irM zfAjP>Ua0AQ(I^&rSD=a+wPE^PQ3is}3#g?7d0n|CYoJ7$c1+-y8#^U?9D!qTPEwK+ z)`@aZ3Ko}vF=s;_Q!>L^3tb+M#RxM)-JlgnmZ1(>pMPQ2=n|4+#TD+=eg2UYhtG=p zuowL$%FrYeC|<-P$S*sFZwL^6E$k_`&w&BA$%J;3%4-(B(Sj&8JcHLw{E!tju}}4^ z!RV4YMS*mR2vWoXR0KRrV&Dhh1v$b-RdmRm=oiInecLselOK09iKmK*jiIH3b%XuX zjDGm9lA{C9zTxgT;{O}za%Zm@%De~fgB$51VPrDi0r}*u3zRXHT64L|-#xpoUKrPixMiVQygqfqPx~sB>ll`{N9gNwU}Rlld`tB1dt+Z*4avU)o=!W#Lg4tS zu6TR{Iq$FK6s>a)=&@beD*{)?cgAJXt2d`()ghy-(5b#%bPV$o_8wL;E8`#2qG{_EnJ7uQ zSjJ83x=rQVwkZf}Y{2sFuGXs4QCCg;q%Um$XF2v{JyHu}?c;e}>M5kUr4iDIX?BjT z2JvM`wEA4!a)zbpOMR!@6H$J*Dv8vt@7Pg7_G{Fz;r;22Gt{^Ju*(rejkM=)*EZ*v z&Mkp8F_d43tTqnJ7zTX=Qx_djyxvGF!5YzoKK;T1YnzcwA&kw}wLG1sk8Q^`c30c? zs^ee>)f!9nN`pcBt31d&y8_@)E|j% zxil`Patk^eSfeyOS+fD%LPMbhQc-4pPE9#18em3Bd@=XW4-8_Xe+-x^57Ju4X&$~U zK9{bUQil=4t3koXifM31t$1Z*@Z;$NKb1m-#&EfRgqYgr7|q8V8r~)!Qx$2*#>4Lj z_1mptygp?;a`r&444CVAPY{BH2#Df@uGHD(J=M_7KwHo@OZpqy>H1)ZpHI0pV5+XH+e@i&mzy1Y>F# zL1Y8Fs7=elez=MCW?_V8QTzAi|5xDQ?zH}TRhwA1xCmC9MFEX{jc`k zDax{K+ZN4;4BNJCJHwHYVcWKC+qN@o+qUhTVVf`3xv!MF*V(J=b3WeZn;)Z;Ip&x@ zee~LT?X|a7TT&v6c#g*gVFoSA5>Uq0Mhro~ly6dJl|vX4rdIydY#2I=e0wp_MsJXB z=9t4lh`y(xs-QdY_<g_h6XJ$!YZQ!t z8XqENi+s&Bm$q4{HmWn~6l z`aGR-u~klfm*dCsb9>;3NM|jh-f#wv(&PSuk+t|0F74`)$WpU%tD2AjS_RBV?VEK~Co=W3SQ9c1R_f8-;l{1S?JMm~MZ9qYKm zmnFr^6WqF3iRCFj?_L`V(!E?pXjk{G1EJcsZ1Z>&_sp{u`Y9W3FXhD>o}es_UN`tI zY1tV5u)ynR%<$(`pHjI_*~7?1H1N=s&hjYe4Xr}3qzVr`?HDhTmG34wgSx$>YZc8X`M>uIunKhXjC4_iEa3@frlHrN5ZfN-Q^O-k&H>l9xEez?x>gQT@O`JL?$*2CC2>LM875W)%um}5 z-N*XNzImpR){}~<;pj}WO5;XL2b1E)x~+LodrsTsIZk@(9@FzoD8`5LnuViyT+scx zVty0m(2R$R^#*V;)}Y6X!^_Sozhp(8P7p&wKb`g6u!@51vxB!1d2)XjLJG{)DCP8F z-rX?Zx&;pNIA3C>mzRUHpC0DtSoSXaEm{(jLgec@^xkK?diPp!$ zB^m=YKjKbf4b1Mp6#w5!@d9f!4y(~PS;jAHV2wgPdlN0XYUL<0FghgZHz}H{-+dx4KkB#1| zJe=rNs8W|7#_OQjo;1@HOlmw6iDHw9q%+YZG^Iq2N0v+K>ckA7cI|BnEjWg-vFO>! zRBm08+QY-JQ-K=Jwer9@%n}oi-0itGB5Q23GCrRZAdVM144>^6CV37jB?78tATytK z2ZDLWFP5e?h&yip)yTnXXlM*BXwAo>i+l+PXdkv)ep=i$cAt)Qs_ka(n(k8AjahhG z-Ga|xFEF!dj>ROwZ4%YVuqAVS;G*+!Lz1w47hCW+s}HjIlt0l;rD{Myom*UDX}Vk# zRmTS5HrrXZ4%gWZ{T?u8wt`mX=4=e}a+|CfsJVN_0V>&X7A$7t_5HAR_r%4|a1gMM z^k#HeF=EA+Qsr8mK)2Yz_|qGh?jHqqA^W_89|gcM|l2oJ`+u@VgUcw_Q*?vWOnET)HOf-STaua%#nmy!F zhl!4O+1S17Adck0PD8lL>epe!W#SN$_2=uIgCS;QEgH@^rb;R-&-*!5nZu+l4akFn z^V0A9Hs0Wo=y?yCJ($N256sBvN?Z|dY)21ieu~J?){L`074fu#Gy@wbMd<1Tl6#NP z=14uUk2nFQn#q&KzMdys+p)mis%}IqXbW*Pi3^5qj$;6cfMjkl|Ta&L5P6oX4SpZr7Mr${||UU+#Wc zmI&)q)E*o%Q$FmgF5kN0lX45n%RN|++f*JYlB#^bKDd^L$P6`X6eB5>)z})-X zrmIzHE8T&l1G!yvu^%=-Om6PUT&?MXk8Dzw>D~-_*Y@=)qR-=A+>N^|I*FR0NJLT{ z^y+*!9X$x9t4Iwucl0wkid;2jpMH&xaggO~T;Vbeq;MM`EDN<+k>XBznG>lPIC(94 ziu+=8Z4)hl*uTy_I<+}_tVMES<6xc2U${G?o0@sLGG5UwUW-cn?Kgi1HK3M z+9k9?VBBTph}tf{qP$bGVsPV+R8jcn_WY3gaX4m}Ya@b3a6^CGwU=1D^T+h{}!y`P5Xfg2>a}1-_eb^ z{<6~&Ax7m_hWRqzz<~ggs|U~^yzhz{*i*K>rGHdUxw)l2&?5dY({|u1l-T=F>~%r( zFPDDP#pzK?RntKA1%3wWJQlw!3xuK%S1U}BaQlse`EQV0TduA#`3T0iY~RC+uv4{C znQ|J|)6e?OxhzLBJ&3l#E(aCGEesn#j63eAv=tMM6XV58&$m%0`S+y`-6n=BrscS9D;037T*pqrv3V~FVXp|GblM`qRGlR>2p(7 z$SDZelXghK;yT_8{U>v0+amBj`cI_e9c>V+H>@!wwjG9%u-kDA!TZOUw%F0nrZ@ti zI`Zr5^cZ!su}A4VCM(qLESHWJWM}&kKHt=@C(WDEu*%#U9Mo6G3+En!X{O=bNv@-&XDPb#Bs6vK!oENz};%guEmQrYBiji0X>l)It#&>GSOHt1D1Xma`I9LyZCTJgmPw|$ zcbC+fQ>^66yL>EfB5rSzowBOlR-`aUc)n z2?Z|j+Pm(dA47_LxNwV8+j;y_-OOKouG<<}k|t#os1#CbOl+Y{kBr@IxX@iZ+C3hD z9^7-QDEFlg0v=H=`(Z*7oS4IAqO&!853eHWpEp$VTrm-=I%iL1We17MIEY;#t7T)% zo*GsBrMAIQ`0Y1XJbXg9&#rDP_=Gp*Z|53!@-JzPCI1oyNzoJwP8Ld>oUACyE0mq@ z1%3Fbuaql%-LRtudwCXa{?g60wW6xPrD0vtF58*4JO9$ndD{5aZClJ=$rEeAnAq>% z>azj+F$t2B<0SzR&b%`4so1}Klyllfs`Rp&;ZXQlWuEpsm&_^I4TYp0>A|v!6bc01_L7L+BmOLo{`RQY8RoosOKwzaFzus zx$zUY>?Vb&NkxYQfe904gJ{u{I9fc7-}`aIW^dl5OAeF)B^r`8VML-LW|fK-rSB!& z&a=}7bFRb{O}Z~;fY2E%Lz3R-;i=AviCvBYyM*@nyEOe@-JusuM%n~2Hq>js5wk0p zhgMCfT@GXQSWdH2S!`5wow3Uu&5j?T9F{3A#dko?TU2z`P76S(6$t;kcrN@|dh(?4 zitqva!JCKlAB4fLpZfR%K0c?bwWhf1{sAv+|9}?`#{WSbWM*Ug+W`}oo2H`JoYwo+ zgf9_Ok1a+3OuybtF2e)@E2Fk6H3&Q%EWrpV6ciOgfiHRdPEZ}ewZP$=D`>!w_)Lo^ zZ0dk`9X9*yh&fB{>_7$huUn;Qn4zcfl(+Zu$1?BxHXRoyCnqH*r*0tNuTw^n5C{_2 zyQ9(s`*nJ6Us@~i&f!QX9INI*Yj{G%fycgF%I|N{hb(4%C8q}1N-`%BM(p}2u}hGF#vdnn1&mpe}=XUpXJ0fUsQ0-pF0 zmlt6U%uGqB*+7A=iHP~$oTRL5Wb_{Q5Q1lR9h1?mjA>?Q+^9jUGLS~y@{ zKSL3stc=V`k1=U7e)YTSY{ifuKnsS&3z(2E)dtueYJVi5ADoSx{nYk;6|JDXtc3t( zXG#`EY7M)Dgm|J$Ewf}yTO$Ry_|o2(3qe*jmiF)4HJxqM^Pv40LOap4FL&HX{+<0> z7(c?~?zd-Sf4?kId&v_yzHfR406jyOHx7}AI?2qg3?Nm6AhcIn85Bt+Ecw$&5yjzG zI_zosyvtK!+|!>EH1dz0WP4aMK^DK5Djoq|@NwWM8s=@BYG*BYF1ISp>uYHex6~5dtnz>^EToAw-}?jNoZRLn1!%1k5K-Ttsme zB8n=?syqd9jbCgUQCOAH(d;KZl5seLbJze>33gBO5olpau!6HVlGsSXA0!~W$P&i! zB=I6!1U7B%ZG(Bs&rqAP-5(>bP?#SuHYjPg*fu8PI{WbW+Mn0(9iJw-1c;ufMBC`E zBWMn|q0pl&%akt!tD5Mc{IeOSdkQ>Ea}jrJ%ZIAOrlTZG5ffz4;xMB+eb-;@YzoUt}aY9r(q28f~iLEJ*H7~{qA(? zce}d6WzJJNH^e4DU81NqFS{+<`nvxTjAcG990Ar zwWVw()Y$;n)Yr!g3tLiNCe7}tDSf^w#a4Wdm+6bA3rc>mY1G2~LpS3X@o4xFf2H)g z6N_e*({e*!e26R+CI?#5EWK8zM7h_H3xRLD_PfW$-cY9{R;jXil5I2UORaIXB${JW z*sdryve--?@DG=keP)l<#o;Q?9QxiZ0yPk0OtA*~W!3#5^zw&#)V33Kehv66Wl}KS zm&24-7jZzrx|2OnTROqt~`M7d-(sgXm4= z$0^VNZ;h6-{+Wddb^x^npBYE0H?PCvk{!03JzSLs@Xv=K*7b2#>kDllICK8ascrBV zKk3s)aOP4?fy^1Z^K$9ECqM2x&3e{48Zz&tJ10Ww6#6Ygr-6^&Qpw8on}}20Uht?A zK1Dw5bXG#bTua%#)Pup*(xuC6F74JQT^l(!41?_>TDuy`ht_V8glY9CZDjX!PcaFw zz_2{QjIilw4;)G%(_>d{Q^daTa12VSJ-#V|BK_a00jpskM>IdcM^I+v2pbZU1*j&d zjjGwxq*-bay>fSI!#qT0e3y{i_o9&otSw-ro=g1Rk-fjXlR(k>O)mT7lgXhyJvDju zoWg?oy%zFuNDp!Y@m%FwZO6l~8s*Q+gil+(J$XEQFIrlhQq5-^#Bm1jokAPgy1d7a zaXFW?JdHihd^YxjbLX!6@zxYRRM5oYj&Ge?dSIyk5w0^$ZYM}+HQEjTQ{aP^=Ytc; zEqhu5Z(C@tlHVZLzDDW!?K?8qB`=4^WA1szs@_AWR*5u&ky_F)H>AhojQIodE2U4r z?E@M~KZ<)OE*75;r@@B2CuqGb=q6)6!_$H>6?!rq)sYU2@s?H?r`$c^1bBkGtf9sps<29!>Z{Yg1CV2|x`vSIQ$n#=e!{xab;5WA6;n@egI6-lo;hWgw9-WWczJY3 zqj-r}JkQg2Qp22b`f|JGU&H@-X( z5;=rMCI_n+oa&B@Tmq%{t1s%cR*noa%H zF7p%?R&{p!+a@}od&e#-gt;@2`ZCDPUo0-5z%C^f6r4kSe}e@6=>2isoqgqiX06F= zo|`gg>;*4E#^ilX;y1K=@3}ynli}R)_%mV+8Mpyg z4IN9+oW9*}Xi{0zS-9dcw1_S4Whdmkd0Q4WOK$oO80wg!J~;hCeL&@6KR?y0{b78V zz)`P3_?3!)M%EUJJkyO~=i7_6;|~auFL{PGQVtnq zC3O|={~CANtxwqYb}V`luvNsZy2Pl8@fgS3&!ry7WCTt9qhub9TZJSgDzwB_dl(GVU_BX5tO09t%bFLtO3@#mgN)abB!#^QefjMUjpq0&7|I#Ro71DI zrbgD5vSVeYhqz{vr;K}=6RLqW+tjrK0(*4`wXxCmS&|Kou7|!}wXC9+IQT*LJIvsd z8t4f9RF55R7_2w#n*Aj2S22_fy$rp9yREIG?VbJI93*jX8tT-e3&IMjDH!=;E=b0F zhG9_tyJ?Ej-gwaY2|Df>O@TGKv8|Q9oH@==96G;FJ6iSKV_#+-YxYreao>vhNwITR ze+a13IQw|1#@<1VEcY6_(^~!$)%Gqi+3b6f4a*tZ8A@^Hl{OoLx3Qz9{XiG}jQNVT zQjqfZ8@y(osiV8INV)~Na&~Fb0nAFeY9K0E!sD%WRxM?Kz0o^3C@U_%5k}8`_$9!O zB5HfwyFF`AK&W7+%l&-ebxR~PAn6icmc@An766MvI!RNLY6v(Sf^xXq`RIsUE= z*JtYNDAnH_f-s_CSab?eqa(Xv(PZ^++HVla&pBpH*P(CULYZc5PBToTy<$5%To=nG zsq0HKcHiV8tY|dJ7LG)$Diq{VrZ())4-PNWr!xU=-%UpEK(i%GZMXN<=rzg z>~3HHy4GTTF5dA;1GH4Qy?x*Nfn^(73VRJzAtQyGj7y=4p;&pL4ZcNo_I)V`TBp2u zy;BWlsouCgBnX*4i`%tU#qDuouv&p0e;x>blI-{QDoC`dO!DVS~gZ9fsi^EZ6X_t3T zS4doKsTY@}b6-n=q`5}H1&h?VzhwxT_|8uSsNrtj7|ju{BJl%@zJ@*mtlT7QcWAB4`WHlYvl84zr*PNpBbM+L`7)^QkNAFQSIfQ4<76|kNW=@Z4LL4+KJ={? zr$5|L1!L*cq|1tIVqM9|_&LZt1(-A{7Y7L!Q~i?4u1vy1=$vu@T1JF86gkL<7= zbWkE2eK!oNbE-3_YqzI=26;B;#2dX?4+o$Kq1c6zhN6=s$_he~{>$a+6H@z=eo(qL zjOO&k4fymAlA`J||Emny3Lh@zQ>n{`j!;6}1<@yPA*r~C_-UN|N3)Oe?wibHTJNzL zZtPgxpN#GC$-_7PS8$B87@x8&R^z>#BUy90N+}}~?j?5n&bTe1F(LJchB?hDA+90G z`ZKnJ%V?uBg}2KzHy_b5;WG9ZNZ~y;FI=-gH_wRH3KAr70zZ)d;qtcp?E}{}&oJmy z;p68gKG@6d`9DZ3?EfyYFtIZ*{HOdYcILll@G9wuE6Cw|Om(_k#Pq98b&Jy^OoA({ zD>eu>=Nt1ih)Y0ImQ1y;{t4DN$pt42b*B!caF1|D4)`6COW=M~lw0zfpmrGVWD*Z0 z-Rtq9j8qU|Vct8$Ve0XH>hf}G6$~gBBQ#YqP%kxMqFtj`VJlhtd4z8(ziu; zizOMNlXJ!98`KlUvuq_dA@%eG^6`oLrmHEs9L#XqF`hP1{y!bz`jYqNynRgbE>~geKr& z;wp~AXj7Z*4W;F-3^mFl^GrsEUz_pftC@|h<%991o7l1oiiYF;Tsxc6ddhbzo86Hh ztHnP*&*AQ2+;7>~nKNA3`Q~whR};1mA5&wU*Kx@tMg-?bZ1}oYl^IZ=`Y!Fu z?P~QGfTGOUL8+hBt$WuL@4#3Ea;MG_FPfGfoO7D<>wX=1h_uergdenZSsGoq*|}7- z(`(!d*0R96ycqD(T1fV8^vy6n-){}!HNIM24NDp|&VO-Dc&AD_xI6+M(zlCi>*B`c zBGoYR>#-hxoAWq7(hvO&_L@kz0oaOMyiPE-YwZje%|Nb|+tlfvs%|}4z&;F-KZc@k zTrVU^E14Y%WL3h#3FBbOfl)GA^!JReDxD6k^^WEWc*0Nw zvDl+&H6IghBd$K)AQ~BV_gV9Krso+`UHXs+%mf@N1Cf#F`UpnlLqXs47VT@%zy>)V z^-H(lQw~}}pIx&bP&Y9_pPCj;V5L_E>am|!OMy4Yw@{6=|9}~ef4~egE8~CcJiyM% z`8QQ^O-dX^4+H$l#M#(cc~XtkVJSwcawCjQfn8%9o~Dr%hAEzBg>ZRukyX=00~Lje zM>vH}aAxo*AdwFO5=I6|aTzXHcZ+DGEl9%8znqu3J#m5Z=;^a}>Xqm6b%Ob_=kxeQlg~*qR0g0rHXa5=!_hLA=O2{g!JA{^23!!EsBT z9zl(xsC_q<|ES~Rgs?h<^D)@1k5Q>%sx{!&Yje- zC%woUaczZlyIlJ~&K{5tiJ^|OW21c8QSM2S)K;%hbICYy%0;|Iqyp7EYB!%`^SN;K&(GjQN zCFPqJBO<39JE`h!%|{w-k6G=m+Le*10j1l zx#^+TmjmmeyIcI&Q>|c?vC@P3xxb`}W1LW@x{9XfwQ@w2eV)1nLPJ(-db6T}J!`AC zLN`-_azd7Zto&EI#{1_#j(*ls)iiepNx%5J+S(c08z<;JUq9W2ABOe7}Qov+0UsqJtKgESP21B@Wr~ zyD8KPP@T(sss6VLwJlcKYiQ$|KT223r(9%1_gKWeP>O)~B7TJXZy+r!O(172JpAr|NYBn+?%gIKag8cj@Umy ziSz$L=|8pYVENk_m{#Is_&x(d;H5Vdt_gg=y@eD3NOGDa+-tw_BM^?{6b4z7U8c~M zw~?1V+ET2*?BeF)V*JVS{<d#Y77D->bLgiTVz4zW#iz@% zbJa;`w9#N0NWDlUry>wI0_nz0$h@q#fo^gzL+#`_CuG96*d~)IOD= z6Gte*tKyLcw~tM=+ydwMGlsmUD*iMj0Du3{AX?C*0%2Z?aO+P~ge14Neq+ipcc@oR zmk{dkl}^WMobyHdqx0%fv6Lfa7Qf|4HnX|50Ium*wVCCJhCcm(rs$kup^+%qfxuFq zq9Wm+Fi!BfgY%KEHo+r3iHeYzW(}$hre&&SzhmXFXoiohkF7@>4md?7A@tR!8`3YD z9G4k9$OKIk9QOkIA;l(;9Tcw>YkA#padp*G(Tt?hD3G`nf^eVfh-JUqJ3cx#bIk4p zk{L%ZIiHSO%Pr+j?4d9xGwFa6^BnMd9I|YQ95!fv8+Nv%Rnu@AZxv;Hb?;SA0w&@1k*(Ym$IX@>6Z zl+3A?aM4M=i#-^o?j~4YXNtvZ-x{>BmT;tK@i!=RY~v{)k2`eS3F#GCz8X(ayt6I4)38Zi9lPc1-*Y%N z>vz+$^P}&3WoI9b10I~CbHXPdOQRI{HOEr6ZAw6M@b;g`wf7lPiaBt`*qkNvtM8vG zyN2u&`m|%t^)CT{Q>5GJ{bh&9A`(l=;?~A%z9p&5lU<~UPBp#|`)%4BA;gde{#Uw? ztoA>h+41dk7ek?(72<1wnpv?afn{)pKNtxC0%#iI6daZ&mKrl?2iFzXFe_Y^li`Bx zM2yO%K{U}cU4+))`6mL4TLC15gMaELduGTCuNMw5xpJNYKN3H1-k)(fsZU6%{CoRy z8B|{sUh0qXkMu_=4;I9YOkRkmjI@r}3@53cWc&6puN*1l@&fmcuN@svkC$Q!qW5FTQTPrXVs!|*3B;hJ)J=4iy?J@Img?@;*<(Nn= z(mjIb^pG#?-N5OW*Qi;hmlreTP#XYvxouX&&1&*G#N70}qn|ltLF=w4J+m|YGST>g zWCZBykpT`_zZ2Qzk2Vi04bsP0XccC8h|8Q6t$si;El`$DAvAX|YHM!z-Jl_$hu%&Q z9!M1v7ZC?-nm8BdKS%YJ9PRx8QryM7syOPLYW@~X#b`MXjnKYdB&uDu>DPknb`eg8 zrxl1^(!tUHHh^q0vNvZo!EVbNe@gq=(lcq0p1hv{=iSzJ+Ib%Gbe+SnHNavab3tRzW<02MAT(eZ_~65+}4tlnWjnZ zukcuh&6U~%_Wa3zt#LEgeZNH-n}c%a3t4d+Yzs+k7Klg=5ZUy#<2u)By|2iRg@Ip> zx4r0u&A0lorK*W!W$VhSH&76ZZy1}IvvcKy*@C}>;bSYBlZ&JU4n_;_u5w*w45C(} zV590{#?5#OXt6=eqx5=*CBWyyO)+P71wE#7!jN^uK-5RDfYrczy#`hP7G#cZ_rViw)=hKTiK)oc<*PkKG=#X{PWh~7FHHD&~0qqfm}ewOp^D|1dmGb^HYXl}1#9}_6%;ybn% ztzwVj&7Ry|oKgMxWB6uJCTfcOk&L5OZ6yR{3F8cCNpf}&?Hj?hYOQ!8k4ix0PW#r}7?jX<`5s&nnlaj1f%ORUE%k^EdvmiJK2cpnzv}7jWG@bE7d0eY&2;IeEF- z%&S`q(Ji;R@m7r}5db~Iy0sI&0t0q%GJ!B#@D%&a9MP(P^DJgo@la-1kIZ9(BNY2s zZIcsNGk@)yICrbGfGzm*{Qlm#({tvXllA<@htCx=H9J2)KQnltC4i@xkW2c5I2SHj zp`BNz#Tby<%NRgND*2Q30`4)L1V;kHUEN*TPRI#b>?ZjDmMfylSbeww47{Ct=fal# zu0cZve;J#(EkDB8x&JJfht|#fV#&4{j6Wi$curO=P7Th&tL}AO6aTW!06zoqg7UCv z1cw8claD(#d0-tseYWDra(i<3clVjJ-l=e`u<}Ifn>F;kM$@H)ZFx807 zp`i<7kxv~VE9Rf}Y5>uWj(;PR$>U2C)#2YrV)(q% zJZ;KX>jz;-Vcu3!HePvi|KyDGB9g=psO20KaeAMbx@Ow~kd^%xR}gOWs<)*0{Nf5! zV~SslG})CGZd3K_vFt5rWSxnBVSd?}XP92(*}!^#10MbU4sf4%b<8J)e~lF)a=d3N zYjq}HU~_-h*(Iv)!1_1sw{(Jxt|oxBP%;#&E=RA z#m@9|Kda;RAWe?40qF*E1~EWZEaDz{auMeCQS(!}rllA@B_wfwcc zO`EM*BsSWK31((785x;j65s?|Q%F-r+U8~31B8*@_^-b2A<4eU;jcK4W{Hqv@u-6k z_UQ69pr zw^ASZO+*~<2z3btzxW5pkQvJEo}j$eKlFC={8NGo>9)@@qqb_QZ8yhbbpyazY4`zi zprzSdgTDtz(Z?!jH23-h_U(r}{fBw-zg?JF|3)hRzAz(086f}{v>(h=(yV{yXv)Mw zTc?;JwBcWc#r>rl@b;PWMC*3E#q|XzY|u@Ezu=VFFNmerGwUX|_ys48FF2Kc!Reig zJd^7SPFP=XnhH~1=}c$X(H?MX`eZTA16)o~p>c7G@!n zuZ5YH;TG_uPAsVPa)Kxd9^jyeCGm_X;&5P|SF2emVDoW|&8q69lDJ79zIoruh5GoA z-TKyPLeYs}mt7ZtL?UL3mtzr5XiSx(qw>|B<==;ik>US~g*gGQO$Z|LB3`Ju&JuoS z<*WIKWdTP>$GaiBBZD98&8N>^^FfivX@?|%4T@vTe-RisXtO{dF@T6tfE}0B7lA4K zA~27qnYJe#u=!>J-gO-0>UxjneZC%q{38UM{CyX%a}L||j2}CJe7$lFv=|LMM`e5M z=?I3sa9MhrMfW`uIr|n-5KhPCVi>qEQ*ve`hhiv`apgfn_Km=u&m}esy7yX$?RFfs z?R+ZqgPolwbUv4uKk0=tOU`mZMiDcI0qOdoX=4C%lLuRbe;TNNnNW&uwnhZ>GWzBU zj#l*21nhtLDcCzW60k9|{wH0QnU#U#UuWB*#=8AFM|#(GNzbBkTN@FgiEcTyl62-T z9!@t)@(ZgJdebmx4(c#EiEN37sh<;TC^Vov{oi%u)Q^`~>s*v4^g8FveSiQXGF79*yxJuS{^Ag}G@v1e5v~C}6-y zu+RiE4T2<7#DFDZ^u((Q>XIaLP^S$bUGgOSwg;^^>h9(E(1_s=%+Cc=nEURLYmM`B zI+2)*skoyTf6CEV4yu^+Ua~kT4o1Tqmg8rDI6*bcJf1kkIE>4Y%LGOA9J9fYPO z0sP900Hy}kTg;NdiSIx(`dj|er)iJcOf?vS4#}MH*&qNvmT8uQJ#pdWua(INPTT+s zkOo8{9x;y-bGYG)VL95`A3MxP#2m|%3?1z)BdDS_B)-unbb~i;1eWX`6D|;BkSa+k zAywlCA_IxLP?5)XM%qb=BH`oN%RYSlY5{47;vDAreTLvahnZsULMnugxN$eR%4ryeX_3IX`e+(*9+zn+sW* zi(4{%Qw~rFebs;C>()Qw_*Aco*e@GMMqeGlCJiN|!RE~Sv3&SMqO*Z=w#PXd=k0WO zGZ-=6g6qoIkTsQ9eEvk?z&CIG>WnzJyuIYZsRcOoe;BWgACs7Rf_}sP{5U?j95jsL zV~M`KZ&Rhy-@miqqeW-0y)H{v8lPYHl<&C41}I1I0H$E!D?uYSNS2#LqWMW2Gt+P!BYJYVIkUh(r4Qe*MNO z=J4yw&u8%X%DC$h!&qs`h_WB<8H>mMiTmt_NYuirezs9!u@a@CWUardKlb)oBo`O3pU(UjP-UTPPZ}LPXqx$!kN= z%Ok5`t%@=uVMj=OXxmGPTKvLyG~I%7TJ`F58jLfd(S=juN8>Ugn-l*oMnY}S$6$yvs6g7#3 zF_yCdFfxV=aSB>F%Jrg0Y~R8hW^FVxWvoOj$ew_@^1NIvLu!c{OYFQ?*d@K{R)uu= zll}y@F0)z+h-!@pCw%oa&x8|)_sW>Fqnu-5&T|YYd{eDA?QRdP(8_k;}us5#B0<*3% z96NKNor98HX|{>OQHIh^l(h_*qr!Z#ALrRO!z>22@t4Qa0`|B5QU0<{nb&11jdBZl z7vjgA!nH|>@%w)He1oPMgV{~$*h=d+Z(i4dNnH18o>Vj-%81{$_cu}s^LY^~` zqU``=eRgtNXMD~|(oK1ugFw84z|BF(%{}4lJ3`rPL+$n|@HXdaXg0NEnWD`Dl*&@A zh0gnQao%%$&D4lParaaj638Ns!win1y1cIf9r$!{H0kgR`^M;}8ut2=9PU_o`$S%WGP~0%Q2JB9 zklgpCr(~Jc*PFIf&zym%NFmx8PQu=MjvlFUP z^3nvWKVT}x*e=CJlK1K!3t-lVool6HqLb=>&OzkDk|z#xC%E(O(}>dW(GtnydSoPS zW_-!yPQ{>)5(sRg)`=f}t3wu^0a1!iTHo4_MwhE`u0l%rK?<5KC!mw{u@gLoz5DDZ zIU*%llbW7Gtc1Hqfv?jkyB5Vg^fgRlNmPY-NL00vF82zBo^TwRq#t12?Ccx0rY1Jm zYyBEBK*-c;z3q}IECZc9)3xXRaoTd9NVdUT5BFL;s#s~&YCT`!(?VvLRXKxT>AZ#% zYi?P0xHY+7KTiiKe@G8VS+4fb_%Ib15CcLpbA8_LtNbue(>+Nn`Uyh{j zA$H=73TiTQWo&y}V3u3ou}epW_pH?_yQZs09_s=M&SRa7MaS;5N1fOp)bD(iQO<5x zUrs@wSGwyB4ulqge{l>j*PM%s2b8CQi``*C@`VPT{Sb^v*fS4}k%(Cih(O;4c(G4}m^v-HzMlKgD@omwoGpcZbvF0X@vD30B#ij2wz3|EWGzrsRaZ9*9mA zxM0rBhO=Tz*?B@#03LTdB!btsr_a=`f%uEyjw)wzrA;3kyV*&fVf~N;m|}fUpMmSN ziYpUkc`&}?mr%o+43OhAs3J&QRg{t!^ZFQpFo2MXR}&OC!N@J5U)Y&SPJV06WoO=b zO1pit_XOEik^XXK)fHvqQ;j2?gJ2|HFyqUP+U}wq%c!FvrcQ!^f2ehYhKbnXhY2R#tKT9zo_mPwCeFah<8h2S4@VS#VEG375M@wcHudcW>3(q5zxZ1i4SFf zT}H9e!We$L5vgb~l^Qj_823o9+|N8|3faQyWTVGZC5cC_C}|^(Nn#yJJ7Z~u{&8$M z0H=g{Er}Kt@)P~OVy*Dn7Bfuk>>|B+ZMq|5p}N+R`mxsX-F4ji_1R`PXUgYYwXz&9 zqxj-{>Fjt@NI(Dj`_eBVT8D=^qbWVha^48?Rfm)^e2dW*Zo9vj`)U=?XA&h0f^(f>c|Cm>XR)m0_= zi(x1N0@V~Ud;2ecXfe>1lnn&arJlF#WEVoOo4xjSp>FRIhD<1`&>yWgcA^+?<3`2L zxFp3aBfZ#G$tAeb|HKEIgS*BG;sY;K<`JJ4^IZrZrg*uCQD-zMlzIP+di(vQEk=E| zKJK$nS@H4;R!4vd0hMp9-DJqlwY^PsZ}Npm$E!pIX9?kh;&M~(@R(z8o3f^Fc2^@aB9)35NI$GFcUCo(JPp_8~t^UUf#yW zk>D?XKKZX_jcvZ}{Z&jQ|DX1cOq_q~O2s-eO7~qCC3x2%axqNy;psx=%R5+FlVs=IGlYRKlXoihTbf`F05@8*=e>i4pDUnE% zN$HbwWxWzGoQ>~B=+`r;~B@L%k7-pYbEA8RhZX0_*Z47($?oE{vaZb((_7#7tFAo z;Nk7cgG6SZSyDCex`{0sK(UM7GG`bxvh;0Zr7Kjjr6S(5(F;_?4tCj2=x{lwOOY2b kjz6kZrr_USn1iF9y`!tWkuek-I~x-_6e+2Qj40Iq16fw@T>t<8 diff --git a/verisimdb/admin/deno.json b/verisimdb/admin/deno.json deleted file mode 100644 index b193a0d3..00000000 --- a/verisimdb/admin/deno.json +++ /dev/null @@ -1,24 +0,0 @@ -{ - "name": "@hyperpolymath/verisimdb-admin", - "version": "0.1.0", - "license": "MPL-2.0", - "tasks": { - "gossamer:dev": "deno task res:build && deno task css:build && deno run --allow-net --allow-read --allow-env dev-server.ts", - "gossamer:build": "deno task res:build && deno task css:build && deno run --allow-all build.ts", - "res:build": "npx rescript build", - "res:watch": "npx rescript build -w", - "bundle": "deno task gossamer:build && gossamer bundle", - "css:build": "cp src/styles.css public/styles.css", - "check": "deno check dev-server.ts build.ts", - "clean": "npx rescript clean && rm -rf public/*.js public/*.css" - }, - "imports": { - "@gossamer/api": "https://deno.land/x/gossamer_api@0.1.0/mod.ts", - "rescript": "npm:rescript@12", - "@rescript/core": "npm:@rescript/core@1" - }, - "compilerOptions": { - "strict": true, - "noUnusedLocals": true - } -} diff --git a/verisimdb/admin/gossamer.conf.json b/verisimdb/admin/gossamer.conf.json deleted file mode 100644 index 83003d87..00000000 --- a/verisimdb/admin/gossamer.conf.json +++ /dev/null @@ -1,141 +0,0 @@ -{ - "$schema": "https://gossamer.dev/schemas/config/v1", - "productName": "VeriSimDB Admin", - "version": "0.1.0", - "identifier": "com.hyperpolymath.verisimdb-admin", - "build": { - "frontendDist": "../public", - "devUrl": "http://localhost:8200/", - "beforeDevCommand": "deno task dev", - "beforeBuildCommand": "deno task build" - }, - "app": { - "windows": [ - { - "label": "main", - "title": "VeriSimDB Admin", - "width": 1400, - "height": 900, - "minWidth": 1000, - "minHeight": 700, - "maxWidth": null, - "maxHeight": null, - "resizable": true, - "fullscreen": false, - "decorations": true, - "transparent": false, - "center": true, - "alwaysOnTop": false, - "visible": true, - "url": "/" - } - ], - "security": { - "csp": null, - "capabilities": [ - "network", - "filesystem", - "clipboard" - ], - "capabilityTokens": { - "enabled": true, - "issuer": "gossamer-runtime", - "ttl": 3600 - }, - "sandbox": { - "enabled": true, - "allowExec": false, - "allowNetwork": true, - "allowFilesystem": true - } - }, - "ipc": { - "protocol": "json", - "bridgeInjection": true, - "maxMessageSize": 16777216, - "timeout": 30000 - }, - "tray": { - "enabled": false, - "icon": null, - "tooltip": null, - "menuOnLeftClick": true - } - }, - "plugins": {}, - "bundle": { - "active": true, - "targets": [ - "deb", - "appimage" - ], - "icon": [ - "icons/32x32.png", - "icons/128x128.png", - "icons/128x128@2x.png", - "icons/icon.icns", - "icons/icon.ico" - ], - "resources": [], - "copyright": "Copyright (c) 2026 Jonathan D.A. Jewell (hyperpolymath)", - "license": "PMPL-1.0-or-later", - "shortDescription": "Admin panel for VeriSimDB cross-modal entity database", - "longDescription": "Gossamer-native admin panel for VeriSimDB. Provides VCL console, octad browser, entity detail with full modality view, drift detection indicators, and telemetry dashboard. Second app built natively for the Gossamer webview shell.", - "linux": { - "desktopEntry": true, - "category": "Development", - "section": "database", - "depends": [], - "appimage": { - "bundleMediaFramework": false - } - }, - "macos": { - "bundleIdentifier": "com.hyperpolymath.verisimdb-admin", - "minimumSystemVersion": "11.0", - "frameworks": [], - "entitlements": null, - "signingIdentity": null, - "notarization": { - "enabled": false, - "teamId": null - } - }, - "windows": { - "webview2": "embed", - "certificateThumbprint": null, - "wix": { - "language": "en-US" - }, - "nsis": { - "displayLanguageSelector": false, - "installerIcon": null - } - }, - "mobile": { - "ios": { - "minimumVersion": "14.0", - "deviceFamily": [ - "iphone", - "ipad" - ], - "capabilities": [], - "frameworks": [], - "entitlements": {} - }, - "android": { - "minimumSdk": 24, - "targetSdk": 34, - "permissions": [], - "features": [] - } - } - }, - "ephapax": { - "modules": [], - "region": "admin", - "linearVerification": true, - "regionBoundary": "strict", - "preload": [] - } -} diff --git a/verisimdb/admin/panels/manifest.json b/verisimdb/admin/panels/manifest.json deleted file mode 100644 index 3b9c728e..00000000 --- a/verisimdb/admin/panels/manifest.json +++ /dev/null @@ -1,97 +0,0 @@ -{ - "$schema": "panll-harness/v2", - "service_id": "verisimdb-admin", - "default_endpoint": "verisimdb://admin", - "runtime_endpoints": { - "gossamer": "gossamer://verisimdb-admin/bridge" - }, - "health_check": { - "path": "/health", - "interval_ms": 30000, - "timeout_ms": 10000, - "unhealthy_threshold": 3 - }, - "data_sources": { - "verisimdb://admin/health": { - "path": "/invoke/verisimdb_check_health", - "method": "GET", - "returns": "HealthStatus", - "description": "VeriSimDB Rust core health check" - }, - "verisimdb://admin/vcl": { - "path": "/invoke/verisimdb_query_vcl", - "method": "POST", - "body": { "query": "{vcl_query}" }, - "returns": "VclResult", - "description": "Execute a VCL query against the database" - }, - "verisimdb://admin/octads": { - "path": "/invoke/verisimdb_list_octads", - "method": "GET", - "returns": "OctadList", - "description": "Paginated list of octad entity summaries" - }, - "verisimdb://admin/octads/{id}": { - "path": "/invoke/verisimdb_get_entity", - "method": "GET", - "returns": "OctadSnapshot", - "description": "Full 8-modality octad entity detail" - }, - "verisimdb://admin/octads/create": { - "path": "/invoke/verisimdb_create_octad", - "method": "POST", - "body": { "entity": "{entity_json}" }, - "returns": "OctadSnapshot", - "description": "Create a new octad entity" - }, - "verisimdb://admin/octads/{id}/delete": { - "path": "/invoke/verisimdb_delete_octad", - "method": "DELETE", - "returns": "DeleteResult", - "description": "Delete an octad entity" - }, - "verisimdb://admin/drift/{id}": { - "path": "/invoke/verisimdb_get_drift", - "method": "GET", - "returns": "DriftInfo", - "description": "Per-modality drift scores for an entity" - }, - "verisimdb://admin/normalise/{id}": { - "path": "/invoke/verisimdb_trigger_normalise", - "method": "POST", - "returns": "NormaliseResult", - "description": "Trigger self-normalisation for a drifted entity" - }, - "verisimdb://admin/telemetry": { - "path": "/invoke/verisimdb_get_telemetry", - "method": "GET", - "returns": "TelemetryReport", - "description": "Aggregate telemetry from the Elixir orchestration layer" - }, - "verisimdb://admin/orch-status": { - "path": "/invoke/verisimdb_get_orch_status", - "method": "GET", - "returns": "OrchestrationStatus", - "description": "Elixir orchestration layer status (consensus, federation)" - } - }, - "panels": [], - "capabilities": [ - "network", - "filesystem", - "clipboard" - ], - "clade": "infrastructure/database", - "integrations": { - "verisimdb-rust-core": { - "role": "admin-interface", - "protocol": "gossamer-ipc", - "description": "Gossamer-native admin panel for VeriSimDB Rust core (port 8080)" - }, - "verisimdb-elixir-orch": { - "role": "admin-interface", - "protocol": "gossamer-ipc", - "description": "Gossamer-native admin panel for VeriSimDB Elixir orchestration (port 4080)" - } - } -} diff --git a/verisimdb/admin/public/index.html b/verisimdb/admin/public/index.html deleted file mode 100644 index ed40dab8..00000000 --- a/verisimdb/admin/public/index.html +++ /dev/null @@ -1,15 +0,0 @@ - - - - - - - - VeriSimDB Admin - - - -

- - - diff --git a/verisimdb/admin/rescript.json b/verisimdb/admin/rescript.json deleted file mode 100644 index 40496707..00000000 --- a/verisimdb/admin/rescript.json +++ /dev/null @@ -1,30 +0,0 @@ -{ - "name": "@hyperpolymath/verisimdb-admin", - "sources": [ - { - "dir": "src", - "subdirs": true - } - ], - "package-specs": [ - { - "module": "esmodule", - "in-source": true - } - ], - "suffix": ".res.mjs", - "bs-dependencies": [ - "@rescript/core" - ], - "bsc-flags": [ - "-open RescriptCore" - ], - "warnings": { - "number": "+a-4-9-20-40-41-42-50-61", - "error": "+5+6+101+109" - }, - "jsx": { - "version": 4, - "mode": "automatic" - } -} diff --git a/verisimdb/admin/src/App.res b/verisimdb/admin/src/App.res deleted file mode 100644 index eae07abd..00000000 --- a/verisimdb/admin/src/App.res +++ /dev/null @@ -1,904 +0,0 @@ -// SPDX-License-Identifier: MPL-2.0 -// Copyright (c) 2026 Jonathan D.A. Jewell (hyperpolymath) - -/// App -- TEA (The Elm Architecture) entry point for VeriSimDB Admin. -/// -/// This is the SECOND application built natively for the Gossamer webview -/// shell (after Burble Admin). It showcases Gossamer's capability token -/// system applied to a database administration context. -/// -/// Architecture: -/// - Model.res -- State types (octads, VCL, drift, telemetry, caps) -/// - Msg.res -- Message types -/// - App.res -- init, update, view (this file) -/// - VeriSimDbCmd -- IPC commands to the VeriSimDB backend -/// - Capabilities -- Gossamer capability token management -/// - RuntimeBridge -- Gossamer-native IPC bridge -/// -/// Panel layout: -/// +---------------------------------------------+ -/// | Header: status, runtime, health, caps | -/// +----------+----------------------------------+ -/// | VCL Console (top bar, full width) | -/// | [input area] [results] | -/// +----------+----------------------------------+ -/// | Sidebar | Entity Detail (main) | -/// | Octad | [modality tabs] | -/// | Browser | [content per modality] | -/// | (paged) | [drift indicator + normalise] | -/// +----------+----------------------------------+ -/// | Telemetry Dashboard (bottom bar) | -/// +---------------------------------------------+ - -// --------------------------------------------------------------------------- -// TEA command helpers -// --------------------------------------------------------------------------- - -/// Wrap an async operation as a TEA command. -/// Runs the promise and dispatches the resulting message. -let cmdFromPromise = ( - promiseFn: unit => promise, - onOk: string => Msg.msg, - onErr: string => Msg.msg, -): Tea_Cmd.t => { - Tea_Cmd.call(dispatch => { - promiseFn() - ->Promise.thenResolve(result => dispatch(onOk(result))) - ->Promise.catch(err => { - let errMsg = switch err { - | JsExn(jsErr) => - switch JsExn.message(jsErr) { - | Some(m) => m - | None => "Unknown error" - } - | _ => "Unknown error" - } - dispatch(onErr(errMsg)) - Promise.resolve() - }) - ->ignore - }) -} - -/// Extract a network capability token from the model. -/// Returns None if the network capability has not been granted. -let getNetworkToken = (model: Model.model): option => { - switch model.networkCap { - | Granted(token) => Some(token) - | _ => None - } -} - -/// Extract a clipboard capability token from the model. -let getClipboardToken = (model: Model.model): option => { - switch model.clipboardCap { - | Granted(token) => Some(token) - | _ => None - } -} - -// --------------------------------------------------------------------------- -// Init -// --------------------------------------------------------------------------- - -/// Initialise the application. Starts with the capability grant panel -/// visible and no active connections. -let init = (): (Model.model, Tea_Cmd.t) => { - (Model.initial, Tea_Cmd.none) -} - -// --------------------------------------------------------------------------- -// Update -// --------------------------------------------------------------------------- - -/// Process a message and return the new state plus any commands to execute. -let update = (model: Model.model, msg: Msg.msg): (Model.model, Tea_Cmd.t) => { - switch msg { - // --- Server health --- - | CheckHealth => - switch getNetworkToken(model) { - | Some(token) => - let cmd = cmdFromPromise( - () => VeriSimDbCmd.checkHealth(token), - result => Msg.HealthResult(Ok(result)), - err => Msg.HealthResult(Error(err)), - ) - ({...model, status: Connecting}, cmd) - | None => ({...model, error: Some("Network capability required. Grant it in the capability panel.")}, Tea_Cmd.none) - } - - | HealthResult(Ok(_response)) => - ({...model, status: Connected, error: None}, Tea_Cmd.none) - - | HealthResult(Error(err)) => - ({...model, status: Disconnected, error: Some(`Health check failed: ${err}`)}, Tea_Cmd.none) - - // --- VCL console --- - | VclInputChanged(input) => - ({...model, vclInput: input}, Tea_Cmd.none) - - | ExecuteVcl => - switch getNetworkToken(model) { - | Some(token) => - if String.trim(model.vclInput) == "" { - ({...model, error: Some("VCL query cannot be empty.")}, Tea_Cmd.none) - } else { - let cmd = cmdFromPromise( - () => VeriSimDbCmd.queryVcl(model.vclInput, token), - result => Msg.VclResult(Ok(result)), - err => Msg.VclResult(Error(err)), - ) - ({...model, vclExecuting: true, vclResult: None, error: None}, cmd) - } - | None => ({...model, error: Some("Network capability required.")}, Tea_Cmd.none) - } - - | VclResult(Ok(result)) => - ({...model, vclExecuting: false, vclResult: Some(result), error: None}, Tea_Cmd.none) - - | VclResult(Error(err)) => - ({...model, vclExecuting: false, error: Some(`VCL query failed: ${err}`)}, Tea_Cmd.none) - - | CopyVcl => - switch getClipboardToken(model) { - | Some(token) => - let cmd = cmdFromPromise( - () => Capabilities.copyToClipboard(model.vclInput, token)->Promise.thenResolve(_ => "copied"), - _result => Msg.NoOp, - _err => Msg.NoOp, - ) - (model, cmd) - | None => ({...model, error: Some("Clipboard capability required.")}, Tea_Cmd.none) - } - - // --- Octad browser --- - | LoadOctads => - switch getNetworkToken(model) { - | Some(token) => - let cmd = cmdFromPromise( - () => VeriSimDbCmd.listOctads(model.octadLimit, model.octadOffset, token), - result => Msg.OctadsLoaded(Ok(result)), - err => Msg.OctadsLoaded(Error(err)), - ) - (model, cmd) - | None => ({...model, error: Some("Network capability required.")}, Tea_Cmd.none) - } - - | OctadsLoaded(Ok(_response)) => - // In a full implementation, parse the JSON response into octadSummary array. - // For now, store success and clear errors. - ({...model, error: None}, Tea_Cmd.none) - - | OctadsLoaded(Error(err)) => - ({...model, error: Some(`Failed to load octads: ${err}`)}, Tea_Cmd.none) - - | SelectEntity(entityId) => - switch getNetworkToken(model) { - | Some(token) => - let cmd = cmdFromPromise( - () => VeriSimDbCmd.getEntity(entityId, token), - result => Msg.EntityLoaded(Ok(result)), - err => Msg.EntityLoaded(Error(err)), - ) - ({...model, selectedEntity: Some(entityId), detailTab: Overview}, cmd) - | None => ({...model, error: Some("Network capability required.")}, Tea_Cmd.none) - } - - | EntityLoaded(Ok(detail)) => - ({...model, entityDetail: Some(detail), error: None}, Tea_Cmd.none) - - | EntityLoaded(Error(err)) => - ({...model, error: Some(`Failed to load entity: ${err}`)}, Tea_Cmd.none) - - | SwitchDetailTab(tab) => - ({...model, detailTab: tab}, Tea_Cmd.none) - - | NextPage => - let newOffset = model.octadOffset + model.octadLimit - let updated = {...model, octadOffset: newOffset} - update(updated, LoadOctads) - - | PrevPage => - let newOffset = Math.Int.max(0, model.octadOffset - model.octadLimit) - let updated = {...model, octadOffset: newOffset} - update(updated, LoadOctads) - - // --- Entity CRUD --- - | CreateOctad(entityJson) => - switch getNetworkToken(model) { - | Some(token) => - let cmd = cmdFromPromise( - () => VeriSimDbCmd.createOctad(entityJson, token), - result => Msg.OctadCreated(Ok(result)), - err => Msg.OctadCreated(Error(err)), - ) - (model, cmd) - | None => ({...model, error: Some("Network capability required.")}, Tea_Cmd.none) - } - - | OctadCreated(Ok(_response)) => - update(model, LoadOctads) - - | OctadCreated(Error(err)) => - ({...model, error: Some(`Failed to create octad: ${err}`)}, Tea_Cmd.none) - - | DeleteOctad(entityId) => - switch getNetworkToken(model) { - | Some(token) => - let cmd = cmdFromPromise( - () => VeriSimDbCmd.deleteOctad(entityId, token), - result => Msg.OctadDeleted(Ok(result)), - err => Msg.OctadDeleted(Error(err)), - ) - (model, cmd) - | None => ({...model, error: Some("Network capability required.")}, Tea_Cmd.none) - } - - | OctadDeleted(Ok(_response)) => - let updated = {...model, selectedEntity: None, entityDetail: None, driftStatus: None} - update(updated, LoadOctads) - - | OctadDeleted(Error(err)) => - ({...model, error: Some(`Failed to delete octad: ${err}`)}, Tea_Cmd.none) - - // --- Drift detection --- - | LoadDrift(entityId) => - switch getNetworkToken(model) { - | Some(token) => - let cmd = cmdFromPromise( - () => VeriSimDbCmd.getDrift(entityId, token), - result => Msg.DriftLoaded(Ok(result)), - err => Msg.DriftLoaded(Error(err)), - ) - (model, cmd) - | None => ({...model, error: Some("Network capability required.")}, Tea_Cmd.none) - } - - | DriftLoaded(Ok(_response)) => - // In a full implementation, parse the JSON into driftInfo. - ({...model, error: None}, Tea_Cmd.none) - - | DriftLoaded(Error(err)) => - ({...model, error: Some(`Failed to load drift: ${err}`)}, Tea_Cmd.none) - - | TriggerNormalise(entityId) => - switch getNetworkToken(model) { - | Some(token) => - let cmd = cmdFromPromise( - () => VeriSimDbCmd.triggerNormalise(entityId, token), - result => Msg.NormaliseResult(Ok(result)), - err => Msg.NormaliseResult(Error(err)), - ) - (model, cmd) - | None => ({...model, error: Some("Network capability required.")}, Tea_Cmd.none) - } - - | NormaliseResult(Ok(_response)) => - // Reload entity and drift after normalisation. - switch model.selectedEntity { - | Some(entityId) => - let (m1, cmd1) = update(model, SelectEntity(entityId)) - let cmd2 = cmdFromPromise( - () => { - switch getNetworkToken(m1) { - | Some(token) => VeriSimDbCmd.getDrift(entityId, token) - | None => Promise.reject(JsError.throwWithMessage("No token")) - } - }, - result => Msg.DriftLoaded(Ok(result)), - err => Msg.DriftLoaded(Error(err)), - ) - (m1, Tea_Cmd.batch([cmd1, cmd2])) - | None => (model, Tea_Cmd.none) - } - - | NormaliseResult(Error(err)) => - ({...model, error: Some(`Normalisation failed: ${err}`)}, Tea_Cmd.none) - - // --- Telemetry --- - | LoadTelemetry => - switch getNetworkToken(model) { - | Some(token) => - let cmd = cmdFromPromise( - () => VeriSimDbCmd.getTelemetry(token), - result => Msg.TelemetryLoaded(Ok(result)), - err => Msg.TelemetryLoaded(Error(err)), - ) - (model, cmd) - | None => ({...model, error: Some("Network capability required.")}, Tea_Cmd.none) - } - - | TelemetryLoaded(Ok(raw)) => - ({...model, telemetry: Some({raw, enabled: true}), error: None}, Tea_Cmd.none) - - | TelemetryLoaded(Error(err)) => - ({...model, error: Some(`Failed to load telemetry: ${err}`)}, Tea_Cmd.none) - - // --- Orchestration status --- - | LoadOrchStatus => - switch getNetworkToken(model) { - | Some(token) => - let cmd = cmdFromPromise( - () => VeriSimDbCmd.getOrchStatus(token), - result => Msg.OrchStatusLoaded(Ok(result)), - err => Msg.OrchStatusLoaded(Error(err)), - ) - (model, cmd) - | None => ({...model, error: Some("Network capability required.")}, Tea_Cmd.none) - } - - | OrchStatusLoaded(Ok(status)) => - ({...model, orchStatus: Some(status), error: None}, Tea_Cmd.none) - - | OrchStatusLoaded(Error(err)) => - ({...model, error: Some(`Failed to load orchestration status: ${err}`)}, Tea_Cmd.none) - - // --- Gossamer capability tokens --- - | RequestCapability(kind) => - let kindInt = switch kind { - | "network" => Capabilities.Kind.network - | "filesystem" => Capabilities.Kind.filesystem - | "clipboard" => Capabilities.Kind.clipboard - | _ => 0 - } - let updatedModel = switch kind { - | "network" => {...model, networkCap: Pending} - | "filesystem" => {...model, filesystemCap: Pending} - | "clipboard" => {...model, clipboardCap: Pending} - | _ => model - } - let cmd = cmdFromPromise( - () => Capabilities.requestCapability(kindInt)->Promise.thenResolve(token => Float.toString(token)), - tokenStr => { - switch Float.fromString(tokenStr) { - | Some(token) => Msg.CapGranted(kind, token) - | None => Msg.ClearError - } - }, - _err => Msg.CapRevoked(kind), - ) - (updatedModel, cmd) - - | CapGranted(kind, token) => - switch kind { - | "network" => ({...model, networkCap: Granted(token), error: None}, Tea_Cmd.none) - | "filesystem" => ({...model, filesystemCap: Granted(token), error: None}, Tea_Cmd.none) - | "clipboard" => ({...model, clipboardCap: Granted(token), error: None}, Tea_Cmd.none) - | _ => (model, Tea_Cmd.none) - } - - | CapRevoked(kind) => - switch kind { - | "network" => ({...model, networkCap: Denied}, Tea_Cmd.none) - | "filesystem" => ({...model, filesystemCap: Denied}, Tea_Cmd.none) - | "clipboard" => ({...model, clipboardCap: Denied}, Tea_Cmd.none) - | _ => (model, Tea_Cmd.none) - } - - | DismissCapPanel => - ({...model, showCapPanel: false}, Tea_Cmd.none) - - | ShowCapPanel => - ({...model, showCapPanel: true}, Tea_Cmd.none) - - // --- UI --- - | ClearError => - ({...model, error: None}, Tea_Cmd.none) - - | NoOp => - (model, Tea_Cmd.none) - } -} - -// --------------------------------------------------------------------------- -// View helpers -// --------------------------------------------------------------------------- - -/// Render the status indicator with appropriate colour. -let statusIndicator = (status: Model.serverStatus): Tea_Html.t => { - let (label, className) = switch status { - | Connected => ("Connected", "status-connected") - | Disconnected => ("Disconnected", "status-disconnected") - | Connecting => ("Connecting...", "status-connecting") - } - Tea_Html.span( - [Tea_Html.Attributes.class(className)], - [Tea_Html.text(label)], - ) -} - -/// Render a capability row in the grant panel. -let capabilityRow = ( - kindName: string, - kindInt: int, - status: Model.capabilityStatus, -): Tea_Html.t => { - let statusText = switch status { - | NotRequested => "Not requested" - | Pending => "Requesting..." - | Granted(_) => "Granted" - | Denied => "Denied" - } - let statusClass = switch status { - | NotRequested => "cap-not-requested" - | Pending => "cap-pending" - | Granted(_) => "cap-granted" - | Denied => "cap-denied" - } - let button = switch status { - | NotRequested | Denied => - Tea_Html.button( - [Tea_Html.Events.onClick(Msg.RequestCapability(kindName))], - [Tea_Html.text("Grant")], - ) - | Pending => - Tea_Html.button( - [Tea_Html.Attributes.disabled(true)], - [Tea_Html.text("Pending...")], - ) - | Granted(_) => - Tea_Html.button( - [Tea_Html.Attributes.disabled(true)], - [Tea_Html.text("Active")], - ) - } - Tea_Html.div( - [Tea_Html.Attributes.class("cap-row")], - [ - Tea_Html.div( - [Tea_Html.Attributes.class("cap-info")], - [ - Tea_Html.strong([], [Tea_Html.text(Capabilities.Kind.toString(kindInt))]), - Tea_Html.p([], [Tea_Html.text(Capabilities.Kind.description(kindInt))]), - Tea_Html.span([Tea_Html.Attributes.class(statusClass)], [Tea_Html.text(statusText)]), - ], - ), - button, - ], - ) -} - -/// Render an octad summary card in the sidebar. -let octadCard = (octad: Model.octadSummary): Tea_Html.t => { - let driftClass = if octad.driftScore > 0.5 { - "drift-high" - } else if octad.driftScore > 0.2 { - "drift-medium" - } else { - "drift-low" - } - Tea_Html.div( - [ - Tea_Html.Attributes.class("octad-card"), - Tea_Html.Events.onClick(Msg.SelectEntity(octad.id)), - ], - [ - Tea_Html.h3([], [Tea_Html.text(octad.title)]), - Tea_Html.p( - [Tea_Html.Attributes.class("octad-meta")], - [Tea_Html.text(`${Int.toString(octad.activeModalities)}/8 modalities`)], - ), - Tea_Html.span( - [Tea_Html.Attributes.class(driftClass)], - [Tea_Html.text(`Drift: ${Float.toFixedWithPrecision(octad.driftScore, ~digits=3)}`)], - ), - ], - ) -} - -/// Render a modality tab button. -let modalityTab = ( - tab: Model.detailTab, - activeTab: Model.detailTab, - label: string, -): Tea_Html.t => { - let className = if tab == activeTab { "tab-active" } else { "tab-inactive" } - Tea_Html.button( - [ - Tea_Html.Attributes.class(`modality-tab ${className}`), - Tea_Html.Events.onClick(Msg.SwitchDetailTab(tab)), - ], - [Tea_Html.text(label)], - ) -} - -/// Render the drift indicator for the currently selected entity. -let driftIndicator = (driftOpt: option): Tea_Html.t => { - switch driftOpt { - | Some(drift) => - let driftRow = (label: string, value: float) => { - let cls = if value > 0.5 { - "drift-bar-high" - } else if value > 0.2 { - "drift-bar-medium" - } else { - "drift-bar-low" - } - Tea_Html.div( - [Tea_Html.Attributes.class("drift-row")], - [ - Tea_Html.span([Tea_Html.Attributes.class("drift-label")], [Tea_Html.text(label)]), - Tea_Html.div( - [Tea_Html.Attributes.class(`drift-bar ${cls}`)], - [Tea_Html.text(Float.toFixedWithPrecision(value, ~digits=4))], - ), - ], - ) - } - Tea_Html.div( - [Tea_Html.Attributes.class("drift-indicator")], - [ - Tea_Html.h3([], [Tea_Html.text("Drift Status")]), - driftRow("Semantic-Vector", drift.semanticVectorDrift), - driftRow("Graph-Document", drift.graphDocumentDrift), - driftRow("Temporal Consistency", drift.temporalConsistencyDrift), - driftRow("Tensor", drift.tensorDrift), - driftRow("Schema", drift.schemaDrift), - driftRow("Quality", drift.qualityDrift), - Tea_Html.button( - [ - Tea_Html.Attributes.class("normalise-button"), - Tea_Html.Events.onClick(Msg.TriggerNormalise(drift.entityId)), - ], - [Tea_Html.text("Normalise")], - ), - ], - ) - | None => - Tea_Html.div( - [Tea_Html.Attributes.class("drift-indicator drift-empty")], - [Tea_Html.text("Select an entity to view drift status")], - ) - } -} - -// --------------------------------------------------------------------------- -// View -// --------------------------------------------------------------------------- - -/// Render the complete VeriSimDB Admin panel UI. -let view = (model: Model.model): Tea_Html.t => { - Tea_Html.div( - [Tea_Html.Attributes.class("verisimdb-admin")], - [ - // --- Header --- - Tea_Html.header( - [Tea_Html.Attributes.class("admin-header")], - [ - Tea_Html.h1([], [Tea_Html.text("VeriSimDB Admin")]), - Tea_Html.div( - [Tea_Html.Attributes.class("header-controls")], - [ - Tea_Html.span( - [Tea_Html.Attributes.class("runtime-badge")], - [Tea_Html.text(`Runtime: ${RuntimeBridge.runtimeName()}`)], - ), - statusIndicator(model.status), - Tea_Html.button( - [Tea_Html.Events.onClick(Msg.CheckHealth)], - [Tea_Html.text("Check Health")], - ), - Tea_Html.button( - [Tea_Html.Events.onClick(Msg.LoadOrchStatus)], - [Tea_Html.text("Orch Status")], - ), - Tea_Html.button( - [Tea_Html.Events.onClick(Msg.ShowCapPanel)], - [Tea_Html.text("Capabilities")], - ), - ], - ), - ], - ), - - // --- Error bar --- - switch model.error { - | Some(err) => - Tea_Html.div( - [Tea_Html.Attributes.class("error-bar")], - [ - Tea_Html.text(err), - Tea_Html.button( - [Tea_Html.Events.onClick(Msg.ClearError)], - [Tea_Html.text("Dismiss")], - ), - ], - ) - | None => Tea_Html.noNode - }, - - // --- Capability grant panel --- - if model.showCapPanel { - Tea_Html.div( - [Tea_Html.Attributes.class("cap-panel")], - [ - Tea_Html.h2([], [Tea_Html.text("Gossamer Capability Tokens")]), - Tea_Html.p( - [Tea_Html.Attributes.class("cap-description")], - [ - Tea_Html.text( - "VeriSimDB Admin runs in a sandboxed Gossamer webview. " ++ - "Grant capabilities below to enable database management features. " ++ - "Each token is time-limited and can be revoked at any time.", - ), - ], - ), - capabilityRow("network", Capabilities.Kind.network, model.networkCap), - capabilityRow("filesystem", Capabilities.Kind.filesystem, model.filesystemCap), - capabilityRow("clipboard", Capabilities.Kind.clipboard, model.clipboardCap), - Tea_Html.button( - [ - Tea_Html.Attributes.class("cap-dismiss"), - Tea_Html.Events.onClick(Msg.DismissCapPanel), - ], - [Tea_Html.text("Continue to Admin Panel")], - ), - ], - ) - } else { - Tea_Html.noNode - }, - - // --- VCL Console (top panel, full width) --- - Tea_Html.section( - [Tea_Html.Attributes.class("vcl-console")], - [ - Tea_Html.div( - [Tea_Html.Attributes.class("vcl-header")], - [ - Tea_Html.h2([], [Tea_Html.text("VCL Console")]), - Tea_Html.div( - [Tea_Html.Attributes.class("vcl-actions")], - [ - Tea_Html.button( - [ - Tea_Html.Attributes.class(model.vclExecuting ? "vcl-executing" : "vcl-execute"), - Tea_Html.Attributes.disabled(model.vclExecuting), - Tea_Html.Events.onClick(Msg.ExecuteVcl), - ], - [Tea_Html.text(model.vclExecuting ? "Executing..." : "Execute")], - ), - Tea_Html.button( - [Tea_Html.Events.onClick(Msg.CopyVcl)], - [Tea_Html.text("Copy VCL")], - ), - ], - ), - ], - ), - Tea_Html.div( - [Tea_Html.Attributes.class("vcl-body")], - [ - Tea_Html.textarea( - [ - Tea_Html.Attributes.class("vcl-input"), - Tea_Html.Attributes.placeholder("Enter VCL query... e.g. FETCH entity WHERE type = 'Document' PROOF EXISTENCE"), - Tea_Html.Attributes.value(model.vclInput), - Tea_Html.Events.onInput(value => Msg.VclInputChanged(value)), - ], - [], - ), - Tea_Html.div( - [Tea_Html.Attributes.class("vcl-result")], - [ - switch model.vclResult { - | Some(result) => - Tea_Html.pre( - [Tea_Html.Attributes.class("vcl-result-content")], - [Tea_Html.text(result)], - ) - | None => - Tea_Html.p( - [Tea_Html.Attributes.class("vcl-placeholder")], - [Tea_Html.text("Query results will appear here")], - ) - }, - ], - ), - ], - ), - ], - ), - - // --- Main content: Sidebar + Detail --- - Tea_Html.main( - [Tea_Html.Attributes.class("admin-main")], - [ - // Sidebar: octad browser (paginated) - Tea_Html.aside( - [Tea_Html.Attributes.class("octad-sidebar")], - [ - Tea_Html.div( - [Tea_Html.Attributes.class("sidebar-header")], - [ - Tea_Html.h2([], [Tea_Html.text("Octads")]), - Tea_Html.button( - [Tea_Html.Events.onClick(Msg.LoadOctads)], - [Tea_Html.text("Refresh")], - ), - ], - ), - Tea_Html.div( - [Tea_Html.Attributes.class("octad-list")], - Array.map(model.octads, octadCard)->Array.toList->List.toArray, - ), - // Pagination controls - Tea_Html.div( - [Tea_Html.Attributes.class("pagination")], - [ - Tea_Html.button( - [ - Tea_Html.Attributes.disabled(model.octadOffset == 0), - Tea_Html.Events.onClick(Msg.PrevPage), - ], - [Tea_Html.text("Prev")], - ), - Tea_Html.span( - [Tea_Html.Attributes.class("page-info")], - [ - Tea_Html.text( - `${Int.toString(model.octadOffset + 1)}-${Int.toString( - Math.Int.min( - model.octadOffset + model.octadLimit, - model.octadTotal, - ), - )} of ${Int.toString(model.octadTotal)}`, - ), - ], - ), - Tea_Html.button( - [ - Tea_Html.Attributes.disabled( - model.octadOffset + model.octadLimit >= model.octadTotal, - ), - Tea_Html.Events.onClick(Msg.NextPage), - ], - [Tea_Html.text("Next")], - ), - ], - ), - ], - ), - - // Main panel: entity detail with modality tabs - Tea_Html.section( - [Tea_Html.Attributes.class("detail-panel")], - [ - switch model.selectedEntity { - | Some(entityId) => - Tea_Html.div( - [], - [ - Tea_Html.div( - [Tea_Html.Attributes.class("entity-header")], - [ - Tea_Html.h2([], [Tea_Html.text(`Entity: ${entityId}`)]), - Tea_Html.div( - [Tea_Html.Attributes.class("entity-actions")], - [ - Tea_Html.button( - [Tea_Html.Events.onClick(Msg.LoadDrift(entityId))], - [Tea_Html.text("Check Drift")], - ), - Tea_Html.button( - [ - Tea_Html.Attributes.class("delete-button"), - Tea_Html.Events.onClick(Msg.DeleteOctad(entityId)), - ], - [Tea_Html.text("Delete")], - ), - ], - ), - ], - ), - // Modality tabs - Tea_Html.nav( - [Tea_Html.Attributes.class("modality-tabs")], - [ - modalityTab(Overview, model.detailTab, "Overview"), - modalityTab(Graph, model.detailTab, "Graph"), - modalityTab(Vector, model.detailTab, "Vector"), - modalityTab(Tensor, model.detailTab, "Tensor"), - modalityTab(Semantic, model.detailTab, "Semantic"), - modalityTab(Document, model.detailTab, "Document"), - modalityTab(Temporal, model.detailTab, "Temporal"), - modalityTab(Provenance, model.detailTab, "Provenance"), - modalityTab(Spatial, model.detailTab, "Spatial"), - ], - ), - // Modality content - Tea_Html.div( - [Tea_Html.Attributes.class("modality-content")], - [ - switch model.entityDetail { - | Some(detail) => - Tea_Html.pre( - [Tea_Html.Attributes.class("entity-detail-json")], - [Tea_Html.text(detail)], - ) - | None => - Tea_Html.p([], [Tea_Html.text("Loading entity detail...")]) - }, - ], - ), - // Drift indicator with normalise button - driftIndicator(model.driftStatus), - ], - ) - | None => - Tea_Html.div( - [Tea_Html.Attributes.class("overview")], - [ - Tea_Html.h2([], [Tea_Html.text("VeriSimDB Overview")]), - Tea_Html.p([], [ - Tea_Html.text( - `${Int.toString(Array.length(model.octads))} octads loaded`, - ), - ]), - switch model.orchStatus { - | Some(status) => - Tea_Html.pre( - [Tea_Html.Attributes.class("orch-status-display")], - [Tea_Html.text(status)], - ) - | None => Tea_Html.noNode - }, - ], - ) - }, - ], - ), - ], - ), - - // --- Telemetry Dashboard (bottom bar) --- - Tea_Html.footer( - [Tea_Html.Attributes.class("telemetry-dashboard")], - [ - Tea_Html.div( - [Tea_Html.Attributes.class("telemetry-header")], - [ - Tea_Html.h3([], [Tea_Html.text("Telemetry")]), - Tea_Html.button( - [Tea_Html.Events.onClick(Msg.LoadTelemetry)], - [Tea_Html.text("Refresh Telemetry")], - ), - ], - ), - switch model.telemetry { - | Some(telemetry) => - if telemetry.enabled { - Tea_Html.pre( - [Tea_Html.Attributes.class("telemetry-data")], - [Tea_Html.text(telemetry.raw)], - ) - } else { - Tea_Html.p( - [Tea_Html.Attributes.class("telemetry-disabled")], - [Tea_Html.text("Telemetry collection disabled. Enable with VERISIM_TELEMETRY=true.")], - ) - } - | None => - Tea_Html.p( - [Tea_Html.Attributes.class("telemetry-placeholder")], - [Tea_Html.text("Click 'Refresh Telemetry' to load aggregate metrics")], - ) - }, - ], - ), - ], - ) -} - -// --------------------------------------------------------------------------- -// Main -- TEA program registration -// --------------------------------------------------------------------------- - -/// Start the VeriSimDB Admin TEA application. -/// Mounts into the #app element in public/index.html. -let main = Tea_App.standardProgram({ - init: () => init(), - update: update, - view: view, - subscriptions: _model => Tea_Sub.none, -}) diff --git a/verisimdb/admin/src/Capabilities.res b/verisimdb/admin/src/Capabilities.res deleted file mode 100644 index 405857df..00000000 --- a/verisimdb/admin/src/Capabilities.res +++ /dev/null @@ -1,123 +0,0 @@ -// SPDX-License-Identifier: MPL-2.0 -// Copyright (c) 2026 Jonathan D.A. Jewell (hyperpolymath) - -/// Capabilities -- Gossamer capability token management for VeriSimDB Admin. -/// -/// VeriSimDB Admin requires three capabilities: -/// 1 = network -- Required for all VeriSimDB API calls (Rust core + Elixir layer) -/// 2 = filesystem -- Required for exporting query results and octad data -/// 5 = clipboard -- Required for copying VCL queries to clipboard -/// -/// Flow: -/// 1. App starts with NO capabilities (sandbox by default) -/// 2. User sees the capability grant panel -/// 3. User clicks "Grant Network" -> Gossamer shows consent dialog -/// 4. Runtime returns a token (float) valid for TTL seconds -/// 5. All subsequent API calls include the token in the IPC payload -/// 6. Token expires -> app must re-request or operations fail - -/// Capability kind identifiers matching the Gossamer runtime's internal enum. -/// These map to the `kind` field in `__gossamer_cap_grant` requests. -module Kind = { - /// Network access -- HTTP to VeriSimDB Rust core (port 8080) and Elixir - /// orchestration layer (port 4080). - let network = 1 - - /// Filesystem access -- exporting query results, octad snapshots, and - /// drift reports to disk. - let filesystem = 2 - - /// Clipboard access -- copying VCL queries and entity IDs to the system - /// clipboard. - let clipboard = 5 - - /// Human-readable name for a capability kind. - let toString = (kind: int): string => { - switch kind { - | 1 => "network" - | 2 => "filesystem" - | 5 => "clipboard" - | k => `unknown(${Int.toString(k)})` - } - } - - /// Description of why VeriSimDB Admin needs this capability. - let description = (kind: int): string => { - switch kind { - | 1 => "Connect to VeriSimDB Rust core and Elixir orchestration layer to manage octads, run VCL queries, and monitor drift." - | 2 => "Export query results, octad snapshots, and drift reports to local files." - | 5 => "Copy VCL queries and entity IDs to the system clipboard." - | _ => "Unknown capability." - } - } -} - -/// Request a capability token from the Gossamer runtime. -/// -/// This triggers Gossamer's consent dialog. The user must approve the -/// request before the runtime issues a token. Returns a promise that -/// resolves to the token value (float) on success. -/// -/// @param kind - The capability kind (use Kind.network, Kind.filesystem, etc.) -let requestCapability = (kind: int): promise => { - RuntimeBridge.invoke("__gossamer_cap_grant", {"kind": kind}) -} - -/// Request network capability -- needed for ALL VeriSimDB API calls. -/// -/// Without this token, no HTTP requests can be made to the Rust core -/// or Elixir orchestration layer. This is the first capability users -/// should grant. -let requestNetworkAccess = (): promise => { - requestCapability(Kind.network) -} - -/// Request filesystem capability -- needed for exporting data. -/// -/// Export operations write query results and octad snapshots to local -/// files in user-chosen directories. -let requestFilesystemAccess = (): promise => { - requestCapability(Kind.filesystem) -} - -/// Request clipboard capability -- needed for VCL copy operations. -/// -/// The VCL console allows copying queries and results to the clipboard -/// for use in other tools. -let requestClipboardAccess = (): promise => { - requestCapability(Kind.clipboard) -} - -/// Revoke a previously granted capability. -/// -/// This is the counterpart to requestCapability. After revocation, any -/// IPC calls using the old token will fail. The app should update its -/// UI to reflect the reduced permissions. -/// -/// @param kind - The capability kind to revoke -let revokeCapability = (kind: int): promise => { - RuntimeBridge.invoke("__gossamer_cap_revoke", {"kind": kind}) -} - -/// Check whether a token is still valid. -/// -/// Tokens expire after the TTL defined in gossamer.conf.json (default: -/// 3600 seconds). This lets the app proactively check and re-request -/// before a critical operation fails. -/// -/// @param token - The capability token to validate -let validateToken = (token: float): promise => { - RuntimeBridge.invoke("__gossamer_cap_validate", {"token": token}) -} - -/// Copy text to the system clipboard (requires clipboard capability). -/// -/// @param text - The text to copy -/// @param token - Valid clipboard capability token -let copyToClipboard = (text: string, token: float): promise => { - RuntimeBridge.invokeWithToken( - "__gossamer_clipboard_write", - {"text": text}, - token, - ) -} diff --git a/verisimdb/admin/src/Model.res b/verisimdb/admin/src/Model.res deleted file mode 100644 index a59f67db..00000000 --- a/verisimdb/admin/src/Model.res +++ /dev/null @@ -1,164 +0,0 @@ -// SPDX-License-Identifier: MPL-2.0 -// Copyright (c) 2026 Jonathan D.A. Jewell (hyperpolymath) - -/// Model -- Application state for the VeriSimDB Admin panel. -/// -/// Holds the complete UI state including server connection status, octad -/// entity listings, VCL console state, drift detection results, telemetry -/// metrics, and Gossamer capability tokens. -/// -/// VeriSimDB's octad model: each entity exists simultaneously across 8 -/// modalities (Graph, Vector, Tensor, Semantic, Document, Temporal, -/// Provenance, Spatial). The admin panel provides visibility into all -/// modalities and their drift status. - -/// Connection status to the VeriSimDB backend (Rust core + Elixir layer). -type serverStatus = - | /// Successfully connected to the VeriSimDB Rust core. - Connected - | /// No connection -- server unreachable or not started. - Disconnected - | /// Connection attempt in progress. - Connecting - -/// Summary of an octad entity for the sidebar list. -/// Contains just enough information for browsing; full detail is loaded -/// on selection via `getEntity`. -type octadSummary = { - /// Unique octad entity identifier (UUID). - id: string, - /// Human-readable title (from the Document modality). - title: string, - /// Number of active modalities (out of 8). - activeModalities: int, - /// Overall drift score (0.0 = no drift, 1.0 = maximum drift). - driftScore: float, -} - -/// Drift information for a specific entity. -/// Each field represents the divergence between two modalities. -type driftInfo = { - /// ID of the entity this drift info belongs to. - entityId: string, - /// Embedding-to-semantic content divergence. - semanticVectorDrift: float, - /// Graph structure vs document content divergence. - graphDocumentDrift: float, - /// Version history consistency issues. - temporalConsistencyDrift: float, - /// Tensor representation divergence. - tensorDrift: float, - /// Type constraint violations. - schemaDrift: float, - /// Overall data quality metric. - qualityDrift: float, -} - -/// Telemetry aggregate from the Elixir orchestration layer. -type telemetryData = { - /// Raw JSON string of the full telemetry report. - raw: string, - /// Whether telemetry collection is enabled on the server. - enabled: bool, -} - -/// Capability token status for Gossamer security. -/// Each capability must be explicitly granted by the runtime before use. -type capabilityStatus = - | /// Not yet requested from the runtime. - NotRequested - | /// Request sent, awaiting runtime grant. - Pending - | /// Granted with a token. The float is the token value. - Granted(float) - | /// Runtime denied the capability request. - Denied - -/// The active tab in the entity detail view, selecting which modality -/// to display prominently. -type detailTab = - | /// Show all modalities in a summary grid. - Overview - | /// Graph triples and property graph edges. - Graph - | /// Vector embedding visualisation. - Vector - | /// Tensor multi-dimensional representation. - Tensor - | /// Semantic type annotations and proof blobs. - Semantic - | /// Full-text searchable content. - Document - | /// Version history and time-series. - Temporal - | /// Origin tracking and transformation chain. - Provenance - | /// Geospatial coordinates and geometries. - Spatial - -/// Complete application state. -type model = { - /// Current server connection status. - status: serverStatus, - /// Paginated list of octad entity summaries for the sidebar. - octads: array, - /// Total octad count (for pagination display). - octadTotal: int, - /// Current pagination offset. - octadOffset: int, - /// Page size for octad listing. - octadLimit: int, - /// Currently selected entity ID for detail view. - selectedEntity: option, - /// Full JSON detail of the selected entity (all 8 modalities). - entityDetail: option, - /// Active tab in the entity detail view. - detailTab: detailTab, - /// VCL console: current input text. - vclInput: string, - /// VCL console: result of the last executed query. - vclResult: option, - /// VCL console: whether a query is currently executing. - vclExecuting: bool, - /// Drift status for the currently selected entity. - driftStatus: option, - /// Aggregate telemetry data from the Elixir layer. - telemetry: option, - /// Orchestration layer status (raw JSON). - orchStatus: option, - /// Network capability token -- required for ALL API calls. - networkCap: capabilityStatus, - /// Filesystem capability token -- required for exports. - filesystemCap: capabilityStatus, - /// Clipboard capability token -- required for VCL copying. - clipboardCap: capabilityStatus, - /// Error message to display in the UI, if any. - error: option, - /// Whether the capability grant panel is visible. - showCapPanel: bool, -} - -/// Initial application state. Starts with no capabilities granted, -/// forcing the user to explicitly authorise network, filesystem, and -/// clipboard access through the Gossamer capability token system. -let initial: model = { - status: Disconnected, - octads: [], - octadTotal: 0, - octadOffset: 0, - octadLimit: 50, - selectedEntity: None, - entityDetail: None, - detailTab: Overview, - vclInput: "", - vclResult: None, - vclExecuting: false, - driftStatus: None, - telemetry: None, - orchStatus: None, - networkCap: NotRequested, - filesystemCap: NotRequested, - clipboardCap: NotRequested, - error: None, - showCapPanel: true, -} diff --git a/verisimdb/admin/src/Msg.res b/verisimdb/admin/src/Msg.res deleted file mode 100644 index 9c42938c..00000000 --- a/verisimdb/admin/src/Msg.res +++ /dev/null @@ -1,92 +0,0 @@ -// SPDX-License-Identifier: MPL-2.0 -// Copyright (c) 2026 Jonathan D.A. Jewell (hyperpolymath) - -/// Msg -- Message type for the VeriSimDB Admin TEA architecture. -/// -/// Every user interaction and async result flows through this type. -/// Messages are dispatched by the view and processed by the update -/// function in App.res. - -/// All messages that can occur in the VeriSimDB Admin panel. -type msg = - // --- Server health --- - | /// User clicked "Check Health" or auto-poll triggered. - CheckHealth - | /// Health check response arrived from the Rust core. - HealthResult(result) - - // --- VCL console --- - | /// User typed in the VCL console input. - VclInputChanged(string) - | /// User pressed Execute or Ctrl+Enter in the VCL console. - ExecuteVcl - | /// VCL query result arrived. - VclResult(result) - | /// User clicked "Copy VCL" to copy the current query to clipboard. - CopyVcl - - // --- Octad browser --- - | /// Load or refresh the octad entity list. - LoadOctads - | /// Octad list response arrived. - OctadsLoaded(result) - | /// User selected an entity in the sidebar. - SelectEntity(string) - | /// Entity detail response arrived (full 8-modality snapshot). - EntityLoaded(result) - | /// User switched the detail tab to a different modality. - SwitchDetailTab(Model.detailTab) - | /// Navigate pagination forward. - NextPage - | /// Navigate pagination backward. - PrevPage - - // --- Entity CRUD --- - | /// User submitted the create octad form. - CreateOctad(string) - | /// Create response arrived. - OctadCreated(result) - | /// User confirmed deletion of an entity. - DeleteOctad(string) - | /// Delete response arrived. - OctadDeleted(result) - - // --- Drift detection --- - | /// Load drift status for the currently selected entity. - LoadDrift(string) - | /// Drift status response arrived. - DriftLoaded(result) - | /// User clicked "Normalise" to trigger self-normalisation. - TriggerNormalise(string) - | /// Normalisation response arrived. - NormaliseResult(result) - - // --- Telemetry --- - | /// Load aggregate telemetry from the Elixir orchestration layer. - LoadTelemetry - | /// Telemetry response arrived. - TelemetryLoaded(result) - - // --- Orchestration status --- - | /// Load orchestration layer status. - LoadOrchStatus - | /// Orchestration status response arrived. - OrchStatusLoaded(result) - - // --- Gossamer capability tokens --- - | /// User clicked "Grant" on a capability in the cap panel. - RequestCapability(string) - | /// Gossamer runtime granted a capability token. - CapGranted(string, float) - | /// Gossamer runtime revoked or denied a capability token. - CapRevoked(string) - | /// User dismissed the capability panel. - DismissCapPanel - | /// User reopened the capability panel. - ShowCapPanel - - // --- UI --- - | /// Clear the current error message. - ClearError - | /// No-op message (used for commands that have no followup). - NoOp diff --git a/verisimdb/admin/src/RuntimeBridge.res b/verisimdb/admin/src/RuntimeBridge.res deleted file mode 100644 index 48a33248..00000000 --- a/verisimdb/admin/src/RuntimeBridge.res +++ /dev/null @@ -1,116 +0,0 @@ -// SPDX-License-Identifier: MPL-2.0 -// Copyright (c) 2026 Jonathan D.A. Jewell (hyperpolymath) - -/// RuntimeBridge -- Gossamer-native IPC bridge for VeriSimDB Admin. -/// -/// Gossamer-only bridge (no Tauri fallback). VeriSimDB Admin is the SECOND -/// app built natively for Gossamer, following Burble Admin. The IPC pattern -/// is identical: all communication flows through `window.__gossamer_invoke` -/// injected by the Gossamer runtime during webview initialisation. -/// -/// Capability tokens: Every privileged operation (network requests to the -/// VeriSimDB Rust core and Elixir orchestration layer, filesystem access -/// for exports, clipboard for VCL copying) requires a valid token obtained -/// via `__gossamer_cap_grant`. Tokens are time-limited (TTL from config) -/// and can be revoked by the runtime at any time. - -// --------------------------------------------------------------------------- -// Gossamer runtime detection -// --------------------------------------------------------------------------- - -/// Check whether the Gossamer runtime is available in this webview. -/// Returns true when `window.__gossamer_invoke` has been injected by -/// the gossamer_channel_open() call during webview initialisation. -%%raw(` -function isGossamerRuntime() { - return typeof window !== 'undefined' - && typeof window.__gossamer_invoke === 'function'; -} -`) -@val external isGossamerRuntime: unit => bool = "isGossamerRuntime" - -/// Raw Gossamer IPC call. Sends a command name and JSON payload to the -/// Gossamer runtime and returns a promise with the response. -%%raw(` -function gossamerInvoke(cmd, args) { - return window.__gossamer_invoke(cmd, args); -} -`) -@val external gossamerInvoke: (string, 'a) => promise<'b> = "gossamerInvoke" - -// --------------------------------------------------------------------------- -// Runtime type (Gossamer-only, no Tauri path) -// --------------------------------------------------------------------------- - -/// The runtime environment. For VeriSimDB Admin, this is always Gossamer -/// or an error state (dev browser without the runtime). -type runtime = - | /// Running inside the Gossamer webview shell (production). - Gossamer - | /// Running in a plain browser (development only -- most features disabled). - BrowserDev - -/// Detect the current runtime environment. -let detectRuntime = (): runtime => { - if isGossamerRuntime() { - Gossamer - } else { - BrowserDev - } -} - -// --------------------------------------------------------------------------- -// Unified invoke -- Gossamer-native with dev fallback -// --------------------------------------------------------------------------- - -/// Invoke a Gossamer IPC command. -/// -/// In production (Gossamer runtime), this calls `window.__gossamer_invoke`. -/// In development (browser), this rejects with a descriptive error so the -/// developer knows to run inside Gossamer. -/// -/// All command modules (VeriSimDbCmd, Capabilities) use this function. -let invoke = (cmd: string, args: 'a): promise<'b> => { - if isGossamerRuntime() { - gossamerInvoke(cmd, args) - } else { - Promise.reject( - JsError.throwWithMessage( - `Gossamer runtime required -- "${cmd}" cannot run in a plain browser. ` ++ - `Launch via: gossamer run --config gossamer.conf.json`, - ), - ) - } -} - -/// Invoke a command that requires a capability token. -/// -/// This is the security-critical path. The token is included in the IPC -/// payload so the Gossamer runtime can verify the caller holds the -/// required capability before executing the command. -/// -/// @param cmd - The IPC command name -/// @param args - The command payload -/// @param token - The capability token (obtained from __gossamer_cap_grant) -let invokeWithToken = (cmd: string, args: 'a, token: float): promise<'b> => { - if isGossamerRuntime() { - gossamerInvoke(cmd, {"__cap_token": token, "payload": args}) - } else { - Promise.reject( - JsError.throwWithMessage( - `Gossamer runtime required -- "${cmd}" needs a capability token`, - ), - ) - } -} - -/// Check whether the Gossamer runtime is available. -let hasRuntime = (): bool => isGossamerRuntime() - -/// Human-readable runtime name for display in the UI. -let runtimeName = (): string => { - switch detectRuntime() { - | Gossamer => "Gossamer" - | BrowserDev => "Browser (dev)" - } -} diff --git a/verisimdb/admin/src/VeriSimDbCmd.res b/verisimdb/admin/src/VeriSimDbCmd.res deleted file mode 100644 index 90b3498c..00000000 --- a/verisimdb/admin/src/VeriSimDbCmd.res +++ /dev/null @@ -1,207 +0,0 @@ -// SPDX-License-Identifier: MPL-2.0 -// Copyright (c) 2026 Jonathan D.A. Jewell (hyperpolymath) - -/// VeriSimDbCmd -- Backend command dispatch for the VeriSimDB Admin panel. -/// -/// Each function wraps a Gossamer IPC call to the VeriSimDB backend. Commands -/// target either the Rust core API (default port 8080) or the Elixir -/// orchestration layer (default port 4080). All commands require a valid -/// network capability token obtained from the Gossamer runtime via -/// `Capabilities.requestNetworkAccess()`. -/// -/// The commands map to VeriSimDB's REST API: -/// Rust core (port 8080, prefix /api/v1): -/// - Health: GET /health -/// - VCL: POST /vcl/execute -/// - Octads: GET /octads, GET /octads/{id}, POST /octads, DELETE /octads/{id} -/// - Drift: GET /drift/entity/{id} -/// - Normaliser: POST /normalizer/trigger/{id} -/// -/// Elixir orchestration (port 4080): -/// - Telemetry: GET /telemetry -/// - Orch status: GET /status -/// -/// Gossamer acts as the network proxy -- the webview never makes direct -/// HTTP calls. Instead, each command goes through IPC to the Gossamer -/// Zig runtime, which holds the network capability and forwards the -/// request to the VeriSimDB backend. - -/// Base URL for the VeriSimDB Rust core API. -/// In production this comes from server config; defaults to local dev port. -let _rustBaseUrl = "http://localhost:8080/api/v1" - -/// Base URL for the VeriSimDB Elixir orchestration layer. -/// Runs on a separate port from the Rust core. -let _elixirBaseUrl = "http://localhost:4080" - -// --------------------------------------------------------------------------- -// Health -// --------------------------------------------------------------------------- - -/// Check the VeriSimDB Rust core health endpoint. -/// -/// Maps to: GET /health -/// Returns the server's health status including uptime and version. -let checkHealth = (token: float): promise => { - RuntimeBridge.invokeWithToken( - "verisimdb_check_health", - {"url": `${_rustBaseUrl}/health`}, - token, - ) -} - -// --------------------------------------------------------------------------- -// VCL Console -// --------------------------------------------------------------------------- - -/// Execute a VCL (VeriSim Consonance Language) query against the database. -/// -/// Maps to: POST /vcl/execute -/// VCL supports octad queries across all 8 modalities with proof -/// generation and drift-aware consistency. -/// -/// @param query - The VCL query string to execute -let queryVcl = (query: string, token: float): promise => { - RuntimeBridge.invokeWithToken( - "verisimdb_query_vcl", - {"url": `${_rustBaseUrl}/vcl/execute`, "query": query}, - token, - ) -} - -// --------------------------------------------------------------------------- -// Octad management -// --------------------------------------------------------------------------- - -/// List octad entities with pagination. -/// -/// Maps to: GET /octads?limit=N&offset=M -/// Returns a JSON array of octad entity summaries (ID, title, modality -/// status flags, drift score). -/// -/// @param limit - Maximum number of entities to return -/// @param offset - Offset for pagination -let listOctads = (limit: int, offset: int, token: float): promise => { - RuntimeBridge.invokeWithToken( - "verisimdb_list_octads", - {"url": `${_rustBaseUrl}/octads?limit=${Int.toString(limit)}&offset=${Int.toString(offset)}`}, - token, - ) -} - -/// Get a single octad entity with full detail across all 8 modalities. -/// -/// Maps to: GET /octads/{id} -/// Returns the complete octad snapshot: graph triples, vector embedding, -/// tensor data, semantic annotations, document content, temporal versions, -/// provenance chain, and spatial coordinates. -/// -/// @param id - The octad entity UUID -let getEntity = (id: string, token: float): promise => { - RuntimeBridge.invokeWithToken( - "verisimdb_get_entity", - {"url": `${_rustBaseUrl}/octads/${id}`}, - token, - ) -} - -/// Create a new octad entity. -/// -/// Maps to: POST /octads -/// Creates an entity with the provided modality data. At minimum, -/// a document title is required. Other modalities are populated -/// automatically or can be provided explicitly. -/// -/// @param entityJson - JSON string with octad input fields -let createOctad = (entityJson: string, token: float): promise => { - RuntimeBridge.invokeWithToken( - "verisimdb_create_octad", - {"url": `${_rustBaseUrl}/octads`, "body": entityJson}, - token, - ) -} - -/// Delete an octad entity. -/// -/// Maps to: DELETE /octads/{id} -/// Removes the entity and all its modality data. This is irreversible -/// (unless temporal versioning provides recovery). -/// -/// @param id - The octad entity UUID to delete -let deleteOctad = (id: string, token: float): promise => { - RuntimeBridge.invokeWithToken( - "verisimdb_delete_octad", - {"url": `${_rustBaseUrl}/octads/${id}`, "method": "DELETE"}, - token, - ) -} - -// --------------------------------------------------------------------------- -// Drift detection -// --------------------------------------------------------------------------- - -/// Get the drift status for a specific entity. -/// -/// Maps to: GET /drift/entity/{id} -/// Returns per-modality drift scores: semantic_vector_drift, -/// graph_document_drift, temporal_consistency_drift, tensor_drift, -/// schema_drift, quality_drift. -/// -/// @param id - The octad entity UUID -let getDrift = (id: string, token: float): promise => { - RuntimeBridge.invokeWithToken( - "verisimdb_get_drift", - {"url": `${_rustBaseUrl}/drift/entity/${id}`}, - token, - ) -} - -/// Trigger normalisation for a drifted entity. -/// -/// Maps to: POST /normalizer/trigger/{id} -/// The normaliser identifies the most authoritative modality, -/// regenerates drifted modalities from it, validates consistency, -/// and updates all modalities atomically. -/// -/// @param id - The octad entity UUID to normalise -let triggerNormalise = (id: string, token: float): promise => { - RuntimeBridge.invokeWithToken( - "verisimdb_trigger_normalise", - {"url": `${_rustBaseUrl}/normalizer/trigger/${id}`, "method": "POST"}, - token, - ) -} - -// --------------------------------------------------------------------------- -// Telemetry (Elixir orchestration layer) -// --------------------------------------------------------------------------- - -/// Get aggregate telemetry from the Elixir orchestration layer. -/// -/// Maps to: GET /telemetry (port 4080) -/// Returns opt-in aggregate metrics: modality heatmap, query patterns, -/// drift reports, performance summary, federation health. No PII. -let getTelemetry = (token: float): promise => { - RuntimeBridge.invokeWithToken( - "verisimdb_get_telemetry", - {"url": `${_elixirBaseUrl}/telemetry`}, - token, - ) -} - -// --------------------------------------------------------------------------- -// Orchestration status (Elixir layer) -// --------------------------------------------------------------------------- - -/// Get the orchestration layer status. -/// -/// Maps to: GET /status (port 4080) -/// Returns consensus state, federation adapter count, and telemetry -/// enabled flag from the Elixir OTP supervision tree. -let getOrchStatus = (token: float): promise => { - RuntimeBridge.invokeWithToken( - "verisimdb_get_orch_status", - {"url": `${_elixirBaseUrl}/status`}, - token, - ) -} diff --git a/verisimdb/admin/src/styles.css b/verisimdb/admin/src/styles.css deleted file mode 100644 index d114048a..00000000 --- a/verisimdb/admin/src/styles.css +++ /dev/null @@ -1,534 +0,0 @@ -/* SPDX-License-Identifier: MPL-2.0 */ -/* Copyright (c) 2026 Jonathan D.A. Jewell (hyperpolymath) */ - -/* VeriSimDB Admin -- Gossamer-native admin panel stylesheet. - * - * Layout: - * Header (top) - * VCL Console (full width, below header) - * Main: Sidebar (octad browser) | Detail (entity + modality tabs + drift) - * Footer: Telemetry dashboard (bottom bar) - */ - -:root { - --bg-primary: #0d1117; - --bg-secondary: #161b22; - --bg-tertiary: #21262d; - --text-primary: #c9d1d9; - --text-secondary: #8b949e; - --accent-blue: #58a6ff; - --accent-green: #3fb950; - --accent-yellow: #d29922; - --accent-red: #f85149; - --accent-purple: #bc8cff; - --border-color: #30363d; - --radius: 6px; - --font-mono: 'JetBrains Mono', 'Fira Code', 'Cascadia Code', monospace; - --font-sans: -apple-system, BlinkMacSystemFont, 'Segoe UI', Helvetica, Arial, sans-serif; -} - -* { - margin: 0; - padding: 0; - box-sizing: border-box; -} - -body { - background: var(--bg-primary); - color: var(--text-primary); - font-family: var(--font-sans); - font-size: 14px; - line-height: 1.5; -} - -/* -- Layout ----------------------------------------------------------- */ - -.verisimdb-admin { - display: flex; - flex-direction: column; - height: 100vh; - overflow: hidden; -} - -/* -- Header ----------------------------------------------------------- */ - -.admin-header { - display: flex; - justify-content: space-between; - align-items: center; - padding: 12px 20px; - background: var(--bg-secondary); - border-bottom: 1px solid var(--border-color); -} - -.admin-header h1 { - font-size: 18px; - font-weight: 600; - color: var(--accent-blue); -} - -.header-controls { - display: flex; - align-items: center; - gap: 12px; -} - -.runtime-badge { - padding: 4px 8px; - background: var(--bg-tertiary); - border-radius: var(--radius); - font-size: 12px; - color: var(--text-secondary); -} - -/* -- Status indicators ------------------------------------------------ */ - -.status-connected { - color: var(--accent-green); - font-weight: 600; -} - -.status-disconnected { - color: var(--accent-red); - font-weight: 600; -} - -.status-connecting { - color: var(--accent-yellow); - font-weight: 600; -} - -/* -- Error bar -------------------------------------------------------- */ - -.error-bar { - display: flex; - justify-content: space-between; - align-items: center; - padding: 8px 20px; - background: #2d1215; - border-bottom: 1px solid var(--accent-red); - color: var(--accent-red); - font-size: 13px; -} - -/* -- Buttons ---------------------------------------------------------- */ - -button { - padding: 6px 12px; - background: var(--bg-tertiary); - color: var(--text-primary); - border: 1px solid var(--border-color); - border-radius: var(--radius); - cursor: pointer; - font-size: 13px; - transition: background 0.15s; -} - -button:hover:not(:disabled) { - background: var(--border-color); -} - -button:disabled { - opacity: 0.5; - cursor: not-allowed; -} - -.delete-button { - color: var(--accent-red); - border-color: var(--accent-red); -} - -.delete-button:hover:not(:disabled) { - background: #2d1215; -} - -.normalise-button { - background: #0d2818; - color: var(--accent-green); - border-color: var(--accent-green); - margin-top: 8px; -} - -/* -- Capability panel ------------------------------------------------- */ - -.cap-panel { - padding: 20px; - background: var(--bg-secondary); - border-bottom: 1px solid var(--border-color); -} - -.cap-panel h2 { - font-size: 16px; - margin-bottom: 8px; - color: var(--accent-purple); -} - -.cap-description { - color: var(--text-secondary); - margin-bottom: 16px; - font-size: 13px; -} - -.cap-row { - display: flex; - justify-content: space-between; - align-items: center; - padding: 10px 12px; - background: var(--bg-tertiary); - border-radius: var(--radius); - margin-bottom: 8px; -} - -.cap-info strong { - color: var(--text-primary); -} - -.cap-info p { - color: var(--text-secondary); - font-size: 12px; - margin: 2px 0; -} - -.cap-not-requested { color: var(--text-secondary); } -.cap-pending { color: var(--accent-yellow); } -.cap-granted { color: var(--accent-green); } -.cap-denied { color: var(--accent-red); } - -.cap-dismiss { - margin-top: 12px; - background: var(--accent-blue); - color: #fff; - border: none; - padding: 8px 16px; - font-weight: 600; -} - -/* -- VCL Console ------------------------------------------------------ */ - -.vcl-console { - border-bottom: 1px solid var(--border-color); - background: var(--bg-secondary); -} - -.vcl-header { - display: flex; - justify-content: space-between; - align-items: center; - padding: 8px 20px; - border-bottom: 1px solid var(--border-color); -} - -.vcl-header h2 { - font-size: 14px; - color: var(--accent-purple); -} - -.vcl-actions { - display: flex; - gap: 8px; -} - -.vcl-execute { - background: #0d2818; - color: var(--accent-green); - border-color: var(--accent-green); - font-weight: 600; -} - -.vcl-executing { - background: var(--bg-tertiary); - color: var(--accent-yellow); - border-color: var(--accent-yellow); -} - -.vcl-body { - display: flex; - height: 160px; -} - -.vcl-input { - flex: 1; - background: var(--bg-primary); - color: var(--text-primary); - border: none; - border-right: 1px solid var(--border-color); - padding: 12px; - font-family: var(--font-mono); - font-size: 13px; - resize: none; - outline: none; -} - -.vcl-input::placeholder { - color: var(--text-secondary); -} - -.vcl-result { - flex: 1; - overflow: auto; - padding: 12px; -} - -.vcl-result-content { - font-family: var(--font-mono); - font-size: 12px; - white-space: pre-wrap; - word-break: break-word; -} - -.vcl-placeholder { - color: var(--text-secondary); - font-style: italic; -} - -/* -- Main layout ------------------------------------------------------ */ - -.admin-main { - display: flex; - flex: 1; - overflow: hidden; -} - -/* -- Octad sidebar ---------------------------------------------------- */ - -.octad-sidebar { - width: 280px; - min-width: 280px; - background: var(--bg-secondary); - border-right: 1px solid var(--border-color); - display: flex; - flex-direction: column; -} - -.sidebar-header { - display: flex; - justify-content: space-between; - align-items: center; - padding: 12px 16px; - border-bottom: 1px solid var(--border-color); -} - -.sidebar-header h2 { - font-size: 14px; -} - -.octad-list { - flex: 1; - overflow-y: auto; - padding: 8px; -} - -.octad-card { - padding: 10px 12px; - background: var(--bg-tertiary); - border-radius: var(--radius); - margin-bottom: 6px; - cursor: pointer; - border: 1px solid transparent; - transition: border-color 0.15s; -} - -.octad-card:hover { - border-color: var(--accent-blue); -} - -.octad-card h3 { - font-size: 13px; - font-weight: 600; - margin-bottom: 4px; - white-space: nowrap; - overflow: hidden; - text-overflow: ellipsis; -} - -.octad-meta { - font-size: 11px; - color: var(--text-secondary); -} - -.drift-low { color: var(--accent-green); font-size: 11px; } -.drift-medium { color: var(--accent-yellow); font-size: 11px; } -.drift-high { color: var(--accent-red); font-size: 11px; } - -/* -- Pagination ------------------------------------------------------- */ - -.pagination { - display: flex; - justify-content: space-between; - align-items: center; - padding: 8px 12px; - border-top: 1px solid var(--border-color); -} - -.page-info { - font-size: 11px; - color: var(--text-secondary); -} - -/* -- Detail panel ----------------------------------------------------- */ - -.detail-panel { - flex: 1; - overflow-y: auto; - padding: 16px 20px; -} - -.entity-header { - display: flex; - justify-content: space-between; - align-items: center; - margin-bottom: 12px; -} - -.entity-header h2 { - font-size: 16px; - font-family: var(--font-mono); - color: var(--accent-blue); -} - -.entity-actions { - display: flex; - gap: 8px; -} - -/* -- Modality tabs ---------------------------------------------------- */ - -.modality-tabs { - display: flex; - gap: 4px; - margin-bottom: 12px; - flex-wrap: wrap; -} - -.modality-tab { - padding: 4px 10px; - font-size: 12px; - border-radius: var(--radius); -} - -.tab-active { - background: var(--accent-blue); - color: #fff; - border-color: var(--accent-blue); -} - -.tab-inactive { - background: var(--bg-tertiary); -} - -/* -- Entity detail ---------------------------------------------------- */ - -.modality-content { - margin-bottom: 16px; -} - -.entity-detail-json { - font-family: var(--font-mono); - font-size: 12px; - background: var(--bg-primary); - padding: 12px; - border-radius: var(--radius); - border: 1px solid var(--border-color); - white-space: pre-wrap; - word-break: break-word; - max-height: 400px; - overflow-y: auto; -} - -/* -- Drift indicator -------------------------------------------------- */ - -.drift-indicator { - background: var(--bg-secondary); - padding: 12px; - border-radius: var(--radius); - border: 1px solid var(--border-color); -} - -.drift-indicator h3 { - font-size: 14px; - margin-bottom: 8px; - color: var(--accent-yellow); -} - -.drift-row { - display: flex; - align-items: center; - margin-bottom: 4px; -} - -.drift-label { - width: 180px; - font-size: 12px; - color: var(--text-secondary); -} - -.drift-bar { - flex: 1; - padding: 2px 8px; - border-radius: 3px; - font-family: var(--font-mono); - font-size: 11px; -} - -.drift-bar-low { background: #0d2818; color: var(--accent-green); } -.drift-bar-medium { background: #2d2200; color: var(--accent-yellow); } -.drift-bar-high { background: #2d1215; color: var(--accent-red); } - -.drift-empty { - color: var(--text-secondary); - font-style: italic; - font-size: 13px; -} - -/* -- Overview --------------------------------------------------------- */ - -.overview h2 { - margin-bottom: 8px; -} - -.orch-status-display { - font-family: var(--font-mono); - font-size: 12px; - background: var(--bg-primary); - padding: 12px; - border-radius: var(--radius); - border: 1px solid var(--border-color); - margin-top: 12px; -} - -/* -- Telemetry dashboard (bottom bar) --------------------------------- */ - -.telemetry-dashboard { - background: var(--bg-secondary); - border-top: 1px solid var(--border-color); - padding: 8px 20px; - max-height: 200px; - overflow-y: auto; -} - -.telemetry-header { - display: flex; - justify-content: space-between; - align-items: center; - margin-bottom: 6px; -} - -.telemetry-header h3 { - font-size: 13px; - color: var(--accent-purple); -} - -.telemetry-data { - font-family: var(--font-mono); - font-size: 11px; - white-space: pre-wrap; - word-break: break-word; - color: var(--text-secondary); -} - -.telemetry-disabled, -.telemetry-placeholder { - color: var(--text-secondary); - font-style: italic; - font-size: 12px; -} diff --git a/verisimdb/benches/Cargo.toml b/verisimdb/benches/Cargo.toml deleted file mode 100644 index fd619903..00000000 --- a/verisimdb/benches/Cargo.toml +++ /dev/null @@ -1,35 +0,0 @@ -# SPDX-License-Identifier: MPL-2.0 - -[package] -name = "verisimdb-benchmarks" -version = "0.1.0" -edition = "2021" -publish = false - -[[bench]] -name = "modality_benchmarks" -path = "modality_benchmarks.rs" -harness = false - -[[bench]] -name = "throughput_benchmarks" -path = "throughput_benchmarks.rs" -harness = false - -[dependencies] -criterion = { version = "0.5", features = ["async_tokio", "html_reports"] } -tokio = { version = "1", features = ["full"] } -uuid = { version = "1.11", features = ["v4"] } -futures = "0.3" - -verisim-document = { path = "../rust-core/verisim-document" } -verisim-graph = { path = "../rust-core/verisim-graph" } -verisim-vector = { path = "../rust-core/verisim-vector" } -verisim-tensor = { path = "../rust-core/verisim-tensor" } -verisim-semantic = { path = "../rust-core/verisim-semantic" } -verisim-temporal = { path = "../rust-core/verisim-temporal" } -verisim-provenance = { path = "../rust-core/verisim-provenance" } -verisim-spatial = { path = "../rust-core/verisim-spatial" } -verisim-octad = { path = "../rust-core/verisim-octad" } -verisim-drift = { path = "../rust-core/verisim-drift" } -verisim-normalizer = { path = "../rust-core/verisim-normalizer" } diff --git a/verisimdb/benches/modality_benchmarks.rs b/verisimdb/benches/modality_benchmarks.rs deleted file mode 100644 index e280641f..00000000 --- a/verisimdb/benches/modality_benchmarks.rs +++ /dev/null @@ -1,624 +0,0 @@ -// SPDX-License-Identifier: MPL-2.0 -//! Performance benchmarks for VeriSimDB modality stores - -use criterion::{black_box, criterion_group, criterion_main, BenchmarkId, Criterion, Throughput}; -use std::collections::HashMap; -use std::sync::Arc; -use tokio::runtime::Runtime; - -use verisim_document::{Document, DocumentStore, TantivyDocumentStore}; -use verisim_drift::{DriftDetector, DriftThresholds, DriftType}; -use verisim_graph::{GraphEdge, GraphNode, GraphObject, GraphStore, SimpleGraphStore}; -use verisim_octad::{ - OctadConfig, OctadDocumentInput, OctadId, OctadInput, OctadSnapshot, OctadStore, - OctadVectorInput, OctadSemanticInput, InMemoryOctadStore, -}; -use verisim_provenance::InMemoryProvenanceStore; -use verisim_semantic::{ - InMemorySemanticStore, ProofBlob, ProofType, SemanticStore, SemanticType, -}; -use verisim_spatial::InMemorySpatialStore; -use verisim_temporal::{InMemoryVersionStore, TemporalStore}; -use verisim_tensor::{InMemoryTensorStore, ReduceOp, Tensor, TensorStore}; -use verisim_vector::{DistanceMetric, Embedding, HnswConfig, HnswVectorStore, VectorStore}; - -// ============================================================================ -// Document Store Benchmarks -// ============================================================================ - -fn bench_document_create(c: &mut Criterion) { - let rt = Runtime::new().unwrap(); - let mut group = c.benchmark_group("document"); - - group.bench_function("create_document", |b| { - let store = TantivyDocumentStore::in_memory().unwrap(); - b.to_async(&rt).iter(|| async { - let doc = Document::new("test-id", "Benchmark Title", "Benchmark body content for testing indexing performance."); - black_box(store.index(&doc).await.unwrap()) - }); - }); - - group.finish(); -} - -fn bench_document_search(c: &mut Criterion) { - let rt = Runtime::new().unwrap(); - let store = TantivyDocumentStore::in_memory().unwrap(); - - // Index 1000 documents - rt.block_on(async { - for i in 0..1000 { - let doc = Document::new( - format!("doc-{}", i), - format!("Document {}", i), - format!("This is document number {} with searchable content about machine learning and databases.", i), - ); - store.index(&doc).await.unwrap(); - } - store.commit().await.unwrap(); - }); - - let mut group = c.benchmark_group("document"); - group.throughput(Throughput::Elements(1000)); - - group.bench_function("search_text", |b| { - b.to_async(&rt).iter(|| async { - black_box(store.search("machine learning", 10).await.unwrap()) - }); - }); - - group.finish(); -} - -// ============================================================================ -// Vector Store Benchmarks -// ============================================================================ - -fn bench_vector_insert(c: &mut Criterion) { - let rt = Runtime::new().unwrap(); - let mut group = c.benchmark_group("vector"); - - for dim in [128, 384, 768].iter() { - group.bench_with_input(BenchmarkId::new("upsert", dim), dim, |b, &dim| { - let store = HnswVectorStore::new(dim, DistanceMetric::Cosine, HnswConfig::default()); - let mut counter = 0u64; - - b.to_async(&rt).iter(|| { - counter += 1; - let embedding = Embedding { - id: format!("vec-{}", counter), - vector: vec![0.5; dim], - metadata: HashMap::new(), - }; - let store_ref = &store; - async move { - black_box(store_ref.upsert(&embedding).await.unwrap()) - } - }); - }); - } - - group.finish(); -} - -fn bench_vector_search(c: &mut Criterion) { - let rt = Runtime::new().unwrap(); - let mut group = c.benchmark_group("vector"); - - for dim in [128, 384, 768].iter() { - group.bench_with_input(BenchmarkId::new("search", dim), dim, |b, &dim| { - let store = HnswVectorStore::new(dim, DistanceMetric::Cosine, HnswConfig::default()); - - // Insert 10000 vectors - rt.block_on(async { - for i in 0..10000 { - let mut vec_data = vec![0.0f32; dim]; - vec_data[0] = (i as f32) / 10000.0; - let embedding = Embedding { - id: format!("vec-{}", i), - vector: vec_data, - metadata: HashMap::new(), - }; - store.upsert(&embedding).await.unwrap(); - } - }); - - let query = vec![0.5f32; dim]; - - b.to_async(&rt).iter(|| async { - black_box(store.search(&query, 10).await.unwrap()) - }); - }); - } - - group.throughput(Throughput::Elements(10000)); - group.finish(); -} - -// ============================================================================ -// Graph Store Benchmarks -// ============================================================================ - -fn bench_graph_operations(c: &mut Criterion) { - let rt = Runtime::new().unwrap(); - let mut group = c.benchmark_group("graph"); - - group.bench_function("insert_edge", |b| { - let store = SimpleGraphStore::in_memory().unwrap(); - let mut counter = 0u64; - - b.to_async(&rt).iter(|| { - counter += 1; - let edge = GraphEdge { - subject: GraphNode::new(format!("https://example.org/node/{}", counter)), - predicate: GraphNode::new("https://example.org/relates_to"), - object: GraphObject::Node(GraphNode::new(format!("https://example.org/target/{}", counter))), - }; - let store_ref = &store; - async move { - black_box(store_ref.insert(&edge).await.unwrap()) - } - }); - }); - - // Pre-populate for query benchmark - let query_store = SimpleGraphStore::in_memory().unwrap(); - let query_node = GraphNode::new("https://example.org/hub"); - rt.block_on(async { - for i in 0..100 { - let edge = GraphEdge { - subject: query_node.clone(), - predicate: GraphNode::new("https://example.org/connects"), - object: GraphObject::Node(GraphNode::new(format!("https://example.org/target/{}", i))), - }; - query_store.insert(&edge).await.unwrap(); - } - }); - - group.bench_function("query_outgoing", |b| { - b.to_async(&rt).iter(|| async { - black_box(query_store.outgoing(&query_node).await.unwrap()) - }); - }); - - group.finish(); -} - -// ============================================================================ -// Octad Store Benchmarks -// ============================================================================ - -fn bench_octad_operations(c: &mut Criterion) { - let rt = Runtime::new().unwrap(); - let mut group = c.benchmark_group("octad"); - - let graph_store = Arc::new(SimpleGraphStore::in_memory().unwrap()); - let vector_store = Arc::new(HnswVectorStore::new(384, DistanceMetric::Cosine, HnswConfig::default())); - let document_store = Arc::new(TantivyDocumentStore::in_memory().unwrap()); - let tensor_store = Arc::new(InMemoryTensorStore::new()); - let semantic_store = Arc::new(InMemorySemanticStore::new()); - let temporal_store: Arc> = Arc::new(InMemoryVersionStore::new()); - let provenance_store = Arc::new(InMemoryProvenanceStore::new()); - let spatial_store = Arc::new(InMemorySpatialStore::new()); - - let config = OctadConfig::default(); - - let store = InMemoryOctadStore::new( - config, - graph_store, - vector_store, - document_store, - tensor_store, - semantic_store, - temporal_store, - provenance_store, - spatial_store, - ); - - group.bench_function("create_octad", |b| { - b.to_async(&rt).iter(|| async { - let input = OctadInput { - document: Some(OctadDocumentInput { - title: "Benchmark Octad".to_string(), - body: "Testing octad creation performance.".to_string(), - fields: HashMap::new(), - }), - vector: Some(OctadVectorInput { - embedding: vec![0.5; 384], - model: None, - }), - ..Default::default() - }; - black_box(store.create(input).await.unwrap()) - }); - }); - - // Create octads for retrieval benchmark - let mut octad_ids = vec![]; - rt.block_on(async { - for i in 0..100 { - let input = OctadInput { - document: Some(OctadDocumentInput { - title: format!("Octad {}", i), - body: format!("Content {}", i), - fields: HashMap::new(), - }), - vector: Some(OctadVectorInput { - embedding: vec![i as f32 / 100.0; 384], - model: None, - }), - ..Default::default() - }; - let octad = store.create(input).await.unwrap(); - octad_ids.push(octad.id.clone()); - } - }); - - group.bench_function("get_octad", |b| { - let id = octad_ids[0].clone(); - b.to_async(&rt).iter(|| async { - black_box(store.get(&id).await.unwrap()) - }); - }); - - group.finish(); -} - -// ============================================================================ -// Drift Detection Benchmarks -// ============================================================================ - -fn bench_drift_detection(c: &mut Criterion) { - let rt = Runtime::new().unwrap(); - let mut group = c.benchmark_group("drift"); - - let detector = DriftDetector::new(DriftThresholds::default()); - - group.bench_function("record_drift", |b| { - b.to_async(&rt).iter(|| async { - black_box( - detector - .record( - DriftType::SemanticVectorDrift, - 0.15, - vec!["entity-bench".to_string()], - ) - .await - .unwrap() - ) - }); - }); - - group.bench_function("health_check", |b| { - b.iter(|| { - black_box(detector.health_check().unwrap()) - }); - }); - - group.finish(); -} - -// ============================================================================ -// Cross-Modal Query Benchmarks -// ============================================================================ - -fn bench_cross_modal_query(c: &mut Criterion) { - let rt = Runtime::new().unwrap(); - let mut group = c.benchmark_group("cross_modal"); - - let graph_store = Arc::new(SimpleGraphStore::in_memory().unwrap()); - let vector_store = Arc::new(HnswVectorStore::new(384, DistanceMetric::Cosine, HnswConfig::default())); - let document_store = Arc::new(TantivyDocumentStore::in_memory().unwrap()); - let tensor_store = Arc::new(InMemoryTensorStore::new()); - let semantic_store = Arc::new(InMemorySemanticStore::new()); - let temporal_store: Arc> = Arc::new(InMemoryVersionStore::new()); - let provenance_store = Arc::new(InMemoryProvenanceStore::new()); - let spatial_store = Arc::new(InMemorySpatialStore::new()); - - let config = OctadConfig::default(); - - let store = InMemoryOctadStore::new( - config, - graph_store, - vector_store, - document_store, - tensor_store, - semantic_store, - temporal_store, - provenance_store, - spatial_store, - ); - - // Create 1000 octads with multiple modalities - rt.block_on(async { - for i in 0..1000 { - let mut embedding = vec![0.0f32; 384]; - embedding[0] = (i as f32) / 1000.0; - - let input = OctadInput { - document: Some(OctadDocumentInput { - title: format!("Multi-modal Document {}", i), - body: format!("Content about machine learning topic {}", i), - fields: HashMap::new(), - }), - vector: Some(OctadVectorInput { - embedding, - model: None, - }), - semantic: Some(OctadSemanticInput { - types: vec!["https://example.org/Document".to_string()], - properties: HashMap::new(), - }), - ..Default::default() - }; - store.create(input).await.unwrap(); - } - }); - - group.throughput(Throughput::Elements(1000)); - - group.bench_function("vector_similarity_search", |b| { - let query = vec![0.5f32; 384]; - b.to_async(&rt).iter(|| async { - black_box(store.search_similar(&query, 10).await.unwrap()) - }); - }); - - group.bench_function("fulltext_search", |b| { - b.to_async(&rt).iter(|| async { - black_box(store.search_text("machine learning", 10).await.unwrap()) - }); - }); - - group.finish(); -} - -// ============================================================================ -// Tensor Store Benchmarks -// ============================================================================ - -fn bench_tensor_operations(c: &mut Criterion) { - let rt = Runtime::new().unwrap(); - let mut group = c.benchmark_group("tensor"); - - group.bench_function("store_create_64x64", |b| { - let store = InMemoryTensorStore::new(); - b.to_async(&rt).iter(|| async { - let data: Vec = (0..4096).map(|i| (i as f64) * 0.001).collect(); - let tensor = Tensor::new("bench-tensor", vec![64, 64], data).unwrap(); - black_box(store.put(&tensor).await.unwrap()) - }); - }); - - // Pre-populate for get benchmark - let get_store = InMemoryTensorStore::new(); - rt.block_on(async { - for i in 0..100 { - let data: Vec = (0..4096).map(|j| ((i * 4096 + j) as f64) * 0.001).collect(); - let tensor = Tensor::new(format!("tensor-{}", i), vec![64, 64], data).unwrap(); - get_store.put(&tensor).await.unwrap(); - } - }); - - group.bench_function("store_get", |b| { - b.to_async(&rt).iter(|| async { - black_box(get_store.get("tensor-50").await.unwrap()) - }); - }); - - // Reduce benchmark: sum along axis 0 of a 64x64 tensor - let reduce_store = InMemoryTensorStore::new(); - rt.block_on(async { - let data: Vec = (0..4096).map(|i| (i as f64) * 0.001).collect(); - let tensor = Tensor::new("reduce-tensor", vec![64, 64], data).unwrap(); - reduce_store.put(&tensor).await.unwrap(); - }); - - group.bench_function("reduce_sum_axis0", |b| { - b.to_async(&rt).iter(|| async { - black_box(reduce_store.reduce("reduce-tensor", 0, ReduceOp::Sum).await.unwrap()) - }); - }); - - group.finish(); -} - -// ============================================================================ -// Semantic Store Benchmarks -// ============================================================================ - -fn bench_semantic_operations(c: &mut Criterion) { - let rt = Runtime::new().unwrap(); - let mut group = c.benchmark_group("semantic"); - - group.bench_function("register_type", |b| { - let store = InMemorySemanticStore::new(); - let mut counter = 0u64; - b.to_async(&rt).iter(|| { - counter += 1; - let iri = format!("https://example.org/Type{}", counter); - let typ = SemanticType::new(&iri, "BenchType"); - let store_ref = &store; - async move { - black_box(store_ref.register_type(&typ).await.unwrap()) - } - }); - }); - - // Pre-populate for get_type benchmark - let type_store = InMemorySemanticStore::new(); - rt.block_on(async { - for i in 0..100 { - let typ = SemanticType::new( - format!("https://example.org/Type{}", i), - format!("Type {}", i), - ); - type_store.register_type(&typ).await.unwrap(); - } - }); - - group.bench_function("get_type", |b| { - b.to_async(&rt).iter(|| async { - black_box(type_store.get_type("https://example.org/Type50").await.unwrap()) - }); - }); - - // Proof creation + CBOR serialization - group.bench_function("proof_create_cbor", |b| { - let store = InMemorySemanticStore::new(); - b.to_async(&rt).iter(|| async { - let proof = ProofBlob::new( - "entity:bench is-a Document", - ProofType::TypeAssignment, - vec![1, 2, 3, 4, 5, 6, 7, 8], - ); - let cbor = black_box(proof.to_cbor().unwrap()); - black_box(store.store_proof(&proof).await.unwrap()); - cbor - }); - }); - - // Proof verification - let verify_store = InMemorySemanticStore::new(); - rt.block_on(async { - for i in 0..50 { - let proof = ProofBlob::new( - "entity:verify-bench is-a Document", - ProofType::Attestation, - vec![i as u8; 32], - ); - verify_store.store_proof(&proof).await.unwrap(); - } - }); - - group.bench_function("proof_verify", |b| { - b.to_async(&rt).iter(|| async { - black_box(verify_store.verify_proofs("entity:verify-bench is-a Document").await.unwrap()) - }); - }); - - group.finish(); -} - -// ============================================================================ -// Temporal Store Benchmarks -// ============================================================================ - -fn bench_temporal_operations(c: &mut Criterion) { - let rt = Runtime::new().unwrap(); - let mut group = c.benchmark_group("temporal"); - - group.bench_function("version_create", |b| { - let store: InMemoryVersionStore = InMemoryVersionStore::new(); - let mut counter = 0u64; - b.to_async(&rt).iter(|| { - counter += 1; - let entity = format!("entity-{}", counter % 10); - let data = format!("version data {}", counter); - let store_ref = &store; - async move { - black_box(store_ref.append(&entity, data, "bench-author", Some("bench commit")).await.unwrap()) - } - }); - }); - - // Pre-populate for retrieval benchmarks - let version_store: InMemoryVersionStore = InMemoryVersionStore::new(); - rt.block_on(async { - for v in 0..100 { - version_store - .append("bench-entity", format!("data v{}", v), "bench-author", Some(&format!("commit {}", v))) - .await - .unwrap(); - } - }); - - group.bench_function("version_get_by_number", |b| { - b.to_async(&rt).iter(|| async { - black_box(version_store.at_version("bench-entity", 50).await.unwrap()) - }); - }); - - group.bench_function("version_get_latest", |b| { - b.to_async(&rt).iter(|| async { - black_box(version_store.latest("bench-entity").await.unwrap()) - }); - }); - - group.bench_function("history_10", |b| { - b.to_async(&rt).iter(|| async { - black_box(version_store.history("bench-entity", 10).await.unwrap()) - }); - }); - - group.bench_function("history_100", |b| { - b.to_async(&rt).iter(|| async { - black_box(version_store.history("bench-entity", 100).await.unwrap()) - }); - }); - - group.finish(); -} - -// ============================================================================ -// Benchmark Groups -// ============================================================================ - -criterion_group!( - document_benches, - bench_document_create, - bench_document_search -); - -criterion_group!( - vector_benches, - bench_vector_insert, - bench_vector_search -); - -criterion_group!( - graph_benches, - bench_graph_operations -); - -criterion_group!( - octad_benches, - bench_octad_operations -); - -criterion_group!( - drift_benches, - bench_drift_detection -); - -criterion_group!( - cross_modal_benches, - bench_cross_modal_query -); - -criterion_group!( - tensor_benches, - bench_tensor_operations -); - -criterion_group!( - semantic_benches, - bench_semantic_operations -); - -criterion_group!( - temporal_benches, - bench_temporal_operations -); - -criterion_main!( - document_benches, - vector_benches, - graph_benches, - octad_benches, - drift_benches, - cross_modal_benches, - tensor_benches, - semantic_benches, - temporal_benches -); diff --git a/verisimdb/benches/src/lib.rs b/verisimdb/benches/src/lib.rs deleted file mode 100644 index 5acd0f0e..00000000 --- a/verisimdb/benches/src/lib.rs +++ /dev/null @@ -1,2 +0,0 @@ -#![forbid(unsafe_code)] -// SPDX-License-Identifier: MPL-2.0 diff --git a/verisimdb/benches/throughput_benchmarks.rs b/verisimdb/benches/throughput_benchmarks.rs deleted file mode 100644 index 63eb1ef1..00000000 --- a/verisimdb/benches/throughput_benchmarks.rs +++ /dev/null @@ -1,369 +0,0 @@ -// SPDX-License-Identifier: MPL-2.0 -// Author: Jonathan D.A. Jewell -//! Write throughput, read latency, and VCL complexity benchmarks for VeriSimDB. -//! -//! This file augments `modality_benchmarks.rs` with system-level throughput -//! and latency measurements that correspond to the missing benchmarks listed -//! in `TEST-NEEDS.md`: -//! -//! - Write throughput — N octad inserts/second on the OctadStore hot path. -//! - Read latency — hot path (cached entity), cold path (uncached entity). -//! - VCL execution time — by query complexity (simple, moderate, complex). -//! -//! The benchmarks use in-memory stores only (no persistent disk I/O) to give -//! reproducible baseline numbers across environments. -//! -//! ## Store construction -//! -//! `InMemoryOctadStore::new` takes 9 arguments (in this order): -//! config, graph, vector, document, tensor, semantic, temporal, provenance, spatial -//! -//! We use `SimpleGraphStore` and `BruteForceVectorStore` — the same combination -//! that `verisim-api` uses in its `ConcreteOctadStore` type alias. - -use criterion::{ - black_box, criterion_group, criterion_main, BenchmarkId, Criterion, Throughput, -}; -use std::collections::HashMap; -use std::sync::Arc; -use tokio::runtime::Runtime; - -use verisim_document::TantivyDocumentStore; -use verisim_graph::SimpleGraphStore; -use verisim_octad::{ - InMemoryOctadStore, OctadConfig, OctadDocumentInput, OctadInput, OctadSnapshot, OctadStore, - OctadVectorInput, -}; -use verisim_provenance::InMemoryProvenanceStore; -use verisim_semantic::InMemorySemanticStore; -use verisim_spatial::InMemorySpatialStore; -use verisim_temporal::InMemoryVersionStore; -use verisim_tensor::InMemoryTensorStore; -use verisim_vector::{BruteForceVectorStore, DistanceMetric}; - -// ============================================================================ -// Concrete type alias -// -// Matches the `ConcreteOctadStore` type in `verisim-api/src/lib.rs` (the -// in-memory / non-persistent configuration). -// ============================================================================ - -type BenchOctadStore = InMemoryOctadStore< - SimpleGraphStore, - BruteForceVectorStore, - TantivyDocumentStore, - InMemoryTensorStore, - InMemorySemanticStore, - InMemoryVersionStore, - InMemoryProvenanceStore, - InMemorySpatialStore, ->; - -// ============================================================================ -// Store factory helpers -// ============================================================================ - -/// Create a fresh in-memory OctadStore for benchmarking. -/// -/// All 9 modality stores are in-memory. This is the standard VeriSimDB -/// configuration deployed by consuming projects (IDApTIK, Burble, Hypatia). -fn make_octad_store() -> BenchOctadStore { - let graph = Arc::new(SimpleGraphStore::new()); - let vector = Arc::new(BruteForceVectorStore::new(384, DistanceMetric::Cosine)); - let document = Arc::new(TantivyDocumentStore::in_memory().unwrap()); - let tensor = Arc::new(InMemoryTensorStore::new()); - let semantic = Arc::new(InMemorySemanticStore::new()); - let temporal = Arc::new(InMemoryVersionStore::new()); - let provenance = Arc::new(InMemoryProvenanceStore::new()); - let spatial = Arc::new(InMemorySpatialStore::new()); - - let config = OctadConfig::default(); - - InMemoryOctadStore::new( - config, graph, vector, document, tensor, semantic, temporal, provenance, spatial, - ) -} - -/// Build an OctadInput with document + vector modalities. -/// -/// Each call produces a structurally unique entity via the counter `i` -/// to prevent deduplication from masking real insertion cost. -fn make_octad_input(i: usize) -> OctadInput { - OctadInput { - document: Some(OctadDocumentInput { - title: format!("Throughput Benchmark Entity {}", i), - body: format!( - "Benchmark entity {} measuring write throughput and latency in VeriSimDB.", - i - ), - fields: HashMap::new(), - }), - vector: Some(OctadVectorInput { - embedding: { - let mut v = vec![0.0f32; 384]; - v[0] = (i as f32) / 100_000.0; - v - }, - model: None, - }), - ..Default::default() - } -} - -// ============================================================================ -// Write Throughput Benchmarks -// -// Measures the number of octad inserts per second on the hot path. -// Three batch sizes: 1, 10, 100 inserts per iteration. -// ============================================================================ - -fn bench_write_throughput(c: &mut Criterion) { - let rt = Runtime::new().unwrap(); - let mut group = c.benchmark_group("write_throughput"); - - for batch_size in [1usize, 10, 100].iter() { - let n = *batch_size; - group.throughput(Throughput::Elements(n as u64)); - - group.bench_with_input( - BenchmarkId::new("octad_insert_batch", n), - &n, - |b, &n| { - b.to_async(&rt).iter_batched( - // Fresh store per iteration to prevent write hot-caching. - make_octad_store, - |store| async move { - for i in 0..n { - black_box(store.create(make_octad_input(i)).await.unwrap()); - } - }, - criterion::BatchSize::SmallInput, - ); - }, - ); - } - - group.finish(); -} - -/// Single-entity write latency — wall-clock time for one `OctadStore::create` -/// with document + vector modalities. -fn bench_single_write_latency(c: &mut Criterion) { - let rt = Runtime::new().unwrap(); - let mut group = c.benchmark_group("write_latency"); - - let store = make_octad_store(); - let mut counter = 0usize; - - group.bench_function("single_octad_create", |b| { - b.to_async(&rt).iter(|| { - counter += 1; - let input = make_octad_input(counter); - async { black_box(store.create(input).await.unwrap()) } - }); - }); - - group.finish(); -} - -// ============================================================================ -// Read Latency Benchmarks -// -// Hot path: entity was just written; stores are warm. -// Cold path: entity written first, followed by 10,000 subsequent writes. -// ============================================================================ - -fn bench_read_latency_hot(c: &mut Criterion) { - let rt = Runtime::new().unwrap(); - let mut group = c.benchmark_group("read_latency"); - - let (hot_id, store) = rt.block_on(async { - let s = make_octad_store(); - let octad = s.create(make_octad_input(0)).await.unwrap(); - (octad.id.clone(), s) - }); - - group.bench_function("hot_path_get_by_id", |b| { - b.to_async(&rt).iter(|| async { - black_box(store.get(&hot_id).await.unwrap()) - }); - }); - - group.finish(); -} - -fn bench_read_latency_cold(c: &mut Criterion) { - let rt = Runtime::new().unwrap(); - let mut group = c.benchmark_group("read_latency"); - - // Write 10,000 entities; retrieve the first one (cold access). - let (cold_id, store) = rt.block_on(async { - let s = make_octad_store(); - let first = s.create(make_octad_input(0)).await.unwrap(); - let id = first.id.clone(); - for i in 1..10_000 { - s.create(make_octad_input(i)).await.unwrap(); - } - (id, s) - }); - - group.throughput(Throughput::Elements(10_000)); - - group.bench_function("cold_path_get_by_id_after_10k_writes", |b| { - b.to_async(&rt).iter(|| async { - black_box(store.get(&cold_id).await.unwrap()) - }); - }); - - group.finish(); -} - -// ============================================================================ -// VCL Query Execution Time by Complexity -// -// Since the VCL executor runs in Elixir, we proxy three complexity tiers -// via direct OctadStore operations that a VCL query would invoke: -// -// Simple: single get-by-ID (1 store lookup) -// Moderate: get-by-ID + vector similarity (2 store operations) -// Complex: full-text search + vector similarity over 1,000 entities -// ============================================================================ - -fn bench_vcl_simple_get(c: &mut Criterion) { - let rt = Runtime::new().unwrap(); - let mut group = c.benchmark_group("vcl_complexity"); - - let (entity_id, store) = rt.block_on(async { - let s = make_octad_store(); - let octad = s.create(make_octad_input(42)).await.unwrap(); - (octad.id.clone(), s) - }); - - group.bench_function("simple_get_by_id", |b| { - b.to_async(&rt).iter(|| async { - black_box(store.get(&entity_id).await.unwrap()) - }); - }); - - group.finish(); -} - -fn bench_vcl_moderate_multimodal(c: &mut Criterion) { - let rt = Runtime::new().unwrap(); - let mut group = c.benchmark_group("vcl_complexity"); - - let (entity_id, store) = rt.block_on(async { - let s = make_octad_store(); - let mut target = None; - for i in 0..100 { - let o = s.create(make_octad_input(i)).await.unwrap(); - if i == 50 { - target = Some(o.id.clone()); - } - } - (target.unwrap(), s) - }); - - // Moderate: get + vector similarity search (2 store operations). - group.bench_function("moderate_get_plus_vector_search", |b| { - let query_vec = vec![0.0005f32; 384]; - b.to_async(&rt).iter(|| async { - let octad = store.get(&entity_id).await.unwrap(); - let similar = store.search_similar(&query_vec, 5).await.unwrap(); - black_box((octad, similar)) - }); - }); - - group.finish(); -} - -fn bench_vcl_complex_cross_modal(c: &mut Criterion) { - let rt = Runtime::new().unwrap(); - let mut group = c.benchmark_group("vcl_complexity"); - - let store = rt.block_on(async { - let s = make_octad_store(); - for i in 0..1_000 { - s.create(make_octad_input(i)).await.unwrap(); - } - s - }); - - group.throughput(Throughput::Elements(1_000)); - - // Complex: full-text + vector search over 1,000 entities. - group.bench_function("complex_fulltext_plus_vector_over_1k", |b| { - let query_vec = vec![0.0005f32; 384]; - b.to_async(&rt).iter(|| async { - let text_results = store.search_text("benchmark", 10).await.unwrap(); - let vec_results = store.search_similar(&query_vec, 10).await.unwrap(); - black_box((text_results, vec_results)) - }); - }); - - group.finish(); -} - -// ============================================================================ -// Write-then-Read Round-Trip Latency -// -// Single-node proxy for replication lag: combined cost of one write + one -// read of the same entity. -// ============================================================================ - -fn bench_write_then_read_latency(c: &mut Criterion) { - let rt = Runtime::new().unwrap(); - let mut group = c.benchmark_group("replication_latency"); - - let store = make_octad_store(); - let mut counter = 0usize; - - group.bench_function("write_then_read_roundtrip", |b| { - b.to_async(&rt).iter(|| { - counter += 1; - let input = make_octad_input(counter); - async { - let octad = store.create(input).await.unwrap(); - let retrieved = store.get(&octad.id).await.unwrap(); - black_box((octad, retrieved)) - } - }); - }); - - group.finish(); -} - -// ============================================================================ -// Benchmark Groups -// ============================================================================ - -criterion_group!( - write_throughput_benches, - bench_write_throughput, - bench_single_write_latency, -); - -criterion_group!( - read_latency_benches, - bench_read_latency_hot, - bench_read_latency_cold, -); - -criterion_group!( - vcl_complexity_benches, - bench_vcl_simple_get, - bench_vcl_moderate_multimodal, - bench_vcl_complex_cross_modal, -); - -criterion_group!( - replication_latency_benches, - bench_write_then_read_latency, -); - -criterion_main!( - write_throughput_benches, - read_latency_benches, - vcl_complexity_benches, - replication_latency_benches, -); diff --git a/verisimdb/connectors/README.adoc b/verisimdb/connectors/README.adoc deleted file mode 100644 index 5b2a7026..00000000 --- a/verisimdb/connectors/README.adoc +++ /dev/null @@ -1,315 +0,0 @@ -// SPDX-License-Identifier: CC-BY-SA-4.0 -// Copyright (c) 2026 Jonathan D.A. Jewell (hyperpolymath) -= VeriSimDB Connectors -- Federation Adapters & Client SDKs -:toc: macro -:toc-title: Contents -:toclevels: 3 -:icons: font -:source-highlighter: rouge - -toc::[] - -== Overview - -VeriSimDB connectors provide a two-direction integration architecture: - -**Outbound (Federation Adapters)**:: VeriSimDB reaches _out_ to external -databases, normalising their data into the octad model and maintaining -cross-modal consistency via drift detection. Federation adapters translate -VeriSimDB federation protocol queries into the native query language of each -target store. - -**Inbound (Client SDKs)**:: External applications reach _in_ to VeriSimDB -through idiomatic client libraries. Each SDK wraps the REST / gRPC API, -handles authentication, connection pooling, and provides type-safe access to -octad entities, VCL queries, and drift reports. - -=== Architecture Diagram - -[source,text] ----- - ┌──────────────────────────────────────┐ - │ VeriSimDB Core │ - │ │ - │ ┌──────────────────────────────┐ │ - │ │ Rust Core Engine │ │ - │ │ Graph | Vector | Tensor │ │ - │ │ Semantic | Document │ │ - │ │ Temporal | Provenance | Spatial│ │ - │ └──────────────┬───────────────┘ │ - │ │ │ - │ ┌──────────────┴───────────────┐ │ - │ │ REST / gRPC API Layer │ │ - │ └──┬────────────────────────┬──┘ │ - └─────┼────────────────────────┼───────┘ - │ │ - ┌─────────────────┘ └─────────────────┐ - │ Inbound (Client SDKs) Outbound (Federation) │ - │ │ - ┌──────────┴──────────┐ ┌──────────────────┴─────────┐ - │ │ │ │ - │ External Apps │ │ External Databases │ - │ │ │ │ - │ ┌───────────────┐ │ │ ┌──────────────────────┐ │ - │ │ Rust SDK │ │ Federation Protocol │ │ PostgreSQL Adapter │ │ - │ │ Elixir SDK │ │ (gRPC / REST) │ │ Neo4j Adapter │ │ - │ │ ReScript SDK │ │ ◄─────────────────────► │ │ Elasticsearch Adapter│ │ - │ │ Gleam SDK │ │ │ │ Qdrant Adapter │ │ - │ │ Julia SDK │ │ │ │ MongoDB Adapter │ │ - │ │ Zig SDK │ │ │ │ Redis Adapter │ │ - │ └───────────────┘ │ │ │ DuckDB Adapter │ │ - │ │ │ │ ClickHouse Adapter │ │ - └─────────────────────┘ │ │ Weaviate Adapter │ │ - │ │ Milvus Adapter │ │ - │ │ SurrealDB Adapter │ │ - │ │ InfluxDB Adapter │ │ - │ │ TigerGraph Adapter │ │ - │ │ Pinecone Adapter │ │ - │ └──────────────────────┘ │ - └────────────────────────────┘ ----- - -== Federation Adapters - -Federation adapters allow VeriSimDB to coordinate with external databases, -treating them as modality-specialised peers in a federated mesh. Each adapter -translates the VeriSimDB federation protocol into the native query language of -the target store, maps results back into octad octad entities, and reports drift -scores so the normaliser can maintain cross-modal consistency. - -All adapters implement the `FederationAdapter` trait (Rust) or the -`FederationService` gRPC service (for language-agnostic integration). - -=== Existing Adapters (4) - -[cols="1,1,2,1"] -|=== -| Adapter | Target Store | Primary Modalities | Status - -| `postgres` -| PostgreSQL 15+ -| Document, Temporal, Spatial (PostGIS) -| Implemented - -| `neo4j` -| Neo4j 5.x -| Graph -| Implemented - -| `elasticsearch` -| Elasticsearch 8.x / OpenSearch 2.x -| Document, Vector (kNN) -| Implemented - -| `qdrant` -| Qdrant 1.x -| Vector -| Implemented -|=== - -=== New Adapters (10) - -[cols="1,1,2,1"] -|=== -| Adapter | Target Store | Primary Modalities | Status - -| `mongodb` -| MongoDB 7.x -| Document, Graph (aggregation pipelines), Spatial (GeoJSON) -| Planned - -| `redis` -| Redis 7.x / Valkey -| Vector (RediSearch), Document, Graph (RedisGraph) -| Planned - -| `duckdb` -| DuckDB 1.x -| Tensor (columnar), Document, Temporal -| Planned - -| `clickhouse` -| ClickHouse 24.x -| Temporal (time-series), Tensor (columnar arrays) -| Planned - -| `weaviate` -| Weaviate 1.x -| Vector, Semantic (schema-driven) -| Planned - -| `milvus` -| Milvus 2.x -| Vector, Tensor -| Planned - -| `surrealdb` -| SurrealDB 2.x -| Graph, Document, Spatial, Vector -| Planned - -| `influxdb` -| InfluxDB 3.x -| Temporal (time-series), Provenance (audit events) -| Planned - -| `tigergraph` -| TigerGraph 4.x -| Graph (distributed), Semantic (GSQL type system) -| Planned - -| `pinecone` -| Pinecone (cloud) -| Vector -| Planned -|=== - -== Client SDKs - -Client SDKs provide idiomatic language bindings for applications that need to -read, write, and query VeriSimDB octad entities. Each SDK wraps the REST API -(and optionally gRPC for streaming) with connection pooling, retry logic, -serialisation, and type-safe octad construction. - -NOTE: **Burble** uses VeriSimDB via the **Elixir client** (`Burble.Store` GenServer -wrapping `VeriSimClient`), not the Gleam SDK. The Gleam SDK is a separate effort -for Gleam-native applications. Codec stubs in the Gleam client do not affect -Burble or any other Elixir-based consumer. - -[cols="1,2,1"] -|=== -| SDK | Language / Runtime | Status - -| `verisim-rs` -| Rust (native, via `reqwest` + `tonic`) -| Planned - -| `verisim-ex` -| Elixir (via `Req` + `GRPC`) -| Planned - -| `verisim-rescript` -| ReScript (via `Fetch` API, Deno runtime) -| Planned - -| `verisim-gleam` -| Gleam (via `gleam_http`, BEAM or JS target) -| Planned - -| `verisim-jl` -| Julia (via `HTTP.jl`) -| Planned - -| `verisim-zig` -| Zig (via `std.http`, C ABI compatible) -| Planned -|=== - -== Octad Modality Mapping - -Every octad entity in VeriSimDB maintains up to 8 synchronized modality -representations. Federation adapters map between these canonical modalities and -the capabilities of each external store. - -[cols="1,3,2"] -|=== -| Modality | Description | Canonical Type - -| `graph` -| RDF triples and property-graph edges. Relationships between entities are - stored as directed labelled edges. -| `GraphNode` / `GraphEdge` - -| `vector` -| Dense embedding vector for similarity search (cosine, dot-product, L2). - Dimensionality is configurable per store. -| `Embedding` (Vec) - -| `tensor` -| Multi-dimensional numeric array. Shape and data are stored together. - Used for ML feature tensors, matrices, and higher-order data. -| `Tensor` (shape: Vec, data: Vec) - -| `semantic` -| Ontological type annotations and property maps. Types are IRIs; properties - are key-value pairs. Proof blobs attach formal verification evidence. -| `SemanticAnnotation` - -| `document` -| Full-text searchable content with title, body, and arbitrary fields. - Backed by Tantivy for indexing. -| `Document` - -| `temporal` -| Version history and time-travel queries. Every mutation creates a new - version with a timestamp, enabling point-in-time reconstruction. -| `Version` / `TimeRange` - -| `provenance` -| Origin tracking, transformation chain, and actor trail. Records form a - SHA-256 hash chain for tamper-evident audit logging. -| `ProvenanceRecord` / `ProvenanceChain` - -| `spatial` -| Geospatial coordinates (WGS84), geometry types (Point, Polygon, etc.), - SRID, and proximity queries (radius, bounding box, k-nearest). -| `SpatialData` / `Coordinates` -|=== - -== Shared Definitions - -The `shared/` directory contains canonical type definitions used by all -connectors. These are the single source of truth for the VeriSimDB wire -format. - -[source,text] ----- -shared/ -├── json-schema/ # JSON Schema (2020-12) definitions -│ ├── octad.json # Full octad entity -│ ├── octad-input.json # Create/update input -│ ├── octad-status.json # Entity status across modalities -│ ├── modality.json # Modality enum -│ ├── query-params.json # Federation query parameters -│ ├── federation-result.json # Normalised federation result -│ ├── drift-score.json # Drift measurement per entity -│ ├── provenance-event.json # Single provenance event -│ └── error.json # Error response envelope -├── openapi/ -│ └── verisim-api-v1.yaml # OpenAPI 3.1 specification -└── proto/ - └── verisim_federation.proto # Protobuf/gRPC service definitions ----- - -All JSON schemas use the `https://verisim.db/schema/` namespace. OpenAPI and -Protobuf definitions reference the same canonical types to ensure consistency -across REST, gRPC, and SDK code generation. - -== Development - -=== Adding a New Federation Adapter - -1. Create a new directory under `connectors/federation//`. -2. Implement the `FederationAdapter` trait (Rust) or generate a gRPC client - from `verisim_federation.proto`. -3. Map the target store's native types to the appropriate octad modalities. -4. Implement drift detection by comparing local and remote representations. -5. Add integration tests using a containerised instance of the target store. -6. Register the adapter in the federation registry. - -=== Adding a New Client SDK - -1. Create a new directory under `connectors/sdks//`. -2. Generate types from the JSON schemas in `shared/json-schema/`. -3. Implement HTTP client calls against the OpenAPI spec. -4. Optionally implement gRPC streaming from `verisim_federation.proto`. -5. Add connection pooling, retry logic, and error handling. -6. Publish to the appropriate package registry. - -== Related Documentation - -* link:../README.adoc[VeriSimDB README] -- project overview and quick start -* link:../WHITEPAPER.md[Whitepaper] -- formal description of the octad model -* link:../docs/[docs/] -- design documents and architecture decisions -* link:../rust-core/verisim-api/src/federation.rs[federation.rs] -- Rust federation implementation -* link:../rust-core/verisim-octad/src/lib.rs[octad lib.rs] -- canonical Octad type definitions diff --git a/verisimdb/connectors/clients/elixir/.formatter.exs b/verisimdb/connectors/clients/elixir/.formatter.exs deleted file mode 100644 index be6d97e7..00000000 --- a/verisimdb/connectors/clients/elixir/.formatter.exs +++ /dev/null @@ -1,6 +0,0 @@ -# SPDX-License-Identifier: MPL-2.0 -# Copyright (c) 2026 Jonathan D.A. Jewell (hyperpolymath) - -[ - inputs: ["{mix,.formatter}.exs", "{config,lib,test}/**/*.{ex,exs}"] -] diff --git a/verisimdb/connectors/clients/elixir/.gitignore b/verisimdb/connectors/clients/elixir/.gitignore deleted file mode 100644 index e65bb150..00000000 --- a/verisimdb/connectors/clients/elixir/.gitignore +++ /dev/null @@ -1,22 +0,0 @@ -# SPDX-License-Identifier: MPL-2.0 - -# The directory Mix will write compiled artifacts to. -/_build/ - -# If you run "mix test --cover", coverage assets end up here. -/cover/ - -# The directory Mix downloads your dependencies sources to. -/deps/ - -# Where third-party dependencies like ExDoc output generated docs. -/doc/ - -# If the VM crashes, it generates a dump; ignore these. -erl_crash.dump - -# Also ignore archive artifacts (built via "mix archive.build"). -*.ez - -# Ignore package tarball (built via "mix hex.build"). -verisim_client-*.tar diff --git a/verisimdb/connectors/clients/elixir/lib/verisim_client.ex b/verisimdb/connectors/clients/elixir/lib/verisim_client.ex deleted file mode 100644 index 237fdbfb..00000000 --- a/verisimdb/connectors/clients/elixir/lib/verisim_client.ex +++ /dev/null @@ -1,220 +0,0 @@ -# SPDX-License-Identifier: MPL-2.0 -# Copyright (c) 2026 Jonathan D.A. Jewell (hyperpolymath) - -defmodule VeriSimClient do - @moduledoc """ - Main client module for connecting to a VeriSimDB instance. - - Holds connection configuration (base URL, authentication, timeout) and - provides low-level HTTP helpers that the domain modules (`Octad`, `Search`, - `Drift`, `Provenance`, `Vcl`, `Federation`) delegate to. - - ## Quick Start - - {:ok, client} = VeriSimClient.new("http://localhost:8080") - {:ok, true} = VeriSimClient.health(client) - - ## Authentication - - Four authentication modes are supported: - - * `:none` — No authentication (local development, trusted networks). - * `{:api_key, key}` — API key via the `X-API-Key` header. - * `{:bearer, token}` — Bearer token via the `Authorization` header. - * `{:basic, username, password}` — HTTP Basic authentication. - - ## Examples - - # Unauthenticated - {:ok, client} = VeriSimClient.new("http://localhost:8080") - - # API key - {:ok, client} = VeriSimClient.new("http://localhost:8080", auth: {:api_key, "my-key"}) - - # Bearer token - {:ok, client} = VeriSimClient.new("http://localhost:8080", auth: {:bearer, "my-token"}) - """ - - @type auth :: - :none - | {:api_key, String.t()} - | {:bearer, String.t()} - | {:basic, String.t(), String.t()} - - @type t :: %__MODULE__{ - base_url: String.t(), - auth: auth(), - timeout: pos_integer() - } - - defstruct [:base_url, auth: :none, timeout: 30_000] - - # --------------------------------------------------------------------------- - # Constructors - # --------------------------------------------------------------------------- - - @doc """ - Create a new VeriSimDB client. - - ## Options - - * `:auth` — Authentication mode (default: `:none`). See module docs. - * `:timeout` — Per-request timeout in milliseconds (default: 30_000). - - ## Examples - - {:ok, client} = VeriSimClient.new("http://localhost:8080") - {:ok, client} = VeriSimClient.new("http://localhost:8080", auth: {:api_key, "key"}, timeout: 10_000) - """ - @spec new(String.t(), keyword()) :: {:ok, t()} | {:error, String.t()} - def new(base_url, opts \\ []) do - auth = Keyword.get(opts, :auth, :none) - timeout = Keyword.get(opts, :timeout, 30_000) - - # Validate the base URL is parseable. - case URI.parse(base_url) do - %URI{scheme: scheme} when scheme in ["http", "https"] -> - # Strip trailing slash for consistent path joining. - base = String.trim_trailing(base_url, "/") - - {:ok, - %__MODULE__{ - base_url: base, - auth: auth, - timeout: timeout - }} - - _ -> - {:error, "Invalid base URL: #{base_url}. Must use http:// or https:// scheme."} - end - end - - # --------------------------------------------------------------------------- - # Health check - # --------------------------------------------------------------------------- - - @doc """ - Ping the VeriSimDB health endpoint. - - Returns `{:ok, true}` if the server is reachable and healthy, or an error tuple. - """ - @spec health(t()) :: {:ok, boolean()} | {:error, term()} - def health(%__MODULE__{} = client) do - case do_get(client, "/health") do - {:ok, _body} -> {:ok, true} - {:error, reason} -> {:error, reason} - end - end - - # --------------------------------------------------------------------------- - # Internal HTTP helpers (used by domain modules) - # --------------------------------------------------------------------------- - - @doc false - @spec do_get(t(), String.t()) :: {:ok, term()} | {:error, term()} - def do_get(%__MODULE__{} = client, path) do - url = build_url(client, path) - - Req.new(url: url, receive_timeout: client.timeout) - |> apply_auth(client.auth) - |> Req.get() - |> handle_response() - end - - @doc false - @spec do_post(t(), String.t(), term()) :: {:ok, term()} | {:error, term()} - def do_post(%__MODULE__{} = client, path, body) do - url = build_url(client, path) - - Req.new(url: url, receive_timeout: client.timeout, json: body) - |> apply_auth(client.auth) - |> Req.post() - |> handle_response() - end - - @doc false - @spec do_put(t(), String.t(), term()) :: {:ok, term()} | {:error, term()} - def do_put(%__MODULE__{} = client, path, body) do - url = build_url(client, path) - - Req.new(url: url, receive_timeout: client.timeout, json: body) - |> apply_auth(client.auth) - |> Req.put() - |> handle_response() - end - - @doc false - @spec do_delete(t(), String.t()) :: :ok | {:error, term()} - def do_delete(%__MODULE__{} = client, path) do - url = build_url(client, path) - - case Req.new(url: url, receive_timeout: client.timeout) - |> apply_auth(client.auth) - |> Req.delete() do - {:ok, %Req.Response{status: status}} when status in 200..299 -> - :ok - - {:ok, %Req.Response{status: 404, body: body}} -> - {:error, {:not_found, extract_message(body)}} - - {:ok, %Req.Response{status: status, body: body}} when status in [401, 403] -> - {:error, {:unauthorized, extract_message(body)}} - - {:ok, %Req.Response{status: status, body: body}} -> - {:error, {:server_error, status, extract_message(body)}} - - {:error, reason} -> - {:error, {:network_error, reason}} - end - end - - # --------------------------------------------------------------------------- - # Private helpers - # --------------------------------------------------------------------------- - - defp build_url(%__MODULE__{base_url: base}, path) do - base <> path - end - - defp apply_auth(req, :none), do: req - - defp apply_auth(req, {:api_key, key}) do - Req.Request.put_header(req, "x-api-key", key) - end - - defp apply_auth(req, {:bearer, token}) do - Req.Request.put_header(req, "authorization", "Bearer #{token}") - end - - defp apply_auth(req, {:basic, username, password}) do - encoded = Base.encode64("#{username}:#{password}") - Req.Request.put_header(req, "authorization", "Basic #{encoded}") - end - - defp handle_response({:ok, %Req.Response{status: status, body: body}}) - when status in 200..299 do - {:ok, body} - end - - defp handle_response({:ok, %Req.Response{status: 404, body: body}}) do - {:error, {:not_found, extract_message(body)}} - end - - defp handle_response({:ok, %Req.Response{status: status, body: body}}) - when status in [401, 403] do - {:error, {:unauthorized, extract_message(body)}} - end - - defp handle_response({:ok, %Req.Response{status: status, body: body}}) do - {:error, {:server_error, status, extract_message(body)}} - end - - defp handle_response({:error, reason}) do - {:error, {:network_error, reason}} - end - - defp extract_message(%{"message" => msg}) when is_binary(msg), do: msg - defp extract_message(%{"error" => err}) when is_binary(err), do: err - defp extract_message(body) when is_binary(body), do: body - defp extract_message(body), do: inspect(body) -end diff --git a/verisimdb/connectors/clients/elixir/lib/verisim_client/drift.ex b/verisimdb/connectors/clients/elixir/lib/verisim_client/drift.ex deleted file mode 100644 index ed1b57d6..00000000 --- a/verisimdb/connectors/clients/elixir/lib/verisim_client/drift.ex +++ /dev/null @@ -1,90 +0,0 @@ -# SPDX-License-Identifier: MPL-2.0 -# Copyright (c) 2026 Jonathan D.A. Jewell (hyperpolymath) - -defmodule VeriSimClient.Drift do - @moduledoc """ - Drift detection and normalization operations for VeriSimDB. - - VeriSimDB continuously monitors how far each octad's modality data has - diverged from its normalised baseline. When drift exceeds a configurable - threshold the entity is flagged for re-normalisation. This module exposes - drift score retrieval, system-wide status, and manual normalization triggers. - - ## Examples - - {:ok, client} = VeriSimClient.new("http://localhost:8080") - - {:ok, score} = VeriSimClient.Drift.score(client, "entity-uuid") - IO.puts("Overall drift: \#{score["overall_score"]}") - - {:ok, status} = VeriSimClient.Drift.status(client) - """ - - alias VeriSimClient.Types - - @doc """ - Retrieve the drift score for a single octad entity. - - The score aggregates per-modality drift metrics into an overall value - (0.0 = perfectly normalised, higher = more drift). - - ## Parameters - - * `client` — A `VeriSimClient.t()` connection. - * `id` — The octad entity identifier. - """ - @spec score(VeriSimClient.t(), String.t()) :: - {:ok, Types.drift_score()} | {:error, term()} - def score(%VeriSimClient{} = client, id) when is_binary(id) do - VeriSimClient.do_get(client, "/api/v1/drift/#{id}") - end - - @doc """ - Retrieve system-wide drift status. - - Returns a map summarising total entities monitored, number exceeding drift - threshold, average drift, and last sweep timestamp. - - ## Parameters - - * `client` — A `VeriSimClient.t()` connection. - """ - @spec status(VeriSimClient.t()) :: {:ok, map()} | {:error, term()} - def status(%VeriSimClient{} = client) do - VeriSimClient.do_get(client, "/api/v1/drift/status") - end - - @doc """ - Trigger re-normalisation for a specific octad entity. - - This enqueues the entity for the normaliser pipeline, which will recompute - cross-modality consistency and update the baseline. - - ## Parameters - - * `client` — A `VeriSimClient.t()` connection. - * `id` — The octad entity identifier. - """ - @spec normalize(VeriSimClient.t(), String.t()) :: :ok | {:error, term()} - def normalize(%VeriSimClient{} = client, id) when is_binary(id) do - case VeriSimClient.do_post(client, "/api/v1/drift/#{id}/normalize", %{}) do - {:ok, _body} -> :ok - {:error, reason} -> {:error, reason} - end - end - - @doc """ - Retrieve the normaliser pipeline status. - - Returns a map with queue depth, active workers, throughput metrics, and last - error (if any). - - ## Parameters - - * `client` — A `VeriSimClient.t()` connection. - """ - @spec normalizer_status(VeriSimClient.t()) :: {:ok, map()} | {:error, term()} - def normalizer_status(%VeriSimClient{} = client) do - VeriSimClient.do_get(client, "/api/v1/drift/normalizer/status") - end -end diff --git a/verisimdb/connectors/clients/elixir/lib/verisim_client/error.ex b/verisimdb/connectors/clients/elixir/lib/verisim_client/error.ex deleted file mode 100644 index 8d67ed34..00000000 --- a/verisimdb/connectors/clients/elixir/lib/verisim_client/error.ex +++ /dev/null @@ -1,103 +0,0 @@ -# SPDX-License-Identifier: MPL-2.0 -# Copyright (c) 2026 Jonathan D.A. Jewell (hyperpolymath) - -defmodule VeriSimClient.Error do - @moduledoc """ - Error types for the VeriSimDB Elixir client SDK. - - Errors are represented as exception structs so they can be raised with - `raise/1` or matched in `{:error, reason}` tuples. Each error type carries - enough context for callers to decide whether to retry, surface a user-facing - message, or escalate. - """ - - # --------------------------------------------------------------------------- - # NotFound - # --------------------------------------------------------------------------- - - defmodule NotFound do - @moduledoc "The requested entity (octad, peer, provenance record) was not found." - defexception [:message] - - @impl true - def exception(id) do - %__MODULE__{message: "Entity not found: #{id}"} - end - end - - # --------------------------------------------------------------------------- - # Unauthorized - # --------------------------------------------------------------------------- - - defmodule Unauthorized do - @moduledoc "Authentication or authorization failed." - defexception [:message] - - @impl true - def exception(reason) do - %__MODULE__{message: "Unauthorized: #{reason}"} - end - end - - # --------------------------------------------------------------------------- - # NetworkError - # --------------------------------------------------------------------------- - - defmodule NetworkError do - @moduledoc "An underlying HTTP / network transport error." - defexception [:message, :reason] - - @impl true - def exception(reason) do - %__MODULE__{message: "Network error: #{inspect(reason)}", reason: reason} - end - end - - # --------------------------------------------------------------------------- - # ServerError - # --------------------------------------------------------------------------- - - defmodule ServerError do - @moduledoc "The server returned an HTTP error status." - defexception [:message, :status] - - @impl true - def exception({status, message}) do - %__MODULE__{ - message: "Server error (#{status}): #{message}", - status: status - } - end - end - - # --------------------------------------------------------------------------- - # ValidationError - # --------------------------------------------------------------------------- - - defmodule ValidationError do - @moduledoc "Client-side validation failed before the request was sent." - defexception [:message] - - @impl true - def exception(reason) do - %__MODULE__{message: "Validation error: #{reason}"} - end - end - - # --------------------------------------------------------------------------- - # Timeout - # --------------------------------------------------------------------------- - - defmodule Timeout do - @moduledoc "The request exceeded the configured timeout duration." - defexception [:message, :timeout_ms] - - @impl true - def exception(timeout_ms) do - %__MODULE__{ - message: "Timeout after #{timeout_ms}ms", - timeout_ms: timeout_ms - } - end - end -end diff --git a/verisimdb/connectors/clients/elixir/lib/verisim_client/federation.ex b/verisimdb/connectors/clients/elixir/lib/verisim_client/federation.ex deleted file mode 100644 index f02905a0..00000000 --- a/verisimdb/connectors/clients/elixir/lib/verisim_client/federation.ex +++ /dev/null @@ -1,108 +0,0 @@ -# SPDX-License-Identifier: MPL-2.0 -# Copyright (c) 2026 Jonathan D.A. Jewell (hyperpolymath) - -defmodule VeriSimClient.Federation do - @moduledoc """ - Federation operations for cross-instance VeriSimDB queries. - - VeriSimDB supports a federated architecture where multiple instances can be - registered as peers. Federated queries fan out to all (or selected) peers - and merge results transparently. This module provides peer management and - federated query execution. - - ## Examples - - {:ok, client} = VeriSimClient.new("http://localhost:8080") - - {:ok, peer} = VeriSimClient.Federation.register_peer(client, - "us-west-replica", - "https://peer.example.com:8080", - "verisimdb", - %{api_key: "peer-key-123"} - ) - - {:ok, peers} = VeriSimClient.Federation.list_peers(client) - - {:ok, results} = VeriSimClient.Federation.query(client, - [:vector, :graph], - %{vector: [0.1, 0.2, 0.3], k: 5} - ) - """ - - alias VeriSimClient.Types - - @doc """ - Register a remote VeriSimDB instance (or compatible adapter) as a - federation peer. - - ## Parameters - - * `client` — A `VeriSimClient.t()` connection. - * `store_id` — Unique logical name for the peer (e.g. "us-west-replica"). - * `endpoint` — Base URL of the peer's API. - * `adapter_type` — Adapter kind: "verisimdb", "quandledb", "lithoglyph", or custom. - * `config` — Adapter-specific configuration (auth tokens, timeouts, etc.). - - ## Returns - - `{:ok, peer_record}` on success, where `peer_record` is a map. - """ - @spec register_peer(VeriSimClient.t(), String.t(), String.t(), String.t(), map()) :: - {:ok, map()} | {:error, term()} - def register_peer(%VeriSimClient{} = client, store_id, endpoint, adapter_type, config) - when is_binary(store_id) and is_binary(endpoint) and - is_binary(adapter_type) and is_map(config) do - body = %{ - store_id: store_id, - endpoint: endpoint, - adapter_type: adapter_type, - config: config - } - - VeriSimClient.do_post(client, "/api/v1/federation/peers", body) - end - - @doc """ - List all registered federation peers. - - Returns a list of peer records, each containing the store ID, endpoint, - adapter type, health status, and last-seen timestamp. - - ## Parameters - - * `client` — A `VeriSimClient.t()` connection. - """ - @spec list_peers(VeriSimClient.t()) :: {:ok, [map()]} | {:error, term()} - def list_peers(%VeriSimClient{} = client) do - VeriSimClient.do_get(client, "/api/v1/federation/peers") - end - - @doc """ - Execute a federated query that fans out to all registered peers. - - The query targets the specified modalities and passes `params` to each - peer's local query engine. Results are merged and returned with per-peer - attribution. - - ## Parameters - - * `client` — A `VeriSimClient.t()` connection. - * `modalities` — List of modality atoms to query across peers. - * `params` — Query parameters map (modality-specific filters, limits, etc.). - - ## Returns - - `{:ok, results}` where `results` is a list of `federation_result()` maps. - """ - @spec query(VeriSimClient.t(), [Types.modality()], map()) :: - {:ok, [Types.federation_result()]} | {:error, term()} - def query(%VeriSimClient{} = client, modalities, params) - when is_list(modalities) and is_map(params) do - body = %{ - modalities: modalities, - params: params - } - - VeriSimClient.do_post(client, "/api/v1/federation/query", body) - end -end diff --git a/verisimdb/connectors/clients/elixir/lib/verisim_client/octad.ex b/verisimdb/connectors/clients/elixir/lib/verisim_client/octad.ex deleted file mode 100644 index 04739414..00000000 --- a/verisimdb/connectors/clients/elixir/lib/verisim_client/octad.ex +++ /dev/null @@ -1,116 +0,0 @@ -# SPDX-License-Identifier: MPL-2.0 -# Copyright (c) 2026 Jonathan D.A. Jewell (hyperpolymath) - -defmodule VeriSimClient.Octad do - @moduledoc """ - Octad CRUD operations for VeriSimDB. - - Octads are the fundamental multi-modal entities in VeriSimDB. Each octad can - carry data across all eight modalities (graph, vector, tensor, semantic, - document, temporal, provenance, spatial). This module provides create, read, - update, delete, and list operations. - - All functions take a `VeriSimClient.t()` as their first argument and return - `{:ok, result}` or `{:error, reason}` tuples. - - ## Examples - - {:ok, client} = VeriSimClient.new("http://localhost:8080") - - {:ok, octad} = VeriSimClient.Octad.create(client, %{ - name: "My Entity", - description: "A test octad", - vector: %{embedding: [0.1, 0.2, 0.3], model: "ada-002"} - }) - - {:ok, fetched} = VeriSimClient.Octad.get(client, octad["id"]) - """ - - alias VeriSimClient.Types - - @doc """ - Create a new octad entity. - - The server assigns a UUID and timestamps; the returned map contains the - fully-populated record. - - ## Parameters - - * `client` — A `VeriSimClient.t()` connection. - * `input` — A `octad_input()` map. At minimum, `:name` should be set. - """ - @spec create(VeriSimClient.t(), Types.octad_input()) :: - {:ok, Types.octad()} | {:error, term()} - def create(%VeriSimClient{} = client, input) when is_map(input) do - VeriSimClient.do_post(client, "/api/v1/octads", input) - end - - @doc """ - Retrieve a single octad by its unique identifier. - - ## Parameters - - * `client` — A `VeriSimClient.t()` connection. - * `id` — The octad UUID string. - """ - @spec get(VeriSimClient.t(), String.t()) :: - {:ok, Types.octad()} | {:error, term()} - def get(%VeriSimClient{} = client, id) when is_binary(id) do - VeriSimClient.do_get(client, "/api/v1/octads/#{id}") - end - - @doc """ - Update an existing octad entity (partial update / merge semantics). - - Only the fields present in `input` are modified; omitted fields retain their - current values. - - ## Parameters - - * `client` — A `VeriSimClient.t()` connection. - * `id` — The octad UUID string. - * `input` — A `octad_input()` map with the fields to update. - """ - @spec update(VeriSimClient.t(), String.t(), Types.octad_input()) :: - {:ok, Types.octad()} | {:error, term()} - def update(%VeriSimClient{} = client, id, input) - when is_binary(id) and is_map(input) do - VeriSimClient.do_put(client, "/api/v1/octads/#{id}", input) - end - - @doc """ - Delete a octad entity by its unique identifier. - - This is a hard delete — the entity and all associated modality data are - removed. Provenance records are retained for auditability. - - ## Parameters - - * `client` — A `VeriSimClient.t()` connection. - * `id` — The octad UUID string. - """ - @spec delete(VeriSimClient.t(), String.t()) :: :ok | {:error, term()} - def delete(%VeriSimClient{} = client, id) when is_binary(id) do - VeriSimClient.do_delete(client, "/api/v1/octads/#{id}") - end - - @doc """ - List octad entities with pagination. - - ## Parameters - - * `client` — A `VeriSimClient.t()` connection. - * `opts` — Keyword list with optional `:limit` (default 20) and `:offset` (default 0). - - ## Examples - - {:ok, page} = VeriSimClient.Octad.list(client, limit: 50, offset: 100) - """ - @spec list(VeriSimClient.t(), keyword()) :: - {:ok, Types.paginated_response()} | {:error, term()} - def list(%VeriSimClient{} = client, opts \\ []) do - limit = Keyword.get(opts, :limit, 20) - offset = Keyword.get(opts, :offset, 0) - VeriSimClient.do_get(client, "/api/v1/octads?limit=#{limit}&offset=#{offset}") - end -end diff --git a/verisimdb/connectors/clients/elixir/lib/verisim_client/provenance.ex b/verisimdb/connectors/clients/elixir/lib/verisim_client/provenance.ex deleted file mode 100644 index 79081c8a..00000000 --- a/verisimdb/connectors/clients/elixir/lib/verisim_client/provenance.ex +++ /dev/null @@ -1,84 +0,0 @@ -# SPDX-License-Identifier: MPL-2.0 -# Copyright (c) 2026 Jonathan D.A. Jewell (hyperpolymath) - -defmodule VeriSimClient.Provenance do - @moduledoc """ - Provenance chain operations for VeriSimDB. - - Every octad entity in VeriSimDB maintains an append-only provenance chain - recording creation, transformation, derivation, and access events. This - module provides methods to read, append to, and cryptographically verify - provenance chains. - - ## Examples - - {:ok, client} = VeriSimClient.new("http://localhost:8080") - - {:ok, chain} = VeriSimClient.Provenance.chain(client, "entity-uuid") - - {:ok, event} = VeriSimClient.Provenance.record(client, "entity-uuid", %{ - entity_id: "entity-uuid", - event_type: "transformation", - agent: "etl-pipeline-v2", - description: "Re-embedded with ada-003 model" - }) - - {:ok, verification} = VeriSimClient.Provenance.verify(client, "entity-uuid") - """ - - alias VeriSimClient.Types - - @doc """ - Retrieve the full provenance chain for a octad entity. - - Events are returned in chronological order (oldest first). - - ## Parameters - - * `client` — A `VeriSimClient.t()` connection. - * `id` — The octad entity identifier. - """ - @spec chain(VeriSimClient.t(), String.t()) :: - {:ok, [Types.provenance_event()]} | {:error, term()} - def chain(%VeriSimClient{} = client, id) when is_binary(id) do - VeriSimClient.do_get(client, "/api/v1/provenance/#{id}") - end - - @doc """ - Append a new event to a octad's provenance chain. - - The event is immutably recorded; its timestamp and identifier are assigned - by the server. The returned map contains the server-assigned fields. - - ## Parameters - - * `client` — A `VeriSimClient.t()` connection. - * `id` — The octad entity identifier. - * `event` — A `provenance_event()` map describing what happened. - """ - @spec record(VeriSimClient.t(), String.t(), Types.provenance_event()) :: - {:ok, Types.provenance_event()} | {:error, term()} - def record(%VeriSimClient{} = client, id, event) - when is_binary(id) and is_map(event) do - VeriSimClient.do_post(client, "/api/v1/provenance/#{id}", event) - end - - @doc """ - Verify the integrity of a octad's provenance chain. - - The server checks that the chain is contiguous, that no events have been - tampered with, and that cryptographic hashes (if enabled) are consistent. - - Returns a map with `"valid"` (boolean), `"chain_length"` (integer), and - optional `"errors"` list describing any integrity violations. - - ## Parameters - - * `client` — A `VeriSimClient.t()` connection. - * `id` — The octad entity identifier. - """ - @spec verify(VeriSimClient.t(), String.t()) :: {:ok, map()} | {:error, term()} - def verify(%VeriSimClient{} = client, id) when is_binary(id) do - VeriSimClient.do_get(client, "/api/v1/provenance/#{id}/verify") - end -end diff --git a/verisimdb/connectors/clients/elixir/lib/verisim_client/search.ex b/verisimdb/connectors/clients/elixir/lib/verisim_client/search.ex deleted file mode 100644 index fa6c5ad0..00000000 --- a/verisimdb/connectors/clients/elixir/lib/verisim_client/search.ex +++ /dev/null @@ -1,165 +0,0 @@ -# SPDX-License-Identifier: MPL-2.0 -# Copyright (c) 2026 Jonathan D.A. Jewell (hyperpolymath) - -defmodule VeriSimClient.Search do - @moduledoc """ - Search operations across VeriSimDB's multi-modal octad entities. - - Supports full-text search, vector similarity (k-NN), graph-relational - traversal, and geospatial queries (radius, bounding box, nearest-neighbour). - - All functions take a `VeriSimClient.t()` as their first argument and return - `{:ok, results}` or `{:error, reason}` tuples. - - ## Examples - - {:ok, client} = VeriSimClient.new("http://localhost:8080") - - # Full-text search - {:ok, results} = VeriSimClient.Search.text(client, "machine learning", limit: 10) - - # Vector similarity - {:ok, results} = VeriSimClient.Search.vector(client, [0.1, 0.2, 0.3], k: 5) - - # Spatial radius search - {:ok, results} = VeriSimClient.Search.spatial_radius(client, 51.5074, -0.1278, 10.0, limit: 20) - """ - - alias VeriSimClient.Types - - @doc """ - Full-text search across octad names, descriptions, and document content. - - ## Parameters - - * `client` — A `VeriSimClient.t()` connection. - * `query` — The search query string. - * `opts` — Keyword list with optional `:limit` (default 20). - """ - @spec text(VeriSimClient.t(), String.t(), keyword()) :: - {:ok, [Types.octad()]} | {:error, term()} - def text(%VeriSimClient{} = client, query, opts \\ []) when is_binary(query) do - limit = Keyword.get(opts, :limit, 20) - encoded_query = URI.encode_www_form(query) - VeriSimClient.do_get(client, "/api/v1/search/text?q=#{encoded_query}&limit=#{limit}") - end - - @doc """ - Vector similarity search (k-nearest neighbours). - - Finds the `k` octads whose stored vector embeddings are closest to the - provided vector (cosine similarity by default). - - ## Parameters - - * `client` — A `VeriSimClient.t()` connection. - * `vector` — The query embedding (list of floats). - * `opts` — Keyword list with optional `:k` (default 10). - """ - @spec vector(VeriSimClient.t(), [float()], keyword()) :: - {:ok, [Types.octad()]} | {:error, term()} - def vector(%VeriSimClient{} = client, vector, opts \\ []) when is_list(vector) do - k = Keyword.get(opts, :k, 10) - body = %{vector: vector, k: k} - VeriSimClient.do_post(client, "/api/v1/search/vector", body) - end - - @doc """ - Find octads related to the given entity via graph edges. - - Traverses one hop of the graph modality and returns all directly connected - octads. - - ## Parameters - - * `client` — A `VeriSimClient.t()` connection. - * `id` — The octad identifier to find relations for. - """ - @spec related(VeriSimClient.t(), String.t()) :: - {:ok, [Types.octad()]} | {:error, term()} - def related(%VeriSimClient{} = client, id) when is_binary(id) do - VeriSimClient.do_get(client, "/api/v1/search/related/#{id}") - end - - @doc """ - Spatial search: find octads within a given radius of a point. - - ## Parameters - - * `client` — A `VeriSimClient.t()` connection. - * `lat` — Centre latitude (WGS 84 decimal degrees). - * `lon` — Centre longitude (WGS 84 decimal degrees). - * `radius_km` — Search radius in kilometres. - * `opts` — Keyword list with optional `:limit` (default 20). - """ - @spec spatial_radius(VeriSimClient.t(), float(), float(), float(), keyword()) :: - {:ok, [Types.octad()]} | {:error, term()} - def spatial_radius(%VeriSimClient{} = client, lat, lon, radius_km, opts \\ []) - when is_number(lat) and is_number(lon) and is_number(radius_km) do - limit = Keyword.get(opts, :limit, 20) - - body = %{ - latitude: lat, - longitude: lon, - radius_km: radius_km, - limit: limit - } - - VeriSimClient.do_post(client, "/api/v1/search/spatial/radius", body) - end - - @doc """ - Spatial search: find octads within a rectangular bounding box. - - ## Parameters - - * `client` — A `VeriSimClient.t()` connection. - * `min_lat` — Southern boundary latitude. - * `min_lon` — Western boundary longitude. - * `max_lat` — Northern boundary latitude. - * `max_lon` — Eastern boundary longitude. - * `opts` — Keyword list with optional `:limit` (default 20). - """ - @spec spatial_bounds(VeriSimClient.t(), float(), float(), float(), float(), keyword()) :: - {:ok, [Types.octad()]} | {:error, term()} - def spatial_bounds(%VeriSimClient{} = client, min_lat, min_lon, max_lat, max_lon, opts \\ []) - when is_number(min_lat) and is_number(min_lon) and - is_number(max_lat) and is_number(max_lon) do - limit = Keyword.get(opts, :limit, 20) - - body = %{ - min_lat: min_lat, - min_lon: min_lon, - max_lat: max_lat, - max_lon: max_lon, - limit: limit - } - - VeriSimClient.do_post(client, "/api/v1/search/spatial/bounds", body) - end - - @doc """ - Spatial search: find the `k` nearest octads to a given point. - - ## Parameters - - * `client` — A `VeriSimClient.t()` connection. - * `lat` — Query point latitude (WGS 84 decimal degrees). - * `lon` — Query point longitude (WGS 84 decimal degrees). - * `opts` — Keyword list with optional `:k` (default 10). - """ - @spec nearest(VeriSimClient.t(), float(), float(), keyword()) :: - {:ok, [Types.octad()]} | {:error, term()} - def nearest(%VeriSimClient{} = client, lat, lon, opts \\ []) - when is_number(lat) and is_number(lon) do - k = Keyword.get(opts, :k, 10) - - body = %{ - latitude: lat, - longitude: lon, - k: k - } - - VeriSimClient.do_post(client, "/api/v1/search/spatial/nearest", body) - end -end diff --git a/verisimdb/connectors/clients/elixir/lib/verisim_client/types.ex b/verisimdb/connectors/clients/elixir/lib/verisim_client/types.ex deleted file mode 100644 index d5b56348..00000000 --- a/verisimdb/connectors/clients/elixir/lib/verisim_client/types.ex +++ /dev/null @@ -1,291 +0,0 @@ -# SPDX-License-Identifier: MPL-2.0 -# Copyright (c) 2026 Jonathan D.A. Jewell (hyperpolymath) - -defmodule VeriSimClient.Types do - @moduledoc """ - Type definitions for the VeriSimDB Elixir client SDK. - - These types mirror the VeriSimDB JSON Schema and cover the full octad of - modalities: Graph, Vector, Tensor, Semantic, Document, Temporal, Provenance, - and Spatial. - - All types are plain maps (not structs) to maintain wire-format fidelity with - the REST API. Type specs are provided for documentation and Dialyzer analysis. - """ - - # --------------------------------------------------------------------------- - # Modality - # --------------------------------------------------------------------------- - - @typedoc """ - The eight modalities supported by VeriSimDB's octad data model. - """ - @type modality :: - :graph - | :vector - | :tensor - | :semantic - | :document - | :temporal - | :provenance - | :spatial - - @doc "List of all supported modality atoms." - @spec all_modalities() :: [modality()] - def all_modalities do - [:graph, :vector, :tensor, :semantic, :document, :temporal, :provenance, :spatial] - end - - # --------------------------------------------------------------------------- - # ModalityStatus - # --------------------------------------------------------------------------- - - @typedoc """ - Boolean flags indicating which modalities are currently active for a octad. - """ - @type modality_status :: %{ - graph: boolean(), - vector: boolean(), - tensor: boolean(), - semantic: boolean(), - document: boolean(), - temporal: boolean(), - provenance: boolean(), - spatial: boolean() - } - - # --------------------------------------------------------------------------- - # OctadStatus (lightweight summary) - # --------------------------------------------------------------------------- - - @typedoc """ - Lightweight status summary for a octad entity. - """ - @type octad_status :: %{ - id: String.t(), - created_at: String.t(), - modified_at: String.t(), - version: non_neg_integer(), - modality_status: modality_status() - } - - # --------------------------------------------------------------------------- - # Octad (full entity) - # --------------------------------------------------------------------------- - - @typedoc """ - A complete octad entity encompassing all eight modality payloads. - """ - @type octad :: %{ - id: String.t(), - name: String.t(), - description: String.t() | nil, - created_at: String.t(), - modified_at: String.t(), - version: non_neg_integer(), - modality_status: modality_status(), - metadata: map() | nil, - graph: map() | nil, - vector: map() | nil, - tensor: map() | nil, - semantic: map() | nil, - document: map() | nil, - temporal: map() | nil, - provenance: map() | nil, - spatial: map() | nil - } - - # --------------------------------------------------------------------------- - # OctadInput (create / update payload) - # --------------------------------------------------------------------------- - - @typedoc """ - Input payload for creating or updating a octad entity. - All fields are optional to support partial updates. - """ - @type octad_input :: %{ - optional(:name) => String.t(), - optional(:description) => String.t(), - optional(:metadata) => map(), - optional(:graph) => octad_graph_input(), - optional(:vector) => octad_vector_input(), - optional(:tensor) => octad_tensor_input(), - optional(:semantic) => octad_semantic_input(), - optional(:document) => octad_document_input(), - optional(:temporal) => octad_temporal_input(), - optional(:provenance) => octad_provenance_input(), - optional(:spatial) => octad_spatial_input() - } - - # --------------------------------------------------------------------------- - # Per-modality input types - # --------------------------------------------------------------------------- - - @typedoc "Graph modality input: nodes, edges, and properties." - @type octad_graph_input :: %{ - optional(:nodes) => [String.t()], - optional(:edges) => [graph_edge()], - optional(:properties) => map() - } - - @typedoc "A single directed edge in the graph modality." - @type graph_edge :: %{ - required(:source) => String.t(), - required(:target) => String.t(), - required(:label) => String.t(), - optional(:properties) => map() - } - - @typedoc "Vector modality input: dense embeddings for similarity search." - @type octad_vector_input :: %{ - required(:embedding) => [float()], - optional(:dimensions) => non_neg_integer(), - optional(:model) => String.t() - } - - @typedoc "Tensor modality input: multi-dimensional numeric data." - @type octad_tensor_input :: %{ - required(:data) => [float()], - required(:shape) => [non_neg_integer()], - optional(:dtype) => String.t() - } - - @typedoc "Semantic modality input: RDF-style triples and ontology references." - @type octad_semantic_input :: %{ - optional(:triples) => [semantic_triple()], - optional(:ontology) => String.t(), - optional(:annotations) => map() - } - - @typedoc "A single semantic triple (subject, predicate, object)." - @type semantic_triple :: %{ - subject: String.t(), - predicate: String.t(), - object: String.t() - } - - @typedoc "Document modality input: unstructured / semi-structured content." - @type octad_document_input :: %{ - required(:content) => String.t(), - optional(:content_type) => String.t(), - optional(:language) => String.t(), - optional(:metadata) => map() - } - - @typedoc "Temporal modality input: time-series events and temporal metadata." - @type octad_temporal_input :: %{ - required(:timestamp) => String.t(), - optional(:duration_ms) => non_neg_integer(), - optional(:recurrence) => String.t(), - optional(:timezone) => String.t(), - optional(:metadata) => map() - } - - @typedoc "Provenance modality input: lineage and audit trail events." - @type octad_provenance_input :: %{ - required(:event_type) => String.t(), - required(:agent) => String.t(), - optional(:description) => String.t(), - optional(:source_ids) => [String.t()], - optional(:metadata) => map() - } - - @typedoc "Spatial modality input: geospatial coordinates and geometries." - @type octad_spatial_input :: %{ - optional(:latitude) => float(), - optional(:longitude) => float(), - optional(:altitude) => float(), - optional(:geometry) => map(), - optional(:crs) => String.t() - } - - # --------------------------------------------------------------------------- - # DriftScore - # --------------------------------------------------------------------------- - - @typedoc """ - Drift score for a octad entity, measuring divergence from normalised baseline. - """ - @type drift_score :: %{ - entity_id: String.t(), - overall_score: float(), - modality_scores: %{String.t() => float()}, - last_checked: String.t(), - needs_normalization: boolean() - } - - # --------------------------------------------------------------------------- - # ProvenanceEvent - # --------------------------------------------------------------------------- - - @typedoc """ - A single immutable event in a octad's provenance chain. - """ - @type provenance_event :: %{ - optional(:id) => String.t(), - required(:entity_id) => String.t(), - required(:event_type) => String.t(), - required(:agent) => String.t(), - optional(:description) => String.t(), - optional(:timestamp) => String.t(), - optional(:source_ids) => [String.t()], - optional(:metadata) => map() - } - - # --------------------------------------------------------------------------- - # FederationResult - # --------------------------------------------------------------------------- - - @typedoc """ - A single result from a federated cross-instance query. - """ - @type federation_result :: %{ - required(:store_id) => String.t(), - required(:entity) => octad(), - optional(:score) => float(), - optional(:latency_ms) => non_neg_integer() - } - - # --------------------------------------------------------------------------- - # VclResponse - # --------------------------------------------------------------------------- - - @typedoc """ - Response from a VCL query execution or explain request. - """ - @type vcl_response :: %{ - required(:success) => boolean(), - required(:statement_type) => String.t(), - required(:row_count) => non_neg_integer(), - required(:data) => term(), - optional(:message) => String.t() - } - - # --------------------------------------------------------------------------- - # ErrorResponse - # --------------------------------------------------------------------------- - - @typedoc """ - Standard error response body from the VeriSimDB REST API. - """ - @type error_response :: %{ - required(:error) => String.t(), - required(:message) => String.t(), - optional(:details) => term() - } - - # --------------------------------------------------------------------------- - # PaginatedResponse - # --------------------------------------------------------------------------- - - @typedoc """ - Generic wrapper for paginated list responses. - """ - @type paginated_response :: %{ - data: [term()], - total: non_neg_integer(), - limit: non_neg_integer(), - offset: non_neg_integer(), - has_more: boolean() - } -end diff --git a/verisimdb/connectors/clients/elixir/lib/verisim_client/vcl.ex b/verisimdb/connectors/clients/elixir/lib/verisim_client/vcl.ex deleted file mode 100644 index f5457c8d..00000000 --- a/verisimdb/connectors/clients/elixir/lib/verisim_client/vcl.ex +++ /dev/null @@ -1,60 +0,0 @@ -# SPDX-License-Identifier: MPL-2.0 -# Copyright (c) 2026 Jonathan D.A. Jewell (hyperpolymath) - -defmodule VeriSimClient.Vcl do - @moduledoc """ - VeriSim Consonance Language (VCL) execution for VeriSimDB. - - VCL is VeriSimDB's native query language, supporting SQL-like syntax extended - with multi-modal operations (vector similarity, graph traversal, spatial - predicates, drift thresholds, etc.). This module provides methods to execute - VCL statements and retrieve explain / query plans. - - ## Examples - - {:ok, client} = VeriSimClient.new("http://localhost:8080") - - {:ok, result} = VeriSimClient.Vcl.execute(client, "SELECT * FROM octads WHERE drift > 0.5") - IO.puts("Rows returned: \#{result["row_count"]}") - - {:ok, plan} = VeriSimClient.Vcl.explain(client, "SELECT * FROM octads WHERE drift > 0.5") - """ - - alias VeriSimClient.Types - - @doc """ - Execute a VCL statement against the VeriSimDB instance. - - Supports SELECT, INSERT, UPDATE, DELETE, and VeriSimDB-specific statements - like `DRIFT CHECK`, `NORMALIZE`, and `FEDERATE`. - - ## Parameters - - * `client` — A `VeriSimClient.t()` connection. - * `query` — The VCL statement string. - """ - @spec execute(VeriSimClient.t(), String.t()) :: - {:ok, Types.vcl_response()} | {:error, term()} - def execute(%VeriSimClient{} = client, query) when is_binary(query) do - body = %{query: query} - VeriSimClient.do_post(client, "/api/v1/vcl/execute", body) - end - - @doc """ - Request an explain / query plan for a VCL statement without executing it. - - Useful for understanding which modalities, indices, and federation peers - would be involved in a query. - - ## Parameters - - * `client` — A `VeriSimClient.t()` connection. - * `query` — The VCL statement string to explain. - """ - @spec explain(VeriSimClient.t(), String.t()) :: - {:ok, Types.vcl_response()} | {:error, term()} - def explain(%VeriSimClient{} = client, query) when is_binary(query) do - body = %{query: query} - VeriSimClient.do_post(client, "/api/v1/vcl/explain", body) - end -end diff --git a/verisimdb/connectors/clients/elixir/mix.exs b/verisimdb/connectors/clients/elixir/mix.exs deleted file mode 100644 index a96e36d9..00000000 --- a/verisimdb/connectors/clients/elixir/mix.exs +++ /dev/null @@ -1,52 +0,0 @@ -# SPDX-License-Identifier: MPL-2.0 -# Copyright (c) 2026 Jonathan D.A. Jewell (hyperpolymath) - -defmodule VeriSimClient.MixProject do - @moduledoc """ - Mix project configuration for the VeriSimDB Elixir client SDK. - - Provides octad entity management, multi-modal search (text, vector, spatial), - drift detection, provenance chain operations, VCL query execution, and - federation across distributed VeriSimDB instances. - """ - - use Mix.Project - - def project do - [ - app: :verisim_client, - version: "0.1.0", - elixir: "~> 1.17", - start_permanent: Mix.env() == :prod, - deps: deps(), - name: "VeriSimDB Client", - source_url: "https://gitlab.com/hyperpolymath/verisimdb", - description: "VeriSimDB client SDK for Elixir — octad entity management, drift detection, and federation", - package: package() - ] - end - - def application do - [ - extra_applications: [:logger] - ] - end - - defp deps do - [ - {:req, "~> 0.5"}, - {:jason, "~> 1.4"}, - {:ex_doc, "~> 0.34", only: :dev, runtime: false} - ] - end - - defp package do - [ - licenses: ["PMPL-1.0-or-later"], - links: %{ - "GitLab" => "https://gitlab.com/hyperpolymath/verisimdb", - "GitHub" => "https://github.com/hyperpolymath/verisimdb" - } - ] - end -end diff --git a/verisimdb/connectors/clients/elixir/test/test_helper.exs b/verisimdb/connectors/clients/elixir/test/test_helper.exs deleted file mode 100644 index 549fa5c5..00000000 --- a/verisimdb/connectors/clients/elixir/test/test_helper.exs +++ /dev/null @@ -1,4 +0,0 @@ -# SPDX-License-Identifier: MPL-2.0 -# Copyright (c) 2026 Jonathan D.A. Jewell (hyperpolymath) - -ExUnit.start() diff --git a/verisimdb/connectors/clients/elixir/test/verisim_client_test.exs b/verisimdb/connectors/clients/elixir/test/verisim_client_test.exs deleted file mode 100644 index f8ef254a..00000000 --- a/verisimdb/connectors/clients/elixir/test/verisim_client_test.exs +++ /dev/null @@ -1,76 +0,0 @@ -# SPDX-License-Identifier: MPL-2.0 -# Copyright (c) 2026 Jonathan D.A. Jewell (hyperpolymath) - -defmodule VeriSimClientTest do - @moduledoc """ - Basic tests for the VeriSimDB Elixir client SDK. - - These tests exercise client construction and validate that the client struct - is correctly populated. Network-dependent tests (health check, CRUD, etc.) - would require a running VeriSimDB instance or a mock server. - """ - - use ExUnit.Case, async: true - - describe "VeriSimClient.new/2" do - test "creates a client with default options" do - assert {:ok, client} = VeriSimClient.new("http://localhost:8080") - assert client.base_url == "http://localhost:8080" - assert client.auth == :none - assert client.timeout == 30_000 - end - - test "creates a client with API key auth" do - assert {:ok, client} = - VeriSimClient.new("https://verisim.example.com", auth: {:api_key, "test-key"}) - - assert client.base_url == "https://verisim.example.com" - assert client.auth == {:api_key, "test-key"} - end - - test "creates a client with bearer auth" do - assert {:ok, client} = - VeriSimClient.new("http://localhost:8080", auth: {:bearer, "my-token"}) - - assert client.auth == {:bearer, "my-token"} - end - - test "creates a client with basic auth" do - assert {:ok, client} = - VeriSimClient.new("http://localhost:8080", - auth: {:basic, "user", "pass"} - ) - - assert client.auth == {:basic, "user", "pass"} - end - - test "creates a client with custom timeout" do - assert {:ok, client} = VeriSimClient.new("http://localhost:8080", timeout: 5_000) - assert client.timeout == 5_000 - end - - test "strips trailing slash from base URL" do - assert {:ok, client} = VeriSimClient.new("http://localhost:8080/") - assert client.base_url == "http://localhost:8080" - end - - test "rejects invalid URL schemes" do - assert {:error, _reason} = VeriSimClient.new("ftp://localhost:8080") - end - end - - describe "VeriSimClient.Types" do - test "all_modalities returns eight modalities" do - modalities = VeriSimClient.Types.all_modalities() - assert length(modalities) == 8 - assert :graph in modalities - assert :vector in modalities - assert :tensor in modalities - assert :semantic in modalities - assert :document in modalities - assert :temporal in modalities - assert :provenance in modalities - assert :spatial in modalities - end - end -end diff --git a/verisimdb/connectors/clients/gleam/gleam.toml b/verisimdb/connectors/clients/gleam/gleam.toml deleted file mode 100644 index 3bdfe8bf..00000000 --- a/verisimdb/connectors/clients/gleam/gleam.toml +++ /dev/null @@ -1,18 +0,0 @@ -# SPDX-License-Identifier: MPL-2.0 -# (PMPL-1.0-or-later preferred; MPL-2.0 required for Gleam/Hex ecosystem) -# Copyright (c) 2026 Jonathan D.A. Jewell (hyperpolymath) - -name = "verisimdb_client" -version = "0.1.0" -description = "VeriSimDB client SDK for Gleam — octad entity management, drift detection, and federation" -licences = ["MPL-2.0"] -repository = { type = "gitlab", url = "https://gitlab.com/hyperpolymath/verisimdb" } - -[dependencies] -gleam_stdlib = ">= 0.69.0 and < 2.0.0" -gleam_http = ">= 3.7.0 and < 4.0.0" -gleam_json = ">= 2.3.0 and < 3.0.0" -gleam_hackney = ">= 1.2.0 and < 2.0.0" - -[dev-dependencies] -gleeunit = ">= 1.0.0 and < 2.0.0" diff --git a/verisimdb/connectors/clients/gleam/src/verisimdb_client.gleam b/verisimdb/connectors/clients/gleam/src/verisimdb_client.gleam deleted file mode 100644 index 512b18f7..00000000 --- a/verisimdb/connectors/clients/gleam/src/verisimdb_client.gleam +++ /dev/null @@ -1,228 +0,0 @@ -//// SPDX-License-Identifier: MPL-2.0 -//// (PMPL-1.0-or-later preferred; MPL-2.0 required for Gleam/Hex ecosystem) -//// Copyright (c) 2026 Jonathan D.A. Jewell (hyperpolymath) -//// -//// VeriSimDB Gleam Client — Main module. -//// -//// This module provides the core Client type and constructor functions for -//// connecting to a VeriSimDB server. It supports multiple authentication -//// methods (API key, Basic, Bearer token, or none) and manages base URL -//// routing, request timeouts, and standard HTTP verb helpers. -//// -//// Usage: -//// let client = verisimdb_client.new("http://localhost:8080") -//// let assert Ok(True) = verisimdb_client.health(client) - -import gleam/http -import gleam/http/request -import gleam/http/response -import gleam/hackney -import gleam/json -import gleam/option.{type Option, None, Some} -import gleam/result -import gleam/string -import gleam/bit_array -import verisimdb_client/error.{type VeriSimError} - -// --------------------------------------------------------------------------- -// Authentication types -// --------------------------------------------------------------------------- - -/// Authentication method for connecting to VeriSimDB. -pub type Auth { - /// API key authentication via X-API-Key header. - ApiKey(key: String) - /// HTTP Basic Authentication (username:password). - Basic(username: String, password: String) - /// Bearer token authentication via Authorization header. - Bearer(token: String) - /// No authentication required. - NoAuth -} - -// --------------------------------------------------------------------------- -// Client type -// --------------------------------------------------------------------------- - -/// Client holds the connection configuration for a VeriSimDB server. -/// -/// Fields: -/// base_url — Root URL of the VeriSimDB API (e.g. "http://localhost:8080"). -/// timeout — Request timeout in milliseconds. Defaults to 30000 (30 seconds). -/// auth — Authentication method. Defaults to NoAuth. -pub type Client { - Client(base_url: String, timeout: Int, auth: Auth) -} - -/// Create a new unauthenticated client with the given base URL. -pub fn new(base_url: String) -> Client { - Client( - base_url: string.trim_end(base_url, "/"), - timeout: 30_000, - auth: NoAuth, - ) -} - -/// Create a client with a specific authentication method. -pub fn new_with_auth(base_url: String, auth: Auth) -> Client { - Client( - base_url: string.trim_end(base_url, "/"), - timeout: 30_000, - auth: auth, - ) -} - -/// Create a client authenticated with an API key. -pub fn new_with_api_key(base_url: String, key: String) -> Client { - new_with_auth(base_url, ApiKey(key)) -} - -/// Create a client authenticated with a Bearer token. -pub fn new_with_bearer(base_url: String, token: String) -> Client { - new_with_auth(base_url, Bearer(token)) -} - -// --------------------------------------------------------------------------- -// Internal HTTP helpers -// --------------------------------------------------------------------------- - -/// Build an HTTP request with authentication headers applied. -fn build_request( - client: Client, - method: http.Method, - path: String, -) -> Result(request.Request(String), VeriSimError) { - let url = client.base_url <> path - case request.to(url) { - Ok(req) -> { - let req = request.set_method(req, method) - let req = apply_auth(req, client.auth) - Ok(req) - } - Error(_) -> Error(error.ConnectionError("Invalid URL: " <> url)) - } -} - -/// Apply authentication headers to a request. -fn apply_auth( - req: request.Request(String), - auth: Auth, -) -> request.Request(String) { - case auth { - ApiKey(key) -> request.set_header(req, "x-api-key", key) - Basic(username, password) -> { - let credentials = username <> ":" <> password - let encoded = - credentials - |> bit_array.from_string - |> bit_array.base64_encode(True) - request.set_header(req, "authorization", "Basic " <> encoded) - } - Bearer(token) -> - request.set_header(req, "authorization", "Bearer " <> token) - NoAuth -> req - } -} - -/// Send a GET request to the given API path. -pub fn do_get( - client: Client, - path: String, -) -> Result(response.Response(String), VeriSimError) { - case build_request(client, http.Get, path) { - Ok(req) -> - case hackney.send(req) { - Ok(resp) -> Ok(resp) - Error(_) -> - Error(error.ConnectionError( - "Failed to connect to VeriSimDB server", - )) - } - Error(err) -> Error(err) - } -} - -/// Send a POST request with a JSON body. -pub fn do_post( - client: Client, - path: String, - body: String, -) -> Result(response.Response(String), VeriSimError) { - case build_request(client, http.Post, path) { - Ok(req) -> { - let req = - req - |> request.set_header("content-type", "application/json") - |> request.set_body(body) - case hackney.send(req) { - Ok(resp) -> Ok(resp) - Error(_) -> - Error(error.ConnectionError( - "Failed to connect to VeriSimDB server", - )) - } - } - Error(err) -> Error(err) - } -} - -/// Send a PUT request with a JSON body. -pub fn do_put( - client: Client, - path: String, - body: String, -) -> Result(response.Response(String), VeriSimError) { - case build_request(client, http.Put, path) { - Ok(req) -> { - let req = - req - |> request.set_header("content-type", "application/json") - |> request.set_body(body) - case hackney.send(req) { - Ok(resp) -> Ok(resp) - Error(_) -> - Error(error.ConnectionError( - "Failed to connect to VeriSimDB server", - )) - } - } - Error(err) -> Error(err) - } -} - -/// Send a DELETE request to the given API path. -pub fn do_delete( - client: Client, - path: String, -) -> Result(response.Response(String), VeriSimError) { - case build_request(client, http.Delete, path) { - Ok(req) -> - case hackney.send(req) { - Ok(resp) -> Ok(resp) - Error(_) -> - Error(error.ConnectionError( - "Failed to connect to VeriSimDB server", - )) - } - Error(err) -> Error(err) - } -} - -// --------------------------------------------------------------------------- -// Health check -// --------------------------------------------------------------------------- - -/// Check whether the VeriSimDB server is reachable and healthy. -/// -/// Sends a GET request to /health and expects a 200 OK response. -/// Returns Ok(True) if healthy, or an error on failure. -pub fn health(client: Client) -> Result(Bool, VeriSimError) { - case do_get(client, "/health") { - Ok(resp) -> - case resp.status { - 200 -> Ok(True) - status -> Error(error.from_status(status)) - } - Error(err) -> Error(err) - } -} diff --git a/verisimdb/connectors/clients/gleam/src/verisimdb_client/codec.gleam b/verisimdb/connectors/clients/gleam/src/verisimdb_client/codec.gleam deleted file mode 100644 index 8f02f321..00000000 --- a/verisimdb/connectors/clients/gleam/src/verisimdb_client/codec.gleam +++ /dev/null @@ -1,858 +0,0 @@ -//// SPDX-License-Identifier: MPL-2.0 -//// (PMPL-1.0-or-later preferred; MPL-2.0 required for Gleam/Hex ecosystem) -//// Copyright (c) 2026 Jonathan D.A. Jewell (hyperpolymath) -//// -//// VeriSimDB Gleam Client — JSON codec module. -//// -//// Centralised JSON encoding and decoding for all VeriSimDB types. -//// Uses gleam/json for encoding and gleam/dynamic/decode for type-safe -//// deserialization of API responses. Every modality's data is fully -//// serialized and deserialized — no stubs. - -import gleam/dict.{type Dict} -import gleam/dynamic/decode -import gleam/json -import gleam/list -import gleam/option.{type Option, None, Some} -import gleam/result -import verisimdb_client/error.{type VeriSimError} -import verisimdb_client/types.{ - type DriftLevel, type DriftScore, type DriftStatusReport, - type DocumentContent, type FederatedQueryResult, type FederationPeer, - type GraphData, type GraphEdge, type Modality, type ModalityStatus, type Octad, - type OctadInput, type OctadStatus, type PaginatedResponse, - type PeerQueryResult, type ProvenanceChain, type ProvenanceEvent, - type SearchResult, type SpatialData, type TensorData, type VectorData, - type VclExplanation, type VclResult, -} - -// ========================================================================== -// Shared helpers -// ========================================================================== - -/// Encode a Dict(String, String) as a JSON object of string values. -pub fn encode_string_dict(d: Dict(String, String)) -> json.Json { - json.object( - d - |> dict.to_list - |> list.map(fn(pair) { #(pair.0, json.string(pair.1)) }), - ) -} - -/// Encode a Dict(String, Float) as a JSON object of float values. -fn encode_float_dict(d: Dict(String, Float)) -> json.Json { - json.object( - d - |> dict.to_list - |> list.map(fn(pair) { #(pair.0, json.float(pair.1)) }), - ) -} - -/// Encode an optional value: produces json.null() when None. -fn encode_optional( - opt: Option(a), - encoder: fn(a) -> json.Json, -) -> json.Json { - case opt { - Some(val) -> encoder(val) - None -> json.null() - } -} - -/// Decode a JSON string body using a decoder, wrapping errors as SerializationError. -fn parse_json( - body: String, - decoder: decode.Decoder(a), -) -> Result(a, VeriSimError) { - case json.parse(body, decoder) { - Ok(value) -> Ok(value) - Error(err) -> - Error(error.SerializationError( - "JSON decode error: " <> decode_error_to_string(err), - )) - } -} - -/// Convert a decode error to a human-readable string. -fn decode_error_to_string(err: json.DecodeError) -> String { - case err { - json.UnexpectedFormat(_decode_errors) -> "unexpected JSON format" - json.UnexpectedEndOfInput -> "unexpected end of JSON input" - json.UnexpectedByte(byte) -> "unexpected byte: " <> byte - json.UnexpectedSequence(seq) -> "unexpected sequence: " <> seq - } -} - -/// Decoder for a Dict(String, String) from a JSON object. -fn string_dict_decoder() -> decode.Decoder(Dict(String, String)) { - decode.dict(decode.string, decode.string) -} - -/// Decoder for a Dict(String, Float) from a JSON object. -fn float_dict_decoder() -> decode.Decoder(Dict(String, Float)) { - decode.dict(decode.string, decode.float) -} - -/// Decoder for an optional field that may be null or absent. -fn optional_field( - name: String, - inner: decode.Decoder(a), -) -> decode.Decoder(Option(a)) { - decode.optional_field(name, None, decode.optional(inner)) -} - -// ========================================================================== -// Modality encoding/decoding -// ========================================================================== - -/// Encode a Modality to its JSON string. -fn encode_modality(modality: Modality) -> json.Json { - json.string(types.modality_to_string(modality)) -} - -/// Decoder for a Modality from a JSON string. -fn modality_decoder() -> decode.Decoder(Modality) { - decode.string - |> decode.then(fn(s) { - case types.modality_from_string(s) { - Some(m) -> decode.success(m) - None -> decode.failure(types.Graph, "Modality") - } - }) -} - -// ========================================================================== -// OctadStatus encoding/decoding -// ========================================================================== - -/// Encode an OctadStatus to its JSON string representation. -fn encode_octad_status(status: OctadStatus) -> json.Json { - json.string(case status { - types.Active -> "active" - types.Archived -> "archived" - types.Draft -> "draft" - types.Deleted -> "deleted" - }) -} - -/// Decoder for an OctadStatus from a JSON string. -fn octad_status_decoder() -> decode.Decoder(OctadStatus) { - decode.string - |> decode.then(fn(s) { - case s { - "active" -> decode.success(types.Active) - "archived" -> decode.success(types.Archived) - "draft" -> decode.success(types.Draft) - "deleted" -> decode.success(types.Deleted) - _ -> decode.failure(types.Active, "OctadStatus") - } - }) -} - -// ========================================================================== -// ModalityStatus encoding/decoding -// ========================================================================== - -/// Encode a ModalityStatus as a JSON object with boolean fields. -fn encode_modality_status(ms: ModalityStatus) -> json.Json { - json.object([ - #("graph", json.bool(ms.graph)), - #("vector", json.bool(ms.vector)), - #("tensor", json.bool(ms.tensor)), - #("semantic", json.bool(ms.semantic)), - #("document", json.bool(ms.document)), - #("temporal", json.bool(ms.temporal)), - #("provenance", json.bool(ms.provenance)), - #("spatial", json.bool(ms.spatial)), - ]) -} - -/// Decoder for a ModalityStatus from a JSON object. -fn modality_status_decoder() -> decode.Decoder(ModalityStatus) { - decode.into({ - use graph <- decode.parameter - use vector <- decode.parameter - use tensor <- decode.parameter - use semantic <- decode.parameter - use document <- decode.parameter - use temporal <- decode.parameter - use provenance <- decode.parameter - use spatial <- decode.parameter - types.ModalityStatus( - graph: graph, - vector: vector, - tensor: tensor, - semantic: semantic, - document: document, - temporal: temporal, - provenance: provenance, - spatial: spatial, - ) - }) - |> decode.field("graph", decode.bool) - |> decode.field("vector", decode.bool) - |> decode.field("tensor", decode.bool) - |> decode.field("semantic", decode.bool) - |> decode.field("document", decode.bool) - |> decode.field("temporal", decode.bool) - |> decode.field("provenance", decode.bool) - |> decode.field("spatial", decode.bool) -} - -// ========================================================================== -// GraphEdge encoding/decoding -// ========================================================================== - -/// Encode a GraphEdge as a JSON object. -fn encode_graph_edge(edge: GraphEdge) -> json.Json { - json.object([ - #("source", json.string(edge.source)), - #("target", json.string(edge.target)), - #("rel_type", json.string(edge.rel_type)), - #("weight", json.float(edge.weight)), - #("metadata", encode_string_dict(edge.metadata)), - ]) -} - -/// Decoder for a GraphEdge from a JSON object. -fn graph_edge_decoder() -> decode.Decoder(GraphEdge) { - decode.into({ - use source <- decode.parameter - use target <- decode.parameter - use rel_type <- decode.parameter - use weight <- decode.parameter - use metadata <- decode.parameter - types.GraphEdge( - source: source, - target: target, - rel_type: rel_type, - weight: weight, - metadata: metadata, - ) - }) - |> decode.field("source", decode.string) - |> decode.field("target", decode.string) - |> decode.field("rel_type", decode.string) - |> decode.field("weight", decode.float) - |> decode.field("metadata", string_dict_decoder()) -} - -// ========================================================================== -// GraphData encoding/decoding -// ========================================================================== - -/// Encode GraphData as a JSON object. -fn encode_graph_data(data: GraphData) -> json.Json { - json.object([ - #("edges", json.array(data.edges, encode_graph_edge)), - #("properties", encode_string_dict(data.properties)), - ]) -} - -/// Decoder for GraphData from a JSON object. -fn graph_data_decoder() -> decode.Decoder(GraphData) { - decode.into({ - use edges <- decode.parameter - use properties <- decode.parameter - types.GraphData(edges: edges, properties: properties) - }) - |> decode.field("edges", decode.list(graph_edge_decoder())) - |> decode.field("properties", string_dict_decoder()) -} - -// ========================================================================== -// VectorData encoding/decoding -// ========================================================================== - -/// Encode VectorData as a JSON object. -fn encode_vector_data(data: VectorData) -> json.Json { - json.object([ - #("embedding", json.array(data.embedding, json.float)), - #("model", json.string(data.model)), - #("dimensions", json.int(data.dimensions)), - ]) -} - -/// Decoder for VectorData from a JSON object. -fn vector_data_decoder() -> decode.Decoder(VectorData) { - decode.into({ - use embedding <- decode.parameter - use model <- decode.parameter - use dimensions <- decode.parameter - types.VectorData(embedding: embedding, model: model, dimensions: dimensions) - }) - |> decode.field("embedding", decode.list(decode.float)) - |> decode.field("model", decode.string) - |> decode.field("dimensions", decode.int) -} - -// ========================================================================== -// TensorData encoding/decoding -// ========================================================================== - -/// Encode TensorData as a JSON object. -fn encode_tensor_data(data: TensorData) -> json.Json { - json.object([ - #("shape", json.array(data.shape, json.int)), - #("dtype", json.string(data.dtype)), - #("data_ref", json.string(data.data_ref)), - ]) -} - -/// Decoder for TensorData from a JSON object. -fn tensor_data_decoder() -> decode.Decoder(TensorData) { - decode.into({ - use shape <- decode.parameter - use dtype <- decode.parameter - use data_ref <- decode.parameter - types.TensorData(shape: shape, dtype: dtype, data_ref: data_ref) - }) - |> decode.field("shape", decode.list(decode.int)) - |> decode.field("dtype", decode.string) - |> decode.field("data_ref", decode.string) -} - -// ========================================================================== -// DocumentContent encoding/decoding -// ========================================================================== - -/// Encode DocumentContent as a JSON object. -fn encode_document_content(doc: DocumentContent) -> json.Json { - json.object([ - #("text", json.string(doc.text)), - #("format", json.string(doc.format)), - #("language", json.string(doc.language)), - #("metadata", encode_string_dict(doc.metadata)), - ]) -} - -/// Decoder for DocumentContent from a JSON object. -fn document_content_decoder() -> decode.Decoder(DocumentContent) { - decode.into({ - use text <- decode.parameter - use format <- decode.parameter - use language <- decode.parameter - use metadata <- decode.parameter - types.DocumentContent( - text: text, - format: format, - language: language, - metadata: metadata, - ) - }) - |> decode.field("text", decode.string) - |> decode.field("format", decode.string) - |> decode.field("language", decode.string) - |> decode.field("metadata", string_dict_decoder()) -} - -// ========================================================================== -// SpatialData encoding/decoding -// ========================================================================== - -/// Encode SpatialData as a JSON object. -fn encode_spatial_data(data: SpatialData) -> json.Json { - json.object([ - #("latitude", json.float(data.latitude)), - #("longitude", json.float(data.longitude)), - #("altitude", encode_optional(data.altitude, json.float)), - #("geometry", encode_optional(data.geometry, json.string)), - #("crs", json.string(data.crs)), - ]) -} - -/// Decoder for SpatialData from a JSON object. -fn spatial_data_decoder() -> decode.Decoder(SpatialData) { - decode.into({ - use latitude <- decode.parameter - use longitude <- decode.parameter - use altitude <- decode.parameter - use geometry <- decode.parameter - use crs <- decode.parameter - types.SpatialData( - latitude: latitude, - longitude: longitude, - altitude: altitude, - geometry: geometry, - crs: crs, - ) - }) - |> decode.field("latitude", decode.float) - |> decode.field("longitude", decode.float) - |> optional_field("altitude", decode.float) - |> optional_field("geometry", decode.string) - |> decode.field("crs", decode.string) -} - -// ========================================================================== -// Octad encoding/decoding -// ========================================================================== - -/// Encode a full Octad as a JSON string. -pub fn encode_octad(octad: Octad) -> String { - json.to_string(json.object([ - #("id", json.string(octad.id)), - #("status", encode_octad_status(octad.status)), - #("modalities", encode_modality_status(octad.modalities)), - #("created_at", json.string(octad.created_at)), - #("updated_at", json.string(octad.updated_at)), - #("metadata", encode_string_dict(octad.metadata)), - #("graph_data", encode_optional(octad.graph_data, encode_graph_data)), - #("vector_data", encode_optional(octad.vector_data, encode_vector_data)), - #("tensor_data", encode_optional(octad.tensor_data, encode_tensor_data)), - #("content", encode_optional(octad.content, encode_document_content)), - #("spatial_data", encode_optional(octad.spatial_data, encode_spatial_data)), - ])) -} - -/// Decoder for a full Octad from a JSON object. -fn octad_decoder() -> decode.Decoder(Octad) { - decode.into({ - use id <- decode.parameter - use status <- decode.parameter - use modalities <- decode.parameter - use created_at <- decode.parameter - use updated_at <- decode.parameter - use metadata <- decode.parameter - use graph_data <- decode.parameter - use vector_data <- decode.parameter - use tensor_data <- decode.parameter - use content <- decode.parameter - use spatial_data <- decode.parameter - types.Octad( - id: id, - status: status, - modalities: modalities, - created_at: created_at, - updated_at: updated_at, - metadata: metadata, - graph_data: graph_data, - vector_data: vector_data, - tensor_data: tensor_data, - content: content, - spatial_data: spatial_data, - ) - }) - |> decode.field("id", decode.string) - |> decode.field("status", octad_status_decoder()) - |> decode.field("modalities", modality_status_decoder()) - |> decode.field("created_at", decode.string) - |> decode.field("updated_at", decode.string) - |> decode.field("metadata", string_dict_decoder()) - |> optional_field("graph_data", graph_data_decoder()) - |> optional_field("vector_data", vector_data_decoder()) - |> optional_field("tensor_data", tensor_data_decoder()) - |> optional_field("content", document_content_decoder()) - |> optional_field("spatial_data", spatial_data_decoder()) -} - -/// Decode an Octad from a JSON response body string. -pub fn decode_octad(body: String) -> Result(Octad, VeriSimError) { - parse_json(body, octad_decoder()) -} - -// ========================================================================== -// OctadInput encoding -// ========================================================================== - -/// Encode an OctadInput as a JSON string for create/update requests. -/// Serializes all 8 modality data fields when present, plus metadata -/// and the active modality list. -pub fn encode_octad_input(input: OctadInput) -> String { - json.to_string(json.object([ - #( - "modalities", - json.array(input.modalities, encode_modality), - ), - #("metadata", encode_string_dict(input.metadata)), - #("graph_data", encode_optional(input.graph_data, encode_graph_data)), - #("vector_data", encode_optional(input.vector_data, encode_vector_data)), - #("tensor_data", encode_optional(input.tensor_data, encode_tensor_data)), - #("content", encode_optional(input.content, encode_document_content)), - #("spatial_data", encode_optional(input.spatial_data, encode_spatial_data)), - ])) -} - -// ========================================================================== -// PaginatedResponse decoding -// ========================================================================== - -/// Decoder for a PaginatedResponse from a JSON object. -fn paginated_response_decoder() -> decode.Decoder(PaginatedResponse) { - decode.into({ - use items <- decode.parameter - use total <- decode.parameter - use page <- decode.parameter - use per_page <- decode.parameter - use total_pages <- decode.parameter - types.PaginatedResponse( - items: items, - total: total, - page: page, - per_page: per_page, - total_pages: total_pages, - ) - }) - |> decode.field("items", decode.list(octad_decoder())) - |> decode.field("total", decode.int) - |> decode.field("page", decode.int) - |> decode.field("per_page", decode.int) - |> decode.field("total_pages", decode.int) -} - -/// Decode a PaginatedResponse from a JSON response body string. -pub fn decode_paginated_response( - body: String, -) -> Result(PaginatedResponse, VeriSimError) { - parse_json(body, paginated_response_decoder()) -} - -// ========================================================================== -// SearchResult decoding -// ========================================================================== - -/// Decoder for a SearchResult from a JSON object. -fn search_result_decoder() -> decode.Decoder(SearchResult) { - decode.into({ - use octad <- decode.parameter - use score <- decode.parameter - types.SearchResult(octad: octad, score: score) - }) - |> decode.field("octad", octad_decoder()) - |> decode.field("score", decode.float) -} - -/// Decode a list of SearchResult from a JSON response body string. -pub fn decode_search_results( - body: String, -) -> Result(List(SearchResult), VeriSimError) { - parse_json(body, decode.list(search_result_decoder())) -} - -// ========================================================================== -// DriftScore encoding/decoding -// ========================================================================== - -/// Decoder for a DriftScore from a JSON object. -fn drift_score_decoder() -> decode.Decoder(DriftScore) { - decode.into({ - use octad_id <- decode.parameter - use score <- decode.parameter - use components <- decode.parameter - use measured_at <- decode.parameter - use baseline_at <- decode.parameter - types.DriftScore( - octad_id: octad_id, - score: score, - components: components, - measured_at: measured_at, - baseline_at: baseline_at, - ) - }) - |> decode.field("octad_id", decode.string) - |> decode.field("score", decode.float) - |> decode.field("components", float_dict_decoder()) - |> decode.field("measured_at", decode.string) - |> decode.field("baseline_at", decode.string) -} - -/// Decode a DriftScore from a JSON response body string. -pub fn decode_drift_score( - body: String, -) -> Result(DriftScore, VeriSimError) { - parse_json(body, drift_score_decoder()) -} - -// ========================================================================== -// DriftLevel decoding -// ========================================================================== - -/// Decoder for a DriftLevel from a JSON string. -fn drift_level_decoder() -> decode.Decoder(DriftLevel) { - decode.string - |> decode.then(fn(s) { - case s { - "stable" -> decode.success(types.DriftStable) - "low" -> decode.success(types.DriftLow) - "moderate" -> decode.success(types.DriftModerate) - "high" -> decode.success(types.DriftHigh) - "critical" -> decode.success(types.DriftCritical) - _ -> decode.failure(types.DriftStable, "DriftLevel") - } - }) -} - -// ========================================================================== -// DriftStatusReport decoding -// ========================================================================== - -/// Decoder for a DriftStatusReport from a JSON object. -fn drift_status_report_decoder() -> decode.Decoder(DriftStatusReport) { - decode.into({ - use octad_id <- decode.parameter - use level <- decode.parameter - use score <- decode.parameter - use message <- decode.parameter - types.DriftStatusReport( - octad_id: octad_id, - level: level, - score: score, - message: message, - ) - }) - |> decode.field("octad_id", decode.string) - |> decode.field("level", drift_level_decoder()) - |> decode.field("score", drift_score_decoder()) - |> decode.field("message", decode.string) -} - -/// Decode a DriftStatusReport from a JSON response body string. -pub fn decode_drift_status_report( - body: String, -) -> Result(DriftStatusReport, VeriSimError) { - parse_json(body, drift_status_report_decoder()) -} - -// ========================================================================== -// ProvenanceEvent encoding/decoding -// ========================================================================== - -/// Decoder for a ProvenanceEvent from a JSON object. -fn provenance_event_decoder() -> decode.Decoder(types.ProvenanceEvent) { - decode.into({ - use event_id <- decode.parameter - use octad_id <- decode.parameter - use event_type <- decode.parameter - use actor <- decode.parameter - use timestamp <- decode.parameter - use details <- decode.parameter - use parent_id <- decode.parameter - types.ProvenanceEvent( - event_id: event_id, - octad_id: octad_id, - event_type: event_type, - actor: actor, - timestamp: timestamp, - details: details, - parent_id: parent_id, - ) - }) - |> decode.field("event_id", decode.string) - |> decode.field("octad_id", decode.string) - |> decode.field("event_type", decode.string) - |> decode.field("actor", decode.string) - |> decode.field("timestamp", decode.string) - |> decode.field("details", string_dict_decoder()) - |> optional_field("parent_id", decode.string) -} - -/// Decode a ProvenanceEvent from a JSON response body string. -pub fn decode_provenance_event( - body: String, -) -> Result(types.ProvenanceEvent, VeriSimError) { - parse_json(body, provenance_event_decoder()) -} - -// ========================================================================== -// ProvenanceChain decoding -// ========================================================================== - -/// Decoder for a ProvenanceChain from a JSON object. -fn provenance_chain_decoder() -> decode.Decoder(types.ProvenanceChain) { - decode.into({ - use octad_id <- decode.parameter - use events <- decode.parameter - use verified <- decode.parameter - types.ProvenanceChain( - octad_id: octad_id, - events: events, - verified: verified, - ) - }) - |> decode.field("octad_id", decode.string) - |> decode.field("events", decode.list(provenance_event_decoder())) - |> decode.field("verified", decode.bool) -} - -/// Decode a ProvenanceChain from a JSON response body string. -pub fn decode_provenance_chain( - body: String, -) -> Result(types.ProvenanceChain, VeriSimError) { - parse_json(body, provenance_chain_decoder()) -} - -// ========================================================================== -// VclResult decoding -// ========================================================================== - -/// Decoder for a VclResult from a JSON object. -fn vcl_result_decoder() -> decode.Decoder(VclResult) { - decode.into({ - use columns <- decode.parameter - use rows <- decode.parameter - use count <- decode.parameter - use elapsed_ms <- decode.parameter - types.VclResult( - columns: columns, - rows: rows, - count: count, - elapsed_ms: elapsed_ms, - ) - }) - |> decode.field("columns", decode.list(decode.string)) - |> decode.field("rows", decode.list(decode.list(decode.string))) - |> decode.field("count", decode.int) - |> decode.field("elapsed_ms", decode.float) -} - -/// Decode a VclResult from a JSON response body string. -pub fn decode_vcl_result( - body: String, -) -> Result(VclResult, VeriSimError) { - parse_json(body, vcl_result_decoder()) -} - -// ========================================================================== -// VclExplanation decoding -// ========================================================================== - -/// Decoder for a VclExplanation from a JSON object. -fn vcl_explanation_decoder() -> decode.Decoder(VclExplanation) { - decode.into({ - use query <- decode.parameter - use plan <- decode.parameter - use cost <- decode.parameter - use warnings <- decode.parameter - types.VclExplanation( - query: query, - plan: plan, - cost: cost, - warnings: warnings, - ) - }) - |> decode.field("query", decode.string) - |> decode.field("plan", decode.string) - |> decode.field("cost", decode.float) - |> decode.field("warnings", decode.list(decode.string)) -} - -/// Decode a VclExplanation from a JSON response body string. -pub fn decode_vcl_explanation( - body: String, -) -> Result(VclExplanation, VeriSimError) { - parse_json(body, vcl_explanation_decoder()) -} - -// ========================================================================== -// FederationPeer decoding -// ========================================================================== - -/// Decoder for a FederationPeer from a JSON object. -fn federation_peer_decoder() -> decode.Decoder(FederationPeer) { - decode.into({ - use peer_id <- decode.parameter - use name <- decode.parameter - use url <- decode.parameter - use status <- decode.parameter - use last_seen <- decode.parameter - use metadata <- decode.parameter - types.FederationPeer( - peer_id: peer_id, - name: name, - url: url, - status: status, - last_seen: last_seen, - metadata: metadata, - ) - }) - |> decode.field("peer_id", decode.string) - |> decode.field("name", decode.string) - |> decode.field("url", decode.string) - |> decode.field("status", decode.string) - |> decode.field("last_seen", decode.string) - |> decode.field("metadata", string_dict_decoder()) -} - -/// Decode a FederationPeer from a JSON response body string. -pub fn decode_federation_peer( - body: String, -) -> Result(FederationPeer, VeriSimError) { - parse_json(body, federation_peer_decoder()) -} - -/// Decode a list of FederationPeer from a JSON response body string. -pub fn decode_federation_peers( - body: String, -) -> Result(List(FederationPeer), VeriSimError) { - parse_json(body, decode.list(federation_peer_decoder())) -} - -// ========================================================================== -// PeerQueryResult decoding -// ========================================================================== - -/// Decoder for a PeerQueryResult from a JSON object. -fn peer_query_result_decoder() -> decode.Decoder(PeerQueryResult) { - decode.into({ - use peer_id <- decode.parameter - use peer_name <- decode.parameter - use result <- decode.parameter - use elapsed_ms <- decode.parameter - use error <- decode.parameter - types.PeerQueryResult( - peer_id: peer_id, - peer_name: peer_name, - result: result, - elapsed_ms: elapsed_ms, - error: error, - ) - }) - |> decode.field("peer_id", decode.string) - |> decode.field("peer_name", decode.string) - |> decode.field("result", vcl_result_decoder()) - |> decode.field("elapsed_ms", decode.float) - |> optional_field("error", decode.string) -} - -// ========================================================================== -// FederatedQueryResult decoding -// ========================================================================== - -/// Decoder for a FederatedQueryResult from a JSON object. -fn federated_query_result_decoder() -> decode.Decoder(FederatedQueryResult) { - decode.into({ - use results <- decode.parameter - use total <- decode.parameter - use elapsed_ms <- decode.parameter - types.FederatedQueryResult( - results: results, - total: total, - elapsed_ms: elapsed_ms, - ) - }) - |> decode.field("results", decode.list(peer_query_result_decoder())) - |> decode.field("total", decode.int) - |> decode.field("elapsed_ms", decode.float) -} - -/// Decode a FederatedQueryResult from a JSON response body string. -pub fn decode_federated_query_result( - body: String, -) -> Result(FederatedQueryResult, VeriSimError) { - parse_json(body, federated_query_result_decoder()) -} - -// ========================================================================== -// ProvenanceEventInput encoding -// ========================================================================== - -/// Encode a ProvenanceEventInput as a JSON string for POST requests. -pub fn encode_provenance_event_input( - input: types.ProvenanceEventInput, -) -> String { - json.to_string(json.object([ - #("event_type", json.string(input.event_type)), - #("actor", json.string(input.actor)), - #("details", encode_string_dict(input.details)), - ])) -} diff --git a/verisimdb/connectors/clients/gleam/src/verisimdb_client/drift.gleam b/verisimdb/connectors/clients/gleam/src/verisimdb_client/drift.gleam deleted file mode 100644 index c97b43f2..00000000 --- a/verisimdb/connectors/clients/gleam/src/verisimdb_client/drift.gleam +++ /dev/null @@ -1,92 +0,0 @@ -//// SPDX-License-Identifier: MPL-2.0 -//// (PMPL-1.0-or-later preferred; MPL-2.0 required for Gleam/Hex ecosystem) -//// Copyright (c) 2026 Jonathan D.A. Jewell (hyperpolymath) -//// -//// VeriSimDB Gleam Client — Drift detection operations. -//// -//// Drift measures how much a octad's embeddings, relationships, or content -//// have diverged from a baseline state (0.0 = no drift, 1.0 = maximum drift). -//// This module provides functions to query drift scores, check classified -//// status, and trigger re-normalisation. -//// -//// JSON decoding uses the shared codec module for type-safe deserialization. - -import verisimdb_client.{type Client} -import verisimdb_client/codec -import verisimdb_client/error.{type VeriSimError} -import verisimdb_client/types.{type DriftScore, type DriftStatusReport} - -/// Retrieve the current drift score for a specific octad. -/// -/// The drift score is a floating-point value between 0.0 (no drift — fully -/// aligned with baseline) and 1.0 (maximum drift — completely diverged). -/// -/// Parameters: -/// client — The authenticated client. -/// octad_id — The unique identifier of the octad. -/// -/// Returns a DriftScore with component breakdown, or an error. -pub fn get_score( - client: Client, - octad_id: String, -) -> Result(DriftScore, VeriSimError) { - let path = "/api/v1/octads/" <> octad_id <> "/drift" - case verisimdb_client.do_get(client, path) { - Ok(resp) -> - case resp.status { - 200 -> codec.decode_drift_score(resp.body) - status -> Error(error.from_status(status)) - } - Error(err) -> Error(err) - } -} - -/// Retrieve a classified drift status report for a octad. -/// -/// The report includes the drift level (Stable, Low, Moderate, High, Critical), -/// the underlying score, and a human-readable explanation. -/// -/// Parameters: -/// client — The authenticated client. -/// octad_id — The unique identifier of the octad. -/// -/// Returns a DriftStatusReport, or an error. -pub fn status( - client: Client, - octad_id: String, -) -> Result(DriftStatusReport, VeriSimError) { - let path = "/api/v1/octads/" <> octad_id <> "/drift/status" - case verisimdb_client.do_get(client, path) { - Ok(resp) -> - case resp.status { - 200 -> codec.decode_drift_status_report(resp.body) - status -> Error(error.from_status(status)) - } - Error(err) -> Error(err) - } -} - -/// Trigger re-normalisation of a drifted octad. -/// -/// Normalisation recomputes the octad's embeddings and relationship weights -/// against the current baseline, effectively resetting the drift score. -/// -/// Parameters: -/// client — The authenticated client. -/// octad_id — The unique identifier of the octad. -/// -/// Returns the updated DriftScore after normalisation, or an error. -pub fn normalize( - client: Client, - octad_id: String, -) -> Result(DriftScore, VeriSimError) { - let path = "/api/v1/octads/" <> octad_id <> "/drift/normalize" - case verisimdb_client.do_post(client, path, "{}") { - Ok(resp) -> - case resp.status { - 200 -> codec.decode_drift_score(resp.body) - status -> Error(error.from_status(status)) - } - Error(err) -> Error(err) - } -} diff --git a/verisimdb/connectors/clients/gleam/src/verisimdb_client/error.gleam b/verisimdb/connectors/clients/gleam/src/verisimdb_client/error.gleam deleted file mode 100644 index c96d6aff..00000000 --- a/verisimdb/connectors/clients/gleam/src/verisimdb_client/error.gleam +++ /dev/null @@ -1,144 +0,0 @@ -//// SPDX-License-Identifier: MPL-2.0 -//// (PMPL-1.0-or-later preferred; MPL-2.0 required for Gleam/Hex ecosystem) -//// Copyright (c) 2026 Jonathan D.A. Jewell (hyperpolymath) -//// -//// VeriSimDB Gleam Client — Error types and handling. -//// -//// This module defines typed error variants for all failure modes that can -//// occur when communicating with a VeriSimDB server. Gleam's custom types -//// enable exhaustive pattern matching on all error variants. -//// -//// The VeriSimDB server returns errors in a standard JSON envelope: -//// { "error": { "code": "HEXAD_NOT_FOUND", "message": "...", "details": {...} } } - -import gleam/int - -/// Typed error variants for VeriSimDB client operations. -/// -/// Client errors (4xx), server errors (5xx), domain-specific errors, and -/// client-side errors are all represented as a single custom type for -/// exhaustive pattern matching. -pub type VeriSimError { - // --- Client errors (4xx) --- - /// HTTP 400 — Malformed request. - BadRequest(message: String) - /// HTTP 401 — Missing or invalid authentication. - Unauthorized(message: String) - /// HTTP 403 — Insufficient permissions. - Forbidden(message: String) - /// HTTP 404 — Resource does not exist. - NotFound(message: String) - /// HTTP 409 — Resource version conflict. - Conflict(message: String) - /// HTTP 422 — Input validation failure. - ValidationFailed(message: String) - /// HTTP 429 — Too many requests. - RateLimited(message: String) - - // --- Server errors (5xx) --- - /// HTTP 500 — Unexpected server error. - InternalError(message: String) - /// HTTP 503 — Server temporarily unavailable. - ServiceUnavailable(message: String) - - // --- Domain-specific errors --- - /// Octad with given ID does not exist. - OctadNotFound(message: String) - /// Requested modality is not enabled. - ModalityUnavailable(message: String) - /// Drift score computation failed. - DriftComputationError(message: String) - /// Provenance chain integrity failure. - ProvenanceInvalid(message: String) - /// VCL syntax error. - VclParseError(message: String) - /// VCL runtime error. - VclExecutionError(message: String) - /// Federation peer communication failure. - FederationError(message: String) - - // --- Client-side errors --- - /// Network connectivity failure. - ConnectionError(message: String) - /// Request exceeded timeout. - TimeoutError(message: String) - /// JSON serialization/deserialization failure. - SerializationError(message: String) - /// Unrecognised error. - UnknownError(message: String) -} - -/// Extract a human-readable message from any error variant. -pub fn message(err: VeriSimError) -> String { - case err { - BadRequest(msg) -> "Bad request: " <> msg - Unauthorized(msg) -> "Unauthorized: " <> msg - Forbidden(msg) -> "Forbidden: " <> msg - NotFound(msg) -> "Not found: " <> msg - Conflict(msg) -> "Conflict: " <> msg - ValidationFailed(msg) -> "Validation failed: " <> msg - RateLimited(msg) -> "Rate limited: " <> msg - InternalError(msg) -> "Internal server error: " <> msg - ServiceUnavailable(msg) -> "Service unavailable: " <> msg - OctadNotFound(msg) -> "Octad not found: " <> msg - ModalityUnavailable(msg) -> "Modality unavailable: " <> msg - DriftComputationError(msg) -> "Drift computation error: " <> msg - ProvenanceInvalid(msg) -> "Provenance invalid: " <> msg - VclParseError(msg) -> "VCL parse error: " <> msg - VclExecutionError(msg) -> "VCL execution error: " <> msg - FederationError(msg) -> "Federation error: " <> msg - ConnectionError(msg) -> "Connection error: " <> msg - TimeoutError(msg) -> "Timeout error: " <> msg - SerializationError(msg) -> "Serialization error: " <> msg - UnknownError(msg) -> "Unknown error: " <> msg - } -} - -/// Construct an error variant from an HTTP status code. -/// -/// Maps standard HTTP status codes to the appropriate error variant. -/// Used internally by other SDK modules when the server returns a non-success -/// status code. -pub fn from_status(status: Int) -> VeriSimError { - case status { - 400 -> BadRequest("Bad request") - 401 -> Unauthorized("Authentication required") - 403 -> Forbidden("Insufficient permissions") - 404 -> NotFound("Resource not found") - 409 -> Conflict("Resource conflict") - 422 -> ValidationFailed("Input validation failed") - 429 -> RateLimited("Too many requests") - 500 -> InternalError("Internal server error") - 503 -> ServiceUnavailable("Server temporarily unavailable") - _ -> UnknownError("Unexpected HTTP status: " <> int.to_string(status)) - } -} - -/// Check whether an error is retryable. -/// -/// Server errors (5xx), rate limiting (429), connection errors, and timeouts -/// are generally retryable. Client errors (4xx) are not. -pub fn is_retryable(err: VeriSimError) -> Bool { - case err { - RateLimited(_) -> True - InternalError(_) -> True - ServiceUnavailable(_) -> True - ConnectionError(_) -> True - TimeoutError(_) -> True - BadRequest(_) -> False - Unauthorized(_) -> False - Forbidden(_) -> False - NotFound(_) -> False - Conflict(_) -> False - ValidationFailed(_) -> False - OctadNotFound(_) -> False - ModalityUnavailable(_) -> False - DriftComputationError(_) -> False - ProvenanceInvalid(_) -> False - VclParseError(_) -> False - VclExecutionError(_) -> False - FederationError(_) -> False - SerializationError(_) -> False - UnknownError(_) -> False - } -} diff --git a/verisimdb/connectors/clients/gleam/src/verisimdb_client/federation.gleam b/verisimdb/connectors/clients/gleam/src/verisimdb_client/federation.gleam deleted file mode 100644 index bf627bed..00000000 --- a/verisimdb/connectors/clients/gleam/src/verisimdb_client/federation.gleam +++ /dev/null @@ -1,121 +0,0 @@ -//// SPDX-License-Identifier: MPL-2.0 -//// (PMPL-1.0-or-later preferred; MPL-2.0 required for Gleam/Hex ecosystem) -//// Copyright (c) 2026 Jonathan D.A. Jewell (hyperpolymath) -//// -//// VeriSimDB Gleam Client — Federation operations. -//// -//// VeriSimDB supports federated operation where multiple instances form a -//// cluster, sharing and synchronising octad data across peers. This module -//// provides functions to register and manage peers and to execute cross-node -//// queries. -//// -//// JSON decoding uses the shared codec module for type-safe deserialization. - -import gleam/dict.{type Dict} -import gleam/json -import gleam/list -import verisimdb_client.{type Client} -import verisimdb_client/codec -import verisimdb_client/error.{type VeriSimError} -import verisimdb_client/types.{ - type FederatedQueryResult, type FederationPeer, -} - -/// Peer registration input. -pub type PeerRegistration { - PeerRegistration( - name: String, - url: String, - metadata: Dict(String, String), - ) -} - -/// Federated query request. -pub type FederatedQueryRequest { - FederatedQueryRequest( - query: String, - params: Dict(String, String), - peer_ids: List(String), - timeout: Int, - ) -} - -/// Register a new VeriSimDB instance as a federation peer. -/// -/// Parameters: -/// client — The authenticated client. -/// input — The peer registration details. -/// -/// Returns the registered FederationPeer with server-assigned ID, or an error. -pub fn register_peer( - client: Client, - input: PeerRegistration, -) -> Result(FederationPeer, VeriSimError) { - let body = - json.to_string(json.object([ - #("name", json.string(input.name)), - #("url", json.string(input.url)), - #("metadata", codec.encode_string_dict(input.metadata)), - ])) - case verisimdb_client.do_post(client, "/api/v1/federation/peers", body) { - Ok(resp) -> - case resp.status { - 201 -> codec.decode_federation_peer(resp.body) - status -> Error(error.from_status(status)) - } - Error(err) -> Error(err) - } -} - -/// Retrieve all registered federation peers. -/// -/// Parameters: -/// client — The authenticated client. -/// -/// Returns a list of FederationPeer records, or an error. -pub fn list_peers( - client: Client, -) -> Result(List(FederationPeer), VeriSimError) { - case verisimdb_client.do_get(client, "/api/v1/federation/peers") { - Ok(resp) -> - case resp.status { - 200 -> codec.decode_federation_peers(resp.body) - status -> Error(error.from_status(status)) - } - Error(err) -> Error(err) - } -} - -/// Execute a VCL query across one or more federation peers. -/// -/// If peer_ids is empty, the query is broadcast to all active peers. -/// -/// Parameters: -/// client — The authenticated client. -/// input — The federated query request. -/// -/// Returns aggregated results from all queried peers, or an error. -pub fn federated_query( - client: Client, - input: FederatedQueryRequest, -) -> Result(FederatedQueryResult, VeriSimError) { - let param_pairs = - input.params - |> dict.to_list - |> list.map(fn(pair) { #(pair.0, json.string(pair.1)) }) - let body = - json.to_string(json.object([ - #("query", json.string(input.query)), - #("params", json.object(param_pairs)), - #("peer_ids", json.array(input.peer_ids, json.string)), - #("timeout", json.int(input.timeout)), - ])) - case verisimdb_client.do_post(client, "/api/v1/federation/query", body) { - Ok(resp) -> - case resp.status { - 200 -> codec.decode_federated_query_result(resp.body) - status -> Error(error.from_status(status)) - } - Error(err) -> Error(err) - } -} diff --git a/verisimdb/connectors/clients/gleam/src/verisimdb_client/octad.gleam b/verisimdb/connectors/clients/gleam/src/verisimdb_client/octad.gleam deleted file mode 100644 index ce5a45e3..00000000 --- a/verisimdb/connectors/clients/gleam/src/verisimdb_client/octad.gleam +++ /dev/null @@ -1,136 +0,0 @@ -//// SPDX-License-Identifier: MPL-2.0 -//// (PMPL-1.0-or-later preferred; MPL-2.0 required for Gleam/Hex ecosystem) -//// Copyright (c) 2026 Jonathan D.A. Jewell (hyperpolymath) -//// -//// VeriSimDB Gleam Client — Octad CRUD operations. -//// -//// This module provides create, read, update, delete, and paginated list -//// operations for VeriSimDB octad entities. All functions communicate with -//// the VeriSimDB REST API via the main client module's HTTP helpers. -//// -//// JSON encoding serializes all 8 modality data fields when present. -//// JSON decoding uses gleam/dynamic/decode for type-safe deserialization. - -import gleam/dict -import gleam/dynamic/decode -import gleam/int -import gleam/json -import gleam/list -import gleam/option.{type Option, None, Some} -import gleam/result -import verisimdb_client.{type Client} -import verisimdb_client/codec -import verisimdb_client/error.{type VeriSimError} -import verisimdb_client/types.{type Octad, type OctadInput, type PaginatedResponse} - -/// Create a new octad on the VeriSimDB server. -/// -/// Parameters: -/// client — The authenticated client. -/// input — The octad input describing modalities and data. -/// -/// Returns the newly created Octad with server-assigned ID, or an error. -pub fn create( - client: Client, - input: OctadInput, -) -> Result(Octad, VeriSimError) { - let body = codec.encode_octad_input(input) - case verisimdb_client.do_post(client, "/api/v1/octads", body) { - Ok(resp) -> - case resp.status { - 201 -> codec.decode_octad(resp.body) - status -> Error(error.from_status(status)) - } - Error(err) -> Error(err) - } -} - -/// Retrieve a single octad by its unique identifier. -/// -/// Parameters: -/// client — The authenticated client. -/// id — The octad's unique identifier. -/// -/// Returns the requested Octad, or an error if not found. -pub fn get(client: Client, id: String) -> Result(Octad, VeriSimError) { - case verisimdb_client.do_get(client, "/api/v1/octads/" <> id) { - Ok(resp) -> - case resp.status { - 200 -> codec.decode_octad(resp.body) - status -> Error(error.from_status(status)) - } - Error(err) -> Error(err) - } -} - -/// Update an existing octad with the given input fields. -/// Only the fields present in the input are modified. -/// -/// Parameters: -/// client — The authenticated client. -/// id — The octad's unique identifier. -/// input — The fields to update. -/// -/// Returns the updated Octad, or an error. -pub fn update( - client: Client, - id: String, - input: OctadInput, -) -> Result(Octad, VeriSimError) { - let body = codec.encode_octad_input(input) - case verisimdb_client.do_put(client, "/api/v1/octads/" <> id, body) { - Ok(resp) -> - case resp.status { - 200 -> codec.decode_octad(resp.body) - status -> Error(error.from_status(status)) - } - Error(err) -> Error(err) - } -} - -/// Delete a octad by its unique identifier. -/// -/// Parameters: -/// client — The authenticated client. -/// id — The octad's unique identifier. -/// -/// Returns Ok(True) if deletion succeeded, or an error. -pub fn delete(client: Client, id: String) -> Result(Bool, VeriSimError) { - case verisimdb_client.do_delete(client, "/api/v1/octads/" <> id) { - Ok(resp) -> - case resp.status { - 200 -> Ok(True) - 204 -> Ok(True) - status -> Error(error.from_status(status)) - } - Error(err) -> Error(err) - } -} - -/// Retrieve a paginated list of octads. -/// -/// Parameters: -/// client — The authenticated client. -/// page — Page number (1-indexed). -/// per_page — Number of octads per page. -/// -/// Returns a PaginatedResponse, or an error. -pub fn list( - client: Client, - page: Int, - per_page: Int, -) -> Result(PaginatedResponse, VeriSimError) { - let path = - "/api/v1/octads?page=" - <> int.to_string(page) - <> "&per_page=" - <> int.to_string(per_page) - case verisimdb_client.do_get(client, path) { - Ok(resp) -> - case resp.status { - 200 -> codec.decode_paginated_response(resp.body) - status -> Error(error.from_status(status)) - } - Error(err) -> Error(err) - } -} diff --git a/verisimdb/connectors/clients/gleam/src/verisimdb_client/provenance.gleam b/verisimdb/connectors/clients/gleam/src/verisimdb_client/provenance.gleam deleted file mode 100644 index 11f26a9f..00000000 --- a/verisimdb/connectors/clients/gleam/src/verisimdb_client/provenance.gleam +++ /dev/null @@ -1,100 +0,0 @@ -//// SPDX-License-Identifier: MPL-2.0 -//// (PMPL-1.0-or-later preferred; MPL-2.0 required for Gleam/Hex ecosystem) -//// Copyright (c) 2026 Jonathan D.A. Jewell (hyperpolymath) -//// -//// VeriSimDB Gleam Client — Provenance operations. -//// -//// Every octad maintains an immutable provenance chain — a cryptographically -//// linked sequence of events recording every mutation applied to it. This -//// module provides functions to query chains, record new events, and verify -//// chain integrity. -//// -//// JSON encoding/decoding uses the shared codec module. - -import verisimdb_client.{type Client} -import verisimdb_client/codec -import verisimdb_client/error.{type VeriSimError} -import verisimdb_client/types.{ - type ProvenanceChain, type ProvenanceEvent, type ProvenanceEventInput, -} - -/// Retrieve the complete provenance chain for a octad. -/// -/// The chain is returned in chronological order (oldest first) and includes -/// the verification status. -/// -/// Parameters: -/// client — The authenticated client. -/// octad_id — The unique identifier of the octad. -/// -/// Returns the ProvenanceChain with all events, or an error. -pub fn get_chain( - client: Client, - octad_id: String, -) -> Result(ProvenanceChain, VeriSimError) { - let path = "/api/v1/octads/" <> octad_id <> "/provenance" - case verisimdb_client.do_get(client, path) { - Ok(resp) -> - case resp.status { - 200 -> codec.decode_provenance_chain(resp.body) - status -> Error(error.from_status(status)) - } - Error(err) -> Error(err) - } -} - -/// Record a new provenance event on a octad's chain. -/// -/// The event is cryptographically linked to the previous event. -/// The server assigns the event ID and timestamp. -/// -/// Parameters: -/// client — The authenticated client. -/// octad_id — The unique identifier of the octad. -/// input — The event details to record. -/// -/// Returns the newly created ProvenanceEvent, or an error. -pub fn record_event( - client: Client, - octad_id: String, - input: ProvenanceEventInput, -) -> Result(ProvenanceEvent, VeriSimError) { - let path = "/api/v1/octads/" <> octad_id <> "/provenance" - let body = codec.encode_provenance_event_input(input) - case verisimdb_client.do_post(client, path, body) { - Ok(resp) -> - case resp.status { - 201 -> codec.decode_provenance_event(resp.body) - status -> Error(error.from_status(status)) - } - Error(err) -> Error(err) - } -} - -/// Verify the cryptographic integrity of a octad's provenance chain. -/// -/// Returns Ok(True) if the chain is intact, Ok(False) if tampered, -/// or an error on failure. -/// -/// Parameters: -/// client — The authenticated client. -/// octad_id — The unique identifier of the octad. -pub fn verify( - client: Client, - octad_id: String, -) -> Result(Bool, VeriSimError) { - let path = "/api/v1/octads/" <> octad_id <> "/provenance/verify" - case verisimdb_client.do_post(client, path, "{}") { - Ok(resp) -> - case resp.status { - 200 -> { - case codec.decode_provenance_chain(resp.body) { - Ok(chain) -> Ok(chain.verified) - Error(err) -> Error(err) - } - } - status -> Error(error.from_status(status)) - } - Error(err) -> Error(err) - } -} diff --git a/verisimdb/connectors/clients/gleam/src/verisimdb_client/search.gleam b/verisimdb/connectors/clients/gleam/src/verisimdb_client/search.gleam deleted file mode 100644 index 5c7de393..00000000 --- a/verisimdb/connectors/clients/gleam/src/verisimdb_client/search.gleam +++ /dev/null @@ -1,226 +0,0 @@ -//// SPDX-License-Identifier: MPL-2.0 -//// (PMPL-1.0-or-later preferred; MPL-2.0 required for Gleam/Hex ecosystem) -//// Copyright (c) 2026 Jonathan D.A. Jewell (hyperpolymath) -//// -//// VeriSimDB Gleam Client — Search operations. -//// -//// This module provides multi-modal search capabilities against VeriSimDB, -//// including full-text search, vector similarity search, spatial radius and -//// bounding-box queries, nearest-neighbour lookups, and relationship traversal. -//// -//// JSON decoding uses the shared codec module for type-safe deserialization. - -import gleam/json -import gleam/option.{type Option} -import verisimdb_client.{type Client} -import verisimdb_client/codec -import verisimdb_client/error.{type VeriSimError} -import verisimdb_client/types.{type Modality, type SearchResult} - -/// Parameters for a full-text search query. -pub type TextSearchParams { - TextSearchParams( - query: String, - modalities: List(Modality), - limit: Int, - offset: Int, - ) -} - -/// Parameters for a vector similarity search. -pub type VectorSearchParams { - VectorSearchParams( - vector: List(Float), - model: String, - top_k: Int, - threshold: Float, - ) -} - -/// Parameters for a spatial radius search (point + distance). -pub type SpatialRadiusParams { - SpatialRadiusParams( - latitude: Float, - longitude: Float, - radius_km: Float, - limit: Int, - ) -} - -/// Parameters for a spatial bounding-box search. -pub type SpatialBoundsParams { - SpatialBoundsParams( - min_lat: Float, - min_lon: Float, - max_lat: Float, - max_lon: Float, - limit: Int, - ) -} - -/// Parameters for a nearest-neighbour search by octad ID. -pub type NearestParams { - NearestParams(octad_id: String, top_k: Int, modality: Modality) -} - -/// Parameters for a relationship traversal search. -pub type RelatedParams { - RelatedParams( - octad_id: String, - rel_type: Option(String), - depth: Int, - limit: Int, - ) -} - -/// Perform a full-text search across octad content. -/// -/// Returns a list of SearchResult items ranked by relevance, or an error. -pub fn text( - client: Client, - params: TextSearchParams, -) -> Result(List(SearchResult), VeriSimError) { - let body = - json.to_string(json.object([ - #("query", json.string(params.query)), - #( - "modalities", - json.array(params.modalities, fn(m) { - json.string(types.modality_to_string(m)) - }), - ), - #("limit", json.int(params.limit)), - #("offset", json.int(params.offset)), - ])) - case verisimdb_client.do_post(client, "/api/v1/search/text", body) { - Ok(resp) -> - case resp.status { - 200 -> codec.decode_search_results(resp.body) - status -> Error(error.from_status(status)) - } - Error(err) -> Error(err) - } -} - -/// Perform a vector similarity search using a query embedding. -/// -/// Returns a list of SearchResult items ranked by cosine similarity, or an error. -pub fn vector( - client: Client, - params: VectorSearchParams, -) -> Result(List(SearchResult), VeriSimError) { - let body = - json.to_string(json.object([ - #("vector", json.array(params.vector, json.float)), - #("model", json.string(params.model)), - #("top_k", json.int(params.top_k)), - #("threshold", json.float(params.threshold)), - ])) - case verisimdb_client.do_post(client, "/api/v1/search/vector", body) { - Ok(resp) -> - case resp.status { - 200 -> codec.decode_search_results(resp.body) - status -> Error(error.from_status(status)) - } - Error(err) -> Error(err) - } -} - -/// Find octads within a given radius of a geographic point. -/// -/// Returns a list of SearchResult items within the radius, or an error. -pub fn spatial_radius( - client: Client, - params: SpatialRadiusParams, -) -> Result(List(SearchResult), VeriSimError) { - let body = - json.to_string(json.object([ - #("latitude", json.float(params.latitude)), - #("longitude", json.float(params.longitude)), - #("radius_km", json.float(params.radius_km)), - #("limit", json.int(params.limit)), - ])) - case verisimdb_client.do_post(client, "/api/v1/search/spatial/radius", body) { - Ok(resp) -> - case resp.status { - 200 -> codec.decode_search_results(resp.body) - status -> Error(error.from_status(status)) - } - Error(err) -> Error(err) - } -} - -/// Find octads within a rectangular bounding box. -/// -/// Returns a list of SearchResult items within the bounds, or an error. -pub fn spatial_bounds( - client: Client, - params: SpatialBoundsParams, -) -> Result(List(SearchResult), VeriSimError) { - let body = - json.to_string(json.object([ - #("min_lat", json.float(params.min_lat)), - #("min_lon", json.float(params.min_lon)), - #("max_lat", json.float(params.max_lat)), - #("max_lon", json.float(params.max_lon)), - #("limit", json.int(params.limit)), - ])) - case verisimdb_client.do_post(client, "/api/v1/search/spatial/bounds", body) { - Ok(resp) -> - case resp.status { - 200 -> codec.decode_search_results(resp.body) - status -> Error(error.from_status(status)) - } - Error(err) -> Error(err) - } -} - -/// Find the nearest neighbours of a given octad. -/// -/// Returns a list of SearchResult items ordered by proximity, or an error. -pub fn nearest( - client: Client, - params: NearestParams, -) -> Result(List(SearchResult), VeriSimError) { - let body = - json.to_string(json.object([ - #("octad_id", json.string(params.octad_id)), - #("top_k", json.int(params.top_k)), - #("modality", json.string(types.modality_to_string(params.modality))), - ])) - case verisimdb_client.do_post(client, "/api/v1/search/nearest", body) { - Ok(resp) -> - case resp.status { - 200 -> codec.decode_search_results(resp.body) - status -> Error(error.from_status(status)) - } - Error(err) -> Error(err) - } -} - -/// Traverse relationships from a given octad. -/// -/// Returns a list of SearchResult items connected by relationships, or an error. -pub fn related( - client: Client, - params: RelatedParams, -) -> Result(List(SearchResult), VeriSimError) { - let base_fields = [ - #("octad_id", json.string(params.octad_id)), - #("depth", json.int(params.depth)), - #("limit", json.int(params.limit)), - ] - let fields = case params.rel_type { - option.Some(rt) -> [#("rel_type", json.string(rt)), ..base_fields] - option.None -> base_fields - } - let body = json.to_string(json.object(fields)) - case verisimdb_client.do_post(client, "/api/v1/search/related", body) { - Ok(resp) -> - case resp.status { - 200 -> codec.decode_search_results(resp.body) - status -> Error(error.from_status(status)) - } - Error(err) -> Error(err) - } -} diff --git a/verisimdb/connectors/clients/gleam/src/verisimdb_client/types.gleam b/verisimdb/connectors/clients/gleam/src/verisimdb_client/types.gleam deleted file mode 100644 index 8e403c17..00000000 --- a/verisimdb/connectors/clients/gleam/src/verisimdb_client/types.gleam +++ /dev/null @@ -1,363 +0,0 @@ -//// SPDX-License-Identifier: MPL-2.0 -//// (PMPL-1.0-or-later preferred; MPL-2.0 required for Gleam/Hex ecosystem) -//// Copyright (c) 2026 Jonathan D.A. Jewell (hyperpolymath) -//// -//// VeriSimDB Gleam Client — Core type definitions. -//// -//// This module defines all data structures exchanged between the Gleam client -//// SDK and the VeriSimDB server. Types are defined as Gleam custom types, -//// designed for use with gleam_json for serialization/deserialization. -//// -//// The central entity is the Octad — a six-faceted data object unifying graph, -//// vector, tensor, semantic, document, temporal, provenance, and spatial modalities. - -import gleam/dict.{type Dict} -import gleam/option.{type Option} - -// --------------------------------------------------------------------------- -// Modality -// --------------------------------------------------------------------------- - -/// The eight data modalities supported by VeriSimDB octads. -/// A single octad can participate in multiple modalities simultaneously. -pub type Modality { - Graph - Vector - Tensor - Semantic - Document - Temporal - Provenance - Spatial -} - -/// Convert a modality to its JSON string representation. -pub fn modality_to_string(modality: Modality) -> String { - case modality { - Graph -> "graph" - Vector -> "vector" - Tensor -> "tensor" - Semantic -> "semantic" - Document -> "document" - Temporal -> "temporal" - Provenance -> "provenance" - Spatial -> "spatial" - } -} - -/// Parse a modality from its JSON string representation. -pub fn modality_from_string(s: String) -> Option(Modality) { - case s { - "graph" -> option.Some(Graph) - "vector" -> option.Some(Vector) - "tensor" -> option.Some(Tensor) - "semantic" -> option.Some(Semantic) - "document" -> option.Some(Document) - "temporal" -> option.Some(Temporal) - "provenance" -> option.Some(Provenance) - "spatial" -> option.Some(Spatial) - _ -> option.None - } -} - -// --------------------------------------------------------------------------- -// Modality status -// --------------------------------------------------------------------------- - -/// Which modalities are active on a given octad. -pub type ModalityStatus { - ModalityStatus( - graph: Bool, - vector: Bool, - tensor: Bool, - semantic: Bool, - document: Bool, - temporal: Bool, - provenance: Bool, - spatial: Bool, - ) -} - -/// Default modality status with all modalities disabled. -pub fn default_modality_status() -> ModalityStatus { - ModalityStatus( - graph: False, - vector: False, - tensor: False, - semantic: False, - document: False, - temporal: False, - provenance: False, - spatial: False, - ) -} - -// --------------------------------------------------------------------------- -// Octad status -// --------------------------------------------------------------------------- - -/// Lifecycle state of a octad. -pub type OctadStatus { - Active - Archived - Draft - Deleted -} - -// --------------------------------------------------------------------------- -// Graph modality data -// --------------------------------------------------------------------------- - -/// A directed edge between two octads in the graph modality. -pub type GraphEdge { - GraphEdge( - source: String, - target: String, - rel_type: String, - weight: Float, - metadata: Dict(String, String), - ) -} - -/// Graph-modality data: edges and node properties. -pub type GraphData { - GraphData(edges: List(GraphEdge), properties: Dict(String, String)) -} - -// --------------------------------------------------------------------------- -// Vector modality data -// --------------------------------------------------------------------------- - -/// Embedding vector for vector-modality operations. -pub type VectorData { - VectorData(embedding: List(Float), model: String, dimensions: Int) -} - -// --------------------------------------------------------------------------- -// Tensor modality data -// --------------------------------------------------------------------------- - -/// Multi-dimensional tensor data reference. -pub type TensorData { - TensorData(shape: List(Int), dtype: String, data_ref: String) -} - -// --------------------------------------------------------------------------- -// Document modality data -// --------------------------------------------------------------------------- - -/// Document-modality content: text, format, and language metadata. -pub type DocumentContent { - DocumentContent( - text: String, - format: String, - language: String, - metadata: Dict(String, String), - ) -} - -// --------------------------------------------------------------------------- -// Spatial modality data -// --------------------------------------------------------------------------- - -/// Spatial-modality coordinates and geometry. -pub type SpatialData { - SpatialData( - latitude: Float, - longitude: Float, - altitude: Option(Float), - geometry: Option(String), - crs: String, - ) -} - -// --------------------------------------------------------------------------- -// Octad (core entity) -// --------------------------------------------------------------------------- - -/// The core entity in VeriSimDB — a multi-modal data object. -pub type Octad { - Octad( - id: String, - status: OctadStatus, - modalities: ModalityStatus, - created_at: String, - updated_at: String, - metadata: Dict(String, String), - graph_data: Option(GraphData), - vector_data: Option(VectorData), - tensor_data: Option(TensorData), - content: Option(DocumentContent), - spatial_data: Option(SpatialData), - ) -} - -// --------------------------------------------------------------------------- -// Octad input (for create/update) -// --------------------------------------------------------------------------- - -/// Input structure for creating or updating a octad. -pub type OctadInput { - OctadInput( - graph_data: Option(GraphData), - vector_data: Option(VectorData), - tensor_data: Option(TensorData), - content: Option(DocumentContent), - spatial_data: Option(SpatialData), - metadata: Dict(String, String), - modalities: List(Modality), - ) -} - -// --------------------------------------------------------------------------- -// Drift types -// --------------------------------------------------------------------------- - -/// Drift score measurement. Score ranges from 0.0 (no drift) to 1.0 (maximum). -pub type DriftScore { - DriftScore( - octad_id: String, - score: Float, - components: Dict(String, Float), - measured_at: String, - baseline_at: String, - ) -} - -/// Drift level classification. -pub type DriftLevel { - DriftStable - DriftLow - DriftModerate - DriftHigh - DriftCritical -} - -/// Drift status report with classification and score. -pub type DriftStatusReport { - DriftStatusReport( - octad_id: String, - level: DriftLevel, - score: DriftScore, - message: String, - ) -} - -// --------------------------------------------------------------------------- -// Provenance types -// --------------------------------------------------------------------------- - -/// A single event in a octad's provenance chain. -pub type ProvenanceEvent { - ProvenanceEvent( - event_id: String, - octad_id: String, - event_type: String, - actor: String, - timestamp: String, - details: Dict(String, String), - parent_id: Option(String), - ) -} - -/// Complete provenance chain for a octad. -pub type ProvenanceChain { - ProvenanceChain( - octad_id: String, - events: List(ProvenanceEvent), - verified: Bool, - ) -} - -/// Input for recording a new provenance event. -pub type ProvenanceEventInput { - ProvenanceEventInput( - event_type: String, - actor: String, - details: Dict(String, String), - ) -} - -// --------------------------------------------------------------------------- -// Pagination -// --------------------------------------------------------------------------- - -/// Paginated response wrapping a list of octads. -pub type PaginatedResponse { - PaginatedResponse( - items: List(Octad), - total: Int, - page: Int, - per_page: Int, - total_pages: Int, - ) -} - -// --------------------------------------------------------------------------- -// Search types -// --------------------------------------------------------------------------- - -/// A search result pairing a octad with a relevance score. -pub type SearchResult { - SearchResult(octad: Octad, score: Float) -} - -// --------------------------------------------------------------------------- -// VCL types -// --------------------------------------------------------------------------- - -/// Result of a VCL query execution. -pub type VclResult { - VclResult( - columns: List(String), - rows: List(List(String)), - count: Int, - elapsed_ms: Float, - ) -} - -/// Query execution plan for a VCL statement. -pub type VclExplanation { - VclExplanation( - query: String, - plan: String, - cost: Float, - warnings: List(String), - ) -} - -// --------------------------------------------------------------------------- -// Federation types -// --------------------------------------------------------------------------- - -/// A remote VeriSimDB node in a federated cluster. -pub type FederationPeer { - FederationPeer( - peer_id: String, - name: String, - url: String, - status: String, - last_seen: String, - metadata: Dict(String, String), - ) -} - -/// Result from a single peer in a federated query. -pub type PeerQueryResult { - PeerQueryResult( - peer_id: String, - peer_name: String, - result: VclResult, - elapsed_ms: Float, - error: Option(String), - ) -} - -/// Aggregated result from a federated query. -pub type FederatedQueryResult { - FederatedQueryResult( - results: List(PeerQueryResult), - total: Int, - elapsed_ms: Float, - ) -} diff --git a/verisimdb/connectors/clients/gleam/src/verisimdb_client/vcl.gleam b/verisimdb/connectors/clients/gleam/src/verisimdb_client/vcl.gleam deleted file mode 100644 index 2d2f693b..00000000 --- a/verisimdb/connectors/clients/gleam/src/verisimdb_client/vcl.gleam +++ /dev/null @@ -1,88 +0,0 @@ -//// SPDX-License-Identifier: MPL-2.0 -//// (PMPL-1.0-or-later preferred; MPL-2.0 required for Gleam/Hex ecosystem) -//// Copyright (c) 2026 Jonathan D.A. Jewell (hyperpolymath) -//// -//// VeriSimDB Gleam Client — VCL (VeriSimDB Query Language) operations. -//// -//// VCL is VeriSimDB's native query language for multi-modal queries that span -//// graph traversals, vector similarity, spatial filters, and temporal constraints -//// in a single statement. This module provides execution and explain functions. -//// -//// JSON decoding uses the shared codec module for type-safe deserialization. - -import gleam/dict.{type Dict} -import gleam/json -import gleam/list -import verisimdb_client.{type Client} -import verisimdb_client/codec -import verisimdb_client/error.{type VeriSimError} -import verisimdb_client/types.{type VclExplanation, type VclResult} - -/// Execute a VCL query and return the result set. -/// -/// VCL queries can combine modalities — for example: -/// FIND octads WHERE vector_similar($embedding, 0.8) -/// AND spatial_within(51.5, -0.1, 10km) -/// AND graph_connected("category:science", depth: 2) -/// -/// Parameters: -/// client — The authenticated client. -/// query — The VCL query string. -/// params — Named parameters for parameterised queries. -/// -/// Returns a VclResult with columns, rows, and timing, or an error. -pub fn execute( - client: Client, - query: String, - params: Dict(String, String), -) -> Result(VclResult, VeriSimError) { - let param_pairs = - params - |> dict.to_list - |> list.map(fn(pair) { #(pair.0, json.string(pair.1)) }) - let body = - json.to_string(json.object([ - #("query", json.string(query)), - #("params", json.object(param_pairs)), - ])) - case verisimdb_client.do_post(client, "/api/v1/vcl/execute", body) { - Ok(resp) -> - case resp.status { - 200 -> codec.decode_vcl_result(resp.body) - status -> Error(error.from_status(status)) - } - Error(err) -> Error(err) - } -} - -/// Explain a VCL query's execution plan without running it. -/// -/// Parameters: -/// client — The authenticated client. -/// query — The VCL query string. -/// params — Named parameters. -/// -/// Returns a VclExplanation with the plan, cost, and warnings, or an error. -pub fn explain( - client: Client, - query: String, - params: Dict(String, String), -) -> Result(VclExplanation, VeriSimError) { - let param_pairs = - params - |> dict.to_list - |> list.map(fn(pair) { #(pair.0, json.string(pair.1)) }) - let body = - json.to_string(json.object([ - #("query", json.string(query)), - #("params", json.object(param_pairs)), - ])) - case verisimdb_client.do_post(client, "/api/v1/vcl/explain", body) { - Ok(resp) -> - case resp.status { - 200 -> codec.decode_vcl_explanation(resp.body) - status -> Error(error.from_status(status)) - } - Error(err) -> Error(err) - } -} diff --git a/verisimdb/connectors/clients/gleam/test/verisimdb_client_test.gleam b/verisimdb/connectors/clients/gleam/test/verisimdb_client_test.gleam deleted file mode 100644 index f1d2fdd3..00000000 --- a/verisimdb/connectors/clients/gleam/test/verisimdb_client_test.gleam +++ /dev/null @@ -1,173 +0,0 @@ -//// SPDX-License-Identifier: MPL-2.0 -//// (PMPL-1.0-or-later preferred; MPL-2.0 required for Gleam/Hex ecosystem) -//// Copyright (c) 2026 Jonathan D.A. Jewell (hyperpolymath) -//// -//// VeriSimDB Gleam Client — Test suite. -//// -//// Basic unit tests for the verisimdb_client package. These tests validate -//// type construction, error handling, and client configuration without -//// requiring a running VeriSimDB server. - -import gleam/dict -import gleam/option -import gleeunit -import gleeunit/should -import verisimdb_client.{ApiKey, Basic, Bearer, Client, NoAuth} -import verisimdb_client/error -import verisimdb_client/types - -pub fn main() { - gleeunit.main() -} - -// --------------------------------------------------------------------------- -// Client construction tests -// --------------------------------------------------------------------------- - -pub fn new_client_test() { - let client = verisimdb_client.new("http://localhost:8080") - should.equal(client.base_url, "http://localhost:8080") - should.equal(client.timeout, 30_000) - should.equal(client.auth, NoAuth) -} - -pub fn new_client_strips_trailing_slash_test() { - let client = verisimdb_client.new("http://localhost:8080/") - should.equal(client.base_url, "http://localhost:8080") -} - -pub fn new_client_with_api_key_test() { - let client = - verisimdb_client.new_with_api_key("http://localhost:8080", "test-key") - should.equal(client.auth, ApiKey("test-key")) -} - -pub fn new_client_with_bearer_test() { - let client = - verisimdb_client.new_with_bearer("http://localhost:8080", "my-token") - should.equal(client.auth, Bearer("my-token")) -} - -pub fn new_client_with_basic_auth_test() { - let client = - verisimdb_client.new_with_auth( - "http://localhost:8080", - Basic("user", "pass"), - ) - should.equal(client.auth, Basic("user", "pass")) -} - -// --------------------------------------------------------------------------- -// Modality tests -// --------------------------------------------------------------------------- - -pub fn modality_to_string_test() { - should.equal(types.modality_to_string(types.Graph), "graph") - should.equal(types.modality_to_string(types.Vector), "vector") - should.equal(types.modality_to_string(types.Tensor), "tensor") - should.equal(types.modality_to_string(types.Semantic), "semantic") - should.equal(types.modality_to_string(types.Document), "document") - should.equal(types.modality_to_string(types.Temporal), "temporal") - should.equal(types.modality_to_string(types.Provenance), "provenance") - should.equal(types.modality_to_string(types.Spatial), "spatial") -} - -pub fn modality_from_string_test() { - should.equal(types.modality_from_string("graph"), option.Some(types.Graph)) - should.equal(types.modality_from_string("vector"), option.Some(types.Vector)) - should.equal(types.modality_from_string("unknown"), option.None) -} - -// --------------------------------------------------------------------------- -// Error tests -// --------------------------------------------------------------------------- - -pub fn error_from_status_test() { - should.equal(error.from_status(400), error.BadRequest("Bad request")) - should.equal( - error.from_status(401), - error.Unauthorized("Authentication required"), - ) - should.equal( - error.from_status(403), - error.Forbidden("Insufficient permissions"), - ) - should.equal(error.from_status(404), error.NotFound("Resource not found")) - should.equal(error.from_status(409), error.Conflict("Resource conflict")) - should.equal( - error.from_status(422), - error.ValidationFailed("Input validation failed"), - ) - should.equal(error.from_status(429), error.RateLimited("Too many requests")) - should.equal( - error.from_status(500), - error.InternalError("Internal server error"), - ) - should.equal( - error.from_status(503), - error.ServiceUnavailable("Server temporarily unavailable"), - ) -} - -pub fn error_is_retryable_test() { - // Retryable - should.be_true(error.is_retryable(error.RateLimited("slow down"))) - should.be_true(error.is_retryable(error.InternalError("oops"))) - should.be_true(error.is_retryable(error.ServiceUnavailable("busy"))) - should.be_true(error.is_retryable(error.ConnectionError("disconnected"))) - should.be_true(error.is_retryable(error.TimeoutError("too slow"))) - - // Not retryable - should.be_false(error.is_retryable(error.BadRequest("bad"))) - should.be_false(error.is_retryable(error.Unauthorized("no auth"))) - should.be_false(error.is_retryable(error.NotFound("missing"))) - should.be_false(error.is_retryable(error.Conflict("conflict"))) -} - -pub fn error_message_test() { - should.equal( - error.message(error.BadRequest("invalid")), - "Bad request: invalid", - ) - should.equal( - error.message(error.ConnectionError("refused")), - "Connection error: refused", - ) -} - -// --------------------------------------------------------------------------- -// Type construction tests -// --------------------------------------------------------------------------- - -pub fn default_modality_status_test() { - let ms = types.default_modality_status() - should.be_false(ms.graph) - should.be_false(ms.vector) - should.be_false(ms.tensor) - should.be_false(ms.semantic) - should.be_false(ms.document) - should.be_false(ms.temporal) - should.be_false(ms.provenance) - should.be_false(ms.spatial) -} - -pub fn octad_input_construction_test() { - let input = - types.OctadInput( - graph_data: option.None, - vector_data: option.None, - tensor_data: option.None, - content: option.None, - spatial_data: option.None, - metadata: dict.new(), - modalities: [types.Graph, types.Vector], - ) - should.equal(input.modalities, [types.Graph, types.Vector]) -} - -pub fn provenance_event_input_construction_test() { - let details = dict.from_list([#("key", "value")]) - let input = types.ProvenanceEventInput("annotation", "test-user", details) - should.equal(input.event_type, "annotation") - should.equal(input.actor, "test-user") -} diff --git a/verisimdb/connectors/clients/julia/Project.toml b/verisimdb/connectors/clients/julia/Project.toml deleted file mode 100644 index 546d979d..00000000 --- a/verisimdb/connectors/clients/julia/Project.toml +++ /dev/null @@ -1,19 +0,0 @@ -# SPDX-License-Identifier: MPL-2.0 -# Copyright (c) 2026 Jonathan D.A. Jewell (hyperpolymath) - -name = "VeriSimDBClient" -uuid = "a1b2c3d4-e5f6-7890-abcd-ef1234567890" -authors = ["Jonathan D.A. Jewell "] -version = "0.1.0" - -[deps] -HTTP = "cd3eb016-35fb-5094-929b-558a96fad6f3" -JSON3 = "0f8b85d8-7281-11e9-16c2-39a750bddbf1" -URIs = "5c2747f8-b7ea-4ff2-ba2e-563bfd36b1d4" -Dates = "ade2ca70-3891-5945-98fb-dc099432e06a" - -[compat] -julia = "1.10" -HTTP = "1" -JSON3 = "1" -URIs = "1" diff --git a/verisimdb/connectors/clients/julia/src/VeriSimDBClient.jl b/verisimdb/connectors/clients/julia/src/VeriSimDBClient.jl deleted file mode 100644 index 46b9e758..00000000 --- a/verisimdb/connectors/clients/julia/src/VeriSimDBClient.jl +++ /dev/null @@ -1,72 +0,0 @@ -# SPDX-License-Identifier: MPL-2.0 -# Copyright (c) 2026 Jonathan D.A. Jewell (hyperpolymath) -# -# VeriSimDB Julia Client — Main module. -# -# This is the top-level module for the VeriSimDB Julia client SDK. It aggregates -# all submodules (types, error, client, octad, search, drift, provenance, vcl, -# federation) and re-exports the public API. -# -# Usage: -# using VeriSimDBClient -# client = Client("http://localhost:8080") -# octad = create_octad(client, OctadInput(modalities=[Graph, Vector])) - -module VeriSimDBClient - -using HTTP -using JSON3 -using URIs -using Dates - -# Include submodules in dependency order: -# types.jl and error.jl have no internal dependencies; -# client.jl depends on types and error; -# operation modules depend on client, types, and error. -include("types.jl") -include("error.jl") -include("client.jl") -include("octad.jl") -include("search.jl") -include("drift.jl") -include("provenance.jl") -include("vcl.jl") -include("federation.jl") - -# --- Public exports --- - -# Client -export Client, health - -# Types -export Octad, OctadInput, Modality, ModalityStatus, OctadStatus -export GraphData, GraphEdge, VectorData, TensorData, DocumentContent, SpatialData -export DriftScore, DriftLevel, DriftStatusReport -export ProvenanceEvent, ProvenanceChain, ProvenanceEventInput -export PaginatedResponse, SearchResult -export VclResult, VclExplanation -export FederationPeer - -# Octad CRUD -export create_octad, get_octad, update_octad, delete_octad, list_octads - -# Search -export search_text, search_vector, search_spatial_radius, search_spatial_bounds -export search_nearest, search_related - -# Drift -export get_drift_score, drift_status, normalize_drift - -# Provenance -export get_provenance_chain, record_provenance, verify_provenance - -# VCL -export execute_vcl, explain_vcl - -# Federation -export register_peer, list_peers, federated_query - -# Errors -export VeriSimError, is_retryable - -end # module VeriSimDBClient diff --git a/verisimdb/connectors/clients/julia/src/client.jl b/verisimdb/connectors/clients/julia/src/client.jl deleted file mode 100644 index 70f16e77..00000000 --- a/verisimdb/connectors/clients/julia/src/client.jl +++ /dev/null @@ -1,187 +0,0 @@ -# SPDX-License-Identifier: MPL-2.0 -# Copyright (c) 2026 Jonathan D.A. Jewell (hyperpolymath) -# -# VeriSimDB Julia Client — Connection configuration, authentication, and HTTP transport. -# -# This file defines the Client struct and the internal HTTP helper functions -# used by all other SDK modules to communicate with a VeriSimDB server instance. -# It supports multiple authentication methods (API key, Basic, Bearer token, or -# none) and manages base URL routing, request timeouts, and standard HTTP verbs. - -using HTTP -using JSON3 -using Base64 - -# --------------------------------------------------------------------------- -# Authentication types -# --------------------------------------------------------------------------- - -""" - Auth - -Abstract type for authentication methods. Concrete subtypes: -- `ApiKeyAuth` — X-API-Key header -- `BasicAuth` — HTTP Basic Authentication -- `BearerAuth` — Bearer token -- `NoAuth` — No authentication -""" -abstract type Auth end - -"""API key authentication via X-API-Key header.""" -struct ApiKeyAuth <: Auth - key::String -end - -"""HTTP Basic Authentication (username:password).""" -struct BasicAuth <: Auth - username::String - password::String -end - -"""Bearer token authentication.""" -struct BearerAuth <: Auth - token::String -end - -"""No authentication.""" -struct NoAuth <: Auth end - -# --------------------------------------------------------------------------- -# Client -# --------------------------------------------------------------------------- - -""" - Client - -Holds connection configuration for a VeriSimDB server. - -# Fields -- `base_url::String` — Root URL of the VeriSimDB API (e.g. "http://localhost:8080"). -- `timeout::Int` — Request timeout in seconds. Defaults to 30. -- `auth::Auth` — Authentication method. Defaults to `NoAuth()`. - -# Constructors -- `Client(base_url)` — Unauthenticated client. -- `Client(base_url, auth)` — Client with specific authentication. -- `Client(base_url; timeout=30, auth=NoAuth())` — Full keyword constructor. -""" -struct Client - base_url::String - timeout::Int - auth::Auth -end - -# Convenience constructors. -Client(base_url::String) = Client(rstrip(base_url, '/'), 30, NoAuth()) -Client(base_url::String, auth::Auth) = Client(rstrip(base_url, '/'), 30, auth) -function Client(base_url::String; timeout::Int=30, auth::Auth=NoAuth()) - Client(rstrip(base_url, '/'), timeout, auth) -end - -# --------------------------------------------------------------------------- -# Internal HTTP helpers -# --------------------------------------------------------------------------- - -""" - auth_headers(client::Client) -> Vector{Pair{String,String}} - -Build authentication headers from the client's auth configuration. -Returns a vector of header pairs suitable for passing to HTTP.jl. -""" -function auth_headers(client::Client)::Vector{Pair{String,String}} - headers = Pair{String,String}[] - if client.auth isa ApiKeyAuth - push!(headers, "X-API-Key" => client.auth.key) - elseif client.auth isa BasicAuth - encoded = base64encode("$(client.auth.username):$(client.auth.password)") - push!(headers, "Authorization" => "Basic $encoded") - elseif client.auth isa BearerAuth - push!(headers, "Authorization" => "Bearer $(client.auth.token)") - end - return headers -end - -""" - do_get(client::Client, path::String) -> HTTP.Response - -Send an authenticated GET request to the given API path. -""" -function do_get(client::Client, path::String)::HTTP.Response - url = client.base_url * path - headers = auth_headers(client) - return HTTP.get(url; headers=headers, readtimeout=client.timeout, status_exception=false) -end - -""" - do_post(client::Client, path::String, body) -> HTTP.Response - -Send an authenticated POST request with a JSON body. -""" -function do_post(client::Client, path::String, body)::HTTP.Response - url = client.base_url * path - headers = auth_headers(client) - push!(headers, "Content-Type" => "application/json") - json_body = JSON3.write(body) - return HTTP.post(url; headers=headers, body=json_body, readtimeout=client.timeout, status_exception=false) -end - -""" - do_put(client::Client, path::String, body) -> HTTP.Response - -Send an authenticated PUT request with a JSON body. -""" -function do_put(client::Client, path::String, body)::HTTP.Response - url = client.base_url * path - headers = auth_headers(client) - push!(headers, "Content-Type" => "application/json") - json_body = JSON3.write(body) - return HTTP.put(url; headers=headers, body=json_body, readtimeout=client.timeout, status_exception=false) -end - -""" - do_delete(client::Client, path::String) -> HTTP.Response - -Send an authenticated DELETE request to the given API path. -""" -function do_delete(client::Client, path::String)::HTTP.Response - url = client.base_url * path - headers = auth_headers(client) - return HTTP.delete(url; headers=headers, readtimeout=client.timeout, status_exception=false) -end - -""" - parse_response(::Type{T}, resp::HTTP.Response) -> T - -Parse an HTTP response body as JSON into the specified type T. -Throws a VeriSimError if the response indicates failure. -""" -function parse_response(::Type{T}, resp::HTTP.Response) where T - status = resp.status - body = String(resp.body) - if status >= 400 - throw(error_from_status(status, body)) - end - return JSON3.read(body, T) -end - -# --------------------------------------------------------------------------- -# Health check -# --------------------------------------------------------------------------- - -""" - health(client::Client) -> Bool - -Check whether the VeriSimDB server is reachable and healthy. -Sends a GET request to /health and expects a 200 OK response. -""" -function health(client::Client)::Bool - try - resp = do_get(client, "/health") - return resp.status == 200 - catch e - if e isa VeriSimError - rethrow() - end - throw(ConnectionError("Failed to connect to VeriSimDB server: $(sprint(showerror, e))")) - end -end diff --git a/verisimdb/connectors/clients/julia/src/drift.jl b/verisimdb/connectors/clients/julia/src/drift.jl deleted file mode 100644 index 6a5bf8f3..00000000 --- a/verisimdb/connectors/clients/julia/src/drift.jl +++ /dev/null @@ -1,70 +0,0 @@ -# SPDX-License-Identifier: MPL-2.0 -# Copyright (c) 2026 Jonathan D.A. Jewell (hyperpolymath) -# -# VeriSimDB Julia Client — Drift detection operations. -# -# Drift measures how much a octad's embeddings, relationships, or content -# have diverged from a baseline state (0.0 = no drift, 1.0 = maximum drift). -# This file provides functions to query drift scores, check classified status, -# and trigger re-normalisation. - -""" - get_drift_score(client::Client, octad_id::String) -> DriftScore - -Retrieve the current drift score for a specific octad. - -The drift score is a floating-point value between 0.0 (no drift — fully -aligned with baseline) and 1.0 (maximum drift — completely diverged). - -# Arguments -- `client::Client` — The authenticated client. -- `octad_id::String` — The unique identifier of the octad. - -# Returns -A `DriftScore` with overall score, per-modality components, and timestamps. -""" -function get_drift_score(client::Client, octad_id::String)::DriftScore - resp = do_get(client, "/api/v1/octads/$octad_id/drift") - return parse_response(DriftScore, resp) -end - -""" - drift_status(client::Client, octad_id::String) -> DriftStatusReport - -Retrieve a classified drift status report for a octad. - -The report includes the drift level (Stable, Low, Moderate, High, Critical), -the underlying score, and a human-readable message. - -# Arguments -- `client::Client` — The authenticated client. -- `octad_id::String` — The unique identifier of the octad. - -# Returns -A `DriftStatusReport` with classification and score. -""" -function drift_status(client::Client, octad_id::String)::DriftStatusReport - resp = do_get(client, "/api/v1/octads/$octad_id/drift/status") - return parse_response(DriftStatusReport, resp) -end - -""" - normalize_drift(client::Client, octad_id::String) -> DriftScore - -Trigger re-normalisation of a drifted octad. - -Normalisation recomputes the octad's embeddings and relationship weights -against the current baseline, effectively resetting the drift score. -This is a potentially expensive operation for octads with many modalities. - -# Arguments -- `client::Client` — The authenticated client. -- `octad_id::String` — The unique identifier of the octad to normalise. - -# Returns -The updated `DriftScore` after normalisation. -""" -function normalize_drift(client::Client, octad_id::String)::DriftScore - resp = do_post(client, "/api/v1/octads/$octad_id/drift/normalize", Dict()) - return parse_response(DriftScore, resp) -end diff --git a/verisimdb/connectors/clients/julia/src/error.jl b/verisimdb/connectors/clients/julia/src/error.jl deleted file mode 100644 index 19547e84..00000000 --- a/verisimdb/connectors/clients/julia/src/error.jl +++ /dev/null @@ -1,196 +0,0 @@ -# SPDX-License-Identifier: MPL-2.0 -# Copyright (c) 2026 Jonathan D.A. Jewell (hyperpolymath) -# -# VeriSimDB Julia Client — Error types and handling. -# -# This file defines custom exception types for all failure modes that can occur -# when communicating with a VeriSimDB server. Each error category is a concrete -# subtype of VeriSimError, allowing callers to catch specific error classes or -# handle all VeriSimDB errors generically. - -""" - VeriSimError <: Exception - -Abstract base type for all VeriSimDB client errors. Subtypes represent -specific failure categories (HTTP errors, domain errors, client-side errors). -""" -abstract type VeriSimError <: Exception end - -# --------------------------------------------------------------------------- -# HTTP client errors (4xx) -# --------------------------------------------------------------------------- - -"""HTTP 400 — Malformed request.""" -struct BadRequestError <: VeriSimError - message::String - details::Dict{String,String} -end -BadRequestError(msg::String) = BadRequestError(msg, Dict{String,String}()) - -"""HTTP 401 — Missing or invalid authentication.""" -struct UnauthorizedError <: VeriSimError - message::String -end - -"""HTTP 403 — Insufficient permissions.""" -struct ForbiddenError <: VeriSimError - message::String -end - -"""HTTP 404 — Resource does not exist.""" -struct NotFoundError <: VeriSimError - message::String -end - -"""HTTP 409 — Resource version conflict.""" -struct ConflictError <: VeriSimError - message::String -end - -"""HTTP 422 — Input validation failure.""" -struct ValidationError <: VeriSimError - message::String - details::Dict{String,String} -end -ValidationError(msg::String) = ValidationError(msg, Dict{String,String}()) - -"""HTTP 429 — Too many requests.""" -struct RateLimitedError <: VeriSimError - message::String - retry_after::Union{Int,Nothing} -end -RateLimitedError(msg::String) = RateLimitedError(msg, nothing) - -# --------------------------------------------------------------------------- -# HTTP server errors (5xx) -# --------------------------------------------------------------------------- - -"""HTTP 500 — Unexpected server error.""" -struct InternalServerError <: VeriSimError - message::String -end - -"""HTTP 503 — Server temporarily unavailable.""" -struct ServiceUnavailableError <: VeriSimError - message::String -end - -# --------------------------------------------------------------------------- -# Domain-specific errors -# --------------------------------------------------------------------------- - -"""Octad with given ID does not exist.""" -struct OctadNotFoundError <: VeriSimError - octad_id::String - message::String -end - -"""Requested modality is not enabled on the octad.""" -struct ModalityUnavailableError <: VeriSimError - modality::String - message::String -end - -"""Drift score computation failed.""" -struct DriftComputationError <: VeriSimError - message::String -end - -"""Provenance chain integrity failure.""" -struct ProvenanceInvalidError <: VeriSimError - octad_id::String - message::String -end - -"""VCL syntax error.""" -struct VclParseError <: VeriSimError - query::String - message::String -end - -"""VCL runtime error.""" -struct VclExecutionError <: VeriSimError - query::String - message::String -end - -"""Federation peer communication failure.""" -struct FederationError <: VeriSimError - peer_id::Union{String,Nothing} - message::String -end -FederationError(msg::String) = FederationError(nothing, msg) - -# --------------------------------------------------------------------------- -# Client-side errors -# --------------------------------------------------------------------------- - -"""Network connectivity failure.""" -struct ConnectionError <: VeriSimError - message::String -end - -"""Request exceeded timeout.""" -struct TimeoutError <: VeriSimError - message::String - timeout_ms::Int -end - -"""JSON serialization/deserialization failure.""" -struct SerializationError <: VeriSimError - message::String -end - -# --------------------------------------------------------------------------- -# Utility functions -# --------------------------------------------------------------------------- - -""" - is_retryable(err::VeriSimError) -> Bool - -Check whether an error is retryable. Server errors (5xx), rate limiting (429), -connection errors, and timeouts are generally retryable. Client errors (4xx) -are not, as they indicate a problem with the request itself. -""" -function is_retryable(err::VeriSimError)::Bool - return err isa RateLimitedError || - err isa InternalServerError || - err isa ServiceUnavailableError || - err isa ConnectionError || - err isa TimeoutError -end - -""" - error_from_status(status::Int, body::String) -> VeriSimError - -Construct an appropriate VeriSimError subtype from an HTTP status code and -response body. Used internally by other SDK modules when the server returns -a non-success status code. -""" -function error_from_status(status::Int, body::String)::VeriSimError - msg = isempty(body) ? "HTTP $status" : body - if status == 400 - return BadRequestError(msg) - elseif status == 401 - return UnauthorizedError(msg) - elseif status == 403 - return ForbiddenError(msg) - elseif status == 404 - return NotFoundError(msg) - elseif status == 409 - return ConflictError(msg) - elseif status == 422 - return ValidationError(msg) - elseif status == 429 - return RateLimitedError(msg) - elseif status == 500 - return InternalServerError(msg) - elseif status == 503 - return ServiceUnavailableError(msg) - else - return InternalServerError("Unexpected HTTP status $status: $msg") - end -end - -# Implement Base.showerror for pretty-printing VeriSimDB errors. -Base.showerror(io::IO, e::VeriSimError) = print(io, "VeriSimDB Error: ", e.message) diff --git a/verisimdb/connectors/clients/julia/src/federation.jl b/verisimdb/connectors/clients/julia/src/federation.jl deleted file mode 100644 index d004a2ad..00000000 --- a/verisimdb/connectors/clients/julia/src/federation.jl +++ /dev/null @@ -1,127 +0,0 @@ -# SPDX-License-Identifier: MPL-2.0 -# Copyright (c) 2026 Jonathan D.A. Jewell (hyperpolymath) -# -# VeriSimDB Julia Client — Federation operations. -# -# VeriSimDB supports federated operation where multiple instances form a cluster, -# sharing and synchronising octad data across peers. This file provides functions -# to register and manage federation peers and to execute cross-node queries. - -""" - PeerRegistration - -Input for registering a new federation peer. -""" -struct PeerRegistration - name::String - url::String - metadata::Dict{String,String} -end - -PeerRegistration(name::String, url::String) = PeerRegistration(name, url, Dict{String,String}()) - -JSON3.StructTypes.StructType(::Type{PeerRegistration}) = JSON3.StructTypes.Struct() - -""" - FederatedQueryRequest - -Wraps a VCL query intended for federated execution across cluster peers. -""" -struct FederatedQueryRequest - query::String - params::Dict{String,String} - peer_ids::Vector{String} - timeout::Int -end - -function FederatedQueryRequest( - query::String; - params::Dict{String,String}=Dict{String,String}(), - peer_ids::Vector{String}=String[], - timeout::Int=30000 -) - FederatedQueryRequest(query, params, peer_ids, timeout) -end - -JSON3.StructTypes.StructType(::Type{FederatedQueryRequest}) = JSON3.StructTypes.Struct() - -""" - PeerQueryResult - -Result from a single peer in a federated query. -""" -struct PeerQueryResult - peer_id::String - peer_name::String - result::VclResult - elapsed_ms::Float64 - error::Union{String,Nothing} -end - -JSON3.StructTypes.StructType(::Type{PeerQueryResult}) = JSON3.StructTypes.Struct() - -""" - FederatedQueryResult - -Aggregated result from a federated query across multiple peers. -""" -struct FederatedQueryResult - results::Vector{PeerQueryResult} - total::Int - elapsed_ms::Float64 -end - -JSON3.StructTypes.StructType(::Type{FederatedQueryResult}) = JSON3.StructTypes.Struct() - -""" - register_peer(client, input) -> FederationPeer - -Register a new VeriSimDB instance as a federation peer. - -# Arguments -- `client::Client` — The authenticated client. -- `input::PeerRegistration` — The peer registration details. - -# Returns -The registered `FederationPeer` with server-assigned ID. -""" -function register_peer(client::Client, input::PeerRegistration)::FederationPeer - resp = do_post(client, "/api/v1/federation/peers", input) - return parse_response(FederationPeer, resp) -end - -""" - list_peers(client) -> Vector{FederationPeer} - -Retrieve all registered federation peers. - -# Arguments -- `client::Client` — The authenticated client. - -# Returns -A vector of `FederationPeer` records. -""" -function list_peers(client::Client)::Vector{FederationPeer} - resp = do_get(client, "/api/v1/federation/peers") - return parse_response(Vector{FederationPeer}, resp) -end - -""" - federated_query(client, input) -> FederatedQueryResult - -Execute a VCL query across one or more federation peers. - -If `peer_ids` is empty, the query is broadcast to all active peers. Results -are aggregated with per-peer timing and error information. - -# Arguments -- `client::Client` — The authenticated client. -- `input::FederatedQueryRequest` — The query request. - -# Returns -A `FederatedQueryResult` aggregating all peer responses. -""" -function federated_query(client::Client, input::FederatedQueryRequest)::FederatedQueryResult - resp = do_post(client, "/api/v1/federation/query", input) - return parse_response(FederatedQueryResult, resp) -end diff --git a/verisimdb/connectors/clients/julia/src/octad.jl b/verisimdb/connectors/clients/julia/src/octad.jl deleted file mode 100644 index d0f63df7..00000000 --- a/verisimdb/connectors/clients/julia/src/octad.jl +++ /dev/null @@ -1,114 +0,0 @@ -# SPDX-License-Identifier: MPL-2.0 -# Copyright (c) 2026 Jonathan D.A. Jewell (hyperpolymath) -# -# VeriSimDB Julia Client — Octad CRUD operations. -# -# This file provides create, read, update, delete, and paginated list -# operations for VeriSimDB octad entities. All functions communicate with -# the VeriSimDB REST API via the Client's HTTP helpers. - -""" - create_octad(client::Client, input::OctadInput) -> Octad - -Create a new octad on the VeriSimDB server. - -# Arguments -- `client::Client` — The authenticated client. -- `input::OctadInput` — The octad input describing modalities and data. - -# Returns -The newly created `Octad` with server-assigned ID and timestamps. - -# Throws -`VeriSimError` on HTTP or server failure. -""" -function create_octad(client::Client, input::OctadInput)::Octad - resp = do_post(client, "/api/v1/octads", input) - return parse_response(Octad, resp) -end - -""" - get_octad(client::Client, id::String) -> Octad - -Retrieve a single octad by its unique identifier. - -# Arguments -- `client::Client` — The authenticated client. -- `id::String` — The octad's unique identifier. - -# Returns -The requested `Octad`. - -# Throws -`NotFoundError` if the octad does not exist. -""" -function get_octad(client::Client, id::String)::Octad - resp = do_get(client, "/api/v1/octads/$id") - return parse_response(Octad, resp) -end - -""" - update_octad(client::Client, id::String, input::OctadInput) -> Octad - -Update an existing octad with the given input fields. -Only the fields present in the input are modified; others remain unchanged. - -# Arguments -- `client::Client` — The authenticated client. -- `id::String` — The octad's unique identifier. -- `input::OctadInput` — The fields to update. - -# Returns -The updated `Octad`. - -# Throws -`VeriSimError` on failure. -""" -function update_octad(client::Client, id::String, input::OctadInput)::Octad - resp = do_put(client, "/api/v1/octads/$id", input) - return parse_response(Octad, resp) -end - -""" - delete_octad(client::Client, id::String) -> Bool - -Delete a octad by its unique identifier. - -# Arguments -- `client::Client` — The authenticated client. -- `id::String` — The octad's unique identifier. - -# Returns -`true` if the octad was successfully deleted. - -# Throws -`VeriSimError` on failure. -""" -function delete_octad(client::Client, id::String)::Bool - resp = do_delete(client, "/api/v1/octads/$id") - status = resp.status - if status == 204 || status == 200 - return true - end - throw(error_from_status(status, String(resp.body))) -end - -""" - list_octads(client::Client; page::Int=1, per_page::Int=20) -> PaginatedResponse - -Retrieve a paginated list of octads. - -# Keyword Arguments -- `page::Int` — Page number (1-indexed). Defaults to 1. -- `per_page::Int` — Number of octads per page. Defaults to 20. - -# Returns -A `PaginatedResponse` containing octads and pagination metadata. - -# Throws -`VeriSimError` on failure. -""" -function list_octads(client::Client; page::Int=1, per_page::Int=20)::PaginatedResponse - resp = do_get(client, "/api/v1/octads?page=$page&per_page=$per_page") - return parse_response(PaginatedResponse, resp) -end diff --git a/verisimdb/connectors/clients/julia/src/provenance.jl b/verisimdb/connectors/clients/julia/src/provenance.jl deleted file mode 100644 index 2e235f02..00000000 --- a/verisimdb/connectors/clients/julia/src/provenance.jl +++ /dev/null @@ -1,76 +0,0 @@ -# SPDX-License-Identifier: MPL-2.0 -# Copyright (c) 2026 Jonathan D.A. Jewell (hyperpolymath) -# -# VeriSimDB Julia Client — Provenance operations. -# -# Every octad in VeriSimDB maintains an immutable provenance chain — a -# cryptographically linked sequence of events recording every mutation -# applied to the octad. This file provides functions to query chains, -# record new events, and verify chain integrity. - -""" - get_provenance_chain(client::Client, octad_id::String) -> ProvenanceChain - -Retrieve the complete provenance chain for a octad. - -The chain is returned in chronological order (oldest event first) and -includes the verification status. - -# Arguments -- `client::Client` — The authenticated client. -- `octad_id::String` — The unique identifier of the octad. - -# Returns -A `ProvenanceChain` containing all events and verification status. -""" -function get_provenance_chain(client::Client, octad_id::String)::ProvenanceChain - resp = do_get(client, "/api/v1/octads/$octad_id/provenance") - return parse_response(ProvenanceChain, resp) -end - -""" - record_provenance(client, octad_id, input) -> ProvenanceEvent - -Record a new provenance event on a octad's chain. - -The event is cryptographically linked to the previous event in the chain. -The server assigns the event ID and timestamp. - -# Arguments -- `client::Client` — The authenticated client. -- `octad_id::String` — The unique identifier of the octad. -- `input::ProvenanceEventInput` — The event details to record. - -# Returns -The newly created `ProvenanceEvent` with server-assigned fields. -""" -function record_provenance( - client::Client, - octad_id::String, - input::ProvenanceEventInput -)::ProvenanceEvent - resp = do_post(client, "/api/v1/octads/$octad_id/provenance", input) - return parse_response(ProvenanceEvent, resp) -end - -""" - verify_provenance(client::Client, octad_id::String) -> Bool - -Verify the cryptographic integrity of a octad's provenance chain. - -The server traverses the entire chain, checking each event's hash link to -its parent. Returns `true` if the chain is intact, `false` if tampering -is detected. - -# Arguments -- `client::Client` — The authenticated client. -- `octad_id::String` — The unique identifier of the octad. - -# Returns -`true` if the provenance chain is verified intact. -""" -function verify_provenance(client::Client, octad_id::String)::Bool - resp = do_post(client, "/api/v1/octads/$octad_id/provenance/verify", Dict()) - chain = parse_response(ProvenanceChain, resp) - return chain.verified -end diff --git a/verisimdb/connectors/clients/julia/src/search.jl b/verisimdb/connectors/clients/julia/src/search.jl deleted file mode 100644 index 04c7dad8..00000000 --- a/verisimdb/connectors/clients/julia/src/search.jl +++ /dev/null @@ -1,163 +0,0 @@ -# SPDX-License-Identifier: MPL-2.0 -# Copyright (c) 2026 Jonathan D.A. Jewell (hyperpolymath) -# -# VeriSimDB Julia Client — Search operations. -# -# This file provides multi-modal search capabilities against VeriSimDB, -# including full-text search, vector similarity search, spatial radius and -# bounding-box queries, nearest-neighbour lookups, and relationship traversal. - -""" - search_text(client, query; modalities=Modality[], limit=20, offset=0) -> Vector{SearchResult} - -Perform a full-text search across octad content. - -# Arguments -- `client::Client` — The authenticated client. -- `query::String` — The text query string. - -# Keyword Arguments -- `modalities::Vector{Modality}` — Filter by specific modalities. -- `limit::Int` — Maximum results to return. -- `offset::Int` — Number of results to skip. - -# Returns -A vector of `SearchResult` items ranked by relevance. -""" -function search_text( - client::Client, - query::String; - modalities::Vector{Modality}=Modality[], - limit::Int=20, - offset::Int=0 -)::Vector{SearchResult} - body = Dict( - "query" => query, - "modalities" => [string(m) for m in modalities], - "limit" => limit, - "offset" => offset - ) - resp = do_post(client, "/api/v1/search/text", body) - return parse_response(Vector{SearchResult}, resp) -end - -""" - search_vector(client, vector; model="", top_k=10, threshold=0.0) -> Vector{SearchResult} - -Perform a vector similarity search using a query embedding. - -# Arguments -- `client::Client` — The authenticated client. -- `vector::Vector{Float64}` — The query embedding vector. - -# Keyword Arguments -- `model::String` — Name of the embedding model. -- `top_k::Int` — Number of nearest results. -- `threshold::Float64` — Minimum similarity threshold. -""" -function search_vector( - client::Client, - vector::Vector{Float64}; - model::String="", - top_k::Int=10, - threshold::Float64=0.0 -)::Vector{SearchResult} - body = Dict( - "vector" => vector, - "model" => model, - "top_k" => top_k, - "threshold" => threshold - ) - resp = do_post(client, "/api/v1/search/vector", body) - return parse_response(Vector{SearchResult}, resp) -end - -""" - search_spatial_radius(client; latitude, longitude, radius_km, limit=20) -> Vector{SearchResult} - -Find octads within a given radius of a geographic point. -""" -function search_spatial_radius( - client::Client; - latitude::Float64, - longitude::Float64, - radius_km::Float64, - limit::Int=20 -)::Vector{SearchResult} - body = Dict( - "latitude" => latitude, - "longitude" => longitude, - "radius_km" => radius_km, - "limit" => limit - ) - resp = do_post(client, "/api/v1/search/spatial/radius", body) - return parse_response(Vector{SearchResult}, resp) -end - -""" - search_spatial_bounds(client; min_lat, min_lon, max_lat, max_lon, limit=20) -> Vector{SearchResult} - -Find octads within a rectangular bounding box. -""" -function search_spatial_bounds( - client::Client; - min_lat::Float64, - min_lon::Float64, - max_lat::Float64, - max_lon::Float64, - limit::Int=20 -)::Vector{SearchResult} - body = Dict( - "min_lat" => min_lat, - "min_lon" => min_lon, - "max_lat" => max_lat, - "max_lon" => max_lon, - "limit" => limit - ) - resp = do_post(client, "/api/v1/search/spatial/bounds", body) - return parse_response(Vector{SearchResult}, resp) -end - -""" - search_nearest(client, octad_id; top_k=10, modality=Vector) -> Vector{SearchResult} - -Find the nearest neighbours of a given octad. -""" -function search_nearest( - client::Client, - octad_id::String; - top_k::Int=10, - modality::Modality=Vector -)::Vector{SearchResult} - body = Dict( - "octad_id" => octad_id, - "top_k" => top_k, - "modality" => string(modality) - ) - resp = do_post(client, "/api/v1/search/nearest", body) - return parse_response(Vector{SearchResult}, resp) -end - -""" - search_related(client, octad_id; rel_type=nothing, depth=1, limit=20) -> Vector{SearchResult} - -Traverse relationships from a given octad. -""" -function search_related( - client::Client, - octad_id::String; - rel_type::Union{String,Nothing}=nothing, - depth::Int=1, - limit::Int=20 -)::Vector{SearchResult} - body = Dict{String,Any}( - "octad_id" => octad_id, - "depth" => depth, - "limit" => limit - ) - if !isnothing(rel_type) - body["rel_type"] = rel_type - end - resp = do_post(client, "/api/v1/search/related", body) - return parse_response(Vector{SearchResult}, resp) -end diff --git a/verisimdb/connectors/clients/julia/src/types.jl b/verisimdb/connectors/clients/julia/src/types.jl deleted file mode 100644 index f8558739..00000000 --- a/verisimdb/connectors/clients/julia/src/types.jl +++ /dev/null @@ -1,424 +0,0 @@ -# SPDX-License-Identifier: MPL-2.0 -# Copyright (c) 2026 Jonathan D.A. Jewell (hyperpolymath) -# -# VeriSimDB Julia Client — Core type definitions. -# -# This file defines all data structures exchanged between the Julia client SDK -# and the VeriSimDB server. Types are Julia structs with JSON3.StructTypes -# registration for automatic serialization/deserialization. -# -# The central entity in VeriSimDB is the Octad — a six-faceted data object that -# unifies graph, vector, tensor, semantic, document, temporal, provenance, and -# spatial modalities into a single addressable record. - -using JSON3 -using Dates - -# --------------------------------------------------------------------------- -# Modality -# --------------------------------------------------------------------------- - -""" - Modality - -Enumeration of the eight data modalities supported by VeriSimDB octads. -A single octad can participate in multiple modalities simultaneously. -""" -@enum Modality begin - Graph - Vector - Tensor - Semantic - Document - Temporal - Provenance - Spatial -end - -""" - ModalityStatus - -Indicates which modalities are active on a given octad. -Each field is a Bool; true means the modality is enabled. -""" -struct ModalityStatus - graph::Bool - vector::Bool - tensor::Bool - semantic::Bool - document::Bool - temporal::Bool - provenance::Bool - spatial::Bool -end - -# Default constructor with all modalities disabled. -ModalityStatus() = ModalityStatus(false, false, false, false, false, false, false, false) - -JSON3.StructTypes.StructType(::Type{ModalityStatus}) = JSON3.StructTypes.Struct() - -# --------------------------------------------------------------------------- -# Octad status -# --------------------------------------------------------------------------- - -""" - OctadStatus - -Lifecycle state of a octad: active, archived, draft, or deleted. -""" -@enum OctadStatus begin - Active - Archived - Draft - Deleted -end - -# --------------------------------------------------------------------------- -# Graph modality data -# --------------------------------------------------------------------------- - -""" - GraphEdge - -A directed relationship between two octads in the graph modality. -""" -struct GraphEdge - source::String - target::String - rel_type::String - weight::Float64 - metadata::Dict{String,String} -end - -JSON3.StructTypes.StructType(::Type{GraphEdge}) = JSON3.StructTypes.Struct() - -""" - GraphData - -Graph-modality data for a octad: edges and node properties. -""" -struct GraphData - edges::Vector{GraphEdge} - properties::Dict{String,String} -end - -JSON3.StructTypes.StructType(::Type{GraphData}) = JSON3.StructTypes.Struct() - -# --------------------------------------------------------------------------- -# Vector modality data -# --------------------------------------------------------------------------- - -""" - VectorData - -Embedding vector data for vector-modality operations such as similarity -search and nearest-neighbour queries. -""" -struct VectorData - embedding::Vector{Float64} - model::String - dimensions::Int -end - -JSON3.StructTypes.StructType(::Type{VectorData}) = JSON3.StructTypes.Struct() - -# --------------------------------------------------------------------------- -# Tensor modality data -# --------------------------------------------------------------------------- - -""" - TensorData - -Multi-dimensional tensor data reference. The actual tensor data is stored -externally; `data_ref` is a URI pointing to the storage location. -""" -struct TensorData - shape::Vector{Int} - dtype::String - data_ref::String -end - -JSON3.StructTypes.StructType(::Type{TensorData}) = JSON3.StructTypes.Struct() - -# --------------------------------------------------------------------------- -# Document modality data -# --------------------------------------------------------------------------- - -""" - DocumentContent - -Document-modality content: raw text, structured format, and language metadata. -""" -struct DocumentContent - text::String - format::String # e.g. "plain", "markdown", "html" - language::String # ISO 639-1 language code - metadata::Dict{String,String} -end - -JSON3.StructTypes.StructType(::Type{DocumentContent}) = JSON3.StructTypes.Struct() - -# --------------------------------------------------------------------------- -# Spatial modality data -# --------------------------------------------------------------------------- - -""" - SpatialData - -Spatial-modality coordinates and geometry. Supports WGS-84 and other CRS. -""" -struct SpatialData - latitude::Float64 - longitude::Float64 - altitude::Union{Float64,Nothing} - geometry::Union{String,Nothing} # GeoJSON geometry string - crs::String # e.g. "EPSG:4326" -end - -JSON3.StructTypes.StructType(::Type{SpatialData}) = JSON3.StructTypes.Struct() - -# --------------------------------------------------------------------------- -# Octad (core entity) -# --------------------------------------------------------------------------- - -""" - Octad - -The core entity in VeriSimDB — a multi-modal data object unifying graph, -vector, tensor, semantic, document, temporal, provenance, and spatial modalities -into a single addressable record. -""" -struct Octad - id::String - status::OctadStatus - modalities::ModalityStatus - created_at::String # ISO 8601 timestamp - updated_at::String # ISO 8601 timestamp - metadata::Dict{String,String} - graph_data::Union{GraphData,Nothing} - vector_data::Union{VectorData,Nothing} - tensor_data::Union{TensorData,Nothing} - content::Union{DocumentContent,Nothing} - spatial_data::Union{SpatialData,Nothing} -end - -JSON3.StructTypes.StructType(::Type{Octad}) = JSON3.StructTypes.Struct() - -# --------------------------------------------------------------------------- -# Octad input (for create/update) -# --------------------------------------------------------------------------- - -""" - OctadInput - -Input structure for creating or updating a octad. Optional fields use -`Union{T, Nothing}` (Julia's equivalent of Option/Maybe). -""" -struct OctadInput - graph_data::Union{GraphData,Nothing} - vector_data::Union{VectorData,Nothing} - tensor_data::Union{TensorData,Nothing} - content::Union{DocumentContent,Nothing} - spatial_data::Union{SpatialData,Nothing} - metadata::Dict{String,String} - modalities::Vector{Modality} -end - -# Convenience constructor with keyword arguments. -function OctadInput(; - graph_data::Union{GraphData,Nothing}=nothing, - vector_data::Union{VectorData,Nothing}=nothing, - tensor_data::Union{TensorData,Nothing}=nothing, - content::Union{DocumentContent,Nothing}=nothing, - spatial_data::Union{SpatialData,Nothing}=nothing, - metadata::Dict{String,String}=Dict{String,String}(), - modalities::Vector{Modality}=Modality[] -) - OctadInput(graph_data, vector_data, tensor_data, content, spatial_data, metadata, modalities) -end - -JSON3.StructTypes.StructType(::Type{OctadInput}) = JSON3.StructTypes.Struct() - -# --------------------------------------------------------------------------- -# Drift types -# --------------------------------------------------------------------------- - -""" - DriftScore - -Drift measurement for a octad. The `score` field ranges from 0.0 (no drift, -fully aligned with baseline) to 1.0 (maximum drift, completely diverged). -The `components` dictionary breaks down the score by modality. -""" -struct DriftScore - octad_id::String - score::Float64 - components::Dict{String,Float64} - measured_at::String # ISO 8601 - baseline_at::String # ISO 8601 -end - -JSON3.StructTypes.StructType(::Type{DriftScore}) = JSON3.StructTypes.Struct() - -""" - DriftLevel - -Classification of drift severity. -""" -@enum DriftLevel begin - DriftStable - DriftLow - DriftModerate - DriftHigh - DriftCritical -end - -""" - DriftStatusReport - -Classified drift status for a octad, combining the numeric score with -a human-readable level and message. -""" -struct DriftStatusReport - octad_id::String - level::DriftLevel - score::DriftScore - message::String -end - -JSON3.StructTypes.StructType(::Type{DriftStatusReport}) = JSON3.StructTypes.Struct() - -# --------------------------------------------------------------------------- -# Provenance types -# --------------------------------------------------------------------------- - -""" - ProvenanceEvent - -A single event in a octad's provenance chain. Each event is cryptographically -linked to its parent, forming an immutable audit trail. -""" -struct ProvenanceEvent - event_id::String - octad_id::String - event_type::String # e.g. "created", "updated", "merged", "split" - actor::String - timestamp::String # ISO 8601 - details::Dict{String,String} - parent_id::Union{String,Nothing} -end - -JSON3.StructTypes.StructType(::Type{ProvenanceEvent}) = JSON3.StructTypes.Struct() - -""" - ProvenanceChain - -Complete provenance history for a octad, including verification status. -""" -struct ProvenanceChain - octad_id::String - events::Vector{ProvenanceEvent} - verified::Bool -end - -JSON3.StructTypes.StructType(::Type{ProvenanceChain}) = JSON3.StructTypes.Struct() - -""" - ProvenanceEventInput - -Input for recording a new provenance event. -""" -struct ProvenanceEventInput - event_type::String - actor::String - details::Dict{String,String} -end - -JSON3.StructTypes.StructType(::Type{ProvenanceEventInput}) = JSON3.StructTypes.Struct() - -# --------------------------------------------------------------------------- -# Pagination -# --------------------------------------------------------------------------- - -""" - PaginatedResponse - -Paginated response wrapping a list of octads with page metadata. -""" -struct PaginatedResponse - items::Vector{Octad} - total::Int - page::Int - per_page::Int - total_pages::Int -end - -JSON3.StructTypes.StructType(::Type{PaginatedResponse}) = JSON3.StructTypes.Struct() - -# --------------------------------------------------------------------------- -# Search types -# --------------------------------------------------------------------------- - -""" - SearchResult - -A search result pairing a octad with a relevance score (0.0 to 1.0). -""" -struct SearchResult - octad::Octad - score::Float64 -end - -JSON3.StructTypes.StructType(::Type{SearchResult}) = JSON3.StructTypes.Struct() - -# --------------------------------------------------------------------------- -# VCL types -# --------------------------------------------------------------------------- - -""" - VclResult - -Result of a VCL query execution, containing columnar data and timing. -""" -struct VclResult - columns::Vector{String} - rows::Vector{Vector{String}} - count::Int - elapsed_ms::Float64 -end - -JSON3.StructTypes.StructType(::Type{VclResult}) = JSON3.StructTypes.Struct() - -""" - VclExplanation - -Query execution plan for a VCL statement, showing cost estimates and warnings. -""" -struct VclExplanation - query::String - plan::String - cost::Float64 - warnings::Vector{String} -end - -JSON3.StructTypes.StructType(::Type{VclExplanation}) = JSON3.StructTypes.Struct() - -# --------------------------------------------------------------------------- -# Federation types -# --------------------------------------------------------------------------- - -""" - FederationPeer - -A remote VeriSimDB node in a federated cluster. -""" -struct FederationPeer - peer_id::String - name::String - url::String - status::String # "active", "inactive", "syncing" - last_seen::String # ISO 8601 - metadata::Dict{String,String} -end - -JSON3.StructTypes.StructType(::Type{FederationPeer}) = JSON3.StructTypes.Struct() diff --git a/verisimdb/connectors/clients/julia/src/vcl.jl b/verisimdb/connectors/clients/julia/src/vcl.jl deleted file mode 100644 index 42578b9b..00000000 --- a/verisimdb/connectors/clients/julia/src/vcl.jl +++ /dev/null @@ -1,66 +0,0 @@ -# SPDX-License-Identifier: MPL-2.0 -# Copyright (c) 2026 Jonathan D.A. Jewell (hyperpolymath) -# -# VeriSimDB Julia Client — VCL (VeriSimDB Query Language) operations. -# -# VCL is VeriSimDB's native query language for multi-modal queries that span -# graph traversals, vector similarity, spatial filters, and temporal constraints -# in a single statement. This file provides execution and explain functions. - -""" - execute_vcl(client, query; params=Dict()) -> VclResult - -Execute a VCL query and return the result set. - -VCL queries can combine modalities — for example: -``` -FIND octads WHERE vector_similar(\$embedding, 0.8) - AND spatial_within(51.5, -0.1, 10km) - AND graph_connected("category:science", depth: 2) -``` - -# Arguments -- `client::Client` — The authenticated client. -- `query::String` — The VCL query string. - -# Keyword Arguments -- `params::Dict{String,String}` — Named parameters for parameterised queries. - -# Returns -A `VclResult` containing columns, rows, count, and execution time. -""" -function execute_vcl( - client::Client, - query::String; - params::Dict{String,String}=Dict{String,String}() -)::VclResult - body = Dict("query" => query, "params" => params) - resp = do_post(client, "/api/v1/vcl/execute", body) - return parse_response(VclResult, resp) -end - -""" - explain_vcl(client, query; params=Dict()) -> VclExplanation - -Return the query execution plan for a VCL statement without running it. -Useful for debugging and optimising queries. - -# Arguments -- `client::Client` — The authenticated client. -- `query::String` — The VCL query string. - -# Keyword Arguments -- `params::Dict{String,String}` — Named parameters. - -# Returns -A `VclExplanation` containing the plan, estimated cost, and warnings. -""" -function explain_vcl( - client::Client, - query::String; - params::Dict{String,String}=Dict{String,String}() -)::VclExplanation - body = Dict("query" => query, "params" => params) - resp = do_post(client, "/api/v1/vcl/explain", body) - return parse_response(VclExplanation, resp) -end diff --git a/verisimdb/connectors/clients/julia/test/runtests.jl b/verisimdb/connectors/clients/julia/test/runtests.jl deleted file mode 100644 index c1821046..00000000 --- a/verisimdb/connectors/clients/julia/test/runtests.jl +++ /dev/null @@ -1,126 +0,0 @@ -# SPDX-License-Identifier: MPL-2.0 -# Copyright (c) 2026 Jonathan D.A. Jewell (hyperpolymath) -# -# VeriSimDB Julia Client — Test suite. -# -# Basic unit tests for the VeriSimDBClient package. These tests validate -# type construction, error handling, and client configuration without -# requiring a running VeriSimDB server. - -using Test -using VeriSimDBClient - -@testset "VeriSimDBClient" begin - - @testset "Client construction" begin - # Unauthenticated client - c = Client("http://localhost:8080") - @test c.base_url == "http://localhost:8080" - @test c.timeout == 30 - @test c.auth isa NoAuth - - # Client with API key - c_api = Client("http://localhost:8080", ApiKeyAuth("test-key")) - @test c_api.auth isa ApiKeyAuth - - # Client with Bearer token - c_bearer = Client("http://localhost:8080", BearerAuth("my-token")) - @test c_bearer.auth isa BearerAuth - - # Client with Basic auth - c_basic = Client("http://localhost:8080", BasicAuth("user", "pass")) - @test c_basic.auth isa BasicAuth - - # Trailing slash is stripped - c_slash = Client("http://localhost:8080/") - @test c_slash.base_url == "http://localhost:8080" - - # Keyword constructor - c_kw = Client("http://localhost:8080"; timeout=60) - @test c_kw.timeout == 60 - end - - @testset "Type construction" begin - # ModalityStatus defaults - ms = ModalityStatus() - @test ms.graph == false - @test ms.vector == false - - # OctadInput keyword constructor - hi = OctadInput(modalities=[Graph, Vector]) - @test length(hi.modalities) == 2 - @test hi.graph_data === nothing - @test hi.metadata == Dict{String,String}() - - # ProvenanceEventInput - pei = ProvenanceEventInput("annotation", "test-user", Dict("key" => "value")) - @test pei.event_type == "annotation" - @test pei.actor == "test-user" - - # PeerRegistration - pr = PeerRegistration("peer-1", "http://peer1:8080") - @test pr.name == "peer-1" - @test pr.metadata == Dict{String,String}() - - # FederatedQueryRequest keyword constructor - fqr = FederatedQueryRequest("FIND octads"; timeout=5000) - @test fqr.query == "FIND octads" - @test fqr.timeout == 5000 - @test fqr.peer_ids == String[] - end - - @testset "Error types" begin - # Error construction - e_bad = BadRequestError("invalid input") - @test e_bad.message == "invalid input" - @test e_bad isa VeriSimError - - e_notfound = NotFoundError("not found") - @test e_notfound isa VeriSimError - - # Retryable errors - @test is_retryable(RateLimitedError("slow down")) == true - @test is_retryable(InternalServerError("oops")) == true - @test is_retryable(ServiceUnavailableError("busy")) == true - @test is_retryable(ConnectionError("disconnected")) == true - @test is_retryable(TimeoutError("too slow", 30000)) == true - - # Non-retryable errors - @test is_retryable(BadRequestError("bad")) == false - @test is_retryable(UnauthorizedError("no auth")) == false - @test is_retryable(NotFoundError("missing")) == false - @test is_retryable(ConflictError("conflict")) == false - - # error_from_status - @test error_from_status(400, "bad") isa BadRequestError - @test error_from_status(401, "unauth") isa UnauthorizedError - @test error_from_status(403, "forbidden") isa ForbiddenError - @test error_from_status(404, "missing") isa NotFoundError - @test error_from_status(409, "conflict") isa ConflictError - @test error_from_status(422, "invalid") isa ValidationError - @test error_from_status(429, "limit") isa RateLimitedError - @test error_from_status(500, "error") isa InternalServerError - @test error_from_status(503, "unavailable") isa ServiceUnavailableError - @test error_from_status(999, "unknown") isa InternalServerError - end - - @testset "Modality enum" begin - @test Graph isa Modality - @test Vector isa Modality - @test Tensor isa Modality - @test Semantic isa Modality - @test Document isa Modality - @test Temporal isa Modality - @test Provenance isa Modality - @test Spatial isa Modality - end - - @testset "DriftLevel enum" begin - @test DriftStable isa DriftLevel - @test DriftLow isa DriftLevel - @test DriftModerate isa DriftLevel - @test DriftHigh isa DriftLevel - @test DriftCritical isa DriftLevel - end - -end # @testset "VeriSimDBClient" diff --git a/verisimdb/connectors/clients/rescript/.gitignore b/verisimdb/connectors/clients/rescript/.gitignore deleted file mode 100644 index 14260731..00000000 --- a/verisimdb/connectors/clients/rescript/.gitignore +++ /dev/null @@ -1,3 +0,0 @@ -# ReScript build output -/lib/ -node_modules/ diff --git a/verisimdb/connectors/clients/rescript/deno.json b/verisimdb/connectors/clients/rescript/deno.json deleted file mode 100644 index 21d1b5c0..00000000 --- a/verisimdb/connectors/clients/rescript/deno.json +++ /dev/null @@ -1,29 +0,0 @@ -{ - "name": "@hyperpolymath/verisimdb-client", - "version": "0.1.1", - "license": "MPL-2.0", - "description": "ReScript client SDK for VeriSimDB — the 8-modality database with drift detection and self-normalisation. CRUD, search, drift scoring, provenance tracking, VCL queries, and federation. Zero npm dependencies.", - "exports": { - ".": "./src/VeriSimClient.res.mjs", - "./types": "./src/VeriSimTypes.res.mjs", - "./hexad": "./src/VeriSimHexad.res.mjs", - "./search": "./src/VeriSimSearch.res.mjs", - "./drift": "./src/VeriSimDrift.res.mjs", - "./provenance": "./src/VeriSimProvenance.res.mjs", - "./vcl": "./src/VeriSimVcl.res.mjs", - "./federation": "./src/VeriSimFederation.res.mjs", - "./error": "./src/VeriSimError.res.mjs" - }, - "publish": { - "include": [ - "src/*.res.mjs", - "src/*.res", - "LICENSE", - "README.adoc", - "deno.json" - ], - "exclude": [ - "!src/*.res.mjs" - ] - } -} diff --git a/verisimdb/connectors/clients/rescript/rescript.json b/verisimdb/connectors/clients/rescript/rescript.json deleted file mode 100644 index 33eb2087..00000000 --- a/verisimdb/connectors/clients/rescript/rescript.json +++ /dev/null @@ -1,7 +0,0 @@ -{ - "name": "verisimdb-client", - "sources": [{ "dir": "src", "subdirs": true }], - "package-specs": [{ "module": "esmodule", "in-source": true }], - "suffix": ".res.mjs", - "dependencies": [] -} diff --git a/verisimdb/connectors/clients/rescript/src/VeriSimClient.res b/verisimdb/connectors/clients/rescript/src/VeriSimClient.res deleted file mode 100644 index e5fd182b..00000000 --- a/verisimdb/connectors/clients/rescript/src/VeriSimClient.res +++ /dev/null @@ -1,162 +0,0 @@ -// SPDX-License-Identifier: PMPL-1.0-or-later -// Copyright (c) 2026 Jonathan D.A. Jewell (hyperpolymath) -// -// VeriSimDB ReScript Client — Connection configuration, authentication, and HTTP transport. -// -// This module provides the core client configuration used by all other SDK modules. -// It wraps the Fetch API via external bindings (no npm packages) and supports -// multiple authentication methods: API key, Basic, Bearer token, or none. -// -// Usage: -// let client = VeriSimClient.make(~baseUrl="http://localhost:8080") -// let healthResult = await VeriSimClient.health(client) - -// -------------------------------------------------------------------------- -// Fetch API bindings (no npm dependencies — uses browser/Deno global fetch) -// -------------------------------------------------------------------------- - -/** External binding to the global fetch function. */ -@val external fetch: (string, {..}) => promise = "fetch" - -/** Read the response body as a JSON value. */ -@send external jsonBody: VeriSimTypes.fetchResponse => promise = "json" - -/** Read the response body as a UTF-8 string. */ -@send external textBody: VeriSimTypes.fetchResponse => promise = "text" - -/** External binding to btoa for Base64 encoding (available in browser and Deno). */ -@val external btoa: string => string = "btoa" - -// -------------------------------------------------------------------------- -// Authentication types -// -------------------------------------------------------------------------- - -/** Authentication method for connecting to VeriSimDB. */ -type auth = - | ApiKey(string) - | Basic({username: string, password: string}) - | Bearer(string) - | NoAuth - -// -------------------------------------------------------------------------- -// Client configuration -// -------------------------------------------------------------------------- - -/** Client holds the connection configuration for a VeriSimDB server. */ -type t = { - baseUrl: string, - timeout: int, - auth: auth, -} - -/** Create a new unauthenticated client with the given base URL. */ -let make = (~baseUrl: string, ~timeout: int=30000, ~auth: auth=NoAuth): t => { - { - baseUrl: baseUrl, - timeout: timeout, - auth: auth, - } -} - -/** Create a client authenticated with an API key. */ -let makeWithApiKey = (~baseUrl: string, ~apiKey: string, ~timeout: int=30000): t => { - make(~baseUrl, ~timeout, ~auth=ApiKey(apiKey)) -} - -/** Create a client authenticated with a Bearer token. */ -let makeWithBearer = (~baseUrl: string, ~token: string, ~timeout: int=30000): t => { - make(~baseUrl, ~timeout, ~auth=Bearer(token)) -} - -// -------------------------------------------------------------------------- -// Internal helpers -// -------------------------------------------------------------------------- - -/** Build authentication headers from the client's auth configuration. */ -let authHeaders = (client: t): Dict.t => { - let headers = Dict.make() - switch client.auth { - | ApiKey(key) => Dict.set(headers, "X-API-Key", key) - | Basic({username, password}) => { - let encoded = btoa(`${username}:${password}`) - Dict.set(headers, "Authorization", `Basic ${encoded}`) - } - | Bearer(token) => Dict.set(headers, "Authorization", `Bearer ${token}`) - | NoAuth => () - } - headers -} - -/** Perform a GET request to the given path on the client's base URL. */ -let doGet = async (client: t, path: string): VeriSimTypes.fetchResponse => { - let headers = authHeaders(client) - await fetch( - `${client.baseUrl}${path}`, - { - "method": "GET", - "headers": headers, - }, - ) -} - -/** Perform a POST request with a JSON body. */ -let doPost = async (client: t, path: string, body: JSON.t): VeriSimTypes.fetchResponse => { - let headers = authHeaders(client) - Dict.set(headers, "Content-Type", "application/json") - await fetch( - `${client.baseUrl}${path}`, - { - "method": "POST", - "headers": headers, - "body": JSON.stringify(body), - }, - ) -} - -/** Perform a PUT request with a JSON body. */ -let doPut = async (client: t, path: string, body: JSON.t): VeriSimTypes.fetchResponse => { - let headers = authHeaders(client) - Dict.set(headers, "Content-Type", "application/json") - await fetch( - `${client.baseUrl}${path}`, - { - "method": "PUT", - "headers": headers, - "body": JSON.stringify(body), - }, - ) -} - -/** Perform a DELETE request to the given path. */ -let doDelete = async (client: t, path: string): VeriSimTypes.fetchResponse => { - let headers = authHeaders(client) - await fetch( - `${client.baseUrl}${path}`, - { - "method": "DELETE", - "headers": headers, - }, - ) -} - -// -------------------------------------------------------------------------- -// Health check -// -------------------------------------------------------------------------- - -/** Check whether the VeriSimDB server is reachable and healthy. - * - * Sends a GET request to /health and expects a 200 OK response. - * Returns a result indicating success or an error message. - */ -let health = async (client: t): result => { - try { - let resp = await doGet(client, "/health") - if resp.ok { - Ok(true) - } else { - Error(VeriSimError.fromStatus(resp.status)) - } - } catch { - | _ => Error(VeriSimError.ConnectionError("Failed to connect to VeriSimDB server")) - } -} diff --git a/verisimdb/connectors/clients/rescript/src/VeriSimDrift.res b/verisimdb/connectors/clients/rescript/src/VeriSimDrift.res deleted file mode 100644 index 1051bb55..00000000 --- a/verisimdb/connectors/clients/rescript/src/VeriSimDrift.res +++ /dev/null @@ -1,98 +0,0 @@ -// SPDX-License-Identifier: PMPL-1.0-or-later -// Copyright (c) 2026 Jonathan D.A. Jewell (hyperpolymath) -// -// VeriSimDB ReScript Client — Drift detection operations. -// -// Drift measures how much a hexad's embeddings, relationships, or content -// have diverged from a baseline state. This module provides functions to -// query drift scores, check drift status classifications, and trigger -// re-normalisation of drifted hexads. - -/// JSON boundary cast — used at the HTTP response boundary where we trust -/// the VeriSimDB server's JSON schema matches our ReScript types. -/// This replaces Obj.magic with an explicit, auditable cast point. -external fromJson: JSON.t => 'a = "%identity" -external toJson: 'a => JSON.t = "%identity" - -/** Retrieve the current drift score for a specific hexad. - * - * The drift score is a floating-point value between 0.0 (no drift) and - * 1.0 (maximum drift). - * - * @param client The authenticated client. - * @param hexadId The unique identifier of the hexad. - * @returns The drift score with component breakdown, or an error. - */ -let getScore = async ( - client: VeriSimClient.t, - hexadId: string, -): result => { - try { - let resp = await VeriSimClient.doGet(client, `/api/v1/hexads/${hexadId}/drift`) - if resp.ok { - let json = await VeriSimClient.jsonBody(resp) - Ok(json->fromJson) - } else { - Error(VeriSimError.fromStatus(resp.status)) - } - } catch { - | _ => Error(VeriSimError.ConnectionError("Failed to get drift score")) - } -} - -/** Retrieve a classified drift status report for a hexad. - * - * The report includes the drift level (Stable, Low, Moderate, High, Critical), - * the underlying score, and a human-readable explanation. - * - * @param client The authenticated client. - * @param hexadId The unique identifier of the hexad. - * @returns The drift status report, or an error. - */ -let status = async ( - client: VeriSimClient.t, - hexadId: string, -): result => { - try { - let resp = await VeriSimClient.doGet(client, `/api/v1/hexads/${hexadId}/drift/status`) - if resp.ok { - let json = await VeriSimClient.jsonBody(resp) - Ok(json->fromJson) - } else { - Error(VeriSimError.fromStatus(resp.status)) - } - } catch { - | _ => Error(VeriSimError.ConnectionError("Failed to get drift status")) - } -} - -/** Trigger re-normalisation of a drifted hexad. - * - * Normalisation recomputes the hexad's embeddings and relationship weights - * against the current baseline, effectively resetting the drift score. - * - * @param client The authenticated client. - * @param hexadId The unique identifier of the hexad. - * @returns The updated drift score after normalisation, or an error. - */ -let normalize = async ( - client: VeriSimClient.t, - hexadId: string, -): result => { - try { - let emptyBody = JSON.parseExn("{}") - let resp = await VeriSimClient.doPost( - client, - `/api/v1/hexads/${hexadId}/drift/normalize`, - emptyBody, - ) - if resp.ok { - let json = await VeriSimClient.jsonBody(resp) - Ok(json->fromJson) - } else { - Error(VeriSimError.fromStatus(resp.status)) - } - } catch { - | _ => Error(VeriSimError.ConnectionError("Failed to normalize drift")) - } -} diff --git a/verisimdb/connectors/clients/rescript/src/VeriSimError.res b/verisimdb/connectors/clients/rescript/src/VeriSimError.res deleted file mode 100644 index 5ed94e0a..00000000 --- a/verisimdb/connectors/clients/rescript/src/VeriSimError.res +++ /dev/null @@ -1,128 +0,0 @@ -// SPDX-License-Identifier: PMPL-1.0-or-later -// Copyright (c) 2026 Jonathan D.A. Jewell (hyperpolymath) -// -// VeriSimDB ReScript Client — Error types and handling. -// -// This module defines typed error variants for all failure modes that can -// occur when communicating with a VeriSimDB server. Errors are structured -// to provide actionable information: HTTP status codes, server-side error -// codes, and human-readable messages. - -/** Typed error variants for VeriSimDB client operations. - * - * Client errors (4xx), server errors (5xx), domain-specific errors, and - * client-side errors are all represented as a single variant type for - * exhaustive pattern matching. - */ -type t = - // --- Client errors (4xx) --- - | BadRequest(string) - | Unauthorized(string) - | Forbidden(string) - | NotFound(string) - | Conflict(string) - | ValidationFailed(string) - | RateLimited(string) - // --- Server errors (5xx) --- - | InternalError(string) - | ServiceUnavailable(string) - // --- Domain-specific errors --- - | HexadNotFound(string) - | ModalityUnavailable(string) - | DriftComputationError(string) - | ProvenanceInvalid(string) - | VclParseError(string) - | VclExecutionError(string) - | FederationError(string) - // --- Client-side errors --- - | ConnectionError(string) - | TimeoutError(string) - | SerializationError(string) - | UnknownError(string) - -/** Extract a human-readable message from any error variant. - * - * @param err The error variant. - * @returns A string describing the error. - */ -let message = (err: t): string => { - switch err { - | BadRequest(msg) => `Bad request: ${msg}` - | Unauthorized(msg) => `Unauthorized: ${msg}` - | Forbidden(msg) => `Forbidden: ${msg}` - | NotFound(msg) => `Not found: ${msg}` - | Conflict(msg) => `Conflict: ${msg}` - | ValidationFailed(msg) => `Validation failed: ${msg}` - | RateLimited(msg) => `Rate limited: ${msg}` - | InternalError(msg) => `Internal server error: ${msg}` - | ServiceUnavailable(msg) => `Service unavailable: ${msg}` - | HexadNotFound(msg) => `Hexad not found: ${msg}` - | ModalityUnavailable(msg) => `Modality unavailable: ${msg}` - | DriftComputationError(msg) => `Drift computation error: ${msg}` - | ProvenanceInvalid(msg) => `Provenance invalid: ${msg}` - | VclParseError(msg) => `VCL parse error: ${msg}` - | VclExecutionError(msg) => `VCL execution error: ${msg}` - | FederationError(msg) => `Federation error: ${msg}` - | ConnectionError(msg) => `Connection error: ${msg}` - | TimeoutError(msg) => `Timeout error: ${msg}` - | SerializationError(msg) => `Serialization error: ${msg}` - | UnknownError(msg) => `Unknown error: ${msg}` - } -} - -/** Construct an error variant from an HTTP status code. - * - * Maps standard HTTP status codes to the appropriate error variant. - * Used internally by other SDK modules when the server returns a non-success - * status code. - * - * @param status The HTTP status code. - * @returns The appropriate error variant with a default message. - */ -let fromStatus = (status: int): t => { - switch status { - | 400 => BadRequest("Bad request") - | 401 => Unauthorized("Authentication required") - | 403 => Forbidden("Insufficient permissions") - | 404 => NotFound("Resource not found") - | 409 => Conflict("Resource conflict") - | 422 => ValidationFailed("Input validation failed") - | 429 => RateLimited("Too many requests") - | 500 => InternalError("Internal server error") - | 503 => ServiceUnavailable("Server temporarily unavailable") - | code => UnknownError(`Unexpected HTTP status: ${Int.toString(code)}`) - } -} - -/** Check whether an error is retryable. - * - * Server errors (5xx) and rate limiting (429) are generally retryable. - * Client errors (4xx) are not, as they indicate a problem with the request. - * - * @param err The error to check. - * @returns true if the operation can be safely retried. - */ -let isRetryable = (err: t): bool => { - switch err { - | RateLimited(_) => true - | InternalError(_) => true - | ServiceUnavailable(_) => true - | ConnectionError(_) => true - | TimeoutError(_) => true - | BadRequest(_) => false - | Unauthorized(_) => false - | Forbidden(_) => false - | NotFound(_) => false - | Conflict(_) => false - | ValidationFailed(_) => false - | HexadNotFound(_) => false - | ModalityUnavailable(_) => false - | DriftComputationError(_) => false - | ProvenanceInvalid(_) => false - | VclParseError(_) => false - | VclExecutionError(_) => false - | FederationError(_) => false - | SerializationError(_) => false - | UnknownError(_) => false - } -} diff --git a/verisimdb/connectors/clients/rescript/src/VeriSimFederation.res b/verisimdb/connectors/clients/rescript/src/VeriSimFederation.res deleted file mode 100644 index 79857d1a..00000000 --- a/verisimdb/connectors/clients/rescript/src/VeriSimFederation.res +++ /dev/null @@ -1,106 +0,0 @@ -// SPDX-License-Identifier: PMPL-1.0-or-later -// Copyright (c) 2026 Jonathan D.A. Jewell (hyperpolymath) -// -// VeriSimDB ReScript Client — Federation operations. -// -// VeriSimDB supports federated operation where multiple instances form a cluster, -// sharing and synchronising hexad data across peers. This module provides functions -// to register and manage peers and to execute cross-node queries. - -/// JSON boundary cast — used at the HTTP response boundary where we trust -/// the VeriSimDB server's JSON schema matches our ReScript types. -/// This replaces Obj.magic with an explicit, auditable cast point. -external fromJson: JSON.t => 'a = "%identity" -external toJson: 'a => JSON.t = "%identity" - -/** Peer registration input. */ -type peerRegistration = { - name: string, - url: string, - metadata: Dict.t, -} - -/** Federated query request. */ -type federatedQueryRequest = { - query: string, - params: Dict.t, - peerIds: array, - timeout: int, -} - -/** Register a new VeriSimDB instance as a federation peer. - * - * @param client The authenticated client. - * @param input The peer registration details. - * @returns The registered peer with server-assigned ID, or an error. - */ -let registerPeer = async ( - client: VeriSimClient.t, - input: peerRegistration, -): result => { - try { - let body = switch JSON.stringifyAny(input) { - | Some(s) => JSON.parseExn(s) - | None => JSON.parseExn("{}") - } - let resp = await VeriSimClient.doPost(client, "/api/v1/federation/peers", body) - if resp.status == 201 { - let json = await VeriSimClient.jsonBody(resp) - Ok(json->fromJson) - } else { - Error(VeriSimError.fromStatus(resp.status)) - } - } catch { - | _ => Error(VeriSimError.ConnectionError("Failed to register peer")) - } -} - -/** Retrieve all registered federation peers. - * - * @param client The authenticated client. - * @returns A list of federation peers, or an error. - */ -let listPeers = async ( - client: VeriSimClient.t, -): result, VeriSimError.t> => { - try { - let resp = await VeriSimClient.doGet(client, "/api/v1/federation/peers") - if resp.ok { - let json = await VeriSimClient.jsonBody(resp) - Ok(json->fromJson) - } else { - Error(VeriSimError.fromStatus(resp.status)) - } - } catch { - | _ => Error(VeriSimError.ConnectionError("Failed to list peers")) - } -} - -/** Execute a VCL query across one or more federation peers. - * - * If peerIds is empty, the query is broadcast to all active peers. - * - * @param client The authenticated client. - * @param input The federated query request. - * @returns Aggregated results from all queried peers, or an error. - */ -let federatedQuery = async ( - client: VeriSimClient.t, - input: federatedQueryRequest, -): result => { - try { - let body = switch JSON.stringifyAny(input) { - | Some(s) => JSON.parseExn(s) - | None => JSON.parseExn("{}") - } - let resp = await VeriSimClient.doPost(client, "/api/v1/federation/query", body) - if resp.ok { - let json = await VeriSimClient.jsonBody(resp) - Ok(json->fromJson) - } else { - Error(VeriSimError.fromStatus(resp.status)) - } - } catch { - | _ => Error(VeriSimError.ConnectionError("Federated query failed")) - } -} diff --git a/verisimdb/connectors/clients/rescript/src/VeriSimHexad.res b/verisimdb/connectors/clients/rescript/src/VeriSimHexad.res deleted file mode 100644 index 2327f9fa..00000000 --- a/verisimdb/connectors/clients/rescript/src/VeriSimHexad.res +++ /dev/null @@ -1,147 +0,0 @@ -// SPDX-License-Identifier: PMPL-1.0-or-later -// Copyright (c) 2026 Jonathan D.A. Jewell (hyperpolymath) -// -// VeriSimDB ReScript Client — Hexad CRUD operations. -// -// This module provides create, read, update, delete, and paginated list -// operations for VeriSimDB hexad entities. All functions are async and -// communicate with the VeriSimDB REST API via VeriSimClient's HTTP helpers. - -/// JSON boundary cast — used at the HTTP response boundary where we trust -/// the VeriSimDB server's JSON schema matches our ReScript types. -/// This replaces Obj.magic with an explicit, auditable cast point. -external fromJson: JSON.t => 'a = "%identity" -external toJson: 'a => JSON.t = "%identity" - -/** Create a new hexad on the VeriSimDB server. - * - * @param client The authenticated client configuration. - * @param input The hexad input describing modalities and data. - * @returns The newly created hexad with server-assigned ID, or an error. - */ -let create = async ( - client: VeriSimClient.t, - input: VeriSimTypes.hexadInput, -): result => { - try { - let body = switch JSON.stringifyAny(input) { - | Some(s) => JSON.parseExn(s) - | None => JSON.parseExn("{}") - } - let resp = await VeriSimClient.doPost(client, "/api/v1/hexads", body) - if resp.status == 201 { - let json = await VeriSimClient.jsonBody(resp) - Ok(json->fromJson) - } else { - Error(VeriSimError.fromStatus(resp.status)) - } - } catch { - | _ => Error(VeriSimError.ConnectionError("Failed to create hexad")) - } -} - -/** Retrieve a single hexad by its unique identifier. - * - * @param client The authenticated client configuration. - * @param id The hexad's unique identifier. - * @returns The requested hexad, or an error if not found. - */ -let get = async ( - client: VeriSimClient.t, - id: string, -): result => { - try { - let resp = await VeriSimClient.doGet(client, `/api/v1/hexads/${id}`) - if resp.ok { - let json = await VeriSimClient.jsonBody(resp) - Ok(json->fromJson) - } else { - Error(VeriSimError.fromStatus(resp.status)) - } - } catch { - | _ => Error(VeriSimError.ConnectionError("Failed to get hexad")) - } -} - -/** Update an existing hexad with the given input fields. - * - * Only the fields present in the input are modified; others remain unchanged. - * - * @param client The authenticated client configuration. - * @param id The hexad's unique identifier. - * @param input The fields to update. - * @returns The updated hexad, or an error on failure. - */ -let update = async ( - client: VeriSimClient.t, - id: string, - input: VeriSimTypes.hexadInput, -): result => { - try { - let body = switch JSON.stringifyAny(input) { - | Some(s) => JSON.parseExn(s) - | None => JSON.parseExn("{}") - } - let resp = await VeriSimClient.doPut(client, `/api/v1/hexads/${id}`, body) - if resp.ok { - let json = await VeriSimClient.jsonBody(resp) - Ok(json->fromJson) - } else { - Error(VeriSimError.fromStatus(resp.status)) - } - } catch { - | _ => Error(VeriSimError.ConnectionError("Failed to update hexad")) - } -} - -/** Delete a hexad by its unique identifier. - * - * @param client The authenticated client configuration. - * @param id The hexad's unique identifier. - * @returns true if deletion succeeded, or an error on failure. - */ -let delete = async ( - client: VeriSimClient.t, - id: string, -): result => { - try { - let resp = await VeriSimClient.doDelete(client, `/api/v1/hexads/${id}`) - if resp.status == 204 || resp.status == 200 { - Ok(true) - } else { - Error(VeriSimError.fromStatus(resp.status)) - } - } catch { - | _ => Error(VeriSimError.ConnectionError("Failed to delete hexad")) - } -} - -/** Retrieve a paginated list of hexads. - * - * @param client The authenticated client configuration. - * @param page The page number (1-indexed). - * @param perPage The number of hexads per page. - * @returns A paginated response containing hexads and metadata, or an error. - */ -let list = async ( - client: VeriSimClient.t, - ~page: int=1, - ~perPage: int=20, -): result => { - try { - let pageStr = Int.toString(page) - let perPageStr = Int.toString(perPage) - let resp = await VeriSimClient.doGet( - client, - `/api/v1/hexads?page=${pageStr}&per_page=${perPageStr}`, - ) - if resp.ok { - let json = await VeriSimClient.jsonBody(resp) - Ok(json->fromJson) - } else { - Error(VeriSimError.fromStatus(resp.status)) - } - } catch { - | _ => Error(VeriSimError.ConnectionError("Failed to list hexads")) - } -} diff --git a/verisimdb/connectors/clients/rescript/src/VeriSimProvenance.res b/verisimdb/connectors/clients/rescript/src/VeriSimProvenance.res deleted file mode 100644 index 9a3dd694..00000000 --- a/verisimdb/connectors/clients/rescript/src/VeriSimProvenance.res +++ /dev/null @@ -1,104 +0,0 @@ -// SPDX-License-Identifier: PMPL-1.0-or-later -// Copyright (c) 2026 Jonathan D.A. Jewell (hyperpolymath) -// -// VeriSimDB ReScript Client — Provenance operations. -// -// Every hexad maintains an immutable provenance chain — a cryptographically -// linked sequence of events recording every mutation applied to it. This module -// provides functions to query chains, record new events, and verify integrity. - -/// JSON boundary cast — used at the HTTP response boundary where we trust -/// the VeriSimDB server's JSON schema matches our ReScript types. -/// This replaces Obj.magic with an explicit, auditable cast point. -external fromJson: JSON.t => 'a = "%identity" -external toJson: 'a => JSON.t = "%identity" - -/** Retrieve the complete provenance chain for a hexad. - * - * The chain is returned in chronological order (oldest first) and includes - * the verification status. - * - * @param client The authenticated client. - * @param hexadId The unique identifier of the hexad. - * @returns The provenance chain with all events, or an error. - */ -let getChain = async ( - client: VeriSimClient.t, - hexadId: string, -): result => { - try { - let resp = await VeriSimClient.doGet(client, `/api/v1/hexads/${hexadId}/provenance`) - if resp.ok { - let json = await VeriSimClient.jsonBody(resp) - Ok(json->fromJson) - } else { - Error(VeriSimError.fromStatus(resp.status)) - } - } catch { - | _ => Error(VeriSimError.ConnectionError("Failed to get provenance chain")) - } -} - -/** Record a new provenance event on a hexad's chain. - * - * The event is cryptographically linked to the previous event in the chain. - * The server assigns the event ID and timestamp. - * - * @param client The authenticated client. - * @param hexadId The unique identifier of the hexad. - * @param input The event details to record. - * @returns The newly created provenance event, or an error. - */ -let recordEvent = async ( - client: VeriSimClient.t, - hexadId: string, - input: VeriSimTypes.provenanceEventInput, -): result => { - try { - let body = switch JSON.stringifyAny(input) { - | Some(s) => JSON.parseExn(s) - | None => JSON.parseExn("{}") - } - let resp = await VeriSimClient.doPost(client, `/api/v1/hexads/${hexadId}/provenance`, body) - if resp.status == 201 { - let json = await VeriSimClient.jsonBody(resp) - Ok(json->fromJson) - } else { - Error(VeriSimError.fromStatus(resp.status)) - } - } catch { - | _ => Error(VeriSimError.ConnectionError("Failed to record provenance event")) - } -} - -/** Verify the cryptographic integrity of a hexad's provenance chain. - * - * The server traverses the entire chain, checking each event's hash link. - * Returns true if the chain is intact, false if tampering is detected. - * - * @param client The authenticated client. - * @param hexadId The unique identifier of the hexad. - * @returns true if the chain is verified intact, or an error. - */ -let verify = async ( - client: VeriSimClient.t, - hexadId: string, -): result => { - try { - let emptyBody = JSON.parseExn("{}") - let resp = await VeriSimClient.doPost( - client, - `/api/v1/hexads/${hexadId}/provenance/verify`, - emptyBody, - ) - if resp.ok { - let json = await VeriSimClient.jsonBody(resp) - let chain: VeriSimTypes.provenanceChain = json->fromJson - Ok(chain.verified) - } else { - Error(VeriSimError.fromStatus(resp.status)) - } - } catch { - | _ => Error(VeriSimError.ConnectionError("Failed to verify provenance")) - } -} diff --git a/verisimdb/connectors/clients/rescript/src/VeriSimSearch.res b/verisimdb/connectors/clients/rescript/src/VeriSimSearch.res deleted file mode 100644 index 2e3bd849..00000000 --- a/verisimdb/connectors/clients/rescript/src/VeriSimSearch.res +++ /dev/null @@ -1,232 +0,0 @@ -// SPDX-License-Identifier: PMPL-1.0-or-later -// Copyright (c) 2026 Jonathan D.A. Jewell (hyperpolymath) -// -// VeriSimDB ReScript Client — Search operations. -// -// This module provides multi-modal search capabilities against VeriSimDB, -// including full-text search, vector similarity search, spatial radius and -// bounding-box queries, nearest-neighbour lookups, and relationship traversal. - -/// JSON boundary cast — used at the HTTP response boundary where we trust -/// the VeriSimDB server's JSON schema matches our ReScript types. -/// This replaces Obj.magic with an explicit, auditable cast point. -external fromJson: JSON.t => 'a = "%identity" -external toJson: 'a => JSON.t = "%identity" - -// -------------------------------------------------------------------------- -// Search parameter types -// -------------------------------------------------------------------------- - -/** Parameters for a full-text search query. */ -type textSearchParams = { - query: string, - modalities: array, - limit: int, - offset: int, -} - -/** Parameters for a vector similarity search. */ -type vectorSearchParams = { - vector: array, - model: string, - topK: int, - threshold: float, -} - -/** Parameters for a spatial radius search (point + distance). */ -type spatialRadiusParams = { - latitude: float, - longitude: float, - radiusKm: float, - limit: int, -} - -/** Parameters for a spatial bounding-box search. */ -type spatialBoundsParams = { - minLat: float, - minLon: float, - maxLat: float, - maxLon: float, - limit: int, -} - -/** Parameters for a nearest-neighbour search by hexad ID. */ -type nearestParams = { - hexadId: string, - topK: int, - modality: VeriSimTypes.modality, -} - -/** Parameters for a relationship traversal search. */ -type relatedParams = { - hexadId: string, - relType: option, - depth: int, - limit: int, -} - -// -------------------------------------------------------------------------- -// Search functions -// -------------------------------------------------------------------------- - -/** Perform a full-text search across hexad content. - * - * @param client The authenticated client. - * @param params Text search parameters including query string and filters. - * @returns A list of search results ranked by relevance, or an error. - */ -let text = async ( - client: VeriSimClient.t, - params: textSearchParams, -): result, VeriSimError.t> => { - try { - let body = switch JSON.stringifyAny(params) { - | Some(s) => JSON.parseExn(s) - | None => JSON.parseExn("{}") - } - let resp = await VeriSimClient.doPost(client, "/api/v1/search/text", body) - if resp.ok { - let json = await VeriSimClient.jsonBody(resp) - Ok(json->fromJson) - } else { - Error(VeriSimError.fromStatus(resp.status)) - } - } catch { - | _ => Error(VeriSimError.ConnectionError("Text search failed")) - } -} - -/** Perform a vector similarity search using a query embedding. - * - * @param client The authenticated client. - * @param params Vector search parameters including the query vector. - * @returns A list of search results ranked by cosine similarity, or an error. - */ -let vector = async ( - client: VeriSimClient.t, - params: vectorSearchParams, -): result, VeriSimError.t> => { - try { - let body = switch JSON.stringifyAny(params) { - | Some(s) => JSON.parseExn(s) - | None => JSON.parseExn("{}") - } - let resp = await VeriSimClient.doPost(client, "/api/v1/search/vector", body) - if resp.ok { - let json = await VeriSimClient.jsonBody(resp) - Ok(json->fromJson) - } else { - Error(VeriSimError.fromStatus(resp.status)) - } - } catch { - | _ => Error(VeriSimError.ConnectionError("Vector search failed")) - } -} - -/** Find hexads within a given radius of a geographic point. - * - * @param client The authenticated client. - * @param params Latitude, longitude, and radius in kilometres. - * @returns A list of search results within the radius, or an error. - */ -let spatialRadius = async ( - client: VeriSimClient.t, - params: spatialRadiusParams, -): result, VeriSimError.t> => { - try { - let body = switch JSON.stringifyAny(params) { - | Some(s) => JSON.parseExn(s) - | None => JSON.parseExn("{}") - } - let resp = await VeriSimClient.doPost(client, "/api/v1/search/spatial/radius", body) - if resp.ok { - let json = await VeriSimClient.jsonBody(resp) - Ok(json->fromJson) - } else { - Error(VeriSimError.fromStatus(resp.status)) - } - } catch { - | _ => Error(VeriSimError.ConnectionError("Spatial radius search failed")) - } -} - -/** Find hexads within a rectangular bounding box. - * - * @param client The authenticated client. - * @param params The bounding box defined by min/max latitude and longitude. - * @returns A list of search results within the bounds, or an error. - */ -let spatialBounds = async ( - client: VeriSimClient.t, - params: spatialBoundsParams, -): result, VeriSimError.t> => { - try { - let body = switch JSON.stringifyAny(params) { - | Some(s) => JSON.parseExn(s) - | None => JSON.parseExn("{}") - } - let resp = await VeriSimClient.doPost(client, "/api/v1/search/spatial/bounds", body) - if resp.ok { - let json = await VeriSimClient.jsonBody(resp) - Ok(json->fromJson) - } else { - Error(VeriSimError.fromStatus(resp.status)) - } - } catch { - | _ => Error(VeriSimError.ConnectionError("Spatial bounds search failed")) - } -} - -/** Find the nearest neighbours of a given hexad. - * - * @param client The authenticated client. - * @param params The hexad ID and number of neighbours to return. - * @returns A list of search results ordered by proximity, or an error. - */ -let nearest = async ( - client: VeriSimClient.t, - params: nearestParams, -): result, VeriSimError.t> => { - try { - let body = switch JSON.stringifyAny(params) { - | Some(s) => JSON.parseExn(s) - | None => JSON.parseExn("{}") - } - let resp = await VeriSimClient.doPost(client, "/api/v1/search/nearest", body) - if resp.ok { - let json = await VeriSimClient.jsonBody(resp) - Ok(json->fromJson) - } else { - Error(VeriSimError.fromStatus(resp.status)) - } - } catch { - | _ => Error(VeriSimError.ConnectionError("Nearest search failed")) - } -} - -/** Traverse relationships from a given hexad. - * - * @param client The authenticated client. - * @param params The source hexad ID, optional relationship type filter, and depth. - * @returns A list of search results connected by relationships, or an error. - */ -let related = async ( - client: VeriSimClient.t, - params: relatedParams, -): result, VeriSimError.t> => { - try { - let body = switch JSON.stringifyAny(params) { - | Some(s) => JSON.parseExn(s) - | None => JSON.parseExn("{}") - } - let resp = await VeriSimClient.doPost(client, "/api/v1/search/related", body) - if resp.ok { - let json = await VeriSimClient.jsonBody(resp) - Ok(json->fromJson) - } else { - Error(VeriSimError.fromStatus(resp.status)) - } - } catch { - | _ => Error(VeriSimError.ConnectionError("Related search failed")) - } -} diff --git a/verisimdb/connectors/clients/rescript/src/VeriSimTypes.res b/verisimdb/connectors/clients/rescript/src/VeriSimTypes.res deleted file mode 100644 index ce5a8115..00000000 --- a/verisimdb/connectors/clients/rescript/src/VeriSimTypes.res +++ /dev/null @@ -1,324 +0,0 @@ -// SPDX-License-Identifier: PMPL-1.0-or-later -// Copyright (c) 2026 Jonathan D.A. Jewell (hyperpolymath) -// -// VeriSimDB ReScript Client — Core type definitions. -// -// This module defines all data structures exchanged between the ReScript client -// SDK and the VeriSimDB server. Types are defined as ReScript records and -// variants, designed for direct JSON serialization via the JSON module. -// -// The central entity is the Hexad — a six-faceted data object unifying graph, -// vector, tensor, semantic, document, temporal, provenance, and spatial modalities. - -// -------------------------------------------------------------------------- -// Fetch API response type (used by VeriSimClient) -// -------------------------------------------------------------------------- - -/** Minimal representation of the Fetch API Response object. */ -type fetchResponse = { - ok: bool, - status: int, - statusText: string, -} - -// -------------------------------------------------------------------------- -// Modality -// -------------------------------------------------------------------------- - -/** The eight data modalities supported by VeriSimDB hexads. */ -type modality = - | Graph - | Vector - | Tensor - | Semantic - | Document - | Temporal - | Provenance - | Spatial - -/** Convert a modality variant to its JSON string representation. */ -let modalityToString = (m: modality): string => { - switch m { - | Graph => "graph" - | Vector => "vector" - | Tensor => "tensor" - | Semantic => "semantic" - | Document => "document" - | Temporal => "temporal" - | Provenance => "provenance" - | Spatial => "spatial" - } -} - -/** Parse a modality from its JSON string representation. */ -let modalityFromString = (s: string): option => { - switch s { - | "graph" => Some(Graph) - | "vector" => Some(Vector) - | "tensor" => Some(Tensor) - | "semantic" => Some(Semantic) - | "document" => Some(Document) - | "temporal" => Some(Temporal) - | "provenance" => Some(Provenance) - | "spatial" => Some(Spatial) - | _ => None - } -} - -// -------------------------------------------------------------------------- -// Modality status -// -------------------------------------------------------------------------- - -/** Which modalities are active on a given hexad. */ -type modalityStatus = { - graph: bool, - vector: bool, - tensor: bool, - semantic: bool, - document: bool, - temporal: bool, - provenance: bool, - spatial: bool, -} - -// -------------------------------------------------------------------------- -// Hexad status -// -------------------------------------------------------------------------- - -/** The lifecycle state of a hexad. */ -type hexadStatus = - | Active - | Archived - | Draft - | Deleted - -// -------------------------------------------------------------------------- -// Graph modality data -// -------------------------------------------------------------------------- - -/** A directed edge between two hexads in the graph modality. */ -type graphEdge = { - source: string, - target: string, - relType: string, - weight: float, - metadata: Dict.t, -} - -/** Graph-modality data for a hexad: edges and node properties. */ -type graphData = { - edges: array, - properties: Dict.t, -} - -// -------------------------------------------------------------------------- -// Vector modality data -// -------------------------------------------------------------------------- - -/** Embedding vector data for vector-modality operations. */ -type vectorData = { - embedding: array, - model: string, - dimensions: int, -} - -// -------------------------------------------------------------------------- -// Tensor modality data -// -------------------------------------------------------------------------- - -/** Multi-dimensional tensor data reference. */ -type tensorData = { - shape: array, - dtype: string, - dataRef: string, -} - -// -------------------------------------------------------------------------- -// Document modality data -// -------------------------------------------------------------------------- - -/** Document-modality content: text, format, and language metadata. */ -type documentContent = { - text: string, - format: string, - language: string, - metadata: Dict.t, -} - -// -------------------------------------------------------------------------- -// Spatial modality data -// -------------------------------------------------------------------------- - -/** Spatial-modality coordinates and geometry. */ -type spatialData = { - latitude: float, - longitude: float, - altitude: option, - geometry: option, - crs: string, -} - -// -------------------------------------------------------------------------- -// Hexad (core entity) -// -------------------------------------------------------------------------- - -/** The core entity in VeriSimDB — a multi-modal data object. */ -type hexad = { - id: string, - status: hexadStatus, - modalities: modalityStatus, - createdAt: string, - updatedAt: string, - metadata: Dict.t, - graphData: option, - vectorData: option, - tensorData: option, - content: option, - spatialData: option, -} - -// -------------------------------------------------------------------------- -// Hexad input (for create/update) -// -------------------------------------------------------------------------- - -/** Input structure for creating or updating a hexad. */ -type hexadInput = { - graphData: option, - vectorData: option, - tensorData: option, - content: option, - spatialData: option, - metadata: Dict.t, - modalities: array, -} - -// -------------------------------------------------------------------------- -// Drift types -// -------------------------------------------------------------------------- - -/** Drift score measurement for a hexad. Score ranges from 0.0 to 1.0. */ -type driftScore = { - hexadId: string, - score: float, - components: Dict.t, - measuredAt: string, - baselineAt: string, -} - -/** Drift level classification. */ -type driftLevel = - | Stable - | Low - | Moderate - | High - | Critical - -/** Drift status report with classification and score. */ -type driftStatusReport = { - hexadId: string, - level: driftLevel, - score: driftScore, - message: string, -} - -// -------------------------------------------------------------------------- -// Provenance types -// -------------------------------------------------------------------------- - -/** A single event in a hexad's provenance chain. */ -type provenanceEvent = { - eventId: string, - hexadId: string, - eventType: string, - actor: string, - timestamp: string, - details: Dict.t, - parentId: option, -} - -/** The complete provenance chain for a hexad. */ -type provenanceChain = { - hexadId: string, - events: array, - verified: bool, -} - -/** Input for recording a new provenance event. */ -type provenanceEventInput = { - eventType: string, - actor: string, - details: Dict.t, -} - -// -------------------------------------------------------------------------- -// Pagination -// -------------------------------------------------------------------------- - -/** Paginated response wrapping a list of hexads. */ -type paginatedResponse = { - items: array, - total: int, - page: int, - perPage: int, - totalPages: int, -} - -// -------------------------------------------------------------------------- -// Search types -// -------------------------------------------------------------------------- - -/** A search result pairing a hexad with a relevance score. */ -type searchResult = { - hexad: hexad, - score: float, -} - -// -------------------------------------------------------------------------- -// VCL types -// -------------------------------------------------------------------------- - -/** Result of a VCL query execution. */ -type vclResult = { - columns: array, - rows: array>, - count: int, - elapsedMs: float, -} - -/** Query execution plan explanation for a VCL statement. */ -type vclExplanation = { - query: string, - plan: string, - cost: float, - warnings: array, -} - -// -------------------------------------------------------------------------- -// Federation types -// -------------------------------------------------------------------------- - -/** A remote VeriSimDB node in a federated cluster. */ -type federationPeer = { - peerId: string, - name: string, - url: string, - status: string, - lastSeen: string, - metadata: Dict.t, -} - -/** Result from a single peer in a federated query. */ -type peerQueryResult = { - peerId: string, - peerName: string, - result: vclResult, - elapsedMs: float, - error: option, -} - -/** Aggregated result from a federated query across multiple peers. */ -type federatedQueryResult = { - results: array, - total: int, - elapsedMs: float, -} diff --git a/verisimdb/connectors/clients/rescript/src/VeriSimVcl.res b/verisimdb/connectors/clients/rescript/src/VeriSimVcl.res deleted file mode 100644 index 78018f78..00000000 --- a/verisimdb/connectors/clients/rescript/src/VeriSimVcl.res +++ /dev/null @@ -1,80 +0,0 @@ -// SPDX-License-Identifier: PMPL-1.0-or-later -// Copyright (c) 2026 Jonathan D.A. Jewell (hyperpolymath) -// -// VeriSimDB ReScript Client — VCL (VeriSimDB Query Language) operations. -// -// VCL is VeriSimDB's native query language for multi-modal queries that span -// graph traversals, vector similarity, spatial filters, and temporal constraints -// in a single statement. This module provides execution and explain functions. - -/// JSON boundary cast — used at the HTTP response boundary where we trust -/// the VeriSimDB server's JSON schema matches our ReScript types. -/// This replaces Obj.magic with an explicit, auditable cast point. -external fromJson: JSON.t => 'a = "%identity" -external toJson: 'a => JSON.t = "%identity" - -/** VCL request payload for executing or explaining a query. */ -type vclRequest = { - query: string, - params: Dict.t, -} - -/** Execute a VCL query and return the result set. - * - * @param client The authenticated client. - * @param query The VCL query string. - * @param params Optional named parameters for parameterised queries. - * @returns The query result with columns, rows, and timing, or an error. - */ -let execute = async ( - client: VeriSimClient.t, - query: string, - ~params: Dict.t=Dict.make(), -): result => { - try { - let req: vclRequest = {query, params} - let body = switch JSON.stringifyAny(req) { - | Some(s) => JSON.parseExn(s) - | None => JSON.parseExn("{}") - } - let resp = await VeriSimClient.doPost(client, "/api/v1/vcl/execute", body) - if resp.ok { - let json = await VeriSimClient.jsonBody(resp) - Ok(json->fromJson) - } else { - Error(VeriSimError.fromStatus(resp.status)) - } - } catch { - | _ => Error(VeriSimError.ConnectionError("VCL execution failed")) - } -} - -/** Explain a VCL query's execution plan without running it. - * - * @param client The authenticated client. - * @param query The VCL query string. - * @param params Optional named parameters. - * @returns The query plan, estimated cost, and any warnings, or an error. - */ -let explain = async ( - client: VeriSimClient.t, - query: string, - ~params: Dict.t=Dict.make(), -): result => { - try { - let req: vclRequest = {query, params} - let body = switch JSON.stringifyAny(req) { - | Some(s) => JSON.parseExn(s) - | None => JSON.parseExn("{}") - } - let resp = await VeriSimClient.doPost(client, "/api/v1/vcl/explain", body) - if resp.ok { - let json = await VeriSimClient.jsonBody(resp) - Ok(json->fromJson) - } else { - Error(VeriSimError.fromStatus(resp.status)) - } - } catch { - | _ => Error(VeriSimError.ConnectionError("VCL explain failed")) - } -} diff --git a/verisimdb/connectors/clients/rust/Cargo.toml b/verisimdb/connectors/clients/rust/Cargo.toml deleted file mode 100644 index 2e6866c4..00000000 --- a/verisimdb/connectors/clients/rust/Cargo.toml +++ /dev/null @@ -1,24 +0,0 @@ -# SPDX-License-Identifier: MPL-2.0 -# Copyright (c) 2026 Jonathan D.A. Jewell (hyperpolymath) - -[package] -name = "verisimdb-client" -version = "0.1.0" -edition = "2021" -authors = ["Jonathan D.A. Jewell "] -license = "PMPL-1.0-or-later" -description = "VeriSimDB client SDK for Rust — octad entity management, drift detection, and federation" -repository = "https://gitlab.com/hyperpolymath/verisimdb" - -[dependencies] -reqwest = { version = "0.12", features = ["json"] } -serde = { version = "1", features = ["derive"] } -serde_json = "1" -tokio = { version = "1", features = ["full"] } -thiserror = "2" -url = "2" -uuid = { version = "1", features = ["v4"] } -chrono = { version = "0.4", features = ["serde"] } - -[dev-dependencies] -tokio-test = "0.4" diff --git a/verisimdb/connectors/clients/rust/src/client.rs b/verisimdb/connectors/clients/rust/src/client.rs deleted file mode 100644 index 07d6afba..00000000 --- a/verisimdb/connectors/clients/rust/src/client.rs +++ /dev/null @@ -1,272 +0,0 @@ -// SPDX-License-Identifier: MPL-2.0 -// Copyright (c) 2026 Jonathan D.A. Jewell (hyperpolymath) - -//! VeriSimDB client configuration, authentication, and HTTP transport layer. -//! -//! [`VeriSimClient`] is the primary entry point for all SDK operations. It owns -//! the base URL, HTTP client, authentication credentials, and timeout settings. -//! Domain-specific methods (octad CRUD, search, drift, etc.) are defined as -//! `impl VeriSimClient` blocks in their respective modules. - -use std::time::Duration; - -use reqwest::header::{HeaderMap, HeaderValue, AUTHORIZATION, CONTENT_TYPE}; -use serde::de::DeserializeOwned; -use serde::Serialize; -use url::Url; - -use crate::error::{Result, VeriSimError}; -use crate::types::ErrorResponse; - -// --------------------------------------------------------------------------- -// Auth -// --------------------------------------------------------------------------- - -/// Authentication method for connecting to a VeriSimDB instance. -#[derive(Debug, Clone)] -pub enum Auth { - /// No authentication (local development, trusted networks). - None, - /// API key passed via the `X-API-Key` header. - ApiKey(String), - /// Bearer token passed via the `Authorization: Bearer ` header. - Bearer(String), - /// HTTP Basic authentication. - Basic { - /// Username. - username: String, - /// Password. - password: String, - }, -} - -// --------------------------------------------------------------------------- -// VeriSimClient -// --------------------------------------------------------------------------- - -/// The main VeriSimDB client. -/// -/// Holds connection parameters and provides low-level HTTP helpers that the -/// higher-level module methods (`octad`, `search`, `drift`, etc.) delegate to. -/// -/// # Examples -/// -/// ```rust,no_run -/// use verisimdb_client::client::VeriSimClient; -/// -/// # #[tokio::main] -/// # async fn main() -> verisimdb_client::error::Result<()> { -/// let client = VeriSimClient::new("http://localhost:8080")?; -/// assert!(client.health().await?); -/// # Ok(()) -/// # } -/// ``` -pub struct VeriSimClient { - /// Parsed base URL of the VeriSimDB instance (e.g. `http://localhost:8080`). - base_url: Url, - /// Underlying `reqwest` HTTP client (connection-pooled, TLS-capable). - http: reqwest::Client, - /// Authentication credentials. - auth: Auth, - /// Per-request timeout. - timeout: Duration, -} - -impl VeriSimClient { - // -- Constructors ------------------------------------------------------- - - /// Create a new unauthenticated client pointing at `base_url`. - /// - /// # Errors - /// - /// Returns [`VeriSimError::Validation`] if `base_url` cannot be parsed. - pub fn new(base_url: &str) -> Result { - Self::build(base_url, Auth::None) - } - - /// Create a client that authenticates via an API key header. - pub fn with_api_key(base_url: &str, key: &str) -> Result { - Self::build(base_url, Auth::ApiKey(key.to_owned())) - } - - /// Create a client that authenticates via a bearer token. - pub fn with_bearer(base_url: &str, token: &str) -> Result { - Self::build(base_url, Auth::Bearer(token.to_owned())) - } - - /// Create a client that authenticates via HTTP Basic credentials. - pub fn with_basic(base_url: &str, username: &str, password: &str) -> Result { - Self::build( - base_url, - Auth::Basic { - username: username.to_owned(), - password: password.to_owned(), - }, - ) - } - - /// Internal builder shared by all constructors. - fn build(base_url: &str, auth: Auth) -> Result { - let base_url = Url::parse(base_url) - .map_err(|e| VeriSimError::Validation(format!("Invalid base URL: {e}")))?; - - let timeout = Duration::from_secs(30); - - let http = reqwest::Client::builder() - .timeout(timeout) - .build() - .map_err(VeriSimError::Network)?; - - Ok(Self { - base_url, - http, - auth, - timeout, - }) - } - - // -- Health check ------------------------------------------------------- - - /// Ping the VeriSimDB health endpoint. - /// - /// Returns `true` if the server is reachable and reports healthy. - pub async fn health(&self) -> Result { - let url = self.url("/health"); - let response = self.apply_auth(self.http.get(url)).send().await?; - Ok(response.status().is_success()) - } - - // -- Public timeout accessor -------------------------------------------- - - /// Return the configured per-request timeout. - pub fn timeout(&self) -> Duration { - self.timeout - } - - /// Set a custom per-request timeout. - pub fn set_timeout(&mut self, timeout: Duration) { - self.timeout = timeout; - } - - // -- Internal HTTP helpers ---------------------------------------------- - - /// Build a full URL by joining `path` onto the base URL. - pub(crate) fn url(&self, path: &str) -> Url { - // Unwrap is safe: path is always a well-formed relative segment. - self.base_url.join(path).expect("valid path join") - } - - /// Attach authentication headers to an outgoing request builder. - pub(crate) fn apply_auth( - &self, - builder: reqwest::RequestBuilder, - ) -> reqwest::RequestBuilder { - match &self.auth { - Auth::None => builder, - Auth::ApiKey(key) => builder.header("X-API-Key", key.as_str()), - Auth::Bearer(token) => { - let value = format!("Bearer {token}"); - builder.header(AUTHORIZATION, value) - } - Auth::Basic { username, password } => { - builder.basic_auth(username, Some(password)) - } - } - } - - /// Perform a GET request and deserialize the JSON response body. - pub(crate) async fn get(&self, path: &str) -> Result { - let url = self.url(path); - let response = self - .apply_auth(self.http.get(url)) - .send() - .await - .map_err(VeriSimError::Network)?; - - self.handle_response(response).await - } - - /// Perform a POST request with a JSON body and deserialize the response. - pub(crate) async fn post( - &self, - path: &str, - body: &B, - ) -> Result { - let url = self.url(path); - let response = self - .apply_auth(self.http.post(url)) - .json(body) - .send() - .await - .map_err(VeriSimError::Network)?; - - self.handle_response(response).await - } - - /// Perform a PUT request with a JSON body and deserialize the response. - pub(crate) async fn put( - &self, - path: &str, - body: &B, - ) -> Result { - let url = self.url(path); - let response = self - .apply_auth(self.http.put(url)) - .json(body) - .send() - .await - .map_err(VeriSimError::Network)?; - - self.handle_response(response).await - } - - /// Perform a DELETE request. Returns `()` on success. - pub(crate) async fn delete(&self, path: &str) -> Result<()> { - let url = self.url(path); - let response = self - .apply_auth(self.http.delete(url)) - .send() - .await - .map_err(VeriSimError::Network)?; - - if response.status().is_success() { - Ok(()) - } else { - Err(self.extract_error(response).await) - } - } - - // -- Response handling -------------------------------------------------- - - /// Deserialize a successful response or extract an error from the body. - async fn handle_response( - &self, - response: reqwest::Response, - ) -> Result { - let status = response.status(); - - if status.is_success() { - let body = response.text().await.map_err(VeriSimError::Network)?; - serde_json::from_str(&body).map_err(VeriSimError::Serialization) - } else { - Err(self.extract_error(response).await) - } - } - - /// Turn a non-2xx response into the appropriate [`VeriSimError`] variant. - async fn extract_error(&self, response: reqwest::Response) -> VeriSimError { - let status = response.status().as_u16(); - - // Attempt to parse a structured error body. - let message = match response.json::().await { - Ok(err_body) => err_body.message, - Err(_) => format!("HTTP {status}"), - }; - - match status { - 404 => VeriSimError::NotFound(message), - 401 | 403 => VeriSimError::Unauthorized(message), - _ => VeriSimError::Server { status, message }, - } - } -} diff --git a/verisimdb/connectors/clients/rust/src/drift.rs b/verisimdb/connectors/clients/rust/src/drift.rs deleted file mode 100644 index a200c4b1..00000000 --- a/verisimdb/connectors/clients/rust/src/drift.rs +++ /dev/null @@ -1,63 +0,0 @@ -// SPDX-License-Identifier: MPL-2.0 -// Copyright (c) 2026 Jonathan D.A. Jewell (hyperpolymath) - -//! Drift detection and normalization operations. -//! -//! VeriSimDB continuously monitors how far each octad's modality data has -//! diverged from its normalised baseline. When drift exceeds a configurable -//! threshold the entity is flagged for re-normalisation. This module exposes -//! drift score retrieval, system-wide status, and manual normalization triggers. - -use crate::client::VeriSimClient; -use crate::error::Result; -use crate::types::DriftScore; - -impl VeriSimClient { - /// Retrieve the drift score for a single octad entity. - /// - /// The score aggregates per-modality drift metrics into an overall value - /// (0.0 = perfectly normalised, higher = more drift). - /// - /// # Arguments - /// - /// * `id` — The octad entity identifier. - /// - /// # Errors - /// - /// Returns [`VeriSimError::NotFound`] if the entity does not exist. - pub async fn get_drift_score(&self, id: &str) -> Result { - let path = format!("/api/v1/drift/{id}"); - self.get(&path).await - } - - /// Retrieve system-wide drift status. - /// - /// Returns a JSON object summarising total entities monitored, number - /// exceeding drift threshold, average drift, and last sweep timestamp. - pub async fn drift_status(&self) -> Result { - self.get("/api/v1/drift/status").await - } - - /// Trigger re-normalisation for a specific octad entity. - /// - /// This enqueues the entity for the normaliser pipeline, which will - /// recompute cross-modality consistency and update the baseline. - /// - /// # Arguments - /// - /// * `id` — The octad entity identifier. - pub async fn trigger_normalization(&self, id: &str) -> Result<()> { - let path = format!("/api/v1/drift/{id}/normalize"); - let empty: serde_json::Value = serde_json::json!({}); - let _: serde_json::Value = self.post(&path, &empty).await?; - Ok(()) - } - - /// Retrieve the normaliser pipeline status. - /// - /// Returns a JSON object with queue depth, active workers, throughput - /// metrics, and last error (if any). - pub async fn normalizer_status(&self) -> Result { - self.get("/api/v1/drift/normalizer/status").await - } -} diff --git a/verisimdb/connectors/clients/rust/src/error.rs b/verisimdb/connectors/clients/rust/src/error.rs deleted file mode 100644 index 40bea4fa..00000000 --- a/verisimdb/connectors/clients/rust/src/error.rs +++ /dev/null @@ -1,54 +0,0 @@ -// SPDX-License-Identifier: MPL-2.0 -// Copyright (c) 2026 Jonathan D.A. Jewell (hyperpolymath) - -//! Error types for the VeriSimDB client SDK. -//! -//! All fallible operations in this crate return [`Result`], which is an alias -//! for `std::result::Result`. The [`VeriSimError`] enum covers -//! network failures, serialization issues, server-side errors, validation -//! problems, and timeouts. - -use thiserror::Error; - -/// Comprehensive error type for VeriSimDB client operations. -/// -/// Each variant carries enough context for callers to decide whether to retry, -/// surface a user-facing message, or escalate. -#[derive(Error, Debug)] -pub enum VeriSimError { - /// The requested entity (octad, peer, provenance record) was not found. - #[error("Entity not found: {0}")] - NotFound(String), - - /// Authentication or authorization failed. Check API key / bearer token. - #[error("Unauthorized: {0}")] - Unauthorized(String), - - /// An underlying HTTP / network transport error from `reqwest`. - #[error("Network error: {0}")] - Network(#[from] reqwest::Error), - - /// JSON serialization or deserialization failed. - #[error("Serialization error: {0}")] - Serialization(#[from] serde_json::Error), - - /// The server returned an HTTP error status with a message body. - #[error("Server error ({status}): {message}")] - Server { - /// HTTP status code (e.g. 500, 502, 503). - status: u16, - /// Human-readable error message from the server response body. - message: String, - }, - - /// Client-side validation failed before the request was sent. - #[error("Validation error: {0}")] - Validation(String), - - /// The request exceeded the configured timeout duration. - #[error("Timeout after {0}ms")] - Timeout(u64), -} - -/// Crate-level result alias using [`VeriSimError`]. -pub type Result = std::result::Result; diff --git a/verisimdb/connectors/clients/rust/src/federation.rs b/verisimdb/connectors/clients/rust/src/federation.rs deleted file mode 100644 index 684402b8..00000000 --- a/verisimdb/connectors/clients/rust/src/federation.rs +++ /dev/null @@ -1,97 +0,0 @@ -// SPDX-License-Identifier: MPL-2.0 -// Copyright (c) 2026 Jonathan D.A. Jewell (hyperpolymath) - -//! Federation operations for cross-instance VeriSimDB queries. -//! -//! VeriSimDB supports a federated architecture where multiple instances can be -//! registered as peers. Federated queries fan out to all (or selected) peers -//! and merge results transparently. This module provides peer management and -//! federated query execution. - -use serde::Serialize; - -use crate::client::VeriSimClient; -use crate::error::Result; -use crate::types::{FederationResult, Modality}; - -/// Internal request body for registering a federation peer. -#[derive(Debug, Serialize)] -struct RegisterPeerRequest { - store_id: String, - endpoint: String, - adapter_type: String, - config: serde_json::Value, -} - -/// Internal request body for executing a federated query. -#[derive(Debug, Serialize)] -struct FederatedQueryRequest { - modalities: Vec, - params: serde_json::Value, -} - -impl VeriSimClient { - /// Register a remote VeriSimDB instance (or compatible adapter) as a - /// federation peer. - /// - /// # Arguments - /// - /// * `store_id` — Unique logical name for the peer (e.g. "us-west-replica"). - /// * `endpoint` — Base URL of the peer's API (e.g. "https://peer.example.com:8080"). - /// * `adapter_type` — Adapter kind: "verisimdb", "quandledb", "lithoglyph", or custom. - /// * `config` — Adapter-specific configuration (auth tokens, timeouts, etc.). - /// - /// # Returns - /// - /// A JSON object representing the registered peer record. - pub async fn register_peer( - &self, - store_id: &str, - endpoint: &str, - adapter_type: &str, - config: serde_json::Value, - ) -> Result { - let body = RegisterPeerRequest { - store_id: store_id.to_owned(), - endpoint: endpoint.to_owned(), - adapter_type: adapter_type.to_owned(), - config, - }; - self.post("/api/v1/federation/peers", &body).await - } - - /// List all registered federation peers. - /// - /// Returns a vector of peer records, each containing the store ID, endpoint, - /// adapter type, health status, and last-seen timestamp. - pub async fn list_peers(&self) -> Result> { - self.get("/api/v1/federation/peers").await - } - - /// Execute a federated query that fans out to all registered peers. - /// - /// The query targets the specified modalities and passes `params` to each - /// peer's local query engine. Results are merged and returned with per-peer - /// attribution. - /// - /// # Arguments - /// - /// * `modalities` — Which modalities to query across peers. - /// * `params` — Query parameters (modality-specific filters, limits, etc.). - /// - /// # Returns - /// - /// A vector of [`FederationResult`] items, one per matched entity across - /// all responding peers. - pub async fn federated_query( - &self, - modalities: &[Modality], - params: serde_json::Value, - ) -> Result> { - let body = FederatedQueryRequest { - modalities: modalities.to_vec(), - params, - }; - self.post("/api/v1/federation/query", &body).await - } -} diff --git a/verisimdb/connectors/clients/rust/src/lib.rs b/verisimdb/connectors/clients/rust/src/lib.rs deleted file mode 100644 index 8adbcaea..00000000 --- a/verisimdb/connectors/clients/rust/src/lib.rs +++ /dev/null @@ -1,46 +0,0 @@ -// SPDX-License-Identifier: MPL-2.0 -// Copyright (c) 2026 Jonathan D.A. Jewell (hyperpolymath) - -//! # VeriSimDB Client SDK -//! -//! A Rust client library for interacting with VeriSimDB — a multi-modal database -//! supporting octad entity management, drift detection, provenance tracking, -//! VCL query execution, and federation across distributed instances. -//! -//! ## Quick Start -//! -//! ```rust,no_run -//! use verisimdb_client::client::VeriSimClient; -//! use verisimdb_client::types::OctadInput; -//! -//! #[tokio::main] -//! async fn main() -> verisimdb_client::error::Result<()> { -//! let client = VeriSimClient::new("http://localhost:8080")?; -//! let healthy = client.health().await?; -//! println!("VeriSimDB healthy: {healthy}"); -//! Ok(()) -//! } -//! ``` -//! -//! ## Modules -//! -//! - [`client`] — Connection configuration, authentication, and HTTP transport. -//! - [`types`] — Data types mirroring the VeriSimDB JSON Schema (Octad, Modality, etc.). -//! - [`octad`] — CRUD operations for octad entities. -//! - [`search`] — Text, vector, graph-relational, and spatial search operations. -//! - [`drift`] — Drift score retrieval and normalization triggers. -//! - [`provenance`] — Immutable provenance chain management. -//! - [`vcl`] — VeriSim Consonance Language execution and explain plans. -//! - [`federation`] — Peer registration and federated cross-instance queries. -//! - [`error`] — Error types and the crate-level `Result` alias. - -#![forbid(unsafe_code)] -pub mod client; -pub mod types; -pub mod octad; -pub mod search; -pub mod drift; -pub mod provenance; -pub mod vcl; -pub mod federation; -pub mod error; diff --git a/verisimdb/connectors/clients/rust/src/octad.rs b/verisimdb/connectors/clients/rust/src/octad.rs deleted file mode 100644 index 8deb153b..00000000 --- a/verisimdb/connectors/clients/rust/src/octad.rs +++ /dev/null @@ -1,87 +0,0 @@ -// SPDX-License-Identifier: MPL-2.0 -// Copyright (c) 2026 Jonathan D.A. Jewell (hyperpolymath) - -//! Octad CRUD operations. -//! -//! Octads are the fundamental multi-modal entities in VeriSimDB. Each octad can -//! carry data across all eight modalities (graph, vector, tensor, semantic, -//! document, temporal, provenance, spatial). This module provides create, read, -//! update, delete, and list operations as methods on [`VeriSimClient`]. - -use crate::client::VeriSimClient; -use crate::error::Result; -use crate::types::{Octad, OctadInput, PaginatedResponse}; - -impl VeriSimClient { - /// Create a new octad entity. - /// - /// The server assigns a UUID and timestamps; the returned [`Octad`] contains - /// the fully-populated record. - /// - /// # Arguments - /// - /// * `input` — The octad payload. At minimum, `name` should be set. - /// - /// # Errors - /// - /// Returns [`VeriSimError::Validation`] if required fields are missing, or a - /// network / server error on transport failure. - pub async fn create_octad(&self, input: &OctadInput) -> Result { - self.post("/api/v1/octads", input).await - } - - /// Retrieve a single octad by its unique identifier. - /// - /// # Errors - /// - /// Returns [`VeriSimError::NotFound`] if no octad exists with the given `id`. - pub async fn get_octad(&self, id: &str) -> Result { - let path = format!("/api/v1/octads/{id}"); - self.get(&path).await - } - - /// Update an existing octad entity. - /// - /// Only the fields present in `input` are modified; omitted fields retain - /// their current values (partial update / merge semantics). - /// - /// # Errors - /// - /// Returns [`VeriSimError::NotFound`] if the octad does not exist. - pub async fn update_octad(&self, id: &str, input: &OctadInput) -> Result { - let path = format!("/api/v1/octads/{id}"); - self.put(&path, input).await - } - - /// Delete a octad entity by its unique identifier. - /// - /// This is a hard delete — the entity and all associated modality data are - /// removed. Provenance records are retained for auditability. - /// - /// # Errors - /// - /// Returns [`VeriSimError::NotFound`] if the octad does not exist. - pub async fn delete_octad(&self, id: &str) -> Result<()> { - let path = format!("/api/v1/octads/{id}"); - self.delete(&path).await - } - - /// List octad entities with pagination. - /// - /// # Arguments - /// - /// * `limit` — Maximum number of results to return (server may cap this). - /// * `offset` — Zero-based offset for pagination. - /// - /// # Returns - /// - /// A [`PaginatedResponse`] containing the requested page of octads. - pub async fn list_octads( - &self, - limit: usize, - offset: usize, - ) -> Result> { - let path = format!("/api/v1/octads?limit={limit}&offset={offset}"); - self.get(&path).await - } -} diff --git a/verisimdb/connectors/clients/rust/src/provenance.rs b/verisimdb/connectors/clients/rust/src/provenance.rs deleted file mode 100644 index 90f76bb5..00000000 --- a/verisimdb/connectors/clients/rust/src/provenance.rs +++ /dev/null @@ -1,68 +0,0 @@ -// SPDX-License-Identifier: MPL-2.0 -// Copyright (c) 2026 Jonathan D.A. Jewell (hyperpolymath) - -//! Provenance chain operations. -//! -//! Every octad entity in VeriSimDB maintains an append-only provenance chain -//! recording creation, transformation, derivation, and access events. This -//! module provides methods to read, append to, and cryptographically verify -//! provenance chains. - -use crate::client::VeriSimClient; -use crate::error::Result; -use crate::types::ProvenanceEvent; - -impl VeriSimClient { - /// Retrieve the full provenance chain for a octad entity. - /// - /// Events are returned in chronological order (oldest first). - /// - /// # Arguments - /// - /// * `id` — The octad entity identifier. - /// - /// # Errors - /// - /// Returns [`VeriSimError::NotFound`] if the entity does not exist. - pub async fn get_provenance_chain(&self, id: &str) -> Result> { - let path = format!("/api/v1/provenance/{id}"); - self.get(&path).await - } - - /// Append a new event to a octad's provenance chain. - /// - /// The event is immutably recorded; its timestamp and identifier are - /// assigned by the server. The returned [`ProvenanceEvent`] contains the - /// server-assigned fields. - /// - /// # Arguments - /// - /// * `id` — The octad entity identifier. - /// * `event` — The provenance event to record. - pub async fn record_provenance( - &self, - id: &str, - event: &ProvenanceEvent, - ) -> Result { - let path = format!("/api/v1/provenance/{id}"); - self.post(&path, event).await - } - - /// Verify the integrity of a octad's provenance chain. - /// - /// The server checks that the chain is contiguous, that no events have been - /// tampered with, and that cryptographic hashes (if enabled) are consistent. - /// - /// # Returns - /// - /// A JSON object with `valid: bool`, `chain_length: usize`, and optional - /// `errors` array describing any integrity violations. - /// - /// # Arguments - /// - /// * `id` — The octad entity identifier. - pub async fn verify_provenance(&self, id: &str) -> Result { - let path = format!("/api/v1/provenance/{id}/verify"); - self.get(&path).await - } -} diff --git a/verisimdb/connectors/clients/rust/src/search.rs b/verisimdb/connectors/clients/rust/src/search.rs deleted file mode 100644 index 9d713984..00000000 --- a/verisimdb/connectors/clients/rust/src/search.rs +++ /dev/null @@ -1,192 +0,0 @@ -// SPDX-License-Identifier: MPL-2.0 -// Copyright (c) 2026 Jonathan D.A. Jewell (hyperpolymath) - -//! Search operations across VeriSimDB's multi-modal octad entities. -//! -//! Supports full-text search, vector similarity (k-NN), graph-relational -//! traversal, and geospatial queries (radius, bounding box, nearest-neighbour). - -use serde::Serialize; - -use crate::client::VeriSimClient; -use crate::error::Result; -use crate::types::Octad; - -// --------------------------------------------------------------------------- -// Internal request bodies -// --------------------------------------------------------------------------- - -/// Body for vector similarity search requests. -#[derive(Debug, Serialize)] -struct VectorSearchRequest { - vector: Vec, - k: usize, -} - -/// Body for radius-based spatial search. -#[derive(Debug, Serialize)] -struct SpatialRadiusRequest { - latitude: f64, - longitude: f64, - radius_km: f64, - limit: usize, -} - -/// Body for bounding-box spatial search. -#[derive(Debug, Serialize)] -struct SpatialBoundsRequest { - min_lat: f64, - min_lon: f64, - max_lat: f64, - max_lon: f64, - limit: usize, -} - -/// Body for nearest-neighbour spatial search. -#[derive(Debug, Serialize)] -struct SpatialNearestRequest { - latitude: f64, - longitude: f64, - k: usize, -} - -impl VeriSimClient { - /// Full-text search across octad names, descriptions, and document content. - /// - /// # Arguments - /// - /// * `query` — The search query string (supports simple keyword matching). - /// * `limit` — Maximum number of results to return. - pub async fn search_text(&self, query: &str, limit: usize) -> Result> { - let path = format!( - "/api/v1/search/text?q={}&limit={limit}", - urlencoding_encode(query) - ); - self.get(&path).await - } - - /// Vector similarity search (k-nearest neighbours). - /// - /// Finds the `k` octads whose stored vector embeddings are closest to the - /// provided `vector` (cosine similarity by default). - /// - /// # Arguments - /// - /// * `vector` — The query embedding. - /// * `k` — Number of nearest neighbours to return. - pub async fn search_vector(&self, vector: &[f32], k: usize) -> Result> { - let body = VectorSearchRequest { - vector: vector.to_vec(), - k, - }; - self.post("/api/v1/search/vector", &body).await - } - - /// Find octads related to the given entity via graph edges. - /// - /// Traverses one hop of the graph modality and returns all directly - /// connected octads. - /// - /// # Arguments - /// - /// * `id` — The octad identifier to find relations for. - pub async fn search_related(&self, id: &str) -> Result> { - let path = format!("/api/v1/search/related/{id}"); - self.get(&path).await - } - - /// Spatial search: find octads within a given radius of a point. - /// - /// # Arguments - /// - /// * `lat` — Centre latitude (WGS 84 decimal degrees). - /// * `lon` — Centre longitude (WGS 84 decimal degrees). - /// * `radius_km` — Search radius in kilometres. - /// * `limit` — Maximum number of results. - pub async fn search_spatial_radius( - &self, - lat: f64, - lon: f64, - radius_km: f64, - limit: usize, - ) -> Result> { - let body = SpatialRadiusRequest { - latitude: lat, - longitude: lon, - radius_km, - limit, - }; - self.post("/api/v1/search/spatial/radius", &body).await - } - - /// Spatial search: find octads within a rectangular bounding box. - /// - /// # Arguments - /// - /// * `min_lat` — Southern boundary latitude. - /// * `min_lon` — Western boundary longitude. - /// * `max_lat` — Northern boundary latitude. - /// * `max_lon` — Eastern boundary longitude. - /// * `limit` — Maximum number of results. - pub async fn search_spatial_bounds( - &self, - min_lat: f64, - min_lon: f64, - max_lat: f64, - max_lon: f64, - limit: usize, - ) -> Result> { - let body = SpatialBoundsRequest { - min_lat, - min_lon, - max_lat, - max_lon, - limit, - }; - self.post("/api/v1/search/spatial/bounds", &body).await - } - - /// Spatial search: find the `k` nearest octads to a given point. - /// - /// # Arguments - /// - /// * `lat` — Query point latitude (WGS 84 decimal degrees). - /// * `lon` — Query point longitude (WGS 84 decimal degrees). - /// * `k` — Number of nearest neighbours to return. - pub async fn search_spatial_nearest( - &self, - lat: f64, - lon: f64, - k: usize, - ) -> Result> { - let body = SpatialNearestRequest { - latitude: lat, - longitude: lon, - k, - }; - self.post("/api/v1/search/spatial/nearest", &body).await - } -} - -// --------------------------------------------------------------------------- -// Helpers -// --------------------------------------------------------------------------- - -/// Minimal percent-encoding for query string values. -/// -/// This avoids pulling in a full `percent-encoding` crate for a single use. -fn urlencoding_encode(input: &str) -> String { - let mut output = String::with_capacity(input.len()); - for byte in input.bytes() { - match byte { - b'A'..=b'Z' | b'a'..=b'z' | b'0'..=b'9' | b'-' | b'_' | b'.' | b'~' => { - output.push(byte as char); - } - _ => { - output.push('%'); - output.push_str(&format!("{byte:02X}")); - } - } - } - output -} diff --git a/verisimdb/connectors/clients/rust/src/types.rs b/verisimdb/connectors/clients/rust/src/types.rs deleted file mode 100644 index cea0ce8a..00000000 --- a/verisimdb/connectors/clients/rust/src/types.rs +++ /dev/null @@ -1,392 +0,0 @@ -// SPDX-License-Identifier: MPL-2.0 -// Copyright (c) 2026 Jonathan D.A. Jewell (hyperpolymath) - -//! Core data types for the VeriSimDB client SDK. -//! -//! These types mirror the VeriSimDB JSON Schema and cover the full octad of -//! modalities: Graph, Vector, Tensor, Semantic, Document, Temporal, Provenance, -//! and Spatial. Every struct derives `Serialize` and `Deserialize` so it can be -//! round-tripped through the REST API transparently. - -use std::collections::HashMap; - -use chrono::{DateTime, Utc}; -use serde::{Deserialize, Serialize}; - -// --------------------------------------------------------------------------- -// Modality enum -// --------------------------------------------------------------------------- - -/// The eight modalities supported by VeriSimDB's octad data model. -/// -/// Each octad entity can participate in any combination of these modalities, -/// enabling truly multi-modal storage and querying. -#[derive(Debug, Clone, Copy, PartialEq, Eq, Hash, Serialize, Deserialize)] -#[serde(rename_all = "snake_case")] -pub enum Modality { - /// Graph relationships (nodes, edges, properties). - Graph, - /// Dense vector embeddings for similarity search. - Vector, - /// Multi-dimensional tensor data (ML feature stores, etc.). - Tensor, - /// Semantic triples and ontology-backed knowledge. - Semantic, - /// Unstructured or semi-structured document content. - Document, - /// Time-series and temporal event data. - Temporal, - /// Immutable provenance / lineage chains. - Provenance, - /// Geospatial coordinates, regions, and geometries. - Spatial, -} - -// --------------------------------------------------------------------------- -// ModalityStatus -// --------------------------------------------------------------------------- - -/// Boolean flags indicating which modalities are currently active for a octad. -#[derive(Debug, Clone, Default, Serialize, Deserialize)] -pub struct ModalityStatus { - /// Whether graph data is present. - pub graph: bool, - /// Whether vector embeddings are present. - pub vector: bool, - /// Whether tensor data is present. - pub tensor: bool, - /// Whether semantic triples are present. - pub semantic: bool, - /// Whether document content is present. - pub document: bool, - /// Whether temporal events are present. - pub temporal: bool, - /// Whether provenance records are present. - pub provenance: bool, - /// Whether spatial geometry is present. - pub spatial: bool, -} - -// --------------------------------------------------------------------------- -// OctadStatus -// --------------------------------------------------------------------------- - -/// Lightweight status summary for a octad entity, returned by list/search -/// endpoints that do not need full payloads. -#[derive(Debug, Clone, Serialize, Deserialize)] -pub struct OctadStatus { - /// Unique identifier. - pub id: String, - /// Creation timestamp. - pub created_at: DateTime, - /// Last modification timestamp. - pub modified_at: DateTime, - /// Optimistic concurrency version counter. - pub version: u64, - /// Per-modality activation flags. - pub modality_status: ModalityStatus, -} - -// --------------------------------------------------------------------------- -// Octad (full entity) -// --------------------------------------------------------------------------- - -/// A complete octad entity encompassing all eight modality payloads. -/// -/// This is the primary read model returned by `get_octad` and single-entity -/// search results. -#[derive(Debug, Clone, Serialize, Deserialize)] -pub struct Octad { - /// Unique identifier (UUID v4). - pub id: String, - /// Human-readable label. - pub name: String, - /// Optional free-text description. - pub description: Option, - /// Creation timestamp. - pub created_at: DateTime, - /// Last modification timestamp. - pub modified_at: DateTime, - /// Optimistic concurrency version counter. - pub version: u64, - /// Per-modality activation flags. - pub modality_status: ModalityStatus, - /// Arbitrary user-defined metadata. - pub metadata: Option, - /// Graph modality payload. - pub graph: Option, - /// Vector modality payload. - pub vector: Option, - /// Tensor modality payload. - pub tensor: Option, - /// Semantic modality payload. - pub semantic: Option, - /// Document modality payload. - pub document: Option, - /// Temporal modality payload. - pub temporal: Option, - /// Provenance modality payload. - pub provenance: Option, - /// Spatial modality payload. - pub spatial: Option, -} - -// --------------------------------------------------------------------------- -// OctadInput (create / update payload) -// --------------------------------------------------------------------------- - -/// Input payload for creating or updating a octad entity. -/// -/// All fields are optional so that partial updates are possible. The server -/// merges the provided fields into the existing entity on update. -#[derive(Debug, Clone, Default, Serialize, Deserialize)] -pub struct OctadInput { - /// Human-readable label. - pub name: Option, - /// Free-text description. - pub description: Option, - /// Arbitrary user-defined metadata. - pub metadata: Option, - /// Graph modality input. - pub graph: Option, - /// Vector modality input. - pub vector: Option, - /// Tensor modality input. - pub tensor: Option, - /// Semantic modality input. - pub semantic: Option, - /// Document modality input. - pub document: Option, - /// Temporal modality input. - pub temporal: Option, - /// Provenance modality input. - pub provenance: Option, - /// Spatial modality input. - pub spatial: Option, -} - -// --------------------------------------------------------------------------- -// Per-modality input structs -// --------------------------------------------------------------------------- - -/// Graph modality input: nodes, edges, and properties. -#[derive(Debug, Clone, Serialize, Deserialize)] -pub struct OctadGraphInput { - /// List of node identifiers to associate with this octad. - pub nodes: Option>, - /// List of edges, each a (source, target, label) triple. - pub edges: Option>, - /// Arbitrary graph-level properties. - pub properties: Option, -} - -/// A single directed edge in the graph modality. -#[derive(Debug, Clone, Serialize, Deserialize)] -pub struct GraphEdge { - /// Source node identifier. - pub source: String, - /// Target node identifier. - pub target: String, - /// Edge label / relationship type. - pub label: String, - /// Optional edge-level properties. - pub properties: Option, -} - -/// Vector modality input: dense embeddings for similarity search. -#[derive(Debug, Clone, Serialize, Deserialize)] -pub struct OctadVectorInput { - /// The embedding vector (list of f32 values). - pub embedding: Vec, - /// Dimensionality (inferred from `embedding.len()` if omitted). - pub dimensions: Option, - /// The model or algorithm that produced this embedding. - pub model: Option, -} - -/// Tensor modality input: multi-dimensional numeric data. -#[derive(Debug, Clone, Serialize, Deserialize)] -pub struct OctadTensorInput { - /// Flattened tensor data. - pub data: Vec, - /// Shape of the tensor (e.g. `[3, 224, 224]`). - pub shape: Vec, - /// Data type label (e.g. "float32", "int64"). - pub dtype: Option, -} - -/// Semantic modality input: RDF-style triples and ontology references. -#[derive(Debug, Clone, Serialize, Deserialize)] -pub struct OctadSemanticInput { - /// Semantic triples (subject, predicate, object). - pub triples: Option>, - /// Ontology URI this entity conforms to. - pub ontology: Option, - /// Free-form semantic annotations. - pub annotations: Option, -} - -/// A single semantic triple. -#[derive(Debug, Clone, Serialize, Deserialize)] -pub struct SemanticTriple { - /// Subject URI or identifier. - pub subject: String, - /// Predicate URI or identifier. - pub predicate: String, - /// Object URI, identifier, or literal value. - pub object: String, -} - -/// Document modality input: unstructured / semi-structured content. -#[derive(Debug, Clone, Serialize, Deserialize)] -pub struct OctadDocumentInput { - /// The document body (plain text, HTML, Markdown, etc.). - pub content: String, - /// MIME type of the content (e.g. "text/plain", "application/json"). - pub content_type: Option, - /// Language code (e.g. "en", "fr"). - pub language: Option, - /// Document-level metadata (author, tags, etc.). - pub metadata: Option, -} - -/// Provenance modality input: lineage and audit trail events. -#[derive(Debug, Clone, Serialize, Deserialize)] -pub struct OctadProvenanceInput { - /// The type of provenance event (e.g. "creation", "transformation", "derivation"). - pub event_type: String, - /// The agent (user, service, pipeline) that triggered the event. - pub agent: String, - /// Human-readable description of what happened. - pub description: Option, - /// References to source entities this was derived from. - pub source_ids: Option>, - /// Arbitrary event-level metadata. - pub metadata: Option, -} - -/// Spatial modality input: geospatial coordinates and geometries. -#[derive(Debug, Clone, Serialize, Deserialize)] -pub struct OctadSpatialInput { - /// Latitude in decimal degrees (WGS 84). - pub latitude: Option, - /// Longitude in decimal degrees (WGS 84). - pub longitude: Option, - /// Altitude in metres above mean sea level. - pub altitude: Option, - /// GeoJSON geometry object for complex shapes. - pub geometry: Option, - /// Coordinate reference system identifier (default: "EPSG:4326"). - pub crs: Option, -} - -/// Temporal modality input: time-series events and temporal metadata. -#[derive(Debug, Clone, Serialize, Deserialize)] -pub struct OctadTemporalInput { - /// The timestamp of the event. - pub timestamp: DateTime, - /// Duration in milliseconds (for interval events). - pub duration_ms: Option, - /// Recurrence rule (iCal RRULE format). - pub recurrence: Option, - /// Timezone identifier (e.g. "Europe/London"). - pub timezone: Option, - /// Arbitrary temporal metadata. - pub metadata: Option, -} - -// --------------------------------------------------------------------------- -// DriftScore -// --------------------------------------------------------------------------- - -/// Drift score for a octad entity, measuring how far its modality data has -/// diverged from its normalised baseline. -#[derive(Debug, Clone, Serialize, Deserialize)] -pub struct DriftScore { - /// The octad entity identifier. - pub entity_id: String, - /// Aggregate drift score across all active modalities (0.0 = no drift). - pub overall_score: f64, - /// Per-modality drift scores keyed by modality name. - pub modality_scores: HashMap, - /// When the drift was last computed. - pub last_checked: DateTime, - /// Whether the entity should be re-normalised. - pub needs_normalization: bool, -} - -// --------------------------------------------------------------------------- -// ProvenanceEvent -// --------------------------------------------------------------------------- - -/// A single immutable event in a octad's provenance chain. -#[derive(Debug, Clone, Serialize, Deserialize)] -pub struct ProvenanceEvent { - /// Unique event identifier. - pub id: Option, - /// The octad entity this event belongs to. - pub entity_id: String, - /// Event type (e.g. "creation", "transformation", "derivation"). - pub event_type: String, - /// The agent that triggered the event. - pub agent: String, - /// Human-readable description. - pub description: Option, - /// When the event occurred. - pub timestamp: Option>, - /// References to predecessor events or source entities. - pub source_ids: Option>, - /// Arbitrary event metadata. - pub metadata: Option, -} - -// --------------------------------------------------------------------------- -// FederationResult -// --------------------------------------------------------------------------- - -/// A single result from a federated cross-instance query. -#[derive(Debug, Clone, Serialize, Deserialize)] -pub struct FederationResult { - /// The peer store that produced this result. - pub store_id: String, - /// The matched octad entity. - pub entity: Octad, - /// Relevance or similarity score (interpretation depends on query type). - pub score: Option, - /// Latency in milliseconds for this peer's response. - pub latency_ms: Option, -} - -// --------------------------------------------------------------------------- -// ErrorResponse -// --------------------------------------------------------------------------- - -/// Standard error response body from the VeriSimDB REST API. -#[derive(Debug, Clone, Serialize, Deserialize)] -pub struct ErrorResponse { - /// Machine-readable error code. - pub error: String, - /// Human-readable error message. - pub message: String, - /// Optional details (validation errors, stack traces in debug mode, etc.). - pub details: Option, -} - -// --------------------------------------------------------------------------- -// PaginatedResponse -// --------------------------------------------------------------------------- - -/// Generic wrapper for paginated list responses. -#[derive(Debug, Clone, Serialize, Deserialize)] -pub struct PaginatedResponse { - /// The result items for the current page. - pub data: Vec, - /// Total number of matching items across all pages. - pub total: usize, - /// Number of items per page. - pub limit: usize, - /// Zero-based offset of the current page. - pub offset: usize, - /// Whether more pages are available after this one. - pub has_more: bool, -} diff --git a/verisimdb/connectors/clients/rust/src/vcl.rs b/verisimdb/connectors/clients/rust/src/vcl.rs deleted file mode 100644 index 014bdda9..00000000 --- a/verisimdb/connectors/clients/rust/src/vcl.rs +++ /dev/null @@ -1,72 +0,0 @@ -// SPDX-License-Identifier: MPL-2.0 -// Copyright (c) 2026 Jonathan D.A. Jewell (hyperpolymath) - -//! VeriSim Consonance Language (VCL) execution. -//! -//! VCL is VeriSimDB's native query language, supporting SQL-like syntax extended -//! with multi-modal operations (vector similarity, graph traversal, spatial -//! predicates, drift thresholds, etc.). This module provides methods to execute -//! VCL statements and retrieve explain / query plans. - -use serde::{Deserialize, Serialize}; - -use crate::client::VeriSimClient; -use crate::error::Result; - -/// Response from a VCL query execution or explain request. -#[derive(Debug, Clone, Serialize, Deserialize)] -pub struct VclResponse { - /// Whether the query executed successfully. - pub success: bool, - /// The type of VCL statement ("SELECT", "INSERT", "UPDATE", "DELETE", "EXPLAIN", etc.). - pub statement_type: String, - /// Number of rows affected or returned. - pub row_count: usize, - /// The result data (rows for SELECT, affected IDs for mutations, plan for EXPLAIN). - pub data: serde_json::Value, - /// Optional human-readable message (warnings, notices, etc.). - pub message: Option, -} - -/// Internal request body for VCL execution. -#[derive(Debug, Serialize)] -struct VclRequest { - query: String, -} - -impl VeriSimClient { - /// Execute a VCL statement against the VeriSimDB instance. - /// - /// Supports SELECT, INSERT, UPDATE, DELETE, and VeriSimDB-specific - /// statements like `DRIFT CHECK`, `NORMALIZE`, and `FEDERATE`. - /// - /// # Arguments - /// - /// * `query` — The VCL statement string. - /// - /// # Errors - /// - /// Returns [`VeriSimError::Server`] if the query has syntax errors or - /// the server rejects it for semantic reasons. - pub async fn execute_vcl(&self, query: &str) -> Result { - let body = VclRequest { - query: query.to_owned(), - }; - self.post("/api/v1/vcl/execute", &body).await - } - - /// Request an explain / query plan for a VCL statement without executing it. - /// - /// Useful for understanding which modalities, indices, and federation peers - /// would be involved in a query. - /// - /// # Arguments - /// - /// * `query` — The VCL statement string to explain. - pub async fn explain_vcl(&self, query: &str) -> Result { - let body = VclRequest { - query: query.to_owned(), - }; - self.post("/api/v1/vcl/explain", &body).await - } -} diff --git a/verisimdb/connectors/clients/vlang/MIGRATION.adoc b/verisimdb/connectors/clients/vlang/MIGRATION.adoc deleted file mode 100644 index 63f87149..00000000 --- a/verisimdb/connectors/clients/vlang/MIGRATION.adoc +++ /dev/null @@ -1,47 +0,0 @@ -// SPDX-License-Identifier: CC-BY-SA-4.0 -// Copyright (c) 2026 Jonathan D.A. Jewell (hyperpolymath) -= VeriSimDB V Client — Deprecated (2026-04-12) -:toc: - -== Status - -The V-lang VeriSimDB client (`src/`) is *deprecated* as of 2026-04-12 -following the estate-wide V-lang ban (2026-04-10). - -== Migration Target: Rust Client - -The canonical Rust client at `connectors/clients/rust/` provides identical -coverage: - -[cols="1,1"] -|=== -| V module (deprecated) | Rust module - -| `src/client.v` | `src/client.rs` -| `src/octad.v` | `src/octad.rs` -| `src/drift.v` | `src/drift.rs` -| `src/search.v` | `src/search.rs` -| `src/federation.v` | `src/federation.rs` -| `src/provenance.v` | `src/provenance.rs` -| `src/vcl.v` | `src/vcl.rs` -| `src/error.v` | `src/error.rs` -| `src/types.v` | `src/types.rs` -|=== - -[source,toml] ----- -[dependencies] -verisimdb-client = { path = "connectors/clients/rust" } ----- - -== Auth Migration - -The V `Auth` discriminated union maps directly to the Rust enum: - -[source,rust] ----- -use verisimdb_client::{Auth, Client}; - -let client = Client::new("http://localhost:8200") - .with_auth(Auth::ApiKey("your-key".to_string())); ----- diff --git a/verisimdb/connectors/shared/json-schema/drift-score.json b/verisimdb/connectors/shared/json-schema/drift-score.json deleted file mode 100644 index 0f4b001a..00000000 --- a/verisimdb/connectors/shared/json-schema/drift-score.json +++ /dev/null @@ -1,87 +0,0 @@ -{ - "$schema": "https://json-schema.org/draft/2020-12/schema", - "$id": "https://verisim.db/schema/drift-score.json", - "$comment": "SPDX-License-Identifier: MPL-2.0", - "title": "DriftScore", - "description": "Drift measurement for a single octad entity. Drift quantifies the divergence between an entity's modality representations -- when modalities fall out of sync (e.g. the vector embedding no longer matches the document content), the drift score increases. Scores are per-modality and aggregated into an overall score.", - "type": "object", - "required": ["entity_id", "overall_score", "modality_scores", "last_checked", "needs_normalization"], - "properties": { - "entity_id": { - "type": "string", - "description": "Octad entity identifier for which drift was measured." - }, - "overall_score": { - "type": "number", - "minimum": 0.0, - "maximum": 1.0, - "description": "Aggregated drift score across all modalities, normalised to [0.0, 1.0]. 0.0 means perfect cross-modal consistency; 1.0 means complete divergence." - }, - "modality_scores": { - "type": "object", - "description": "Per-modality drift scores. Each key is a modality name; each value is the drift score for that modality pair comparison.", - "properties": { - "graph": { - "type": "number", - "minimum": 0.0, - "maximum": 1.0, - "description": "Drift score for graph-vs-other modality consistency." - }, - "vector": { - "type": "number", - "minimum": 0.0, - "maximum": 1.0, - "description": "Drift score for vector-vs-semantic (semantic_vector_drift)." - }, - "tensor": { - "type": "number", - "minimum": 0.0, - "maximum": 1.0, - "description": "Drift score for tensor representation divergence." - }, - "semantic": { - "type": "number", - "minimum": 0.0, - "maximum": 1.0, - "description": "Drift score for semantic annotation consistency." - }, - "document": { - "type": "number", - "minimum": 0.0, - "maximum": 1.0, - "description": "Drift score for document-vs-graph (graph_document_drift)." - }, - "temporal": { - "type": "number", - "minimum": 0.0, - "maximum": 1.0, - "description": "Drift score for temporal version consistency." - }, - "provenance": { - "type": "number", - "minimum": 0.0, - "maximum": 1.0, - "description": "Drift score for provenance chain integrity." - }, - "spatial": { - "type": "number", - "minimum": 0.0, - "maximum": 1.0, - "description": "Drift score for spatial-vs-document/graph location consistency." - } - }, - "required": ["graph", "vector", "tensor", "semantic", "document", "temporal", "provenance", "spatial"], - "additionalProperties": false - }, - "last_checked": { - "type": "string", - "format": "date-time", - "description": "ISO 8601 timestamp of when this drift measurement was last performed." - }, - "needs_normalization": { - "type": "boolean", - "description": "Whether the entity's drift score exceeds the configured threshold and requires normalisation. When true, the normaliser should be triggered to repair cross-modal consistency." - } - }, - "additionalProperties": false -} diff --git a/verisimdb/connectors/shared/json-schema/error.json b/verisimdb/connectors/shared/json-schema/error.json deleted file mode 100644 index 8c7096ff..00000000 --- a/verisimdb/connectors/shared/json-schema/error.json +++ /dev/null @@ -1,62 +0,0 @@ -{ - "$schema": "https://json-schema.org/draft/2020-12/schema", - "$id": "https://verisim.db/schema/error.json", - "$comment": "SPDX-License-Identifier: MPL-2.0", - "title": "ErrorResponse", - "description": "Standard error response envelope for the VeriSimDB API. All error responses follow this structure, providing a machine-readable error code, a human-readable message, the HTTP status, and optional structured details for programmatic handling.", - "type": "object", - "required": ["error", "message", "status"], - "properties": { - "error": { - "type": "string", - "description": "Machine-readable error code. Uses snake_case naming convention. Standard codes include: 'not_found', 'validation_error', 'consistency_violation', 'modality_error', 'drift_detected', 'chain_corrupted', 'unauthorized', 'forbidden', 'rate_limited', 'internal_error', 'federation_error', 'timeout'.", - "examples": [ - "not_found", - "validation_error", - "consistency_violation", - "modality_error", - "drift_detected", - "chain_corrupted", - "unauthorized", - "internal_error" - ] - }, - "message": { - "type": "string", - "description": "Human-readable error message providing context about what went wrong. Suitable for display in developer tools and logs. Should not contain sensitive information.", - "examples": [ - "Entity not found: 550e8400-e29b-41d4-a716-446655440000", - "Tensor shape [3, 4] does not match data length 15 (expected 12)", - "Provenance chain corrupted for entity abc-123: record 5 parent_hash mismatch" - ] - }, - "status": { - "type": "integer", - "description": "HTTP status code that was returned with this error response.", - "minimum": 400, - "maximum": 599, - "examples": [400, 401, 403, 404, 409, 422, 429, 500, 502, 503, 504] - }, - "details": { - "type": "object", - "description": "Optional structured details for programmatic error handling. Content varies by error type. Examples include affected modalities, validation failures, drift scores, or federation peer information.", - "additionalProperties": true, - "examples": [ - { - "entity_id": "550e8400-e29b-41d4-a716-446655440000", - "modality": "vector", - "expected_dimension": 384, - "actual_dimension": 256 - }, - { - "field": "spatial.latitude", - "constraint": "range", - "min": -90.0, - "max": 90.0, - "actual": 91.5 - } - ] - } - }, - "additionalProperties": false -} diff --git a/verisimdb/connectors/shared/json-schema/federation-result.json b/verisimdb/connectors/shared/json-schema/federation-result.json deleted file mode 100644 index 0d5573ba..00000000 --- a/verisimdb/connectors/shared/json-schema/federation-result.json +++ /dev/null @@ -1,40 +0,0 @@ -{ - "$schema": "https://json-schema.org/draft/2020-12/schema", - "$id": "https://verisim.db/schema/federation-result.json", - "$comment": "SPDX-License-Identifier: MPL-2.0", - "title": "FederationResult", - "description": "A normalised result from a federated query. Each result represents a single octad entity returned by a specific peer store in the federation mesh, along with its relevance score, drift status, and response timing.", - "type": "object", - "required": ["source_store", "octad_id", "score", "drifted", "data", "response_time_ms"], - "properties": { - "source_store": { - "type": "string", - "description": "Identifier of the peer store that provided this result (e.g. 'store-1', '/universities/oxford')." - }, - "octad_id": { - "type": "string", - "description": "Octad entity identifier within the source store." - }, - "score": { - "type": "number", - "minimum": 0.0, - "maximum": 1.0, - "description": "Relevance score for this result, normalised to [0.0, 1.0]. Higher is more relevant. Scoring method depends on the query type (cosine similarity for vector, BM25 for text, etc.)." - }, - "drifted": { - "type": "boolean", - "description": "Whether the source store has known drift issues for this entity. When true, the entity's modality representations may be inconsistent. The federation coordinator annotates this based on the peer's reported drift scores and the query's drift policy." - }, - "data": { - "type": "object", - "description": "Result data payload. Structure depends on the queried modalities. May contain document content, graph edges, embedding vectors, or any combination.", - "additionalProperties": true - }, - "response_time_ms": { - "type": "integer", - "minimum": 0, - "description": "Time in milliseconds that the source store took to respond to this query." - } - }, - "additionalProperties": false -} diff --git a/verisimdb/connectors/shared/json-schema/modality.json b/verisimdb/connectors/shared/json-schema/modality.json deleted file mode 100644 index 4ec7e4a5..00000000 --- a/verisimdb/connectors/shared/json-schema/modality.json +++ /dev/null @@ -1,28 +0,0 @@ -{ - "$schema": "https://json-schema.org/draft/2020-12/schema", - "$id": "https://verisim.db/schema/modality.json", - "$comment": "SPDX-License-Identifier: MPL-2.0", - "title": "Modality", - "description": "VeriSimDB octad modality identifier. Each modality represents a distinct representational dimension of a octad entity. Together the 8 modalities form the octad -- the fundamental multi-modal unit of VeriSimDB.", - "type": "string", - "enum": [ - "graph", - "vector", - "tensor", - "semantic", - "document", - "temporal", - "provenance", - "spatial" - ], - "x-enum-descriptions": { - "graph": "RDF triples and property-graph edges. Captures relationships between entities.", - "vector": "Dense embedding for similarity search (cosine, dot-product, L2).", - "tensor": "Multi-dimensional numeric array. Shape + flattened data.", - "semantic": "Ontological type annotations, properties, and proof blobs.", - "document": "Full-text searchable content (title, body, fields).", - "temporal": "Version history and time-travel queries.", - "provenance": "Origin tracking, transformation chain, and actor trail (SHA-256 hash chain).", - "spatial": "Geospatial coordinates, geometry types, and proximity queries." - } -} diff --git a/verisimdb/connectors/shared/json-schema/octad-input.json b/verisimdb/connectors/shared/json-schema/octad-input.json deleted file mode 100644 index e7bf098a..00000000 --- a/verisimdb/connectors/shared/json-schema/octad-input.json +++ /dev/null @@ -1,215 +0,0 @@ -{ - "$schema": "https://json-schema.org/draft/2020-12/schema", - "$id": "https://verisim.db/schema/octad-input.json", - "$comment": "SPDX-License-Identifier: MPL-2.0", - "title": "OctadInput", - "description": "Input payload for creating or updating a VeriSimDB octad entity. All modality fields are optional -- only the modalities provided will be populated or updated. Omitted modalities are left unchanged on update, or unpopulated on create.", - "type": "object", - "properties": { - "graph": { - "oneOf": [ - { "type": "null" }, - { - "type": "object", - "description": "Graph modality input. Defines outgoing relationships from this entity to other entities.", - "required": ["relationships"], - "properties": { - "relationships": { - "type": "array", - "description": "List of outgoing edges. Each relationship is a (predicate, target_id) pair.", - "items": { - "type": "object", - "required": ["predicate", "target_id"], - "properties": { - "predicate": { - "type": "string", - "description": "Edge label or relationship type (IRI or plain string, e.g. 'relates_to', 'https://schema.org/knows')." - }, - "target_id": { - "type": "string", - "description": "Target entity octad ID or IRI." - } - } - } - } - } - } - ], - "description": "Graph relationships to create. Null or omitted to skip." - }, - "vector": { - "oneOf": [ - { "type": "null" }, - { - "type": "object", - "description": "Vector modality input. Provides a dense embedding for similarity search.", - "required": ["embedding"], - "properties": { - "embedding": { - "type": "array", - "items": { "type": "number" }, - "description": "Dense embedding vector (float32 values). Length must match the store's configured vector_dimension." - }, - "model": { - "type": "string", - "description": "Name or identifier of the embedding model used (e.g. 'all-MiniLM-L6-v2'). Optional, for provenance tracking." - } - } - } - ], - "description": "Vector embedding to set. Null or omitted to skip." - }, - "tensor": { - "oneOf": [ - { "type": "null" }, - { - "type": "object", - "description": "Tensor modality input. Multi-dimensional numeric data with shape metadata.", - "required": ["shape", "data"], - "properties": { - "shape": { - "type": "array", - "items": { "type": "integer", "minimum": 0 }, - "description": "Tensor dimensions (e.g. [3, 4] for a 3x4 matrix). The product of all values must equal the length of the data array.", - "minItems": 1 - }, - "data": { - "type": "array", - "items": { "type": "number" }, - "description": "Flattened tensor values in row-major order (float64)." - } - } - } - ], - "description": "Tensor data to set. Null or omitted to skip." - }, - "semantic": { - "oneOf": [ - { "type": "null" }, - { - "type": "object", - "description": "Semantic modality input. Ontological type annotations and property maps.", - "properties": { - "types": { - "type": "array", - "items": { "type": "string" }, - "description": "Type IRIs from ontologies (e.g. ['https://schema.org/Person', 'http://xmlns.com/foaf/0.1/Agent'])." - }, - "properties": { - "type": "object", - "additionalProperties": { "type": "string" }, - "description": "Semantic key-value properties. Keys are property IRIs or plain strings; values are string literals." - } - } - } - ], - "description": "Semantic annotations to set. Null or omitted to skip." - }, - "document": { - "oneOf": [ - { "type": "null" }, - { - "type": "object", - "description": "Document modality input. Full-text content for indexing and search.", - "required": ["title", "body"], - "properties": { - "title": { - "type": "string", - "description": "Document title. Used for display and search ranking." - }, - "body": { - "type": "string", - "description": "Document body. Full-text indexed for search." - }, - "fields": { - "type": "object", - "additionalProperties": { "type": "string" }, - "description": "Additional searchable fields beyond title and body (e.g. 'author', 'category')." - } - } - } - ], - "description": "Document content to set. Null or omitted to skip." - }, - "provenance": { - "oneOf": [ - { "type": "null" }, - { - "type": "object", - "description": "Provenance modality input. Records a lineage event in the entity's hash chain.", - "required": ["event_type", "actor", "description"], - "properties": { - "event_type": { - "type": "string", - "description": "Classification of the provenance event.", - "enum": ["created", "modified", "imported", "normalized", "drift_repaired", "deleted", "merged"] - }, - "actor": { - "type": "string", - "description": "Who or what caused this event (user ID, system component, bot name)." - }, - "source": { - "type": "string", - "description": "Optional source identifier (URL, file path, upstream entity ID)." - }, - "description": { - "type": "string", - "description": "Human-readable description of what happened." - } - } - } - ], - "description": "Provenance event to record. Null or omitted to skip." - }, - "spatial": { - "oneOf": [ - { "type": "null" }, - { - "type": "object", - "description": "Spatial modality input. Geospatial coordinates and geometry.", - "required": ["latitude", "longitude"], - "properties": { - "latitude": { - "type": "number", - "minimum": -90.0, - "maximum": 90.0, - "description": "Latitude in decimal degrees (WGS84). North is positive." - }, - "longitude": { - "type": "number", - "minimum": -180.0, - "maximum": 180.0, - "description": "Longitude in decimal degrees (WGS84). East is positive." - }, - "altitude": { - "type": "number", - "description": "Altitude in metres above the WGS84 ellipsoid. Optional." - }, - "geometry_type": { - "type": "string", - "enum": ["Point", "LineString", "Polygon", "MultiPoint", "MultiPolygon"], - "description": "OGC Simple Features geometry type. Defaults to 'Point' if omitted." - }, - "srid": { - "type": "integer", - "minimum": 0, - "description": "Spatial Reference System Identifier. Defaults to 4326 (WGS84) if omitted." - }, - "properties": { - "type": "object", - "additionalProperties": { "type": "string" }, - "description": "Arbitrary spatial metadata (address, region, accuracy, etc.)." - } - } - } - ], - "description": "Spatial data to set. Null or omitted to skip." - }, - "metadata": { - "type": "object", - "additionalProperties": { "type": "string" }, - "description": "Arbitrary key-value metadata. Not indexed by any specific modality but stored alongside the entity." - } - }, - "additionalProperties": false -} diff --git a/verisimdb/connectors/shared/json-schema/octad-status.json b/verisimdb/connectors/shared/json-schema/octad-status.json deleted file mode 100644 index b2f7e09e..00000000 --- a/verisimdb/connectors/shared/json-schema/octad-status.json +++ /dev/null @@ -1,47 +0,0 @@ -{ - "$schema": "https://json-schema.org/draft/2020-12/schema", - "$id": "https://verisim.db/schema/octad-status.json", - "$comment": "SPDX-License-Identifier: MPL-2.0", - "title": "OctadStatus", - "description": "Status of a octad entity across all eight octad modalities. Tracks creation time, modification time, current version, and which modalities are populated.", - "type": "object", - "required": ["id", "created_at", "modified_at", "version", "modality_status"], - "properties": { - "id": { - "type": "string", - "description": "Entity identifier (matches the parent Octad's id)." - }, - "created_at": { - "type": "string", - "format": "date-time", - "description": "ISO 8601 timestamp of when the entity was first created." - }, - "modified_at": { - "type": "string", - "format": "date-time", - "description": "ISO 8601 timestamp of the most recent modification." - }, - "version": { - "type": "integer", - "minimum": 0, - "description": "Monotonically increasing version number. Incremented on each mutation." - }, - "modality_status": { - "type": "object", - "description": "Boolean flags indicating which of the 8 octad modalities are populated for this entity.", - "required": ["graph", "vector", "tensor", "semantic", "document", "temporal", "provenance", "spatial"], - "properties": { - "graph": { "type": "boolean", "description": "Whether the graph modality is populated." }, - "vector": { "type": "boolean", "description": "Whether the vector modality is populated." }, - "tensor": { "type": "boolean", "description": "Whether the tensor modality is populated." }, - "semantic": { "type": "boolean", "description": "Whether the semantic modality is populated." }, - "document": { "type": "boolean", "description": "Whether the document modality is populated." }, - "temporal": { "type": "boolean", "description": "Whether the temporal modality is populated." }, - "provenance": { "type": "boolean", "description": "Whether the provenance modality is populated." }, - "spatial": { "type": "boolean", "description": "Whether the spatial modality is populated." } - }, - "additionalProperties": false - } - }, - "additionalProperties": false -} diff --git a/verisimdb/connectors/shared/json-schema/octad.json b/verisimdb/connectors/shared/json-schema/octad.json deleted file mode 100644 index 40d33f10..00000000 --- a/verisimdb/connectors/shared/json-schema/octad.json +++ /dev/null @@ -1,230 +0,0 @@ -{ - "$schema": "https://json-schema.org/draft/2020-12/schema", - "$id": "https://verisim.db/schema/octad.json", - "$comment": "SPDX-License-Identifier: MPL-2.0", - "title": "Octad", - "description": "A VeriSimDB octad entity with 8 synchronized modality representations. The octad is the fundamental unit of VeriSimDB -- each entity exists simultaneously across all eight modalities (Graph, Vector, Tensor, Semantic, Document, Temporal, Provenance, Spatial), maintaining cross-modal consistency through drift detection and self-normalisation.", - "type": "object", - "required": ["id", "status"], - "properties": { - "id": { - "type": "string", - "description": "Unique entity identifier (UUID v4). Serves as the canonical key across all modality stores.", - "pattern": "^[0-9a-f]{8}-[0-9a-f]{4}-4[0-9a-f]{3}-[89ab][0-9a-f]{3}-[0-9a-f]{12}$|^.+$", - "examples": ["550e8400-e29b-41d4-a716-446655440000"] - }, - "status": { - "$ref": "octad-status.json", - "description": "Entity status including creation time, modification time, version, and per-modality population flags." - }, - "graph_node": { - "oneOf": [ - { "type": "null" }, - { - "type": "object", - "description": "Graph modality representation. The entity as a node in an RDF / property graph with typed edges to other entities.", - "properties": { - "id": { - "type": "string", - "description": "Graph node IRI (derived from the octad ID and the store's base IRI)." - }, - "labels": { - "type": "array", - "items": { "type": "string" }, - "description": "RDF type IRIs or property-graph labels attached to this node." - }, - "properties": { - "type": "object", - "additionalProperties": { "type": "string" }, - "description": "Key-value properties stored directly on the graph node." - }, - "edges": { - "type": "array", - "description": "Outgoing edges from this node to other entities.", - "items": { - "type": "object", - "required": ["predicate", "target"], - "properties": { - "predicate": { - "type": "string", - "description": "Edge label / relationship type (IRI or plain string)." - }, - "target": { - "type": "string", - "description": "Target node IRI or octad ID." - } - } - } - } - } - } - ], - "description": "Graph node data, or null if the graph modality is not populated." - }, - "embedding": { - "oneOf": [ - { "type": "null" }, - { - "type": "array", - "items": { "type": "number" }, - "description": "Dense embedding vector (float32 values). Dimensionality matches the store's configured vector_dimension (default: 384)." - } - ], - "description": "Vector modality embedding, or null if not populated." - }, - "tensor": { - "oneOf": [ - { "type": "null" }, - { - "type": "object", - "description": "Multi-dimensional tensor data with shape metadata.", - "required": ["shape", "data"], - "properties": { - "shape": { - "type": "array", - "items": { "type": "integer", "minimum": 0 }, - "description": "Tensor dimensions (e.g. [3, 4] for a 3x4 matrix). The product of all shape values must equal the length of the data array." - }, - "data": { - "type": "array", - "items": { "type": "number" }, - "description": "Flattened tensor data in row-major order (float64 values)." - } - } - } - ], - "description": "Tensor modality data, or null if not populated." - }, - "semantic": { - "oneOf": [ - { "type": "null" }, - { - "type": "object", - "description": "Semantic modality annotation. Ontological types and properties with optional proof blobs for formal verification.", - "properties": { - "types": { - "type": "array", - "items": { "type": "string" }, - "description": "Type IRIs from ontologies (e.g. 'https://schema.org/Person')." - }, - "properties": { - "type": "object", - "additionalProperties": { "type": "string" }, - "description": "Semantic key-value properties." - }, - "proof": { - "oneOf": [ - { "type": "null" }, - { - "type": "object", - "description": "Cryptographic proof blob (CBOR-encoded verification evidence).", - "properties": { - "proof_type": { - "type": "string", - "description": "Proof system identifier (e.g. 'zkp', 'proven', 'sanctify')." - }, - "data": { - "type": "string", - "description": "Base64-encoded proof data." - } - } - } - ] - } - } - } - ], - "description": "Semantic annotation, or null if not populated." - }, - "document": { - "oneOf": [ - { "type": "null" }, - { - "type": "object", - "description": "Document modality for full-text search. Indexed by Tantivy.", - "required": ["title", "body"], - "properties": { - "title": { - "type": "string", - "description": "Document title, used for display and search boosting." - }, - "body": { - "type": "string", - "description": "Document body content, full-text indexed." - }, - "fields": { - "type": "object", - "additionalProperties": { "type": "string" }, - "description": "Additional searchable fields beyond title and body." - } - } - } - ], - "description": "Document content, or null if not populated." - }, - "version_count": { - "type": "integer", - "minimum": 0, - "description": "Number of temporal versions stored for this entity. Each mutation creates a new version." - }, - "provenance_chain_length": { - "type": "integer", - "minimum": 0, - "description": "Number of provenance records in the entity's hash chain. Each lifecycle event (created, modified, imported, normalized, etc.) appends a record." - }, - "spatial_data": { - "oneOf": [ - { "type": "null" }, - { - "type": "object", - "description": "Spatial modality data. Geographic coordinates and geometry in a configurable SRS (default WGS84 / EPSG:4326).", - "required": ["coordinates", "geometry_type", "srid"], - "properties": { - "coordinates": { - "type": "object", - "required": ["latitude", "longitude"], - "properties": { - "latitude": { - "type": "number", - "minimum": -90.0, - "maximum": 90.0, - "description": "Latitude in decimal degrees (WGS84). North is positive." - }, - "longitude": { - "type": "number", - "minimum": -180.0, - "maximum": 180.0, - "description": "Longitude in decimal degrees (WGS84). East is positive." - }, - "altitude": { - "oneOf": [ - { "type": "null" }, - { "type": "number" } - ], - "description": "Altitude in metres above the WGS84 ellipsoid. Optional." - } - } - }, - "geometry_type": { - "type": "string", - "enum": ["Point", "LineString", "Polygon", "MultiPoint", "MultiPolygon"], - "description": "OGC Simple Features geometry type." - }, - "srid": { - "type": "integer", - "minimum": 0, - "description": "Spatial Reference System Identifier (default: 4326 for WGS84)." - }, - "properties": { - "type": "object", - "additionalProperties": { "type": "string" }, - "description": "Arbitrary spatial metadata (address, region, accuracy, etc.)." - } - } - } - ], - "description": "Spatial data, or null if not populated." - } - }, - "additionalProperties": false -} diff --git a/verisimdb/connectors/shared/json-schema/provenance-event.json b/verisimdb/connectors/shared/json-schema/provenance-event.json deleted file mode 100644 index b0f07f60..00000000 --- a/verisimdb/connectors/shared/json-schema/provenance-event.json +++ /dev/null @@ -1,59 +0,0 @@ -{ - "$schema": "https://json-schema.org/draft/2020-12/schema", - "$id": "https://verisim.db/schema/provenance-event.json", - "$comment": "SPDX-License-Identifier: MPL-2.0", - "title": "ProvenanceEvent", - "description": "A single provenance event in an entity's lineage chain. Provenance records form a SHA-256 hash chain where each record's previous_hash is the content_hash of the preceding record, creating a tamper-evident audit trail. The first record in a chain uses the SHA-256 of the empty string as its previous_hash.", - "type": "object", - "required": ["event_id", "entity_id", "event_type", "actor", "description", "timestamp", "hash"], - "properties": { - "event_id": { - "type": "string", - "description": "Unique identifier for this provenance event. Typically the content_hash itself or a separate UUID." - }, - "entity_id": { - "type": "string", - "description": "Octad entity identifier that this provenance event belongs to." - }, - "event_type": { - "type": "string", - "description": "Classification of the provenance event. Standard types cover the entity lifecycle; 'custom:*' allows domain-specific extensions.", - "enum": [ - "created", - "modified", - "imported", - "normalized", - "federated", - "deleted" - ] - }, - "actor": { - "type": "string", - "description": "Who or what caused this event. May be a user ID, system component name, bot identifier, or service account." - }, - "source": { - "type": "string", - "description": "Optional source identifier providing context for where the data came from. Can be a URL, file path, upstream entity ID, or external system reference." - }, - "description": { - "type": "string", - "description": "Human-readable description of what happened in this event." - }, - "timestamp": { - "type": "string", - "format": "date-time", - "description": "ISO 8601 timestamp of when this provenance event occurred." - }, - "hash": { - "type": "string", - "pattern": "^[0-9a-f]{64}$", - "description": "SHA-256 hex digest of this record's canonical serialisation (the content_hash). Computed over event_type, actor, timestamp, source, description, and previous_hash." - }, - "previous_hash": { - "type": "string", - "pattern": "^[0-9a-f]{64}$", - "description": "SHA-256 hex digest of the preceding record in the chain (its content_hash). For the first record in a chain, this is the SHA-256 of the empty string: 'e3b0c44298fc1c149afbf4c8996fb92427ae41e4649b934ca495991b7852b855'." - } - }, - "additionalProperties": false -} diff --git a/verisimdb/connectors/shared/json-schema/query-params.json b/verisimdb/connectors/shared/json-schema/query-params.json deleted file mode 100644 index 23396f3f..00000000 --- a/verisimdb/connectors/shared/json-schema/query-params.json +++ /dev/null @@ -1,94 +0,0 @@ -{ - "$schema": "https://json-schema.org/draft/2020-12/schema", - "$id": "https://verisim.db/schema/query-params.json", - "$comment": "SPDX-License-Identifier: MPL-2.0", - "title": "QueryParams", - "description": "Parameters for a VeriSimDB federation query. Specifies which modalities to search, what to search for, and how to filter results. Used by federation adapters to translate queries into the target store's native language.", - "type": "object", - "required": ["modalities"], - "properties": { - "modalities": { - "type": "array", - "items": { "$ref": "modality.json" }, - "minItems": 1, - "uniqueItems": true, - "description": "Which octad modalities to include in the query. At least one modality must be specified." - }, - "limit": { - "type": "integer", - "minimum": 1, - "maximum": 10000, - "default": 100, - "description": "Maximum number of results to return. Defaults to 100, capped at 10000." - }, - "text_query": { - "type": "string", - "description": "Full-text search query string. Used by the document modality (Tantivy) and any adapter that supports text search." - }, - "vector_query": { - "type": "array", - "items": { "type": "number" }, - "description": "Vector for similarity search (cosine / dot-product / L2). Dimensionality must match the target store's configured vector_dimension." - }, - "graph_pattern": { - "type": "string", - "description": "Graph traversal pattern. Syntax depends on the adapter (SPARQL, Cypher, Gremlin, etc.). Federation protocol uses a simplified triple pattern: 'subject predicate object'." - }, - "spatial_bounds": { - "type": "object", - "description": "Bounding box for spatial queries. Defined by south-west and north-east corners in WGS84.", - "required": ["min_lat", "min_lon", "max_lat", "max_lon"], - "properties": { - "min_lat": { - "type": "number", - "minimum": -90.0, - "maximum": 90.0, - "description": "Southern latitude boundary." - }, - "min_lon": { - "type": "number", - "minimum": -180.0, - "maximum": 180.0, - "description": "Western longitude boundary." - }, - "max_lat": { - "type": "number", - "minimum": -90.0, - "maximum": 90.0, - "description": "Northern latitude boundary." - }, - "max_lon": { - "type": "number", - "minimum": -180.0, - "maximum": 180.0, - "description": "Eastern longitude boundary." - } - }, - "additionalProperties": false - }, - "temporal_range": { - "type": "object", - "description": "Time range filter for temporal queries. Both endpoints are inclusive.", - "required": ["start", "end"], - "properties": { - "start": { - "type": "string", - "format": "date-time", - "description": "Start of the time range (ISO 8601)." - }, - "end": { - "type": "string", - "format": "date-time", - "description": "End of the time range (ISO 8601)." - } - }, - "additionalProperties": false - }, - "filters": { - "type": "object", - "description": "Additional key-value filters applied after the primary query. Interpretation is adapter-specific. Common keys include 'status', 'actor', 'event_type', and 'geometry_type'.", - "additionalProperties": true - } - }, - "additionalProperties": false -} diff --git a/verisimdb/connectors/shared/openapi/verisim-api-v1.yaml b/verisimdb/connectors/shared/openapi/verisim-api-v1.yaml deleted file mode 100644 index 9c571055..00000000 --- a/verisimdb/connectors/shared/openapi/verisim-api-v1.yaml +++ /dev/null @@ -1,1024 +0,0 @@ -# SPDX-License-Identifier: MPL-2.0 -# Copyright (c) 2026 Jonathan D.A. Jewell (hyperpolymath) -# -# OpenAPI 3.1 specification for the VeriSimDB REST API. -# Covers all octad CRUD operations, search endpoints, drift monitoring, -# provenance chain management, spatial queries, normalisation triggers, -# and VCL query execution. - -openapi: "3.1.0" - -info: - title: VeriSimDB API - version: "1.0.0" - description: | - VeriSimDB (Veridical Simulacrum Database) REST API. - - Each entity (octad) exists simultaneously across 8 modalities -- the - octad: Graph, Vector, Tensor, Semantic, Document, Temporal, Provenance, - and Spatial. This API provides unified CRUD, search, drift detection, - normalisation, and provenance management for all modalities. - contact: - name: Jonathan D.A. Jewell - email: j.d.a.jewell@open.ac.uk - license: - name: PMPL-1.0-or-later - url: https://github.com/hyperpolymath/verisimdb/blob/main/LICENSE - -servers: - - url: http://localhost:8080 - description: Local development server - -tags: - - name: health - description: Health check and readiness probes - - name: octads - description: Octad entity CRUD operations - - name: search - description: Text, vector, and relationship search - - name: spatial - description: Geospatial search (radius, bounds, nearest) - - name: drift - description: Cross-modal drift detection and monitoring - - name: normalizer - description: Self-normalisation trigger and status - - name: provenance - description: Provenance chain management and verification - - name: vcl - description: VeriSim Consonance Language execution - -paths: - # --------------------------------------------------------------------------- - # Health - # --------------------------------------------------------------------------- - /api/v1/health: - get: - operationId: healthCheck - summary: Health check - description: | - Returns the health status of the VeriSimDB instance, including - storage backend connectivity and current drift health. - tags: [health] - responses: - "200": - description: Service is healthy. - content: - application/json: - schema: - type: object - required: [status, version] - properties: - status: - type: string - enum: [healthy, degraded, critical] - description: Overall health status. - version: - type: string - description: VeriSimDB version string. - uptime_seconds: - type: integer - description: Seconds since the service started. - drift_health: - type: string - enum: [healthy, warning, degraded, critical] - description: Current drift subsystem health. - "503": - description: Service is unhealthy. - content: - application/json: - schema: - $ref: "../json-schema/error.json" - - # --------------------------------------------------------------------------- - # Octad CRUD - # --------------------------------------------------------------------------- - /api/v1/octads: - post: - operationId: createOctad - summary: Create a new octad entity - description: | - Creates a new octad with the provided modality data. A UUID is - generated automatically. Only the modalities included in the - request body are populated; omitted modalities remain empty. - A provenance "created" event is recorded automatically. - tags: [octads] - requestBody: - required: true - content: - application/json: - schema: - $ref: "../json-schema/octad-input.json" - responses: - "201": - description: Octad created successfully. - content: - application/json: - schema: - $ref: "../json-schema/octad.json" - "400": - description: Validation error (e.g. tensor shape mismatch, invalid coordinates). - content: - application/json: - schema: - $ref: "../json-schema/error.json" - "500": - description: Internal server error. - content: - application/json: - schema: - $ref: "../json-schema/error.json" - - get: - operationId: listOctads - summary: List octad entities (paginated) - description: | - Returns a paginated list of octad entities. Results are ordered - by creation time (newest first). - tags: [octads] - parameters: - - name: limit - in: query - description: Maximum number of results (default 100, max 10000). - schema: - type: integer - minimum: 1 - maximum: 10000 - default: 100 - - name: offset - in: query - description: Number of results to skip for pagination. - schema: - type: integer - minimum: 0 - default: 0 - responses: - "200": - description: List of octad entities. - content: - application/json: - schema: - type: array - items: - $ref: "../json-schema/octad.json" - "500": - description: Internal server error. - content: - application/json: - schema: - $ref: "../json-schema/error.json" - - /api/v1/octads/{id}: - parameters: - - name: id - in: path - required: true - description: Octad entity UUID. - schema: - type: string - - get: - operationId: getOctad - summary: Get a octad entity by ID - description: | - Returns the full octad entity including all populated modality - data, status, version count, provenance chain length, and - spatial data. - tags: [octads] - responses: - "200": - description: Octad entity found. - content: - application/json: - schema: - $ref: "../json-schema/octad.json" - "404": - description: Entity not found. - content: - application/json: - schema: - $ref: "../json-schema/error.json" - "500": - description: Internal server error. - content: - application/json: - schema: - $ref: "../json-schema/error.json" - - put: - operationId: updateOctad - summary: Update an existing octad entity - description: | - Updates the specified modalities of an existing octad. Only - the modalities included in the request body are modified; - omitted modalities are left unchanged. A new temporal version - is created, and a provenance "modified" event is recorded. - tags: [octads] - requestBody: - required: true - content: - application/json: - schema: - $ref: "../json-schema/octad-input.json" - responses: - "200": - description: Octad updated successfully. - content: - application/json: - schema: - $ref: "../json-schema/octad.json" - "400": - description: Validation error. - content: - application/json: - schema: - $ref: "../json-schema/error.json" - "404": - description: Entity not found. - content: - application/json: - schema: - $ref: "../json-schema/error.json" - "500": - description: Internal server error. - content: - application/json: - schema: - $ref: "../json-schema/error.json" - - delete: - operationId: deleteOctad - summary: Delete a octad entity - description: | - Deletes the octad and all its modality data. A provenance - "deleted" event is recorded before the entity is removed. - tags: [octads] - responses: - "204": - description: Entity deleted successfully (no content). - "404": - description: Entity not found. - content: - application/json: - schema: - $ref: "../json-schema/error.json" - "500": - description: Internal server error. - content: - application/json: - schema: - $ref: "../json-schema/error.json" - - # --------------------------------------------------------------------------- - # Search - # --------------------------------------------------------------------------- - /api/v1/search/text: - get: - operationId: searchText - summary: Full-text search across octad documents - description: | - Searches the document modality using the Tantivy full-text - engine. Returns octads whose title or body match the query, - ranked by BM25 relevance. - tags: [search] - parameters: - - name: q - in: query - required: true - description: Text search query string. - schema: - type: string - - name: limit - in: query - description: Maximum results to return (default 100). - schema: - type: integer - minimum: 1 - maximum: 10000 - default: 100 - responses: - "200": - description: Search results. - content: - application/json: - schema: - type: array - items: - $ref: "../json-schema/octad.json" - "400": - description: Missing or invalid query parameter. - content: - application/json: - schema: - $ref: "../json-schema/error.json" - "500": - description: Internal server error. - content: - application/json: - schema: - $ref: "../json-schema/error.json" - - /api/v1/search/vector: - post: - operationId: searchVector - summary: Vector similarity search - description: | - Searches the vector modality using HNSW approximate nearest - neighbour search. Returns the k most similar octads ranked - by cosine similarity. - tags: [search] - requestBody: - required: true - content: - application/json: - schema: - type: object - required: [vector, k] - properties: - vector: - type: array - items: - type: number - description: Query embedding vector. Dimensionality must match the store's vector_dimension. - k: - type: integer - minimum: 1 - maximum: 10000 - description: Number of nearest neighbours to return. - responses: - "200": - description: Similar entities. - content: - application/json: - schema: - type: array - items: - $ref: "../json-schema/octad.json" - "400": - description: Invalid vector dimension or missing parameters. - content: - application/json: - schema: - $ref: "../json-schema/error.json" - "500": - description: Internal server error. - content: - application/json: - schema: - $ref: "../json-schema/error.json" - - /api/v1/search/related/{id}: - get: - operationId: searchRelated - summary: Find related entities via graph traversal - description: | - Traverses the graph modality to find entities related to the - given octad. Returns entities connected by outgoing edges. - tags: [search] - parameters: - - name: id - in: path - required: true - description: Source octad entity UUID. - schema: - type: string - - name: predicate - in: query - description: Filter by relationship type / edge label. If omitted, all relationships are returned. - schema: - type: string - responses: - "200": - description: Related entities. - content: - application/json: - schema: - type: array - items: - $ref: "../json-schema/octad.json" - "404": - description: Source entity not found. - content: - application/json: - schema: - $ref: "../json-schema/error.json" - "500": - description: Internal server error. - content: - application/json: - schema: - $ref: "../json-schema/error.json" - - # --------------------------------------------------------------------------- - # Spatial Search - # --------------------------------------------------------------------------- - /api/v1/spatial/search/radius: - post: - operationId: spatialSearchRadius - summary: Search entities within a radius - description: | - Finds octad entities within a given radius (km) of a centre - point. Results are sorted by distance ascending. - tags: [spatial] - requestBody: - required: true - content: - application/json: - schema: - type: object - required: [latitude, longitude, radius_km] - properties: - latitude: - type: number - minimum: -90.0 - maximum: 90.0 - description: Centre latitude (WGS84). - longitude: - type: number - minimum: -180.0 - maximum: 180.0 - description: Centre longitude (WGS84). - radius_km: - type: number - minimum: 0.0 - description: Search radius in kilometres. - limit: - type: integer - minimum: 1 - maximum: 10000 - default: 100 - description: Maximum results. - responses: - "200": - description: Entities within the radius. - content: - application/json: - schema: - type: array - items: - type: object - required: [entity_id, distance_km, spatial_data] - properties: - entity_id: - type: string - distance_km: - type: number - spatial_data: - type: object - "400": - description: Invalid coordinates or radius. - content: - application/json: - schema: - $ref: "../json-schema/error.json" - "500": - description: Internal server error. - content: - application/json: - schema: - $ref: "../json-schema/error.json" - - /api/v1/spatial/search/bounds: - post: - operationId: spatialSearchBounds - summary: Search entities within a bounding box - description: | - Finds octad entities within a geographic bounding box defined - by south-west and north-east corners (WGS84). - tags: [spatial] - requestBody: - required: true - content: - application/json: - schema: - type: object - required: [min_lat, min_lon, max_lat, max_lon] - properties: - min_lat: - type: number - minimum: -90.0 - maximum: 90.0 - description: Southern latitude boundary. - min_lon: - type: number - minimum: -180.0 - maximum: 180.0 - description: Western longitude boundary. - max_lat: - type: number - minimum: -90.0 - maximum: 90.0 - description: Northern latitude boundary. - max_lon: - type: number - minimum: -180.0 - maximum: 180.0 - description: Eastern longitude boundary. - limit: - type: integer - minimum: 1 - maximum: 10000 - default: 100 - description: Maximum results. - responses: - "200": - description: Entities within the bounding box. - content: - application/json: - schema: - type: array - items: - type: object - required: [entity_id, distance_km, spatial_data] - properties: - entity_id: - type: string - distance_km: - type: number - spatial_data: - type: object - "400": - description: Invalid bounding box. - content: - application/json: - schema: - $ref: "../json-schema/error.json" - "500": - description: Internal server error. - content: - application/json: - schema: - $ref: "../json-schema/error.json" - - /api/v1/spatial/search/nearest: - post: - operationId: spatialSearchNearest - summary: Find k nearest entities to a point - description: | - Finds the k nearest octad entities to a given geographic point - using the Haversine formula for distance computation. - tags: [spatial] - requestBody: - required: true - content: - application/json: - schema: - type: object - required: [latitude, longitude, k] - properties: - latitude: - type: number - minimum: -90.0 - maximum: 90.0 - description: Query point latitude (WGS84). - longitude: - type: number - minimum: -180.0 - maximum: 180.0 - description: Query point longitude (WGS84). - k: - type: integer - minimum: 1 - maximum: 10000 - description: Number of nearest neighbours to return. - responses: - "200": - description: Nearest entities. - content: - application/json: - schema: - type: array - items: - type: object - required: [entity_id, distance_km, spatial_data] - properties: - entity_id: - type: string - distance_km: - type: number - spatial_data: - type: object - "400": - description: Invalid coordinates. - content: - application/json: - schema: - $ref: "../json-schema/error.json" - "500": - description: Internal server error. - content: - application/json: - schema: - $ref: "../json-schema/error.json" - - # --------------------------------------------------------------------------- - # Drift Detection - # --------------------------------------------------------------------------- - /api/v1/drift/entity/{id}: - get: - operationId: getDriftScore - summary: Get drift score for an entity - description: | - Returns the current drift measurement for a specific octad - entity, including per-modality scores and whether normalisation - is needed. - tags: [drift] - parameters: - - name: id - in: path - required: true - description: Octad entity UUID. - schema: - type: string - responses: - "200": - description: Drift score for the entity. - content: - application/json: - schema: - $ref: "../json-schema/drift-score.json" - "404": - description: Entity not found. - content: - application/json: - schema: - $ref: "../json-schema/error.json" - "500": - description: Internal server error. - content: - application/json: - schema: - $ref: "../json-schema/error.json" - - /api/v1/drift/status: - get: - operationId: getDriftStatus - summary: Get overall drift status - description: | - Returns the aggregate drift health status for the entire - VeriSimDB instance, including the worst drift type and score. - tags: [drift] - responses: - "200": - description: Overall drift health status. - content: - application/json: - schema: - type: object - required: [status, worst_drift_type, worst_score, checked_at] - properties: - status: - type: string - enum: [healthy, warning, degraded, critical] - description: Overall drift health status. - worst_drift_type: - type: string - description: The drift type with the highest current score. - worst_score: - type: number - minimum: 0.0 - maximum: 1.0 - description: The highest current drift score across all types. - checked_at: - type: string - format: date-time - description: When the health check was performed. - "500": - description: Internal server error. - content: - application/json: - schema: - $ref: "../json-schema/error.json" - - # --------------------------------------------------------------------------- - # Normalizer - # --------------------------------------------------------------------------- - /api/v1/normalizer/trigger/{id}: - post: - operationId: triggerNormalization - summary: Trigger normalisation for an entity - description: | - Manually triggers the self-normalisation process for a specific - octad entity. The normaliser identifies the most authoritative - modality and regenerates drifted modalities from it, validating - consistency and updating all modalities atomically. - tags: [normalizer] - parameters: - - name: id - in: path - required: true - description: Octad entity UUID to normalise. - schema: - type: string - responses: - "200": - description: Normalisation completed. - content: - application/json: - schema: - type: object - required: [entity_id, status, modalities_repaired] - properties: - entity_id: - type: string - description: The normalised entity's ID. - status: - type: string - enum: [completed, no_drift, failed] - description: Normalisation outcome. - modalities_repaired: - type: array - items: - $ref: "../json-schema/modality.json" - description: Which modalities were repaired. - drift_score_before: - type: number - description: Overall drift score before normalisation. - drift_score_after: - type: number - description: Overall drift score after normalisation. - "404": - description: Entity not found. - content: - application/json: - schema: - $ref: "../json-schema/error.json" - "500": - description: Internal server error. - content: - application/json: - schema: - $ref: "../json-schema/error.json" - - /api/v1/normalizer/status: - get: - operationId: getNormalizerStatus - summary: Get normaliser status - description: | - Returns the current status of the normalisation subsystem, - including queue depth, active normalisations, and recent - normalisation history. - tags: [normalizer] - responses: - "200": - description: Normaliser status. - content: - application/json: - schema: - type: object - required: [active, queue_depth] - properties: - active: - type: boolean - description: Whether the normaliser is currently running. - queue_depth: - type: integer - minimum: 0 - description: Number of entities awaiting normalisation. - entities_normalised_total: - type: integer - minimum: 0 - description: Total entities normalised since startup. - last_normalisation_at: - type: string - format: date-time - description: Timestamp of the most recent normalisation. - "500": - description: Internal server error. - content: - application/json: - schema: - $ref: "../json-schema/error.json" - - # --------------------------------------------------------------------------- - # Provenance - # --------------------------------------------------------------------------- - /api/v1/provenance/{id}: - get: - operationId: getProvenanceChain - summary: Get provenance chain for an entity - description: | - Returns the full provenance chain for a octad entity -- an - ordered list of provenance events linked by SHA-256 hashes. - Records are returned oldest-first. - tags: [provenance] - parameters: - - name: id - in: path - required: true - description: Octad entity UUID. - schema: - type: string - responses: - "200": - description: Provenance chain. - content: - application/json: - schema: - type: object - required: [entity_id, records] - properties: - entity_id: - type: string - records: - type: array - items: - $ref: "../json-schema/provenance-event.json" - chain_length: - type: integer - minimum: 0 - "404": - description: Entity or provenance chain not found. - content: - application/json: - schema: - $ref: "../json-schema/error.json" - "500": - description: Internal server error. - content: - application/json: - schema: - $ref: "../json-schema/error.json" - - /api/v1/provenance/{id}/record: - post: - operationId: recordProvenanceEvent - summary: Record a provenance event - description: | - Appends a new provenance event to the entity's hash chain. - The parent_hash is computed automatically from the previous - record's content_hash. - tags: [provenance] - parameters: - - name: id - in: path - required: true - description: Octad entity UUID. - schema: - type: string - requestBody: - required: true - content: - application/json: - schema: - type: object - required: [event_type, actor, description] - properties: - event_type: - type: string - enum: [created, modified, imported, normalized, federated, deleted] - description: Type of provenance event. - actor: - type: string - description: Who or what caused this event. - source: - type: string - description: Optional source identifier. - description: - type: string - description: Human-readable description. - responses: - "201": - description: Provenance event recorded. - content: - application/json: - schema: - $ref: "../json-schema/provenance-event.json" - "404": - description: Entity not found. - content: - application/json: - schema: - $ref: "../json-schema/error.json" - "500": - description: Internal server error. - content: - application/json: - schema: - $ref: "../json-schema/error.json" - - /api/v1/provenance/{id}/verify: - get: - operationId: verifyProvenanceChain - summary: Verify provenance chain integrity - description: | - Verifies the SHA-256 hash chain integrity for an entity's - provenance records. Each record's parent_hash is checked - against the content_hash of the preceding record, and each - record's content_hash is recomputed to detect tampering. - tags: [provenance] - parameters: - - name: id - in: path - required: true - description: Octad entity UUID. - schema: - type: string - responses: - "200": - description: Verification result. - content: - application/json: - schema: - type: object - required: [entity_id, valid, chain_length] - properties: - entity_id: - type: string - valid: - type: boolean - description: Whether the hash chain is intact. - chain_length: - type: integer - minimum: 0 - error: - type: string - description: Description of the integrity violation, if any. - "404": - description: Entity or provenance chain not found. - content: - application/json: - schema: - $ref: "../json-schema/error.json" - "500": - description: Internal server error. - content: - application/json: - schema: - $ref: "../json-schema/error.json" - - # --------------------------------------------------------------------------- - # VCL - # --------------------------------------------------------------------------- - /api/v1/vcl/execute: - post: - operationId: executeVcl - summary: Execute a VCL query - description: | - Executes a VeriSim Consonance Language (VCL) query against the - database. VCL is VeriSimDB's native query language, providing - cross-modal querying with optional proof requirements. - tags: [vcl] - requestBody: - required: true - content: - application/json: - schema: - type: object - required: [query] - properties: - query: - type: string - description: | - VCL query string. Example: - `FIND entities WHERE document CONTAINS "neural" AND vector SIMILAR TO [0.1, 0.2, ...] LIMIT 10` - explain: - type: boolean - default: false - description: If true, return the query execution plan instead of results. - timeout_ms: - type: integer - minimum: 100 - maximum: 300000 - default: 30000 - description: Query timeout in milliseconds. - responses: - "200": - description: Query results or execution plan. - content: - application/json: - schema: - type: object - required: [results, execution_time_ms] - properties: - results: - type: array - items: - $ref: "../json-schema/octad.json" - description: Matching octad entities. - execution_time_ms: - type: integer - description: Time taken to execute the query. - total_count: - type: integer - description: Total matching entities (may exceed returned results). - explain_plan: - type: object - description: Query execution plan (only when explain=true). - "400": - description: VCL syntax error or invalid query. - content: - application/json: - schema: - $ref: "../json-schema/error.json" - "408": - description: Query timed out. - content: - application/json: - schema: - $ref: "../json-schema/error.json" - "500": - description: Internal server error. - content: - application/json: - schema: - $ref: "../json-schema/error.json" diff --git a/verisimdb/connectors/shared/proto/verisim_federation.proto b/verisimdb/connectors/shared/proto/verisim_federation.proto deleted file mode 100644 index 3b44502a..00000000 --- a/verisimdb/connectors/shared/proto/verisim_federation.proto +++ /dev/null @@ -1,414 +0,0 @@ -// SPDX-License-Identifier: MPL-2.0 -// Copyright (c) 2026 Jonathan D.A. Jewell (hyperpolymath) -// -// Protobuf definitions for VeriSimDB federation protocol. -// -// This file defines the gRPC service and message types used for -// inter-instance federation. VeriSimDB instances form a federated -// mesh where each peer can register, discover, query, and monitor -// other peers while maintaining local autonomy. -// -// The federation protocol is drift-aware: query results are annotated -// with drift status, and the coordinator can apply drift policies -// (strict, repair, tolerate, latest) to filter or annotate results. - -syntax = "proto3"; - -package verisim.federation; - -option java_package = "db.verisim.federation"; -option java_outer_classname = "VeriSimFederationProto"; -option go_package = "verisim/federation"; - -// --------------------------------------------------------------------------- -// Enums -// --------------------------------------------------------------------------- - -// Modality identifies one of the 8 octad modality dimensions. -// Each octad entity can have data in any combination of these modalities. -enum Modality { - // Unspecified / unknown modality (protobuf default). - MODALITY_UNSPECIFIED = 0; - - // RDF triples and property-graph edges. - MODALITY_GRAPH = 1; - - // Dense embedding vector for similarity search. - MODALITY_VECTOR = 2; - - // Multi-dimensional numeric tensor (shape + data). - MODALITY_TENSOR = 3; - - // Ontological type annotations and property maps. - MODALITY_SEMANTIC = 4; - - // Full-text searchable document content. - MODALITY_DOCUMENT = 5; - - // Version history and time-travel queries. - MODALITY_TEMPORAL = 6; - - // Origin tracking, transformation chain, actor trail (hash chain). - MODALITY_PROVENANCE = 7; - - // Geospatial coordinates, geometry types, proximity queries. - MODALITY_SPATIAL = 8; -} - -// DriftPolicy controls how the federation coordinator handles -// drift when aggregating results from peer stores. -enum DriftPolicy { - // Unspecified policy (defaults to TOLERATE). - DRIFT_POLICY_UNSPECIFIED = 0; - - // Only return results from stores with drift below threshold. - DRIFT_POLICY_STRICT = 1; - - // Return results and trigger normalisation on drifted stores. - DRIFT_POLICY_REPAIR = 2; - - // Return all results, annotating drifted ones (default). - DRIFT_POLICY_TOLERATE = 3; - - // Return only the most recent version from each store. - DRIFT_POLICY_LATEST = 4; -} - -// --------------------------------------------------------------------------- -// Authentication -// --------------------------------------------------------------------------- - -// BasicAuth credentials for adapter connections to external stores. -message BasicAuth { - // Username for authentication. - string username = 1; - - // Password (transmitted over TLS; stored hashed at rest). - string password = 2; -} - -// BearerToken for OAuth2 / JWT authentication. -message BearerToken { - // The bearer token string. - string token = 1; -} - -// ApiKey for key-based authentication. -message ApiKey { - // The API key string. - string key = 1; - - // Header name for the API key (e.g. "X-API-Key", "Authorization"). - // Defaults to "X-API-Key" if empty. - string header_name = 2; -} - -// AuthConfig wraps the supported authentication methods. -// Exactly one method should be set. -message AuthConfig { - oneof method { - // HTTP Basic authentication. - BasicAuth basic_auth = 1; - - // Bearer token (OAuth2 / JWT). - BearerToken bearer_token = 2; - - // API key in a custom header. - ApiKey api_key = 3; - } -} - -// --------------------------------------------------------------------------- -// Peer Registration -// --------------------------------------------------------------------------- - -// AdapterConfig describes how to connect to a specific external store -// via a federation adapter. -message AdapterConfig { - // Adapter type identifier (e.g. "postgres", "neo4j", "elasticsearch", - // "qdrant", "mongodb", "redis", "duckdb", "clickhouse", "weaviate", - // "milvus", "surrealdb", "influxdb", "tigergraph", "pinecone"). - string adapter_type = 1; - - // Adapter-specific connection parameters (e.g. "host", "port", - // "database", "index_name", "collection", "timeout_ms"). - map params = 2; - - // Authentication configuration for the adapter's connection to - // the external store. - AuthConfig auth = 3; -} - -// PeerInfo describes a peer store registering with the federation. -message PeerInfo { - // Unique store identifier within the federation (e.g. "store-1", - // "/universities/oxford"). Must be alphanumeric plus dash, underscore, - // and slash. Maximum 128 characters. - string store_id = 1; - - // HTTP/gRPC endpoint URL for the peer's VeriSimDB API - // (e.g. "https://store-2.verisimdb.example.com:8080/api/v1"). - string endpoint = 2; - - // Which modalities this peer supports. - repeated Modality supported_modalities = 3; - - // Adapter configuration for connecting to this peer's backing store. - AdapterConfig config = 4; - - // Pre-shared key for federation authentication (sent once at - // registration, stored as SHA-256 hash on the coordinator). - string secret = 5; -} - -// RegisterResponse returned after successful peer registration. -message RegisterResponse { - // Whether the registration was successful. - bool success = 1; - - // The coordinator's own store_id (for the peer to record). - string coordinator_store_id = 2; - - // Human-readable message (e.g. "Peer registered successfully"). - string message = 3; -} - -// --------------------------------------------------------------------------- -// Federation Queries -// --------------------------------------------------------------------------- - -// TextQuery searches the document modality using full-text matching. -message TextQuery { - // The text search string. - string query = 1; -} - -// VectorQuery searches the vector modality using similarity search. -message VectorQuery { - // The query embedding vector (float32). - repeated float vector = 1; - - // Number of nearest neighbours to return. - int32 k = 2; -} - -// GraphQuery searches the graph modality using a pattern. -message GraphQuery { - // Graph traversal pattern. Syntax is adapter-dependent (SPARQL, - // Cypher, simplified triple pattern, etc.). - string pattern = 1; -} - -// SpatialQuery searches the spatial modality. -message SpatialQuery { - // Centre latitude (WGS84). - double latitude = 1; - - // Centre longitude (WGS84). - double longitude = 2; - - // Search radius in kilometres (for radius search). - double radius_km = 3; - - // Alternatively, a bounding box (for bounds search). - BoundingBox bounds = 4; - - // Number of nearest neighbours (for k-nearest search). - int32 k = 5; -} - -// BoundingBox for spatial bounding-box queries. -message BoundingBox { - // Southern latitude boundary. - double min_lat = 1; - - // Western longitude boundary. - double min_lon = 2; - - // Northern latitude boundary. - double max_lat = 3; - - // Eastern longitude boundary. - double max_lon = 4; -} - -// TemporalQuery filters by time range. -message TemporalQuery { - // Start of the time range (RFC 3339). - string start = 1; - - // End of the time range (RFC 3339). - string end = 2; -} - -// FederationQuery is the primary query message sent to the federation -// coordinator. It specifies which modalities to search, the query -// payload, result limits, and drift handling policy. -message FederationQuery { - // Which modalities to include in the query. - repeated Modality modalities = 1; - - // Maximum results to return per peer (capped at 10000). - int32 limit = 2; - - // Drift handling policy (defaults to TOLERATE). - DriftPolicy drift_policy = 3; - - // The query payload. Exactly one should be set. - oneof query { - // Full-text search. - TextQuery text_query = 10; - - // Vector similarity search. - VectorQuery vector_query = 11; - - // Graph pattern traversal. - GraphQuery graph_query = 12; - - // Spatial proximity search. - SpatialQuery spatial_query = 13; - - // Temporal range filter. - TemporalQuery temporal_query = 14; - } - - // Optional glob pattern to match peer store IDs - // (e.g. "*", "/universities/*", "store-1"). - string store_pattern = 20; -} - -// FederationResult is a single result from a federated query. -// Results are streamed from the coordinator to the client as they -// arrive from peer stores. -message FederationResult { - // Which peer store provided this result. - string source_store = 1; - - // Octad entity identifier within the source store. - string octad_id = 2; - - // Relevance score normalised to [0.0, 1.0]. - double score = 3; - - // Whether the source store has drift issues for this entity. - bool drifted = 4; - - // Serialised result data (JSON-encoded octad or modality-specific payload). - bytes data = 5; - - // Time in milliseconds the source store took to respond. - int64 response_time_ms = 6; -} - -// --------------------------------------------------------------------------- -// Health Check -// --------------------------------------------------------------------------- - -// HealthRequest asks a peer for its current health status. -message HealthRequest { - // The requesting store's identifier (for logging and auditing). - string store_id = 1; -} - -// HealthResponse reports a peer's health status. -message HealthResponse { - // Whether the peer considers itself healthy. - bool healthy = 1; - - // Round-trip latency of the health check in milliseconds. - int64 latency_ms = 2; - - // Per-modality drift scores (modality name -> score). - map drift_scores = 3; - - // Overall drift health status. - string drift_status = 4; - - // Number of octad entities stored by this peer. - int64 entity_count = 5; -} - -// --------------------------------------------------------------------------- -// Peer Management -// --------------------------------------------------------------------------- - -// ListPeersRequest asks the coordinator for all registered peers. -message ListPeersRequest { - // Optional modality filter: only return peers supporting this modality. - Modality modality_filter = 1; -} - -// PeerSummary provides a summary of a registered peer. -message PeerSummary { - // Peer's store identifier. - string store_id = 1; - - // Peer's endpoint URL. - string endpoint = 2; - - // Modalities the peer supports. - repeated Modality supported_modalities = 3; - - // Trust level (0.0 - 1.0). - double trust_level = 4; - - // Last time the peer was seen (RFC 3339 timestamp). - string last_seen = 5; - - // Average response time in milliseconds. - int64 avg_response_time_ms = 6; -} - -// ListPeersResponse returns all registered peers. -message ListPeersResponse { - // List of registered peer summaries. - repeated PeerSummary peers = 1; -} - -// DeregisterRequest removes a peer from the federation. -message DeregisterRequest { - // Store identifier of the peer to remove. - string store_id = 1; -} - -// DeregisterResponse confirms peer removal. -message DeregisterResponse { - // Whether the deregistration was successful. - bool success = 1; - - // Human-readable message. - string message = 2; -} - -// --------------------------------------------------------------------------- -// Service Definition -// --------------------------------------------------------------------------- - -// FederationService is the gRPC service for VeriSimDB federation. -// -// It provides peer registration, discovery, federated querying with -// streaming results, and health monitoring. All RPCs require TLS -// and federation PSK authentication (via the X-Federation-PSK header -// or the PeerInfo.secret field at registration time). -service FederationService { - // RegisterPeer registers a new peer store with the federation - // coordinator. Requires a valid pre-shared key matching the - // coordinator's VERISIM_FEDERATION_KEYS configuration. - rpc RegisterPeer(PeerInfo) returns (RegisterResponse); - - // Query executes a federated query across matching peer stores. - // Results are streamed as they arrive from each peer, allowing - // the client to process results incrementally. - rpc Query(FederationQuery) returns (stream FederationResult); - - // HealthCheck performs a health probe against a specific peer. - // Returns the peer's self-reported health status and drift scores. - rpc HealthCheck(HealthRequest) returns (HealthResponse); - - // ListPeers returns all peers registered with the coordinator, - // optionally filtered by supported modality. - rpc ListPeers(ListPeersRequest) returns (ListPeersResponse); - - // DeregisterPeer removes a peer from the federation. - rpc DeregisterPeer(DeregisterRequest) returns (DeregisterResponse); -} diff --git a/verisimdb/connectors/test-infra/.gatekeeper.yaml b/verisimdb/connectors/test-infra/.gatekeeper.yaml deleted file mode 100644 index e86a57c8..00000000 --- a/verisimdb/connectors/test-infra/.gatekeeper.yaml +++ /dev/null @@ -1,68 +0,0 @@ -# SPDX-License-Identifier: MPL-2.0 -# -# Svalinn gatekeeper policy for VeriSimDB test infrastructure -# -# Permissive policy for local integration testing. No authentication -# required, no rate limiting, all origins allowed. This policy is -# intentionally relaxed — do NOT use in production. -# -# See: stapeln/container-stack/svalinn/ - -version: "1.0" - -# Authentication requirements — NONE for testing -auth: - # All endpoints are public (no auth required) - public: - - path: "/*" - methods: ["GET", "POST", "PUT", "DELETE", "PATCH", "OPTIONS", "HEAD"] - - # No authenticated endpoints — everything is open - authenticated: [] - - # No federation PSK requirements - federation: [] - -# Rate limiting — DISABLED for testing -rate_limits: - global: - requests_per_second: 0 - burst: 0 - -# Container trust policy — RELAXED for testing -trust: - # No trusted signers required for test images - trusted_signers: [] - - # No attestations required - required_attestations: [] - - # Accept unsigned images (test images are built with --no-sign) - reject_unsigned: false - -# Request validation — RELAXED for testing -validation: - # Large body size for bulk test data uploads - max_body_size: "128MB" - - # Do not reject NaN/Inf (useful for testing edge cases) - reject_nan_inf: false - - # Large vector dimension limit for testing - max_vector_dimension: 16384 - - # Large result limit for testing - max_result_limit: 100000 - -# CORS policy — WIDE OPEN for testing -cors: - allow_origins: ["*"] - allow_methods: ["GET", "POST", "PUT", "DELETE", "PATCH", "OPTIONS", "HEAD"] - allow_headers: ["*"] - max_age: 86400 - -# Logging — verbose for testing -logging: - format: "text" - level: "debug" - audit_paths: [] diff --git a/verisimdb/connectors/test-infra/0-AI-MANIFEST.a2ml b/verisimdb/connectors/test-infra/0-AI-MANIFEST.a2ml deleted file mode 100644 index 27105318..00000000 --- a/verisimdb/connectors/test-infra/0-AI-MANIFEST.a2ml +++ /dev/null @@ -1,157 +0,0 @@ -; SPDX-License-Identifier: MPL-2.0 -; -; 0-AI-MANIFEST.a2ml — AI-readable manifest for VeriSimDB test infrastructure -; -; This file is the universal entry point for ALL AI agents working with -; the VeriSimDB integration test database stack. -; -; Author: Jonathan D.A. Jewell - -(ai-manifest - (version "1.0") - (name "verisimdb-test-infra") - (purpose "Integration test database stack for VeriSimDB federation adapters") - - ;; ========================================================================= - ;; Canonical file locations - ;; ========================================================================= - - (canonical-locations - (compose-file "connectors/test-infra/compose.toml") - (containerfiles "connectors/test-infra/images/Containerfile.*") - (seed-scripts "connectors/test-infra/seed/") - (build-script "connectors/test-infra/ct-build.sh") - (cerro-torre-manifest "connectors/test-infra/manifest.toml") - (gatekeeper-policy "connectors/test-infra/.gatekeeper.yaml") - (vordr-config "connectors/test-infra/vordr.toml") - (k9-deploy "connectors/test-infra/deploy.k9.ncl")) - - ;; ========================================================================= - ;; Critical invariants — rules that must NEVER be violated - ;; ========================================================================= - - (invariants - (rule "All custom Containerfiles use cgr.dev/chainguard/wolfi-base:latest as base") - (rule "All containers run as non-root users") - (rule "Test images are NEVER signed (--no-sign flag in ct-build.sh)") - (rule "No production credentials in seed scripts — test-only passwords") - (rule "DuckDB and SQLite are NOT in this stack — they are embedded databases tested via unit tests") - (rule "MinIO API port is remapped to 9002 to avoid conflict with ClickHouse native port 9000") - (rule "Use Containerfile NOT Dockerfile") - (rule "Use Podman NOT Docker")) - - ;; ========================================================================= - ;; Available databases and ports - ;; ========================================================================= - - (services - (service - (name "mongodb") - (image "cgr.dev/chainguard/mongodb:latest") - (ports (port 27017 "MongoDB wire protocol")) - (purpose "Document store federation adapter") - (notes "Replica set rs0 for change streams. Seeds automatically on first start.")) - - (service - (name "redis-stack") - (image "verisimdb-test-redis-stack:latest (custom build)") - (ports (port 6379 "Redis protocol")) - (purpose "In-memory store with RediSearch + RedisJSON + RedisTimeSeries") - (notes "Modules compiled from source on wolfi-base.")) - - (service - (name "neo4j") - (image "verisimdb-test-neo4j:latest (custom build)") - (ports - (port 7474 "HTTP API + browser") - (port 7687 "Bolt protocol")) - (purpose "Graph database for graph modality federation") - (notes "Community edition with APOC plugin. Auth disabled for testing.")) - - (service - (name "clickhouse") - (image "verisimdb-test-clickhouse:latest (custom build)") - (ports - (port 8123 "HTTP interface") - (port 9000 "Native TCP protocol")) - (purpose "Columnar analytics for aggregate drift queries") - (notes "Static binary build. Default user has full access management.")) - - (service - (name "surrealdb") - (image "verisimdb-test-surrealdb:latest (custom build)") - (ports (port 8000 "HTTP + WebSocket API")) - (purpose "Multi-model database (document + graph)") - (notes "Runs in memory mode for tests — no persistence overhead.")) - - (service - (name "influxdb") - (image "verisimdb-test-influxdb:latest (custom build)") - (ports (port 8086 "HTTP API")) - (purpose "Time-series database for temporal modality federation") - (notes "Pre-configured org verisimdb, bucket metrics. Token: verisim-test-token-do-not-use-in-production")) - - (service - (name "minio") - (image "cgr.dev/chainguard/minio:latest") - (ports - (port 9002 "S3 API (remapped from 9000)") - (port 9001 "Web console")) - (purpose "S3-compatible object storage for binary blob modalities") - (notes "Credentials: verisim / verisim-test-password. Console at http://localhost:9001"))) - - ;; ========================================================================= - ;; How to start and stop the stack - ;; ========================================================================= - - (operations - (start - (step 1 "Build custom images" "cd connectors/test-infra && ./ct-build.sh") - (step 2 "Start stack" "cd connectors/test-infra && selur-compose up --detach") - (step 3 "Wait for health" "Wait ~30 seconds for all services to pass health checks") - (step 4 "Seed data" "Run seed scripts (see below)")) - - (stop - (step 1 "Stop stack" "cd connectors/test-infra && selur-compose down") - (step 2 "Remove volumes" "cd connectors/test-infra && selur-compose down --volumes")) - - (fallback - (note "If selur-compose is not installed, use podman-compose as fallback"))) - - ;; ========================================================================= - ;; How to run seed scripts - ;; ========================================================================= - - (seeding - (note "MongoDB seeds automatically via /docker-entrypoint-initdb.d/init.js") - (script "redis-init.sh" "shell" "cd connectors/test-infra/seed && ./redis-init.sh") - (script "neo4j-init.cypher" "cypher" "cat connectors/test-infra/seed/neo4j-init.cypher | cypher-shell -a bolt://localhost:7687") - (script "clickhouse-init.sql" "sql" "clickhouse-client --multiquery < connectors/test-infra/seed/clickhouse-init.sql") - (script "surrealdb-init.surql" "surql" "cat connectors/test-infra/seed/surrealdb-init.surql | surreal sql --endpoint http://localhost:8000 --ns verisimdb --db test") - (script "influxdb-init.sh" "shell" "cd connectors/test-infra/seed && ./influxdb-init.sh") - (script "minio-init.sh" "shell" "cd connectors/test-infra/seed && ./minio-init.sh")) - - ;; ========================================================================= - ;; How to run integration tests - ;; ========================================================================= - - (testing - (rust "cd rust-core && cargo test --test integration -- --test-threads=1") - (elixir "cd elixir-orchestration && mix test test/integration") - (note "Integration tests expect the full stack to be running and seeded") - (note "Use --test-threads=1 for Rust to avoid port contention")) - - ;; ========================================================================= - ;; Dependencies - ;; ========================================================================= - - (dependencies - (required "Podman" "Container runtime — 4.0+ recommended") - (required "selur-compose OR podman-compose" "Container orchestration") - (optional "ct" "Cerro-torre CLI for .ctp bundle packaging") - (optional "redis-cli" "For running redis-init.sh seed script") - (optional "cypher-shell" "For running neo4j-init.cypher seed script") - (optional "clickhouse-client" "For running clickhouse-init.sql seed script") - (optional "surreal" "For running surrealdb-init.surql seed script") - (optional "influx" "For running influxdb-init.sh seed script") - (optional "mc" "MinIO client for running minio-init.sh seed script"))) diff --git a/verisimdb/connectors/test-infra/README.adoc b/verisimdb/connectors/test-infra/README.adoc deleted file mode 100644 index 16faed78..00000000 --- a/verisimdb/connectors/test-infra/README.adoc +++ /dev/null @@ -1,344 +0,0 @@ -// SPDX-License-Identifier: CC-BY-SA-4.0 -// -// Author: Jonathan D.A. Jewell - -= VeriSimDB Test Infrastructure -:toc: macro -:toc-title: Contents -:toclevels: 3 -:icons: font -:source-highlighter: rouge - -Integration test database stack for VeriSimDB's 10 federation adapters. - -toc::[] - -== Overview - -VeriSimDB federates across 10 external database systems. This test infrastructure -provides a containerised stack of 7 databases for integration testing. The remaining -2 adapters (DuckDB, SQLite) are embedded databases tested via unit tests. - -.Database services -[cols="1,2,1,2",options="header"] -|=== -| Service | Image | Ports | Purpose - -| MongoDB -| `cgr.dev/chainguard/mongodb:latest` -| 27017 -| Document store federation (replica set for change streams) - -| Redis Stack -| Custom (wolfi-base + modules) -| 6379 -| In-memory store: RediSearch, RedisJSON, RedisTimeSeries - -| Neo4j -| Custom (wolfi-base + Neo4j CE) -| 7474, 7687 -| Graph database with APOC plugin - -| ClickHouse -| Custom (wolfi-base + static binary) -| 8123, 9000 -| Columnar analytics for aggregate drift queries - -| SurrealDB -| Custom (wolfi-base + binary) -| 8000 -| Multi-model database (document + graph) - -| InfluxDB 2 -| Custom (wolfi-base + binary) -| 8086 -| Time-series for temporal modality federation - -| MinIO -| `cgr.dev/chainguard/minio:latest` -| 9002, 9001 -| S3-compatible object storage -|=== - -NOTE: MinIO's S3 API is mapped to port **9002** (not 9000) to avoid conflict -with ClickHouse's native TCP port on 9000. - -== Prerequisites - -.Required -* https://podman.io/[Podman] 4.0+ (container runtime) -* https://github.com/hyperpolymath/stapeln[selur-compose] or - https://github.com/containers/podman-compose[podman-compose] (container orchestration) - -.Optional (for seeding and verification) -* `redis-cli` -- seed Redis Stack -* `cypher-shell` -- seed Neo4j -* `clickhouse-client` -- seed ClickHouse -* `surreal` CLI -- seed SurrealDB -* `influx` CLI -- seed InfluxDB -* `mc` (MinIO Client) -- seed MinIO -* `ct` (cerro-torre CLI) -- pack `.ctp` bundles - -== Quick Start - -=== 1. Build custom images - -[source,bash] ----- -cd connectors/test-infra -./ct-build.sh ----- - -Or build in parallel for faster results: - -[source,bash] ----- -./ct-build.sh --parallel ----- - -=== 2. Start the stack - -[source,bash] ----- -# Preferred: selur-compose -selur-compose up --detach - -# Fallback: podman-compose -podman-compose up --detach ----- - -=== 3. Wait for health checks - -All services include health checks. Wait approximately 30 seconds for all -services to become healthy, then verify: - -[source,bash] ----- -selur-compose ps # or: podman-compose ps ----- - -=== 4. Seed test data - -MongoDB seeds automatically on first startup via `/docker-entrypoint-initdb.d/init.js`. -For the other databases: - -[source,bash] ----- -# Redis Stack -cd seed && ./redis-init.sh && cd .. - -# Neo4j (requires cypher-shell) -cat seed/neo4j-init.cypher | cypher-shell -a bolt://localhost:7687 - -# ClickHouse (requires clickhouse-client) -clickhouse-client --multiquery < seed/clickhouse-init.sql - -# SurrealDB (requires surreal CLI) -cat seed/surrealdb-init.surql | surreal sql \ - --endpoint http://localhost:8000 --ns verisimdb --db test - -# InfluxDB 2 -cd seed && ./influxdb-init.sh && cd .. - -# MinIO (requires mc) -cd seed && ./minio-init.sh && cd .. ----- - -=== 5. Run integration tests - -[source,bash] ----- -# Rust integration tests (sequential to avoid port contention) -cd ../../rust-core -cargo test --test integration -- --test-threads=1 - -# Elixir integration tests -cd ../../elixir-orchestration -mix test test/integration ----- - -=== 6. Tear down - -[source,bash] ----- -# Stop services (preserve volumes) -selur-compose down - -# Stop services and remove volumes -selur-compose down --volumes ----- - -== Architecture - -=== Network Topology - -All 7 services communicate over a single network. The default network driver -is `selur` (zero-copy IPC for same-host communication) with automatic fallback -to the standard `bridge` driver when selur is unavailable. - -[source] ----- - selur network (bridge fallback) - +--------------------------------------------------+ - | | - +-------+-------+ +-----------+ +----------+ +---------+--------+ - | MongoDB :27017| | Redis | | Neo4j | | ClickHouse | - | (rs0) | | :6379 | | :7474 | | :8123 (HTTP) | - +---------------+ +-----------+ | :7687 | | :9000 (native) | - +----------+ +------------------+ - +---------------+ +-----------+ +-----------+ - | SurrealDB | | InfluxDB | | MinIO | - | :8000 | | :8086 | | :9002 (S3)| - +---------------+ +-----------+ | :9001 (UI)| - +-----------+ ----- - -=== Volumes - -Each service has its own persistent volume for data storage. Volumes survive -container restarts but are removed with `selur-compose down --volumes`. - -[cols="1,2",options="header"] -|=== -| Volume | Purpose -| `mongodb-data` | MongoDB data directory -| `redis-data` | Redis RDB/AOF persistence -| `neo4j-data` | Neo4j graph data -| `neo4j-logs` | Neo4j server logs -| `clickhouse-data` | ClickHouse MergeTree data -| `influxdb-data` | InfluxDB 2 storage engine -| `minio-data` | MinIO object storage -|=== - -=== Custom Containerfiles - -All custom images use multi-stage builds from `cgr.dev/chainguard/wolfi-base:latest`: - -* **Stage 1 (builder)**: Downloads and/or compiles the database and its plugins -* **Stage 2 (runtime)**: Minimal image with only runtime dependencies - -All containers: - -* Run as non-root users -* Include `HEALTHCHECK` instructions -* Expose only necessary ports -* Use the smallest possible set of runtime dependencies - -== Test Data - -Each seed script creates consistent test data across all databases: - -.Test octads -[cols="1,2,1,1",options="header"] -|=== -| ID | Title | Drift Status | Modalities - -| `octad-test-001` -| Introduction to Cross-Modal Consistency -| healthy (0.045) -| document, vector, spatial, temporal, graph, provenance - -| `octad-test-002` -| Drift Detection Algorithms -| drifted (0.213) -| document, vector, spatial, graph - -| `octad-test-003` -| Self-Normalisation Process -| healthy (0.0) -| document, vector, semantic -|=== - -Each database's seed script includes these same entities in the appropriate -format for that database system, ensuring cross-database consistency that -the federation adapter tests can verify. - -== Troubleshooting - -=== Service fails to start - -Check the service logs: - -[source,bash] ----- -selur-compose logs mongodb -selur-compose logs redis-stack ----- - -=== Port conflicts - -If a port is already in use on the host, modify the port mapping in -`compose.toml`. For example, to remap MongoDB to port 27018: - -[source,toml] ----- -[services.mongodb] -ports = ["27018:27017"] ----- - -=== MongoDB replica set not initialising - -The replica set initialisation is handled by `mongodb-init.js`. If it fails, -initialise manually: - -[source,bash] ----- -mongosh --eval 'rs.initiate({_id: "rs0", members: [{_id: 0, host: "localhost:27017"}]})' ----- - -=== Redis modules not loading - -If RediSearch or other modules fail to load, check that the module `.so` files -exist in the image: - -[source,bash] ----- -podman exec -it ls -la /opt/redis/lib/ ----- - -=== ClickHouse or Neo4j out of memory - -These databases can be memory-hungry. Ensure at least 4 GB of RAM is -available for the full test stack. Reduce the ClickHouse -`mark_cache_size` in the Containerfile if needed. - -=== Seed script failures - -All seed scripts are idempotent -- they can be safely re-run. If a script -fails partway through, fix the issue and run it again. - -== Customising Seed Data - -Seed scripts are in `seed/` and can be modified to add additional test data. -When adding new test octads, ensure: - -1. The octad ID is unique across all seed scripts -2. The same octad is created in all relevant databases (for federation testing) -3. Drift scores and provenance events reference valid octad IDs -4. Spatial coordinates use valid GeoJSON (for MongoDB's `2dsphere` index) - -== File Reference - -[cols="1,3",options="header"] -|=== -| File | Purpose -| `compose.toml` | selur-compose service definitions (7 services) -| `manifest.toml` | Cerro-torre `.ctp` bundle manifest -| `.gatekeeper.yaml` | Svalinn gatekeeper policy (permissive for testing) -| `ct-build.sh` | Build pipeline for custom images -| `vordr.toml` | Runtime health monitoring configuration -| `deploy.k9.ncl` | k9-svc deployment component (Hunt level) -| `0-AI-MANIFEST.a2ml` | AI-readable manifest -| `images/Containerfile.redis-stack` | Redis + modules custom image -| `images/Containerfile.neo4j` | Neo4j + APOC custom image -| `images/Containerfile.clickhouse` | ClickHouse static binary custom image -| `images/Containerfile.surrealdb` | SurrealDB custom image -| `images/Containerfile.influxdb` | InfluxDB 2 custom image -| `seed/mongodb-init.js` | MongoDB seed (runs automatically) -| `seed/redis-init.sh` | Redis Stack seed (RediSearch + JSON + TimeSeries) -| `seed/neo4j-init.cypher` | Neo4j seed (constraints, indexes, nodes, relationships) -| `seed/clickhouse-init.sql` | ClickHouse seed (tables, materialized views, data) -| `seed/surrealdb-init.surql` | SurrealDB seed (schema, edges, data) -| `seed/influxdb-init.sh` | InfluxDB seed (buckets, metrics, federation health) -| `seed/minio-init.sh` | MinIO seed (buckets, objects, metadata) -|=== diff --git a/verisimdb/connectors/test-infra/compose.toml b/verisimdb/connectors/test-infra/compose.toml deleted file mode 100644 index cee39f58..00000000 --- a/verisimdb/connectors/test-infra/compose.toml +++ /dev/null @@ -1,152 +0,0 @@ -# SPDX-License-Identifier: MPL-2.0 -# -# VeriSimDB Test Infrastructure — selur-compose configuration -# -# Orchestrates 7 containerised databases for integration testing of the -# 10 federation adapters. DuckDB and SQLite are embedded and tested via -# unit tests; the remaining 8 adapters (MongoDB, Redis, Neo4j, ClickHouse, -# SurrealDB, InfluxDB, object storage, vector DB) are tested here. -# -# MinIO provides S3-compatible object storage. The vector DB adapter is -# tested against the native verisim-vector store (no external service needed). -# -# Usage: -# selur-compose up # Start all services -# selur-compose up --detach # Start in background -# selur-compose ps # Check status -# selur-compose logs -f mongodb # Stream logs -# selur-compose down # Stop all services -# selur-compose down --volumes # Stop and remove test volumes -# -# Fallback (without selur): -# podman-compose up --detach -# podman-compose down - -version = "1.0" - -# ============================================================================ -# Services -# ============================================================================ - -# MongoDB — document store federation adapter -# Replica set required for change streams (used by drift monitor) -[services.mongodb] -image = "cgr.dev/chainguard/mongodb:latest" -ports = ["27017:27017"] -environment = { - MONGO_INITDB_DATABASE = "verisimdb", -} -volumes = [ - "mongodb-data:/data/db", - "./seed/mongodb-init.js:/docker-entrypoint-initdb.d/init.js:ro", -] -command = ["mongod", "--replSet", "rs0", "--bind_ip_all"] -restart = "unless-stopped" -healthcheck = { test = "mongosh --eval 'db.adminCommand({ping:1})' --quiet", interval = "10s", timeout = "5s", retries = 10, start_period = "30s" } - -# Redis Stack — in-memory store with RediSearch, RedisJSON, RedisTimeSeries -[services.redis-stack] -build = { context = "./images", containerfile = "Containerfile.redis-stack" } -ports = ["6379:6379"] -volumes = ["redis-data:/data"] -restart = "unless-stopped" -healthcheck = { test = "redis-cli ping | grep -q PONG", interval = "10s", timeout = "5s", retries = 5, start_period = "10s" } - -# Neo4j Community — graph database for graph modality federation -[services.neo4j] -build = { context = "./images", containerfile = "Containerfile.neo4j" } -ports = ["7474:7474", "7687:7687"] -environment = { - NEO4J_AUTH = "none", - NEO4J_PLUGINS = '["apoc"]', - NEO4J_dbms_security_auth__enabled = "false", -} -volumes = [ - "neo4j-data:/data", - "neo4j-logs:/logs", -] -restart = "unless-stopped" -healthcheck = { test = "curl -sf http://localhost:7474/ || exit 1", interval = "10s", timeout = "5s", retries = 10, start_period = "30s" } - -# ClickHouse — columnar analytics for aggregate drift queries -[services.clickhouse] -build = { context = "./images", containerfile = "Containerfile.clickhouse" } -ports = ["8123:8123", "9000:9000"] -environment = { - CLICKHOUSE_DB = "verisimdb", - CLICKHOUSE_DEFAULT_ACCESS_MANAGEMENT = "1", -} -volumes = ["clickhouse-data:/var/lib/clickhouse"] -restart = "unless-stopped" -healthcheck = { test = "curl -sf http://localhost:8123/ping || exit 1", interval = "10s", timeout = "5s", retries = 5, start_period = "15s" } - -# SurrealDB — multi-model database (document + graph in one) -[services.surrealdb] -build = { context = "./images", containerfile = "Containerfile.surrealdb" } -ports = ["8000:8000"] -command = ["start", "--log", "info", "memory"] -restart = "unless-stopped" -healthcheck = { test = "curl -sf http://localhost:8000/health || exit 1", interval = "10s", timeout = "5s", retries = 5, start_period = "10s" } - -# InfluxDB 2 — time-series database for temporal modality federation -[services.influxdb] -build = { context = "./images", containerfile = "Containerfile.influxdb" } -ports = ["8086:8086"] -environment = { - DOCKER_INFLUXDB_INIT_MODE = "setup", - DOCKER_INFLUXDB_INIT_USERNAME = "verisim", - DOCKER_INFLUXDB_INIT_PASSWORD = "verisim-test-password", - DOCKER_INFLUXDB_INIT_ORG = "verisimdb", - DOCKER_INFLUXDB_INIT_BUCKET = "metrics", - DOCKER_INFLUXDB_INIT_ADMIN_TOKEN = "verisim-test-token-do-not-use-in-production", -} -volumes = ["influxdb-data:/var/lib/influxdb2"] -restart = "unless-stopped" -healthcheck = { test = "curl -sf http://localhost:8086/health || exit 1", interval = "10s", timeout = "5s", retries = 5, start_period = "15s" } - -# MinIO — S3-compatible object storage for binary blob modalities -[services.minio] -image = "cgr.dev/chainguard/minio:latest" -ports = ["9002:9000", "9001:9001"] -environment = { - MINIO_ROOT_USER = "verisim", - MINIO_ROOT_PASSWORD = "verisim-test-password", -} -command = ["server", "/data", "--console-address", ":9001"] -volumes = ["minio-data:/data"] -restart = "unless-stopped" -healthcheck = { test = "curl -sf http://localhost:9000/minio/health/live || exit 1", interval = "10s", timeout = "5s", retries = 5, start_period = "10s" } - -# ============================================================================ -# Volumes -# ============================================================================ - -[volumes.mongodb-data] -driver = "local" - -[volumes.redis-data] -driver = "local" - -[volumes.neo4j-data] -driver = "local" - -[volumes.neo4j-logs] -driver = "local" - -[volumes.clickhouse-data] -driver = "local" - -[volumes.influxdb-data] -driver = "local" - -[volumes.minio-data] -driver = "local" - -# ============================================================================ -# Networks -# ============================================================================ - -# Use selur zero-copy IPC for inter-service communication on the same host. -# Falls back to standard bridge networking when selur driver is unavailable. -[networks.default] -driver = "selur" diff --git a/verisimdb/connectors/test-infra/ct-build.sh b/verisimdb/connectors/test-infra/ct-build.sh deleted file mode 100755 index a66fc2cb..00000000 --- a/verisimdb/connectors/test-infra/ct-build.sh +++ /dev/null @@ -1,230 +0,0 @@ -#!/usr/bin/env bash -# SPDX-License-Identifier: MPL-2.0 -# -# VeriSimDB Test Infrastructure — Build Pipeline -# -# Builds all 5 custom container images for the test database stack. -# Unlike the production ct-build.sh, test images are NOT signed -# (--no-sign) and do not require cerro-torre attestations. -# -# The 2 remaining services (mongodb, minio) use upstream Chainguard -# images directly and do not need local builds. -# -# Prerequisites: -# - podman (container build) -# - ct (optional — cerro-torre CLI for .ctp bundles) -# -# Usage: -# ./ct-build.sh # Build all test images -# ./ct-build.sh --parallel # Build in parallel (faster) -# ./ct-build.sh --clean # Remove old test images first -# -# Author: Jonathan D.A. Jewell - -set -euo pipefail - -# --------------------------------------------------------------------------- -# Configuration -# --------------------------------------------------------------------------- - -SCRIPT_DIR="$(cd "$(dirname "$0")" && pwd)" -IMAGES_DIR="${SCRIPT_DIR}/images" - -# Image name prefix for test images -PREFIX="verisimdb-test" - -# Images to build (name:containerfile pairs) -# Ordered by typical build time (fastest first) -IMAGES=( - "surrealdb:Containerfile.surrealdb" - "influxdb:Containerfile.influxdb" - "clickhouse:Containerfile.clickhouse" - "neo4j:Containerfile.neo4j" - "redis-stack:Containerfile.redis-stack" -) - -PARALLEL=false -CLEAN=false - -for arg in "$@"; do - case "$arg" in - --parallel) PARALLEL=true ;; - --clean) CLEAN=true ;; - --help|-h) - echo "Usage: $0 [--parallel] [--clean]" - echo "" - echo "Options:" - echo " --parallel Build images in parallel (faster, more memory)" - echo " --clean Remove existing test images before building" - echo "" - echo "Builds 5 custom images:" - echo " ${PREFIX}-surrealdb" - echo " ${PREFIX}-influxdb" - echo " ${PREFIX}-clickhouse" - echo " ${PREFIX}-neo4j" - echo " ${PREFIX}-redis-stack" - exit 0 - ;; - *) - echo "Unknown option: ${arg}" - exit 1 - ;; - esac -done - -echo "=== VeriSimDB Test Infrastructure — Build Pipeline ===" -echo " Images dir: ${IMAGES_DIR}" -echo " Prefix: ${PREFIX}" -echo " Parallel: ${PARALLEL}" -echo " Clean: ${CLEAN}" -echo " Images: ${#IMAGES[@]}" -echo "" - -# --------------------------------------------------------------------------- -# Step 0: Clean (optional) -# --------------------------------------------------------------------------- - -if [ "$CLEAN" = true ]; then - echo "--- Step 0: Cleaning existing test images ---" - for entry in "${IMAGES[@]}"; do - name="${entry%%:*}" - podman rmi "${PREFIX}-${name}:latest" 2>/dev/null || true - echo " Removed: ${PREFIX}-${name}:latest (if it existed)" - done - echo "" -fi - -# --------------------------------------------------------------------------- -# Step 1: Build custom images -# --------------------------------------------------------------------------- - -echo "--- Step 1: Building ${#IMAGES[@]} custom images ---" - -build_image() { - local name="$1" - local containerfile="$2" - local full_name="${PREFIX}-${name}:latest" - - echo " Building ${full_name} from ${containerfile}..." - - podman build \ - --no-cache=false \ - -t "${full_name}" \ - -f "${IMAGES_DIR}/${containerfile}" \ - "${IMAGES_DIR}" - - echo " Built: ${full_name}" -} - -if [ "$PARALLEL" = true ]; then - # Build in parallel (requires enough memory for concurrent builds) - PIDS=() - for entry in "${IMAGES[@]}"; do - name="${entry%%:*}" - containerfile="${entry#*:}" - build_image "${name}" "${containerfile}" & - PIDS+=($!) - done - - # Wait for all builds to complete - FAILED=0 - for pid in "${PIDS[@]}"; do - if ! wait "$pid"; then - FAILED=$((FAILED + 1)) - fi - done - - if [ "$FAILED" -gt 0 ]; then - echo " ERROR: ${FAILED} image build(s) failed." - exit 1 - fi -else - # Build sequentially - for entry in "${IMAGES[@]}"; do - name="${entry%%:*}" - containerfile="${entry#*:}" - build_image "${name}" "${containerfile}" - echo "" - done -fi - -echo "" - -# --------------------------------------------------------------------------- -# Step 2: Pull upstream images -# --------------------------------------------------------------------------- - -echo "--- Step 2: Pulling upstream images ---" - -podman pull cgr.dev/chainguard/mongodb:latest -echo " Pulled: cgr.dev/chainguard/mongodb:latest" - -podman pull cgr.dev/chainguard/minio:latest -echo " Pulled: cgr.dev/chainguard/minio:latest" - -echo "" - -# --------------------------------------------------------------------------- -# Step 3: Pack as .ctp bundles (optional, --no-sign for test images) -# --------------------------------------------------------------------------- - -echo "--- Step 3: Packing .ctp bundles (if cerro-torre available) ---" - -if command -v ct &>/dev/null; then - for entry in "${IMAGES[@]}"; do - name="${entry%%:*}" - full_name="${PREFIX}-${name}:latest" - ctp_file="${SCRIPT_DIR}/${PREFIX}-${name}-latest.ctp" - - ct pack "${full_name}" -o "${ctp_file}" --no-sign - echo " Packed (unsigned): ${ctp_file}" - done -else - echo " SKIP: ct not found (install cerro-torre CLI from stapeln/container-stack/cerro-torre)" - echo " Images are built and tagged but not packed as .ctp bundles." -fi - -echo "" - -# --------------------------------------------------------------------------- -# Step 4: Verify images -# --------------------------------------------------------------------------- - -echo "--- Step 4: Listing built images ---" - -echo "" -echo " Custom images:" -for entry in "${IMAGES[@]}"; do - name="${entry%%:*}" - podman image inspect "${PREFIX}-${name}:latest" --format ' {{.Id | printf "%.12s"}} {{.Size | printf "%10d"}} {{index .RepoTags 0}}' 2>/dev/null \ - || echo " MISSING: ${PREFIX}-${name}:latest" -done - -echo "" -echo " Upstream images:" -podman image inspect cgr.dev/chainguard/mongodb:latest --format ' {{.Id | printf "%.12s"}} {{.Size | printf "%10d"}} {{index .RepoTags 0}}' 2>/dev/null \ - || echo " MISSING: cgr.dev/chainguard/mongodb:latest" -podman image inspect cgr.dev/chainguard/minio:latest --format ' {{.Id | printf "%.12s"}} {{.Size | printf "%10d"}} {{index .RepoTags 0}}' 2>/dev/null \ - || echo " MISSING: cgr.dev/chainguard/minio:latest" - -echo "" - -# --------------------------------------------------------------------------- -# Summary -# --------------------------------------------------------------------------- - -echo "=== Build pipeline complete ===" -echo "" -echo " To start the test stack:" -echo " cd ${SCRIPT_DIR}" -echo " selur-compose up --detach # or: podman-compose up --detach" -echo "" -echo " To seed test data:" -echo " ./seed/redis-init.sh" -echo " ./seed/influxdb-init.sh" -echo " ./seed/minio-init.sh" -echo " cat ./seed/neo4j-init.cypher | cypher-shell -a bolt://localhost:7687" -echo " clickhouse-client --multiquery < ./seed/clickhouse-init.sql" -echo " cat ./seed/surrealdb-init.surql | surreal sql --endpoint http://localhost:8000 --ns verisimdb --db test" -echo "" -echo " MongoDB seeds automatically via /docker-entrypoint-initdb.d/init.js" diff --git a/verisimdb/connectors/test-infra/deploy.k9.ncl b/verisimdb/connectors/test-infra/deploy.k9.ncl deleted file mode 100644 index eff4acfe..00000000 --- a/verisimdb/connectors/test-infra/deploy.k9.ncl +++ /dev/null @@ -1,256 +0,0 @@ -K9! -# SPDX-License-Identifier: MPL-2.0 -# -# deploy.k9.ncl — k9-svc deployment component for VeriSimDB test infrastructure -# -# Security Level: 'Hunt (requires cryptographic handshake) -# -# This is a LEGITIMATE Hunt-level component: it spawns database containers, -# writes to the filesystem (volumes), and binds network ports. It requires -# explicit authorization before execution. -# -# Author: Jonathan D.A. Jewell - -let pedigree = import "../../standards/k9-svc/pedigree.ncl" in -let leash = import "../../standards/k9-svc/leash.ncl" in - -# --------------------------------------------------------------------------- -# Component pedigree (self-description) -# --------------------------------------------------------------------------- - -let component_pedigree = { - metadata = { - name = "verisimdb-test-infra", - version = "0.1.0", - breed = "application/vnd.k9+nickel", - magic_number = "K9!", - description = "Integration test database stack for VeriSimDB federation adapters (7 services)", - }, - target = { - os = 'Linux, - is_edge = false, - requires_podman = true, - min_memory_mb = 4096, - }, - security = { - trust_level = 'Hunt, - allow_network = true, - allow_filesystem_write = true, - allow_subprocess = true, - # In production, this would be a real Ed25519 signature - signature = "PLACEHOLDER-SIGNATURE-REQUIRED-FOR-HUNT", - }, - validation = { - checksum = "sha256:placeholder-compose-and-seed-checksum", - pedigree_version = "1.0.0", - hunt_authorized = false, # Must be set true after handshake - }, - recipes = { - install = "just test-infra-build", - validate = "just test-infra-validate", - deploy = "just test-infra-up", - migrate = "just test-infra-seed", - }, -} in - -# --------------------------------------------------------------------------- -# Deployment configuration -# --------------------------------------------------------------------------- - -let deployment = { - # Test infrastructure only has one environment: local testing - environments = { - test = { - replicas = 1, - memory = "4Gi", - cpu = "2000m", - image_tag = "latest", - }, - }, - - # Services (7 databases) - services = { - mongodb = { - image = "cgr.dev/chainguard/mongodb:latest", - ports = [27017], - purpose = "Document store federation (replica set for change streams)", - }, - redis_stack = { - image = "verisimdb-test-redis-stack:latest", - ports = [6379], - purpose = "In-memory store with RediSearch, RedisJSON, RedisTimeSeries", - }, - neo4j = { - image = "verisimdb-test-neo4j:latest", - ports = [7474, 7687], - purpose = "Graph database with APOC plugin", - }, - clickhouse = { - image = "verisimdb-test-clickhouse:latest", - ports = [8123, 9000], - purpose = "Columnar analytics for aggregate drift queries", - }, - surrealdb = { - image = "verisimdb-test-surrealdb:latest", - ports = [8000], - purpose = "Multi-model database (document + graph)", - }, - influxdb = { - image = "verisimdb-test-influxdb:latest", - ports = [8086], - purpose = "Time-series database for temporal modality", - }, - minio = { - image = "cgr.dev/chainguard/minio:latest", - ports = [9002, 9001], - purpose = "S3-compatible object storage for binary blobs", - }, - }, - - # Network - network = { - driver = "selur", - fallback = "bridge", - }, -} in - -# --------------------------------------------------------------------------- -# Deployment scripts -# --------------------------------------------------------------------------- - -let scripts = { - # Build all custom images - install = m%" -#!/bin/sh -set -eu -echo "K9: Building VeriSimDB test infrastructure images..." -cd connectors/test-infra -./ct-build.sh -echo "K9: Build complete." -"%, - - # Validate compose.toml and seed scripts exist - validate = m%" -#!/bin/sh -set -eu -echo "K9: Validating VeriSimDB test infrastructure..." -cd connectors/test-infra - -for file in compose.toml manifest.toml .gatekeeper.yaml vordr.toml; do - if [ ! -f "$file" ]; then - echo "K9: FAIL — missing $file" - exit 1 - fi -done - -for seed in seed/mongodb-init.js seed/redis-init.sh seed/neo4j-init.cypher \ - seed/clickhouse-init.sql seed/surrealdb-init.surql \ - seed/influxdb-init.sh seed/minio-init.sh; do - if [ ! -f "$seed" ]; then - echo "K9: FAIL — missing $seed" - exit 1 - fi -done - -echo "K9: Validation passed (all files present)." -"%, - - # Start the full stack - deploy = m%" -#!/bin/sh -set -eu -echo "K9: Deploying VeriSimDB test infrastructure..." -cd connectors/test-infra - -if command -v selur-compose >/dev/null 2>&1; then - selur-compose up --detach -elif command -v podman-compose >/dev/null 2>&1; then - podman-compose up --detach -else - echo "K9: FAIL — neither selur-compose nor podman-compose found" - exit 1 -fi - -echo "K9: Test infrastructure deployed." -echo "K9: Waiting for services to become healthy..." -sleep 15 -echo "K9: Services should be ready. Run seed scripts next." -"%, - - # Run all seed scripts - migrate = m%" -#!/bin/sh -set -eu -echo "K9: Seeding VeriSimDB test databases..." -cd connectors/test-infra/seed - -echo "K9: MongoDB seeds automatically via entrypoint." - -echo "K9: Seeding Redis..." -./redis-init.sh - -echo "K9: Seeding Neo4j..." -cat neo4j-init.cypher | cypher-shell -a bolt://localhost:7687 2>/dev/null || \ - echo "K9: WARN — cypher-shell not found, skip Neo4j seed (seed manually)" - -echo "K9: Seeding ClickHouse..." -clickhouse-client --multiquery < clickhouse-init.sql 2>/dev/null || \ - echo "K9: WARN — clickhouse-client not found, skip ClickHouse seed (seed manually)" - -echo "K9: Seeding SurrealDB..." -cat surrealdb-init.surql | surreal sql --endpoint http://localhost:8000 --ns verisimdb --db test 2>/dev/null || \ - echo "K9: WARN — surreal CLI not found, skip SurrealDB seed (seed manually)" - -echo "K9: Seeding InfluxDB..." -./influxdb-init.sh - -echo "K9: Seeding MinIO..." -./minio-init.sh - -echo "K9: All seeds complete." -"%, - - # Tear down the stack - teardown = m%" -#!/bin/sh -set -eu -echo "K9: Tearing down VeriSimDB test infrastructure..." -cd connectors/test-infra - -if command -v selur-compose >/dev/null 2>&1; then - selur-compose down --volumes -elif command -v podman-compose >/dev/null 2>&1; then - podman-compose down --volumes -fi - -echo "K9: Test infrastructure removed." -"%, -} in - -# --------------------------------------------------------------------------- -# Export the component -# --------------------------------------------------------------------------- - -{ - pedigree = component_pedigree, - deployment = deployment, - scripts = scripts, - - # Security check: this component requires Hunt level - required_level = 'Hunt, - - # Warning for users - warning = m%" -WARNING: This is a Hunt-level component. - -It spawns 7 database containers, binds 11 network ports, and writes to -local volumes. Before running, ensure you have: - -1. Reviewed the compose.toml and Containerfiles -2. Verified the signature (when implemented) -3. Explicitly authorized Hunt-level execution -4. At least 4 GB RAM available for the database stack - -Run with: just authorize connectors/test-infra/deploy.k9.ncl && just test-infra-up -"%, -} diff --git a/verisimdb/connectors/test-infra/images/Containerfile.clickhouse b/verisimdb/connectors/test-infra/images/Containerfile.clickhouse deleted file mode 100644 index 206d5c0f..00000000 --- a/verisimdb/connectors/test-infra/images/Containerfile.clickhouse +++ /dev/null @@ -1,100 +0,0 @@ -# SPDX-License-Identifier: MPL-2.0 -# -# VeriSimDB Test Infrastructure — ClickHouse Server -# -# Builds ClickHouse server on Chainguard wolfi-base. Used for columnar -# analytics federation — aggregate drift queries, modality statistics, -# and cross-entity analytics. -# -# Provides: -# - HTTP interface on port 8123 (queries, health) -# - Native TCP interface on port 9000 (client connections) -# -# Author: Jonathan D.A. Jewell - -# --------------------------------------------------------------------------- -# Stage 1: Download ClickHouse static binary -# --------------------------------------------------------------------------- -FROM cgr.dev/chainguard/wolfi-base:latest AS builder - -RUN apk add --no-cache \ - curl \ - tar - -# ClickHouse distributes static binaries — no build needed -ARG CLICKHOUSE_VERSION=24.12.3.47 -RUN curl -fsSL \ - "https://packages.clickhouse.com/tgz/stable/clickhouse-common-static-${CLICKHOUSE_VERSION}-amd64.tgz" \ - | tar xz -C /tmp \ - && install -m 0755 /tmp/clickhouse-common-static-${CLICKHOUSE_VERSION}/usr/bin/clickhouse /usr/local/bin/clickhouse - -# --------------------------------------------------------------------------- -# Stage 2: Runtime image -# --------------------------------------------------------------------------- -FROM cgr.dev/chainguard/wolfi-base:latest - -RUN apk add --no-cache \ - libstdc++ \ - curl - -# Copy ClickHouse binary -COPY --from=builder /usr/local/bin/clickhouse /usr/bin/clickhouse - -# Symlinks for server/client/local sub-commands -RUN ln -s /usr/bin/clickhouse /usr/bin/clickhouse-server \ - && ln -s /usr/bin/clickhouse /usr/bin/clickhouse-client \ - && ln -s /usr/bin/clickhouse /usr/bin/clickhouse-local - -# Create non-root user -RUN addgroup -g 1000 clickhouse \ - && adduser -u 1000 -G clickhouse -D -h /var/lib/clickhouse clickhouse - -# Create directories -RUN mkdir -p \ - /var/lib/clickhouse \ - /var/log/clickhouse-server \ - /etc/clickhouse-server/config.d \ - /etc/clickhouse-server/users.d \ - && chown -R clickhouse:clickhouse \ - /var/lib/clickhouse \ - /var/log/clickhouse-server \ - /etc/clickhouse-server - -# Minimal server config for testing -RUN printf '%s\n' \ - '' \ - ' ' \ - ' information' \ - ' 1' \ - ' ' \ - ' 8123' \ - ' 9000' \ - ' 0.0.0.0' \ - ' /var/lib/clickhouse/' \ - ' /var/lib/clickhouse/tmp/' \ - ' /var/lib/clickhouse/access/' \ - ' 5368709120' \ - ' 100' \ - '' \ - > /etc/clickhouse-server/config.xml - -# Allow default user full access for testing -RUN printf '%s\n' \ - '' \ - ' ' \ - ' ' \ - ' 1' \ - ' ' \ - ' ' \ - '' \ - > /etc/clickhouse-server/users.d/test-access.xml - -USER clickhouse -WORKDIR /var/lib/clickhouse - -EXPOSE 8123 9000 - -HEALTHCHECK --interval=10s --timeout=5s --retries=5 --start-period=15s \ - CMD curl -sf http://localhost:8123/ping || exit 1 - -ENTRYPOINT ["clickhouse-server", "--config-file=/etc/clickhouse-server/config.xml"] diff --git a/verisimdb/connectors/test-infra/images/Containerfile.influxdb b/verisimdb/connectors/test-infra/images/Containerfile.influxdb deleted file mode 100644 index c1d57be6..00000000 --- a/verisimdb/connectors/test-infra/images/Containerfile.influxdb +++ /dev/null @@ -1,66 +0,0 @@ -# SPDX-License-Identifier: MPL-2.0 -# -# VeriSimDB Test Infrastructure — InfluxDB 2 -# -# Builds InfluxDB 2 on Chainguard wolfi-base. Used for temporal modality -# federation — time-series drift scores, query latency metrics, and -# federation health telemetry. -# -# Provides: -# - HTTP API on port 8086 (Flux queries, Line Protocol writes) -# - /health endpoint for readiness checks -# - Pre-configured org "verisimdb" and bucket "metrics" -# -# Author: Jonathan D.A. Jewell - -# --------------------------------------------------------------------------- -# Stage 1: Download InfluxDB 2 binary -# --------------------------------------------------------------------------- -FROM cgr.dev/chainguard/wolfi-base:latest AS builder - -RUN apk add --no-cache \ - curl \ - tar - -ARG INFLUXDB_VERSION=2.7.11 -RUN curl -fsSL \ - "https://download.influxdata.com/influxdb/releases/influxdb2-${INFLUXDB_VERSION}_linux_amd64.tar.gz" \ - | tar xz -C /tmp \ - && install -m 0755 "/tmp/influxdb2-${INFLUXDB_VERSION}/usr/bin/influxd" /usr/local/bin/influxd - -# Also grab the influx CLI for seed scripts -ARG INFLUX_CLI_VERSION=2.7.5 -RUN curl -fsSL \ - "https://download.influxdata.com/influxdb/releases/influxdb2-client-${INFLUX_CLI_VERSION}-linux-amd64.tar.gz" \ - | tar xz -C /tmp \ - && install -m 0755 /tmp/influx /usr/local/bin/influx - -# --------------------------------------------------------------------------- -# Stage 2: Runtime image -# --------------------------------------------------------------------------- -FROM cgr.dev/chainguard/wolfi-base:latest - -RUN apk add --no-cache \ - curl - -# Copy InfluxDB server and CLI -COPY --from=builder /usr/local/bin/influxd /usr/bin/influxd -COPY --from=builder /usr/local/bin/influx /usr/bin/influx - -# Create non-root user -RUN addgroup -g 1000 influxdb \ - && adduser -u 1000 -G influxdb -D -h /var/lib/influxdb2 influxdb - -# Create directories -RUN mkdir -p /var/lib/influxdb2 /etc/influxdb2 \ - && chown -R influxdb:influxdb /var/lib/influxdb2 /etc/influxdb2 - -USER influxdb -WORKDIR /var/lib/influxdb2 - -EXPOSE 8086 - -HEALTHCHECK --interval=10s --timeout=5s --retries=5 --start-period=15s \ - CMD curl -sf http://localhost:8086/health || exit 1 - -ENTRYPOINT ["influxd"] diff --git a/verisimdb/connectors/test-infra/images/Containerfile.neo4j b/verisimdb/connectors/test-infra/images/Containerfile.neo4j deleted file mode 100644 index 91229c09..00000000 --- a/verisimdb/connectors/test-infra/images/Containerfile.neo4j +++ /dev/null @@ -1,82 +0,0 @@ -# SPDX-License-Identifier: MPL-2.0 -# -# VeriSimDB Test Infrastructure — Neo4j Community Edition -# -# Builds Neo4j Community with APOC plugin on Chainguard wolfi-base. -# Auth is disabled for testing (NEO4J_AUTH=none). -# -# Provides: -# - Bolt protocol on port 7687 (driver access) -# - HTTP API on port 7474 (browser + REST) -# - APOC plugin for advanced graph procedures -# -# Author: Jonathan D.A. Jewell - -# --------------------------------------------------------------------------- -# Stage 1: Download and prepare Neo4j + APOC -# --------------------------------------------------------------------------- -FROM cgr.dev/chainguard/wolfi-base:latest AS builder - -RUN apk add --no-cache \ - curl \ - tar - -ARG NEO4J_VERSION=5.26.3 -ARG APOC_VERSION=5.26.0 - -# Download Neo4j Community -RUN curl -fsSL \ - "https://dist.neo4j.org/neo4j-community-${NEO4J_VERSION}-unix.tar.gz" \ - | tar xz -C /opt \ - && mv "/opt/neo4j-community-${NEO4J_VERSION}" /opt/neo4j - -# Download APOC Core plugin -RUN curl -fsSL \ - "https://github.com/neo4j/apoc/releases/download/${APOC_VERSION}/apoc-${APOC_VERSION}-core.jar" \ - -o /opt/neo4j/plugins/apoc-core.jar - -# --------------------------------------------------------------------------- -# Stage 2: Runtime image -# --------------------------------------------------------------------------- -FROM cgr.dev/chainguard/wolfi-base:latest - -RUN apk add --no-cache \ - openjdk-17-jre-headless \ - curl \ - bash - -# Copy Neo4j installation -COPY --from=builder /opt/neo4j /opt/neo4j - -# Create non-root user -RUN addgroup -g 7474 neo4j \ - && adduser -u 7474 -G neo4j -D -h /var/lib/neo4j neo4j - -# Create data and log directories -RUN mkdir -p /data /logs /var/lib/neo4j \ - && chown -R neo4j:neo4j /opt/neo4j /data /logs /var/lib/neo4j - -# Configure Neo4j for test use -RUN printf '%s\n' \ - "server.default_listen_address=0.0.0.0" \ - "server.bolt.listen_address=:7687" \ - "server.http.listen_address=:7474" \ - "dbms.security.auth_enabled=false" \ - "server.directories.data=/data" \ - "server.directories.logs=/logs" \ - "dbms.security.procedures.unrestricted=apoc.*" \ - "dbms.security.procedures.allowlist=apoc.*" \ - > /opt/neo4j/conf/neo4j.conf - -ENV PATH="/opt/neo4j/bin:${PATH}" -ENV NEO4J_HOME="/opt/neo4j" - -USER neo4j -WORKDIR /var/lib/neo4j - -EXPOSE 7474 7687 - -HEALTHCHECK --interval=10s --timeout=5s --retries=10 --start-period=30s \ - CMD curl -sf http://localhost:7474/ || exit 1 - -ENTRYPOINT ["neo4j", "console"] diff --git a/verisimdb/connectors/test-infra/images/Containerfile.redis-stack b/verisimdb/connectors/test-infra/images/Containerfile.redis-stack deleted file mode 100644 index a593bbb0..00000000 --- a/verisimdb/connectors/test-infra/images/Containerfile.redis-stack +++ /dev/null @@ -1,116 +0,0 @@ -# SPDX-License-Identifier: MPL-2.0 -# -# VeriSimDB Test Infrastructure — Redis Stack -# -# Multi-stage build of Redis with RediSearch, RedisJSON, RedisTimeSeries, -# and RedisGraph modules on top of Chainguard wolfi-base. -# -# Provides: -# - RediSearch: full-text search over hexad document modalities -# - RedisJSON: native JSON storage for hexad documents -# - RedisTimeSeries: temporal modality time-series data -# - RedisGraph: lightweight graph queries (deprecated upstream but -# still useful for test coverage of graph federation paths) -# -# Author: Jonathan D.A. Jewell - -# --------------------------------------------------------------------------- -# Stage 1: Build Redis and modules from source -# --------------------------------------------------------------------------- -FROM cgr.dev/chainguard/wolfi-base:latest AS builder - -RUN apk add --no-cache \ - build-base \ - cmake \ - git \ - curl \ - clang \ - llvm \ - rust \ - cargo \ - openssl-dev \ - libstdc++-dev - -# Build Redis -ARG REDIS_VERSION=7.4.2 -RUN curl -fsSL "https://download.redis.io/releases/redis-${REDIS_VERSION}.tar.gz" \ - | tar xz -C /tmp \ - && cd "/tmp/redis-${REDIS_VERSION}" \ - && make -j"$(nproc)" BUILD_TLS=yes \ - && make install PREFIX=/opt/redis - -# Build RediSearch -ARG REDISEARCH_VERSION=2.10.7 -RUN git clone --depth 1 --branch "v${REDISEARCH_VERSION}" \ - --recurse-submodules \ - https://github.com/RediSearch/RediSearch.git /tmp/RediSearch \ - && cd /tmp/RediSearch \ - && mkdir build && cd build \ - && cmake .. -DCMAKE_BUILD_TYPE=Release \ - && make -j"$(nproc)" \ - && cp /tmp/RediSearch/build/redisearch.so /opt/redis/lib/ - -# Build RedisJSON -ARG REDISJSON_VERSION=2.8.5 -RUN git clone --depth 1 --branch "v${REDISJSON_VERSION}" \ - https://github.com/RedisJSON/RedisJSON.git /tmp/RedisJSON \ - && cd /tmp/RedisJSON \ - && cargo build --release \ - && cp target/release/librejson.so /opt/redis/lib/rejson.so - -# Build RedisTimeSeries -ARG REDISTIMESERIES_VERSION=1.12.2 -RUN git clone --depth 1 --branch "v${REDISTIMESERIES_VERSION}" \ - --recurse-submodules \ - https://github.com/RedisTimeSeries/RedisTimeSeries.git /tmp/RedisTimeSeries \ - && cd /tmp/RedisTimeSeries \ - && mkdir build && cd build \ - && cmake .. -DCMAKE_BUILD_TYPE=Release \ - && make -j"$(nproc)" \ - && cp /tmp/RedisTimeSeries/build/redistimeseries.so /opt/redis/lib/ - -# --------------------------------------------------------------------------- -# Stage 2: Runtime image (minimal) -# --------------------------------------------------------------------------- -FROM cgr.dev/chainguard/wolfi-base:latest - -RUN apk add --no-cache \ - libstdc++ \ - openssl \ - curl - -# Copy Redis binaries and modules -COPY --from=builder /opt/redis /opt/redis - -# Create non-root user -RUN addgroup -g 1000 redis \ - && adduser -u 1000 -G redis -D -h /data redis - -# Create directories -RUN mkdir -p /data /etc/redis \ - && chown -R redis:redis /data /etc/redis - -# Redis configuration with modules loaded -RUN printf '%s\n' \ - "bind 0.0.0.0" \ - "port 6379" \ - "dir /data" \ - "loadmodule /opt/redis/lib/redisearch.so" \ - "loadmodule /opt/redis/lib/rejson.so" \ - "loadmodule /opt/redis/lib/redistimeseries.so" \ - "protected-mode no" \ - "save ''" \ - "appendonly no" \ - > /etc/redis/redis.conf - -ENV PATH="/opt/redis/bin:${PATH}" - -USER redis -WORKDIR /data - -EXPOSE 6379 - -HEALTHCHECK --interval=10s --timeout=5s --retries=5 --start-period=10s \ - CMD redis-cli ping | grep -q PONG - -ENTRYPOINT ["redis-server", "/etc/redis/redis.conf"] diff --git a/verisimdb/connectors/test-infra/images/Containerfile.surrealdb b/verisimdb/connectors/test-infra/images/Containerfile.surrealdb deleted file mode 100644 index efdbd038..00000000 --- a/verisimdb/connectors/test-infra/images/Containerfile.surrealdb +++ /dev/null @@ -1,63 +0,0 @@ -# SPDX-License-Identifier: MPL-2.0 -# -# VeriSimDB Test Infrastructure — SurrealDB -# -# Builds SurrealDB on Chainguard wolfi-base. SurrealDB is a multi-model -# database that combines document, graph, and key-value capabilities — -# useful for testing federation across document and graph modalities -# simultaneously. -# -# Runs in memory mode for tests (no persistence overhead). -# -# Provides: -# - HTTP API on port 8000 (REST + WebSocket) -# - /health endpoint for readiness checks -# -# Author: Jonathan D.A. Jewell - -# --------------------------------------------------------------------------- -# Stage 1: Download SurrealDB binary -# --------------------------------------------------------------------------- -FROM cgr.dev/chainguard/wolfi-base:latest AS builder - -RUN apk add --no-cache \ - curl \ - tar - -ARG SURREALDB_VERSION=2.2.1 -RUN curl -fsSL \ - "https://download.surrealdb.com/v${SURREALDB_VERSION}/surreal-v${SURREALDB_VERSION}.linux-amd64.tgz" \ - | tar xz -C /tmp \ - && install -m 0755 /tmp/surreal /usr/local/bin/surreal - -# --------------------------------------------------------------------------- -# Stage 2: Runtime image -# --------------------------------------------------------------------------- -FROM cgr.dev/chainguard/wolfi-base:latest - -RUN apk add --no-cache \ - libstdc++ \ - curl - -# Copy SurrealDB binary -COPY --from=builder /usr/local/bin/surreal /usr/bin/surreal - -# Create non-root user -RUN addgroup -g 1000 surrealdb \ - && adduser -u 1000 -G surrealdb -D -h /var/lib/surrealdb surrealdb - -# Create data directory (used if persistent mode ever needed) -RUN mkdir -p /var/lib/surrealdb \ - && chown -R surrealdb:surrealdb /var/lib/surrealdb - -USER surrealdb -WORKDIR /var/lib/surrealdb - -EXPOSE 8000 - -HEALTHCHECK --interval=10s --timeout=5s --retries=5 --start-period=10s \ - CMD curl -sf http://localhost:8000/health || exit 1 - -# Default: memory mode, no auth, all interfaces -ENTRYPOINT ["surreal"] -CMD ["start", "--log", "info", "--bind", "0.0.0.0:8000", "memory"] diff --git a/verisimdb/connectors/test-infra/manifest.toml b/verisimdb/connectors/test-infra/manifest.toml deleted file mode 100644 index 9dadadb5..00000000 --- a/verisimdb/connectors/test-infra/manifest.toml +++ /dev/null @@ -1,136 +0,0 @@ -# SPDX-License-Identifier: MPL-2.0 -# -# Cerro Torre manifest for VeriSimDB test infrastructure bundle -# -# Describes the test database stack for verified container packaging. -# Unlike the production manifest, test images are NOT signed (--no-sign) -# and use relaxed security policies. -# -# Used by `ct pack` to create .ctp bundles of the test infrastructure. - -[metadata] -name = "verisimdb-test-infra" -version = "0.1.0" -revision = 1 -summary = "Integration test database stack for VeriSimDB federation adapters" -description = """ -Containerised stack of 7 databases for integration testing of VeriSimDB's -10 federation adapters. Includes MongoDB (replica set), Redis Stack -(RediSearch + RedisJSON + RedisTimeSeries), Neo4j Community (APOC), -ClickHouse, SurrealDB, InfluxDB 2, and MinIO (S3-compatible). - -DuckDB and SQLite are embedded databases tested via unit tests and are -not included in this stack. -""" -license = "PMPL-1.0-or-later" -homepage = "https://github.com/hyperpolymath/verisimdb" -maintainer = "Jonathan D.A. Jewell " - -[provenance] -upstream = "https://github.com/hyperpolymath/verisimdb" -import_date = 2026-02-28T00:00:00Z - -[dependencies] -runtime = ["podman", "curl"] -build = ["podman"] - -[build] -system = "podman" - -[build.environment] -COMPOSE_FILE = "compose.toml" - -# --------------------------------------------------------------------------- -# Services — 7 database containers -# --------------------------------------------------------------------------- - -[services.mongodb] -image = "cgr.dev/chainguard/mongodb:latest" -build_context = false -ports = [27017] -purpose = "Document store federation adapter (replica set for change streams)" - -[services.redis-stack] -image = "verisimdb-test-redis-stack:latest" -build_context = "./images" -containerfile = "Containerfile.redis-stack" -ports = [6379] -purpose = "In-memory store with RediSearch, RedisJSON, RedisTimeSeries" - -[services.neo4j] -image = "verisimdb-test-neo4j:latest" -build_context = "./images" -containerfile = "Containerfile.neo4j" -ports = [7474, 7687] -purpose = "Graph database for graph modality federation (APOC plugin)" - -[services.clickhouse] -image = "verisimdb-test-clickhouse:latest" -build_context = "./images" -containerfile = "Containerfile.clickhouse" -ports = [8123, 9000] -purpose = "Columnar analytics for aggregate drift queries" - -[services.surrealdb] -image = "verisimdb-test-surrealdb:latest" -build_context = "./images" -containerfile = "Containerfile.surrealdb" -ports = [8000] -purpose = "Multi-model database (document + graph) federation" - -[services.influxdb] -image = "verisimdb-test-influxdb:latest" -build_context = "./images" -containerfile = "Containerfile.influxdb" -ports = [8086] -purpose = "Time-series database for temporal modality federation" - -[services.minio] -image = "cgr.dev/chainguard/minio:latest" -build_context = false -ports = [9002, 9001] -purpose = "S3-compatible object storage for binary blob modalities" - -# --------------------------------------------------------------------------- -# Outputs -# --------------------------------------------------------------------------- - -[outputs] -primary = "verisimdb-test-infra" -split = [ - "verisimdb-test-redis-stack", - "verisimdb-test-neo4j", - "verisimdb-test-clickhouse", - "verisimdb-test-surrealdb", - "verisimdb-test-influxdb", -] - -# --------------------------------------------------------------------------- -# Attestations — relaxed for test images -# --------------------------------------------------------------------------- - -[attestations] -require = [] -recommend = ["sbom-complete"] - -# --------------------------------------------------------------------------- -# Security — relaxed for local testing -# --------------------------------------------------------------------------- - -[security] -user = "test" -group = "test" -read_only_root = false -no_new_privileges = false - -[security.capabilities] -drop = [] -add = [] - -[security.network] -listen_tcp = [6379, 7474, 7687, 8000, 8086, 8123, 9000, 9001, 9002, 27017] - -[security.filesystem] -read = ["/"] -write = ["/data/", "/tmp/", "/var/lib/"] -execute = [] diff --git a/verisimdb/connectors/test-infra/seed/clickhouse-init.sql b/verisimdb/connectors/test-infra/seed/clickhouse-init.sql deleted file mode 100644 index a74b60cd..00000000 --- a/verisimdb/connectors/test-infra/seed/clickhouse-init.sql +++ /dev/null @@ -1,247 +0,0 @@ --- SPDX-License-Identifier: MPL-2.0 --- --- VeriSimDB Test Infrastructure — ClickHouse Seed Script --- --- Creates the verisimdb database with tables for hexads, modalities, and --- drift scores. ClickHouse is used for columnar analytics over large --- modality datasets — aggregate drift queries, modality distribution --- statistics, and cross-entity analytics at scale. --- --- Execute with: --- clickhouse-client --host localhost --port 9000 --multiquery < clickhouse-init.sql --- curl -s http://localhost:8123/ --data-binary @clickhouse-init.sql --- --- Author: Jonathan D.A. Jewell - --- --------------------------------------------------------------------------- --- Database --- --------------------------------------------------------------------------- - -CREATE DATABASE IF NOT EXISTS verisimdb; - --- --------------------------------------------------------------------------- --- Hexads table — core entity metadata (MergeTree for fast OLAP scans) --- --------------------------------------------------------------------------- - -CREATE TABLE IF NOT EXISTS verisimdb.hexads -( - id String, - title String, - content String, - entity_type String, - primary_modality String, - version UInt32, - drift_status Enum8('healthy' = 0, 'drifted' = 1, 'normalising' = 2, 'stale' = 3), - drift_score Float64, - modality_count UInt8, - created_at DateTime64(3), - updated_at DateTime64(3), - -- Spatial data (nullable — not all hexads have spatial modalities) - spatial_lat Nullable(Float64), - spatial_lon Nullable(Float64), - -- Vector metadata - embedding_model Nullable(String), - embedding_dimensions Nullable(UInt16), - -- Tags for fast filtering - tags Array(String) -) -ENGINE = MergeTree() -ORDER BY (created_at, id) -PARTITION BY toYYYYMM(created_at) -SETTINGS index_granularity = 8192; - --- --------------------------------------------------------------------------- --- Modalities table — individual modality records per hexad --- --------------------------------------------------------------------------- - -CREATE TABLE IF NOT EXISTS verisimdb.modalities -( - hexad_id String, - modality_type Enum8('graph' = 0, 'vector' = 1, 'tensor' = 2, 'semantic' = 3, - 'document' = 4, 'temporal' = 5, 'provenance' = 6, 'spatial' = 7), - data_size_bytes UInt64, - last_updated DateTime64(3), - is_authoritative UInt8, -- Boolean: 1 = authoritative source of truth - quality_score Float64, - metadata String -- JSON blob for modality-specific data -) -ENGINE = MergeTree() -ORDER BY (hexad_id, modality_type) -SETTINGS index_granularity = 8192; - --- --------------------------------------------------------------------------- --- Drift scores table — time-series of drift measurements --- --------------------------------------------------------------------------- - -CREATE TABLE IF NOT EXISTS verisimdb.drift_scores -( - hexad_id String, - measured_at DateTime64(3), - semantic_vector_drift Float64, - graph_document_drift Float64, - temporal_consistency_drift Float64, - tensor_drift Float64, - schema_drift Float64, - quality_drift Float64, - overall Float64, - status Enum8('healthy' = 0, 'drifted' = 1, 'normalising' = 2, 'stale' = 3) -) -ENGINE = MergeTree() -ORDER BY (hexad_id, measured_at) -PARTITION BY toYYYYMM(measured_at) -SETTINGS index_granularity = 8192; - --- --------------------------------------------------------------------------- --- Provenance events table — lineage tracking --- --------------------------------------------------------------------------- - -CREATE TABLE IF NOT EXISTS verisimdb.provenance_events -( - event_id String, - hexad_id String, - event_type String, - actor String, - timestamp DateTime64(3), - hash String, - parent_hash Nullable(String), - details String -) -ENGINE = MergeTree() -ORDER BY (hexad_id, timestamp) -SETTINGS index_granularity = 8192; - --- --------------------------------------------------------------------------- --- Materialized views for fast aggregate queries --- --------------------------------------------------------------------------- - --- Drift status distribution (how many hexads per drift status?) -CREATE MATERIALIZED VIEW IF NOT EXISTS verisimdb.mv_drift_status_counts -ENGINE = SummingMergeTree() -ORDER BY (status, day) -AS SELECT - status, - toDate(measured_at) AS day, - count() AS hexad_count -FROM verisimdb.drift_scores -GROUP BY status, day; - --- Modality distribution (which modalities are most common?) -CREATE MATERIALIZED VIEW IF NOT EXISTS verisimdb.mv_modality_distribution -ENGINE = SummingMergeTree() -ORDER BY modality_type -AS SELECT - modality_type, - count() AS occurrence_count, - avg(quality_score) AS avg_quality -FROM verisimdb.modalities -GROUP BY modality_type; - --- Average drift by entity type -CREATE MATERIALIZED VIEW IF NOT EXISTS verisimdb.mv_avg_drift_by_type -ENGINE = AggregatingMergeTree() -ORDER BY entity_type -AS SELECT - h.entity_type AS entity_type, - avgState(d.overall) AS avg_drift, - countState() AS entity_count -FROM verisimdb.drift_scores d -INNER JOIN verisimdb.hexads h ON d.hexad_id = h.id -GROUP BY h.entity_type; - --- --------------------------------------------------------------------------- --- Insert test data — hexads --- --------------------------------------------------------------------------- - -INSERT INTO verisimdb.hexads VALUES -( - 'hexad-test-001', - 'Introduction to Cross-Modal Consistency', - 'VeriSimDB maintains consistency across 8 modality representations.', - 'Article', - 'document', - 3, - 'healthy', - 0.045, - 6, - now() - INTERVAL 1 DAY, - now(), - 51.5074, - -0.1278, - 'test-embedding-v1', - 128, - ['consistency', 'cross-modal', 'verisimdb'] -), -( - 'hexad-test-002', - 'Drift Detection Algorithms', - 'Drift is measured as divergence between modalities using cosine similarity.', - 'TechArticle', - 'document', - 1, - 'drifted', - 0.213, - 3, - now() - INTERVAL 1 DAY, - now() - INTERVAL 1 HOUR, - 40.7484, - -73.9857, - 'test-embedding-v1', - 128, - ['drift', 'algorithms', 'cosine-similarity'] -), -( - 'hexad-test-003', - 'Self-Normalisation Process', - 'When drift exceeds configurable thresholds, the normaliser regenerates modalities.', - 'Article', - 'document', - 1, - 'healthy', - 0.0, - 2, - now(), - now(), - NULL, - NULL, - NULL, - NULL, - ['normalisation', 'consistency', 'drift'] -); - --- --------------------------------------------------------------------------- --- Insert test data — modalities --- --------------------------------------------------------------------------- - -INSERT INTO verisimdb.modalities VALUES -('hexad-test-001', 'document', 1024, now(), 1, 0.95, '{"format": "text/plain"}'), -('hexad-test-001', 'vector', 512, now(), 0, 0.88, '{"model": "test-embedding-v1", "dim": 128}'), -('hexad-test-001', 'graph', 256, now(), 0, 0.92, '{"triples": 5, "types": 2}'), -('hexad-test-001', 'temporal', 128, now(), 0, 0.97, '{"versions": 3}'), -('hexad-test-001', 'spatial', 64, now(), 0, 0.90, '{"type": "Point", "srid": 4326}'), -('hexad-test-001', 'provenance', 96, now(), 0, 0.99, '{"chain_length": 2}'), -('hexad-test-002', 'document', 896, now() - INTERVAL 1 HOUR, 1, 0.82, '{"format": "text/plain"}'), -('hexad-test-002', 'vector', 512, now() - INTERVAL 1 HOUR, 0, 0.55, '{"model": "test-embedding-v1", "dim": 128}'), -('hexad-test-002', 'spatial', 64, now() - INTERVAL 1 HOUR, 0, 0.78, '{"type": "Point", "srid": 4326}'), -('hexad-test-003', 'document', 768, now(), 1, 0.91, '{"format": "text/plain"}'), -('hexad-test-003', 'semantic', 128, now(), 0, 0.92, '{"categories": 3, "confidence": 0.92}'); - --- --------------------------------------------------------------------------- --- Insert test data — drift scores (time-series) --- --------------------------------------------------------------------------- - -INSERT INTO verisimdb.drift_scores VALUES -('hexad-test-001', now() - INTERVAL 1 HOUR, 0.08, 0.03, 0.01, 0.0, 0.0, 0.05, 0.028, 'healthy'), -('hexad-test-001', now() - INTERVAL 30 MINUTE, 0.10, 0.04, 0.02, 0.0, 0.0, 0.06, 0.037, 'healthy'), -('hexad-test-001', now(), 0.12, 0.05, 0.02, 0.0, 0.0, 0.08, 0.045, 'healthy'), -('hexad-test-002', now() - INTERVAL 1 HOUR, 0.38, 0.28, 0.12, 0.0, 0.06, 0.22, 0.177, 'drifted'), -('hexad-test-002', now(), 0.45, 0.32, 0.15, 0.0, 0.08, 0.28, 0.213, 'drifted'); - --- --------------------------------------------------------------------------- --- Insert test data — provenance events --- --------------------------------------------------------------------------- - -INSERT INTO verisimdb.provenance_events VALUES -('prov-001-create', 'hexad-test-001', 'created', 'test-seed-script', now() - INTERVAL 1 DAY, 'sha256:a1b2c3d4e5f6', NULL, 'Initial creation via clickhouse-init.sql'), -('prov-001-update', 'hexad-test-001', 'modality_updated', 'test-seed-script', now() - INTERVAL 1 HOUR, 'sha256:f6e5d4c3b2a1', 'sha256:a1b2c3d4e5f6', 'Vector embedding regenerated'), -('prov-002-create', 'hexad-test-002', 'created', 'test-seed-script', now() - INTERVAL 1 DAY, 'sha256:b2c3d4e5f6a1', NULL, 'Initial creation'), -('prov-003-create', 'hexad-test-003', 'created', 'test-seed-script', now(), 'sha256:c3d4e5f6a1b2', NULL, 'Initial creation'); diff --git a/verisimdb/connectors/test-infra/seed/influxdb-init.sh b/verisimdb/connectors/test-infra/seed/influxdb-init.sh deleted file mode 100755 index 59e93844..00000000 --- a/verisimdb/connectors/test-infra/seed/influxdb-init.sh +++ /dev/null @@ -1,176 +0,0 @@ -#!/bin/sh -# SPDX-License-Identifier: MPL-2.0 -# -# VeriSimDB Test Infrastructure — InfluxDB 2 Seed Script -# -# Creates the verisimdb org and metrics bucket, then writes test data -# points for drift scores, query latency, and federation health metrics. -# -# Prerequisites: -# - InfluxDB 2 container running on localhost:8086 -# - influx CLI available in PATH -# - InfluxDB already initialised (auto-setup via environment variables) -# -# Usage: -# ./influxdb-init.sh # Default: localhost:8086 -# INFLUX_HOST=http://influxdb:8086 INFLUX_TOKEN=my-token ./influxdb-init.sh -# -# Author: Jonathan D.A. Jewell - -set -eu - -INFLUX_HOST="${INFLUX_HOST:-http://localhost:8086}" -INFLUX_TOKEN="${INFLUX_TOKEN:-verisim-test-token-do-not-use-in-production}" -INFLUX_ORG="${INFLUX_ORG:-verisimdb}" - -echo "=== VeriSimDB InfluxDB 2 Seed ===" -echo " Host: ${INFLUX_HOST}" -echo " Org: ${INFLUX_ORG}" - -# --------------------------------------------------------------------------- -# Wait for InfluxDB to be ready -# --------------------------------------------------------------------------- - -echo "--- Waiting for InfluxDB..." -for attempt in $(seq 1 30); do - if curl -sf "${INFLUX_HOST}/health" >/dev/null 2>&1; then - echo " InfluxDB is ready." - break - fi - if [ "$attempt" -eq 30 ]; then - echo " ERROR: InfluxDB did not become ready within 30 seconds." - exit 1 - fi - sleep 1 -done - -# --------------------------------------------------------------------------- -# Create additional buckets -# --------------------------------------------------------------------------- - -echo "--- Creating additional buckets..." - -# 'metrics' bucket is created by auto-setup; create additional ones -influx bucket create \ - --host "${INFLUX_HOST}" \ - --token "${INFLUX_TOKEN}" \ - --org "${INFLUX_ORG}" \ - --name "drift_scores" \ - --retention 30d \ - 2>/dev/null || echo " Bucket 'drift_scores' already exists (OK)" - -influx bucket create \ - --host "${INFLUX_HOST}" \ - --token "${INFLUX_TOKEN}" \ - --org "${INFLUX_ORG}" \ - --name "federation_health" \ - --retention 7d \ - 2>/dev/null || echo " Bucket 'federation_health' already exists (OK)" - -echo " Buckets created: metrics (auto), drift_scores, federation_health" - -# --------------------------------------------------------------------------- -# Helper: generate timestamps relative to now (seconds precision) -# --------------------------------------------------------------------------- - -NOW=$(date +%s) - -# Write data using Line Protocol via the API -write_lines() { - BUCKET="$1" - shift - printf '%s\n' "$@" | curl -sf -XPOST \ - "${INFLUX_HOST}/api/v2/write?org=${INFLUX_ORG}&bucket=${BUCKET}&precision=s" \ - -H "Authorization: Token ${INFLUX_TOKEN}" \ - -H "Content-Type: text/plain" \ - --data-binary @- \ - || echo " WARNING: Failed to write to bucket ${BUCKET}" -} - -# --------------------------------------------------------------------------- -# Write drift score metrics -# --------------------------------------------------------------------------- - -echo "--- Writing drift score metrics..." - -# Drift scores for hexad-test-001 (healthy, gradual drift) -write_lines "drift_scores" \ - "drift,hexad_id=hexad-test-001,status=healthy semantic_vector=0.08,graph_document=0.03,temporal=0.01,overall=0.028 $((NOW - 3600))" \ - "drift,hexad_id=hexad-test-001,status=healthy semantic_vector=0.10,graph_document=0.04,temporal=0.02,overall=0.037 $((NOW - 1800))" \ - "drift,hexad_id=hexad-test-001,status=healthy semantic_vector=0.12,graph_document=0.05,temporal=0.02,overall=0.045 ${NOW}" - -# Drift scores for hexad-test-002 (drifted, increasing) -write_lines "drift_scores" \ - "drift,hexad_id=hexad-test-002,status=drifted semantic_vector=0.30,graph_document=0.22,temporal=0.10,overall=0.150 $((NOW - 7200))" \ - "drift,hexad_id=hexad-test-002,status=drifted semantic_vector=0.38,graph_document=0.28,temporal=0.12,overall=0.177 $((NOW - 3600))" \ - "drift,hexad_id=hexad-test-002,status=drifted semantic_vector=0.45,graph_document=0.32,temporal=0.15,overall=0.213 ${NOW}" - -# Drift scores for hexad-test-003 (healthy, zero drift) -write_lines "drift_scores" \ - "drift,hexad_id=hexad-test-003,status=healthy semantic_vector=0.0,graph_document=0.0,temporal=0.0,overall=0.0 ${NOW}" - -echo " Drift scores written (7 data points across 3 hexads)." - -# --------------------------------------------------------------------------- -# Write query latency metrics -# --------------------------------------------------------------------------- - -echo "--- Writing query latency metrics..." - -# Simulate query latency over the last hour (1-minute intervals) -for offset in $(seq 0 5 60); do - TS=$((NOW - offset * 60)) - # Vary latency between 3ms and 45ms - LATENCY=$(( (offset * 7 + 13) % 45 + 3 )) - write_lines "metrics" \ - "query_latency,service=verisimdb,query_type=search duration_ms=${LATENCY}i ${TS}" -done - -# VCL query latency (separate measurement) -for offset in $(seq 0 10 60); do - TS=$((NOW - offset * 60)) - LATENCY=$(( (offset * 11 + 7) % 80 + 5 )) - write_lines "metrics" \ - "query_latency,service=verisimdb,query_type=vcl duration_ms=${LATENCY}i ${TS}" -done - -echo " Query latency metrics written." - -# --------------------------------------------------------------------------- -# Write federation health metrics -# --------------------------------------------------------------------------- - -echo "--- Writing federation health metrics..." - -# Federation adapter health (1 = up, 0 = down) -write_lines "federation_health" \ - "adapter_health,adapter=mongodb,host=mongodb:27017 status=1i,latency_ms=12i ${NOW}" \ - "adapter_health,adapter=redis,host=redis:6379 status=1i,latency_ms=3i ${NOW}" \ - "adapter_health,adapter=neo4j,host=neo4j:7687 status=1i,latency_ms=18i ${NOW}" \ - "adapter_health,adapter=clickhouse,host=clickhouse:8123 status=1i,latency_ms=8i ${NOW}" \ - "adapter_health,adapter=surrealdb,host=surrealdb:8000 status=1i,latency_ms=15i ${NOW}" \ - "adapter_health,adapter=influxdb,host=influxdb:8086 status=1i,latency_ms=5i ${NOW}" \ - "adapter_health,adapter=minio,host=minio:9000 status=1i,latency_ms=7i ${NOW}" - -# Federation sync events -write_lines "federation_health" \ - "federation_sync,source=verisimdb,target=mongodb hexads_synced=150i,errors=0i,duration_ms=1200i $((NOW - 300))" \ - "federation_sync,source=verisimdb,target=redis hexads_synced=150i,errors=2i,duration_ms=350i $((NOW - 300))" \ - "federation_sync,source=verisimdb,target=neo4j hexads_synced=148i,errors=2i,duration_ms=2100i $((NOW - 300))" - -echo " Federation health metrics written." - -# --------------------------------------------------------------------------- -# Summary -# --------------------------------------------------------------------------- - -echo "" -echo "=== InfluxDB 2 seed complete ===" -echo " Buckets: metrics, drift_scores, federation_health" -echo " Drift: 7 data points across 3 hexads" -echo " Latency: ~20 data points (search + VCL)" -echo " Federation: 7 adapter health checks + 3 sync events" -echo "" -echo " Verify with:" -echo " influx query --host ${INFLUX_HOST} --token ${INFLUX_TOKEN} --org ${INFLUX_ORG} \\" -echo " 'from(bucket: \"drift_scores\") |> range(start: -1h) |> limit(n: 5)'" diff --git a/verisimdb/connectors/test-infra/seed/minio-init.sh b/verisimdb/connectors/test-infra/seed/minio-init.sh deleted file mode 100755 index 2e1b8391..00000000 --- a/verisimdb/connectors/test-infra/seed/minio-init.sh +++ /dev/null @@ -1,193 +0,0 @@ -#!/bin/sh -# SPDX-License-Identifier: MPL-2.0 -# -# VeriSimDB Test Infrastructure — MinIO (S3) Seed Script -# -# Creates the verisimdb-objects bucket and uploads test objects including -# sample binary blobs and JSON metadata files. MinIO provides S3-compatible -# object storage for the object storage federation adapter. -# -# Prerequisites: -# - MinIO container running (API on port 9002, console on port 9001) -# - mc (MinIO client) available in PATH -# -# Usage: -# ./minio-init.sh # Default: localhost:9002 -# MINIO_HOST=http://minio:9000 MINIO_USER=admin MINIO_PASS=secret ./minio-init.sh -# -# Author: Jonathan D.A. Jewell - -set -eu - -MINIO_HOST="${MINIO_HOST:-http://localhost:9002}" -MINIO_USER="${MINIO_USER:-verisim}" -MINIO_PASS="${MINIO_PASS:-verisim-test-password}" -MINIO_ALIAS="verisim-test" - -echo "=== VeriSimDB MinIO (S3) Seed ===" -echo " Host: ${MINIO_HOST}" -echo " User: ${MINIO_USER}" - -# --------------------------------------------------------------------------- -# Wait for MinIO to be ready -# --------------------------------------------------------------------------- - -echo "--- Waiting for MinIO..." -for attempt in $(seq 1 30); do - if curl -sf "${MINIO_HOST}/minio/health/live" >/dev/null 2>&1; then - echo " MinIO is ready." - break - fi - if [ "$attempt" -eq 30 ]; then - echo " ERROR: MinIO did not become ready within 30 seconds." - exit 1 - fi - sleep 1 -done - -# --------------------------------------------------------------------------- -# Configure mc alias -# --------------------------------------------------------------------------- - -echo "--- Configuring MinIO client alias..." - -mc alias set "${MINIO_ALIAS}" "${MINIO_HOST}" "${MINIO_USER}" "${MINIO_PASS}" --api S3v4 \ - 2>/dev/null || true - -echo " Alias '${MINIO_ALIAS}' configured." - -# --------------------------------------------------------------------------- -# Create buckets -# --------------------------------------------------------------------------- - -echo "--- Creating buckets..." - -mc mb "${MINIO_ALIAS}/verisimdb-objects" 2>/dev/null || echo " Bucket 'verisimdb-objects' already exists (OK)" -mc mb "${MINIO_ALIAS}/verisimdb-backups" 2>/dev/null || echo " Bucket 'verisimdb-backups' already exists (OK)" -mc mb "${MINIO_ALIAS}/verisimdb-embeddings" 2>/dev/null || echo " Bucket 'verisimdb-embeddings' already exists (OK)" - -echo " Buckets: verisimdb-objects, verisimdb-backups, verisimdb-embeddings" - -# --------------------------------------------------------------------------- -# Create temporary directory for test objects -# --------------------------------------------------------------------------- - -TMPDIR=$(mktemp -d) -trap 'rm -rf "${TMPDIR}"' EXIT - -# --------------------------------------------------------------------------- -# Generate and upload test objects -# --------------------------------------------------------------------------- - -echo "--- Uploading test objects..." - -# Hexad metadata (JSON documents) -cat > "${TMPDIR}/hexad-test-001.json" <<'ENDJSON' -{ - "id": "hexad-test-001", - "title": "Introduction to Cross-Modal Consistency", - "entity_type": "Article", - "version": 3, - "modality_count": 6, - "storage_class": "STANDARD", - "created_at": "2026-02-27T00:00:00Z", - "updated_at": "2026-02-28T00:00:00Z", - "object_refs": { - "document": "hexads/hexad-test-001/document.txt", - "embedding": "embeddings/hexad-test-001.bin", - "provenance": "hexads/hexad-test-001/provenance.cbor" - } -} -ENDJSON - -cat > "${TMPDIR}/hexad-test-002.json" <<'ENDJSON' -{ - "id": "hexad-test-002", - "title": "Drift Detection Algorithms", - "entity_type": "TechArticle", - "version": 1, - "modality_count": 3, - "storage_class": "STANDARD", - "created_at": "2026-02-27T00:00:00Z", - "updated_at": "2026-02-27T23:00:00Z", - "object_refs": { - "document": "hexads/hexad-test-002/document.txt", - "embedding": "embeddings/hexad-test-002.bin" - } -} -ENDJSON - -# Upload hexad metadata -mc cp "${TMPDIR}/hexad-test-001.json" "${MINIO_ALIAS}/verisimdb-objects/hexads/hexad-test-001/metadata.json" -mc cp "${TMPDIR}/hexad-test-002.json" "${MINIO_ALIAS}/verisimdb-objects/hexads/hexad-test-002/metadata.json" - -echo " Uploaded 2 hexad metadata JSON files." - -# Document modality content (plain text) -printf '%s' "VeriSimDB maintains consistency across 8 modality representations. Each entity exists simultaneously as graph, vector, tensor, semantic, document, temporal, provenance, and spatial data." \ - > "${TMPDIR}/document-001.txt" - -printf '%s' "Drift is measured as divergence between modalities using cosine similarity for vectors, Jaccard distance for sets, and temporal decay functions for time-series data." \ - > "${TMPDIR}/document-002.txt" - -mc cp "${TMPDIR}/document-001.txt" "${MINIO_ALIAS}/verisimdb-objects/hexads/hexad-test-001/document.txt" -mc cp "${TMPDIR}/document-002.txt" "${MINIO_ALIAS}/verisimdb-objects/hexads/hexad-test-002/document.txt" - -echo " Uploaded 2 document content files." - -# Binary embedding blobs (simulated — 128 floats = 512 bytes each) -dd if=/dev/urandom bs=512 count=1 2>/dev/null > "${TMPDIR}/embedding-001.bin" -dd if=/dev/urandom bs=512 count=1 2>/dev/null > "${TMPDIR}/embedding-002.bin" - -mc cp "${TMPDIR}/embedding-001.bin" "${MINIO_ALIAS}/verisimdb-embeddings/hexad-test-001.bin" -mc cp "${TMPDIR}/embedding-002.bin" "${MINIO_ALIAS}/verisimdb-embeddings/hexad-test-002.bin" - -echo " Uploaded 2 embedding binary blobs (512 bytes each)." - -# Provenance CBOR placeholder (just a marker file for testing) -printf '{"_note": "CBOR placeholder for testing", "chain_length": 2, "hash": "sha256:a1b2c3d4e5f6"}' \ - > "${TMPDIR}/provenance-001.json" - -mc cp "${TMPDIR}/provenance-001.json" "${MINIO_ALIAS}/verisimdb-objects/hexads/hexad-test-001/provenance.cbor" - -echo " Uploaded 1 provenance placeholder." - -# Backup snapshot (simulated) -printf '{"snapshot_id": "snap-test-001", "timestamp": "2026-02-28T00:00:00Z", "hexad_count": 3, "size_bytes": 4096}' \ - > "${TMPDIR}/snapshot-meta.json" - -mc cp "${TMPDIR}/snapshot-meta.json" "${MINIO_ALIAS}/verisimdb-backups/snapshots/snap-test-001/metadata.json" - -echo " Uploaded 1 backup snapshot metadata." - -# --------------------------------------------------------------------------- -# Set bucket policies (read-only public access for objects bucket) -# --------------------------------------------------------------------------- - -echo "--- Setting bucket policies..." - -# Allow anonymous read on objects bucket (for test convenience) -mc anonymous set download "${MINIO_ALIAS}/verisimdb-objects" 2>/dev/null \ - || echo " Could not set anonymous policy (OK for testing)" - -echo " Policies configured." - -# --------------------------------------------------------------------------- -# Summary -# --------------------------------------------------------------------------- - -echo "" -echo "=== MinIO seed complete ===" -echo " Buckets: verisimdb-objects, verisimdb-backups, verisimdb-embeddings" -echo " Objects: 8 total" -echo " - 2 hexad metadata JSON" -echo " - 2 document content TXT" -echo " - 2 embedding binaries" -echo " - 1 provenance placeholder" -echo " - 1 backup snapshot metadata" -echo "" -echo " Verify with:" -echo " mc ls ${MINIO_ALIAS}/verisimdb-objects/ --recursive" -echo " mc cat ${MINIO_ALIAS}/verisimdb-objects/hexads/hexad-test-001/metadata.json" -echo "" -echo " Console: ${MINIO_HOST%:*}:9001 (user: ${MINIO_USER})" diff --git a/verisimdb/connectors/test-infra/seed/mongodb-init.js b/verisimdb/connectors/test-infra/seed/mongodb-init.js deleted file mode 100644 index ffe6f080..00000000 --- a/verisimdb/connectors/test-infra/seed/mongodb-init.js +++ /dev/null @@ -1,364 +0,0 @@ -// SPDX-License-Identifier: MPL-2.0 -// -// VeriSimDB Test Infrastructure — MongoDB Seed Script -// -// Creates the verisimdb database with a hexads collection and supporting -// indexes. Inserts test hexad documents with multiple modalities (text, -// vector, spatial, temporal) to exercise the MongoDB federation adapter. -// -// This script runs automatically via the /docker-entrypoint-initdb.d/ -// mechanism on first container startup. -// -// Author: Jonathan D.A. Jewell - -// --------------------------------------------------------------------------- -// Initialise replica set (required for change streams) -// --------------------------------------------------------------------------- - -try { - rs.initiate({ - _id: "rs0", - members: [{ _id: 0, host: "localhost:27017" }], - }); - print("Replica set rs0 initiated."); -} catch (e) { - // Already initialised — ignore - print("Replica set already initialised or error: " + e.message); -} - -// Wait for the replica set to become ready -let ready = false; -for (let attempt = 0; attempt < 30; attempt++) { - try { - const status = rs.status(); - if (status.myState === 1) { - ready = true; - break; - } - } catch (_e) { - // Not ready yet - } - sleep(1000); -} - -if (!ready) { - print("WARNING: Replica set did not become PRIMARY within 30 seconds."); -} - -// --------------------------------------------------------------------------- -// Switch to verisimdb database -// --------------------------------------------------------------------------- - -const db = db.getSiblingDB("verisimdb"); - -// --------------------------------------------------------------------------- -// Create collections -// --------------------------------------------------------------------------- - -db.createCollection("hexads"); -db.createCollection("drift_scores"); -db.createCollection("provenance_events"); - -print("Collections created: hexads, drift_scores, provenance_events"); - -// --------------------------------------------------------------------------- -// Create indexes on hexads collection -// --------------------------------------------------------------------------- - -// Unique index on hexad ID -db.hexads.createIndex({ id: 1 }, { unique: true, name: "idx_hexad_id" }); - -// Index on modality type for filtered queries -db.hexads.createIndex( - { "modalities.type": 1 }, - { name: "idx_modality_type" }, -); - -// Text index on document modality content (full-text search) -db.hexads.createIndex( - { "modalities.data.content": "text", "modalities.data.title": "text" }, - { name: "idx_fulltext_content", default_language: "english" }, -); - -// 2dsphere index on spatial modality location (geospatial queries) -db.hexads.createIndex( - { "modalities.data.location": "2dsphere" }, - { name: "idx_spatial_location" }, -); - -// Compound index on created_at + updated_at for temporal queries -db.hexads.createIndex( - { created_at: -1, updated_at: -1 }, - { name: "idx_temporal" }, -); - -// Index on drift scores collection -db.drift_scores.createIndex( - { hexad_id: 1, measured_at: -1 }, - { name: "idx_drift_hexad_time" }, -); - -// Index on provenance events -db.provenance_events.createIndex( - { hexad_id: 1, timestamp: -1 }, - { name: "idx_provenance_hexad_time" }, -); - -print("Indexes created on hexads, drift_scores, provenance_events"); - -// --------------------------------------------------------------------------- -// Insert test hexad documents -// --------------------------------------------------------------------------- - -const now = new Date(); -const oneHourAgo = new Date(now.getTime() - 3600 * 1000); -const oneDayAgo = new Date(now.getTime() - 86400 * 1000); - -db.hexads.insertMany([ - { - id: "hexad-test-001", - created_at: oneDayAgo, - updated_at: now, - version: 3, - modalities: [ - { - type: "document", - data: { - title: "Introduction to Cross-Modal Consistency", - content: - "VeriSimDB maintains consistency across 8 modality representations. " + - "Each entity exists simultaneously as graph, vector, tensor, semantic, " + - "document, temporal, provenance, and spatial data.", - format: "text/plain", - }, - }, - { - type: "vector", - data: { - embedding: Array.from({ length: 128 }, (_, i) => - Math.sin(i * 0.1), - ), - model: "test-embedding-v1", - dimensions: 128, - }, - }, - { - type: "spatial", - data: { - location: { - type: "Point", - coordinates: [-0.1278, 51.5074], // London - }, - radius_km: 5.0, - }, - }, - { - type: "temporal", - data: { - valid_from: oneDayAgo, - valid_to: null, - version_history: [ - { version: 1, timestamp: oneDayAgo, action: "created" }, - { - version: 2, - timestamp: oneHourAgo, - action: "updated_vector", - }, - { - version: 3, - timestamp: now, - action: "updated_document", - }, - ], - }, - }, - { - type: "graph", - data: { - types: [ - "https://schema.org/Article", - "http://verisimdb.org/ontology/Entity", - ], - relationships: [ - { - predicate: "relates_to", - target: "hexad-test-002", - weight: 0.85, - }, - { - predicate: "cites", - target: "hexad-test-003", - weight: 0.72, - }, - ], - }, - }, - { - type: "provenance", - data: { - origin: "manual-import", - actor: "test-seed-script", - chain: [ - { - hash: "sha256:a1b2c3d4e5f6", - action: "created", - timestamp: oneDayAgo, - }, - ], - }, - }, - ], - }, - { - id: "hexad-test-002", - created_at: oneDayAgo, - updated_at: oneHourAgo, - version: 1, - modalities: [ - { - type: "document", - data: { - title: "Drift Detection Algorithms", - content: - "Drift is measured as divergence between modalities using cosine " + - "similarity for vectors, Jaccard distance for sets, and temporal " + - "decay functions for time-series data.", - format: "text/plain", - }, - }, - { - type: "vector", - data: { - embedding: Array.from({ length: 128 }, (_, i) => - Math.cos(i * 0.1), - ), - model: "test-embedding-v1", - dimensions: 128, - }, - }, - { - type: "spatial", - data: { - location: { - type: "Point", - coordinates: [-73.9857, 40.7484], // New York - }, - radius_km: 10.0, - }, - }, - { - type: "graph", - data: { - types: ["https://schema.org/TechArticle"], - relationships: [ - { - predicate: "relates_to", - target: "hexad-test-001", - weight: 0.85, - }, - ], - }, - }, - ], - }, - { - id: "hexad-test-003", - created_at: now, - updated_at: now, - version: 1, - modalities: [ - { - type: "document", - data: { - title: "Self-Normalisation Process", - content: - "When drift exceeds configurable thresholds, the normaliser identifies " + - "the most authoritative modality, regenerates drifted representations, " + - "validates consistency, and updates all modalities atomically.", - format: "text/plain", - }, - }, - { - type: "vector", - data: { - embedding: Array.from({ length: 128 }, (_, i) => - Math.sin(i * 0.2) + Math.cos(i * 0.3), - ), - model: "test-embedding-v1", - dimensions: 128, - }, - }, - { - type: "semantic", - data: { - categories: ["normalisation", "consistency", "drift"], - confidence: 0.92, - proof_blob: "cbor:test-placeholder", - }, - }, - ], - }, -]); - -print("Inserted 3 test hexad documents"); - -// --------------------------------------------------------------------------- -// Insert test drift scores -// --------------------------------------------------------------------------- - -db.drift_scores.insertMany([ - { - hexad_id: "hexad-test-001", - measured_at: now, - scores: { - semantic_vector_drift: 0.12, - graph_document_drift: 0.05, - temporal_consistency_drift: 0.02, - tensor_drift: 0.0, - schema_drift: 0.0, - quality_drift: 0.08, - }, - overall: 0.045, - status: "healthy", - }, - { - hexad_id: "hexad-test-002", - measured_at: now, - scores: { - semantic_vector_drift: 0.45, - graph_document_drift: 0.32, - temporal_consistency_drift: 0.15, - tensor_drift: 0.0, - schema_drift: 0.08, - quality_drift: 0.28, - }, - overall: 0.213, - status: "drifted", - }, -]); - -print("Inserted 2 test drift score documents"); - -// --------------------------------------------------------------------------- -// Insert test provenance events -// --------------------------------------------------------------------------- - -db.provenance_events.insertMany([ - { - hexad_id: "hexad-test-001", - event_type: "created", - timestamp: oneDayAgo, - actor: "test-seed-script", - details: { source: "mongodb-init.js", method: "direct-insert" }, - }, - { - hexad_id: "hexad-test-001", - event_type: "modality_updated", - timestamp: now, - actor: "test-seed-script", - details: { modality: "vector", reason: "embedding regenerated" }, - }, -]); - -print("Inserted 2 test provenance events"); -print("MongoDB seed complete."); diff --git a/verisimdb/connectors/test-infra/seed/neo4j-init.cypher b/verisimdb/connectors/test-infra/seed/neo4j-init.cypher deleted file mode 100644 index 44ae2470..00000000 --- a/verisimdb/connectors/test-infra/seed/neo4j-init.cypher +++ /dev/null @@ -1,205 +0,0 @@ -// SPDX-License-Identifier: MPL-2.0 -// -// VeriSimDB Test Infrastructure — Neo4j Seed Script (Cypher) -// -// Creates constraints, indexes, test nodes, relationships, and provenance -// chains for integration testing of the Neo4j graph federation adapter. -// -// Execute with: -// cat neo4j-init.cypher | cypher-shell -a bolt://localhost:7687 -// -// Or from the Neo4j browser at http://localhost:7474 -// -// Author: Jonathan D.A. Jewell - -// --------------------------------------------------------------------------- -// Constraints — uniqueness guarantees -// --------------------------------------------------------------------------- - -CREATE CONSTRAINT hexad_id_unique IF NOT EXISTS -FOR (h:Hexad) -REQUIRE h.id IS UNIQUE; - -CREATE CONSTRAINT provenance_event_id_unique IF NOT EXISTS -FOR (p:ProvenanceEvent) -REQUIRE p.event_id IS UNIQUE; - -// --------------------------------------------------------------------------- -// Indexes — query performance -// --------------------------------------------------------------------------- - -// Full-text index on hexad document content -CREATE FULLTEXT INDEX hexad_fulltext IF NOT EXISTS -FOR (h:Hexad) -ON EACH [h.title, h.content]; - -// Index on modality type -CREATE INDEX hexad_modality_type IF NOT EXISTS -FOR (h:Hexad) -ON (h.primary_modality); - -// Index on creation timestamp -CREATE INDEX hexad_created_at IF NOT EXISTS -FOR (h:Hexad) -ON (h.created_at); - -// Index on drift status -CREATE INDEX hexad_drift_status IF NOT EXISTS -FOR (h:Hexad) -ON (h.drift_status); - -// Index on entity type -CREATE INDEX hexad_entity_type IF NOT EXISTS -FOR (h:Hexad) -ON (h.entity_type); - -// --------------------------------------------------------------------------- -// Test Hexad Nodes -// --------------------------------------------------------------------------- - -// Hexad 1: Cross-modal consistency article -MERGE (h1:Hexad:Article:Entity {id: 'hexad-test-001'}) -SET h1.title = 'Introduction to Cross-Modal Consistency', - h1.content = 'VeriSimDB maintains consistency across 8 modality representations. Each entity exists simultaneously as graph, vector, tensor, semantic, document, temporal, provenance, and spatial data.', - h1.entity_type = 'Article', - h1.primary_modality = 'document', - h1.version = 3, - h1.created_at = datetime() - duration('P1D'), - h1.updated_at = datetime(), - h1.drift_status = 'healthy', - h1.drift_score = 0.045, - h1.embedding_model = 'test-embedding-v1', - h1.embedding_dimensions = 128, - h1.spatial_lat = 51.5074, - h1.spatial_lon = -0.1278; - -// Hexad 2: Drift detection article -MERGE (h2:Hexad:TechArticle {id: 'hexad-test-002'}) -SET h2.title = 'Drift Detection Algorithms', - h2.content = 'Drift is measured as divergence between modalities using cosine similarity for vectors, Jaccard distance for sets, and temporal decay functions for time-series data.', - h2.entity_type = 'TechArticle', - h2.primary_modality = 'document', - h2.version = 1, - h2.created_at = datetime() - duration('P1D'), - h2.updated_at = datetime() - duration('PT1H'), - h2.drift_status = 'drifted', - h2.drift_score = 0.213, - h2.embedding_model = 'test-embedding-v1', - h2.embedding_dimensions = 128, - h2.spatial_lat = 40.7484, - h2.spatial_lon = -73.9857; - -// Hexad 3: Self-normalisation article -MERGE (h3:Hexad:Article {id: 'hexad-test-003'}) -SET h3.title = 'Self-Normalisation Process', - h3.content = 'When drift exceeds configurable thresholds, the normaliser identifies the most authoritative modality, regenerates drifted representations, validates consistency, and updates all modalities atomically.', - h3.entity_type = 'Article', - h3.primary_modality = 'document', - h3.version = 1, - h3.created_at = datetime(), - h3.updated_at = datetime(), - h3.drift_status = 'healthy', - h3.drift_score = 0.0; - -// Hexad 4: Federation concept -MERGE (h4:Hexad:Concept {id: 'hexad-test-004'}) -SET h4.title = 'Heterogeneous Federation', - h4.content = 'VeriSimDB can coordinate across PostgreSQL, ArangoDB, Elasticsearch, MongoDB, Redis, Neo4j, ClickHouse, SurrealDB, InfluxDB, DuckDB, SQLite, and other VeriSimDB instances.', - h4.entity_type = 'Concept', - h4.primary_modality = 'graph', - h4.version = 1, - h4.created_at = datetime(), - h4.updated_at = datetime(), - h4.drift_status = 'healthy', - h4.drift_score = 0.0; - -// --------------------------------------------------------------------------- -// Relationships -// --------------------------------------------------------------------------- - -// RELATES_TO: bidirectional conceptual link -MATCH (h1:Hexad {id: 'hexad-test-001'}), (h2:Hexad {id: 'hexad-test-002'}) -MERGE (h1)-[r1:RELATES_TO]->(h2) -SET r1.weight = 0.85, r1.since = datetime() - duration('P1D'); - -MATCH (h2:Hexad {id: 'hexad-test-002'}), (h1:Hexad {id: 'hexad-test-001'}) -MERGE (h2)-[r2:RELATES_TO]->(h1) -SET r2.weight = 0.85, r2.since = datetime() - duration('P1D'); - -// CITES: directed citation -MATCH (h1:Hexad {id: 'hexad-test-001'}), (h3:Hexad {id: 'hexad-test-003'}) -MERGE (h1)-[c1:CITES]->(h3) -SET c1.weight = 0.72, c1.context = 'normalisation reference'; - -// PART_OF: concept containment -MATCH (h1:Hexad {id: 'hexad-test-001'}), (h4:Hexad {id: 'hexad-test-004'}) -MERGE (h1)-[p1:PART_OF]->(h4) -SET p1.role = 'consistency-component'; - -MATCH (h2:Hexad {id: 'hexad-test-002'}), (h4:Hexad {id: 'hexad-test-004'}) -MERGE (h2)-[p2:PART_OF]->(h4) -SET p2.role = 'drift-detection-component'; - -MATCH (h3:Hexad {id: 'hexad-test-003'}), (h4:Hexad {id: 'hexad-test-004'}) -MERGE (h3)-[p3:PART_OF]->(h4) -SET p3.role = 'normalisation-component'; - -// --------------------------------------------------------------------------- -// Provenance chain -// --------------------------------------------------------------------------- - -// Provenance events for hexad-test-001 -MERGE (pe1:ProvenanceEvent {event_id: 'prov-001-create'}) -SET pe1.hexad_id = 'hexad-test-001', - pe1.event_type = 'created', - pe1.actor = 'test-seed-script', - pe1.timestamp = datetime() - duration('P1D'), - pe1.hash = 'sha256:a1b2c3d4e5f6', - pe1.details = 'Initial creation via neo4j-init.cypher'; - -MERGE (pe2:ProvenanceEvent {event_id: 'prov-001-update-vector'}) -SET pe2.hexad_id = 'hexad-test-001', - pe2.event_type = 'modality_updated', - pe2.actor = 'test-seed-script', - pe2.timestamp = datetime() - duration('PT1H'), - pe2.hash = 'sha256:f6e5d4c3b2a1', - pe2.details = 'Vector embedding regenerated'; - -// Chain provenance events in order -MATCH (h1:Hexad {id: 'hexad-test-001'}), (pe1:ProvenanceEvent {event_id: 'prov-001-create'}) -MERGE (h1)-[:HAS_PROVENANCE]->(pe1); - -MATCH (pe1:ProvenanceEvent {event_id: 'prov-001-create'}), - (pe2:ProvenanceEvent {event_id: 'prov-001-update-vector'}) -MERGE (pe1)-[:FOLLOWED_BY]->(pe2); - -MATCH (h1:Hexad {id: 'hexad-test-001'}), (pe2:ProvenanceEvent {event_id: 'prov-001-update-vector'}) -MERGE (h1)-[:HAS_PROVENANCE]->(pe2); - -// --------------------------------------------------------------------------- -// Semantic type nodes (ontology) -// --------------------------------------------------------------------------- - -MERGE (t1:OntologyType {uri: 'http://schema.org/Article'}) -SET t1.label = 'Article'; - -MERGE (t2:OntologyType {uri: 'http://schema.org/TechArticle'}) -SET t2.label = 'TechArticle'; - -MERGE (t3:OntologyType {uri: 'http://verisimdb.org/ontology/Entity'}) -SET t3.label = 'VeriSimDB Entity'; - -// Type relationships -MATCH (h1:Hexad {id: 'hexad-test-001'}), (t1:OntologyType {uri: 'http://schema.org/Article'}) -MERGE (h1)-[:IS_TYPE]->(t1); - -MATCH (h1:Hexad {id: 'hexad-test-001'}), (t3:OntologyType {uri: 'http://verisimdb.org/ontology/Entity'}) -MERGE (h1)-[:IS_TYPE]->(t3); - -MATCH (h2:Hexad {id: 'hexad-test-002'}), (t2:OntologyType {uri: 'http://schema.org/TechArticle'}) -MERGE (h2)-[:IS_TYPE]->(t2); - -// TechArticle is subclass of Article -MATCH (t2:OntologyType {uri: 'http://schema.org/TechArticle'}), - (t1:OntologyType {uri: 'http://schema.org/Article'}) -MERGE (t2)-[:SUBCLASS_OF]->(t1); diff --git a/verisimdb/connectors/test-infra/seed/redis-init.sh b/verisimdb/connectors/test-infra/seed/redis-init.sh deleted file mode 100755 index d2cfde49..00000000 --- a/verisimdb/connectors/test-infra/seed/redis-init.sh +++ /dev/null @@ -1,286 +0,0 @@ -#!/bin/sh -# SPDX-License-Identifier: MPL-2.0 -# -# VeriSimDB Test Infrastructure — Redis Stack Seed Script -# -# Creates RediSearch indexes for hexad data, loads RedisJSON documents, -# and creates RedisTimeSeries keys for temporal modality data. -# -# Prerequisites: -# - Redis Stack container running on localhost:6379 -# - redis-cli available in PATH -# -# Usage: -# ./redis-init.sh # Default: localhost:6379 -# REDIS_HOST=redis REDIS_PORT=6379 ./redis-init.sh -# -# Author: Jonathan D.A. Jewell - -set -eu - -REDIS_HOST="${REDIS_HOST:-localhost}" -REDIS_PORT="${REDIS_PORT:-6379}" - -CLI="redis-cli -h ${REDIS_HOST} -p ${REDIS_PORT}" - -echo "=== VeriSimDB Redis Stack Seed ===" -echo " Host: ${REDIS_HOST}:${REDIS_PORT}" - -# --------------------------------------------------------------------------- -# Wait for Redis to be ready -# --------------------------------------------------------------------------- - -echo "--- Waiting for Redis..." -for attempt in $(seq 1 30); do - if ${CLI} ping 2>/dev/null | grep -q PONG; then - echo " Redis is ready." - break - fi - if [ "$attempt" -eq 30 ]; then - echo " ERROR: Redis did not become ready within 30 seconds." - exit 1 - fi - sleep 1 -done - -# --------------------------------------------------------------------------- -# Create RediSearch indexes -# --------------------------------------------------------------------------- - -echo "--- Creating RediSearch indexes..." - -# Hexad document search index -# Indexes JSON documents stored at hexad:* keys -${CLI} FT.CREATE idx:hexads ON JSON PREFIX 1 "hexad:" SCHEMA \ - '$.id' AS id TAG SORTABLE \ - '$.modalities[?(@.type=="document")].data.title' AS title TEXT WEIGHT 2.0 \ - '$.modalities[?(@.type=="document")].data.content' AS content TEXT WEIGHT 1.0 \ - '$.modalities[?(@.type=="graph")].data.types[*]' AS entity_type TAG \ - '$.version' AS version NUMERIC SORTABLE \ - '$.created_at' AS created_at NUMERIC SORTABLE \ - '$.updated_at' AS updated_at NUMERIC SORTABLE \ - 2>/dev/null || echo " Index idx:hexads already exists (OK)" - -# Drift scores search index -${CLI} FT.CREATE idx:drift ON JSON PREFIX 1 "drift:" SCHEMA \ - '$.hexad_id' AS hexad_id TAG SORTABLE \ - '$.overall' AS overall NUMERIC SORTABLE \ - '$.status' AS status TAG \ - '$.measured_at' AS measured_at NUMERIC SORTABLE \ - 2>/dev/null || echo " Index idx:drift already exists (OK)" - -echo " RediSearch indexes created." - -# --------------------------------------------------------------------------- -# Load RedisJSON documents — test hexads -# --------------------------------------------------------------------------- - -echo "--- Loading RedisJSON hexad documents..." - -NOW=$(date +%s) -ONE_HOUR_AGO=$((NOW - 3600)) -ONE_DAY_AGO=$((NOW - 86400)) - -# Hexad 1: multi-modality entity with document, vector, graph, temporal -${CLI} JSON.SET "hexad:test-001" '$' "$(cat </dev/null || echo " ts:drift:test-001:overall already exists (OK)" - -${CLI} TS.CREATE "ts:drift:test-001:semantic_vector" \ - RETENTION 86400000 \ - LABELS hexad_id hexad-test-001 metric semantic_vector_drift \ - 2>/dev/null || echo " ts:drift:test-001:semantic_vector already exists (OK)" - -# Add sample data points (timestamps in milliseconds) -NOW_MS=$((NOW * 1000)) -for offset in 0 300 600 900 1200 1500 1800; do - TS=$((NOW_MS - offset * 1000)) - # Simulate gradually increasing drift - DRIFT_VAL=$(echo "scale=3; 0.01 + ${offset} * 0.00002" | bc 2>/dev/null || echo "0.045") - ${CLI} TS.ADD "ts:drift:test-001:overall" "${TS}" "${DRIFT_VAL}" 2>/dev/null || true -done - -# Query latency time-series -${CLI} TS.CREATE "ts:query:latency_ms" \ - RETENTION 86400000 \ - LABELS service verisimdb metric query_latency_ms \ - 2>/dev/null || echo " ts:query:latency_ms already exists (OK)" - -for offset in 0 60 120 180 240 300; do - TS=$((NOW_MS - offset * 1000)) - ${CLI} TS.ADD "ts:query:latency_ms" "${TS}" "$(( (offset % 50) + 5 ))" 2>/dev/null || true -done - -echo " RedisTimeSeries keys created with sample data." - -# --------------------------------------------------------------------------- -# Summary -# --------------------------------------------------------------------------- - -echo "" -echo "=== Redis Stack seed complete ===" -echo " Hexads: 3 (hexad:test-001, hexad:test-002, hexad:test-003)" -echo " Drift: 2 (drift:test-001, drift:test-002)" -echo " TimeSeries: 3 keys with sample data points" -echo " Indexes: idx:hexads, idx:drift" -echo "" -echo " Verify with:" -echo " redis-cli -h ${REDIS_HOST} -p ${REDIS_PORT} FT.SEARCH idx:hexads '*'" -echo " redis-cli -h ${REDIS_HOST} -p ${REDIS_PORT} JSON.GET hexad:test-001" diff --git a/verisimdb/connectors/test-infra/seed/surrealdb-init.surql b/verisimdb/connectors/test-infra/seed/surrealdb-init.surql deleted file mode 100644 index 88a95de7..00000000 --- a/verisimdb/connectors/test-infra/seed/surrealdb-init.surql +++ /dev/null @@ -1,260 +0,0 @@ --- SPDX-License-Identifier: MPL-2.0 --- --- VeriSimDB Test Infrastructure — SurrealDB Seed Script --- --- Defines namespace, database, tables with schemafull definitions, --- edge tables for relationships, and inserts test data. SurrealDB --- is tested as a multi-model federation target — it supports both --- document and graph queries natively. --- --- Execute with: --- cat surrealdb-init.surql | surreal sql --endpoint http://localhost:8000 --ns verisimdb --db test --- --- Author: Jonathan D.A. Jewell - --- --------------------------------------------------------------------------- --- Namespace and database --- --------------------------------------------------------------------------- - -DEFINE NAMESPACE IF NOT EXISTS verisimdb; -USE NS verisimdb; - -DEFINE DATABASE IF NOT EXISTS test; -USE DB test; - --- --------------------------------------------------------------------------- --- Table definitions (schemafull) --- --------------------------------------------------------------------------- - --- Hexads: core entity table -DEFINE TABLE IF NOT EXISTS hexads SCHEMAFULL; -DEFINE FIELD IF NOT EXISTS title ON TABLE hexads TYPE string; -DEFINE FIELD IF NOT EXISTS content ON TABLE hexads TYPE string; -DEFINE FIELD IF NOT EXISTS entity_type ON TABLE hexads TYPE string; -DEFINE FIELD IF NOT EXISTS primary_modality ON TABLE hexads TYPE string; -DEFINE FIELD IF NOT EXISTS version ON TABLE hexads TYPE int; -DEFINE FIELD IF NOT EXISTS drift_status ON TABLE hexads TYPE string - ASSERT $value INSIDE ['healthy', 'drifted', 'normalising', 'stale']; -DEFINE FIELD IF NOT EXISTS drift_score ON TABLE hexads TYPE float; -DEFINE FIELD IF NOT EXISTS tags ON TABLE hexads TYPE array; -DEFINE FIELD IF NOT EXISTS tags.* ON TABLE hexads TYPE string; -DEFINE FIELD IF NOT EXISTS created_at ON TABLE hexads TYPE datetime; -DEFINE FIELD IF NOT EXISTS updated_at ON TABLE hexads TYPE datetime; - --- Modalities: individual modality records -DEFINE TABLE IF NOT EXISTS modalities SCHEMAFULL; -DEFINE FIELD IF NOT EXISTS hexad_id ON TABLE modalities TYPE record; -DEFINE FIELD IF NOT EXISTS modality_type ON TABLE modalities TYPE string - ASSERT $value INSIDE ['graph', 'vector', 'tensor', 'semantic', 'document', 'temporal', 'provenance', 'spatial']; -DEFINE FIELD IF NOT EXISTS data_size_bytes ON TABLE modalities TYPE int; -DEFINE FIELD IF NOT EXISTS is_authoritative ON TABLE modalities TYPE bool; -DEFINE FIELD IF NOT EXISTS quality_score ON TABLE modalities TYPE float; -DEFINE FIELD IF NOT EXISTS metadata ON TABLE modalities TYPE object; -DEFINE FIELD IF NOT EXISTS last_updated ON TABLE modalities TYPE datetime; - --- Drift scores: time-series measurements -DEFINE TABLE IF NOT EXISTS drift_scores SCHEMAFULL; -DEFINE FIELD IF NOT EXISTS hexad_id ON TABLE drift_scores TYPE record; -DEFINE FIELD IF NOT EXISTS measured_at ON TABLE drift_scores TYPE datetime; -DEFINE FIELD IF NOT EXISTS semantic_vector_drift ON TABLE drift_scores TYPE float; -DEFINE FIELD IF NOT EXISTS graph_document_drift ON TABLE drift_scores TYPE float; -DEFINE FIELD IF NOT EXISTS temporal_consistency_drift ON TABLE drift_scores TYPE float; -DEFINE FIELD IF NOT EXISTS tensor_drift ON TABLE drift_scores TYPE float; -DEFINE FIELD IF NOT EXISTS schema_drift ON TABLE drift_scores TYPE float; -DEFINE FIELD IF NOT EXISTS quality_drift ON TABLE drift_scores TYPE float; -DEFINE FIELD IF NOT EXISTS overall ON TABLE drift_scores TYPE float; -DEFINE FIELD IF NOT EXISTS status ON TABLE drift_scores TYPE string; - --- Provenance events: lineage tracking -DEFINE TABLE IF NOT EXISTS provenance_events SCHEMAFULL; -DEFINE FIELD IF NOT EXISTS hexad_id ON TABLE provenance_events TYPE record; -DEFINE FIELD IF NOT EXISTS event_type ON TABLE provenance_events TYPE string; -DEFINE FIELD IF NOT EXISTS actor ON TABLE provenance_events TYPE string; -DEFINE FIELD IF NOT EXISTS timestamp ON TABLE provenance_events TYPE datetime; -DEFINE FIELD IF NOT EXISTS hash ON TABLE provenance_events TYPE string; -DEFINE FIELD IF NOT EXISTS parent_hash ON TABLE provenance_events TYPE option; -DEFINE FIELD IF NOT EXISTS details ON TABLE provenance_events TYPE string; - --- --------------------------------------------------------------------------- --- Edge tables (graph relationships between hexads) --- --------------------------------------------------------------------------- - --- relates_to: bidirectional conceptual link -DEFINE TABLE IF NOT EXISTS relates_to SCHEMAFULL TYPE RELATION IN hexads OUT hexads; -DEFINE FIELD IF NOT EXISTS weight ON TABLE relates_to TYPE float; -DEFINE FIELD IF NOT EXISTS since ON TABLE relates_to TYPE datetime; - --- cites: directed citation -DEFINE TABLE IF NOT EXISTS cites SCHEMAFULL TYPE RELATION IN hexads OUT hexads; -DEFINE FIELD IF NOT EXISTS weight ON TABLE cites TYPE float; -DEFINE FIELD IF NOT EXISTS context ON TABLE cites TYPE string; - --- derived_from: provenance derivation chain -DEFINE TABLE IF NOT EXISTS derived_from SCHEMAFULL TYPE RELATION IN hexads OUT hexads; -DEFINE FIELD IF NOT EXISTS derivation_type ON TABLE derived_from TYPE string; -DEFINE FIELD IF NOT EXISTS confidence ON TABLE derived_from TYPE float; -DEFINE FIELD IF NOT EXISTS timestamp ON TABLE derived_from TYPE datetime; - --- part_of: hierarchical containment -DEFINE TABLE IF NOT EXISTS part_of SCHEMAFULL TYPE RELATION IN hexads OUT hexads; -DEFINE FIELD IF NOT EXISTS role ON TABLE part_of TYPE string; - --- --------------------------------------------------------------------------- --- Indexes --- --------------------------------------------------------------------------- - -DEFINE INDEX IF NOT EXISTS idx_hexads_entity_type ON TABLE hexads FIELDS entity_type; -DEFINE INDEX IF NOT EXISTS idx_hexads_drift_status ON TABLE hexads FIELDS drift_status; -DEFINE INDEX IF NOT EXISTS idx_hexads_created_at ON TABLE hexads FIELDS created_at; -DEFINE INDEX IF NOT EXISTS idx_modalities_type ON TABLE modalities FIELDS modality_type; -DEFINE INDEX IF NOT EXISTS idx_drift_hexad ON TABLE drift_scores FIELDS hexad_id; -DEFINE INDEX IF NOT EXISTS idx_provenance_hexad ON TABLE provenance_events FIELDS hexad_id; - --- Full-text search index on hexad content -DEFINE ANALYZER IF NOT EXISTS hexad_analyzer TOKENIZERS blank, class FILTERS lowercase, snowball(english); -DEFINE INDEX IF NOT EXISTS idx_hexads_fulltext ON TABLE hexads FIELDS title, content SEARCH ANALYZER hexad_analyzer BM25; - --- --------------------------------------------------------------------------- --- Insert test data — hexads --- --------------------------------------------------------------------------- - -CREATE hexads:test001 SET - title = 'Introduction to Cross-Modal Consistency', - content = 'VeriSimDB maintains consistency across 8 modality representations. Each entity exists simultaneously as graph, vector, tensor, semantic, document, temporal, provenance, and spatial data.', - entity_type = 'Article', - primary_modality = 'document', - version = 3, - drift_status = 'healthy', - drift_score = 0.045, - tags = ['consistency', 'cross-modal', 'verisimdb'], - created_at = time::now() - 1d, - updated_at = time::now(); - -CREATE hexads:test002 SET - title = 'Drift Detection Algorithms', - content = 'Drift is measured as divergence between modalities using cosine similarity for vectors, Jaccard distance for sets, and temporal decay functions for time-series data.', - entity_type = 'TechArticle', - primary_modality = 'document', - version = 1, - drift_status = 'drifted', - drift_score = 0.213, - tags = ['drift', 'algorithms', 'cosine-similarity'], - created_at = time::now() - 1d, - updated_at = time::now() - 1h; - -CREATE hexads:test003 SET - title = 'Self-Normalisation Process', - content = 'When drift exceeds configurable thresholds, the normaliser identifies the most authoritative modality, regenerates drifted representations, validates consistency, and updates all modalities atomically.', - entity_type = 'Article', - primary_modality = 'document', - version = 1, - drift_status = 'healthy', - drift_score = 0.0, - tags = ['normalisation', 'consistency', 'drift'], - created_at = time::now(), - updated_at = time::now(); - -CREATE hexads:test004 SET - title = 'Heterogeneous Federation', - content = 'VeriSimDB can coordinate across PostgreSQL, ArangoDB, Elasticsearch, MongoDB, Redis, Neo4j, ClickHouse, SurrealDB, InfluxDB, DuckDB, SQLite, and other VeriSimDB instances.', - entity_type = 'Concept', - primary_modality = 'graph', - version = 1, - drift_status = 'healthy', - drift_score = 0.0, - tags = ['federation', 'heterogeneous', 'multi-database'], - created_at = time::now(), - updated_at = time::now(); - --- --------------------------------------------------------------------------- --- Insert test data — edge relationships --- --------------------------------------------------------------------------- - -RELATE hexads:test001 -> relates_to -> hexads:test002 SET - weight = 0.85, - since = time::now() - 1d; - -RELATE hexads:test002 -> relates_to -> hexads:test001 SET - weight = 0.85, - since = time::now() - 1d; - -RELATE hexads:test001 -> cites -> hexads:test003 SET - weight = 0.72, - context = 'normalisation reference'; - -RELATE hexads:test001 -> part_of -> hexads:test004 SET - role = 'consistency-component'; - -RELATE hexads:test002 -> part_of -> hexads:test004 SET - role = 'drift-detection-component'; - -RELATE hexads:test003 -> part_of -> hexads:test004 SET - role = 'normalisation-component'; - -RELATE hexads:test003 -> derived_from -> hexads:test001 SET - derivation_type = 'conceptual-extraction', - confidence = 0.88, - timestamp = time::now(); - --- --------------------------------------------------------------------------- --- Insert test data — modalities --- --------------------------------------------------------------------------- - -CREATE modalities:m001 SET - hexad_id = hexads:test001, modality_type = 'document', data_size_bytes = 1024, - is_authoritative = true, quality_score = 0.95, - metadata = { format: 'text/plain' }, last_updated = time::now(); - -CREATE modalities:m002 SET - hexad_id = hexads:test001, modality_type = 'vector', data_size_bytes = 512, - is_authoritative = false, quality_score = 0.88, - metadata = { model: 'test-embedding-v1', dimensions: 128 }, last_updated = time::now(); - -CREATE modalities:m003 SET - hexad_id = hexads:test001, modality_type = 'graph', data_size_bytes = 256, - is_authoritative = false, quality_score = 0.92, - metadata = { triples: 5, types: 2 }, last_updated = time::now(); - -CREATE modalities:m004 SET - hexad_id = hexads:test002, modality_type = 'document', data_size_bytes = 896, - is_authoritative = true, quality_score = 0.82, - metadata = { format: 'text/plain' }, last_updated = time::now() - 1h; - -CREATE modalities:m005 SET - hexad_id = hexads:test002, modality_type = 'vector', data_size_bytes = 512, - is_authoritative = false, quality_score = 0.55, - metadata = { model: 'test-embedding-v1', dimensions: 128 }, last_updated = time::now() - 1h; - --- --------------------------------------------------------------------------- --- Insert test data — drift scores --- --------------------------------------------------------------------------- - -CREATE drift_scores:ds001 SET - hexad_id = hexads:test001, measured_at = time::now(), - semantic_vector_drift = 0.12, graph_document_drift = 0.05, - temporal_consistency_drift = 0.02, tensor_drift = 0.0, - schema_drift = 0.0, quality_drift = 0.08, - overall = 0.045, status = 'healthy'; - -CREATE drift_scores:ds002 SET - hexad_id = hexads:test002, measured_at = time::now(), - semantic_vector_drift = 0.45, graph_document_drift = 0.32, - temporal_consistency_drift = 0.15, tensor_drift = 0.0, - schema_drift = 0.08, quality_drift = 0.28, - overall = 0.213, status = 'drifted'; - --- --------------------------------------------------------------------------- --- Insert test data — provenance events --- --------------------------------------------------------------------------- - -CREATE provenance_events:pe001 SET - hexad_id = hexads:test001, event_type = 'created', - actor = 'test-seed-script', timestamp = time::now() - 1d, - hash = 'sha256:a1b2c3d4e5f6', parent_hash = NONE, - details = 'Initial creation via surrealdb-init.surql'; - -CREATE provenance_events:pe002 SET - hexad_id = hexads:test001, event_type = 'modality_updated', - actor = 'test-seed-script', timestamp = time::now() - 1h, - hash = 'sha256:f6e5d4c3b2a1', parent_hash = 'sha256:a1b2c3d4e5f6', - details = 'Vector embedding regenerated'; diff --git a/verisimdb/connectors/test-infra/vordr.toml b/verisimdb/connectors/test-infra/vordr.toml deleted file mode 100644 index da40e926..00000000 --- a/verisimdb/connectors/test-infra/vordr.toml +++ /dev/null @@ -1,174 +0,0 @@ -# SPDX-License-Identifier: MPL-2.0 -# -# Vordr runtime monitoring configuration for VeriSimDB test infrastructure -# -# Monitors health endpoints of all 7 database services in the test stack. -# Alerts on container crashes, unhealthy services, and port unreachability. -# Logs to stdout for test-time visibility. -# -# See: stapeln/container-stack/vordr/ -# -# Author: Jonathan D.A. Jewell - -[metadata] -name = "verisimdb-test-infra-monitor" -version = "0.1.0" -description = "Runtime health monitoring for VeriSimDB test database stack" - -# --------------------------------------------------------------------------- -# Global settings -# --------------------------------------------------------------------------- - -[settings] -# Check interval for all probes (seconds) -check_interval = 10 - -# Number of consecutive failures before alerting -failure_threshold = 3 - -# Number of consecutive successes to clear alert -recovery_threshold = 1 - -# Log output target -log_target = "stdout" - -# Log format: "text" for human-readable, "json" for structured -log_format = "text" - -# Log level: "debug", "info", "warn", "error" -log_level = "info" - -# --------------------------------------------------------------------------- -# Service probes — health endpoints for each database -# --------------------------------------------------------------------------- - -[[probes]] -name = "mongodb" -type = "tcp" -host = "localhost" -port = 27017 -timeout = 5 -description = "MongoDB replica set (document store federation)" - -[[probes]] -name = "mongodb-command" -type = "exec" -command = "mongosh --eval 'db.adminCommand({ping:1})' --quiet mongodb://localhost:27017" -timeout = 10 -description = "MongoDB command-level health (replica set status)" - -[[probes]] -name = "redis-stack" -type = "exec" -command = "redis-cli -h localhost -p 6379 ping" -expect = "PONG" -timeout = 5 -description = "Redis Stack (RediSearch + RedisJSON + RedisTimeSeries)" - -[[probes]] -name = "neo4j-http" -type = "http" -url = "http://localhost:7474/" -method = "GET" -expected_status = 200 -timeout = 5 -description = "Neo4j HTTP API (browser + REST)" - -[[probes]] -name = "neo4j-bolt" -type = "tcp" -host = "localhost" -port = 7687 -timeout = 5 -description = "Neo4j Bolt protocol (driver connections)" - -[[probes]] -name = "clickhouse-http" -type = "http" -url = "http://localhost:8123/ping" -method = "GET" -expected_status = 200 -timeout = 5 -description = "ClickHouse HTTP interface (queries + health)" - -[[probes]] -name = "clickhouse-native" -type = "tcp" -host = "localhost" -port = 9000 -timeout = 5 -description = "ClickHouse native TCP protocol (client connections)" - -[[probes]] -name = "surrealdb" -type = "http" -url = "http://localhost:8000/health" -method = "GET" -expected_status = 200 -timeout = 5 -description = "SurrealDB HTTP API (document + graph)" - -[[probes]] -name = "influxdb" -type = "http" -url = "http://localhost:8086/health" -method = "GET" -expected_status = 200 -timeout = 5 -description = "InfluxDB 2 HTTP API (time-series)" - -[[probes]] -name = "minio-api" -type = "http" -url = "http://localhost:9002/minio/health/live" -method = "GET" -expected_status = 200 -timeout = 5 -description = "MinIO S3 API (object storage)" - -[[probes]] -name = "minio-console" -type = "tcp" -host = "localhost" -port = 9001 -timeout = 5 -description = "MinIO web console" - -# --------------------------------------------------------------------------- -# Alerts — actions on failure/recovery -# --------------------------------------------------------------------------- - -[alerts] -# Log to stdout on state change (default for test environments) -on_failure = "log" -on_recovery = "log" - -# In production, these would be: -# on_failure = "webhook" -# webhook_url = "https://hooks.example.com/verisimdb-alerts" -# on_recovery = "webhook" - -# --------------------------------------------------------------------------- -# Container monitoring — detect crashes and restarts -# --------------------------------------------------------------------------- - -[container_monitor] -enabled = true -runtime = "podman" - -# Watch these container name patterns -watch_patterns = [ - "test-infra-mongodb*", - "test-infra-redis*", - "test-infra-neo4j*", - "test-infra-clickhouse*", - "test-infra-surrealdb*", - "test-infra-influxdb*", - "test-infra-minio*", -] - -# Alert on these container events -alert_on = ["die", "oom", "restart"] - -# Ignore these events (too noisy for testing) -ignore = ["start", "stop", "pause", "unpause"] diff --git a/verisimdb/container/.gatekeeper.yaml b/verisimdb/container/.gatekeeper.yaml deleted file mode 100644 index 1b7bbb94..00000000 --- a/verisimdb/container/.gatekeeper.yaml +++ /dev/null @@ -1,100 +0,0 @@ -# SPDX-License-Identifier: MPL-2.0 -# -# Svalinn gatekeeper policy for VeriSimDB -# -# Controls which operations are permitted through the edge gateway. -# See: stapeln/container-stack/svalinn/ - -version: "1.0" - -# Authentication requirements -auth: - # Public endpoints (no auth required) - public: - - path: "/health" - methods: ["GET"] - - path: "/ready" - methods: ["GET"] - - path: "/metrics" - methods: ["GET"] - - # Endpoints requiring JWT/OAuth2 authentication - authenticated: - - path: "/api/v1/*" - methods: ["GET", "POST", "PUT", "DELETE"] - - path: "/graphql" - methods: ["POST"] - - # Federation endpoints require PSK in addition to JWT - federation: - - path: "/api/v1/federation/*" - methods: ["GET", "POST"] - require: ["jwt", "psk"] - -# Rate limiting -rate_limits: - # Global: 1000 req/s per authenticated client - global: - requests_per_second: 1000 - burst: 2000 - - # Write operations: 100 req/s (protects modality stores) - writes: - paths: ["/api/v1/octads"] - methods: ["POST", "PUT", "DELETE"] - requests_per_second: 100 - burst: 200 - - # Search operations: 500 req/s (protects Tantivy/HNSW) - search: - paths: ["/api/v1/search/*"] - methods: ["GET", "POST"] - requests_per_second: 500 - burst: 1000 - -# Container trust policy -trust: - # Only accept .ctp bundles signed by these keys - trusted_signers: - - key_id: "hyperpolymath-release" - algorithm: "Ed25519" - public_key_file: "/etc/svalinn/keys/hyperpolymath-release.pub" - - # Require these attestations on all .ctp bundles - required_attestations: - - "source-signature" - - "sbom-complete" - - # Reject unsigned or untrusted images - reject_unsigned: true - -# Request validation -validation: - # Maximum request body size (16MB — accommodates large tensor uploads) - max_body_size: "16MB" - - # Reject requests with NaN/Inf in numeric fields - reject_nan_inf: true - - # Maximum vector dimension (prevents OOM from oversized embeddings) - max_vector_dimension: 4096 - - # Maximum octad limit per list/search query - max_result_limit: 1000 - -# CORS policy (delegated from Rust core) -cors: - allow_origins: ["*"] - allow_methods: ["GET", "POST", "PUT", "DELETE", "OPTIONS"] - allow_headers: ["Content-Type", "Authorization", "X-Federation-PSK"] - max_age: 3600 - -# Logging -logging: - format: "json" - level: "info" - # Log all write operations for audit trail - audit_paths: - - "/api/v1/octads" - - "/api/v1/federation/*" - - "/api/v1/normalizer/trigger/*" diff --git a/verisimdb/container/Containerfile b/verisimdb/container/Containerfile deleted file mode 100644 index 22088c35..00000000 --- a/verisimdb/container/Containerfile +++ /dev/null @@ -1,125 +0,0 @@ -# SPDX-License-Identifier: MPL-2.0 -# VeriSimDB Container Image -# -# Build with Podman (in-memory, default): -# podman build -t verisimdb:latest -f container/Containerfile . -# -# Build with persistent storage (redb + file-backed Tantivy + WAL): -# podman build -t verisimdb:latest --build-arg FEATURES=persistent -f container/Containerfile . -# -# Run (in-memory): -# podman run -p 8080:8080 verisimdb:latest -# -# Run (persistent, with volume): -# podman run -p 8080:8080 -v verisimdb-data:/data verisimdb:latest - -# Stage 1: Build Rust -FROM cgr.dev/chainguard/wolfi-base:latest AS rust-builder - -# openssl-dev removed: TLS now pure Rust (rustls + ring) -# clang-19 removed: Oxigraph is now optional (oxigraph-backend feature flag, off by default) -# protoc removed: protobuf code pre-generated at verisim-api/src/proto/verisim.rs -RUN apk add --no-cache rust pkgconf build-base - -# Build arg: set to "persistent" for redb + file-backed Tantivy + WAL -ARG FEATURES="" - -WORKDIR /build - -# Copy Rust workspace -COPY Cargo.toml ./ -COPY rust-core/ ./rust-core/ -COPY benches/ ./benches/ - -# Build only the API binary (skip benches). -# When FEATURES=persistent, enables redb graph store + file-backed Tantivy + WAL. -RUN if [ -n "$FEATURES" ]; then \ - cargo build --release -p verisim-api --features "$FEATURES"; \ - else \ - cargo build --release -p verisim-api; \ - fi - -# Stage 2: Build Elixir -FROM cgr.dev/chainguard/wolfi-base:latest AS elixir-builder - -# Pin to OTP 27: mint 1.7.1 uses pkix_verify_hostname/3 removed in OTP 28 -RUN apk add --no-cache erl27-elixir-1.18 erlang-27 erlang-27-dev git build-base - -WORKDIR /build - -# Copy Elixir project -COPY elixir-orchestration/ ./ - -# Install deps, compile, and build release -ENV MIX_ENV=prod -RUN mix local.hex --force && \ - mix local.rebar --force && \ - mix deps.get --only prod && \ - mix compile && \ - mix release verisim - -# Stage 3: Runtime -FROM cgr.dev/chainguard/wolfi-base:latest - -# OCI image labels (compatible with cerro-torre .ctp bundle metadata) -LABEL org.opencontainers.image.title="VeriSimDB" \ - org.opencontainers.image.description="Cross-system entity consistency engine with 8-modality drift detection" \ - org.opencontainers.image.url="https://github.com/hyperpolymath/verisimdb" \ - org.opencontainers.image.source="https://github.com/hyperpolymath/verisimdb" \ - org.opencontainers.image.vendor="hyperpolymath" \ - org.opencontainers.image.licenses="PMPL-1.0-or-later" \ - org.opencontainers.image.authors="Jonathan D.A. Jewell " \ - dev.cerrotorre.manifest="container/manifest.toml" \ - dev.cerrotorre.gatekeeper="container/.gatekeeper.yaml" \ - dev.stapeln.compose="container/compose.toml" - -# libssl3 removed: TLS now pure Rust (rustls + ring) -RUN apk add --no-cache ca-certificates curl libstdc++ ncurses - -# Create non-root user -RUN addgroup -S verisim && adduser -S verisim -G verisim - -WORKDIR /app - -# Copy Rust binary -COPY --from=rust-builder /build/target/release/verisim-api /app/verisim-api - -# Copy Elixir release -COPY --from=elixir-builder /build/_build/prod/rel/verisim /app/elixir/ - -# Copy entrypoint -COPY container/entrypoint.sh /app/entrypoint.sh -RUN chmod +x /app/entrypoint.sh - -# Copy stapeln integration files (svalinn gatekeeper policy, cerro-torre manifest) -COPY container/.gatekeeper.yaml /etc/svalinn/gatekeeper.yaml -COPY container/manifest.toml /app/manifest.toml - -# Create data directory for persistent mode (mountable volume) -RUN mkdir -p /data && chown verisim:verisim /data - -# Set ownership -RUN chown -R verisim:verisim /app - -# Environment — IPv6-only by default -ENV RUST_LOG=info -ENV VERISIM_HOST=[::] -ENV VERISIM_PORT=8080 -ENV VERISIM_RUST_CORE_URL=http://[::1]:8080/api/v1 -ENV VERISIM_LOG_FORMAT=json -ENV VERISIM_PERSISTENCE_DIR=/data - -# Declare /data as a volume for persistent storage -VOLUME ["/data"] - -# Run as non-root -USER verisim - -# Expose API port -EXPOSE 8080 - -# Health check (as non-root, curl still works) -HEALTHCHECK --interval=30s --timeout=5s --start-period=10s --retries=3 \ - CMD curl -sf http://localhost:8080/health || exit 1 - -ENTRYPOINT ["/app/entrypoint.sh"] diff --git a/verisimdb/container/compose.toml b/verisimdb/container/compose.toml deleted file mode 100644 index 1952b4db..00000000 --- a/verisimdb/container/compose.toml +++ /dev/null @@ -1,94 +0,0 @@ -# SPDX-License-Identifier: MPL-2.0 -# -# VeriSimDB selur-compose configuration -# -# Orchestrates the full VeriSimDB stack as verified container bundles (.ctp). -# Uses selur zero-copy IPC between services on the same host. -# -# Usage: -# selur-compose up # Start all services -# selur-compose up --detach # Start in background -# selur-compose verify # Verify all .ctp signatures -# selur-compose ps # Check status -# selur-compose logs -f rust-core # Stream logs -# selur-compose down # Stop all services - -version = "1.0" - -# ============================================================================ -# Services -# ============================================================================ - -# Rust core: modality stores, drift detection, HTTP/gRPC API -[services.rust-core] -image = "ghcr.io/hyperpolymath/verisimdb-rust:latest.ctp" -ports = ["8080:8080", "50051:50051"] -environment = { - RUST_LOG = "info", - VERISIM_HOST = "[::]", - VERISIM_PORT = "8080", - VERISIM_LOG_FORMAT = "json", - VERISIM_PERSISTENCE_DIR = "/data", -} -volumes = ["verisimdb-data:/data"] -restart = "always" -healthcheck = { test = "curl -sf http://localhost:8080/health", interval = "30s", timeout = "5s", retries = 3 } - -# Elixir orchestration: entity servers, drift monitor, federation, query router -[services.elixir-orchestration] -image = "ghcr.io/hyperpolymath/verisimdb-elixir:latest.ctp" -ports = ["4000:4000"] -environment = { - VERISIM_RUST_CORE_URL = "http://rust-core:8080/api/v1", - VERISIM_LOG_FORMAT = "json", - MIX_ENV = "prod", -} -depends_on = ["rust-core"] -restart = "always" -healthcheck = { test = "curl -sf http://localhost:4000/health", interval = "30s", timeout = "5s", retries = 3 } - -# Svalinn edge gateway: validates requests, enforces policies, OAuth2/JWT auth -[services.svalinn] -image = "ghcr.io/hyperpolymath/svalinn:latest.ctp" -ports = ["443:443", "80:80"] -environment = { - SVALINN_BACKEND = "http://rust-core:8080", - SVALINN_ELIXIR_BACKEND = "http://elixir-orchestration:4000", - SVALINN_POLICY_FILE = "/etc/svalinn/gatekeeper.yaml", - SVALINN_TLS_AUTO = "true", -} -volumes = ["svalinn-config:/etc/svalinn:ro"] -depends_on = ["rust-core", "elixir-orchestration"] -restart = "always" -healthcheck = { test = "curl -sf http://localhost:80/health", interval = "30s", timeout = "5s", retries = 3 } - -# Rokur secrets manager (stub — placeholder until panic-attacker containerization completes) -# Uncomment when rokur has a real implementation: -# [services.rokur] -# image = "ghcr.io/hyperpolymath/rokur:latest.ctp" -# ports = ["8443:8443"] -# environment = { ROKUR_POLICY = "strict" } -# volumes = ["rokur-secrets:/var/lib/rokur"] -# restart = "always" - -# ============================================================================ -# Volumes -# ============================================================================ - -[volumes.verisimdb-data] -driver = "local" - -[volumes.svalinn-config] -driver = "local" - -# [volumes.rokur-secrets] -# driver = "local" - -# ============================================================================ -# Networks -# ============================================================================ - -# Use selur zero-copy IPC for inter-service communication on the same host. -# Falls back to standard bridge networking when selur driver is unavailable. -[networks.default] -driver = "selur" diff --git a/verisimdb/container/ct-build.sh b/verisimdb/container/ct-build.sh deleted file mode 100755 index 6e8487c3..00000000 --- a/verisimdb/container/ct-build.sh +++ /dev/null @@ -1,163 +0,0 @@ -#!/usr/bin/env bash -# SPDX-License-Identifier: MPL-2.0 -# -# VeriSimDB — Cerro Torre build, sign, and verify pipeline -# -# Builds the container image, packages it as a verified .ctp bundle, -# signs it with Ed25519, and verifies the result. -# -# Prerequisites: -# - podman (container build) -# - ct (cerro-torre CLI: pack, sign, verify) -# - cerro-sign (Ed25519 signing, from stapeln/container-stack/cerro-torre) -# -# Usage: -# ./ct-build.sh # Build + sign (in-memory) -# ./ct-build.sh persistent # Build + sign (persistent storage) -# ./ct-build.sh persistent --push # Build + sign + push to registry -# CT_KEY_ID=my-key ./ct-build.sh # Use specific signing key -# -# Environment variables: -# CT_KEY_ID — Signing key identifier (default: verisimdb-release) -# CT_REGISTRY — OCI registry to push to (default: ghcr.io/hyperpolymath) -# CT_TAG — Image tag (default: latest) -# FEATURES — Cargo feature flags (default: "" or "persistent") - -set -euo pipefail - -# --------------------------------------------------------------------------- -# Configuration -# --------------------------------------------------------------------------- - -SCRIPT_DIR="$(cd "$(dirname "$0")" && pwd)" -REPO_ROOT="$(cd "$SCRIPT_DIR/.." && pwd)" - -FEATURES="${1:-}" -PUSH="${2:-}" -CT_KEY_ID="${CT_KEY_ID:-verisimdb-release}" -CT_REGISTRY="${CT_REGISTRY:-ghcr.io/hyperpolymath}" -CT_TAG="${CT_TAG:-latest}" - -IMAGE_NAME="verisimdb" -if [ "$FEATURES" = "persistent" ]; then - IMAGE_NAME="verisimdb-persistent" -fi - -FULL_IMAGE="${CT_REGISTRY}/${IMAGE_NAME}:${CT_TAG}" -CTP_FILE="${SCRIPT_DIR}/${IMAGE_NAME}-${CT_TAG}.ctp" - -echo "=== VeriSimDB Cerro Torre Build Pipeline ===" -echo " Image: ${FULL_IMAGE}" -echo " Features: ${FEATURES:-none (in-memory)}" -echo " Key: ${CT_KEY_ID}" -echo " Bundle: ${CTP_FILE}" -echo "" - -# --------------------------------------------------------------------------- -# Step 1: Build container image with Podman -# --------------------------------------------------------------------------- - -echo "--- Step 1: Building container image ---" - -BUILD_ARGS=() -if [ -n "$FEATURES" ]; then - BUILD_ARGS+=(--build-arg "FEATURES=${FEATURES}") -fi - -podman build \ - -t "${FULL_IMAGE}" \ - "${BUILD_ARGS[@]}" \ - -f "${SCRIPT_DIR}/Containerfile" \ - "${REPO_ROOT}" - -echo " Built: ${FULL_IMAGE}" -echo "" - -# --------------------------------------------------------------------------- -# Step 2: Pack into .ctp bundle -# --------------------------------------------------------------------------- - -echo "--- Step 2: Packing into .ctp bundle ---" - -if command -v ct &>/dev/null; then - ct pack "${FULL_IMAGE}" -o "${CTP_FILE}" - echo " Packed: ${CTP_FILE}" -else - echo " SKIP: ct not found (install cerro-torre CLI from stapeln/container-stack/cerro-torre)" - echo " The container image is built and tagged but not packed as a .ctp bundle." - echo " To pack manually: ct pack ${FULL_IMAGE} -o ${CTP_FILE}" - echo "" - echo "=== Build complete (without .ctp signing) ===" - exit 0 -fi - -echo "" - -# --------------------------------------------------------------------------- -# Step 3: Sign the .ctp bundle -# --------------------------------------------------------------------------- - -echo "--- Step 3: Signing .ctp bundle ---" - -if command -v cerro-sign &>/dev/null; then - cerro-sign sign "${CTP_FILE}" --key-id "${CT_KEY_ID}" - echo " Signed: ${CTP_FILE} (key: ${CT_KEY_ID})" -elif command -v ct &>/dev/null; then - ct sign "${CTP_FILE}" --key "${CT_KEY_ID}" - echo " Signed: ${CTP_FILE} (key: ${CT_KEY_ID})" -else - echo " SKIP: cerro-sign not found (install from stapeln/container-stack/cerro-torre)" -fi - -echo "" - -# --------------------------------------------------------------------------- -# Step 4: Verify the .ctp bundle -# --------------------------------------------------------------------------- - -echo "--- Step 4: Verifying .ctp bundle ---" - -if command -v ct &>/dev/null; then - ct verify "${CTP_FILE}" - echo " Verified: ${CTP_FILE}" -else - echo " SKIP: ct not found" -fi - -echo "" - -# --------------------------------------------------------------------------- -# Step 5: Push to registry (optional) -# --------------------------------------------------------------------------- - -if [ "$PUSH" = "--push" ]; then - echo "--- Step 5: Pushing to registry ---" - - if command -v ct &>/dev/null; then - ct push "${CTP_FILE}" "${FULL_IMAGE}" - echo " Pushed: ${FULL_IMAGE}" - else - # Fall back to podman push (unsigned OCI image) - echo " ct not available, falling back to podman push (unsigned)" - podman push "${FULL_IMAGE}" - echo " Pushed: ${FULL_IMAGE} (unsigned OCI — not a .ctp bundle)" - fi - echo "" -fi - -# --------------------------------------------------------------------------- -# Summary -# --------------------------------------------------------------------------- - -echo "=== Build pipeline complete ===" -echo " Image: ${FULL_IMAGE}" -echo " Bundle: ${CTP_FILE}" -echo "" -echo " To deploy with selur-compose:" -echo " cd container && selur-compose up" -echo "" -echo " To verify at any time:" -echo " ct verify ${CTP_FILE}" -echo "" -echo " To explain the verification chain:" -echo " ct explain ${CTP_FILE}" diff --git a/verisimdb/container/entrypoint.sh b/verisimdb/container/entrypoint.sh deleted file mode 100644 index 2106e10f..00000000 --- a/verisimdb/container/entrypoint.sh +++ /dev/null @@ -1,36 +0,0 @@ -#!/bin/sh -# SPDX-License-Identifier: MPL-2.0 -# VeriSimDB entrypoint — starts Rust API then Elixir orchestration -set -e - -# Start Rust API server in background -echo "Starting VeriSimDB Rust API..." -/app/verisim-api & -RUST_PID=$! - -# Wait for Rust API to become ready -echo "Waiting for Rust API to be ready..." -RETRIES=0 -MAX_RETRIES=30 -until curl -sf http://127.0.0.1:${VERISIM_PORT:-8080}/health > /dev/null 2>&1; do - RETRIES=$((RETRIES + 1)) - if [ "$RETRIES" -ge "$MAX_RETRIES" ]; then - echo "ERROR: Rust API failed to start after ${MAX_RETRIES}s" - kill "$RUST_PID" 2>/dev/null || true - exit 1 - fi - sleep 1 -done -echo "Rust API ready on port ${VERISIM_PORT:-8080}" - -# Clean shutdown: kill Rust API when Elixir exits -cleanup() { - echo "Shutting down..." - kill "$RUST_PID" 2>/dev/null || true - wait "$RUST_PID" 2>/dev/null || true -} -trap cleanup EXIT INT TERM - -# Start Elixir release in foreground -echo "Starting VeriSimDB Elixir orchestration..." -exec /app/elixir/bin/verisim start diff --git a/verisimdb/container/manifest.toml b/verisimdb/container/manifest.toml deleted file mode 100644 index 982adb1e..00000000 --- a/verisimdb/container/manifest.toml +++ /dev/null @@ -1,66 +0,0 @@ -# SPDX-License-Identifier: MPL-2.0 -# -# Cerro Torre manifest for VeriSimDB .ctp bundle -# -# This manifest describes the VeriSimDB container image for verified -# container packaging. Used by `ct pack` to create .ctp bundles. - -[metadata] -name = "verisimdb" -version = "0.1.0" -revision = 1 -summary = "Cross-system entity consistency engine with 8-modality drift detection" -description = """ -VeriSimDB (Veridical Simulacrum Database) is a cross-system entity consistency -engine that detects and repairs drift across 8 modality representations of the -same entity: Graph, Vector, Tensor, Semantic, Document, Temporal, Provenance, -and Spatial. Operates as a standalone database or heterogeneous federation -coordinator over PostgreSQL, ArangoDB, Elasticsearch, and other VeriSimDB -instances. -""" -license = "PMPL-1.0-or-later" -homepage = "https://github.com/hyperpolymath/verisimdb" -maintainer = "Jonathan D.A. Jewell " - -[provenance] -upstream = "https://github.com/hyperpolymath/verisimdb" -import_date = 2026-02-28T00:00:00Z - -[dependencies] -runtime = ["ca-certificates", "curl", "libstdc++", "ncurses"] -build = ["rust", "pkgconf", "build-base", "erl27-elixir-1.18", "erlang-27"] - -[build] -system = "cargo" - -[build.environment] -RUST_LOG = "info" -VERISIM_HOST = "[::]" -VERISIM_PORT = "8080" - -[outputs] -primary = "verisimdb" -split = ["verisimdb-rust", "verisimdb-elixir"] - -[attestations] -require = ["source-signature", "sbom-complete"] -recommend = ["security-audit", "reproducible-build"] - -# Runtime security profile -[security] -user = "verisim" -group = "verisim" -read_only_root = false -no_new_privileges = true - -[security.capabilities] -drop = ["ALL"] -add = ["NET_BIND_SERVICE"] - -[security.network] -listen_tcp = [8080, 50051] - -[security.filesystem] -read = ["/app/", "/data/"] -write = ["/data/", "/tmp/"] -execute = ["/app/verisim-api", "/app/elixir/bin/verisim"] diff --git a/verisimdb/contractiles/README.adoc b/verisimdb/contractiles/README.adoc deleted file mode 100644 index d19a3877..00000000 --- a/verisimdb/contractiles/README.adoc +++ /dev/null @@ -1,19 +0,0 @@ -= Contractiles Template Set -:toc: -:sectnums: - -This directory contains the generalized contractiles templates. Copy the `contractiles/` directory into a new repo to establish a consistent operational, validation, trust, recovery, and intent framework. - -== Fill-In Instructions - -1. Update the Mustfile to reflect your real invariants (paths, schema versions, ports). -2. Replace Trustfile.hs placeholders with your actual key paths and verification commands. -3. Adjust Dustfile handlers to match your rollback and recovery tooling. -4. Update Intentfile to mirror the roadmap you want the system to evolve toward. - -== Contents - -* `must/Mustfile` - required invariants and validations. -* `trust/Trustfile.hs` - cryptographic verification steps. -* `dust/Dustfile` - rollback and recovery semantics. -* `lust/Intentfile` - future intent and roadmap direction. diff --git a/verisimdb/contractiles/dust/Dustfile b/verisimdb/contractiles/dust/Dustfile deleted file mode 100644 index aece7295..00000000 --- a/verisimdb/contractiles/dust/Dustfile +++ /dev/null @@ -1 +0,0 @@ -content of dustfile diff --git a/verisimdb/contractiles/must/Mustfile b/verisimdb/contractiles/must/Mustfile deleted file mode 100644 index ee751099..00000000 --- a/verisimdb/contractiles/must/Mustfile +++ /dev/null @@ -1,14 +0,0 @@ -# SPDX-License-Identifier: MPL-2.0 -# Mustfile - mandatory checks -# See: https://github.com/hyperpolymath/mustfile - -version: 1 - -checks: - - name: security - run: just lint - - name: tests - run: just test - - name: format - run: just fmt - diff --git a/verisimdb/contractiles/trust/Trustfile b/verisimdb/contractiles/trust/Trustfile deleted file mode 100644 index 315b82be..00000000 --- a/verisimdb/contractiles/trust/Trustfile +++ /dev/null @@ -1,137 +0,0 @@ -;; SPDX-License-Identifier: MPL-2.0 -;; Trustfile — Cryptographic trust and proof verification policy for VeriSimDB -;; -;; This file defines the zero-knowledge proof (ZKP) scheme selection, -;; cryptographic verification policies, and trust boundaries. - -;; ============================================================================ -;; ZKP Scheme Selection for VCL-UT PROOF Clause -;; ============================================================================ -;; -;; Decision: PLONK (Permutations over Lagrange-bases for Oecumenical -;; Noninteractive arguments of Knowledge) -;; -;; Rationale: PLONK is the best fit for VeriSimDB's proof obligations: -;; -;; | Criterion | PLONK | STARKs | Bulletproofs | Groth16 | -;; |--------------------------|-------------|--------------|--------------|--------------| -;; | Trusted setup | Universal¹ | None (trans) | None (trans) | Per-circuit | -;; | Proof size | ~400 bytes | ~45 KB | ~700 bytes | ~200 bytes | -;; | Verification time | ~5ms | ~10ms | ~50ms | ~3ms | -;; | Prover time | ~1s | ~2s | ~5s | ~1s | -;; | Updatable | Yes | N/A | N/A | No | -;; | Post-quantum | No | Yes | No | No | -;; | Recursive proofs | Yes | Yes (harder) | No | No | -;; | Circuit flexibility | High | High | Limited | Fixed | -;; -;; ¹ PLONK uses a Universal Structured Reference String (SRS) — one trusted -;; setup ceremony for ALL circuits, not per-circuit like Groth16. -;; -;; Why PLONK over alternatives: -;; -;; 1. vs Groth16: Groth16 has the smallest proofs and fastest verification, -;; but requires a NEW trusted setup for every circuit change. VeriSimDB's -;; proof contracts change frequently (new types, new properties). PLONK's -;; universal SRS means one setup covers all VCL-UT contracts forever. -;; -;; 2. vs STARKs: STARKs are transparent (no trusted setup) and post-quantum, -;; but proof sizes are ~100x larger (45KB vs 400B). For VeriSimDB where -;; proofs are stored in the semantic modality and transmitted over gRPC, -;; proof size matters. STARKs are better for on-chain verification where -;; transparency is paramount — that's not our use case. -;; -;; 3. vs Bulletproofs: Bulletproofs are transparent and compact, but -;; verification time is O(n) — too slow for interactive query verification. -;; PLONK verification is constant-time regardless of circuit size. -;; -;; Future: If post-quantum resistance becomes critical, migrate to STARKs -;; for the most sensitive proof types (Integrity, Custom) while keeping -;; PLONK for lightweight proofs (Existence, Citation, Access). The proof -;; infrastructure should abstract over the backend scheme. - -;; ============================================================================ -;; Proof Type → Circuit Mapping -;; ============================================================================ -;; -;; (proof-type "existence" -;; :scheme :plonk -;; :circuit "hexad-exists" -;; :description "Verify hexad exists and is accessible" -;; :estimated-verify-ms 1 -;; :estimated-prove-ms 50) -;; -;; (proof-type "citation" -;; :scheme :plonk -;; :circuit "contract-registry-lookup" -;; :description "Verify contract exists in semantic registry" -;; :estimated-verify-ms 5 -;; :estimated-prove-ms 200) -;; -;; (proof-type "access" -;; :scheme :plonk -;; :circuit "rbac-check" -;; :description "Verify caller has rights to modality/entity" -;; :estimated-verify-ms 15 -;; :estimated-prove-ms 500) -;; -;; (proof-type "integrity" -;; :scheme :plonk -;; :circuit "merkle-membership" -;; :description "Verify data integrity via Merkle proof" -;; :estimated-verify-ms 50 -;; :estimated-prove-ms 1000 -;; :notes "Uses Merkle tree optimization for n>100 results (O(log n) proof size)") -;; -;; (proof-type "provenance" -;; :scheme :plonk -;; :circuit "lineage-chain" -;; :description "Verify data lineage/audit trail" -;; :estimated-verify-ms 30 -;; :estimated-prove-ms 800 -;; :sequential true) ;; Chain walk cannot be parallelized -;; -;; (proof-type "custom" -;; :scheme :plonk -;; :circuit "user-defined" -;; :description "User-defined proof contract" -;; :estimated-verify-ms 200 -;; :estimated-prove-ms 2000) - -;; ============================================================================ -;; Rust Implementation -;; ============================================================================ -;; -;; Recommended crate: `ark-plonk` (arkworks ecosystem) -;; - ark-poly-commit: polynomial commitment schemes -;; - ark-serialize: proof serialization -;; - ark-bn254: BN254 curve (128-bit security) -;; - ark-bls12-381: BLS12-381 curve (alternative, wider adoption) -;; -;; Curve selection: BN254 for speed, BLS12-381 for wider ecosystem compatibility. -;; Start with BN254, offer BLS12-381 as configuration option. -;; -;; SRS ceremony: Use Zcash Powers of Tau ceremony artifacts or generate -;; project-specific SRS with auditable randomness. - -;; ============================================================================ -;; Existing Cryptographic Infrastructure (from Trustfile.hs) -;; ============================================================================ -;; -;; The following are RETAINED from the existing Trustfile: -;; - Policy hash verification: SHA-256 (policy/policy.ncl) -;; - Schema signatures: OpenSSL RSA/ECDSA (schema/schema.json) -;; - Driver signatures: Kyber-1024 post-quantum (drivers/*.bin) -;; - Migration provenance: OpenSSL RSA/ECDSA (migrations/provenance.json) -;; -;; These are ORTHOGONAL to ZKP — they protect infrastructure integrity, -;; while ZKP protects query-time proof obligations. - -;; ============================================================================ -;; Trust Boundaries -;; ============================================================================ -;; -;; 1. Hexad Store → Semantic Store: proof blobs stored as CBOR in semantic modality -;; 2. VCL Parser → Planner: PROOF clause extracted, cost estimated -;; 3. Planner → Executor: proof verification dispatched before/after query -;; 4. Executor → ZKP Backend: circuit evaluated, proof generated/verified -;; 5. Client → API: proof result included in response (pass/fail + proof bytes) diff --git a/verisimdb/debugger/.claude/CLAUDE.md b/verisimdb/debugger/.claude/CLAUDE.md deleted file mode 100644 index 7fc46425..00000000 --- a/verisimdb/debugger/.claude/CLAUDE.md +++ /dev/null @@ -1,126 +0,0 @@ -# CLAUDE.md — VeriSimDB Debugger - -## Purpose - -The VeriSimDB debugger is a Rust TUI (ratatui/crossterm) tool that provides interactive debugging, inspection, and visualisation of VeriSimDB internals. It serves **two simultaneous roles**: - -1. **Production debugger** for VeriSimDB operators and developers -2. **Reference example** demonstrating how PanLL's database design/development/evaluation functions work with a real database - -## Architecture Mandate - -This debugger MUST be designed so that **PanLL can embed and drive it programmatically**. Every UI panel in the TUI corresponds to a PanLL pane concept: - -| TUI Panel | PanLL Pane | Function | -|-----------|-----------|----------| -| Modality Inspector | Pane-W (results) | View all 8 octad modalities for a given entity | -| Drift Heatmap | Pane-W (results) | Colour-coded grid of drift scores across entities | -| VCL Trace | Pane-N (reasoning) | Step-by-step VCL query execution trace | -| Proof Verifier | Pane-L (constraints) | VCL-UT proof obligation checking and certificate display | -| Federation Map | Pane-W (results) | Live peer status, replication lag, adapter types | -| Performance Flamegraph | Pane-W (results) | Query latency breakdown by modality | -| Normalisation Timeline | Pane-N (reasoning) | History of self-normalisation events, before/after states | -| Error Budget Dashboard | Pane-L (constraints) | SLO tracking, error rates, circuit breaker status | - -## Example Use Cases (for PanLL integration) - -### Database Design/Development/Evaluation - -PanLL should use this debugger as its **canonical example** for the database design/development/evaluation module. When implementing PanLL's database evaluation functions, reference this debugger's: - -- **Entity inspection**: How to display multi-modal data (8 modalities) in a coherent view -- **Drift detection visualisation**: How to show data quality metrics across a large entity set -- **Query tracing**: How to step through query execution plans and show which modality stores were hit -- **Proof certificates**: How to display formal verification results in a human-readable way -- **Federation monitoring**: How to show distributed system health across multiple database backends - -### Language Design/Development/Evaluation - -PanLL should use **Eclexia** (from `nextgen-languages/eclexia/`) as the canonical example for its language design/development/evaluation module. The VeriSimDB debugger connects to Eclexia via: - -- VCL is defined in Eclexia's grammar format → debugger can show VCL parse trees -- VCL-UT proof obligations are typed in Eclexia's type system → debugger can show type derivation trees -- Eclexia's REPL can be embedded as a debugger panel for interactive VCL exploration - -## Implementation Spec - -### Phase 1: Core Infrastructure (Do This First) - -Build the TUI skeleton with these panels: - -``` -┌──────────────────────────────────┬──────────────────────────────────┐ -│ Entity Inspector │ Drift Heatmap │ -│ ├── octad_id │ ┌────────────────────────────┐ │ -│ ├── graph: [edges] │ │ ■■■■■■■■ entity-001 0.02 │ │ -│ ├── vector: [dims] │ │ ■■■■■■■□ entity-002 0.15 │ │ -│ ├── tensor: [shape] │ │ ■■■□□□□□ entity-003 0.67 │ │ -│ ├── semantic: [types] │ │ ■■■■■■■■ entity-004 0.01 │ │ -│ ├── document: [text preview] │ └────────────────────────────┘ │ -│ ├── temporal: [versions] │ │ -│ ├── provenance: [chain] │ │ -│ └── spatial: [coords] │ │ -├──────────────────────────────────┼──────────────────────────────────┤ -│ VCL Trace │ Proof Verifier │ -│ > SELECT * FROM octads │ PROOF EXISTENCE(entity-001) │ -│ 1. Parse: 2ms │ ✅ octad_id: found │ -│ 2. Type-check: 15ms │ ✅ modality_count: 8 │ -│ 3. Route → graph,vector: 8ms │ ✅ certificate: SHA-256 ok │ -│ 4. Execute graph: 12ms │ │ -│ 5. Execute vector: 9ms │ PROOF PROVENANCE(entity-001) │ -│ 6. Cross-modal merge: 3ms │ ✅ chain_length: 5 │ -│ 7. Return 42 rows: 1ms │ ✅ chain_hash: verified │ -│ Total: 50ms │ ✅ certificate: SHA-256 ok │ -└──────────────────────────────────┴──────────────────────────────────┘ -``` - -### Phase 2: API Client - -Connect to VeriSimDB via: -- **Rust core** at `http://localhost:8080/api/v1/` — entity CRUD, search, drift -- **Elixir orchestration** at `http://localhost:4080/` — telemetry, health, consensus status -- **gRPC** at `localhost:50051` — for high-throughput entity streaming - -### Phase 3: PanLL Protocol - -Expose a JSON-over-stdio protocol so PanLL can drive the debugger headlessly: -- `{"cmd": "inspect", "entity_id": "..."}` → returns entity data for all 8 modalities -- `{"cmd": "drift_scan", "threshold": 0.3}` → returns entities with drift above threshold -- `{"cmd": "trace_vcl", "query": "SELECT ..."}` → returns execution trace -- `{"cmd": "verify_proof", "query": "... PROOF ..."}` → returns proof certificates -- `{"cmd": "health"}` → returns full telemetry snapshot - -### Phase 4: Eclexia Integration - -Add a panel that can: -- Parse VCL using Eclexia's grammar and show the parse tree -- Show VCL-UT type derivation trees using Eclexia's type system -- Provide an interactive VCL REPL powered by Eclexia's evaluator - -## Build Commands - -```bash -cd debugger -cargo build -cargo test -cargo run -- --help - -# Connect to running VeriSimDB -cargo run -- --rust-url http://localhost:8080 --orch-url http://localhost:4080 -``` - -## Key Dependencies - -- `ratatui` 0.29 + `crossterm` 0.28 — TUI framework -- `reqwest` — HTTP client for VeriSimDB API -- `tokio` — async runtime -- `clap` — CLI argument parsing -- `serde_json` — JSON serialization for PanLL protocol - -## What NOT To Do - -- Do NOT duplicate VeriSimDB logic in the debugger — always call the API -- Do NOT add a web UI — TUI only (PanLL handles the web layer) -- Do NOT use Python, Go, TypeScript, or Node.js -- Do NOT use unsafe Rust without `// SAFETY:` comments -- Do NOT store entity data locally — the debugger is stateless diff --git a/verisimdb/debugger/Cargo.toml b/verisimdb/debugger/Cargo.toml deleted file mode 100644 index 470ad5c6..00000000 --- a/verisimdb/debugger/Cargo.toml +++ /dev/null @@ -1,54 +0,0 @@ -# SPDX-License-Identifier: MPL-2.0 - -[package] -name = "verisimdb-debugger" -version = "0.1.0" -edition = "2024" -authors = ["Jonathan D.A. Jewell "] -license = "MPL-2.0" -description = "Interactive debugger and visualization tool for VeriSimDB" -repository = "https://github.com/hyperpolymath/verisimdb-debugger" -keywords = ["database", "debugger", "tui", "multimodal", "zkp"] -categories = ["command-line-utilities", "development-tools::debugging"] - -[dependencies] -# Async runtime -tokio = { version = "1.42", features = ["full"] } -async-trait = "0.1" - -# HTTP client (for VeriSimDB API) -reqwest = { version = "0.12", features = ["json"] } - -# Serialization -serde = { version = "1.0", features = ["derive"] } -serde_json = "1.0" - -# CLI -clap = { version = "4.5", features = ["derive", "cargo"] } - -# Logging -tracing = "0.1" -tracing-subscriber = { version = "0.3", features = ["env-filter"] } - -# Error handling -anyhow = "1.0" -thiserror = "2.0" - -# Time handling -chrono = { version = "0.4", features = ["serde"] } - -# UUID -uuid = { version = "1.11", features = ["serde", "v4"] } - -[dev-dependencies] -tokio-test = "0.4" - -[profile.release] -opt-level = 3 -lto = true -codegen-units = 1 -strip = true - -[[bin]] -name = "verisimdb-debugger" -path = "src/main.rs" diff --git a/verisimdb/debugger/docs/CITATIONS.adoc b/verisimdb/debugger/docs/CITATIONS.adoc deleted file mode 100644 index 6f167bdf..00000000 --- a/verisimdb/debugger/docs/CITATIONS.adoc +++ /dev/null @@ -1,36 +0,0 @@ -= RSR-template-repo - Citation Guide -:toc: - -== BibTeX - -[source,bibtex] ----- -@software{rsr-template-repo_2025, - author = {Polymath, Hyper}, - title = {RSR-template-repo}, - year = {2025}, - url = {https://github.com/hyperpolymath/RSR-template-repo}, - license = {PMPL-1.0-or-later} -} ----- - -== Harvard Style - -Polymath, H. (2025) _RSR-template-repo_ [Computer software]. Available at: https://github.com/hyperpolymath/RSR-template-repo - -== OSCOLA - -Hyper Polymath, 'RSR-template-repo' (2025) - -== MLA - -Polymath, Hyper. "RSR-template-repo." 2025, github.com/hyperpolymath/RSR-template-repo. - -== APA 7 - -Polymath, H. (2025). _RSR-template-repo_ [Computer software]. GitHub. https://github.com/hyperpolymath/RSR-template-repo - -== See Also - -* link:../CITATION.cff[CITATION.cff] -* link:../codemeta.json[codemeta.json] diff --git a/verisimdb/debugger/examples/SafeDOMExample.res b/verisimdb/debugger/examples/SafeDOMExample.res deleted file mode 100644 index e5c90460..00000000 --- a/verisimdb/debugger/examples/SafeDOMExample.res +++ /dev/null @@ -1,109 +0,0 @@ -// SPDX-License-Identifier: MPL-2.0 -// Example: Using SafeDOM for formally verified DOM mounting - -open SafeDOM - -// Example 1: Basic mounting with error handling -let mountApp = () => { - mountSafe( - "#app", - "

Hello, World!

Mounted safely with proofs.

", - ~onSuccess=el => { - Console.log("✓ App mounted successfully!") - Console.log("Element:", el) - }, - ~onError=err => { - Console.error("✗ Mount failed:", err) - } - ) -} - -// Example 2: Wait for DOM ready before mounting -let mountWhenDOMReady = () => { - mountWhenReady( - "#app", - "

App Title

", - ~onSuccess=_ => Console.log("✓ Mounted after DOM ready"), - ~onError=err => Console.error("✗ Failed:", err) - ) -} - -// Example 3: Batch mounting (atomic - all or nothing) -let mountMultiple = () => { - let specs = [ - {selector: "#header", html: "

Site Title

"}, - {selector: "#nav", html: "
"}, - {selector: "#main", html: "

Content here

"}, - {selector: "#footer", html: "
© 2026
"} - ] - - switch mountBatch(specs) { - | Ok(elements) => { - Console.log(`✓ Successfully mounted ${Array.length(elements)} elements`) - elements->Array.forEach(el => Console.log(" -", el)) - } - | Error(err) => { - Console.error("✗ Batch mount failed:", err) - Console.error(" (None were mounted - atomic operation)") - } - } -} - -// Example 4: Explicit validation before mounting -let mountWithValidation = () => { - // Validate selector first - switch ProvenSelector.validate("#my-app") { - | Error(e) => Console.error(`Invalid selector: ${e}`) - | Ok(validSelector) => { - // Validate HTML - switch ProvenHTML.validate("
Content
") { - | Error(e) => Console.error(`Invalid HTML: ${e}`) - | Ok(validHtml) => { - // Now mount with proven safety - switch mount(validSelector, validHtml) { - | Mounted(el) => Console.log("✓ Mounted with validated inputs:", el) - | MountPointNotFound(s) => Console.error(`✗ Element not found: ${s}`) - | InvalidSelector(_) => Console.error("Impossible - already validated") - | InvalidHTML(_) => Console.error("Impossible - already validated") - } - } - } - } -} - -// Example 5: Integration with TEA -module MyApp = { - type model = {message: string} - type msg = NoOp - - let init = () => {message: "Hello from TEA"} - let update = (model, _msg) => model - let view = model => `

${model.message}

` -} - -let mountTEAApp = () => { - let model = MyApp.init() - let html = MyApp.view(model) - - mountWhenReady( - "#tea-app", - html, - ~onSuccess=el => { - Console.log("✓ TEA app mounted") - // Set up event handlers, subscriptions here - }, - ~onError=err => Console.error(`✗ TEA mount failed: ${err}`) - ) -} - -// Entry point -let main = () => { - Console.log("SafeDOM Examples") - Console.log("================\n") - - // Choose which example to run - mountWhenDOMReady() // Run on DOM ready -} - -// Auto-execute when module loads -main() diff --git a/verisimdb/debugger/examples/web-project-deno.json b/verisimdb/debugger/examples/web-project-deno.json deleted file mode 100644 index 5ddd3bd7..00000000 --- a/verisimdb/debugger/examples/web-project-deno.json +++ /dev/null @@ -1,20 +0,0 @@ -{ - "// NOTE": "Example deno.json for ReScript web projects", - "tasks": { - "build": "deno run -A npm:rescript", - "clean": "deno run -A npm:rescript clean", - "watch": "deno run -A npm:rescript -w", - "serve": "deno run -A jsr:@std/http/file-server .", - "test": "deno test --allow-all" - }, - "imports": { - "rescript": "^12.0.0", - "@rescript/core": "npm:@rescript/core@^1.6.0", - "safe-dom/": "https://raw.githubusercontent.com/hyperpolymath/rescript-dom-mounter/main/src/", - "proven/": "../proven/bindings/rescript/src/" - }, - "compilerOptions": { - "allowJs": true, - "checkJs": false - } -} diff --git a/verisimdb/debugger/fuzz/Cargo.toml b/verisimdb/debugger/fuzz/Cargo.toml deleted file mode 100644 index 49d6f905..00000000 --- a/verisimdb/debugger/fuzz/Cargo.toml +++ /dev/null @@ -1,20 +0,0 @@ -[package] -name = "fuzz" -version = "0.0.0" -publish = false -edition = "2021" - -[package.metadata] -cargo-fuzz = true - -[dependencies] -libfuzzer-sys = "0.4" - -[dependencies.verisimdb-debugger] -path = ".." - -[[bin]] -name = "fuzz_main" -path = "fuzz_targets/fuzz_main.rs" -test = false -doc = false diff --git a/verisimdb/debugger/fuzz/fuzz_targets/fuzz_main.rs b/verisimdb/debugger/fuzz/fuzz_targets/fuzz_main.rs deleted file mode 100644 index 1f71ba3b..00000000 --- a/verisimdb/debugger/fuzz/fuzz_targets/fuzz_main.rs +++ /dev/null @@ -1,22 +0,0 @@ -#![no_main] -use libfuzzer_sys::fuzz_target; - -fuzz_target!(|data: &[u8]| { - // Generic fuzzing for memory safety and crash detection - if data.is_empty() || data.len() > 100000 { - return; - } - - // Test UTF-8 validity - if let Ok(text) = std::str::from_utf8(data) { - // Test string operations - let _ = text.to_lowercase(); - let _ = text.chars().count(); - let _ = text.split_whitespace().collect::>(); - } - - // Test binary data handling - if data.len() >= 8 { - let _chunk = &data[..8]; - } -}); diff --git a/verisimdb/debugger/src/main.rs b/verisimdb/debugger/src/main.rs deleted file mode 100644 index f3234c3b..00000000 --- a/verisimdb/debugger/src/main.rs +++ /dev/null @@ -1,14 +0,0 @@ -// SPDX-License-Identifier: MPL-2.0 -//! VeriSimDB Debugger - Interactive debugger for VeriSimDB queries and drift -//! -//! Status: Alpha development (v0.1.0) -//! Not yet functional - scaffolding only - -#![forbid(unsafe_code)] -fn main() { - eprintln!("verisimdb-debugger v0.1.0"); - eprintln!("Status: Alpha development - not yet functional"); - eprintln!(""); - eprintln!("This tool is in early development."); - eprintln!("See https://github.com/hyperpolymath/verisimdb-debugger for status."); -} diff --git a/verisimdb/debugger/tests/integration_test.rs b/verisimdb/debugger/tests/integration_test.rs deleted file mode 100644 index 7237aed0..00000000 --- a/verisimdb/debugger/tests/integration_test.rs +++ /dev/null @@ -1,7 +0,0 @@ -// SPDX-License-Identifier: MPL-2.0 - -#[test] -fn placeholder_test() { - // Placeholder test for CI - assert!(true); -} diff --git a/verisimdb/demos/drift-detection/run_demo.exs b/verisimdb/demos/drift-detection/run_demo.exs deleted file mode 100644 index 63c76496..00000000 --- a/verisimdb/demos/drift-detection/run_demo.exs +++ /dev/null @@ -1,536 +0,0 @@ -# SPDX-License-Identifier: MPL-2.0 -# -# VeriSimDB Drift Detection & Self-Normalisation Demo -# -# Demonstrates VeriSimDB's core value proposition: -# 1. Create entities with 8 modalities (octad) -# 2. Introduce cross-modal corruption -# 3. Detect drift via sweep -# 4. Repair via normalisation -# 5. Verify consistency restored -# -# Usage: -# cd elixir-orchestration && mix run ../demos/drift-detection/run_demo.exs -# -# The demo works in two modes: -# - LIVE mode: Rust core is running — real octads, real drift, real repair -# - LOCAL mode: Rust core unavailable — Elixir-level simulation via DriftMonitor - -defmodule VeriSimDB.Demo.DriftDetection do - @moduledoc """ - Drift detection demo script. - - Creates entities, corrupts a subset, detects drift, repairs, and reports. - Demonstrates the full detect → monitor → repair → verify cycle. - """ - - require Logger - - alias VeriSim.{DriftMonitor, EntityServer, RustClient} - - # ── Configuration ────────────────────────────────────────────────────── - - @entity_count 1_000 - @corrupt_count 50 - @modalities ~w(graph vector tensor semantic document temporal provenance spatial)a - - # Corruption patterns: each introduces a specific kind of cross-modal drift - @corruption_patterns [ - # Vector changed without updating document — semantic_vector drift - %{type: :semantic_vector, score: 0.75, description: "vector/document desync"}, - # Graph edges removed but document still references them — graph_document drift - %{type: :graph_document, score: 0.82, description: "graph/document mismatch"}, - # Temporal version skipped — temporal_consistency drift - %{type: :temporal_consistency, score: 0.65, description: "version gap"}, - # Tensor shape changed — tensor drift - %{type: :tensor, score: 0.70, description: "tensor shape mismatch"}, - # Required modality missing — schema drift - %{type: :schema, score: 0.90, description: "missing modality"}, - # Overall quality degradation — quality drift - %{type: :quality, score: 0.60, description: "quality degradation"} - ] - - # ── Entry Point ──────────────────────────────────────────────────────── - - def run do - print_banner() - - # Ensure required processes are running - ensure_infrastructure() - - # Determine mode based on Rust core availability - mode = detect_mode() - print_mode(mode) - - # Phase 1: Create entities - {create_us, entities} = :timer.tc(fn -> create_entities(mode) end) - print_phase_complete(1, "Create #{@entity_count} entities", create_us) - - # Phase 2: Corrupt a subset - {corrupt_us, corrupted_ids} = :timer.tc(fn -> corrupt_entities(mode, entities) end) - print_phase_complete(2, "Corrupt #{@corrupt_count} entities", corrupt_us) - - # Phase 3: Detect drift via sweep - {detect_us, detections} = :timer.tc(fn -> detect_drift(mode) end) - print_phase_complete(3, "Drift detection sweep", detect_us) - - # Phase 4: Repair via normalisation - {repair_us, repairs} = :timer.tc(fn -> repair_drift(mode, corrupted_ids) end) - print_phase_complete(4, "Normalisation repair", repair_us) - - # Phase 5: Verify consistency - {verify_us, verification} = :timer.tc(fn -> verify_consistency(mode, corrupted_ids) end) - print_phase_complete(5, "Consistency verification", verify_us) - - # Print summary - print_summary(%{ - mode: mode, - entity_count: @entity_count, - corrupt_count: @corrupt_count, - detected: detections, - repaired: repairs, - verified: verification, - timings: %{ - create_us: create_us, - corrupt_us: corrupt_us, - detect_us: detect_us, - repair_us: repair_us, - verify_us: verify_us, - total_us: create_us + corrupt_us + detect_us + repair_us + verify_us - } - }) - end - - # ── Infrastructure ───────────────────────────────────────────────────── - - defp ensure_infrastructure do - IO.puts(" Starting infrastructure...") - - # Start DriftMonitor if not already running - case DriftMonitor.start_link(config: %{ - sweep_interval_ms: 600_000, # Long interval — we trigger sweeps manually - max_concurrent_normalizations: 50, - thresholds: %{ - semantic_vector: %{warning: 0.3, critical: 0.7}, - graph_document: %{warning: 0.4, critical: 0.8}, - temporal_consistency: %{warning: 0.2, critical: 0.6}, - tensor: %{warning: 0.35, critical: 0.75}, - schema: %{warning: 0.1, critical: 0.5}, - quality: %{warning: 0.25, critical: 0.65} - } - }) do - {:ok, _pid} -> IO.puts(" DriftMonitor started") - {:error, {:already_started, _pid}} -> IO.puts(" DriftMonitor already running") - end - - # Initialize ETS cache for RustClient - RustClient.init_cache() - IO.puts(" RustClient cache initialised") - IO.puts("") - end - - defp detect_mode do - # Verify the Rust core is actually running — not just any server on port 8080. - # The health endpoint returns a JSON map with a "status" field when the Rust - # core is running. If we get HTML or a non-map response, fall back to local. - case RustClient.health() do - {:ok, body} when is_map(body) -> :live - {:ok, _non_map} -> :local # e.g. HTML from nginx — not the Rust core - {:error, _} -> :local - end - end - - # ── Phase 1: Create Entities ─────────────────────────────────────────── - - defp create_entities(:live) do - IO.puts(" Creating #{@entity_count} octads via Rust core...") - - entities = - 1..@entity_count - |> Enum.map(fn i -> - input = build_octad_input(i) - case RustClient.create_octad(input) do - {:ok, octad} -> - id = octad["id"] || "entity-#{String.pad_leading(Integer.to_string(i), 6, "0")}" - if rem(i, 200) == 0, do: IO.write(" ... #{i}/#{@entity_count}\r") - id - {:error, _reason} -> - # Fall back to generating a local ID - "entity-#{String.pad_leading(Integer.to_string(i), 6, "0")}" - end - end) - - IO.puts(" Created #{length(entities)} octads ") - entities - end - - defp create_entities(:local) do - IO.puts(" Creating #{@entity_count} entities (local mode)...") - - entities = - 1..@entity_count - |> Enum.map(fn i -> - id = "entity-#{String.pad_leading(Integer.to_string(i), 6, "0")}" - - # Start an EntityServer for each entity - case EntityServer.start_link(id) do - {:ok, _pid} -> :ok - {:error, {:already_started, _pid}} -> :ok - end - - # Mark all 8 modalities as populated - EntityServer.update(id, Enum.map(@modalities, fn m -> {:modality, m, true} end)) - - if rem(i, 200) == 0, do: IO.write(" ... #{i}/#{@entity_count}\r") - id - end) - - IO.puts(" Created #{length(entities)} entities (local) ") - entities - end - - defp build_octad_input(i) do - # Generate realistic-looking data for all 8 modalities - lat = 51.5 + :rand.uniform() * 0.1 - 0.05 - lon = -0.12 + :rand.uniform() * 0.1 - 0.05 - - %{ - title: "Entity ##{i}: #{Enum.random(entity_names())}", - body: "Cross-modal entity #{i} with full octad representation. " <> - "Created for drift detection demo. Category: #{Enum.random(categories())}.", - embedding: Enum.map(1..384, fn _ -> :rand.uniform() * 2 - 1 end), - types: ["https://verisimdb.dev/ontology/#{Enum.random(categories())}"], - relationships: [ - {"relatesTo", "entity-#{String.pad_leading(Integer.to_string(max(1, i - 1)), 6, "0")}"}, - {"inCategory", "category-#{Enum.random(1..20)}"} - ], - provenance: %{ - event_type: "created", - actor: "drift-demo@verisimdb.dev", - source: "demo-script", - description: "Created by drift detection demo" - }, - spatial: %{ - latitude: lat, - longitude: lon, - geometry_type: "Point", - properties: %{"region" => "demo-region"} - }, - metadata: %{ - demo: true, - batch: "drift-detection-#{Date.utc_today()}" - } - } - end - - # ── Phase 2: Corrupt Entities ────────────────────────────────────────── - - defp corrupt_entities(mode, entities) do - IO.puts(" Corrupting #{@corrupt_count} entities with cross-modal drift...") - - # Select random entities to corrupt - corrupted_ids = - entities - |> Enum.shuffle() - |> Enum.take(@corrupt_count) - - corrupted_ids - |> Enum.with_index(1) - |> Enum.each(fn {entity_id, idx} -> - # Pick a random corruption pattern - pattern = Enum.random(@corruption_patterns) - - case mode do - :live -> - corrupt_live(entity_id, pattern) - :local -> - corrupt_local(entity_id, pattern) - end - - if rem(idx, 10) == 0 do - IO.write(" ... #{idx}/#{@corrupt_count} corrupted\r") - end - end) - - IO.puts(" Corrupted #{@corrupt_count} entities ") - corrupted_ids - end - - defp corrupt_live(entity_id, pattern) do - # In live mode, update the octad to introduce inconsistency - corruption_payload = case pattern.type do - :semantic_vector -> - # Change embedding without updating document - %{embedding: Enum.map(1..384, fn _ -> :rand.uniform() end)} - - :graph_document -> - # Change document without updating graph - %{body: "CORRUPTED: This content no longer matches the graph edges."} - - :temporal_consistency -> - # Force a version skip (metadata-level corruption) - %{metadata: %{force_version_skip: true}} - - :tensor -> - # Change tensor shape - %{metadata: %{corrupted_tensor_shape: true}} - - :schema -> - # Mark a required modality as missing - %{metadata: %{removed_modality: "provenance"}} - - :quality -> - # General degradation — scramble multiple fields - %{body: "", embedding: Enum.map(1..384, fn _ -> 0.0 end)} - end - - RustClient.update_octad(entity_id, corruption_payload) - - # Report the drift to DriftMonitor - DriftMonitor.report_drift(entity_id, pattern.score, pattern.type) - end - - defp corrupt_local(entity_id, pattern) do - # In local mode, simulate corruption via DriftMonitor - DriftMonitor.report_drift(entity_id, pattern.score, pattern.type) - - # Also update entity modality status for schema drift - if pattern.type == :schema do - EntityServer.update(entity_id, [{:modality, :provenance, false}]) - end - end - - # ── Phase 3: Detect Drift ────────────────────────────────────────────── - - defp detect_drift(_mode) do - IO.puts(" Running drift detection sweep...") - - # Trigger a manual sweep - DriftMonitor.sweep() - - # Give async tasks a moment to propagate - Process.sleep(500) - - # Collect drift status - status = DriftMonitor.status() - - detected_count = status.entities_with_drift - - IO.puts(" Detected #{detected_count} entities with drift") - IO.puts(" Overall health: #{status.overall_health}") - - Enum.each(status.drift_by_type, fn {type, stats} -> - IO.puts(" #{type}: avg=#{Float.round(stats.average, 3)}, " <> - "max=#{Float.round(stats.max, 3)}, count=#{stats.count}") - end) - - %{ - detected_count: detected_count, - health: status.overall_health, - drift_by_type: status.drift_by_type, - pending_normalizations: status.pending_normalizations - } - end - - # ── Phase 4: Repair via Normalisation ────────────────────────────────── - - defp repair_drift(mode, corrupted_ids) do - IO.puts(" Triggering normalisation for #{length(corrupted_ids)} corrupted entities...") - - results = - corrupted_ids - |> Enum.with_index(1) - |> Enum.map(fn {entity_id, idx} -> - result = case mode do - :live -> - case RustClient.normalize(entity_id) do - :ok -> :repaired - {:error, _} -> :failed - end - :local -> - # In local mode, simulate normalisation by clearing drift - DriftMonitor.report_drift(entity_id, 0.0, :quality) - DriftMonitor.report_drift(entity_id, 0.0, :semantic_vector) - DriftMonitor.report_drift(entity_id, 0.0, :graph_document) - DriftMonitor.report_drift(entity_id, 0.0, :temporal_consistency) - DriftMonitor.report_drift(entity_id, 0.0, :tensor) - DriftMonitor.report_drift(entity_id, 0.0, :schema) - - # Restore any removed modalities - EntityServer.update(entity_id, [{:modality, :provenance, true}]) - - :repaired - end - - if rem(idx, 10) == 0 do - IO.write(" ... #{idx}/#{length(corrupted_ids)} normalised\r") - end - - {entity_id, result} - end) - - repaired = Enum.count(results, fn {_, r} -> r == :repaired end) - failed = Enum.count(results, fn {_, r} -> r == :failed end) - - IO.puts(" Normalisation complete: #{repaired} repaired, #{failed} failed") - - %{repaired: repaired, failed: failed, results: results} - end - - # ── Phase 5: Verify Consistency ──────────────────────────────────────── - - defp verify_consistency(mode, corrupted_ids) do - IO.puts(" Verifying consistency post-repair...") - - # Give normalisation a moment to complete - Process.sleep(500) - - consistent_count = - corrupted_ids - |> Enum.count(fn entity_id -> - case mode do - :live -> - case RustClient.get_drift_score(entity_id) do - {:ok, score} when score < 0.3 -> true - _ -> false - end - :local -> - history = DriftMonitor.entity_history(entity_id) - max_drift = history |> Map.values() |> Enum.max(fn -> 0.0 end) - max_drift < 0.3 - end - end) - - # Check overall system health after repair - status = DriftMonitor.status() - - IO.puts(" Consistent entities: #{consistent_count}/#{length(corrupted_ids)}") - IO.puts(" Post-repair system health: #{status.overall_health}") - - %{ - consistent: consistent_count, - total: length(corrupted_ids), - post_repair_health: status.overall_health - } - end - - # ── Output Formatting ────────────────────────────────────────────────── - - defp print_banner do - IO.puts(""" - - ╔══════════════════════════════════════════════════════════════════╗ - ║ ║ - ║ VeriSimDB — Drift Detection & Self-Normalisation Demo ║ - ║ ║ - ║ Cross-modal consistency for the octad (8 modalities): ║ - ║ Graph | Vector | Tensor | Semantic | Document | ║ - ║ Temporal | Provenance | Spatial ║ - ║ ║ - ╚══════════════════════════════════════════════════════════════════╝ - """) - end - - defp print_mode(:live) do - IO.puts(""" - Mode: LIVE (Rust core connected) - Octads will be created, corrupted, and repaired via the Rust API. - """) - end - - defp print_mode(:local) do - IO.puts(""" - Mode: LOCAL (Rust core unavailable) - Simulating via Elixir DriftMonitor and EntityServer. - Start the Rust core (cargo run -p verisim-api) for live mode. - """) - end - - defp print_phase_complete(phase, description, microseconds) do - ms = Float.round(microseconds / 1_000, 1) - IO.puts("") - IO.puts(" Phase #{phase} complete: #{description} (#{ms}ms)") - IO.puts(" " <> String.duplicate("─", 60)) - IO.puts("") - end - - defp print_summary(summary) do - total_ms = Float.round(summary.timings.total_us / 1_000, 1) - detection_rate = if summary.corrupt_count > 0 do - Float.round(summary.detected.detected_count / summary.corrupt_count * 100, 1) - else - 0.0 - end - repair_rate = if summary.corrupt_count > 0 do - Float.round(summary.repaired.repaired / summary.corrupt_count * 100, 1) - else - 0.0 - end - consistency_rate = if summary.verified.total > 0 do - Float.round(summary.verified.consistent / summary.verified.total * 100, 1) - else - 0.0 - end - - IO.puts(""" - - ╔══════════════════════════════════════════════════════════════════╗ - ║ DEMO RESULTS ║ - ╠══════════════════════════════════════════════════════════════════╣ - ║ ║ - ║ Mode: #{String.pad_trailing(Atom.to_string(summary.mode), 42)}║ - ║ Entities created: #{String.pad_trailing(Integer.to_string(summary.entity_count), 42)}║ - ║ Entities corrupted: #{String.pad_trailing(Integer.to_string(summary.corrupt_count), 40)}║ - ║ ║ - ╠──────────────────────────────────────────────────────────────────╣ - ║ DETECTION ║ - ║ Entities with drift: #{String.pad_trailing(Integer.to_string(summary.detected.detected_count), 36)}║ - ║ Detection rate: #{String.pad_trailing("#{detection_rate}%", 36)}║ - ║ ║ - ║ REPAIR ║ - ║ Repaired: #{String.pad_trailing(Integer.to_string(summary.repaired.repaired), 36)}║ - ║ Failed: #{String.pad_trailing(Integer.to_string(summary.repaired.failed), 36)}║ - ║ Repair rate: #{String.pad_trailing("#{repair_rate}%", 36)}║ - ║ ║ - ║ VERIFICATION ║ - ║ Consistent post-repair: #{String.pad_trailing("#{summary.verified.consistent}/#{summary.verified.total}", 33)}║ - ║ Consistency rate: #{String.pad_trailing("#{consistency_rate}%", 33)}║ - ║ System health: #{String.pad_trailing(Atom.to_string(summary.verified.post_repair_health), 33)}║ - ║ ║ - ╠──────────────────────────────────────────────────────────────────╣ - ║ TIMING ║ - ║ Create: #{String.pad_trailing("#{Float.round(summary.timings.create_us / 1_000, 1)}ms", 43)}║ - ║ Corrupt: #{String.pad_trailing("#{Float.round(summary.timings.corrupt_us / 1_000, 1)}ms", 43)}║ - ║ Detect: #{String.pad_trailing("#{Float.round(summary.timings.detect_us / 1_000, 1)}ms", 43)}║ - ║ Repair: #{String.pad_trailing("#{Float.round(summary.timings.repair_us / 1_000, 1)}ms", 43)}║ - ║ Verify: #{String.pad_trailing("#{Float.round(summary.timings.verify_us / 1_000, 1)}ms", 43)}║ - ║ ───────────────────────────── ║ - ║ Total: #{String.pad_trailing("#{total_ms}ms", 43)}║ - ║ ║ - ╚══════════════════════════════════════════════════════════════════╝ - """) - end - - # ── Data Generators ──────────────────────────────────────────────────── - - defp entity_names do - [ - "Research Paper", "Dataset Record", "Person Profile", "Organisation", - "Event Log", "Sensor Reading", "Transaction", "Contract", - "Knowledge Claim", "Audit Trail", "Policy Document", "Spatial Feature", - "Time Series Point", "Graph Fragment", "Embedding Vector", "Tensor Block", - "Provenance Chain", "Citation Link", "Access Control Entry", "Schema Definition" - ] - end - - defp categories do - [ - "Research", "Finance", "Healthcare", "Education", "Government", - "Technology", "Science", "Engineering", "Legal", "Environmental" - ] - end -end - -# ── Run the demo ────────────────────────────────────────────────────────── - -VeriSimDB.Demo.DriftDetection.run() diff --git a/verisimdb/deny.toml b/verisimdb/deny.toml deleted file mode 100644 index c7963c1c..00000000 --- a/verisimdb/deny.toml +++ /dev/null @@ -1,37 +0,0 @@ -# SPDX-License-Identifier: MPL-2.0 -# cargo-deny configuration for VeriSimDB -# Run: cargo deny check - -[advisories] -db-path = "~/.cargo/advisory-db" -db-urls = ["https://github.com/rustsec/advisory-db"] -# Allow known transitive dep vulnerability (lru 0.12.5, LOW severity) -ignore = ["RUSTSEC-2026-0002"] - -[licenses] -allow = [ - "MIT", - "Apache-2.0", - "Apache-2.0 WITH LLVM-exception", - "BSD-2-Clause", - "BSD-3-Clause", - "ISC", - "Zlib", - "BSL-1.0", - "Unicode-3.0", - "Unicode-DFS-2016", - "OpenSSL", - "MPL-2.0", -] -confidence-threshold = 0.8 - -[bans] -multiple-versions = "warn" -wildcards = "deny" -deny = [] - -[sources] -unknown-registry = "deny" -unknown-git = "warn" -allow-registry = ["https://github.com/rust-lang/crates.io-index"] -allow-git = [] diff --git a/verisimdb/docs/.well-known/void.rdf b/verisimdb/docs/.well-known/void.rdf deleted file mode 100644 index a2b3dc53..00000000 --- a/verisimdb/docs/.well-known/void.rdf +++ /dev/null @@ -1,16 +0,0 @@ - - - - - verisimdb Dataset - Linked data from verisimdb project - - - - - - diff --git a/verisimdb/docs/.well-known/void.ttl b/verisimdb/docs/.well-known/void.ttl deleted file mode 100644 index 57ce0131..00000000 --- a/verisimdb/docs/.well-known/void.ttl +++ /dev/null @@ -1,75 +0,0 @@ -@prefix void: . -@prefix rdf: . -@prefix rdfs: . -@prefix owl: . -@prefix dcterms: . -@prefix foaf: . -@prefix xsd: . - -# Dataset Description - a void:Dataset ; - dcterms:title "verisimdb Dataset" ; - dcterms:description "Linked data from verisimdb project" ; - dcterms:creator ; - dcterms:publisher ; - dcterms:license ; - dcterms:created "2026-01-31T16:49:57Z"^^xsd:dateTime ; - dcterms:modified "2026-01-31T16:49:57Z"^^xsd:dateTime ; - - # Dataset statistics (update these based on actual data) - void:triples 0 ; - void:entities 0 ; - void:distinctSubjects 0 ; - void:distinctObjects 0 ; - - # Access methods - void:sparqlEndpoint ; - void:dataDump ; - void:dataDump ; - void:dataDump ; - - # Technical details - void:feature ; - void:feature ; - void:feature ; - - # Vocabulary usage (customize based on your data model) - void:vocabulary ; - void:vocabulary ; - void:vocabulary ; - void:vocabulary ; - - # Example linksets (connections to other datasets) - # void:subset ; - # void:subset ; -. - -# Creator information - a foaf:Person ; - foaf:name "Jonathan D.A. Jewell" ; - foaf:mbox ; - foaf:homepage ; - foaf:account ; -. - -# Publisher information - a foaf:Organization ; - foaf:name "hyperpolymath" ; - foaf:homepage ; -. - -# Example linkset to DBpedia (uncomment and customize) -# a void:Linkset ; -# void:linkPredicate owl:sameAs ; -# void:target ; -# void:target ; -# void:triples 0 ; -# . - -# Example linkset to Wikidata (uncomment and customize) -# a void:Linkset ; -# void:linkPredicate owl:sameAs ; -# void:target ; -# void:target ; -# void:triples 0 ; -# . diff --git a/verisimdb/docs/CITATIONS.adoc b/verisimdb/docs/CITATIONS.adoc deleted file mode 100644 index 6f167bdf..00000000 --- a/verisimdb/docs/CITATIONS.adoc +++ /dev/null @@ -1,36 +0,0 @@ -= RSR-template-repo - Citation Guide -:toc: - -== BibTeX - -[source,bibtex] ----- -@software{rsr-template-repo_2025, - author = {Polymath, Hyper}, - title = {RSR-template-repo}, - year = {2025}, - url = {https://github.com/hyperpolymath/RSR-template-repo}, - license = {PMPL-1.0-or-later} -} ----- - -== Harvard Style - -Polymath, H. (2025) _RSR-template-repo_ [Computer software]. Available at: https://github.com/hyperpolymath/RSR-template-repo - -== OSCOLA - -Hyper Polymath, 'RSR-template-repo' (2025) - -== MLA - -Polymath, Hyper. "RSR-template-repo." 2025, github.com/hyperpolymath/RSR-template-repo. - -== APA 7 - -Polymath, H. (2025). _RSR-template-repo_ [Computer software]. GitHub. https://github.com/hyperpolymath/RSR-template-repo - -== See Also - -* link:../CITATION.cff[CITATION.cff] -* link:../codemeta.json[codemeta.json] diff --git a/verisimdb/docs/VCL-SPEC.adoc b/verisimdb/docs/VCL-SPEC.adoc deleted file mode 100644 index ee97b85e..00000000 --- a/verisimdb/docs/VCL-SPEC.adoc +++ /dev/null @@ -1,3034 +0,0 @@ -// SPDX-License-Identifier: CC-BY-SA-4.0 -// Copyright (c) 2026 Jonathan D.A. Jewell (hyperpolymath) - -= VeriSim Consonance Language (VCL) Specification -:author: Jonathan D.A. Jewell -:email: j.d.a.jewell@open.ac.uk -:revnumber: 2.0 -:revdate: 2026-02-27 -:toc: left -:toclevels: 4 -:sectnums: -:stem: latexmath -:icons: font -:source-highlighter: rouge - -// ============================================================================ -// 1. INTRODUCTION -// ============================================================================ - -== Introduction - -=== Name and Terminology - -[NOTE] -==== -**VCL = VeriSim Consonance Language.** - -The word _consonance_ is deliberate: it means agreement and harmony between -parts. VeriSimDB maintains a single entity across eight modalities simultaneously -(the octad — Graph, Vector, Tensor, Semantic, Document, Temporal, Provenance, -Spatial). Those eight representations must stay in consonance with each other. -The query language takes its name from that guarantee. - -**VCL-UT = VCL Usage-Tracked types** — the type-theoretic extension tier, housed -in `typeql-experimental/`. The "UT" abbreviation expands to _Usage-Tracked_ -because these extensions track resource consumption (linear consumption counts, -session state, side-effects, usage limits). File extension: `.vclut`. - -Prior name **VCL** (VeriSim Consonance Language) is retired. Prior extension name -**VCL-UT** is retired. Use VCL and VCL-UT exclusively. -==== - -=== Purpose and Scope - -The VeriSim Consonance Language (VCL) is the native query interface for VeriSimDB, a cross-system entity consistency engine with drift detection, self-normalisation, and formally verified queries. This document is the **normative language specification** for VCL version 2.0. - -VCL is not SQL. It is a domain-specific language for querying and mutating **octad entities** — data objects that exist simultaneously across up to eight modalities (Graph, Vector, Tensor, Semantic, Document, Temporal, Provenance, Spatial). VCL provides: - -* **Multi-modal querying** — SELECT across any combination of eight modalities in a single statement. -* **Dual-path execution** — Slipstream (fast, unverified) and VCL-UT (formally verified with ZKP proofs). -* **Federation** — Query across distributed VeriSimDB instances with configurable drift policies. -* **Cross-modal conditions** — Filter on relationships _between_ modalities (drift, consistency, existence). -* **Write path** — INSERT, UPDATE, DELETE with optional proof verification. - -This specification unifies and supersedes the information in: - -* `vcl-grammar.ebnf` — Normative EBNF grammar (included in full as <>) -* `vcl-type-system.adoc` — Formal type system -* `vcl-formal-semantics.adoc` — Operational and denotational semantics -* `vcl-examples.adoc` — 63 comprehensive examples -* `vcl-architecture.adoc` — Dual-path routing architecture -* `vcl-vs-vcl-dt.adoc` — Slipstream vs dependent-type comparison -* `vcl-vs-sql.adoc` — SQL differences - -Those documents remain valid for deep dives; this specification is the authoritative reference for language behaviour. - -=== Notational Conventions - -[cols="1,3"] -|=== -| Convention | Meaning - -| `KEYWORD` -| VCL keyword (case-insensitive in source, uppercase in this document) - -| `identifier` -| User-defined name - -| `` -| Grammar production rule (see <> for full EBNF) - -| stem:[\tau] -| Type expression - -| stem:[\Gamma \vdash e : \tau] -| Typing judgment: expression _e_ has type stem:[\tau] in context stem:[\Gamma] - -| `[.implemented]` -| Feature is fully implemented and tested - -| `[.partial]` -| Feature is partially implemented (gaps noted) - -| `[.planned]` -| Feature is specified but not yet implemented -|=== - -Keywords are **case-insensitive** throughout VCL. `SELECT`, `select`, and `Select` are equivalent. - -=== Dual-Path Architecture - -VCL operates in two execution modes, determined by the presence or absence of a `PROOF` clause: - -[cols="1,1,1"] -|=== -| Property | Slipstream (no PROOF) | VCL-UT (with PROOF) - -| Parse -| Same parser, same AST -| Same parser, same AST - -| Type checking -| Basic validation only -| Full dependent type checking - -| Proof generation -| Skipped -| ZKP witness generation per proof obligation - -| Result type -| `List(Octad_M)` — bare results -| `Sigma(QueryResult, Proof)` — results bundled with proof certificate - -| Latency -| ~50-500ms -| ~240-1350ms - -| Use case -| Analytics, exploration, performance-sensitive -| Compliance, audit, provenance, federation trust -|=== - -Both modes use the same parser and produce the same AST. They diverge at the **query router**, which inspects the AST for a `PROOF` node and selects the execution pipeline. - -.Dual-path architecture ----- - ┌──────────────────────────────────────────────────┐ - │ VCL Statement │ - └───────────────────┬──────────────────────────────┘ - │ - ┌─────▼──────┐ - │ Parser │ (ReScript / Elixir fallback) - └─────┬──────┘ - │ AST - ┌─────▼──────┐ - │ Type Check │ (VCLBidir.res) - └─────┬──────┘ - │ - ┌──────────┴──────────┐ - │ PROOF clause? │ - ┌────▼────┐ ┌─────▼─────┐ - │ NO │ │ YES │ - ┌────▼────┐ ┌─────▼─────┐ - │Slipstream│ │ VCL-UT │ - │ Path │ │ Path │ - └────┬────┘ └─────┬─────┘ - │ │ - │ ├─ Dependent type check - │ ├─ Proof obligation gen - │ ├─ ZKP witness gen - │ │ - ┌────▼─────────────────────▼────┐ - │ Elixir Orchestrator │ - │ (QueryRouter GenServer) │ - └────┬──────────────────────┬───┘ - │ │ - ┌────────────▼──────────┐ ┌───────▼───────────┐ - │ Modality Stores │ │ Federation Fan-out │ - │ (Rust: Oxigraph, │ │ (multi-node) │ - │ Milvus, Tantivy, │ └───────────────────┘ - │ Burn, etc.) │ - └───────────────────────┘ ----- - -`[.implemented]` Parser, basic type checking, query routing, store dispatch. + -`[.partial]` Bidirectional type inference (implemented in ReScript, not wired to runtime Lean checker). + -`[.planned]` ZKP proof generation via `proven-library` / `sanctify`. - -=== Conformance Levels - -A VCL implementation MAY support one or both execution paths: - -* **Level 1 — Slipstream**: Parse, validate, execute queries without proof obligations. All SELECT, FROM, WHERE, GROUP BY, HAVING, ORDER BY, LIMIT, OFFSET, and mutation statements MUST be supported. -* **Level 2 — Full (Slipstream + VCL-UT)**: All Level 1 features plus PROOF clause parsing, dependent type checking, proof generation, and proof certificate bundling. - -VeriSimDB's current implementation is Level 1 complete with Level 2 partially implemented (parser and type checker complete; runtime proof generation planned). - - -// ============================================================================ -// 2. LEXICAL STRUCTURE -// ============================================================================ - -== Lexical Structure - -=== Character Set - -VCL source text is encoded in **UTF-8**. Identifiers are restricted to ASCII letters, digits, and underscores. String literals may contain any valid UTF-8 sequence. - -=== Keywords - -Keywords are **case-insensitive**. The following 70 tokens are reserved and MUST NOT be used as identifiers: - -.Reserved keywords (alphabetical) ----- -ACCESS AND AS ASC AVG BETWEEN BY -CITATION CONSISTENT CONTAINS COSINE COUNT CUSTOM DELETE -DESC DOCUMENT DOT_PRODUCT DRIFT EUCLIDEAN EXISTENCE EXISTS -FEDERATION FIELD FROM FULLTEXT GRAPH GROUP HAS -HAVING HEXAD INSERT INTEGRITY JACCARD LATEST LIKE -LIMIT MATCHES MAX MIN MODIFIED NEAREST NOT -OF OFFSET OR ORDER PROOF PROVENANCE RANK -REPAIR SELECT SEMANTIC SET SHAPE SIMILAR STORE -STRICT SUM TEMPORAL TENSOR TO TOLERATE UPDATE -USING VECTOR VERIFIED VERSION WHERE WITH WITHIN -false true ----- - -`[.implemented]` All keywords are recognised by the parser. - -=== Identifiers - ----- -identifier = letter_or_underscore , { letter_or_digit_or_underscore } ; ----- - -Identifiers are **case-sensitive** (unlike keywords). They name stores, contracts, verifiers, actors, fields, and variables. - -.Examples ----- -my_store -- valid -research_papers -- valid -node1 -- valid -123invalid -- INVALID: starts with digit -SELECT -- INVALID: reserved keyword ----- - -=== Literals - -==== Integer Literals - ----- -integer = digit , { digit } ; ----- - -Non-negative integers. Examples: `0`, `42`, `1000`. - -==== Float Literals - ----- -float = digit , { digit } , '.' , digit , { digit } , - [ ('e' | 'E') , ['+' | '-'] , digit , { digit } ] ; ----- - -IEEE 754 double-precision. Examples: `3.14`, `0.001`, `1.5e10`, `2.0E-3`. - -==== String Literals - ----- -string_literal = "'" , { any_char_except_quote_or_backslash - | escaped_quote | escaped_backslash } , "'" ; ----- - -Single-quoted. Escape sequences: `\'` (literal quote), `\\` (literal backslash). - -==== Boolean Literals - ----- -boolean = 'true' | 'false' ; ----- - -Case-sensitive (`true` and `false`, not `TRUE` or `FALSE`). - -==== Array Literals - ----- -array_literal = '[' , literal , { ',' , literal } , ']' ; ----- - -Homogeneous arrays. Used for vector data and tensor slices. - -==== UUID Literals - ----- -uuid = 8*hex_digit , '-' , 4*hex_digit , '-' , 4*hex_digit , '-' , - 4*hex_digit , '-' , 12*hex_digit ; ----- - -Standard RFC 4122 format. Example: `550e8400-e29b-41d4-a716-446655440000`. - -==== Timestamp Literals - ----- -timestamp = ISO-8601 datetime ; ----- - -Full ISO 8601 format with timezone. Example: `2026-02-13T10:30:00Z`. - -==== Regex Literals - ----- -regex_literal = '/' , { any_char_except_slash | '\/' | '\\' } , '/' ; ----- - -Forward-slash delimited regular expressions. Example: `/\b(AI|ML|DL)\b.*ethics/`. - -=== Operators - -[cols="1,2,1"] -|=== -| Operator | Meaning | Valid Types - -| `==` -| Equality -| All types - -| `!=` -| Inequality -| All types - -| `>` -| Greater than -| Int, Float, String, Timestamp - -| `<` -| Less than -| Int, Float, String, Timestamp - -| `>=` -| Greater than or equal -| Int, Float, String, Timestamp - -| `\<=` -| Less than or equal -| Int, Float, String, Timestamp - -| `LIKE` -| Pattern matching (% wildcard) -| String only - -| `CONTAINS` -| Substring containment (via FULLTEXT) -| String only - -| `MATCHES` -| Regex matching (via FULLTEXT) -| String only -|=== - -=== Comments - -VCL supports two comment styles: - ----- --- This is a line comment (extends to end of line) - -/* This is a - block comment - spanning multiple lines */ ----- - -Comments are stripped during lexing and have no semantic effect. - - -// ============================================================================ -// 3. DATA MODEL -// ============================================================================ - -== Data Model - -=== Octad Entities - -The fundamental unit of data in VeriSimDB is the **octad** — a single entity that exists simultaneously across up to eight modalities. Each octad is identified by a UUID and contains zero or more modality slots populated with data. - -NOTE: The name "octad" persists in some API surfaces and code for backward compatibility. The data model is octad (8 modalities). - -.Octad structure ----- -Octad := { - id: UUID, - graph: Option, - vector: Option, - tensor: Option, - semantic: Option, - document: Option, - temporal: Option, - provenance: Option, - spatial: Option -} ----- - -Not every octad populates every modality. An octad might have only Document and Vector data, or all eight. The `EXISTS` / `NOT EXISTS` conditions (<>) filter on modality presence. - -=== Modalities - -[cols="1,2,2,1"] -|=== -| Modality | Storage Engine | Data Model | Status - -| `GRAPH` -| Oxigraph (RDF/Property Graph) -| Triples, edges, property annotations -| `[.implemented]` - -| `VECTOR` -| HNSW (verisim-vector) -| Fixed-dimension float embeddings -| `[.implemented]` - -| `TENSOR` -| Burn / ndarray (verisim-tensor) -| Multi-dimensional numeric arrays with shape and dtype -| `[.implemented]` - -| `SEMANTIC` -| CBOR blobs (verisim-semantic) -| Type annotations, ZKP proof blobs, contract references -| `[.implemented]` - -| `DOCUMENT` -| Tantivy (verisim-document) -| Full-text searchable content with structured fields -| `[.implemented]` - -| `TEMPORAL` -| verisim-temporal (Merkle trees) -| Version history, time-series, actor trail -| `[.implemented]` - -| `PROVENANCE` -| verisim-provenance (hash chains) -| Origin tracking, transformation chain, actor trail, chain integrity verification -| `[.implemented]` - -| `SPATIAL` -| verisim-spatial (Haversine/brute-force) -| WGS84 coordinates, geometries, radius/bounds/nearest queries -| `[.implemented]` -|=== - -=== Primitive Types - -VCL's type system operates over these primitive types: - -[cols="1,2,2"] -|=== -| Type | Description | ReScript Representation - -| `Int` -| Arbitrary-precision integer -| `IntType` - -| `Float` -| IEEE 754 double-precision float -| `FloatType` - -| `String` -| UTF-8 text -| `StringType` - -| `Bool` -| Boolean (true/false) -| `BoolType` - -| `Vector` -| Fixed-size float vector of dimension _N_ -| `VectorType(int)` - -| `Tensor` -| Multi-dimensional array with shape `[d1, ..., dk]` -| `TensorType(array)` - -| `UUID` -| RFC 4122 universally unique identifier -| `UuidType` - -| `Timestamp` -| ISO 8601 datetime -| `TimestampType` -|=== - -=== Cross-Modal Consistency Model - -VeriSimDB continuously monitors **drift** — divergence between modality representations of the same octad. When an octad's vector embedding no longer matches its document content, or its graph structure diverges from its semantic annotations, drift is detected. - -Drift types monitored (8 total): - -* `semantic_vector_drift` — Embedding does not match semantic content -* `graph_document_drift` — Graph structure does not match document -* `temporal_consistency_drift` — Version history issues -* `tensor_drift` — Tensor representation diverged -* `provenance_drift` — Chain integrity broken, missing events, or stale chain -* `spatial_drift` — Coordinates inconsistent with graph/document location mentions -* `schema_drift` — Type constraint violations -* `quality_drift` — Overall data quality metric - -When drift exceeds configurable thresholds, the **self-normalizer** identifies the most authoritative modality, regenerates drifted modalities from it, validates consistency, and updates all modalities atomically. - - -// ============================================================================ -// 4. TYPE SYSTEM -// ============================================================================ - -== Type System - -VCL has a rich type system that operates in two modes depending on the execution path. The slipstream path uses simple types for basic validation. The VCL-UT path adds dependent types, refinement types, and proof types for formal verification. - -=== Type Language - -The full VCL type language stem:[\tau] is defined by: - -[stem] -++++ -\begin{aligned} -\tau ::= &\ \text{UUID} \mid \text{String} \mid \text{Int} \mid \text{Float} \mid \text{Bool} \\ - | &\ \text{Vector}[n] \mid \text{Tensor}[d_1, \ldots, d_k] \\ - | &\ \text{Timestamp} \\ - | &\ \text{Octad} \mid \text{OctadRef} \\ - | &\ \text{Modality} \mid \text{ModalitySet} \\ - | &\ \text{List}(\tau) \mid \text{Option}(\tau) \\ - | &\ \tau_1 \times \tau_2 \quad \text{(product)} \\ - | &\ \tau_1 \to \tau_2 \quad \text{(function)} \\ - | &\ \{x : \tau \mid \phi(x)\} \quad \text{(refinement)} \\ - | &\ \Pi x : \tau_1. \tau_2(x) \quad \text{(dependent product)} \\ - | &\ \Sigma x : \tau_1. \tau_2(x) \quad \text{(dependent sum)} \\ - | &\ \text{Proof}[\phi] \quad \text{(proof type)} -\end{aligned} -++++ - -=== Simple Types (Slipstream Path) - -In slipstream mode, the type checker validates: - -1. **Modality resolution** — Requested modalities exist and are valid -2. **Field types** — Field references (`MODALITY.field`) resolve to known primitive types -3. **Operator compatibility** — Comparison operators match operand types -4. **Aggregate validity** — Aggregate functions operate on appropriate types (e.g., `SUM` requires numeric) - -.ReScript type representation -[source,rescript] ----- -type rec vclType = - | Primitive(primitiveType) - | ArrayType(vclType) - | ModalityType(modalityType) - | OctadType(array) - | QueryResultType(queryResultInfo) - | ProofType(proofKind, string) - | ProvedResultType(queryResultInfo, proofKind, string) - | PiType(string, vclType, vclType) - | SigmaType(string, vclType, vclType) - | UnitType - | NeverType ----- - -`[.implemented]` All simple type checking in `VCLBidir.res`. - -=== Dependent Types (VCL-UT Path) - -When a `PROOF` clause is present, the type checker activates dependent type checking: - -==== Pi Types (Dependent Functions) - -stem:[\Pi x : \tau_1. \tau_2(x)] — The codomain type depends on the argument value. Used internally to type modality projections: - -[stem] -++++ -\text{VECTOR} : \Pi n : \text{Nat}. \text{Octad} \to \text{Option}(\text{Vector}[n]) -++++ - -==== Sigma Types (Dependent Pairs) - -stem:[\Sigma x : \tau_1. \tau_2(x)] — A pair where the type of the second component depends on the value of the first. This is the return type of proved queries: - -[stem] -++++ -\text{ProvedQuery} : \Sigma r : \text{QueryResult}[M]. \text{Proof}[\phi(r)] -++++ - -The result is bundled with a proof that the result satisfies the stated property stem:[\phi]. - -==== Proof Types - -`Proof[φ]` — A type inhabited by evidence that proposition stem:[\phi] holds. There are six proof kinds: - -[cols="1,3,1"] -|=== -| Proof Kind | Verifies | Cost - -| `EXISTENCE` -| Octad exists and is accessible in the store -| Low - -| `CITATION` -| Citation chain is valid (graph traversal verification) -| Medium - -| `ACCESS` -| User has verified access rights (ZKP access control) -| Medium - -| `INTEGRITY` -| Data has not been tampered with (Merkle root verification) -| Medium - -| `PROVENANCE` -| Lineage is verifiable (chain of custody) -| High - -| `CUSTOM` -| User-defined ZKP contract -| Variable -|=== - -`[.partial]` Parser recognises all six proof types. Type checker validates proof clause structure. Runtime proof generation is planned. - -=== Subtyping Rules - -The subtyping relation stem:[\tau_1 <: \tau_2] governs when a value of type stem:[\tau_1] can be used where stem:[\tau_2] is expected: - -[stem] -++++ -\begin{aligned} -& \frac{}{\tau <: \tau} \quad \text{(Reflexivity)} \\[6pt] -& \frac{\tau_1 <: \tau_2 \quad \tau_2 <: \tau_3}{\tau_1 <: \tau_3} \quad \text{(Transitivity)} \\[6pt] -& \frac{}{\{x : \tau \mid \phi(x)\} <: \tau} \quad \text{(Refinement Subsumption)} \\[6pt] -& \frac{\tau_1 <: \tau_2}{\text{List}(\tau_1) <: \text{List}(\tau_2)} \quad \text{(List Covariance)} \\[6pt] -& \frac{\tau_2 <: \tau_1 \quad \sigma_1 <: \sigma_2}{\tau_1 \to \sigma_1 <: \tau_2 \to \sigma_2} \quad \text{(Function Contra/Covariance)} -\end{aligned} -++++ - -`[.implemented]` Structural subtyping in `VCLBidir.res`. - -=== Type Inference (Bidirectional) - -VCL uses **bidirectional type checking** with two modes: - -* **Synthesis** (stem:[\Gamma \vdash e \Rightarrow \tau]) — Infer the type of an expression from its structure. -* **Checking** (stem:[\Gamma \vdash e \Leftarrow \tau]) — Verify that an expression has an expected type. - -The `synthesizeQuery` function in `VCLBidir.res` walks the query AST through nine phases: - -1. Resolve modalities from SELECT clause -2. Check source (HEXAD/FEDERATION/STORE) typing -3. Check WHERE clause conditions -4. Check field projections -5. Check aggregate expressions -6. Check GROUP BY fields -7. Check ORDER BY fields -8. Build result type -9. Handle PROOF clause (if present, switch to dependent-type path) - -.Core typing rules -[stem] -++++ -\frac{\Gamma \vdash q : \text{Query}[M] \quad \text{no PROOF clause}}{\Gamma \vdash q \Rightarrow \text{List}(\text{Octad}_M)} \quad \text{(T-SlipstreamQuery)} -++++ - -[stem] -++++ -\frac{\Gamma \vdash q : \text{Query}[M] \quad \Gamma \vdash p : \text{ProofSpec}[\phi]}{\Gamma \vdash q\ \text{PROOF}\ p \Rightarrow \Sigma r : \text{QueryResult}[M]. \text{Proof}[\phi(r)]} \quad \text{(T-ProvedQuery)} -++++ - -`[.implemented]` Bidirectional type inference in `VCLBidir.res` (841 lines, 9-phase pipeline). - -=== Refinement Types - -Refinement types stem:[\{x : \tau \mid \phi(x)\}] constrain values beyond their base type. In VCL, refinement types appear implicitly in WHERE conditions: - ----- -WHERE FIELD severity > 3 ----- - -This produces a refined query result type: stem:[\{h : \text{Octad} \mid h.\text{severity} > 3\}]. - -`[.partial]` Refinement types are implicit in condition checking. Explicit refinement type syntax is planned. - -=== Type Safety Properties - -The VCL type system satisfies three key properties (proven in `vcl-type-system.adoc`): - -* **Progress** — A well-typed query can always take a step (it is never stuck). -* **Preservation** — If a query steps, the result is still well-typed. -* **Soundness** — The type checker only accepts queries that produce well-typed results. - - -// ============================================================================ -// 5. QUERY STATEMENTS -// ============================================================================ - -== Query Statements - -A query statement retrieves octad data and optionally verifies it: - -.Grammar ----- -query = select_clause , from_clause , [where_clause] , - [group_by_clause] , [having_clause] , [proof_clause] , - [order_by_clause] , [limit_clause] , [offset_clause] ; ----- - -Clause order is **fixed**: SELECT, FROM, WHERE, GROUP BY, HAVING, PROOF, ORDER BY, LIMIT, OFFSET. Omitting optional clauses is valid. The only mandatory clauses are SELECT and FROM. - -=== SELECT Clause - -The SELECT clause specifies which data to retrieve. - -.Grammar ----- -select_clause = 'SELECT' , select_item_list ; -select_item_list = select_item , { ',' , select_item } ; -select_item = aggregate_expr | field_ref | modality_spec ; ----- - -There are three kinds of select items: - -==== Modality Selection - -Select entire modalities by name: - -[source,vcl] ----- -SELECT GRAPH -- single modality -SELECT GRAPH, VECTOR, DOCUMENT -- multiple modalities -SELECT * -- all available modalities ----- - -Each modality may include a projection to limit the data returned: - ----- -SELECT GRAPH(nodes, edges) -- graph with specific projection -SELECT DOCUMENT(title, abstract) -- document fields only ----- - -==== Field Projections - -Select specific fields from a modality using dot notation (`MODALITY.field`): - -[source,vcl] ----- -SELECT DOCUMENT.name, DOCUMENT.severity ----- - -Field projections may be mixed with full modality selections: - -[source,vcl] ----- -SELECT GRAPH, DOCUMENT.name, DOCUMENT.severity ----- - -==== Aggregate Expressions - -Five aggregate functions are supported: - -[cols="1,2,2"] -|=== -| Function | Syntax | Return Type - -| `COUNT(*)` -| Count all matching octads -| `Int` - -| `COUNT(M.field)` -| Count non-null values of a field -| `Int` - -| `SUM(M.field)` -| Sum numeric field values -| `Float` (numeric fields only) - -| `AVG(M.field)` -| Average numeric field values -| `Float` (numeric fields only) - -| `MIN(M.field)` -| Minimum value -| Same as field type (comparable types only) - -| `MAX(M.field)` -| Maximum value -| Same as field type (comparable types only) -|=== - -.Type error: aggregate on non-numeric field ----- -SELECT SUM(DOCUMENT.title) -- ERROR: AggregateTypeMismatch - -- SUM cannot operate on String ----- - -.Example: aggregates -[source,vcl] ----- -SELECT COUNT(*), AVG(DOCUMENT.severity), MAX(DOCUMENT.severity) -FROM FEDERATION /scans/* ----- - -`[.implemented]` All aggregate functions, field projections, and modality selection. - -=== FROM Clause - -The FROM clause specifies the data source. - -.Grammar ----- -from_clause = 'FROM' , source_spec ; -source_spec = octad_source | federation_source | store_source ; ----- - -==== HEXAD Source - -Query a specific octad by UUID: - -[source,vcl] ----- -FROM HEXAD 550e8400-e29b-41d4-a716-446655440000 ----- - -Type: stem:[\text{Octad}(\text{UUID})] - -==== FEDERATION Source - -Query across multiple VeriSimDB instances matching a pattern: - -[source,vcl] ----- -FROM FEDERATION /universities/* -- glob pattern -FROM FEDERATION /universities/* WITH DRIFT STRICT -- with drift policy ----- - -Drift policies control how inter-node inconsistency is handled: - -[cols="1,3,1"] -|=== -| Policy | Behaviour | Cost - -| `STRICT` -| Fail immediately if any node has drifted. Guarantees perfect consistency. -| Low (fail-fast) - -| `REPAIR` -| Detect drift, auto-repair via DriftMonitor, then return consistent results. -| High (repair + retry) - -| `TOLERATE` -| Return data even if nodes are inconsistent. Results may contain stale data. -| Low - -| `LATEST` -| Use the most recent version from each node, ignoring temporal consistency. -| Low -|=== - -.Example: federation with drift policies -[source,vcl] ----- --- Strict: fail on any drift -SELECT * -FROM FEDERATION /legal-records/* WITH DRIFT STRICT -WHERE AS OF 2024-01-01T00:00:00Z -PROOF INTEGRITY(LegalContract) - --- Repair: auto-fix before returning -SELECT GRAPH, DOCUMENT -FROM FEDERATION /collaborative-wiki/* WITH DRIFT REPAIR -WHERE FULLTEXT CONTAINS "VeriSimDB" - --- Tolerate: return potentially inconsistent data -SELECT * -FROM FEDERATION /mirrors/* WITH DRIFT TOLERATE -WHERE FULLTEXT CONTAINS "archived content" -LIMIT 1000 - --- Latest: most recent version wins -SELECT DOCUMENT -FROM FEDERATION /news-feeds/* WITH DRIFT LATEST -WHERE FULLTEXT CONTAINS "breaking news" -LIMIT 10 ----- - -`[.implemented]` All drift modes parsed and routed. Drift detection implemented. Auto-repair partially implemented. - -==== STORE Source - -Query a specific modality store directly, bypassing federation: - -[source,vcl] ----- -FROM STORE milvus-us-east-1 -- direct store access -FROM STORE oxigraph-node-1 -- specific graph store ----- - -Useful for lowest-latency queries when you know which store has your data. - -`[.implemented]` Store-direct queries. - -=== WHERE Clause - -The WHERE clause filters octads based on conditions. - -.Grammar ----- -where_clause = 'WHERE' , condition ; -condition = simple_condition | compound_condition | '(' , condition , ')' ; -compound_condition = condition , 'AND' , condition - | condition , 'OR' , condition - | 'NOT' , condition ; ----- - -Conditions are composed with `AND`, `OR`, and `NOT`. Parentheses control precedence. Each simple condition targets a specific modality or operates across modalities. - -==== Graph Conditions - -Graph conditions use SPARQL-like patterns for traversing the RDF/property graph. - -.Grammar ----- -graph_condition = sparql_pattern | path_pattern ; -sparql_pattern = '(' , node_var , ')' , edge_pattern , '(' , node_var , ')' ; -edge_pattern = '-[' , edge_type , ']->' - | '-[' , edge_type , ']-' - | '<-[' , edge_type , ']-' ; -path_pattern = node_var , path_quantifier , node_var ; -path_quantifier = '-[' , edge_type , ('*' | '+' | '{' , int , ',' , int , '}') , ']->' ; ----- - -Edge patterns support three directions: - -* `-[:TYPE]\->` — Directed edge (outgoing) -* `-[:TYPE]-` — Undirected edge (either direction) -* `<-[:TYPE]-` — Reverse edge (incoming) - -Path quantifiers control traversal depth: - -* `*` — Zero or more hops -* `+` — One or more hops -* `{m,n}` — Between _m_ and _n_ hops - -.Examples: graph conditions -[source,vcl] ----- --- Direct citation -WHERE (h)-[:CITES]->(target) - --- Citation chain (1-5 hops) -WHERE (h)-[:CITES*1..5]->(paper) - --- Bidirectional co-authorship -WHERE (h)-[:CO_AUTHOR]-(colleague) - --- Property conditions on graph nodes -WHERE (author)-[:WROTE]->(paper) - AND paper.citations > 100 ----- - -.Type rule -[stem] -++++ -\frac{\Gamma \vdash v_1 : \text{Node} \quad \Gamma \vdash v_2 : \text{Node} \quad e : \text{EdgeType}}{\Gamma \vdash (v_1)\text{-[}e\text{]->(}v_2) : \text{GraphCondition}} -++++ - -`[.implemented]` SPARQL patterns, path quantifiers, property conditions. - -==== Vector Conditions - -Vector conditions support similarity search and K-nearest-neighbor queries. - -.Grammar ----- -vector_condition = vector_field , 'SIMILAR' , 'TO' , vector_literal , [similarity_threshold] - | vector_field , 'NEAREST' , integer , [metric_type] ; -similarity_threshold = 'WITHIN' , float ; -metric_type = 'USING' , ('COSINE' | 'EUCLIDEAN' | 'DOT_PRODUCT') ; ----- - -.Examples: vector conditions -[source,vcl] ----- --- Approximate nearest neighbor with threshold -WHERE h.embedding SIMILAR TO [0.12, 0.45, 0.78, ...] WITHIN 0.8 - --- K-nearest neighbors with metric -WHERE h.embedding NEAREST 15 USING DOT_PRODUCT - --- Similarity with explicit metric -WHERE h.embedding SIMILAR TO [0.12, 0.45, 0.78, ...] - WITHIN 0.9 - USING COSINE ----- - -.Type rule -[stem] -++++ -\frac{\Gamma \vdash f : \text{Vector}[n] \quad \Gamma \vdash v : \text{Vector}[m] \quad n = m}{\Gamma \vdash f\ \text{SIMILAR TO}\ v : \text{VectorCondition}} \quad \text{(T-VectorSim)} -++++ - -Dimension mismatch produces `VectorDimensionMismatch` error. - -`[.implemented]` Similarity search, KNN, all three metrics (COSINE, EUCLIDEAN, DOT_PRODUCT). - -==== Tensor Conditions - -Tensor conditions filter on shape, rank, and element properties. - -.Grammar ----- -tensor_condition = tensor_field , tensor_op , tensor_literal ; -tensor_op = '==' | '>' | '<' | '>=' | '<=' | 'SHAPE' | 'RANK' ; ----- - -.Examples: tensor conditions -[source,vcl] ----- --- Shape filtering -WHERE tensor.data SHAPE == [256, 256, 3] - --- Rank filtering with statistics -WHERE tensor.data RANK == 3 - AND tensor.statistics.mean > 0.5 - AND tensor.statistics.std < 0.2 ----- - -`[.implemented]` Shape and rank filtering. Element-wise operations partial. - -==== Semantic Conditions - -Semantic conditions verify ZKP contracts and proof properties. - -.Grammar ----- -semantic_condition = 'SATISFIES' , contract_name , [contract_params] - | 'HAS' , 'PROOF' , proof_type - | 'VERIFIED' , 'BY' , verifier_id ; -contract_params = '(' , param_list , ')' ; -param_list = identifier , '=' , literal , { ',' , identifier , '=' , literal } ; ----- - -.Examples: semantic conditions -[source,vcl] ----- --- Contract satisfaction -WHERE SATISFIES AccessControlContract(role=researcher, institution=MIT) - --- Existing proof check -WHERE HAS PROOF ACCESS - --- Verifier check -WHERE VERIFIED BY fda-validator - --- Multi-contract -WHERE SATISFIES GDPRContract(anonymized=true) - AND SATISFIES FAIRContract(findable=true, accessible=true) - AND SATISFIES LicenseContract(license=CC-BY-4.0) ----- - -`[.partial]` Parser and type checker handle semantic conditions. Runtime contract verification via ZKP is planned. - -==== Document Conditions - -Document conditions support full-text search, regex matching, and structured field queries against the Tantivy index. - -.Grammar ----- -document_condition = 'FULLTEXT' , 'CONTAINS' , string_literal - | 'FULLTEXT' , 'MATCHES' , regex_literal - | 'FIELD' , identifier , comparison_op , literal ; ----- - -.Examples: document conditions -[source,vcl] ----- --- Full-text search -WHERE FULLTEXT CONTAINS "quantum computing" - --- Regex pattern matching -WHERE FULLTEXT MATCHES /\b(AI|ML|DL)\b.*ethics/ - --- Structured field query -WHERE FIELD author == "Jane Doe" - AND FIELD year >= 2020 - --- Combining full-text with structured -WHERE FULLTEXT CONTAINS "deep learning" - AND FIELD doi LIKE "10.1234/%" - AND FIELD impact_factor >= 5.0 ----- - -`[.implemented]` Full-text search, regex matching, structured field queries. - -==== Temporal Conditions - -Temporal conditions query version history and time-series data backed by Merkle trees. - -.Grammar ----- -temporal_condition = 'AS' , 'OF' , timestamp - | 'BETWEEN' , timestamp , 'AND' , timestamp - | 'VERSION' , version_id - | 'MODIFIED' , 'BY' , actor_id ; ----- - -.Examples: temporal conditions -[source,vcl] ----- --- Point-in-time query -WHERE AS OF 2024-06-15T10:30:00Z - --- Version range -WHERE BETWEEN 2024-01-01T00:00:00Z AND 2024-12-31T23:59:59Z - --- Specific version -WHERE VERSION v3-final - --- Actor filter -WHERE MODIFIED BY legal-team ----- - -`[.implemented]` All temporal conditions parsed and routed to verisim-temporal. - -[[cross-modal-conditions]] -==== Cross-Modal Conditions - -Cross-modal conditions operate on relationships _between_ modalities rather than within a single modality. These conditions are evaluated **post-fetch** (after modality data has been retrieved from stores), unlike single-modality conditions which are pushed down to stores. - -.Grammar ----- -cross_modal_condition = cross_modal_field_compare - | drift_condition - | consistency_condition - | exists_condition - | not_exists_condition ; ----- - -===== Field Comparison Across Modalities - -Compare field values from different modalities on the same octad: - ----- -cross_modal_field_compare = field_ref , comparison_op , field_ref ; ----- - -[source,vcl] ----- -WHERE DOCUMENT.severity > GRAPH.centrality ----- - -.Type rule -[stem] -++++ -\frac{\Gamma \vdash f_1 : \tau_1 \quad \Gamma \vdash f_2 : \tau_2 \quad \tau_1 \sim \tau_2 \quad \text{op valid for } \tau_1}{\Gamma \vdash f_1\ \text{op}\ f_2 : \text{CrossModalCondition}} -++++ - -Both fields must be comparable types. `CrossModalTypeMismatch` error if they are not. - -===== Drift Detection Between Modalities - -Measure representation drift between two modalities: - ----- -drift_condition = 'DRIFT' , '(' , modality , ',' , modality , ')' , comparison_op , float ; ----- - -[source,vcl] ----- -WHERE DRIFT(VECTOR, DOCUMENT) > 0.3 ----- - -Returns octads where the drift score between the two modalities exceeds the threshold. The drift score is computed as a normalised distance metric (0.0 = perfectly consistent, 1.0 = maximally divergent). - -`DriftRequiresNumeric` error if either modality lacks a numeric/vector representation. - -===== Consistency Check - -Verify that two modalities are consistent using a specific metric: - ----- -consistency_condition = 'CONSISTENT' , '(' , modality , ',' , modality , ')' , - 'USING' , metric_name ; -metric_name = 'COSINE' | 'EUCLIDEAN' | 'DOT_PRODUCT' | 'JACCARD' ; ----- - -[source,vcl] ----- -WHERE CONSISTENT(VECTOR, SEMANTIC) USING COSINE ----- - -`ConsistencyMetricInvalid` error if the metric is not applicable to the modality pair. - -===== Modality Existence - -Filter octads based on which modalities are populated: - ----- -exists_condition = modality_name , 'EXISTS' ; -not_exists_condition = modality_name , 'NOT' , 'EXISTS' ; ----- - -[source,vcl] ----- --- Octads with vectors but without tensors -WHERE VECTOR EXISTS - AND TENSOR NOT EXISTS ----- - -.Combined example: cross-modal with standard conditions -[source,vcl] ----- -SELECT GRAPH, VECTOR, DOCUMENT -FROM FEDERATION /research-db/* WITH DRIFT REPAIR -WHERE FULLTEXT CONTAINS "neural networks" -- pushdown to document store - AND DRIFT(VECTOR, DOCUMENT) < 0.2 -- cross-modal (post-fetch) - AND VECTOR EXISTS -- cross-modal (post-fetch) - AND GRAPH EXISTS -- cross-modal (post-fetch) -PROOF CITATION(NeuralNetworkContract) - AND INTEGRITY(DataIntegrityContract) -LIMIT 50 ----- - -`[.implemented]` All cross-modal conditions: field compare, drift, consistency, exists/not-exists. - -=== PROOF Clause - -The PROOF clause activates the VCL-UT execution path and specifies proof obligations. - -.Grammar ----- -proof_clause = 'PROOF' , proof_spec_list ; -proof_spec_list = proof_spec , { 'AND' , proof_spec } ; -proof_spec = proof_type , '(' , contract_name , ')' , [proof_params] ; -proof_params = 'WITH' , param_list ; ----- - -Multiple proofs are composed with `AND` — all must pass before results are returned. - -.Examples: proof clause -[source,vcl] ----- --- Single proof -PROOF EXISTENCE(ExistenceContract) - --- Dual proof -PROOF EXISTENCE(ExistenceContract) AND INTEGRITY(DataIntegrityContract) - --- Triple proof with parameters -PROOF ACCESS(InstitutionalAccessContract) - AND PROVENANCE(ClinicalTrialContract) - AND INTEGRITY(DataIntegrityContract) - --- Custom proof with verifier -PROOF CUSTOM(ComplianceContract) WITH auditor=legal-team ----- - -.Type rule -[stem] -++++ -\frac{\Gamma \vdash p_1 : \text{Proof}[\phi_1] \quad \Gamma \vdash p_2 : \text{Proof}[\phi_2]}{\Gamma \vdash p_1\ \text{AND}\ p_2 : \text{Proof}[\phi_1 \land \phi_2]} \quad \text{(T-ProofCompose)} -++++ - -The type checker validates proof composability via the context's contract registry. `MultiProofConflict` error if proofs have conflicting requirements. - -`[.partial]` Parser and type checker handle proof clauses. Runtime ZKP generation planned. - -=== GROUP BY / HAVING - -Group results and filter groups. - -.Grammar ----- -group_by_clause = 'GROUP' , 'BY' , field_ref_list ; -having_clause = 'HAVING' , condition ; ----- - -`HAVING` requires `GROUP BY` — using HAVING without GROUP BY produces `HavingWithoutGroupBy` error. - -.Example: grouping with aggregate filter -[source,vcl] ----- -SELECT DOCUMENT.name, COUNT(*), AVG(DOCUMENT.severity) -FROM FEDERATION /universities/* -GROUP BY DOCUMENT.name -HAVING FIELD count > 3 -ORDER BY DOCUMENT.name ASC -LIMIT 50 ----- - -`[.implemented]` GROUP BY, HAVING, aggregate validation. - -=== ORDER BY - -Sort results by one or more fields. - -.Grammar ----- -order_by_clause = 'ORDER' , 'BY' , order_by_list ; -order_by_item = field_ref , [sort_direction] ; -sort_direction = 'ASC' | 'DESC' ; ----- - -Default sort direction is `ASC` when omitted. - -.Example: multi-field sort -[source,vcl] ----- -SELECT DOCUMENT.name, DOCUMENT.severity -FROM FEDERATION /archives/* -ORDER BY DOCUMENT.severity DESC, DOCUMENT.name ASC -LIMIT 100 ----- - -Fields must be comparable types (Int, Float, String, Timestamp). `InvalidOrderByField` error otherwise. - -`[.implemented]` Multi-field ORDER BY with ASC/DESC. - -=== LIMIT / OFFSET - -Paginate results. - -.Grammar ----- -limit_clause = 'LIMIT' , integer ; -offset_clause = 'OFFSET' , integer ; ----- - -.Example: pagination -[source,vcl] ----- -SELECT DOCUMENT -FROM FEDERATION /archives/* -WHERE FULLTEXT CONTAINS "machine learning" -LIMIT 100 -OFFSET 500 ----- - -`[.implemented]` LIMIT and OFFSET. - - -// ============================================================================ -// 6. MUTATION STATEMENTS -// ============================================================================ - -== Mutation Statements - -VCL 2.0 adds a write path with INSERT, UPDATE, and DELETE operations. All mutations optionally accept a PROOF clause for verified writes. - -.Grammar ----- -mutation = insert_mutation | update_mutation | delete_mutation ; ----- - -=== INSERT HEXAD - -Create a new octad with modality data. - -.Grammar ----- -insert_mutation = 'INSERT' , 'HEXAD' , 'WITH' , modality_data_list , [proof_clause] ; -modality_data_list = modality_data , { ',' , modality_data } ; -modality_data = document_data | vector_data | graph_data - | tensor_data | semantic_data | temporal_data ; ----- - -Each modality data type has its own syntax: - -[cols="1,3"] -|=== -| Modality | Syntax - -| `DOCUMENT` -| `DOCUMENT(field = value, ...)` - -| `VECTOR` -| `VECTOR([float, ...])` - -| `GRAPH` -| `GRAPH(edge_type, target_octad_id)` - -| `TENSOR` -| `TENSOR([array_literal])` - -| `SEMANTIC` -| `SEMANTIC(contract_name)` - -| `TEMPORAL` -| `TEMPORAL(timestamp)` -|=== - -.Examples: INSERT -[source,vcl] ----- --- Basic insert with document and vector -INSERT HEXAD WITH - DOCUMENT(title = "New Research Paper", author = "Jane Doe", severity = 3), - VECTOR([0.12, 0.45, 0.78, 0.23, 0.91]) - --- Insert with graph relationship -INSERT HEXAD WITH - DOCUMENT(title = "Follow-Up Study", year = 2026), - VECTOR([0.5, 0.3, 0.2, 0.8]), - GRAPH(CITES, 550e8400-e29b-41d4-a716-446655440000) - --- Verified insert with proof -INSERT HEXAD WITH - DOCUMENT(title = "Clinical Trial Result", status = "verified"), - SEMANTIC(ClinicalTrialContract), - TEMPORAL(2026-02-13T10:30:00Z) -PROOF INTEGRITY(WriteContract) AND PROVENANCE(AuditTrailContract) ----- - -The octad ID is **auto-generated** and returned in the result. Drift detection fires after insert to validate cross-modal consistency. - -`[.implemented]` INSERT parsing and execution via Elixir executor. - -=== UPDATE HEXAD - -Modify fields of an existing octad. - -.Grammar ----- -update_mutation = 'UPDATE' , 'HEXAD' , uuid , 'SET' , set_list , [proof_clause] ; -set_list = set_assignment , { ',' , set_assignment } ; -set_assignment = field_ref , '=' , literal ; ----- - -.Examples: UPDATE -[source,vcl] ----- --- Simple field update -UPDATE HEXAD 550e8400-e29b-41d4-a716-446655440000 -SET DOCUMENT.title = "Corrected Title", DOCUMENT.severity = 5 - --- Verified update with proof -UPDATE HEXAD 550e8400-e29b-41d4-a716-446655440000 -SET DOCUMENT.status = "retracted", - DOCUMENT.retraction_reason = "Data fabrication" -PROOF ACCESS(EditAccessContract) - AND PROVENANCE(RetractionContract) ----- - -`UpdateNotFound` error if the octad UUID does not exist. Drift detection fires after update. - -`[.implemented]` UPDATE parsing and execution. - -=== DELETE HEXAD - -Remove a octad and all its modality data. - -.Grammar ----- -delete_mutation = 'DELETE' , 'HEXAD' , uuid , [proof_clause] ; ----- - -.Examples: DELETE -[source,vcl] ----- --- Simple deletion -DELETE HEXAD 550e8400-e29b-41d4-a716-446655440000 - --- Verified deletion with proof -DELETE HEXAD 550e8400-e29b-41d4-a716-446655440000 -PROOF ACCESS(AdminAccessContract) - AND PROVENANCE(DeletionAuditContract) ----- - -All deletions are logged in verisim-temporal for auditability, even without explicit proof. - -`[.implemented]` DELETE parsing and execution. - - -// ============================================================================ -// 7. FEDERATION QUERIES -// ============================================================================ - -== Federation Queries - -Federation queries distribute execution across multiple VeriSimDB instances. The Elixir orchestration layer handles fan-out, result aggregation, and drift resolution. - -=== Node Patterns - -Two node pattern syntaxes select federation targets: - -.Glob patterns ----- -FROM FEDERATION /universities/* -- all nodes under /universities/ -FROM FEDERATION /research-db/** -- recursive ----- - -.Node lists ----- -FROM FEDERATION [node1, node2, node3] -- explicit list ----- - -=== Drift Policies - -Federation queries MUST decide how to handle inter-node inconsistency. Without an explicit drift policy, the default is `TOLERATE`. - -.Drift policy semantics -[cols="1,3,1"] -|=== -| Policy | Formal Semantics | Consistency Guarantee - -| `STRICT` -| stem:[\forall n \in N : \text{hash}(n) = \text{hash}(n_0)] — all nodes agree -| Full - -| `REPAIR` -| Detect → repair → retry → return consistent results -| Full (eventual) - -| `TOLERATE` -| Return union of results regardless of consistency -| None - -| `LATEST` -| stem:[\forall n \in N : \text{select}(\max(\text{timestamp}(n)))] — most recent wins -| Temporal only -|=== - -.Consistency lattice ----- -STRICT < LATEST < REPAIR < TOLERATE - (increasing tolerance) ----- - -.Small-step drift semantics -[stem] -++++ -\frac{\exists n_i, n_j \in N : \text{hash}(n_i) \neq \text{hash}(n_j) \quad \text{policy} = \text{STRICT}}{\langle N, q \rangle \to \text{Error}(\text{DriftDetected})} \quad \text{(S-DriftStrict)} -++++ - -[stem] -++++ -\frac{\text{divergent}(N) \neq \emptyset \quad \text{policy} = \text{REPAIR} \quad N' = \text{repair}(N)}{\langle N, q \rangle \to \langle N', q \rangle} \quad \text{(S-DriftRepair)} -++++ - -`[.implemented]` Federation fan-out, glob patterns, node lists. Drift detection via DriftMonitor GenServer. + -`[.partial]` Auto-repair strategies (latest_wins implemented, quorum partial). + -`[.planned]` Byzantine fault detection. - - -// ============================================================================ -// 8. PROOF SYSTEM -// ============================================================================ - -== Proof System - -The proof system is the core differentiator between VCL and traditional query languages. It enables cryptographic verification of query results. - -=== Six Proof Types - -[cols="1,3,2"] -|=== -| Type | Contract | Formal Proposition - -| `EXISTENCE` -| Octad exists and is accessible -| stem:[\exists h \in S : h.\text{id} = \text{uuid} \land \text{accessible}(h)] - -| `CITATION` -| Citation chain is valid -| stem:[\forall e \in \text{path} : \text{valid\_edge}(e) \land \text{exists}(\text{target}(e))] - -| `ACCESS` -| User has verified access rights -| stem:[\text{role}(\text{user}) \in \text{allowed\_roles}(\text{resource})] - -| `INTEGRITY` -| Data has not been tampered with -| stem:[\text{hash}(\text{data}) = \text{merkle\_root}(\text{store})] - -| `PROVENANCE` -| Lineage is verifiable -| stem:[\forall t \in \text{chain} : \text{valid\_transform}(t) \land \text{signed}(t.\text{actor})] - -| `CUSTOM` -| User-defined ZKP contract -| stem:[\phi(\text{data}, \text{witness}) = \text{true}] -|=== - -=== Proof Composition - -Multiple proofs are composed with `AND`. Composition is **conjunctive** — all proofs must hold simultaneously: - -[source,vcl] ----- -PROOF ACCESS(InstitutionalAccessContract) - AND PROVENANCE(ClinicalTrialContract) - AND INTEGRITY(DataIntegrityContract) ----- - -This produces a composite proof type: - -[stem] -++++ -\text{Proof}[\text{access} \land \text{provenance} \land \text{integrity}] -++++ - -ZKP witnesses for each proof are generated independently and composed into a single **proof bundle**. The type checker validates that proofs do not conflict (e.g., a contract requiring data deletion with one requiring data preservation would produce `MultiProofConflict`). - -=== ZKP Circuit Definitions - -Each proof type maps to a ZKP circuit via the `proven-library`: - ----- -EXISTENCE → existence_circuit (hash lookup + membership) -CITATION → citation_circuit (graph traversal + edge validity) -ACCESS → access_circuit (role check + credential verification) -INTEGRITY → integrity_circuit (Merkle tree verification, Blake3) -PROVENANCE → provenance_circuit (chain validation + signature check) -CUSTOM → user_circuit (custom circuit from contract definition) ----- - -`[.planned]` ZKP circuit compilation and execution via proven-library / sanctify. - -=== Proof Certificates - -When a VCL-UT query succeeds, the result includes a **proof certificate**: - -[source,json] ----- -{ - "data": [ ... ], // Query results - "proof": { - "type": "INTEGRITY", - "contract": "DataIntegrityContract", - "certificate": "base64-encoded-zkp-certificate", - "verifiable": true, - "generated_at": "2026-02-13T10:30:00Z", - "circuit_id": "integrity_v2" - } -} ----- - -Certificates are independently verifiable — any party with the verification key can confirm the proof without access to the underlying data. - -`[.planned]` Certificate generation and verification. - - -// ============================================================================ -// 9. QUERY EXECUTION MODEL -// ============================================================================ - -== Query Execution Model - -=== Query Routing - -The execution pipeline flows through three layers: - -1. **ReScript Core** — Parsing, type checking, query plan generation -2. **Elixir Orchestrator** — Query routing, federation fan-out, drift handling -3. **Rust Modality Stores** — Per-modality query execution - -.Execution pipeline ----- -VCL String - → Parser (ReScript VCLParser.res / Elixir vcl_bridge.ex fallback) - → AST - → Type Checker (VCLBidir.res) - → Query Plan Generator (VCLExplain.res) - → Condition Classifier (pushdown vs cross-modal) - → Elixir QueryRouter GenServer - → Route to stores by modality: - :text → RustClient.search_text (Tantivy) - :vector → RustClient.search_vector (HNSW) - :graph → RustClient.get_related (Oxigraph) - :semantic → Text search with type: prefix - :temporal → Version lookup via REST API - :multi → Combine text + vector, deduplicate - → Cross-modal evaluation (post-fetch) - → Aggregation / sorting / pagination - → Response ----- - -.Condition classification -The executor separates conditions into two categories: - -* **Pushdown conditions** — Sent directly to modality stores for server-side evaluation (FULLTEXT CONTAINS, FIELD comparisons, graph patterns, vector similarity, tensor operations, temporal conditions, semantic conditions). -* **Cross-modal conditions** — Evaluated in the Elixir layer after all store results are fetched (DRIFT, CONSISTENT, EXISTS/NOT EXISTS, cross-modal field comparisons). - -=== Cross-Modal Joins - -VCL does not have explicit JOIN syntax. Instead, cross-modal querying is implicit — every octad is already the join of all its modalities. When you write: - -[source,vcl] ----- -SELECT GRAPH, VECTOR, DOCUMENT -FROM HEXAD 550e8400-e29b-41d4-a716-446655440000 -WHERE (h)-[:CITES]->(target) - AND h.embedding SIMILAR TO [0.1, 0.2, 0.3] WITHIN 0.8 - AND FULLTEXT CONTAINS "neural networks" ----- - -The executor: - -1. Routes graph condition to Oxigraph -2. Routes vector condition to HNSW -3. Routes document condition to Tantivy -4. Intersects results by octad ID -5. Returns the unified octad with all three modalities populated - -=== EXPLAIN Plans - -Prefix a query with `EXPLAIN` to see the execution plan without executing: - -[source,vcl] ----- -EXPLAIN SELECT GRAPH, VECTOR, DOCUMENT -FROM FEDERATION /research-db/* -WHERE FULLTEXT CONTAINS "AI" - AND h.embedding SIMILAR TO [0.1, 0.2, 0.3] -LIMIT 50 ----- - -.Example EXPLAIN output ----- -VCL Execution Plan -================== -Strategy: Parallel -Steps: - 1. [DOCUMENT] Full-text search: "AI" - Estimated cost: 80ms | Selectivity: 0.3 - Pushed predicates: FULLTEXT CONTAINS "AI" - 2. [VECTOR] Similarity search - Estimated cost: 50ms | Selectivity: 0.1 - Pushed predicates: SIMILAR TO [0.1, 0.2, 0.3] - 3. [GRAPH] Graph retrieval - Estimated cost: 150ms | Selectivity: 1.0 - 4. [INTERSECT] Cross-modal join by octad ID - Estimated cost: 10ms - 5. [LIMIT] Cap results at 50 - Estimated cost: 1ms - -Total estimated cost: 291ms -Optimization hints: - - Consider adding LIMIT to reduce result set - - VECTOR + DOCUMENT queries can execute in parallel ----- - -.Per-modality base costs (from VCLExplain.res) -[cols="1,1"] -|=== -| Modality | Base Cost - -| GRAPH -| 150ms - -| VECTOR -| 50ms - -| TENSOR -| 200ms - -| SEMANTIC -| 300ms - -| DOCUMENT -| 80ms - -| TEMPORAL -| 30ms -|=== - -`[.implemented]` EXPLAIN plan generation with cost estimation and performance hints. - -=== Error Handling and Diagnostics - -VCL provides structured error responses with error codes, messages, and recovery hints. - -==== Error Categories - -[cols="1,2,1"] -|=== -| Category | Error Code Pattern | Recoverable? - -| Parse errors -| `VCL_PARSE_ERROR` -| No - -| Type errors -| `VCL_TYPE_ERROR` -| No - -| Runtime errors -| `VCL_STORE_UNAVAILABLE`, `VCL_QUERY_TIMEOUT`, `VCL_DRIFT_DETECTED`, `VCL_PERMISSION_DENIED`, `VCL_RESOURCE_EXHAUSTED`, `VCL_RUNTIME_ERROR` -| Varies - -| Modality errors -| `VCL_GRAPH_ERROR`, `VCL_VECTOR_ERROR`, `VCL_TENSOR_ERROR`, `VCL_SEMANTIC_ERROR`, `VCL_DOCUMENT_ERROR`, `VCL_TEMPORAL_ERROR` -| No - -| Federation errors -| `VCL_FEDERATION_ERROR` -| `PartialResults` and `RemoteStoreUnreachable` are recoverable -|=== - -==== Error Response Format - -[source,json] ----- -{ - "error_code": "VCL_PARSE_ERROR", - "message": "Parse Error at 3:27-3:27: Expected ']->', found end of input", - "recoverable": false -} ----- - -==== Common Error Conditions - -.Parse-time errors -[cols="1,3"] -|=== -| Error | Cause - -| `UnexpectedToken` -| Parser encountered a token it did not expect - -| `MissingFromClause` -| Query has SELECT but no FROM - -| `HavingWithoutGroupBy` -| HAVING clause used without GROUP BY - -| `AggregateWithoutGroupBy` -| Aggregate function used without GROUP BY -|=== - -.Type-time errors -[cols="1,3"] -|=== -| Error | Cause - -| `VectorDimensionMismatch` -| SIMILAR TO vector has different dimension than stored embeddings - -| `OperatorTypeMismatch` -| Comparison operator used with incompatible types - -| `AggregateTypeMismatch` -| SUM/AVG applied to non-numeric field - -| `CrossModalTypeMismatch` -| Cross-modal field comparison with incompatible types - -| `MultiProofConflict` -| Two proof specifications have conflicting requirements - -| `UnknownField` -| Field reference does not exist for the specified modality -|=== - -.Runtime errors -[cols="1,3"] -|=== -| Error | Cause - -| `StoreUnavailable` -| A modality store is down or unreachable - -| `QueryTimeout` -| Query exceeded the configured timeout - -| `DriftDetected` -| Federation drift detected with STRICT policy - -| `ByzantineFaultDetected` -| Suspicious inconsistencies suggesting Byzantine behaviour -|=== - -`[.implemented]` All error types defined in `VCLError.res` (447 lines). Structured JSON error responses. - - -// ============================================================================ -// 10. IMPLEMENTATION STATUS -// ============================================================================ - -== Implementation Status - -This section provides an honest assessment of what is implemented, partially implemented, and planned as of 2026-02-27. - -=== Implemented - -[cols="2,1,3"] -|=== -| Feature | Version | Notes - -| VCL Parser (ReScript) -| 2.0 -| 1154 lines, monadic parser combinators. Handles all grammar productions. - -| VCL Parser (Elixir fallback) -| 2.0 -| Built into `vcl_bridge.ex` (672 lines). No external Deno/Node dependency needed. - -| Bidirectional Type Checker -| 2.0 -| `VCLBidir.res` (841 lines). 9-phase pipeline. Cross-modal checking, mutation checking. - -| Type System -| 2.0 -| `VCLTypes.res` (296 lines). Pi types, Sigma types, proof types, structural equality. - -| Error System -| 2.0 -| `VCLError.res` (447 lines). 5 error categories, 40+ error kinds, JSON serialization. - -| EXPLAIN Plans -| 2.0 -| `VCLExplain.res` (427 lines). Per-modality cost estimation, performance hints. - -| Query Executor -| 2.0 -| `vcl_executor.ex` (1162 lines). Full pipeline: parse, type check, plan, route, aggregate. - -| Query Router -| 2.0 -| `query_router.ex` (187 lines). GenServer routing by modality type with statistics. - -| SELECT (modalities, projections, aggregates) -| 2.0 -| All three selection modes. - -| FROM (HEXAD, FEDERATION, STORE) -| 2.0 -| All three source types. - -| WHERE (all condition types) -| 2.0 -| Graph, vector, tensor, semantic, document, temporal, cross-modal. - -| GROUP BY / HAVING -| 2.0 -| With aggregate validation. - -| ORDER BY -| 2.0 -| Multi-field, ASC/DESC. - -| LIMIT / OFFSET -| 2.0 -| Pagination. - -| INSERT / UPDATE / DELETE -| 2.0 -| All three mutation types with optional PROOF clause. - -| Drift Detection -| 1.0 -| DriftMonitor GenServer. Merkle tree divergence detection. - -| Modality Stores (Rust) -| 1.0 -| Oxigraph, HNSW, Burn/ndarray, CBOR, Tantivy, verisim-temporal. 510 tests pass. -|=== - -=== Partial - -[cols="2,3"] -|=== -| Feature | Gaps - -| Federation drift auto-repair -| `latest_wins` strategy implemented. `quorum` strategy partial. Byzantine detection planned. - -| Semantic condition runtime -| Parser and type checker handle SATISFIES, HAS PROOF, VERIFIED BY. Runtime contract verification via ZKP is stubbed. - -| Cross-modal write atomicity -| INSERT/UPDATE/DELETE execute but atomic cross-modal writes (all modalities or none) are not guaranteed. - -| VCL Bridge (Elixir ↔ ReScript) -| Port-based communication works. Falls back to built-in Elixir parser when Deno/Node unavailable. -|=== - -=== Planned - -[cols="2,2"] -|=== -| Feature | Timeline - -| VCL-UT runtime (Lean type checker wired to execution) -| Next milestone - -| ZKP proof generation (proven-library / sanctify) -| Next milestone - -| Proof certificates -| After ZKP integration - -| Provenance modality (verisim-provenance) -| Active development - -| Spatial modality (verisim-spatial) -| Backlog - -| Heterogeneous federation (non-VeriSimDB backends) -| Backlog - -| Refinement type syntax (explicit) -| Future - -| Row polymorphism, effect types, linear types for ZKP witnesses -| Research -|=== - - -// ============================================================================ -// APPENDIX A: COMPLETE EBNF GRAMMAR -// ============================================================================ - -[[appendix-a]] -[appendix] -== Complete EBNF Grammar - -This is the normative grammar for VCL 2.0, reproduced from `vcl-grammar.ebnf` (ISO/IEC 14977 EBNF notation). - -[source,ebnf] ----- -(* VeriSim Consonance Language (VCL) Grammar *) -(* Version: 2.0 — Dependent types, cross-modal correlation, write path *) - -(* 1. TOP-LEVEL STRUCTURE *) -statement = query | mutation ; - -query = select_clause, - from_clause, - [where_clause], - [group_by_clause], - [having_clause], - [proof_clause], - [order_by_clause], - [limit_clause], - [offset_clause] ; - -(* 2. SELECT CLAUSE *) -select_clause = 'SELECT', select_item_list ; -select_item_list = select_item, { ',', select_item } ; -select_item = aggregate_expr | field_ref | modality_spec ; - -field_ref = modality_name, '.', identifier ; - -aggregate_expr = count_all | aggregate_field ; -count_all = 'COUNT', '(', '*', ')' ; -aggregate_field = aggregate_func, '(', field_ref, ')' ; -aggregate_func = 'COUNT' | 'SUM' | 'AVG' | 'MIN' | 'MAX' ; - -modality_spec = 'GRAPH', [graph_projection] - | 'VECTOR', [vector_projection] - | 'TENSOR', [tensor_projection] - | 'SEMANTIC', [semantic_projection] - | 'DOCUMENT', [document_projection] - | 'TEMPORAL', [temporal_projection] - | '*' ; - -modality_name = 'GRAPH' | 'VECTOR' | 'TENSOR' - | 'SEMANTIC' | 'DOCUMENT' | 'TEMPORAL' ; - -graph_projection = '(', sparql_pattern, ')' ; -vector_projection = '(', vector_fields, ')' ; -tensor_projection = '(', tensor_slice, ')' ; -semantic_projection = '(', contract_names, ')' ; -document_projection = '(', document_fields, ')' ; -temporal_projection = '(', version_spec, ')' ; - -(* 3. FROM CLAUSE *) -from_clause = 'FROM', source_spec ; -source_spec = octad_source | federation_source | store_source ; - -octad_source = 'HEXAD', uuid ; - -federation_source = 'FEDERATION', node_pattern, [drift_policy] ; -node_pattern = glob_pattern | node_list ; - -drift_policy = 'WITH', 'DRIFT', drift_mode ; -drift_mode = 'STRICT' | 'REPAIR' | 'TOLERATE' | 'LATEST' ; - -store_source = 'STORE', store_id ; - -(* 4. WHERE CLAUSE *) -where_clause = 'WHERE', condition ; - -condition = simple_condition | compound_condition | '(', condition, ')' ; - -simple_condition = graph_condition | vector_condition | tensor_condition - | semantic_condition | document_condition - | temporal_condition | cross_modal_condition ; - -cross_modal_condition = cross_modal_field_compare | drift_condition - | consistency_condition | exists_condition - | not_exists_condition ; - -cross_modal_field_compare = field_ref, comparison_op, field_ref ; - -drift_condition = 'DRIFT', '(', modality_name, ',', modality_name, ')', - comparison_op, float ; - -consistency_condition = 'CONSISTENT', '(', modality_name, ',', modality_name, ')', - 'USING', metric_name ; -metric_name = 'COSINE' | 'EUCLIDEAN' | 'DOT_PRODUCT' | 'JACCARD' ; - -exists_condition = modality_name, 'EXISTS' ; -not_exists_condition = modality_name, 'NOT', 'EXISTS' ; - -compound_condition = condition, 'AND', condition - | condition, 'OR', condition - | 'NOT', condition ; - -(* 4.1. Graph Conditions *) -graph_condition = sparql_pattern | path_pattern ; -sparql_pattern = '(', node_var, ')', edge_pattern, '(', node_var, ')' ; -edge_pattern = '-[', edge_type, ']->' - | '-[', edge_type, ']-' - | '<-[', edge_type, ']-' ; -node_var = identifier | ('?', identifier) ; -edge_type = ':', identifier ; - -path_pattern = node_var, path_quantifier, node_var ; -path_quantifier = '-[', edge_type, - ('*' | '+' | '{', integer, ',', integer, '}'), - ']->' ; - -(* 4.2. Vector Conditions *) -vector_condition = vector_field, 'SIMILAR', 'TO', vector_literal, - [similarity_threshold] - | vector_field, 'NEAREST', integer, [metric_type] ; - -vector_field = identifier, '.', 'embedding' ; -vector_literal = '[', float, { ',', float }, ']' ; -similarity_threshold = 'WITHIN', float ; -metric_type = 'USING', ('COSINE' | 'EUCLIDEAN' | 'DOT_PRODUCT') ; - -(* 4.3. Tensor Conditions *) -tensor_condition = tensor_field, tensor_op, tensor_literal ; -tensor_op = '==' | '>' | '<' | '>=' | '<=' | 'SHAPE' | 'RANK' ; -tensor_literal = array_literal | scalar_literal ; - -(* 4.4. Semantic Conditions *) -semantic_condition = 'SATISFIES', contract_name, [contract_params] - | 'HAS', 'PROOF', proof_type - | 'VERIFIED', 'BY', verifier_id ; - -contract_name = identifier ; -contract_params = '(', param_list, ')' ; -param_list = identifier, '=', literal, - { ',', identifier, '=', literal } ; - -(* 4.5. Document Conditions *) -document_condition = 'FULLTEXT', 'CONTAINS', string_literal - | 'FULLTEXT', 'MATCHES', regex_literal - | 'FIELD', identifier, comparison_op, literal ; - -comparison_op = '==' | '!=' | '>' | '<' | '>=' | '<=' | 'LIKE' ; - -(* 4.6. Temporal Conditions *) -temporal_condition = 'AS', 'OF', timestamp - | 'BETWEEN', timestamp, 'AND', timestamp - | 'VERSION', version_id - | 'MODIFIED', 'BY', actor_id ; - -(* 5. PROOF CLAUSE *) -proof_clause = 'PROOF', proof_spec_list ; -proof_spec_list = proof_spec, { 'AND', proof_spec } ; -proof_spec = proof_type, '(', contract_name, ')', [proof_params] ; -proof_type = 'EXISTENCE' | 'CITATION' | 'ACCESS' - | 'INTEGRITY' | 'PROVENANCE' | 'CUSTOM' ; -proof_params = 'WITH', param_list ; - -(* 6. GROUP BY / HAVING *) -group_by_clause = 'GROUP', 'BY', field_ref_list ; -field_ref_list = field_ref, { ',', field_ref } ; -having_clause = 'HAVING', condition ; - -(* 7. ORDER BY *) -order_by_clause = 'ORDER', 'BY', order_by_list ; -order_by_list = order_by_item, { ',', order_by_item } ; -order_by_item = field_ref, [sort_direction] ; -sort_direction = 'ASC' | 'DESC' ; - -(* 8. PAGINATION *) -limit_clause = 'LIMIT', integer ; -offset_clause = 'OFFSET', integer ; - -(* 9. MUTATIONS *) -mutation = insert_mutation | update_mutation | delete_mutation ; - -insert_mutation = 'INSERT', 'HEXAD', 'WITH', - modality_data_list, [proof_clause] ; -modality_data_list = modality_data, { ',', modality_data } ; -modality_data = document_data | vector_data | graph_data - | tensor_data | semantic_data | temporal_data ; - -document_data = 'DOCUMENT', '(', field_assignment_list, ')' ; -vector_data = 'VECTOR', '(', vector_literal, ')' ; -graph_data = 'GRAPH', '(', identifier, ',', identifier, ')' ; -tensor_data = 'TENSOR', '(', array_literal, ')' ; -semantic_data = 'SEMANTIC', '(', identifier, ')' ; -temporal_data = 'TEMPORAL', '(', timestamp, ')' ; - -field_assignment_list = field_assignment, { ',', field_assignment } ; -field_assignment = identifier, '=', literal ; - -update_mutation = 'UPDATE', 'HEXAD', uuid, 'SET', - set_list, [proof_clause] ; -set_list = set_assignment, { ',', set_assignment } ; -set_assignment = field_ref, '=', literal ; - -delete_mutation = 'DELETE', 'HEXAD', uuid, [proof_clause] ; - -(* 10. LEXICAL ELEMENTS *) -uuid = 8*hex_digit, '-', 4*hex_digit, '-', 4*hex_digit, '-', - 4*hex_digit, '-', 12*hex_digit ; -hex_digit = '0'..'9' | 'a'..'f' | 'A'..'F' ; - -identifier = (letter | '_'), { letter | digit | '_' } ; -store_id = identifier ; -verifier_id = identifier ; -actor_id = identifier ; -version_id = identifier ; - -integer = digit, { digit } ; -float = digit, { digit }, '.', digit, { digit }, - [('e' | 'E'), ['+' | '-'], digit, { digit }] ; - -string_literal = "'", { char - ("'" | '\') | "\'" | "\\" }, "'" ; -regex_literal = '/', { char - ('/' | '\') | '\/' | '\\' }, '/' ; - -timestamp = iso8601_datetime ; - -glob_pattern = '/', path_segment, { '/', path_segment }, - ('/*' | '/**') ; -path_segment = (alphanum | '_' | '-'), - { alphanum | '_' | '-' } ; - -array_literal = '[', literal, { ',', literal }, ']' ; -scalar_literal = integer | float | string_literal | boolean ; -boolean = 'true' | 'false' ; -literal = scalar_literal | array_literal ; - -(* 11. COMMENTS *) -comment = '--', { char - newline }, newline - | '/*', { any }, '*/' ; - -(* 12. RESERVED KEYWORDS *) -keywords = 'SELECT' | 'FROM' | 'WHERE' | 'PROOF' | 'LIMIT' | 'OFFSET' - | 'GRAPH' | 'VECTOR' | 'TENSOR' | 'SEMANTIC' | 'DOCUMENT' - | 'TEMPORAL' | 'HEXAD' | 'FEDERATION' | 'STORE' - | 'WITH' | 'DRIFT' | 'STRICT' | 'REPAIR' | 'TOLERATE' - | 'LATEST' | 'AND' | 'OR' | 'NOT' - | 'SIMILAR' | 'TO' | 'WITHIN' | 'NEAREST' | 'USING' - | 'SATISFIES' | 'HAS' | 'VERIFIED' | 'BY' - | 'FULLTEXT' | 'CONTAINS' | 'MATCHES' | 'FIELD' | 'LIKE' - | 'AS' | 'OF' | 'BETWEEN' | 'VERSION' | 'MODIFIED' - | 'EXISTENCE' | 'CITATION' | 'ACCESS' | 'INTEGRITY' - | 'PROVENANCE' | 'CUSTOM' - | 'COSINE' | 'EUCLIDEAN' | 'DOT_PRODUCT' | 'JACCARD' - | 'SHAPE' | 'RANK' - | 'ORDER' | 'GROUP' | 'HAVING' | 'ASC' | 'DESC' - | 'COUNT' | 'SUM' | 'AVG' | 'MIN' | 'MAX' - | 'INSERT' | 'UPDATE' | 'DELETE' | 'SET' - | 'EXISTS' | 'CONSISTENT' - | 'true' | 'false' ; ----- - - -// ============================================================================ -// APPENDIX B: RESERVED KEYWORDS -// ============================================================================ - -[appendix] -== Reserved Keywords - -Complete alphabetical list of all 70 reserved keywords. Keywords are case-insensitive in VCL source. - -[cols="4*"] -|=== - -| `ACCESS` -| `AND` -| `AS` -| `ASC` - -| `AVG` -| `BETWEEN` -| `BY` -| `CITATION` - -| `CONSISTENT` -| `CONTAINS` -| `COSINE` -| `COUNT` - -| `CUSTOM` -| `DELETE` -| `DESC` -| `DOCUMENT` - -| `DOT_PRODUCT` -| `DRIFT` -| `EUCLIDEAN` -| `EXISTENCE` - -| `EXISTS` -| `FEDERATION` -| `FIELD` -| `FROM` - -| `FULLTEXT` -| `GRAPH` -| `GROUP` -| `HAS` - -| `HAVING` -| `HEXAD` -| `INSERT` -| `INTEGRITY` - -| `JACCARD` -| `LATEST` -| `LIKE` -| `LIMIT` - -| `MATCHES` -| `MAX` -| `MIN` -| `MODIFIED` - -| `NEAREST` -| `NOT` -| `OF` -| `OFFSET` - -| `OR` -| `ORDER` -| `PROOF` -| `PROVENANCE` - -| `RANK` -| `REPAIR` -| `SELECT` -| `SEMANTIC` - -| `SET` -| `SHAPE` -| `SIMILAR` -| `STORE` - -| `STRICT` -| `SUM` -| `TEMPORAL` -| `TENSOR` - -| `TO` -| `TOLERATE` -| `UPDATE` -| `USING` - -| `VECTOR` -| `VERIFIED` -| `VERSION` -| `WHERE` - -| `WITH` -| `WITHIN` -| `false` -| `true` -|=== - - -// ============================================================================ -// APPENDIX C: ERROR CODES -// ============================================================================ - -[appendix] -== Error Codes - -=== Top-Level Error Codes - -[cols="2,3"] -|=== -| Code | Description - -| `VCL_PARSE_ERROR` -| Syntax error during parsing - -| `VCL_TYPE_ERROR` -| Type system violation - -| `VCL_STORE_UNAVAILABLE` -| A modality store is unreachable - -| `VCL_QUERY_TIMEOUT` -| Query exceeded time limit - -| `VCL_DRIFT_DETECTED` -| Cross-node inconsistency with STRICT policy - -| `VCL_PERMISSION_DENIED` -| Insufficient access rights - -| `VCL_RESOURCE_EXHAUSTED` -| System resource limit reached - -| `VCL_RUNTIME_ERROR` -| General runtime error - -| `VCL_GRAPH_ERROR` -| Graph-specific error (Oxigraph) - -| `VCL_VECTOR_ERROR` -| Vector-specific error (HNSW) - -| `VCL_TENSOR_ERROR` -| Tensor-specific error (Burn/ndarray) - -| `VCL_SEMANTIC_ERROR` -| Semantic-specific error (ZKP/CBOR) - -| `VCL_DOCUMENT_ERROR` -| Document-specific error (Tantivy) - -| `VCL_TEMPORAL_ERROR` -| Temporal-specific error (Merkle trees) - -| `VCL_FEDERATION_ERROR` -| Federation-specific error - -| `VCL_MULTIPLE_ERRORS` -| Multiple errors occurred -|=== - -=== Parse Error Kinds - -[cols="2,3"] -|=== -| Kind | Description - -| `UnexpectedToken` -| Parser expected specific tokens but found something else - -| `UnterminatedString` -| String literal missing closing quote - -| `InvalidNumber` -| Malformed numeric literal - -| `InvalidModality` -| Unknown modality name (valid: GRAPH, VECTOR, TENSOR, SEMANTIC, DOCUMENT, TEMPORAL) - -| `InvalidDriftPolicy` -| Unknown drift policy (valid: STRICT, REPAIR, TOLERATE, LATEST) - -| `InvalidProofType` -| Unknown proof type (valid: EXISTENCE, CITATION, ACCESS, INTEGRITY, PROVENANCE, CUSTOM) - -| `MissingFromClause` -| Query has SELECT but no FROM - -| `MissingSelectClause` -| Query has FROM but no SELECT - -| `InvalidGraphPattern` -| Malformed SPARQL-like graph pattern - -| `InvalidVectorExpression` -| Malformed vector similarity or KNN expression - -| `InvalidSemanticContract` -| Malformed SATISFIES contract expression - -| `InvalidAggregateExpression` -| Malformed aggregate function call - -| `InvalidOrderByField` -| ORDER BY field not in MODALITY.field format - -| `InvalidGroupByField` -| GROUP BY field not in MODALITY.field format - -| `HavingWithoutGroupBy` -| HAVING clause used without GROUP BY - -| `AggregateWithoutGroupBy` -| Aggregate function used without GROUP BY -|=== - -=== Type Error Kinds - -[cols="2,3"] -|=== -| Kind | Description - -| `ContractNotFound` -| Referenced contract does not exist in registry - -| `ContractViolation` -| Data violates contract constraints - -| `ProofGenerationFailed` -| ZKP proof could not be generated - -| `ProofVerificationFailed` -| ZKP proof did not verify - -| `TypeMismatch` -| Expression type does not match expected type - -| `MissingTypeAnnotation` -| Field or expression requires a type annotation that was not provided - -| `CircularDependency` -| Circular dependency detected in contract or type definitions - -| `SubtypingFailed` -| Value type is not a subtype of expected type - -| `FieldTypeMismatch` -| Field has wrong type for operation - -| `OperatorTypeMismatch` -| Comparison operator used with incompatible types - -| `VectorDimensionMismatch` -| Vector dimension does not match (e.g., SIMILAR TO with wrong dimension) - -| `ProofObligationFailed` -| Proof obligation could not be discharged - -| `AggregateTypeMismatch` -| Aggregate function applied to incompatible type - -| `MultiProofConflict` -| Multiple proof specifications conflict - -| `UnknownField` -| Field does not exist for specified modality - -| `CrossModalTypeMismatch` -| Cross-modal field comparison with incompatible types - -| `DriftRequiresNumeric` -| DRIFT() condition on non-numeric modalities - -| `ConsistencyMetricInvalid` -| CONSISTENT() metric not applicable to modality pair - -| `InsertConflict` -| INSERT octad already exists - -| `UpdateNotFound` -| UPDATE octad does not exist - -| `DeleteNotFound` -| DELETE octad does not exist - -| `ConstraintViolation` -| Write violates field constraints - -| `WriteProofFailed` -| Write proof obligation failed - -| `ReadOnlyStore` -| Attempted write to read-only store -|=== - -=== Runtime Error Kinds - -[cols="2,3"] -|=== -| Kind | Description - -| `StoreUnavailable` -| A modality store is down or unreachable - -| `QueryTimeout` -| Query exceeded the configured time limit - -| `DriftDetected` -| Cross-node inconsistency detected (with STRICT drift policy) - -| `PermissionDenied` -| Insufficient access rights for the requested resource - -| `ResourceExhausted` -| System resource limit reached (memory, connections, etc.) - -| `InvalidOctadId` -| Octad UUID is malformed or syntactically invalid - -| `NetworkError` -| Network error connecting to a store or federation endpoint - -| `InternalError` -| Unexpected internal error -|=== - -=== Modality-Specific Error Kinds - -==== Graph Errors - -[cols="2,3"] -|=== -| Kind | Description - -| `MalformedRDF` -| Invalid RDF syntax - -| `InvalidTriplePattern` -| Invalid triple pattern in WHERE clause - -| `CycleDetected` -| Graph cycle detected during traversal - -| `PredicateNotFound` -| Edge type does not exist - -| `TraversalDepthExceeded` -| Path quantifier exceeded maximum depth -|=== - -==== Vector Errors - -[cols="2,3"] -|=== -| Kind | Description - -| `DimensionMismatch` -| Query vector dimension differs from stored embeddings - -| `InvalidDistanceMetric` -| Unknown or unsupported distance metric - -| `EmbeddingNotFound` -| Requested embedding does not exist - -| `ANNIndexUnavailable` -| Approximate nearest neighbor index not ready -|=== - -==== Tensor Errors - -[cols="2,3"] -|=== -| Kind | Description - -| `ShapeMismatch` -| Tensor shape does not match expected shape - -| `NumericOverflow` -| Numeric computation overflow - -| `InvalidOperation` -| Unsupported tensor operation - -| `UnsupportedDtype` -| Unknown data type for tensor elements -|=== - -==== Semantic Errors - -[cols="2,3"] -|=== -| Kind | Description - -| `InvalidContract` -| Contract definition is malformed - -| `ZKPVerificationFailed` -| Zero-knowledge proof did not verify - -| `WitnessGenerationFailed` -| Could not generate ZKP witness - -| `ContractExpired` -| Contract has expired and is no longer valid -|=== - -==== Document Errors - -[cols="2,3"] -|=== -| Kind | Description - -| `InvalidFullTextQuery` -| Full-text query syntax is malformed - -| `UnsupportedLanguage` -| Text language not supported by analyser - -| `IndexCorrupted` -| Tantivy index is corrupted -|=== - -==== Temporal Errors - -[cols="2,3"] -|=== -| Kind | Description - -| `InvalidTimestamp` -| Timestamp is not valid ISO 8601 - -| `VersionNotFound` -| Requested version does not exist - -| `MerkleVerificationFailed` -| Merkle tree integrity check failed - -| `TemporalConflict` -| Concurrent modification conflict -|=== - -=== Federation Error Kinds - -[cols="2,3"] -|=== -| Kind | Description - -| `RemoteStoreUnreachable` -| Cannot connect to remote federation node (recoverable) - -| `PartialResults` -| Some federation nodes responded, others failed (recoverable) - -| `CrossOrgAccessDenied` -| Cross-organisation access policy violation - -| `ByzantineFaultDetected` -| Suspicious inconsistencies suggesting Byzantine behaviour - -| `ConsensusTimeout` -| Federation consensus protocol timed out - -| `FederationPolicyViolation` -| Query violates federation-level policy -|=== - - -// ============================================================================ -// APPENDIX D: RELATED SPECIFICATIONS -// ============================================================================ - -[appendix] -== Related Specifications - -This specification is the authoritative reference for VCL language behaviour. The following documents provide deeper coverage of specific topics: - -[cols="2,3"] -|=== -| Document | Scope - -| link:vcl-grammar.ebnf[`vcl-grammar.ebnf`] -| Standalone EBNF grammar file (ISO/IEC 14977). Normative source reproduced in <>. - -| link:vcl-type-system.adoc[`vcl-type-system.adoc`] -| Formal type system with LaTeX notation: type syntax, typing rules, type checking algorithm, subtyping, type equivalence, type safety proofs. 923 lines. - -| link:vcl-formal-semantics.adoc[`vcl-formal-semantics.adoc`] -| Operational semantics (big-step and small-step), denotational semantics, proof obligations, soundness theorems, drift semantics. 655 lines. - -| link:vcl-examples.adoc[`vcl-examples.adoc`] -| 63 comprehensive examples covering all VCL constructs, both execution paths, error cases, and advanced use cases. 1078 lines. - -| link:vcl-architecture.adoc[`vcl-architecture.adoc`] -| Dual-path routing architecture, ReScript/Elixir/Rust stack, metadata storage, audit trail, drift detection, API endpoint, performance comparison. 708 lines. - -| link:vcl-vs-vcl-dt.adoc[`vcl-vs-vcl-dt.adoc`] -| Detailed comparison of Slipstream vs VCL-UT execution paths, six proof types, performance implications, honest implementation status. 282 lines. - -| link:vcl-vs-sql.adoc[`vcl-vs-sql.adoc`] -| Comparative guide between VCL and SQL for developers familiar with relational databases. 241 lines. Note: pre-dates v2.0 GROUP BY/ORDER BY/mutation additions. -|=== - -=== Implementation Files - -[cols="2,3"] -|=== -| File | Role - -| `src/vcl/VCLParser.res` -| ReScript parser (1154 lines). Monadic combinator library + grammar. - -| `src/vcl/VCLTypes.res` -| Core type definitions (296 lines). Modality, primitive, VCL, and proof types. - -| `src/vcl/VCLBidir.res` -| Bidirectional type inference engine (841 lines). 9-phase synthesis pipeline. - -| `src/vcl/VCLError.res` -| Structured error types (447 lines). 5 categories, 40+ error kinds. - -| `src/vcl/VCLExplain.res` -| EXPLAIN plan generation (427 lines). Cost estimation and hints. - -| `elixir-orchestration/lib/verisim/query/query_router.ex` -| Elixir query router GenServer (187 lines). Routes by modality. - -| `elixir-orchestration/lib/verisim/query/vcl_bridge.ex` -| Elixir ↔ ReScript bridge (672 lines). Built-in fallback parser. - -| `elixir-orchestration/lib/verisim/query/vcl_executor.ex` -| Full VCL executor (1162 lines). Parse → type check → plan → route → aggregate. -|=== - -=== Known Inconsistencies in Existing Documents - -For transparency, these inconsistencies exist between the pre-v2.0 documents: - -1. **`vcl-vs-sql.adoc`** states VCL lacks GROUP BY, ORDER BY, aggregates, and mutations. This was true for v1.0 but is incorrect for v2.0. The grammar, parser, and executor all support these features. - -2. **`vcl-vs-vcl-dt.adoc`** lists six proof types (EXISTENCE, INTEGRITY, CONSISTENCY, PROVENANCE, FRESHNESS, AUTHORIZATION) which differ from the grammar's six (EXISTENCE, CITATION, ACCESS, INTEGRITY, PROVENANCE, CUSTOM). This specification follows the grammar and implementation. - -3. **`vcl-vs-sql.adoc`** describes VCL as "read-only" which is no longer accurate after v2.0 mutations. - -This specification resolves all three inconsistencies in favour of the implemented grammar and source code. - -// ============================================================================ -// APPENDIX E: VCL-UT EXTENSIONS (TYPELL INTEGRATION) -// ============================================================================ - -[appendix] -== VCL-UT Extensions (Typell Integration) - -VCL-UT extends VCL v3.0 with six optional clauses that provide maximal type-theoretic strictness. These extensions are consumed by the **Typell verification kernel** (the formal verification substrate for PanLL) and are specified normatively in the grammar delta file: - -* link:../../typeql-experimental/docs/vcl-dtpp-grammar.ebnf[`vcl-dtpp-grammar.ebnf`] — EBNF grammar delta (199 lines) - -All six clauses are optional and composable in any combination. They append after the standard query structure with no keyword conflicts against VCL v3.0's 60+ reserved keywords. - -=== VCL-UT vs VCL-UT Feature Comparison - -[cols="2,1,1,2"] -|=== -| Feature | VCL-UT | VCL-UT | Typell Kernel Mapping - -| Dependent types (Pi, Sigma) -| ✅ -| ✅ -| Bidirectional type checker - -| Proof obligations (EXISTENCE, INTEGRITY, etc.) -| ✅ -| ✅ -| Proof engine (auto-gen + Echidna dispatch) - -| ZKP witness generation -| ✅ -| ✅ -| Proof certificates - -| **Linear types** (resource counting) -| ❌ -| ✅ `CONSUME AFTER n USE` -| Linear resource tracker - -| **Session types** (protocol state machines) -| ❌ -| ✅ `WITH SESSION protocol` -| Session protocol manager - -| **Effect systems** (side-effect tracking) -| Partial -| ✅ `EFFECTS { Read, Write, ... }` -| Compositional effect inference - -| **Modal types** (world-indexed scoping) -| ❌ -| ✅ `IN TRANSACTION state` -| Modal scope checker - -| **Proof-carrying code** (theorem attachment) -| Partial (pre-conditions) -| ✅ `PROOF ATTACHED theorem` -| Cryptographic PCC certificates - -| **QTT** (bounded resource quantities) -| ❌ -| ✅ `USAGE LIMIT n` -| Quantitative type theory tracker -|=== - -=== Individual Feature Syntax Examples - -Each VCL-UT clause is shown in isolation to demonstrate independent usage. - -==== Linear Types — CONSUME AFTER - -Resources used a bounded number of times. Prevents duplicate reads and data leaks. - -[source,sql] ----- --- Single-use query result (exactly-once consumption) -SELECT GRAPH.*, DOCUMENT.* FROM HEXAD 'entity-001' - PROOF EXISTENCE(entity-001) - CONSUME AFTER 1 USE; - --- Federation result consumed at most 3 times (e.g., retry budget) -SELECT * FROM FEDERATION '/universities/*' - WITH DRIFT STRICT - CONSUME AFTER 3 USE; ----- - -Idris2 ABI mapping: `(1 conn : Connection)` for single-use, `(n conn : BoundedConn n)` for multi-use. - -==== Session Types — WITH SESSION - -Protocol safety. Connections follow a state machine: Fresh → Authenticated → InTransaction → Committed → Closed. - -[source,sql] ----- --- Read-only session: can query but not mutate -SELECT GRAPH FROM HEXAD 'entity-001' - WITH SESSION ReadOnlyProtocol; - --- Mutation session: allows INSERT/UPDATE/DELETE -INSERT HEXAD WITH DOCUMENT(title = 'New Entry') - WITH SESSION MutationProtocol; - --- Streaming session: maintains cursor across result batches -SELECT * FROM FEDERATION '/sensors/*' - WITH SESSION StreamProtocol; ----- - -Built-in protocols: `ReadOnlyProtocol`, `MutationProtocol`, `StreamProtocol`, `BatchProtocol`. Custom protocols extensible via schema. - -==== Effect Systems — EFFECTS - -Explicit declaration of side effects. The type checker verifies actual operations are a subset of declared effects. - -[source,sql] ----- --- Pure read: no writes, no audit trail -SELECT GRAPH FROM HEXAD 'entity-001' - EFFECTS { Read }; - --- Mutation with audit: declares write and audit effects -INSERT HEXAD WITH DOCUMENT(title = 'Audited Entry') - PROOF INTEGRITY(schema-v2) - EFFECTS { Read, Write, Audit }; - --- Federation with transformation: cross-node data reshaping -SELECT VECTOR FROM FEDERATION '/cluster/*' - EFFECTS { Read, Federate, Transform }; ----- - -Available effects: `Read`, `Write`, `Cite`, `Audit`, `Transform`, `Federate` (extensible via `identifier`). - -==== Modal Types — IN TRANSACTION - -Data scoping via world-indexed types. Data in one scope cannot leak to another without explicit marshalling. - -[source,sql] ----- --- Only visible within committed transactions -SELECT GRAPH FROM HEXAD 'entity-001' - IN TRANSACTION Committed; - --- Read snapshot isolation: sees consistent view at query time -SELECT * FROM FEDERATION '/analytics/*' - IN TRANSACTION ReadSnapshot; - --- Active transaction: data may change before commit -UPDATE HEXAD 'entity-001' SET DOCUMENT.status = 'reviewed' - IN TRANSACTION Active; ----- - -Transaction states: `Fresh`, `Active`, `Committed`, `RolledBack`, `ReadSnapshot` (extensible). - -==== Proof-Carrying Code — PROOF ATTACHED - -Attaches post-condition theorems to query results as sigma pairs. Different from `PROOF` (which verifies pre-conditions). - -[source,sql] ----- --- Attach integrity theorem to result -SELECT GRAPH FROM HEXAD 'entity-001' - PROOF EXISTENCE(entity-001) - PROOF ATTACHED IntegrityTheorem; - --- Attach freshness guarantee to federation result -SELECT * FROM FEDERATION '/realtime/*' - WITH DRIFT STRICT - PROOF ATTACHED FreshnessGuarantee; - --- Attach custom theorem with parameters -SELECT DOCUMENT FROM HEXAD 'entity-001' - PROOF ATTACHED CrossModalConsistency(tolerance = 0.01); ----- - -Idris2 ABI mapping: `ProvedResult : (result : QueryResult) -> (prf : Theorem) -> Type`. - -==== Quantitative Type Theory — USAGE LIMIT - -Bounded resource consumption across the query plan (connections, API calls, memory). Different from `LIMIT` (which caps result rows). - -[source,sql] ----- --- Cap at 100 resource operations (connections, store reads, etc.) -SELECT GRAPH FROM HEXAD 'entity-001' - USAGE LIMIT 100; - --- Federation with bounded resource budget -SELECT * FROM FEDERATION '/global/*' - WITH DRIFT TOLERATE - USAGE LIMIT 1000; ----- - -Idris2 ABI mapping: `BoundedResource : (n : Nat) -> Type`. Generalises linear types from exact-1 to at-most-n. - -==== Combined: Maximal Strictness - -All six clauses composed together (see also link:../../typell/docs/design/DESIGN-2026-03-01-typell-vision.md[Typell Vision Document]): - -[source,sql] ----- -SELECT GRAPH.*, DOCUMENT.* FROM HEXAD 'entity-001' - PROOF EXISTENCE(entity-001) - CONSUME AFTER 1 USE - WITH SESSION ReadOnlyProtocol - EFFECTS { Read } - IN TRANSACTION Committed - PROOF ATTACHED IntegrityTheorem - USAGE LIMIT 100; ----- - -=== Typell Integration - -VCL-UT queries are verified by the Typell kernel via JSON-RPC: - -[source] ----- -typell.check(query, "vcl-ut") → TypeResult ----- - -The kernel returns: types, proof obligations, linear tracking, session protocol compliance, effect analysis, modal scope verification, and proof certificates. See the link:../../typell/docs/design/DESIGN-2026-03-01-typell-vision.md[Typell Vision Document] for the full verification protocol specification. - -=== Status - -`[.planned]` All six VCL-UT clauses. Grammar formally specified. Implementation pending Typell kernel (Phase 3+). diff --git a/verisimdb/docs/adoption-strategy.adoc b/verisimdb/docs/adoption-strategy.adoc deleted file mode 100644 index a484fde7..00000000 --- a/verisimdb/docs/adoption-strategy.adoc +++ /dev/null @@ -1,405 +0,0 @@ -// SPDX-License-Identifier: CC-BY-SA-4.0 -= VeriSimDB External Adoption Strategy -Jonathan D.A. Jewell -:toc: left -:toclevels: 3 -:icons: font - -== Executive Summary - -VeriSimDB's strongest market position is as a *data quality platform* that -detects and repairs cross-system entity drift -- not as a database competitor. -The lead pitch is: - -[quote] -Your data drifts silently across 8 systems. We detect it continuously and -fix it automatically. Here's a demo: 1000 entities, 50 corrupted, 100% -detection and repair rate. - -This document identifies seven target domains for external adoption, ordered -by alignment strength, and provides a concrete outreach plan for each. - -== Positioning - -=== What VeriSimDB Is - -* A *cross-system entity consistency engine* -- sits alongside your existing - databases, not replacing them -* *Drift detection + auto-repair* across 8 modality representations -* *Heterogeneous federation* -- unified queries across PostgreSQL, ArangoDB, - Elasticsearch, and VeriSimDB instances -* *Formally verifiable queries* (VCL-UT) with proof certificates - -=== What VeriSimDB Is Not - -* Not a replacement for PostgreSQL, MongoDB, or Elasticsearch -* Not an ETL pipeline (it monitors consistency, not transforms data) -* Not a data warehouse or analytics platform - -=== Competitive Landscape - -[cols="1,3,3"] -|=== -| Tool | What It Does | How VeriSimDB Differs - -| *Great Expectations* | Post-hoc data validation (schema, distribution) -| VeriSimDB detects drift *continuously* across *cross-system* representations, - not just single-table validation - -| *Monte Carlo* | Data observability (freshness, volume, schema changes) -| VeriSimDB measures *semantic* drift (embedding vs. text, graph vs. document), - not just statistical drift - -| *Evidently AI* | ML model + data drift detection for tabular/embedding data -| VeriSimDB detects drift across *8 modality types simultaneously* and - *automatically repairs* it, not just reports - -| *Apache Atlas* | Metadata governance and lineage tracking -| VeriSimDB tracks lineage *and* enforces consistency. Atlas catalogues but - doesn't detect cross-modal drift - -| *DataHub* | Data catalog with multi-backend search -| VeriSimDB federates *queries* (not just metadata) across heterogeneous - backends with drift-aware result aggregation -|=== - -== Target Domains - -=== 1. GraphRAG / Hybrid Retrieval Systems - -*Alignment: Strongest* - -==== The Problem - -Production GraphRAG architectures combine Neo4j (graph traversal) with -Qdrant/Weaviate (vector similarity) and document stores. They must maintain -consistency between three representations of the same entity: source -document, embedding, and graph node. When source documents change, embeddings -go stale. LangChain explicitly acknowledges that syncing data sources to -vector stores requires dedicated indexing APIs with cleanup processes. - -==== Why VeriSimDB - -This is VeriSimDB's exact architecture. A single entity exists as a document, -a vector embedding, and a graph node. Drift detection catches when an -embedding no longer corresponds to its source text -- the precise failure -mode that causes RAG hallucinations. Self-normalisation triggers re-embedding -automatically. Federation queries traverse Neo4j relationships AND retrieve -vector store results transparently. - -==== Features Used - -* Drift detection (embedding vs. source text) -* Multi-modal storage (document + vector + graph + semantic) -* Heterogeneous federation -* Self-normalisation - -==== Outreach Targets - -* LangChain community (Python/JS -- integration guide needed) -* LlamaIndex community (similar architecture pattern) -* Qdrant users running hybrid Neo4j+Qdrant setups -* Teams building production RAG pipelines who've hit stale embedding bugs - -==== Getting Started Hook - -[source,bash] ----- -# Demo: Create entity, change document, detect embedding drift -curl -X POST http://localhost:8080/api/v1/octads -d '{"title":"AI Safety","body":"...",...}' -# ... modify document without updating embedding ... -curl http://localhost:8080/api/v1/drift/entity/$ID -# Result: vector drift = 0.87 (embedding no longer matches text) ----- - ---- - -=== 2. Biomedical Multi-Omics Data Integration - -*Alignment: Very Strong* - -==== The Problem - -Integrating genomic, proteomic, transcriptomic, and metabolomic data is one -of the hardest data consistency problems in science. The same gene may appear -in a genomics database (sequence), proteomics database (protein structure), -literature database (text), and pathway database (graph node). Cross-level -consistency -- verifying that ChIP-Seq data aligns with mRNA expression -- -requires specialised platforms. There is no ground truth to calibrate against. - -==== Why VeriSimDB - -A gene entity uses 6 of 8 modalities: graph (pathway relationships), vector -(sequence embeddings), semantic (Gene Ontology terms), document (literature), -tensor (expression matrices), provenance (lab, sequencing run, pipeline -version). Drift detection flags when a gene's pathway position contradicts -its expression profile. VCL-UT provides reproducibility guarantees. Provenance -tracking satisfies FDA/EMA regulatory compliance. - -==== Features Used - -* All 8 modalities -* Drift detection (cross-omics consistency) -* Provenance tracking (regulatory compliance) -* VCL-UT (reproducible queries) - -==== Outreach Targets - -* Bioinformatics labs doing multi-omics integration -* ELIXIR infrastructure network (European bioinformatics) -* NIH-funded data commons projects -* Precision medicine platforms - ---- - -=== 3. Healthcare FHIR Interoperability - -*Alignment: High (Safety-Critical)* - -==== The Problem - -Despite the FHIR standard, healthcare data consistency is severely broken. -Vendor-specific interpretations create translation gaps. The same patient -exists in an EHR (structured records), PACS (images), lab system (results), -genomics database, and pharmacy system. Data from EHRs, remote monitors, and -patient-reported tools arrives with uneven quality and inconsistent codes. - -==== Why VeriSimDB - -A patient entity spans document (clinical notes), tensor (medical images), -graph (care team relationships), temporal (longitudinal records), spatial -(facility locations), and provenance (clinician, device, lab attribution). -Drift detection catches when a medication list in the pharmacy system diverges -from the EHR -- a life-safety issue. VCL-UT proof certificates satisfy audit -requirements. - -==== Features Used - -* Drift detection (medication/allergy discrepancies) -* Multi-modal storage (6+ modalities) -* Heterogeneous federation (EHR + pharmacy + lab databases) -* VCL-UT (audit-grade queries) - -==== Outreach Targets - -* Health IT interoperability working groups -* OpenHIE community -* Academic medical centres with multi-system architectures -* FHIR implementer community - ---- - -=== 4. Digital Twin Synchronisation - -*Alignment: High (Drift Is the Value Proposition)* - -==== The Problem - -Digital twins must maintain real-time consistency between the physical world -(sensor data), simulation model (computed state), and historical record -(time series). Synchronisation and scalability challenges are documented -across 11 distinct patterns. Multi-sensor data faces spatio-temporal -misalignment and domain shifts. - -==== Why VeriSimDB - -A digital twin entity exists as sensor time-series (temporal), 3D model -(tensor/spatial), component graph (graph), simulation parameters (document), -and provenance chain (sensor firmware, calibration, simulation version). -Drift detection IS the purpose -- catching when simulated state diverges -from physical state. Self-normalisation triggers model recalibration. - -==== Features Used - -* Drift detection (simulation vs. physical state) -* Temporal + spatial + tensor modalities -* Provenance tracking (sensor calibration history) -* Self-normalisation (recalibration triggers) - -==== Outreach Targets - -* Eclipse Digital Twin working group -* Azure Digital Twins / AWS IoT TwinMaker users hitting sync issues -* Industrial IoT platforms (Siemens, GE) -* Smart city / building management platforms - ---- - -=== 5. Data Mesh / Data Catalog Consistency - -*Alignment: High* - -==== The Problem - -Data mesh architectures distribute ownership across domain teams. Data -catalogs like DataHub use relational databases for storage, Elasticsearch -for search, and graph databases for relationships -- all connected via -Kafka streams. Entity "customer" might be defined differently in Sales -(PostgreSQL), Analytics (Snowflake), Search (Elasticsearch), and ML (feature -store). Definition drift breaks downstream analytics silently. - -==== Why VeriSimDB - -VeriSimDB sits as a consistency layer above the data mesh. Each shared -entity's definition becomes a VeriSimDB entity with representations across -domains. Drift detection catches when Sales' definition of "active customer" -diverges from Analytics' -- before it corrupts dashboards. Heterogeneous -federation (PostgreSQL + Elasticsearch) queries across domain boundaries. -The semantic modality stores authoritative business glossary definitions. - -==== Features Used - -* Drift detection (cross-domain entity definition divergence) -* Heterogeneous federation (PostgreSQL + Elasticsearch + ArangoDB) -* Semantic modality (business glossary / ontology) -* Provenance (domain team attribution) - -==== Outreach Targets - -* Data mesh practitioners (Zhamak Dehghani's community) -* DataHub / OpenMetadata users -* Enterprise data governance teams -* Data platform teams at mid-large companies - ---- - -=== 6. Knowledge Graph Quality Assurance (Wikidata Scale) - -*Alignment: Strong* - -==== The Problem - -Wikidata has 120M+ entities with documented "large-scale conceptual disarray" --- entities simultaneously treated as both instances and classes. The predicates -P31 (instance of) and P279 (subclass of) suffer systematic inconsistencies. -Research explicitly focuses on detecting inconsistencies, errors, and biases. - -==== Why VeriSimDB - -A Wikidata entity exists as graph relationships, document text (300+ -languages), vector embeddings, semantic annotations, and provenance (editor, -bot, source attribution). Drift detection catches when an entity's graph -classification contradicts its textual description -- precisely the documented -"conceptual disarray". VCL-UT formally verifies ontological constraints. - -==== Features Used - -* Drift detection (graph structure vs. semantic classification) -* Multi-modal storage (graph + document + vector + semantic + provenance) -* VCL-UT (ontological constraint verification) -* Self-normalisation (automated deduplication and constraint repair) - -==== Outreach Targets - -* Wikidata community (Wikidata Workshop contributors) -* DBpedia maintainers -* ConceptNet team -* OpenAlex / Semantic Scholar teams - ---- - -=== 7. Earth Observation / Climate Science - -*Alignment: Strong (Exercises All 8 Modalities)* - -==== The Problem - -ESA's Copernicus programme integrates 18.7M aligned images across multiple -satellite sensors. The same climate observation (e.g., sea surface temperature) -has representations as a satellite image pixel, a numerical value in NetCDF, -a time-series point, a geospatial coordinate, and a provenance chain from raw -Level-0 data to derived Level-3 products. Cross-sensor calibration verification -and processing chain reproducibility are critical challenges. - -==== Why VeriSimDB - -A climate observation entity uses all 8 modalities: tensor (imagery), temporal -(time series), spatial (geolocation), provenance (processing chain), vector -(derived embeddings), graph (sensor/platform relationships), document (metadata), -semantic (climate ontology terms). Drift detection catches cross-sensor -calibration divergence. Heterogeneous federation enables cross-archive queries -(NASA Earthdata, ESA CCI, national meteorological services). - -==== Features Used - -* All 8 modalities -* Drift detection (cross-sensor calibration verification) -* Heterogeneous federation (multi-archive queries) -* Provenance (full processing chain from L0 to L3) - -==== Outreach Targets - -* ESA Climate Change Initiative (CCI) -* Pangeo community (Python earth science stack) -* CERN Rucio team (analogous data management challenges) -* NOAA / NASA Earthdata teams - -== Adoption Priority Matrix - -[cols="1,1,1,1,1,1"] -|=== -| Domain | Drift Detection | Federation | VCL-UT | Readiness | Priority - -| GraphRAG | Critical | High | Medium | *Ship today* | *1* -| Biomedical | Critical | High | Critical | Q2 2026 | 2 -| Healthcare FHIR | Critical (safety) | High | Critical (audit) | Q3 2026 | 3 -| Digital Twin | Critical (purpose) | Medium | Medium | *Ship today* | 4 -| Data Mesh | High | Critical | Medium | *Ship today* | 5 -| Knowledge Graphs | High | Medium | Critical | Q2 2026 | 6 -| Earth Observation | High | High | High | Q3 2026 | 7 -|=== - -== Outreach Plan - -=== Phase 1: GraphRAG Community (Immediate) - -1. Write a blog post: "Detecting Embedding Drift in RAG Pipelines with VeriSimDB" -2. Create a demo integration: VeriSimDB + LangChain + Qdrant -3. Submit to the LangChain community showcase -4. Post in vector database communities (Qdrant Discord, Weaviate Slack) - -=== Phase 2: Data Quality / Data Mesh (Month 2) - -1. Write a comparison post: "VeriSimDB vs. Great Expectations vs. Monte Carlo" -2. Create a DataHub integration showing drift-aware data catalog -3. Present at a data engineering meetup (Data Council, dbt Community) - -=== Phase 3: Scientific Computing (Month 3-4) - -1. Submit an abstract to a bioinformatics conference (ISMB, BOSC) -2. Create a tutorial: "Multi-Omics Entity Consistency with VeriSimDB" -3. Engage with ELIXIR infrastructure network - -=== Phase 4: Conference Presence (Month 4-6) - -1. Submit to Strange Loop / FOSDEM database devroom -2. Prepare a live demo: 1000 entities, 50 corrupted, 100% detection and repair -3. Create a 5-minute screencast for the project README - -== Honest Gaps for External Users - -Any external adopter should know: - -[cols="1,3"] -|=== -| Gap | Impact - -| VCL-UT proofs incomplete | PROOF clauses parse but don't generate verifiable - certificates yet. Lean type checker integration planned Q2 2026. - -| No SQL dialect | VCL is intentionally not SQL. Users wanting familiar syntax - need to learn VCL (comprehensive docs provided). - -| Query planner is basic | Queries using all 8 modalities may not be optimally - planned. Single-modality queries perform well. - -| ~~C++ dependency in graph store~~ | *Resolved.* Oxigraph is optional (feature-flagged, - off by default). Default and persistent graph backends are pure Rust (redb). - Container builds require zero C/C++ toolchain. - -| No distributed consensus | Single-instance deployments are safe. Multi-node - requires application-level coordination. -|=== - -These are documented honestly in `KNOWN-ISSUES.adoc` and the -link:getting-started.adoc[Getting Started Guide]. diff --git a/verisimdb/docs/backwards-compatibility.adoc b/verisimdb/docs/backwards-compatibility.adoc deleted file mode 100644 index bf956877..00000000 --- a/verisimdb/docs/backwards-compatibility.adoc +++ /dev/null @@ -1,669 +0,0 @@ -// SPDX-License-Identifier: CC-BY-SA-4.0 - -= VCL Backwards Compatibility Strategy -:toc: -:toc-placement!: - -Strategy for maintaining backwards compatibility across VCL versions, schema evolution, and API changes. - -toc::[] - -== Overview - -VeriSimDB follows **semantic versioning** (SemVer) for VCL and API compatibility: - -- **Major version** (1.x.x → 2.x.x): Breaking changes allowed -- **Minor version** (1.1.x → 1.2.x): New features, backwards compatible -- **Patch version** (1.1.1 → 1.1.2): Bug fixes only, backwards compatible - -Current VCL version: **1.0.0** - -== Compatibility Guarantees - -=== What We Guarantee - -**Query Syntax (VCL):** - -- Valid VCL 1.x queries will parse and execute correctly in all 1.x releases -- New keywords/operators are additive (won't conflict with existing identifiers) -- Deprecations announced at least 2 minor versions before removal - -**API Endpoints:** - -- `/api/v1/*` endpoints remain stable within major version -- New endpoints added without removing old ones -- Response format changes are backwards compatible (new fields added, never removed) - -**Data Formats:** - -- Octad storage format versioned independently -- Old format readers maintained for at least 1 year after new format introduction -- Automatic migration on read for legacy formats - -=== What We Don't Guarantee - -**Performance:** - -- Query performance may change between versions (optimizations, regressions) -- Federation latency depends on network and remote store versions - -**Error Messages:** - -- Error message text may change (error codes remain stable) -- Additional validation may be added (making previously accepted queries invalid) - -**Internal APIs:** - -- ReScript/Elixir module APIs may change without notice -- Use public HTTP API for external integrations - -== Versioning Strategy - -=== VCL Syntax Versions - -Each query can optionally specify a version: - -```vcl -VERSION 1.0; - -FROM verisim:semantic -WHERE octad.types INCLUDES "https://schema.org/Person" -LIMIT 10; -``` - -**Default Behavior:** - -- If `VERSION` is omitted, use latest stable version (currently 1.0) -- Parser detects version and applies appropriate syntax rules -- Mixing versions in a single query is NOT allowed - -=== API Versioning - -**URL-based versioning:** - -``` -/api/v1/octads # Version 1 (current) -/api/v2/octads # Version 2 (future) -``` - -**Multiple versions supported concurrently:** - -- Old API versions maintained for at least 12 months after new version release -- Deprecation warnings in response headers: `X-API-Deprecated: true; sunset=2026-12-31` - -=== Schema Evolution - -**Octad schema changes:** - -```scheme -;; STATE.scm tracks schema version -(metadata - (schema-version "1.2.0") ; Current Octad schema - (min-compatible-version "1.0.0")) ; Oldest readable format -``` - -**Format Evolution:** - -1. **Adding fields** (backwards compatible): - - Old readers ignore unknown fields - - New readers provide defaults for missing fields - -2. **Removing fields** (breaking change): - - Deprecated for 2+ minor versions first - - Removal only in major version bump - -3. **Changing field types** (breaking change): - - Add new field with new type - - Deprecate old field - - Remove old field in major version bump - -== Breaking vs Non-Breaking Changes - -=== Non-Breaking Changes (Minor Version) - -**Safe to add in 1.x.x releases:** - -- New VCL keywords (if not previously valid identifiers) -- New modality types -- New query operators (`INTERSECT`, `UNION ALL`) -- New API endpoints -- New response fields (optional) -- New error codes -- Performance optimizations - -**Example: Adding `INTERSECT` operator** - -```vcl --- Version 1.2.0 introduces INTERSECT --- Old queries (1.0.0) continue to work - --- New query using INTERSECT (requires 1.2.0+) -VERSION 1.2; -(FROM verisim:graph WHERE octad.id IN @set1) -INTERSECT -(FROM verisim:graph WHERE octad.id IN @set2); -``` - -=== Breaking Changes (Major Version) - -**Requires 2.0.0:** - -- Removing VCL keywords -- Changing query semantics (e.g., `LIMIT` behavior) -- Removing API endpoints -- Removing response fields -- Changing error codes - -**Example: Removing deprecated `USING` clause** - -```vcl --- Version 1.0.0 syntax (deprecated in 1.2.0) -FROM verisim:graph -USING CONTRACT CitationContract -WHERE octad.types INCLUDES "Paper"; - --- Version 2.0.0 syntax (USING removed, VERIFY required) -FROM verisim:graph -VERIFY CitationContract -WHERE octad.types INCLUDES "Paper"; -``` - -== Deprecation Process - -=== Timeline - -1. **Announce deprecation** (version N.x.x): - - Add deprecation warning to docs - - Log warning when deprecated feature is used - - Add `X-Deprecated: true` header to API responses - -2. **Maintain support** (versions N.x.x through N+1.x.x): - - Feature still works - - Warnings continue - -3. **Remove** (version N+2.0.0): - - Feature removed - - Queries using deprecated feature return error - -**Minimum timeline:** 6 months between deprecation and removal - -=== Deprecation Warnings - -**VCL Parser:** - -```rescript -// VCLParser.res -let parseUsingClause = (tokens: array): result => { - Logger.warn("USING clause is deprecated since 1.2.0. Use VERIFY instead. Will be removed in 2.0.0.") - - // Still parse correctly - parseUsingClauseImpl(tokens) -} -``` - -**API Response Headers:** - -```http -HTTP/1.1 200 OK -X-Deprecated: true -X-Deprecated-Since: 1.2.0 -X-Sunset: 2026-12-31 -Sunset: Wed, 31 Dec 2026 23:59:59 GMT -Link: ; rel="sunset" -``` - -== Migration Strategies - -=== Strategy 1: Version Pinning - -**Use case:** Stability-critical systems - -```elixir -# config.exs -config :verisim, - vcl_version: "1.0.0", # Pin to specific version - api_version: "v1" -``` - -**Trade-offs:** - -- ✅ No surprises from version updates -- ❌ Miss performance improvements -- ❌ Miss security fixes - -=== Strategy 2: Conservative Upgrades - -**Use case:** Production systems with good test coverage - -- Upgrade to latest **minor** version automatically (1.1.x → 1.2.x) -- Upgrade to latest **major** version manually (1.x.x → 2.x.x) - -```yaml -# Dockerfile -ENV VCL_VERSION="1.x" # Auto-upgrade within major version -``` - -**Trade-offs:** - -- ✅ Get bug fixes and features -- ✅ No breaking changes -- ⚠️ Require testing before major version upgrades - -=== Strategy 3: Gradual Migration - -**Use case:** Large codebases with many queries - -1. **Identify deprecated features:** - ```bash - grep -r "USING" queries/ - # Find all queries using deprecated syntax - ``` - -2. **Update queries gradually:** - ```vcl - -- Old query (1.0.0) - FROM verisim:graph - USING CONTRACT CitationContract - WHERE ...; - - -- New query (1.2.0+) - VERSION 1.2; - FROM verisim:graph - VERIFY CitationContract - WHERE ...; - ``` - -3. **Test both versions side-by-side:** - ```elixir - old_result = VCL.parse_and_execute(old_query, version: "1.0.0") - new_result = VCL.parse_and_execute(new_query, version: "1.2.0") - - assert old_result.data == new_result.data - ``` - -4. **Switch cutover:** - ```elixir - # Deploy new queries - # Monitor for 1 week - # Remove old queries - ``` - -== Compatibility Testing - -=== Version Compatibility Matrix - -**Test all VCL versions against current parser:** - -```elixir -# test/vcl_compatibility_test.exs -defmodule VeriSim.VCLCompatibilityTest do - use ExUnit.Case - - @vcl_versions ["1.0.0", "1.1.0", "1.2.0"] - @test_queries [ - # Basic SELECT - """ - VERSION 1.0; - FROM verisim:graph - WHERE octad.id = @id; - """, - - # WITH clause (added 1.1.0) - """ - VERSION 1.1; - WITH VERIFICATION CitationContract - FROM verisim:graph - WHERE octad.types INCLUDES "Paper"; - """, - - # INTERSECT (added 1.2.0) - """ - VERSION 1.2; - (FROM verisim:graph WHERE ...) INTERSECT (FROM verisim:vector WHERE ...); - """ - ] - - test "parse all versions" do - for query <- @test_queries do - assert {:ok, _ast} = VCL.parse(query) - end - end - - test "1.0.0 queries work in all versions" do - query_1_0 = """ - VERSION 1.0; - FROM verisim:graph LIMIT 10; - """ - - for version <- @vcl_versions do - config = %{vcl_version: version} - assert {:ok, _result} = VCL.execute(query_1_0, config) - end - end -end -``` - -=== Upgrade Testing - -**Test suite for major version upgrades:** - -```elixir -# test/upgrade_test.exs -defmodule VeriSim.UpgradeTest do - use ExUnit.Case - - @tag :upgrade - test "1.x data readable in 2.x" do - # Create octad with 1.x format - octad_1_x = create_octad_v1() - - # Upgrade to 2.x - Application.put_env(:verisim, :schema_version, "2.0.0") - - # Read octad (should auto-migrate) - assert {:ok, octad_2_x} = VeriSim.Octad.read(octad_1_x.id) - - # Verify data integrity - assert octad_1_x.title == octad_2_x.title - assert octad_1_x.body == octad_2_x.body - end - - @tag :upgrade - test "1.x queries fail gracefully in 2.x if deprecated features used" do - deprecated_query = """ - VERSION 1.0; - FROM verisim:graph - USING CONTRACT CitationContract - WHERE octad.id = @id; - """ - - Application.put_env(:verisim, :vcl_version, "2.0.0") - - assert {:error, {:deprecated_syntax, "USING clause removed in 2.0.0"}} = - VCL.parse(deprecated_query) - end -end -``` - -== Data Format Versioning - -=== Octad Storage Format - -**Format version embedded in each octad:** - -```json -{ - "_format_version": "1.2.0", - "octad_id": "550e8400-e29b-41d4-a716-446655440000", - "title": "Machine Learning Paper", - "body": "...", - "created_at": "2025-01-15T10:30:00Z" -} -``` - -**Format evolution example:** - -```rust -// rust-core/verisim-octad/src/format.rs - -#[derive(Serialize, Deserialize)] -#[serde(tag = "_format_version")] -pub enum OctadFormat { - #[serde(rename = "1.0.0")] - V1_0_0(OctadV1_0_0), - - #[serde(rename = "1.1.0")] - V1_1_0(OctadV1_1_0), - - #[serde(rename = "1.2.0")] - V1_2_0(OctadV1_2_0), -} - -impl OctadFormat { - /// Read octad in any format, auto-migrate to latest - pub fn read_and_migrate(data: &[u8]) -> Result { - let format: OctadFormat = serde_json::from_slice(data)?; - - match format { - OctadFormat::V1_0_0(octad) => octad.migrate_to_1_1_0()?.migrate_to_1_2_0(), - OctadFormat::V1_1_0(octad) => octad.migrate_to_1_2_0(), - OctadFormat::V1_2_0(octad) => Ok(octad), - } - } -} - -impl OctadV1_0_0 { - fn migrate_to_1_1_0(self) -> Result { - Ok(OctadV1_1_0 { - octad_id: self.octad_id, - title: self.title, - body: self.body, - created_at: self.created_at, - // New field in 1.1.0 (default value) - provenance: None, - }) - } -} -``` - -=== Migration on Read vs Write - -**Read migration (preferred):** - -- Old format stored on disk -- Auto-migrate to new format when reading -- No disk writes required -- Lazy migration (only migrates what's accessed) - -**Write migration (for breaking changes):** - -- Background job migrates all octads -- Required when old format becomes unsupported -- Progress tracked: `migration_status.json` - -```elixir -# lib/verisim/migration_worker.ex -defmodule VeriSim.MigrationWorker do - @doc """ - Migrate all octads from old format to new format. - """ - def migrate_all_octads(from_version, to_version) do - total = count_octads(from_version) - - Stream.iterate(0, &(&1 + 100)) - |> Stream.take_while(&(&1 < total)) - |> Enum.each(fn offset -> - octads = fetch_octads(from_version, limit: 100, offset: offset) - - octads - |> Enum.each(fn octad -> - migrated = migrate_octad(octad, from_version, to_version) - write_octad(migrated) - end) - - update_progress(offset + length(octads), total) - end) - end -end -``` - -== Federation Compatibility - -=== Mixed Version Networks - -**Challenge:** Federated stores may run different VeriSimDB versions. - -**Solution: Version negotiation** - -```http -POST /api/v1/federation/query HTTP/1.1 -Host: remote-store.edu -X-VeriSimDB-Version: 1.2.0 -X-VCL-Version: 1.0.0 - -{ - "query": "FROM verisim:graph WHERE octad.id = @id" -} -``` - -**Response:** - -```http -HTTP/1.1 200 OK -X-VeriSimDB-Version: 1.1.0 -X-VCL-Version-Supported: 1.0.0, 1.1.0 - -{ - "data": [...] -} -``` - -**Version compatibility rules:** - -1. **VCL version:** Query must use syntax supported by remote store - - If remote supports `[1.0.0, 1.1.0]` and query uses `1.2.0` → error - - Client downgrades query syntax if possible - -2. **API version:** Client uses highest common API version - - Local has `v1`, remote has `v1, v2` → use `v1` - -3. **Data format:** Responses auto-migrated by receiving store - -=== Feature Detection - -**Query store capabilities before sending complex queries:** - -```elixir -def query_federation(stores, query) do - # Check if all stores support required features - required_features = VCL.detect_features(query) - # [:zkp_verification, :tensor_modality, :drift_detection] - - compatible_stores = Enum.filter(stores, fn store -> - store_features = fetch_capabilities(store) - MapSet.subset?(required_features, store_features) - end) - - if Enum.empty?(compatible_stores) do - {:error, {:insufficient_stores, "No stores support required features"}} - else - execute_federated_query(compatible_stores, query) - end -end -``` - -**Store capabilities endpoint:** - -```http -GET /api/v1/capabilities HTTP/1.1 - -Response: -{ - "verisimdb_version": "1.2.0", - "vcl_versions": ["1.0.0", "1.1.0", "1.2.0"], - "modalities": ["graph", "vector", "semantic", "document", "temporal"], - "features": [ - "zkp_verification", - "drift_detection", - "cache_sharing", - "reversibility" - ], - "limits": { - "max_query_size": 1048576, - "max_results": 10000, - "timeout_ms": 60000 - } -} -``` - -== Client Library Compatibility - -=== Language Bindings - -**Support multiple language clients:** - -- Rust client: `verisim-client-rs` (official) -- Python client: `verisim-client-py` (community) -- JavaScript client: `verisim-client-js` (official, Deno) - -**Version compatibility:** - -```toml -# Cargo.toml -[dependencies] -verisim-client = "1.2.0" # Client version matches server minor version - -# Server compatibility: -# verisim-client 1.x.x works with verisim-server 1.y.z (any y, z) -# verisim-client 2.x.x works with verisim-server 2.y.z (any y, z) -``` - -=== Client API Stability - -**Stable client API (no breaking changes in minor versions):** - -```rust -// Stable API -pub trait VeriSimClient { - fn query(&self, vcl: &str) -> Result; - fn create_octad(&self, octad: Octad) -> Result; - fn get_octad(&self, id: OctadId) -> Result; -} - -// Adding new methods is OK (minor version bump) -pub trait VeriSimClient { - fn query(&self, vcl: &str) -> Result; - fn create_octad(&self, octad: Octad) -> Result; - fn get_octad(&self, id: OctadId) -> Result; - - // New in 1.2.0 - fn batch_query(&self, queries: &[&str]) -> Result>; -} -``` - -== Compatibility Checklist - -=== Before Releasing Minor Version (1.x.x → 1.y.x) - -- [ ] All existing tests pass with new changes -- [ ] New features have tests -- [ ] Deprecation warnings added for features planned for removal -- [ ] Documentation updated with new features -- [ ] `CHANGELOG.md` updated with new features -- [ ] No breaking changes to API, VCL syntax, or data formats -- [ ] Federation compatibility tested with previous minor version - -=== Before Releasing Major Version (1.x.x → 2.x.x) - -- [ ] Migration guide written for all breaking changes -- [ ] Deprecated features removed (with at least 6 months notice) -- [ ] Data migration tools provided -- [ ] Previous major version supported for at least 12 months -- [ ] All examples and tutorials updated -- [ ] Client libraries updated -- [ ] Federation compatibility tested -- [ ] Rollback procedure documented - -== Summary - -VeriSimDB maintains backwards compatibility through: - -1. **Semantic versioning** - Clear breaking vs non-breaking change policy -2. **Deprecation timeline** - At least 6 months notice before removal -3. **Format versioning** - Auto-migration for data formats -4. **API versioning** - Multiple API versions supported concurrently -5. **VCL versioning** - Explicit version markers in queries -6. **Feature detection** - Capability negotiation for federation -7. **Comprehensive testing** - Compatibility test suite - -**For users:** - -- Pin to major version for stability -- Upgrade to minor versions for bug fixes and features -- Plan major version upgrades carefully with migration guide - -**For developers:** - -- Follow deprecation process for all breaking changes -- Maintain compatibility test suite -- Document version changes in CHANGELOG.md diff --git a/verisimdb/docs/business/business-case.adoc b/verisimdb/docs/business/business-case.adoc deleted file mode 100644 index eb553b51..00000000 --- a/verisimdb/docs/business/business-case.adoc +++ /dev/null @@ -1,588 +0,0 @@ -// SPDX-License-Identifier: CC-BY-SA-4.0 -// Copyright (c) 2026 Jonathan D.A. Jewell (hyperpolymath) -= VeriSimDB Business Case -Jonathan D.A. Jewell -:revdate: 2026-02-28 -:revremark: Initial business case document -:toc: macro -:toclevels: 3 -:sectnums: -:icons: font - -[abstract] -VeriSimDB is the world's first multimodal database with cross-modal drift detection -and self-normalisation. This business case document establishes the strategic -rationale, market opportunity, revenue model, financial projections, and risk -analysis for VeriSimDB's commercialisation under an open-core model. - -toc::[] - -== Executive Summary - -VeriSimDB addresses a fundamental unsolved problem in data infrastructure: when the -same entity is represented across multiple data modalities (graph, vector, document, -temporal, etc.), those representations inevitably diverge. No existing database -detects, quantifies, or repairs this divergence. VeriSimDB does -- automatically. - -The database stores entities across 8 simultaneous representations (the "octad": -graph, vector, tensor, semantic, document, temporal, provenance, spatial) and -continuously monitors cross-modal consistency. When representations drift apart, -VeriSimDB detects the divergence and can self-normalise -- repairing data before -downstream systems consume stale or contradictory information. - -This capability positions VeriSimDB at the intersection of three high-growth -markets: data quality tools ($2.3B by 2027), graph databases ($5.6B), and the -emerging multimodal database segment ($800M+). The open-core business model -provides a free, full-featured community edition under the Palimpsest License -(PMPL-1.0-or-later), with revenue generated through commercial support licenses, -enterprise SLAs, and deployment consulting. - -== Problem Statement - -=== The Data Quality Crisis - -According to Gartner, poor data quality costs organisations an average of *$12.9 -million per year*. The IBM Global AI Adoption Index found that 35% of enterprises -cite data quality as their primary barrier to AI adoption. McKinsey estimates that -data workers spend up to 80% of their time on data cleaning and reconciliation -rather than analysis. - -=== Root Cause: Cross-System Inconsistency - -The root cause is not dirty data at the point of entry. It is *cross-system -inconsistency* -- the same real-world entity represented differently across -multiple systems, with no mechanism to detect when those representations diverge. - -Consider a customer record that exists simultaneously as: - -* A *graph node* in a relationship database (Neo4j) -* A *vector embedding* in a search index (Pinecone) -* A *document* in an operational store (MongoDB) -* A *row* in a relational database (PostgreSQL) -* A *temporal snapshot* in an audit trail - -When the customer's address changes in the relational database, the graph node -may be updated, but the vector embedding is not re-computed, the document store -retains the old address, and the temporal trail shows a gap. These systems have -*no shared model of consistency*. The representations have drifted. - -=== Why Existing Solutions Fail - -Current approaches to this problem are reactive and manual: - -[cols="2,3,3"] -|=== -| Approach | Description | Limitation - -| ETL pipelines -| Periodic batch synchronisation between systems -| Latency (hours to days), brittle, no drift detection - -| Change Data Capture (CDC) -| Stream changes from source to downstream systems -| One-directional, no cross-modal awareness - -| Master Data Management (MDM) -| Centralised golden record -| Expensive, slow to implement, does not handle multimodal representations - -| Data observability tools -| Monitor data pipelines for anomalies -| Detect pipeline failures, not semantic drift between modalities - -| Multi-model databases -| Store multiple models in one database (ArangoDB, SurrealDB) -| No cross-modal consistency checking, no drift detection -|=== - -None of these solutions address the fundamental problem: *no system tracks whether -multiple representations of the same entity are consistent with each other*. - -== Solution: VeriSimDB - -=== The Octad Model - -VeriSimDB introduces the *octad model* -- every entity can be simultaneously -represented across 8 modalities: - -[cols="1,3,3"] -|=== -| # | Modality | Purpose - -| 1 | *Graph* | Relationships, connections, topology -| 2 | *Vector* | Semantic similarity, embeddings, nearest-neighbour search -| 3 | *Tensor* | Multi-dimensional numerical data, ML feature stores -| 4 | *Semantic* | Ontological meaning, RDF triples, knowledge representation -| 5 | *Document* | Schemaless structured data, JSON/BSON documents -| 6 | *Temporal* | Time-versioned state, bitemporal queries, audit trails -| 7 | *Provenance* | Origin tracking, transformation lineage, trust chains -| 8 | *Spatial* | Geographic and geometric data, GIS operations -|=== - -=== Cross-Modal Drift Detection - -VeriSimDB continuously monitors the consistency of an entity's representations -across all active modalities. The drift detection engine computes a *coherence -score* for each entity -- a quantified measure of how well its representations -agree with each other. - -When drift exceeds configurable thresholds, VeriSimDB can: - -1. *Alert* -- notify downstream systems that data may be inconsistent -2. *Quarantine* -- mark the entity as potentially drifted until resolved -3. *Self-normalise* -- automatically reconcile representations using - configurable resolution strategies - -=== Formal Verification: VCL-UT - -VeriSimDB's query language, VCL (VeriSim Consonance Language), includes an -optional dependent-type layer (VCL-UT) that provides *compile-time proofs* of -query correctness. This means: - -* Type errors in queries are caught before execution -* Cross-modal joins are verified for semantic compatibility -* Drift thresholds can be encoded as type-level constraints - -=== Federation - -VeriSimDB can federate across heterogeneous database backends (PostgreSQL, -ArangoDB, Elasticsearch, Neo4j, etc.), bringing drift detection to existing -infrastructure without requiring data migration. - -== Market Analysis - -=== Market Sizing - -[cols="2,1,1,1,3"] -|=== -| Segment | TAM | SAM | SOM | Notes - -| Data Quality Tools -| $2.3B -| $460M -| $23M -| MarketsAndMarkets 2027 projection; 20% addressable multimodal segment; 1% serviceable - -| Graph Databases -| $5.6B -| $280M -| $14M -| Grand View Research; 5% multimodal sub-segment; 1% serviceable - -| Vector Databases -| $1.2B -| $120M -| $6M -| Emerging market; 10% enterprise segment; 1% serviceable - -| Multimodal Databases -| $800M -| $160M -| $8M -| Nascent market; 20% addressable early adopters; 1% serviceable - -| *Combined* -| -| -| *$51M* -| Total serviceable obtainable market across segments -|=== - -=== Target Segments - -*Primary (Years 1-2):* - -* GraphRAG / AI knowledge graph engineers building retrieval-augmented generation - systems that combine graph and vector modalities -* Biomedical researchers managing multi-modal clinical and research data - -*Secondary (Years 2-4):* - -* Financial services firms requiring cross-system consistency for regulatory compliance -* Supply chain organisations tracking provenance across complex networks - -*Tertiary (Years 3-5+):* - -* Enterprise platform teams standardising multimodal data infrastructure -* OEM partners embedding VeriSimDB as a backend for vertical SaaS products - -== Competitive Landscape - -=== Feature Comparison - -[cols="3,1,1,1,1,1,1"] -|=== -| Feature | VeriSimDB | Neo4j | Pinecone | Weaviate | ArangoDB | SurrealDB - -| Graph queries -| Yes -| Yes -| No -| No -| Yes -| Yes - -| Vector search -| Yes -| Limited -| Yes -| Yes -| No -| Yes - -| Full-text search -| Yes -| Yes -| No -| Yes -| Yes -| Yes - -| Tensor operations -| Yes -| No -| No -| No -| No -| No - -| Temporal versioning -| Yes -| No -| No -| No -| No -| No - -| Provenance tracking -| Yes -| No -| No -| No -| No -| No - -| Spatial queries -| Yes -| Yes -| No -| No -| Yes -| No - -| *Cross-modal drift detection* -| *Yes* -| No -| No -| No -| No -| No - -| *Self-normalisation* -| *Yes* -| No -| No -| No -| No -| No - -| *Formal proofs (VCL-UT)* -| *Yes* -| No -| No -| No -| No -| No - -| Federation -| Yes -| No -| No -| No -| No -| No - -| Multiple modalities per entity -| 8 -| 1 -| 1 -| 2 -| 3 -| 3 - -| Open source -| PMPL -| GPLv3 -| No -| BSD-3 -| Apache 2.0 -| BSL -|=== - -=== Competitive Positioning - -VeriSimDB does not compete directly with single-modality databases. Instead, it -occupies a new category: *cross-modal consistency infrastructure*. The competitive -dynamics are: - -* *Neo4j* excels at graph traversals but has no vector, tensor, temporal, or - provenance modalities, and no drift detection -* *Pinecone* excels at vector similarity search but is purely vector-only with no - cross-modal awareness -* *Weaviate* combines vector search with some text/object capabilities but lacks - drift detection, temporal versioning, and formal query verification -* *ArangoDB* is the closest multi-model competitor but has no cross-modal - consistency checking, no drift detection, and no self-normalisation -* *SurrealDB* offers multi-model storage but focuses on developer ergonomics - rather than data consistency guarantees - -== Revenue Model - -=== Open Core Structure - -[cols="2,4"] -|=== -| Tier | Description - -| *Community Edition (PMPL)* -| Full-featured VeriSimDB with all 8 modalities, drift detection, - self-normalisation, VCL, and federation. Free, open source under PMPL-1.0-or-later. - -| *Commercial Support License* -| Enterprise SLA (99.9% uptime guarantee), priority security patches (24-hour - response), dedicated support channel, deployment consulting, training. - Annual subscription. - -| *OEM License* -| Embedding VeriSimDB as a component within third-party products. - Per-deployment royalty or flat annual fee. -|=== - -=== Pricing Strategy - -The community edition is intentionally full-featured. Revenue comes from the -*operational value* of enterprise support, not from feature gating. This approach: - -* Maximises adoption and community growth -* Creates a large funnel of users who may convert to commercial support -* Avoids the "open core bait-and-switch" perception that damages developer trust -* Aligns with the PMPL license philosophy - -=== Revenue Streams - -1. *Commercial support subscriptions* -- annual contracts, tiered by - organisation size and SLA level -2. *Deployment consulting* -- architecture review, migration planning, - performance tuning, integration with existing infrastructure -3. *Training and certification* -- VCL/VCL-UT training for development teams -4. *OEM licensing* -- for vendors embedding VeriSimDB in their products - -== Financial Projections - -=== 5-Year Revenue Forecast (Conservative Scenario) - -[cols="1,2,4"] -|=== -| Year | Revenue | Activities - -| 1 | $0 -| Community building, conference presentations, open-source adoption, - documentation, academic paper publication - -| 2 | $150,000 -| First enterprise consulting contracts, early adopter programme, - pilot deployments with 2-3 organisations - -| 3 | $500,000 -| Enterprise support license revenue begins, 5-10 paying customers, - community exceeds 1,000 active users - -| 4 | $1,500,000 -| Market expansion, partnership deals with system integrators, - 15-25 paying customers, OEM discussions - -| 5 | $3,000,000 -| Scale: self-serve support tiers, 40+ enterprise customers, - OEM revenue, managed service pilot -|=== - -=== Cost Structure (Annual) - -[cols="2,1,1,1,1,1"] -|=== -| Category | Year 1 | Year 2 | Year 3 | Year 4 | Year 5 - -| Engineering -| $0 -| $80,000 -| $250,000 -| $600,000 -| $1,200,000 - -| Infrastructure -| $500 -| $2,000 -| $10,000 -| $50,000 -| $150,000 - -| Marketing -| $1,000 -| $5,000 -| $30,000 -| $100,000 -| $250,000 - -| Legal -| $500 -| $2,000 -| $5,000 -| $15,000 -| $30,000 - -| Support -| $0 -| $10,000 -| $50,000 -| $150,000 -| $400,000 - -| Operations -| $0 -| $5,000 -| $20,000 -| $80,000 -| $200,000 - -| *Total* -| *$2,000* -| *$104,000* -| *$365,000* -| *$995,000* -| *$2,230,000* -|=== - -=== Path to Profitability - -Under the conservative scenario, VeriSimDB reaches *cash-flow positive in Year 3* -($500K revenue vs. $365K costs) and generates *$770K operating profit by Year 5* -($3M revenue vs. $2.23M costs). The moderate and aggressive scenarios reach -profitability faster and at higher margins. - -=== Unit Economics - -[cols="2,1,3"] -|=== -| Metric | Value | Notes - -| Customer Acquisition Cost (CAC) -| $8,000 -| Enterprise B2B, developer-led adoption reduces traditional sales costs - -| Lifetime Value (LTV) -| $120,000 -| 3-year average contract at $40K ACV - -| Annual Contract Value (ACV) -| $40,000 -| Enterprise support license - -| LTV:CAC Ratio -| 15:1 -| Target: >3:1 is considered healthy - -| Payback Period -| 3 months -| CAC recovered in approximately 3 months of ACV - -| Gross Margin -| 85% -| Software + support, minimal cost of goods sold - -| Net Revenue Retention -| 115% -| Expansion revenue from existing customers -|=== - -== Risks and Mitigations - -=== Technical Risks - -[cols="3,3,3"] -|=== -| Risk | Impact | Mitigation - -| Performance at scale (8 modalities per entity) -| High latency or resource consumption could limit enterprise adoption -| Rust core provides near-C performance; Elixir/OTP orchestration provides - fault-tolerant concurrency; extensive benchmark suite (510 Rust tests + 152 - Elixir tests, 0 failures) - -| VCL-UT complexity deters adoption -| Developers find dependent-type proofs intimidating -| VCL (without DT) is the default; VCL-UT is opt-in for teams that want formal - guarantees. Progressive disclosure: start simple, add proofs when ready. - -| Federation reliability across heterogeneous backends -| Connector bugs or performance variance across backends -| Each federation connector is independently tested; fallback to local-only - mode if federation fails; comprehensive error handling strategy documented -|=== - -=== Market Risks - -[cols="3,3,3"] -|=== -| Risk | Impact | Mitigation - -| Neo4j or ArangoDB adds drift detection -| Reduces differentiation -| VeriSimDB's octad model and VCL-UT provide deep moat; 8-modality consistency - is architecturally difficult to retrofit onto existing databases - -| Market education required (new category) -| Longer sales cycles, higher CAC -| Developer advocacy, academic publications, conference talks, and open-source - community build awareness; category creation is expensive but defensible - -| Enterprise procurement cycles -| Revenue delayed by 6-12 month procurement processes -| Start with developer-led adoption (bottom-up); consulting engagements provide - revenue during procurement cycles -|=== - -=== Business Risks - -[cols="3,3,3"] -|=== -| Risk | Impact | Mitigation - -| Single founder risk -| Bus factor of 1 -| Community contributions, comprehensive documentation, and CI/CD automation - reduce dependency on any individual; early revenue funds additional contributors - -| Open-source sustainability -| Community expects everything free; conversion to commercial support is low -| Full-featured community edition builds trust; commercial support addresses - operational needs (SLA, priority patches) that free users do not require - -| PMPL license unfamiliarity -| Enterprise legal teams may hesitate on non-OSI license -| PMPL is permissive and well-documented; MPL-2.0 fallback available where - required; commercial license option for enterprises that require it -|=== - -== Conclusion - -VeriSimDB addresses a $51M serviceable market at the intersection of data quality, -multimodal databases, and AI infrastructure. Its unique cross-modal drift detection -capability has no direct competitor. The open-core model provides a clear path to -revenue while maintaining open-source community trust. Conservative financial -projections show profitability by Year 3 with healthy unit economics (15:1 -LTV:CAC ratio, 85% gross margin). - -The primary investment needed is in community building (Year 1), followed by a -small engineering team (Year 2-3) to accelerate development and support early -enterprise customers. The technology foundation is sound: Rust core + Elixir/OTP -orchestration, 662 tests with 0 failures, and a formally verified query language. diff --git a/verisimdb/docs/business/financials/cost-structure.csv b/verisimdb/docs/business/financials/cost-structure.csv deleted file mode 100644 index 1df58977..00000000 --- a/verisimdb/docs/business/financials/cost-structure.csv +++ /dev/null @@ -1,14 +0,0 @@ -# SPDX-License-Identifier: MPL-2.0 -# Copyright (c) 2026 Jonathan D.A. Jewell (hyperpolymath) -# -# VeriSimDB Annual Cost Structure — 5-Year Projection -# All values in USD -# -Category,Year1,Year2,Year3,Year4,Year5,Notes -Engineering,0,80000,250000,600000,1200000,"Open-source contributors → hired team" -Infrastructure,500,2000,10000,50000,150000,"CI/CD, cloud testing, demo instances" -Marketing,1000,5000,30000,100000,250000,"Conference travel, content marketing, developer advocacy" -Legal,500,2000,5000,15000,30000,"PMPL license compliance, patents" -Support,0,10000,50000,150000,400000,"Community → commercial support tiers" -Operations,0,5000,20000,80000,200000,"Sales engineering, onboarding" -Total,2000,104000,365000,995000,2230000,"" diff --git a/verisimdb/docs/business/financials/market-sizing.csv b/verisimdb/docs/business/financials/market-sizing.csv deleted file mode 100644 index 761631fc..00000000 --- a/verisimdb/docs/business/financials/market-sizing.csv +++ /dev/null @@ -1,12 +0,0 @@ -# SPDX-License-Identifier: MPL-2.0 -# Copyright (c) 2026 Jonathan D.A. Jewell (hyperpolymath) -# -# VeriSimDB Market Sizing — TAM/SAM/SOM Analysis -# All values in USD -# -Segment,TAM,SAM,SOM,Notes -Data_Quality_Tools,2300000000,460000000,23000000,"$2.3B total, 20% addressable, 1% serviceable" -Graph_Database,5600000000,280000000,14000000,"$5.6B total, 5% multimodal segment, 1% serviceable" -Vector_Database,1200000000,120000000,6000000,"$1.2B total, 10% enterprise segment, 1% serviceable" -Multimodal_Database,800000000,160000000,8000000,"$800M emerging, 20% addressable, 1% serviceable" -Combined_SOM,,,51000000,"Total serviceable obtainable market" diff --git a/verisimdb/docs/business/financials/revenue-model.csv b/verisimdb/docs/business/financials/revenue-model.csv deleted file mode 100644 index 68f60d38..00000000 --- a/verisimdb/docs/business/financials/revenue-model.csv +++ /dev/null @@ -1,12 +0,0 @@ -# SPDX-License-Identifier: MPL-2.0 -# Copyright (c) 2026 Jonathan D.A. Jewell (hyperpolymath) -# -# VeriSimDB Revenue Model — 3 Scenarios Over 5 Years -# All values in USD -# -Year,Conservative,Moderate,Aggressive,Notes -1,0,0,50000,"Community building, conference talks, open-source adoption" -2,150000,300000,600000,"First enterprise consulting contracts" -3,500000,1000000,2000000,"Enterprise license revenue begins" -4,1500000,3000000,6000000,"Market expansion, partnership deals" -5,3000000,7000000,15000000,"Scale: self-serve + enterprise + OEM" diff --git a/verisimdb/docs/business/financials/unit-economics.csv b/verisimdb/docs/business/financials/unit-economics.csv deleted file mode 100644 index 01e0bad7..00000000 --- a/verisimdb/docs/business/financials/unit-economics.csv +++ /dev/null @@ -1,13 +0,0 @@ -# SPDX-License-Identifier: MPL-2.0 -# Copyright (c) 2026 Jonathan D.A. Jewell (hyperpolymath) -# -# VeriSimDB Unit Economics — Enterprise B2B Model -# -Metric,Value,Notes -CAC,8000,"Customer acquisition cost (enterprise B2B)" -LTV,120000,"Lifetime value (3-year average contract)" -ACV,40000,"Annual contract value (enterprise license)" -LTV_CAC_Ratio,15,"Target: >3x is healthy" -Payback_Period_Months,3,"CAC recovered in ~3 months of ACV" -Gross_Margin_Percent,85,"Software + support, minimal COGS" -Net_Revenue_Retention,115,"Expansion revenue from existing customers" diff --git a/verisimdb/docs/business/marketing/feature-comparison.adoc b/verisimdb/docs/business/marketing/feature-comparison.adoc deleted file mode 100644 index b2d14ee6..00000000 --- a/verisimdb/docs/business/marketing/feature-comparison.adoc +++ /dev/null @@ -1,446 +0,0 @@ -// SPDX-License-Identifier: CC-BY-SA-4.0 -// Copyright (c) 2026 Jonathan D.A. Jewell (hyperpolymath) -= VeriSimDB Feature Comparison Matrix -Jonathan D.A. Jewell -:revdate: 2026-02-28 -:revremark: Competitive feature comparison for marketing and sales enablement -:toc: macro -:toclevels: 2 -:sectnums: -:icons: font - -[abstract] -A detailed feature-by-feature comparison of VeriSimDB against the leading -databases in the graph, vector, multi-model, and multimodal segments: Neo4j, -Pinecone, Weaviate, ArangoDB, and SurrealDB. This document is intended for -technical evaluators, architects, and procurement teams. - -toc::[] - -== Summary Matrix - -Legend: {check} = Full support | {half} = Partial/limited | {cross} = Not supported - -[cols="4,1,1,1,1,1,1"] -|=== -| Feature | VeriSimDB | Neo4j | Pinecone | Weaviate | ArangoDB | SurrealDB - -7+h| *Data Modalities* - -| Graph queries (traversals, path finding) -| {check} -| {check} -| {cross} -| {cross} -| {check} -| {check} - -| Vector search (ANN, similarity) -| {check} -| {half} -| {check} -| {check} -| {cross} -| {check} - -| Full-text search -| {check} -| {check} -| {cross} -| {check} -| {check} -| {check} - -| Tensor operations (multi-dimensional) -| {check} -| {cross} -| {cross} -| {cross} -| {cross} -| {cross} - -| Temporal versioning (bitemporal) -| {check} -| {cross} -| {cross} -| {cross} -| {cross} -| {cross} - -| Provenance tracking (lineage) -| {check} -| {cross} -| {cross} -| {cross} -| {cross} -| {cross} - -| Spatial queries (GIS) -| {check} -| {check} -| {cross} -| {cross} -| {check} -| {cross} - -| Semantic/ontological (RDF, knowledge) -| {check} -| {cross} -| {cross} -| {half} -| {cross} -| {cross} - -| Document storage (schemaless JSON) -| {check} -| {half} -| {cross} -| {check} -| {check} -| {check} - -| *Modalities per entity* -| *8* -| *1-2* -| *1* -| *2* -| *3* -| *3* - -7+h| *Cross-Modal Intelligence* - -| Cross-modal drift detection -| {check} -| {cross} -| {cross} -| {cross} -| {cross} -| {cross} - -| Coherence scoring -| {check} -| {cross} -| {cross} -| {cross} -| {cross} -| {cross} - -| Self-normalisation (auto-repair) -| {check} -| {cross} -| {cross} -| {cross} -| {cross} -| {cross} - -| Cross-modal joins -| {check} -| {cross} -| {cross} -| {cross} -| {half} -| {half} - -| Multi-modal consistency guarantees -| {check} -| {cross} -| {cross} -| {cross} -| {cross} -| {cross} - -7+h| *Query Language* - -| Purpose-built query language -| VCL -| Cypher -| N/A (API) -| GraphQL -| AQL -| SurrealQL - -| Dependent-type proofs (compile-time) -| {check} -| {cross} -| {cross} -| {cross} -| {cross} -| {cross} - -| Cross-modal query expressions -| {check} -| {cross} -| {cross} -| {cross} -| {half} -| {half} - -| Formal semantics specification -| {check} -| {half} -| {cross} -| {cross} -| {half} -| {cross} - -7+h| *Federation & Integration* - -| Federation across heterogeneous DBs -| {check} -| {cross} -| {cross} -| {cross} -| {cross} -| {cross} - -| PostgreSQL connector -| {check} -| {cross} -| {cross} -| {cross} -| {cross} -| {cross} - -| ArangoDB connector -| {check} -| {cross} -| {cross} -| {cross} -| N/A -| {cross} - -| Elasticsearch connector -| {check} -| {cross} -| {cross} -| {cross} -| {cross} -| {cross} - -| Neo4j connector -| {check} -| N/A -| {cross} -| {cross} -| {cross} -| {cross} - -7+h| *Engineering & Operations* - -| Open source -| PMPL -| GPLv3 -| Proprietary -| BSD-3-Clause -| Apache 2.0 -| BSL 1.1 - -| Core language -| Rust -| Java -| Proprietary -| Go -| C++ -| Rust - -| Fault tolerance model -| OTP (Elixir) -| JVM -| Managed -| None -| Foxx -| None - -| Managed cloud offering -| No (roadmap) -| Neo4j Aura -| Yes (only) -| Weaviate Cloud -| ArangoDB Oasis -| SurrealDB Cloud - -| SLSA provenance -| {check} -| {cross} -| {cross} -| {cross} -| {cross} -| {cross} - -| SBOM generation -| {check} -| {cross} -| {cross} -| {cross} -| {cross} -| {cross} -|=== - -== Detailed Analysis by Competitor - -=== Neo4j - -*Category:* Graph database (market leader) - -*Strengths:* - -* Mature graph query language (Cypher), widely adopted -* Large ecosystem of tools, drivers, and integrations -* Neo4j Aura managed cloud service -* Strong community and enterprise customer base -* GDS (Graph Data Science) library for analytics - -*Limitations relative to VeriSimDB:* - -* *Single modality:* Graph-only. Vector search is a recent, limited addition - (Neo4j 5.x vector index) without deep integration into the graph model. -* *No drift detection:* No concept of cross-modal consistency. If a graph node - disagrees with its vector embedding, Neo4j has no mechanism to detect this. -* *No temporal versioning:* No built-in bitemporal queries or time-travel. -* *No provenance:* Lineage tracking requires external tooling. -* *No federation:* Cannot federate with other databases. -* *GPLv3 license:* More restrictive than PMPL for embedding in proprietary - products. - -=== Pinecone - -*Category:* Vector database (managed service) - -*Strengths:* - -* Best-in-class vector similarity search performance -* Fully managed, zero-ops deployment -* Simple API for embedding storage and retrieval -* Strong adoption in AI/ML community -* Serverless pricing model - -*Limitations relative to VeriSimDB:* - -* *Single modality:* Vector-only. No graph, document, temporal, or spatial - capabilities. -* *No drift detection:* No awareness of whether embeddings are consistent with - source data or other representations. -* *No query language:* API-only access; no expressive query language for complex - operations. -* *Proprietary:* Closed-source, managed-only. No self-hosted option. -* *No federation:* Isolated from existing data infrastructure. -* *Vendor lock-in:* Data and operations entirely within Pinecone's managed service. - -=== Weaviate - -*Category:* Vector database with object storage - -*Strengths:* - -* Combines vector search with structured object storage -* GraphQL API for flexible querying -* Built-in vectorisation modules (integrates with OpenAI, Cohere, etc.) -* Open source (BSD-3-Clause) -* Good developer experience and documentation - -*Limitations relative to VeriSimDB:* - -* *Two modalities:* Vector + object/document. No graph traversals, tensor - operations, temporal versioning, provenance, or spatial queries. -* *No drift detection:* No mechanism to detect when vector embeddings diverge - from their source objects. -* *No self-normalisation:* No automatic reconciliation of inconsistent data. -* *No formal verification:* No dependent-type proofs for query correctness. -* *No federation:* Cannot connect to existing databases. -* *Limited cross-modal operations:* Vector-object joins are basic; no coherence - scoring. - -=== ArangoDB - -*Category:* Multi-model database - -*Strengths:* - -* Three models in one database: graph, document, key-value -* AQL query language supports cross-model queries -* Mature, production-tested (since 2014) -* Foxx microservices framework for server-side logic -* Apache 2.0 license -* ArangoSearch for full-text search - -*Limitations relative to VeriSimDB:* - -* *Three modalities:* Graph + document + key-value. No vector search, tensor - operations, temporal versioning, provenance, or spatial queries. -* *No drift detection:* ArangoDB can store data in multiple models but has no - mechanism to verify that those models are consistent with each other. A graph - edge and a document describing the same relationship can diverge without - detection. -* *No self-normalisation:* No automatic reconciliation. -* *No formal verification:* AQL is not formally specified; no dependent-type - proofs. -* *No federation:* Cannot connect to external databases. -* *C++ core:* Higher risk of memory safety issues compared to Rust. - -=== SurrealDB - -*Category:* Multi-model database (newer entrant) - -*Strengths:* - -* Modern multi-model: document, graph, and vector in one database -* SurrealQL combines SQL-like syntax with graph and document operations -* Rust core for performance and memory safety -* Record links and graph traversals -* Real-time subscriptions via WebSocket -* Growing community and developer enthusiasm - -*Limitations relative to VeriSimDB:* - -* *Three modalities:* Document + graph + vector (recent). No tensor, temporal, - provenance, or spatial capabilities. -* *No drift detection:* No cross-modal consistency checking despite supporting - multiple models. -* *No self-normalisation:* No automatic repair of divergent data. -* *No formal verification:* SurrealQL has no formal semantics or dependent-type - proofs. -* *No federation:* Cannot connect to external databases. -* *BSL 1.1 license:* Business Source License restricts commercial use until - source-available conversion date. -* *Maturity:* Younger project; production readiness is less established than - ArangoDB or Neo4j. - -== Why Cross-Modal Drift Detection Cannot Be Retrofitted - -The most common question from technical evaluators is: "Why can't Neo4j/ArangoDB/ -SurrealDB just add drift detection as a feature?" - -The answer is architectural. Cross-modal drift detection requires: - -1. *A unified entity model:* Every representation of an entity must be linked at - the storage layer so the database knows they represent the same real-world - thing. Existing multi-model databases store models independently. - -2. *Coherence metrics per modality pair:* The database must define what - "consistency" means between each pair of modalities (graph-vector, graph- - document, vector-temporal, etc.). This is combinatorial: 8 modalities produce - 28 pairwise consistency functions. - -3. *Continuous monitoring infrastructure:* Drift detection cannot be a batch - process. It must run on every write operation to maintain real-time coherence - scores. This requires deep integration into the write path. - -4. *Self-normalisation strategies:* The database must know how to reconcile each - modality pair when drift is detected. This requires domain-specific resolution - logic per modality. - -Retrofitting this onto an existing database would require rewriting the storage -engine, write path, query planner, and replication layer. It is effectively a new -database. - -== Conclusion - -VeriSimDB occupies a unique position in the database landscape. It is not the -fastest graph database (Neo4j), the most scalable vector database (Pinecone), or -the most developer-friendly multi-model database (SurrealDB). Its differentiation -is orthogonal to these axes: *VeriSimDB is the only database that tracks whether -multiple representations of the same entity agree with each other*. - -For organisations whose primary data challenge is cross-system inconsistency -- -which Gartner identifies as the leading cause of data quality costs -- VeriSimDB -addresses the root cause rather than the symptoms. diff --git a/verisimdb/docs/business/marketing/one-pager.adoc b/verisimdb/docs/business/marketing/one-pager.adoc deleted file mode 100644 index 31f58f3a..00000000 --- a/verisimdb/docs/business/marketing/one-pager.adoc +++ /dev/null @@ -1,79 +0,0 @@ -// SPDX-License-Identifier: CC-BY-SA-4.0 -// Copyright (c) 2026 Jonathan D.A. Jewell (hyperpolymath) -= VeriSimDB — Product One-Pager -Jonathan D.A. Jewell -:revdate: 2026-02-28 -:revremark: Product summary for investors, partners, and prospects - -== What Is VeriSimDB? - -*The world's first multimodal database with cross-modal drift detection and -self-normalisation.* VeriSimDB stores every entity across 8 simultaneous -representations and automatically detects when those representations diverge. - -== The Problem - -Data lives in silos. The same customer, product, or transaction exists as a graph -node, a vector embedding, a JSON document, a relational row, and a temporal -snapshot -- all in different systems. When one representation changes, the others -go stale. *Nobody notices until the damage is done.* - -Poor data quality costs enterprises *$12.9 million per year* (Gartner). The root -cause is not dirty data at entry -- it is *cross-system inconsistency*. No existing -database tracks whether multiple representations of the same entity agree with -each other. - -== The Solution - -VeriSimDB introduces the *octad model* -- 8 modalities per entity, with continuous -cross-modal consistency monitoring. - -=== Key Features - -* *8-modality storage* -- Graph, Vector, Tensor, Semantic, Document, Temporal, - Provenance, and Spatial representations in a single database -* *Cross-modal drift detection* -- Continuous coherence scoring detects when - representations diverge, before downstream systems consume inconsistent data -* *Self-normalisation* -- Automatic reconciliation of drifted data using - configurable resolution strategies (alert, quarantine, or repair) -* *VCL query language* -- Purpose-built query language for cross-modal operations - with optional dependent-type proofs (VCL-UT) for compile-time query verification -* *Federation* -- Bring drift detection to existing infrastructure by federating - across PostgreSQL, ArangoDB, Elasticsearch, Neo4j, and more -* *Open source* -- Full-featured community edition under the Palimpsest License - (PMPL-1.0-or-later); no feature gating - -=== Technical Differentiators - -* *Only database with cross-modal drift detection* -- No competitor (Neo4j, - Pinecone, Weaviate, ArangoDB, SurrealDB) tracks cross-modal consistency -* *Formally verified queries* -- VCL-UT provides dependent-type proofs of query - correctness, catching type errors and semantic mismatches before execution -* *Production-grade engineering* -- Rust core + Elixir/OTP orchestration; 662 tests - (510 Rust + 152 Elixir), 0 failures; built for performance and fault tolerance - -== Pricing - -[cols="2,3"] -|=== -| Tier | Description - -| *Community (Free)* -| Full VeriSimDB with all 8 modalities, drift detection, self-normalisation, - VCL, and federation. PMPL-1.0-or-later license. - -| *Commercial Support* -| Enterprise SLA (99.9%), priority security patches, dedicated support channel, - deployment consulting, training. Annual subscription. -|=== - -== Get Started - -* *GitHub:* https://github.com/hyperpolymath/verisimdb[github.com/hyperpolymath/verisimdb] -* *Documentation:* See `docs/getting-started.adoc` in the repository -* *Contact:* j.d.a.jewell@open.ac.uk - -'''' - -_VeriSimDB is developed by Jonathan D.A. Jewell at hyperpolymath._ -_Licensed under PMPL-1.0-or-later (Palimpsest License)._ diff --git a/verisimdb/docs/business/marketing/pitch-deck-outline.adoc b/verisimdb/docs/business/marketing/pitch-deck-outline.adoc deleted file mode 100644 index b1521c6e..00000000 --- a/verisimdb/docs/business/marketing/pitch-deck-outline.adoc +++ /dev/null @@ -1,454 +0,0 @@ -// SPDX-License-Identifier: CC-BY-SA-4.0 -// Copyright (c) 2026 Jonathan D.A. Jewell (hyperpolymath) -= VeriSimDB Pitch Deck — 12-Slide Structure -Jonathan D.A. Jewell -:revdate: 2026-02-28 -:revremark: Complete pitch deck content for investor and partner presentations -:toc: macro -:toclevels: 2 -:sectnums: -:icons: font - -[abstract] -This document contains the full content for a 12-slide pitch deck presenting -VeriSimDB to investors, enterprise partners, and early adopters. Each slide -includes speaker notes, key messages, and visual guidance. - -toc::[] - -== Slide 1: Title - -=== Content - ----- -VeriSimDB -The World's First Multimodal Database with -Cross-Modal Drift Detection - -[Logo Placeholder] - -Jonathan D.A. Jewell -hyperpolymath -j.d.a.jewell@open.ac.uk ----- - -=== Speaker Notes - -Open with the tagline. VeriSimDB is not another database -- it is a new category -of data infrastructure. The word "verisimilitude" means "the appearance of being -true" -- VeriSimDB ensures that all representations of your data maintain -verisimilitude with reality and with each other. - -== Slide 2: The Problem - -=== Content - -*Data Inconsistency Costs Enterprises $12.9 Million Per Year* - -The same entity exists in 5+ systems. When one changes, the others go stale. - ----- -Customer "Jane Doe" - ├── Graph DB: Connected to 47 accounts ← Updated yesterday - ├── Vector DB: Embedding [0.23, 0.87, ...] ← Computed 3 months ago - ├── Document DB: { "address": "123 Old St" } ← Never updated - ├── RDBMS: Row #4821, address = "456 New Ave" ← Updated today - └── Audit Trail: Last change: ??? ← Gap in history ----- - -*Nobody detects this. Nobody repairs it. Downstream systems consume stale data.* - -=== Key Statistics - -* *$12.9M/year* -- Average cost of poor data quality per organisation (Gartner) -* *80%* -- Time data workers spend on cleaning vs. analysis (McKinsey) -* *35%* -- Enterprises citing data quality as primary barrier to AI adoption (IBM) - -=== Speaker Notes - -Walk through the "Jane Doe" example. Point out that each system is individually -correct -- the graph relationships are valid, the vector embedding is -mathematically sound, the document is well-formed. The problem is that they -disagree with each other. No existing tool detects this disagreement. - -== Slide 3: The Solution - -=== Content - -*VeriSimDB: 8 Modalities, One Coherent Entity* - -The *octad model* -- every entity stored across 8 simultaneous representations -with continuous cross-modal consistency monitoring. - ----- -Entity "Jane Doe" — Coherence Score: 0.94 - ├── Graph: 47 connections ✓ Consistent - ├── Vector: [0.23, 0.87, ...] ⚠ Drift: 0.12 - ├── Tensor: Feature matrix updated ✓ Consistent - ├── Semantic: Ontology: Customer ✓ Consistent - ├── Document: { "addr": "456 New" } ✓ Updated - ├── Temporal: Full history, no gaps ✓ Consistent - ├── Provenance: Source: CRM → sync ✓ Traced - └── Spatial: Geo: 51.5074° N ✓ Consistent ----- - -*When the vector embedding drifts, VeriSimDB detects it and triggers -self-normalisation.* - -=== Speaker Notes - -Emphasise the coherence score. This is the key innovation -- a quantified measure -of cross-modal consistency that no other database computes. The octad model is not -about storing 8 copies of data; it is about maintaining 8 complementary -representations that together provide a complete picture of each entity. - -== Slide 4: How It Works - -=== Content - -*Architecture* - ----- - ┌─────────────────────────┐ - │ VCL Engine │ - │ (Query + VCL-UT) │ - └────────────┬────────────┘ - │ - ┌────────────▼────────────┐ - │ Drift Detection Engine │ - │ (Coherence Scoring) │ - └────────────┬────────────┘ - │ - ┌──────────────────────┼──────────────────────┐ - │ │ │ - ┌───────▼───────┐ ┌─────────▼─────────┐ ┌───────▼───────┐ - │ Rust Core │ │ Elixir/OTP Orch. │ │ ReScript Reg. │ - │ (Storage + │ │ (Concurrency + │ │ (Type-safe │ - │ Indexing) │ │ Fault Tolerance) │ │ Registry) │ - └───────┬───────┘ └─────────┬─────────┘ └───────────────┘ - │ │ - ┌───────▼──────────────────────▼───────┐ - │ Octad Storage Layer │ - │ [Graph][Vector][Tensor][Semantic] │ - │ [Document][Temporal][Provenance] │ - │ [Spatial] │ - └──────────────────────────────────────┘ - │ - ┌───────▼──────────────────────────────┐ - │ Federation Layer │ - │ PostgreSQL │ ArangoDB │ Elastic │ … │ - └──────────────────────────────────────┘ ----- - -=== Speaker Notes - -Three layers: the VCL query engine parses and type-checks queries; the drift -detection engine continuously monitors coherence across modalities; the storage -layer handles the octad per entity. Federation allows VeriSimDB to bring drift -detection to existing databases without data migration. Rust provides performance, -Elixir/OTP provides fault tolerance and concurrency, ReScript provides a type-safe -registry for modality metadata. - -== Slide 5: Demo — Drift Detection in Action - -=== Content - -*Before: Traditional Multi-Model Query* - -[source,sql] ----- --- ArangoDB: Get customer graph + document -FOR c IN customers - FILTER c._key == "jane-doe" - FOR rel IN OUTBOUND c customer_accounts - RETURN { customer: c, accounts: rel } - --- Separately: Pinecone vector search -pinecone.query(vector=[0.23, 0.87, ...], top_k=10) - --- No way to know if these agree! ----- - -*After: VeriSimDB Cross-Modal Query with Drift Detection* - -[source] ----- --- VCL: Query across modalities with coherence check -SELECT entity, coherence_score, drifted_modalities -FROM entities -WHERE entity.id = "jane-doe" - AND modalities IN (graph, vector, document) - AND coherence_score > 0.90 - --- Result: --- entity: jane-doe --- coherence_score: 0.82 ⚠ BELOW THRESHOLD --- drifted_modalities: [vector] --- action: self-normalise triggered ----- - -=== Speaker Notes - -Walk through both examples. In the "before" scenario, a developer queries two -separate databases and has no way to know if the results are consistent. In the -"after" scenario, VeriSimDB's VCL query spans modalities and includes the -coherence score directly in the result set. When the score drops below the -configured threshold, self-normalisation is triggered automatically. The developer -does not need to write reconciliation logic. - -== Slide 6: Technology - -=== Content - -*Built for Production* - -[cols="2,3"] -|=== -| Component | Technology - -| Core storage + indexing -| *Rust* -- Near-C performance, memory safety without GC - -| Orchestration + fault tolerance -| *Elixir/OTP* -- Battle-tested concurrency model (Erlang VM) - -| Type-safe registry -| *ReScript* -- ML-family type system compiling to JavaScript - -| Query language -| *VCL* -- Purpose-built for cross-modal operations - -| Formal verification -| *VCL-UT* -- Dependent-type proofs of query correctness - -| ABI definitions -| *Idris2* -- Dependent types for interface verification - -| FFI implementation -| *Zig* -- C-compatible, zero-overhead foreign function interface -|=== - -*Test Suite:* 510 Rust tests + 152 Elixir tests = *662 total, 0 failures* - -=== Speaker Notes - -Emphasise the engineering quality. Rust + Elixir is an increasingly popular -combination for data infrastructure (used by Discord, Fly.io, and others). The -test suite is comprehensive and runs in CI on every commit. VCL-UT's dependent -types are optional -- teams can start with plain VCL and adopt formal verification -incrementally. - -== Slide 7: Market Opportunity - -=== Content - -*$51M Serviceable Obtainable Market* - -[cols="2,1,1,1"] -|=== -| Segment | TAM | SAM | SOM - -| Data Quality Tools -| $2.3B -| $460M -| $23M - -| Graph Databases -| $5.6B -| $280M -| $14M - -| Vector Databases -| $1.2B -| $120M -| $6M - -| Multimodal Databases -| $800M -| $160M -| $8M - -| *Combined* -| -| -| *$51M* -|=== - -*VeriSimDB creates a new category at the intersection of these markets.* - -=== Speaker Notes - -The SOM estimates are deliberately conservative (1% of SAM). VeriSimDB sits at -the intersection of four high-growth markets, each expanding at 15-25% CAGR. -The multimodal database segment is nascent -- there is an opportunity to define -the category rather than compete within an established one. - -== Slide 8: Business Model - -=== Content - -*Open Core: Full-Featured Free + Commercial Support* - ----- -┌─────────────────────────────────────────────┐ -│ Community Edition (PMPL) │ -│ All 8 modalities, drift detection, │ -│ self-normalisation, VCL, federation │ -│ FREE — forever │ -├─────────────────────────────────────────────┤ -│ Commercial Support License │ -│ ├── Enterprise SLA (99.9% uptime) │ -│ ├── Priority security patches (24hr) │ -│ ├── Dedicated support channel │ -│ ├── Deployment consulting │ -│ └── Training & certification │ -│ Annual subscription: $40K ACV │ -├─────────────────────────────────────────────┤ -│ OEM License │ -│ Embed VeriSimDB in your product │ -│ Per-deployment or flat annual fee │ -└─────────────────────────────────────────────┘ ----- - -*No feature gating. Revenue from operational value, not artificial restrictions.* - -=== Speaker Notes - -This model is inspired by successful open-core companies (GitLab, Elastic, -Confluent). The community edition is intentionally full-featured because feature -gating erodes developer trust. Enterprise customers pay for operational value: -SLAs, priority patches, and deployment expertise. This creates a large adoption -funnel and a natural conversion path. - -== Slide 9: Competitive Landscape - -=== Content - -*Feature Comparison Matrix* - -[cols="3,1,1,1,1,1,1"] -|=== -| Feature | VeriSimDB | Neo4j | Pinecone | Weaviate | ArangoDB | SurrealDB - -| Graph queries | * | * | | | * | * -| Vector search | * | | * | * | | * -| Tensor operations | * | | | | | -| Temporal versioning| * | | | | | -| Provenance tracking| * | | | | | -| Spatial queries | * | * | | | * | -| *Drift detection* | * | | | | | -| *Self-normalisation* | * | | | | | -| *Formal proofs* | * | | | | | -| Federation | * | | | | | -| Modalities/entity | 8 | 1 | 1 | 2 | 3 | 3 -| Open source | PMPL | GPLv3 | No | BSD-3 | Apache | BSL -|=== - -_* = supported_ - -*VeriSimDB is the only database in any category with cross-modal drift detection.* - -=== Speaker Notes - -Walk through the matrix column by column. The key insight is not that VeriSimDB -does everything -- it is that VeriSimDB does something no one else does: cross-modal -consistency monitoring. This is not a feature that can be easily added to existing -architectures; it requires the data model to be designed around multi-modal -coherence from the ground up. - -== Slide 10: Traction - -=== Content - -*Engineering Maturity* - -* *662 tests* (510 Rust + 152 Elixir), *0 failures* -* *8 modalities* fully architected and implemented -* *VCL grammar* specified in EBNF with formal semantics -* *VCL-UT* dependent-type layer with Idris2 proofs -* *Federation connectors* for PostgreSQL, ArangoDB, Elasticsearch -* *Comprehensive documentation:* Architecture, deployment, query language, - drift handling, error handling, safety theory - -*Open Source Readiness* - -* PMPL-1.0-or-later license -* CONTRIBUTING.md, CODE_OF_CONDUCT.md, SECURITY.md -* CI/CD pipeline with Hypatia neurosymbolic scanning -* SBOM generation, SLSA provenance, OpenSSF Scorecard - -=== Speaker Notes - -VeriSimDB is not a whitepaper or a prototype. The test suite, documentation, and -engineering infrastructure are at a level that most open-source databases do not -reach until much later. The 0-failure test suite across 662 tests demonstrates -engineering discipline. The formal semantics and dependent-type proofs are unique -in the database industry. - -== Slide 11: Team - -=== Content - -*Jonathan D.A. Jewell* - -* *Role:* Creator, architect, and lead developer -* *Affiliation:* The Open University (UK) -* *GitHub:* https://github.com/hyperpolymath[hyperpolymath] -* *Expertise:* Multimodal data systems, formal methods, programming language - theory, neurosymbolic AI, database internals - -*hyperpolymath Ecosystem* - -* *265+ repositories* spanning databases, programming languages, developer tools, - formal verification, and AI infrastructure -* *Related projects:* QuandleDB (dependent-type graph database), LithoGlyph - (provenance database), PanLL (neurosymbolic mission control) -* *Community:* gitbot-fleet (6 automated bots), Hypatia CI/CD intelligence, - Rhodium Standard Repository (RSR) compliance framework - -=== Speaker Notes - -VeriSimDB is part of a larger ecosystem of data infrastructure projects, all -designed to interoperate. The breadth of the ecosystem demonstrates deep expertise -in database internals, formal methods, and production engineering. The Open -University affiliation provides academic rigour and access to research -collaborations. - -== Slide 12: The Ask - -=== Content - -*What We Are Looking For* - -[cols="1,3"] -|=== -| Contributors -| Rust, Elixir, and ReScript developers interested in multimodal data systems, - formal methods, or database internals. VeriSimDB is open source -- contributions - are welcome and credited. - -| Early Adopters -| Organisations with multimodal data challenges (GraphRAG, biomedical, financial - services, supply chain) willing to pilot VeriSimDB and provide feedback. - -| Partnerships -| Database vendors, system integrators, and cloud providers interested in - federation connectors or OEM licensing. - -| Academic Collaborators -| Researchers in database theory, formal methods, data quality, or multimodal - AI interested in joint publications or research grants. -|=== - -*Get Started:* - -* GitHub: https://github.com/hyperpolymath/verisimdb -* Email: j.d.a.jewell@open.ac.uk -* License: PMPL-1.0-or-later (full-featured, free forever) - -=== Speaker Notes - -End with a clear call to action. VeriSimDB is not looking for traditional VC -funding at this stage -- it is looking for community, adoption, and partnerships. -The open-core model means the product grows through usage, and commercial revenue -follows adoption naturally. diff --git a/verisimdb/docs/business/marketing/use-cases.adoc b/verisimdb/docs/business/marketing/use-cases.adoc deleted file mode 100644 index 85566704..00000000 --- a/verisimdb/docs/business/marketing/use-cases.adoc +++ /dev/null @@ -1,522 +0,0 @@ -// SPDX-License-Identifier: CC-BY-SA-4.0 -// Copyright (c) 2026 Jonathan D.A. Jewell (hyperpolymath) -= VeriSimDB Use Cases — 7 Target Domains -Jonathan D.A. Jewell -:revdate: 2026-02-28 -:revremark: Detailed use case descriptions for target market segments -:toc: macro -:toclevels: 2 -:sectnums: -:icons: font - -[abstract] -This document describes seven target domains where VeriSimDB's unique -cross-modal drift detection and multimodal storage capabilities address -critical, currently unsolved data infrastructure challenges. Each use case -includes the domain context, the specific problem, how VeriSimDB solves it, -example queries, and the business value. - -toc::[] - -== Use Case 1: GraphRAG / AI Knowledge Graphs - -=== Domain Context - -Retrieval-Augmented Generation (RAG) enhances large language models (LLMs) by -retrieving relevant context from external knowledge bases before generating -responses. *GraphRAG* extends this by using knowledge graphs rather than flat -document chunks, enabling structured reasoning over entity relationships. - -GraphRAG systems typically maintain: - -* A *knowledge graph* (entities and relationships) -* *Vector embeddings* of entity descriptions and document chunks -* *Source documents* from which the graph was extracted -* *Temporal metadata* tracking when entities were extracted or updated - -=== The Problem - -GraphRAG systems suffer from a critical but invisible failure mode: *embedding -drift*. When the knowledge graph is updated (new entities, changed relationships), -the vector embeddings computed from the old graph state are not re-computed. The -graph says one thing; the vectors say another. LLM retrieval uses the stale -embeddings, producing answers that contradict the current graph state. - -No existing GraphRAG framework detects this. Developers discover the problem only -when an end user reports a hallucinated or contradictory response. - -=== How VeriSimDB Solves It - -VeriSimDB stores knowledge graph entities with their graph relationships, vector -embeddings, source documents, and temporal metadata as a single octad entity. When -a graph relationship changes, VeriSimDB's drift detection engine flags that the -vector embedding's coherence score has dropped below the configured threshold. - -The system can then: - -* *Alert* the GraphRAG pipeline to re-embed the affected entities -* *Quarantine* the drifted entities so they are excluded from retrieval until - re-embedded -* *Self-normalise* by triggering an automatic re-embedding pipeline - -=== Example VCL Query - -[source] ----- --- Find entities where graph and vector representations have drifted -SELECT entity.id, entity.graph.relationships, coherence(graph, vector) -FROM knowledge_base -WHERE coherence(graph, vector) < 0.85 -ORDER BY coherence(graph, vector) ASC -LIMIT 100 ----- - -=== Business Value - -* *Eliminate silent hallucinations* caused by stale embeddings -* *Reduce manual debugging time* for GraphRAG pipeline engineers -* *Increase retrieval accuracy* by ensuring graph-vector consistency -* *Enable real-time GraphRAG updates* with confidence that embeddings are current - -== Use Case 2: Biomedical Research Data - -=== Domain Context - -Biomedical research generates data across many modalities: patient records -(documents), drug interaction networks (graphs), genomic feature tensors, spatial -pathology images, temporal clinical trial timelines, and provenance chains -tracking data from sample collection through analysis to publication. - -Regulatory requirements (FDA 21 CFR Part 11, EU MDR, GDPR) mandate complete -traceability and data integrity. A single inconsistency between a clinical trial -result and its provenance chain can trigger a regulatory audit. - -=== The Problem - -Biomedical data integration is currently achieved through bespoke ETL pipelines -that move data between specialised databases (Neo4j for drug interactions, -PostgreSQL for clinical records, image databases for pathology, etc.). These -pipelines are brittle, slow, and have no mechanism to verify that a patient -entity's graph representation (drug interactions) is consistent with their -document representation (clinical notes) or their temporal representation (trial -timeline). - -When a clinical trial protocol amendment changes a patient's treatment, the graph, -document, and timeline representations must all be updated. If the graph update -succeeds but the document update fails, the patient's record is inconsistent, and -no system detects this until a manual audit. - -=== How VeriSimDB Solves It - -VeriSimDB stores each biomedical entity (patient, drug, gene, trial) with all its -representations in a single octad. Drift detection continuously monitors -consistency across modalities. When a protocol amendment updates one representation, -VeriSimDB verifies that all other representations are updated within the configured -consistency window. - -The provenance modality provides an immutable audit trail of every change, who made -it, and from what source -- satisfying regulatory traceability requirements. - -=== Example VCL Query - -[source] ----- --- Find patients whose clinical notes disagree with their trial timeline -SELECT patient.id, - coherence(document, temporal) AS doc_time_coherence, - provenance.last_update -FROM clinical_trial -WHERE trial.id = "NCT-2026-0042" - AND coherence(document, temporal) < 0.95 ----- - -=== Business Value - -* *Regulatory compliance:* Continuous consistency monitoring satisfies FDA/EU MDR - data integrity requirements without manual audits -* *Patient safety:* Prevent treatment decisions based on inconsistent records -* *Audit readiness:* Provenance modality provides complete, immutable lineage -* *Reduced integration cost:* Replace bespoke ETL pipelines with federated - drift detection - -== Use Case 3: Financial Risk and Compliance - -=== Domain Context - -Financial institutions maintain customer and transaction data across dozens of -systems: core banking (relational), fraud detection (graph), risk scoring -(vector/ML), regulatory reporting (documents), transaction history (temporal), -and geographic compliance (spatial -- sanctions, jurisdiction). - -Regulations including Basel III, MiFID II, AML directives, and KYC requirements -demand that customer data be consistent across all systems at all times. A -customer's risk score in the ML system must agree with their transaction graph in -the fraud detection system and their geographic profile in the compliance system. - -=== The Problem - -Financial data inconsistency creates two categories of risk: - -1. *Regulatory risk:* If a customer's risk classification differs between the AML - system and the core banking system, regulators may impose fines for inadequate - controls. Fines for AML violations have exceeded $10 billion globally since 2020. - -2. *Operational risk:* If a customer's transaction graph shows suspicious activity - but their risk score has not been updated, fraud goes undetected. Conversely, if - the risk score is elevated but the graph is stale, legitimate customers are - blocked. - -Current solutions rely on periodic batch reconciliation (nightly or weekly), which -means inconsistencies persist for hours or days. - -=== How VeriSimDB Solves It - -VeriSimDB stores each customer entity with graph (relationships, transaction -networks), vector (risk embeddings), document (KYC records), temporal (transaction -history), provenance (data lineage for audit), and spatial (jurisdiction, sanctions -geography) representations. Drift detection runs in real time, flagging any -customer whose representations disagree. - -Federation allows VeriSimDB to connect to the existing core banking, fraud -detection, and compliance systems, bringing cross-modal consistency checking to -current infrastructure without data migration. - -=== Example VCL Query - -[source] ----- --- Find customers with inconsistent risk profiles across modalities -SELECT customer.id, - coherence(vector, graph) AS risk_graph_coherence, - coherence(document, spatial) AS kyc_geo_coherence, - temporal.last_transaction, - provenance.source_systems -FROM customers -WHERE coherence(vector, graph) < 0.90 - OR coherence(document, spatial) < 0.90 -ORDER BY LEAST(coherence(vector, graph), coherence(document, spatial)) ASC ----- - -=== Business Value - -* *Regulatory compliance:* Real-time cross-system consistency satisfies AML/KYC - requirements with auditable coherence scores -* *Fraud detection:* Eliminate the window between graph update and risk score - update where fraud can go undetected -* *Reduced fines:* Continuous monitoring vs. periodic batch reconciliation - demonstrates "reasonable controls" to regulators -* *Audit efficiency:* Provenance modality provides complete lineage for every - customer data element - -== Use Case 4: Supply Chain Traceability - -=== Domain Context - -Modern supply chains involve hundreds of entities (suppliers, manufacturers, -distributors, retailers) connected by complex logistics networks. Regulations -including the EU Digital Product Passport, the US FSMA (Food Safety Modernization -Act), and ESG reporting requirements demand end-to-end traceability. - -Supply chain data is inherently multimodal: - -* *Graph:* Supplier-manufacturer-distributor relationships -* *Document:* Purchase orders, invoices, certificates of origin -* *Temporal:* Shipment timelines, temperature logs, shelf-life tracking -* *Spatial:* Geographic location of goods in transit -* *Provenance:* Chain of custody, certifications, audits - -=== The Problem - -Supply chain transparency initiatives fail because the data is fragmented across -disconnected systems. A manufacturer's certificate of origin (document) may list -materials from a compliant supplier, but the supplier relationship graph has not -been updated to reflect that the supplier's certification expired. The geographic -tracking shows goods in transit through a sanctioned region, but the compliance -document still shows the old routing. - -These inconsistencies are invisible until a physical audit, a product recall, or -a regulatory investigation. - -=== How VeriSimDB Solves It - -VeriSimDB's octad model stores each supply chain entity (product, shipment, -supplier) with all its representations. When a supplier's certification expires, -the document modality is updated, and drift detection immediately flags that the -graph modality (which still shows the supplier as "certified") is inconsistent. - -The spatial modality tracks goods in real time and cross-checks against compliance -documents. The temporal modality provides a complete timeline of every change. The -provenance modality establishes an immutable chain of custody. - -=== Example VCL Query - -[source] ----- --- Find products whose supplier certifications have drifted -SELECT product.id, - graph.supplier, - document.certification_status, - coherence(graph, document) AS cert_coherence, - spatial.current_location, - temporal.last_audit -FROM supply_chain -WHERE coherence(graph, document) < 0.90 - AND temporal.last_audit < NOW() - INTERVAL '90 days' ----- - -=== Business Value - -* *Regulatory compliance:* Continuous traceability for EU Digital Product Passport, - FSMA, ESG reporting -* *Risk reduction:* Detect expired certifications, non-compliant routing, and - chain-of-custody gaps before they cause product recalls -* *Audit efficiency:* Provenance modality eliminates manual chain-of-custody - reconstruction -* *Consumer trust:* Verified, consistent supply chain data for end-to-end - transparency - -== Use Case 5: Digital Humanities and Cultural Heritage - -=== Domain Context - -Cultural heritage institutions (museums, archives, libraries) manage collections -of artefacts that are described across multiple modalities: - -* *Graph:* Relationships between artefacts, creators, historical periods, and - exhibition histories -* *Document:* Catalogue records, conservation reports, scholarly annotations -* *Semantic:* Ontological classification (e.g., CIDOC-CRM, Dublin Core) -* *Temporal:* Provenance timelines (creation, ownership history, exhibition dates) -* *Spatial:* Geographic origin, current location, excavation site coordinates -* *Vector:* Image embeddings for visual similarity search -* *Provenance:* Acquisition history, legal ownership chain - -=== The Problem - -Cultural heritage data is notoriously inconsistent. A catalogue record (document) -may describe an artefact as "Roman, 2nd century CE" while the ontological -classification (semantic) says "Hellenistic" and the temporal timeline shows a -creation date in the 3rd century BCE. These inconsistencies arise from decades of -cataloguing by different scholars with different conventions, and no system -detects them. - -Digital humanities researchers spend enormous amounts of time manually -reconciling conflicting descriptions across catalogues, ontologies, and timelines -before they can begin analysis. - -=== How VeriSimDB Solves It - -VeriSimDB stores each artefact with all its modalities in a single octad entity. -Drift detection flags inconsistencies between catalogue records and ontological -classifications, between temporal timelines and scholarly annotations, and between -image embeddings and textual descriptions. - -Federation allows VeriSimDB to connect to existing collection management systems -(MuseumPlus, TMS, ArchivesSpace) and bring cross-modal consistency checking to -existing catalogues without requiring data migration. - -=== Example VCL Query - -[source] ----- --- Find artefacts with inconsistent date attributions across modalities -SELECT artefact.id, - document.period_attribution, - semantic.ontology_class, - temporal.creation_date, - coherence(document, semantic, temporal) AS date_coherence -FROM collection -WHERE coherence(document, semantic, temporal) < 0.80 -ORDER BY date_coherence ASC ----- - -=== Business Value - -* *Research efficiency:* Eliminate manual reconciliation of conflicting catalogue - records, freeing researchers for analysis -* *Data quality:* Systematically identify and resolve decades of cataloguing - inconsistencies -* *Cross-institutional collaboration:* Federation enables consistency checking - across multiple institutions' collections -* *Grant competitiveness:* Demonstrable data quality infrastructure strengthens - research grant applications - -== Use Case 6: Scientific Data Management - -=== Domain Context - -Scientific research produces multimodal datasets: experimental measurements -(tensors), instrument metadata (documents), collaboration networks (graphs), -publication timelines (temporal), methodological lineage (provenance), and -geographic data (field sites, observatories). - -Reproducibility is the cornerstone of the scientific method, yet the -"reproducibility crisis" affects 52% of scientific fields (Baker, Nature 2016). -A significant contributor is data inconsistency between published results and -underlying datasets. - -=== The Problem - -A published paper (document) reports results derived from a dataset (tensor). -When the dataset is corrected (e.g., an instrument calibration error is -discovered), the tensor is updated, but the published paper is not retracted or -amended. The methodology provenance chain shows the old calibration, but the -instrument metadata document now shows the corrected calibration. These -representations have drifted, and no system detects it. - -Similarly, when a collaboration network changes (a co-author retracts their -contribution), the graph representation must be updated, but the document -(published paper) and provenance (author contribution records) may not be. - -=== How VeriSimDB Solves It - -VeriSimDB stores each scientific entity (dataset, publication, experiment, -researcher) with all its modalities. When a dataset is corrected, drift detection -flags that the published paper's document representation is now inconsistent with -the tensor (corrected data) and provenance (updated methodology). - -This does not automatically retract papers -- it surfaces the inconsistency so -that researchers, journal editors, and data stewards can take appropriate action. - -=== Example VCL Query - -[source] ----- --- Find publications whose cited datasets have been corrected -SELECT publication.doi, - tensor.dataset_version, - provenance.calibration_date, - coherence(document, tensor, provenance) AS data_coherence -FROM research_outputs -WHERE coherence(document, tensor, provenance) < 0.90 - AND tensor.dataset_version > document.cited_version ----- - -=== Business Value - -* *Reproducibility:* Systematically detect when published results no longer match - their underlying data -* *Research integrity:* Provenance modality provides complete methodological - lineage -* *Data management compliance:* Satisfies FAIR data principles (Findable, - Accessible, Interoperable, Reusable) with cross-modal consistency -* *Institutional reputation:* Proactive detection of data inconsistencies before - external challenges - -== Use Case 7: Cybersecurity Threat Intelligence - -=== Domain Context - -Cybersecurity threat intelligence (CTI) involves aggregating and correlating -indicators of compromise (IoCs), threat actor profiles, attack patterns, and -vulnerability data from multiple sources. This data is inherently multimodal: - -* *Graph:* Threat actor relationships, attack infrastructure networks, malware - family trees -* *Vector:* Similarity embeddings for malware samples, phishing emails, network - traffic patterns -* *Document:* Threat reports, vulnerability advisories (CVEs), incident reports -* *Temporal:* Attack timelines, vulnerability disclosure dates, patch release - schedules -* *Provenance:* Intelligence source attribution, confidence levels, corroboration - chains -* *Spatial:* Geographic attribution, IP geolocation, infrastructure mapping - -=== The Problem - -CTI platforms aggregate intelligence from dozens of feeds (MITRE ATT&CK, STIX/TAXII, -open-source feeds, commercial feeds, internal detections). The same threat actor or -malware family may be described differently across sources. When one feed updates -a threat actor's tactics, techniques, and procedures (TTPs), the graph is updated, -but the vector embeddings (used for similarity matching) still reflect the old -behaviour profile. Temporal timelines may show campaign activity that contradicts -the current graph state. - -These inconsistencies cause false negatives (real threats are missed because -embeddings are stale) and false positives (old indicators trigger alerts for -threats that have evolved). - -=== How VeriSimDB Solves It - -VeriSimDB stores each CTI entity (threat actor, malware, vulnerability, campaign) -with all its modalities. When a STIX feed updates a threat actor's TTPs, drift -detection flags that the vector embeddings and document representations have not -been updated to reflect the new behaviour. - -Federation allows VeriSimDB to connect to existing SIEM, SOAR, and TIP platforms, -bringing cross-modal consistency to current CTI infrastructure. - -=== Example VCL Query - -[source] ----- --- Find threat actors whose behaviour profiles have drifted across sources -SELECT actor.id, - graph.ttp_count, - vector.similarity_cluster, - temporal.last_campaign, - coherence(graph, vector, document) AS intel_coherence, - provenance.source_feeds -FROM threat_actors -WHERE coherence(graph, vector, document) < 0.85 - AND temporal.last_campaign > NOW() - INTERVAL '30 days' -ORDER BY intel_coherence ASC ----- - -=== Business Value - -* *Reduced false negatives:* Ensure detection rules are based on current, consistent - threat intelligence -* *Reduced false positives:* Stale indicators are flagged and updated, reducing - alert fatigue -* *Source correlation:* Provenance modality tracks which feeds contributed to each - intelligence element and whether they agree -* *Audit trail:* Complete lineage of every intelligence assessment for incident - response and post-mortem analysis - -== Cross-Domain Summary - -[cols="1,3,3,2"] -|=== -| Domain | Primary Modalities | Key Drift Risk | Business Impact - -| GraphRAG / AI -| Graph, Vector, Document -| Embedding diverges from graph state -| Silent hallucinations - -| Biomedical -| Document, Temporal, Provenance -| Clinical record inconsistency -| Patient safety, regulatory fines - -| Financial -| Vector, Graph, Spatial -| Risk score vs. transaction graph -| Regulatory fines, fraud exposure - -| Supply Chain -| Graph, Document, Spatial -| Expired certifications undetected -| Product recalls, compliance failures - -| Cultural Heritage -| Document, Semantic, Temporal -| Conflicting catalogue attributions -| Research wasted on reconciliation - -| Scientific -| Document, Tensor, Provenance -| Published results vs. corrected data -| Reproducibility failures - -| Cybersecurity -| Graph, Vector, Document -| Stale threat behaviour profiles -| False negatives, missed threats -|=== - -All seven domains share the same fundamental problem: *data exists in multiple -representations that inevitably diverge, and no existing system detects the -divergence*. VeriSimDB's cross-modal drift detection addresses this root cause -across every domain where multimodal data consistency matters. diff --git a/verisimdb/docs/business/pr/faq.adoc b/verisimdb/docs/business/pr/faq.adoc deleted file mode 100644 index 8df3df76..00000000 --- a/verisimdb/docs/business/pr/faq.adoc +++ /dev/null @@ -1,316 +0,0 @@ -// SPDX-License-Identifier: CC-BY-SA-4.0 -// Copyright (c) 2026 Jonathan D.A. Jewell (hyperpolymath) -= VeriSimDB — Frequently Asked Questions -Jonathan D.A. Jewell -:revdate: 2026-02-28 -:revremark: Comprehensive FAQ for press, analysts, developers, and evaluators -:toc: macro -:toclevels: 2 -:sectnums: -:icons: font - -[abstract] -Answers to the most frequently asked questions about VeriSimDB, covering -product fundamentals, technical architecture, competitive positioning, licensing, -and roadmap. Intended for press, analysts, developers, enterprise evaluators, and -potential contributors. - -toc::[] - -== Product Fundamentals - -=== What is VeriSimDB? - -VeriSimDB is the world's first multimodal database with cross-modal drift -detection and self-normalisation. It stores every entity across up to 8 -simultaneous data representations -- graph, vector, tensor, semantic, document, -temporal, provenance, and spatial (the "octad") -- and continuously monitors -whether those representations are consistent with each other. - -When representations diverge (a condition called "drift"), VeriSimDB detects the -divergence, quantifies it with a coherence score, and can automatically -reconcile the data through configurable self-normalisation strategies. - -=== What problem does VeriSimDB solve? - -The same real-world entity (a customer, a product, a transaction, a research -dataset) is typically represented across 5-10 different systems: a graph database, -a vector index, a document store, a relational database, a temporal audit trail, -and so on. When one representation changes, the others go stale. No existing -database or tool detects this cross-system inconsistency. - -According to Gartner, poor data quality -- primarily caused by cross-system -inconsistency -- costs enterprises an average of $12.9 million per year. VeriSimDB -addresses this by making cross-modal consistency a first-class database feature. - -=== What is the octad model? - -The octad model is VeriSimDB's approach to multimodal data representation. Every -entity can be simultaneously represented across 8 modalities: - -1. *Graph* -- Relationships, connections, topology -2. *Vector* -- Semantic similarity, embeddings, nearest-neighbour search -3. *Tensor* -- Multi-dimensional numerical data, ML feature stores -4. *Semantic* -- Ontological meaning, RDF triples, knowledge representation -5. *Document* -- Schemaless structured data (JSON/BSON) -6. *Temporal* -- Time-versioned state, bitemporal queries, audit trails -7. *Provenance* -- Origin tracking, transformation lineage, trust chains -8. *Spatial* -- Geographic and geometric data, GIS operations - -Not every entity needs all 8 modalities. The octad defines the maximum set; each -entity uses the modalities relevant to its domain. - -=== What is drift detection? - -Drift detection is VeriSimDB's core innovation. It continuously monitors the -consistency of an entity's representations across all active modalities. For -each entity, VeriSimDB computes a *coherence score* -- a quantified measure of -how well its representations agree with each other. - -For example, if a customer's graph representation shows 47 account connections but -their vector embedding was computed when they had 12 connections, the graph and -vector modalities have "drifted." VeriSimDB detects this divergence, reports a -reduced coherence score, and can take corrective action. - -The coherence score is not a binary pass/fail -- it is a continuous value that -reflects the degree of consistency across all active modalities. Configurable -thresholds determine when drift triggers alerts, quarantine, or self-normalisation. - -=== What is self-normalisation? - -Self-normalisation is VeriSimDB's ability to automatically repair drifted data. -When drift is detected and exceeds a configured threshold, VeriSimDB can: - -* *Alert:* Notify downstream systems that data may be inconsistent, without - modifying the data -* *Quarantine:* Mark the entity as potentially drifted, excluding it from query - results until a human or automated process resolves the inconsistency -* *Self-normalise:* Automatically reconcile the representations using - configurable resolution strategies (e.g., "prefer the most recently updated - modality" or "prefer the modality with the highest confidence provenance") - -Self-normalisation strategies are fully configurable per entity type and per -modality pair. Organisations can start with alert-only and gradually adopt -automated repair as they build confidence. - -== Competitive Positioning - -=== How is VeriSimDB different from Neo4j? - -Neo4j is a graph database -- it excels at graph traversals, path finding, and -relationship queries. However, Neo4j stores data in a single modality (graph). -It has recently added limited vector search capabilities, but there is no -cross-modal consistency checking between the graph and vector representations. - -VeriSimDB stores entities across 8 modalities and continuously monitors their -consistency. If a graph node's relationships change but its vector embedding is -not re-computed, VeriSimDB detects the drift. Neo4j does not. - -=== How is VeriSimDB different from Pinecone? - -Pinecone is a managed vector database optimised for similarity search. It stores -vector embeddings only -- no graphs, documents, temporal data, or provenance. It -is a proprietary, cloud-only service with no self-hosted option. - -VeriSimDB includes vector search as one of 8 modalities and detects when vector -embeddings diverge from other representations. It is open source, self-hostable, -and designed for organisations that need more than similarity search. - -=== How is VeriSimDB different from Weaviate? - -Weaviate combines vector search with object/document storage and provides -GraphQL-based querying. It supports 2 modalities (vector + document/object) but -has no drift detection, no temporal versioning, no provenance tracking, no formal -query verification, and no federation. - -VeriSimDB supports 8 modalities with cross-modal drift detection and formal query -proofs (VCL-UT). The fundamental difference is architectural: Weaviate treats -vectors and objects as complementary storage; VeriSimDB treats all modalities as -representations of a single entity that must remain consistent. - -=== How is VeriSimDB different from ArangoDB? - -ArangoDB is a multi-model database supporting graph, document, and key-value -models in a single database. It has a mature query language (AQL) and a strong -production track record. - -However, ArangoDB has no mechanism to verify that its graph, document, and -key-value representations of the same entity are consistent. If a graph edge -and a document describing the same relationship diverge, ArangoDB does not -detect this. VeriSimDB does. - -ArangoDB also lacks vector search, tensor operations, temporal versioning, -provenance tracking, and spatial queries -- all of which are native VeriSimDB -modalities. - -=== How is VeriSimDB different from SurrealDB? - -SurrealDB is a newer multi-model database supporting document, graph, and (recently) -vector models with a modern query language (SurrealQL). It is built in Rust and -has strong developer ergonomics. - -Like ArangoDB, SurrealDB stores multiple models but does not check consistency -between them. It has no drift detection, no self-normalisation, no formal query -verification, and no federation. Its Business Source License (BSL 1.1) restricts -commercial use, while VeriSimDB's PMPL license is permissive. - -== Technical Architecture - -=== What languages is VeriSimDB built in? - -* *Rust:* Core storage engine, indexing, drift detection algorithms. Provides - near-C performance with memory safety guarantees. -* *Elixir/OTP:* Orchestration layer, concurrency management, fault tolerance. - The Erlang VM (BEAM) provides battle-tested supervision trees and hot code - reloading. -* *ReScript:* Type-safe registry for modality metadata. Compiles to JavaScript - with an ML-family type system. -* *Idris2:* ABI (Application Binary Interface) definitions with dependent-type - proofs for interface correctness. -* *Zig:* FFI (Foreign Function Interface) implementation for C-compatible - bindings. - -=== What is VCL? - -VCL (VeriSim Consonance Language) is VeriSimDB's purpose-built query language -for cross-modal operations. VCL allows queries to span multiple modalities in -a single expression, include coherence scores in result sets, and filter by -drift thresholds. - -VCL is designed to be familiar to SQL users while supporting modality-specific -operations (graph traversals, vector similarity, tensor operations, temporal -time-travel, spatial containment, etc.) in a unified syntax. - -=== What are dependent-type proofs (VCL-UT)? - -VCL-UT is an optional extension to VCL that adds dependent-type verification. -When enabled, VCL-UT uses the Idris2 type system to prove properties of queries -at compile time: - -* *Type correctness:* Ensure that query expressions are well-typed across - modalities (e.g., a vector similarity comparison is applied to actual vector - fields, not document fields) -* *Semantic compatibility:* Verify that cross-modal joins are semantically - meaningful -* *Threshold validity:* Prove that coherence thresholds are within valid ranges - -VCL-UT is entirely optional. Teams can use plain VCL for standard operations and -adopt VCL-UT incrementally when they want formal guarantees. - -=== How does federation work? - -VeriSimDB can federate across heterogeneous database backends, bringing drift -detection to existing infrastructure without requiring data migration. Federation -connectors translate between VCL and native query languages: - -* *PostgreSQL:* SQL translation, table-to-document mapping -* *ArangoDB:* AQL translation, graph-to-graph mapping -* *Elasticsearch:* Query DSL translation, search-to-vector mapping -* *Neo4j:* Cypher translation, graph-to-graph mapping - -Each connector is independently tested and includes fallback to local-only mode -if the external database is unavailable. Federation is optional -- VeriSimDB can -operate as a standalone database without any external connections. - -== Licensing and Business - -=== Is VeriSimDB open source? - -Yes. VeriSimDB is released under the *Palimpsest License (PMPL-1.0-or-later)*, -a permissive open-source license. The community edition is full-featured with -no artificial restrictions -- all 8 modalities, drift detection, -self-normalisation, VCL, and federation are included. - -=== What is the PMPL license? - -The Palimpsest License (PMPL-1.0-or-later) is a permissive open-source license -developed by hyperpolymath. It allows free use, modification, and distribution -of the software with minimal restrictions. For projects that require an -OSI-approved license, an MPL-2.0 fallback is available. - -=== Is there a commercial version? - -VeriSimDB follows an open-core model. The community edition is free and -full-featured. A *commercial support license* is available for organisations that -require: - -* Enterprise SLA (99.9% uptime guarantee) -* Priority security patches (24-hour response time) -* Dedicated support channel -* Deployment consulting and architecture review -* Training and certification for development teams - -The commercial license does not add features beyond what the community edition -provides. It provides operational guarantees and expert support. - -=== How much does the commercial license cost? - -The target annual contract value (ACV) for enterprise support is $40,000. -Pricing is tiered by organisation size and SLA level. Contact -j.d.a.jewell@open.ac.uk for a quote. - -== Getting Started - -=== How do I try VeriSimDB? - -1. Clone the repository: `git clone https://github.com/hyperpolymath/verisimdb` -2. Follow the getting-started guide: `docs/getting-started.adoc` -3. Run the test suite: `cargo test` (Rust) and `mix test` (Elixir) -4. Explore the VCL examples: `docs/vcl-examples.adoc` - -=== Is VeriSimDB production-ready? - -VeriSimDB is in *alpha stage*. The core architecture, modality model, drift -detection engine, and VCL language are implemented and tested (662 tests, 0 -failures). However, production deployment requires: - -* Performance benchmarking at enterprise scale -* Security hardening review -* Deployment automation (container images, Helm charts) -* Operational tooling (monitoring, alerting, backup) - -VeriSimDB is suitable for development, evaluation, and pilot deployments. We -recommend contacting j.d.a.jewell@open.ac.uk before production deployment to -discuss architecture review and support options. - -== Project and Community - -=== Who built VeriSimDB? - -VeriSimDB was created by *Jonathan D.A. Jewell*, a researcher at The Open -University (UK) and the lead of the hyperpolymath open-source organisation. -Jonathan's background spans multimodal data systems, formal methods, programming -language theory, and database internals. - -=== What is the roadmap? - -The roadmap includes: - -* *Short-term (0-6 months):* Community building, documentation improvements, - federation connector hardening, performance benchmarking -* *Medium-term (6-18 months):* Container images, managed service pilot, VCL-UT - tooling, additional federation connectors -* *Long-term (18-36 months):* Enterprise features (RBAC, audit logging), - OEM licensing framework, academic paper publications - -See `ROADMAP.adoc` in the repository for the detailed roadmap. - -=== Can I contribute? - -Yes. VeriSimDB welcomes contributions in: - -* *Rust:* Core storage, indexing, drift detection algorithms -* *Elixir:* Orchestration, fault tolerance, federation connectors -* *ReScript:* Registry, type-safe metadata -* *Documentation:* Tutorials, use case guides, API references -* *Testing:* Additional test cases, integration tests, benchmarks - -See `CONTRIBUTING.md` in the repository for contribution guidelines. - -=== Where can I learn more? - -* *Repository:* https://github.com/hyperpolymath/verisimdb -* *Documentation:* `docs/` directory in the repository -* *VCL specification:* `docs/VCL-SPEC.adoc` -* *Architecture:* `docs/vcl-architecture.adoc` -* *Contact:* j.d.a.jewell@open.ac.uk diff --git a/verisimdb/docs/business/pr/key-messages.adoc b/verisimdb/docs/business/pr/key-messages.adoc deleted file mode 100644 index 90b28b87..00000000 --- a/verisimdb/docs/business/pr/key-messages.adoc +++ /dev/null @@ -1,347 +0,0 @@ -// SPDX-License-Identifier: CC-BY-SA-4.0 -// Copyright (c) 2026 Jonathan D.A. Jewell (hyperpolymath) -= VeriSimDB -- Key Messages and Talking Points -Jonathan D.A. Jewell -:revdate: 2026-02-28 -:revremark: Audience-segmented talking points for press, conferences, and outreach -:toc: macro -:toclevels: 2 -:sectnums: -:icons: font - -[abstract] -Core messaging framework for VeriSimDB, segmented by audience. Each section -provides the primary message, supporting points, proof points, and phrases to -use and avoid. Intended for press interviews, conference talks, blog posts, -and community outreach. - -toc::[] - -== Core Messages (All Audiences) - -=== The One-Line Pitch - -[quote] -VeriSimDB is the world's first multimodal database with cross-modal drift -detection and self-normalisation. - -=== The Three-Sentence Pitch - -[quote] -The same entity -- a customer, a gene, a transaction -- is represented across -5 to 10 different systems: a graph database, a vector index, a document store, -a time-series database. When one representation changes, the others silently -go stale. VeriSimDB detects this cross-modal drift and repairs it automatically, -before downstream systems consume contradictory data. - -=== The Six Core Differentiators - -1. *Cross-modal drift detection* -- No other database detects when multiple - representations of the same entity diverge -2. *8-modality octad model* -- Graph, vector, tensor, semantic, document, - temporal, provenance, and spatial in one entity -3. *Self-normalisation* -- Automatic repair of drifted representations with - configurable resolution strategies -4. *Proof-carrying queries (VCL-UT)* -- Optional dependent-type verification - of query correctness using Idris2 -5. *Heterogeneous federation* -- Unified querying across PostgreSQL, ArangoDB, - Elasticsearch, and Neo4j without data migration -6. *Open source (PMPL)* -- Full-featured community edition with no artificial - restrictions - -== Technical Audience (Database Engineers, Backend Developers) - -=== Primary Message - -[quote] -VeriSimDB introduces a new primitive to database systems: the coherence score. -Every entity has a quantified measure of how consistent its representations are -across all active modalities. This makes cross-modal drift a queryable, trackable, -alertable condition -- not an invisible failure mode. - -=== Supporting Points - -* *Architecture*: Rust core for storage and drift algorithms, Elixir/OTP for - fault-tolerant orchestration. The BEAM VM provides supervision trees, hot code - reloading, and battle-tested concurrency. Rust provides near-C performance with - memory safety. - -* *VCL (VeriSim Consonance Language)*: Purpose-built query language for - cross-modal operations. Familiar to SQL users but supports modality-specific - operations (graph traversals, vector similarity, tensor ops, temporal - time-travel, spatial containment) in a unified syntax. - -* *VCL-UT*: Optional dependent-type layer using Idris2. Proves query type - correctness, semantic compatibility of cross-modal joins, and threshold - validity at compile time. Entirely opt-in -- plain VCL works without it. - -* *Federation*: Connectors translate between VCL and native query languages - (SQL, Cypher, AQL, Query DSL). Each connector independently tested with - fallback to local-only mode. No vendor lock-in. - -* *ABI/FFI*: Idris2 for ABI definitions with dependent-type proofs, Zig for - C-compatible FFI bindings. Generated C headers bridge the two. This is the - hyperpolymath standard for all projects, not VeriSimDB-specific. - -=== Proof Points - -* 662 tests (510 Rust + 152 Elixir), 0 failures -* Drift detection algorithm: cosine distance between modality embeddings with - adaptive thresholds -* No C/C++ build dependencies (Oxigraph feature-flagged off; redb pure-Rust - graph backend is default) -* Container images build with zero external toolchain requirements -* SLSA provenance and SBOM generation in CI/CD pipeline - -=== Use These Phrases - -* "Coherence score" (not "consistency check" -- it is continuous, not binary) -* "Cross-modal drift" (not "data inconsistency" -- drift implies gradual - divergence, which is more accurate than sudden failure) -* "Self-normalisation" (not "auto-repair" -- normalisation implies bringing - representations into agreement, not fixing broken data) -* "Octad" (not "8 modalities" -- the octad is a unified model, not a feature list) -* "Proof-carrying queries" (not "type-checked queries" -- proofs travel with the - query result, not just the query) - -=== Avoid These Phrases - -* "Better than Neo4j/Pinecone/Weaviate" -- VeriSimDB occupies a different category, - not a better version of the same category -* "Replaces your database" -- VeriSimDB sits alongside existing databases - (federation) or operates standalone; it does not replace PostgreSQL -* "AI-native database" -- VeriSimDB serves AI workloads (GraphRAG) but is not - limited to AI; the core value is data consistency -* "NoSQL" -- VeriSimDB has its own query language (VCL); calling it NoSQL implies - it competes with document stores - -== Business Audience (CTOs, VP Engineering, Data Leaders) - -=== Primary Message - -[quote] -Gartner estimates that poor data quality costs enterprises $12.9 million per year. -The root cause is not dirty data at entry -- it is cross-system inconsistency. -The same entity represented differently across 5-10 systems, with no mechanism to -detect when those representations disagree. VeriSimDB detects this divergence -continuously and repairs it automatically. - -=== Supporting Points - -* *ROI*: Cross-system inconsistency is the single largest driver of data quality - costs. Existing solutions (ETL, CDC, MDM) are reactive and partial. VeriSimDB - is proactive and comprehensive -- it monitors all representations of every entity - and catches drift before it reaches downstream consumers. - -* *Risk reduction*: In regulated industries (financial services, healthcare, - pharmaceuticals), cross-system inconsistency creates compliance risk. VeriSimDB's - proof certificates provide auditable evidence of query correctness and data - consistency. - -* *No rip-and-replace*: Federation mode brings drift detection to existing - infrastructure (PostgreSQL, Neo4j, ArangoDB, Elasticsearch) without data - migration. The value is additive -- VeriSimDB sits on top of what you already - have. - -* *Open core model*: The community edition is full-featured with no artificial - restrictions. Commercial support licenses provide SLAs, priority patches, and - deployment consulting. No bait-and-switch. - -* *Team size*: VeriSimDB is designed to be operated by existing data teams. It - does not require a dedicated "VeriSimDB team" -- it integrates into existing - data platform practices. - -=== Proof Points - -* Open source under PMPL (permissive licence, MPL-2.0 fallback available) -* Enterprise support target ACV: $40,000/year -* Built on battle-tested foundations: Rust (memory safety), Elixir/OTP - (fault tolerance, 99.9999% uptime heritage from telecom) - -=== Use These Phrases - -* "Data quality infrastructure" (not "database" -- CTOs hear "database" and think - "another system to manage") -* "Cross-system consistency" (not "drift detection" -- "drift" requires explanation; - "consistency" is immediately understood) -* "Proactive detection" (not "monitoring" -- monitoring implies dashboards; - proactive detection implies the system acts before humans notice) -* "Open core" (not "freemium" -- open core is a respected OSS business model; - freemium has negative connotations) - -=== Avoid These Phrases - -* "8 modalities" or "octad" -- too technical for a business audience on first - hearing; lead with the problem (cross-system inconsistency), not the solution - architecture -* "Dependent types" or "Idris2" -- these are implementation details that add - confusion in a business conversation; use "mathematically proven query - correctness" if the topic comes up -* "Replacing your data infrastructure" -- CTOs will immediately push back; - VeriSimDB augments, not replaces - -== Academic Audience (Researchers, Conference Attendees) - -=== Primary Message - -[quote] -VeriSimDB formalises the concept of cross-modal drift in multimodal data systems -and provides an operational system for detecting, quantifying, and repairing -representational divergence across 8 data modalities. The query language includes -an optional dependent-type verification layer (VCL-UT) that provides compile-time -proofs of query correctness. - -=== Supporting Points - -* *Novel contribution*: The coherence score -- a continuous, quantified measure of - cross-modal consistency -- is a new primitive in database theory. Existing - multi-model databases store multiple models but do not track inter-model - consistency. - -* *Formal semantics*: VCL has published formal semantics (denotational and - operational). VCL-UT extends these with dependent-type judgements formalised in - Idris2, providing machine-checkable proofs of query properties. - -* *Reproducibility*: The entire system is open source. Experiments can be - reproduced. The drift detection algorithm (cosine distance between modality - embeddings with adaptive thresholds) is fully documented. - -* *Cross-disciplinary relevance*: The octad model and drift detection have - applications in bioinformatics (multi-omics consistency), digital twins - (simulation-physical divergence), knowledge graph quality assurance (ontological - consistency), and climate science (cross-sensor calibration). - -* *Related work*: Builds on multi-model database theory (Lu & Holubova 2019), - data quality measurement (Batini et al. 2009), and dependent-type verification - (Brady 2013). Extends these with cross-modal drift as a first-class concept. - -=== Proof Points - -* Published technical report: "Cross-Modal Drift Detection and Self-Normalisation - in Heterogeneous Database Federations" -* Case study: IDApTIK game level architecture with Idris2 dependent-type proofs -* 14-module ABI with 5 cross-domain proofs, zero `believe_me` usage -* Formal VCL grammar (EBNF) and type system specification - -=== Venues - -* VLDB (Very Large Data Bases) -- core database systems venue -* SIGMOD (ACM Special Interest Group on Management of Data) -* ICDE (IEEE International Conference on Data Engineering) -* EDBT (Extending Database Technology) -- European venue -* CIDR (Conference on Innovative Data Systems Research) -- systems/vision papers -* POPL / ICFP (dependent types and formal methods angle) -* CIKM (Conference on Information and Knowledge Management) -- knowledge graphs - -=== Use These Phrases - -* "Cross-modal drift" (the central concept; define it precisely) -* "Coherence score" (the quantified measure; explain the algorithm) -* "Dependent-type verification" (not "formal proofs" alone -- specify the - mechanism) -* "Octad model" (the 8-modality framework; relate to existing multi-model - taxonomy) - -=== Avoid These Phrases - -* "World's first" -- academic audiences are sceptical of superlatives; lead with - the contribution, not the claim -* "Enterprise-grade" -- irrelevant in an academic context -* "$12.9 million" -- Gartner statistics are marketing, not research evidence - -== Open Source Community - -=== Primary Message - -[quote] -VeriSimDB is a full-featured multimodal database released under a permissive -open-source licence (PMPL-1.0-or-later, with MPL-2.0 fallback). All 8 -modalities, drift detection, self-normalisation, VCL, and federation are in the -community edition. There is no "enterprise edition" that gates features. - -=== Supporting Points - -* *No bait-and-switch*: The commercial licence provides operational support (SLA, - priority patches, consulting), not additional features. Everything ships in the - open-source release. - -* *Contribution-friendly*: Multiple entry points for contributors -- Rust (core - engine), Elixir (orchestration), ReScript (registry), documentation, testing. - `CONTRIBUTING.md` with clear guidelines. - -* *Principled licensing*: PMPL-1.0-or-later is permissive. MPL-2.0 fallback for - ecosystems requiring OSI-approved licences. Third-party dependencies retain - their original licences. - -* *Part of a larger ecosystem*: VeriSimDB is one of 265+ hyperpolymath - repositories. It integrates with QuandleDB (dependent-type graph database), - LithoGlyph (provenance database), PanLL (unified mission control), and the - gitbot-fleet (automated repository maintenance). - -* *Tech stack choices*: Rust + Elixir is a deliberate architectural choice, not - an accident. Rust for performance-critical paths, Elixir/OTP for fault-tolerant - distributed coordination. No Python, no Go, no Node.js. - -=== Proof Points - -* Repository: https://github.com/hyperpolymath/verisimdb -* 662 tests, 0 failures -* CI/CD with SLSA provenance and SBOM generation -* Containerfile for Podman/Docker deployment -* Comprehensive documentation (40+ `.adoc` files) - -=== Community Channels - -* Elixir community (ElixirForum, Elixir Slack) -- OTP orchestration layer -* Rust community (r/rust, Rust Users Forum) -- core engine, storage -* BEAM community (Erlang/Elixir conferences, Code BEAM) -- fault tolerance -* Database community (r/database, HN, Lobsters) -- novel architecture -* Formal methods community (Idris2 Discourse, TYPES mailing list) -- VCL-UT - -=== Use These Phrases - -* "Full-featured community edition" (emphasise no feature gating) -* "Permissive licence" (PMPL is permissive; say so explicitly) -* "Contributions welcome in [Rust/Elixir/ReScript/docs]" (specific, not generic) - -=== Avoid These Phrases - -* "We" (until there is a team; "the project" or "VeriSimDB" is more honest for - a single-developer project) -* "Production-ready" (the project is in alpha; be honest about maturity) -* "Enterprise" (the OSS community is not the enterprise audience; lead with - technical merit) - -== Message Consistency Checklist - -Before any public communication, verify: - -[cols="1,4"] -|=== -| Check | Requirement - -| Drift definition -| Cross-modal drift is defined as divergence between representations of the *same -entity* across modalities -- not divergence within a single modality - -| Octad completeness -| All 8 modalities are listed correctly: graph, vector, tensor, semantic, -document, temporal, provenance, spatial (in that order) - -| Competitor framing -| Competitors are positioned as "different category," not "inferior product." -VeriSimDB is a new category (cross-modal consistency infrastructure), not a -better version of an existing category - -| Maturity honesty -| Current status is alpha. 662 tests pass. Production deployment requires -additional hardening. Do not oversell readiness. - -| Licence accuracy -| PMPL-1.0-or-later (primary). MPL-2.0 (fallback where required). Not AGPL. -Not MIT. Not Apache. - -| Author attribution -| Jonathan D.A. Jewell. The Open University, UK. hyperpolymath. -Not "the VeriSimDB team" (single developer). -|=== diff --git a/verisimdb/docs/business/pr/press-release.adoc b/verisimdb/docs/business/pr/press-release.adoc deleted file mode 100644 index af6208dc..00000000 --- a/verisimdb/docs/business/pr/press-release.adoc +++ /dev/null @@ -1,102 +0,0 @@ -// SPDX-License-Identifier: CC-BY-SA-4.0 -// Copyright (c) 2026 Jonathan D.A. Jewell (hyperpolymath) -= VeriSimDB Launch Press Release -Jonathan D.A. Jewell -:revdate: 2026-02-28 -:revremark: Launch announcement draft — standard press release format - -== FOR IMMEDIATE RELEASE - -=== VeriSimDB: World's First Multimodal Database with Cross-Modal Drift Detection Launches as Open Source - -_New database category addresses the $12.9M annual cost of data inconsistency -by detecting and repairing divergent data representations automatically_ - -*[City, Date]* -- hyperpolymath today announced the open-source release of -*VeriSimDB*, the world's first multimodal database with cross-modal drift -detection and self-normalisation. VeriSimDB stores entities across 8 simultaneous -data representations -- graph, vector, tensor, semantic, document, temporal, -provenance, and spatial -- and continuously monitors whether those representations -are consistent with each other. - -Unlike existing databases that store data in a single modality (Neo4j for graphs, -Pinecone for vectors) or multiple models without consistency checking (ArangoDB, -SurrealDB), VeriSimDB introduces *coherence scoring*: a quantified measure of -cross-modal consistency for every entity. When an entity's representations -diverge -- a condition called "drift" -- VeriSimDB detects the divergence and can -automatically reconcile the data through configurable self-normalisation -strategies. - -"Every enterprise I have spoken with has the same problem: the same customer, -product, or transaction is represented differently across five or ten systems, and -nobody knows when those representations disagree," said *Jonathan D.A. Jewell*, -creator of VeriSimDB and researcher at The Open University. "Gartner estimates -this costs organisations $12.9 million per year. VeriSimDB is the first database -designed from the ground up to detect and repair this kind of inconsistency." - -=== Technical Foundation - -VeriSimDB is built on a *Rust core* for performance and memory safety, with -*Elixir/OTP orchestration* for fault-tolerant concurrency. The database includes -VCL (VeriSim Consonance Language), a purpose-built query language for -cross-modal operations, and an optional dependent-type layer (VCL-UT) that -provides compile-time proofs of query correctness using Idris2. - -The project's engineering maturity is demonstrated by a test suite of *662 tests* -(510 Rust + 152 Elixir) with *0 failures*, comprehensive documentation -including formal semantics, and a CI/CD pipeline with SLSA provenance and SBOM -generation. - -VeriSimDB also supports *federation* across heterogeneous database backends -including PostgreSQL, ArangoDB, and Elasticsearch, allowing organisations to -bring drift detection to their existing data infrastructure without requiring -data migration. - -=== Availability and Licensing - -VeriSimDB is available immediately as open source under the *Palimpsest License -(PMPL-1.0-or-later)*, a permissive open-source license. The community edition -is full-featured with no artificial restrictions. Commercial support licenses -with enterprise SLAs, priority security patches, and deployment consulting are -available for organisations that require operational guarantees. - -The source code, documentation, and getting-started guide are available at: -https://github.com/hyperpolymath/verisimdb - -=== Target Use Cases - -VeriSimDB addresses data consistency challenges across multiple domains: - -* *GraphRAG / AI Knowledge Graphs:* Detect when vector embeddings drift from - knowledge graph state, preventing LLM hallucinations -* *Biomedical Research:* Ensure clinical records, trial timelines, and - provenance chains remain consistent for regulatory compliance -* *Financial Services:* Real-time cross-system consistency for AML/KYC - compliance and fraud detection -* *Supply Chain:* End-to-end traceability with automatic detection of expired - certifications and routing inconsistencies -* *Cybersecurity:* Maintain consistent threat intelligence across graph, vector, - and document representations - -=== About hyperpolymath - -hyperpolymath is an open-source software organisation led by Jonathan D.A. Jewell, -a researcher at The Open University (UK). The hyperpolymath ecosystem comprises -265+ repositories spanning databases, programming languages, developer tools, -formal verification, and AI infrastructure. VeriSimDB is part of the -hyperpolymath next-generation database portfolio, which also includes QuandleDB -(dependent-type graph database) and LithoGlyph (provenance database). - -=== Contact - -[cols="1,3"] -|=== -| Name | Jonathan D.A. Jewell -| Email | j.d.a.jewell@open.ac.uk -| GitHub | https://github.com/hyperpolymath -| Project | https://github.com/hyperpolymath/verisimdb -|=== - -'''' - -_###_ diff --git a/verisimdb/docs/business/strategy/adoption-roadmap.adoc b/verisimdb/docs/business/strategy/adoption-roadmap.adoc deleted file mode 100644 index 6e6a3e6b..00000000 --- a/verisimdb/docs/business/strategy/adoption-roadmap.adoc +++ /dev/null @@ -1,546 +0,0 @@ -// SPDX-License-Identifier: CC-BY-SA-4.0 -// Copyright (c) 2026 Jonathan D.A. Jewell (hyperpolymath) -= VeriSimDB -- Adoption Roadmap -Jonathan D.A. Jewell -:revdate: 2026-02-28 -:revremark: Three-phase adoption roadmap with milestones and success criteria -:toc: macro -:toclevels: 3 -:sectnums: -:icons: font - -[abstract] -Phased adoption roadmap for VeriSimDB, from open-source community building -through institutional pilots to enterprise deployment. Each phase has specific -milestones, success criteria, exit conditions, and contingency plans. The roadmap -is designed to be achievable by a single developer (Phase 1) scaling to a small -team (Phase 2-3). - -toc::[] - -== Roadmap Overview - -.... -Phase 1: GraphRAG Community Phase 2: Biomedical/Research Phase 3: Enterprise -(Months 0-12) (Months 12-24) (Months 24-36) -┌────────────────────┐ ┌────────────────────┐ ┌────────────────────┐ -│ Open source release│ │ Institutional pilots│ │ Commercial support │ -│ Community building │ ──exit──> │ Academic papers │ ──exit──> │ Managed service │ -│ Content + talks │ criteria │ Consulting revenue │ criteria │ Enterprise sales │ -│ GraphRAG demos │ │ Domain-specific docs│ │ Partner channel │ -└────────────────────┘ └────────────────────┘ └────────────────────┘ -.... - -== Phase 1: GraphRAG Community (Months 0-12) - -=== Objective - -Establish VeriSimDB as the recognised solution for cross-modal drift detection, -starting with the GraphRAG community where the problem is most acutely felt. - -=== Why GraphRAG First - -The GraphRAG community is the ideal starting point for three reasons: - -1. *Immediate pain point*: RAG pipelines that combine vector databases with - knowledge graphs suffer from embedding drift -- the exact problem VeriSimDB - detects. LangChain's documentation explicitly acknowledges the need for - "indexing APIs with cleanup processes" to keep vector stores synchronised - with source documents. - -2. *Technical alignment*: GraphRAG architectures use 3-4 of VeriSimDB's 8 - modalities (graph, vector, document, sometimes semantic). This is enough to - demonstrate value without requiring the full octad. - -3. *Accessible community*: The GraphRAG community is active on GitHub, Discord, - and Reddit. It is developer-led, open-source-friendly, and values technical - demonstrations over marketing claims. - -=== Milestones - -[cols="1,3,2,1"] -|=== -| Month | Milestone | Deliverable | Owner - -| 1 -| Public repository launch -| README with architecture diagram, getting-started guide, VCL examples, -CONTRIBUTING.md -| Founder - -| 2 -| First blog post -| "Introducing VeriSimDB: Cross-Modal Drift Detection for Multimodal Data" -(1500-2000 words, published on GitHub Pages) -| Founder - -| 3 -| GraphRAG demo -| Working demo: create entities with graph + vector + document modalities, -introduce drift, detect and repair. Screencast (5 minutes). -| Founder - -| 4 -| GraphRAG integration guide -| Tutorial: "Detecting Embedding Drift in RAG Pipelines with VeriSimDB" -(step-by-step with code examples) -| Founder - -| 5 -| Community engagement -| Post in LangChain community showcase, Qdrant Discord, Weaviate Slack. -Respond to all GitHub issues within 48 hours. -| Founder - -| 6 -| Conference submission -| Submit talk proposal to Code BEAM EU or EuroRust (whichever has earlier -deadline) -| Founder - -| 8 -| VCL tutorial series -| 3-part tutorial: basic VCL queries, cross-modal queries, drift detection -queries -| Founder - -| 10 -| Comparison content -| Blog post: "VeriSimDB vs. Great Expectations vs. Monte Carlo: Different -Problems, Different Solutions" -| Founder - -| 12 -| Phase 1 retrospective -| Assess metrics, decide Phase 2 entry -| Founder -|=== - -=== Success Criteria (Exit Conditions for Phase 2) - -Phase 2 entry requires meeting *at least 4 of 6* criteria: - -[cols="1,3,1"] -|=== -| # | Criterion | Target - -| 1 | GitHub stars | >= 500 -| 2 | Monthly unique cloners | >= 100 -| 3 | External pull requests (merged) | >= 5 -| 4 | Open issues from external users | >= 15 -| 5 | Blog post cumulative views | >= 10,000 -| 6 | Conference talks given or accepted | >= 2 -|=== - -=== Contingency: Phase 1 Stalls - -If fewer than 3 criteria are met at Month 12: - -* *Diagnose*: Analyse which content performed best, which communities engaged, - which demo resonated. The problem may be messaging, not product. -* *Pivot community*: If GraphRAG is not responding, try the data quality - community (Great Expectations users) or the Elixir community (OTP angle). -* *Extend Phase 1*: Add 6 months. The product is open source and self-funded -- - there is no investor pressure to advance phases. -* *Minimum threshold*: If GitHub stars < 100 at Month 12, reassess whether the - product-market fit hypothesis is correct. Consider a fundamental messaging - change. - -== Phase 2: Biomedical and Research Institutions (Months 12-24) - -=== Objective - -Secure 2-3 pilot deployments with research institutions where data quality is -critical, multi-modal data is the norm, and the academic community provides -credibility leverage. - -=== Why Biomedical/Research Second - -1. *Natural octad fit*: Biomedical data (multi-omics, clinical records, imaging) - uses 6-8 of VeriSimDB's modalities. The octad model is not an academic - abstraction -- it maps directly to how biological data exists. - -2. *Regulatory motivation*: FDA 21 CFR Part 11, EMA GxP requirements, and grant - agency data management plans create institutional demand for data consistency - tools with audit trails (provenance modality) and formal verification - (VCL-UT proof certificates). - -3. *Publication opportunity*: Successful pilot deployments generate academic - papers, which generate citations, which generate credibility. A VLDB or - SIGMOD publication is worth more than 1,000 GitHub stars for enterprise - credibility. - -4. *Budget availability*: Research institutions have grant budgets for data - infrastructure. These budgets are smaller than enterprise IT budgets but have - shorter procurement cycles and fewer compliance hurdles. - -=== Milestones - -[cols="1,3,2,1"] -|=== -| Month | Milestone | Deliverable | Owner - -| 13 -| Academic paper submission -| Submit to VLDB or SIGMOD: "Cross-Modal Drift Detection and Self-Normalisation -in Heterogeneous Database Federations" (12-page research paper) -| Founder - -| 14 -| Domain-specific documentation -| Tutorial: "Multi-Omics Entity Consistency with VeriSimDB" (bioinformatics -audience) -| Founder / DevRel - -| 15 -| Conference presentation -| Present at ISMB, BOSC, or domain-specific conference (bioinformatics, -clinical informatics, or digital humanities) -| Founder - -| 16 -| First pilot engagement -| Identify and engage pilot institution. Begin architecture review for their -specific data consistency challenge. -| Founder - -| 18 -| Pilot deployment -| VeriSimDB deployed in pilot institution's staging environment. Initial data -ingestion and drift detection running. -| Founder + Pilot Partner - -| 20 -| Pilot evaluation -| Quantified results from pilot: drift detection accuracy, false positive rate, -performance metrics. Draft case study. -| Founder + Pilot Partner - -| 22 -| Second pilot -| Begin second pilot deployment (different institution or domain) to validate -generalisability -| Founder - -| 24 -| Phase 2 retrospective -| Assess metrics, publish case studies, decide Phase 3 entry -| Founder -|=== - -=== Target Institutions - -[cols="2,2,3"] -|=== -| Type | Examples | VeriSimDB Value - -| *Bioinformatics labs* -| EBI (Cambridge), Sanger Institute, Broad Institute -| Multi-omics data consistency across genomic, proteomic, and metabolomic -representations - -| *Clinical research centres* -| NHS Digital, academic medical centres -| Cross-system patient record consistency (EHR + pharmacy + lab + imaging) - -| *Digital humanities* -| Wikidata, British Library, Europeana -| Knowledge graph quality assurance, ontological consistency at scale - -| *Climate science* -| CEDA (UK), DKRZ (Germany), ESA CCI -| Cross-sensor calibration verification, multi-archive federation -|=== - -=== Engagement Model - -1. *Initial contact*: Through academic conference, community forum, or direct - outreach to data management teams -2. *Architecture review*: Free 2-hour consultation to understand their data - consistency challenges and assess VeriSimDB fit -3. *Pilot agreement*: Formal agreement covering scope, timeline, success metrics, - and publication rights -4. *Deployment support*: Free deployment support for the pilot (investment in - case study material) -5. *Evaluation and publication*: Joint paper or case study documenting results - -=== Revenue Model (Phase 2) - -Revenue comes from consulting, not software licences: - -[cols="2,1,3"] -|=== -| Service | Price | Description - -| Architecture review -| $5,000-$15,000 -| Assessment of data consistency challenges, VeriSimDB fit, deployment plan - -| Custom federation connector -| $10,000-$25,000 -| Development of a federation connector for the institution's specific backend -(e.g., OpenClinica, REDCap, iRODS) - -| Training workshop -| $2,000-$5,000 -| Half-day or full-day VCL training for the institution's development team - -| Ongoing consulting -| $150-$250/hour -| Architecture guidance, performance tuning, VCL query optimisation -|=== - -*Target:* 3-5 consulting engagements, $50,000-$150,000 total revenue. - -=== Success Criteria (Exit Conditions for Phase 3) - -Phase 3 entry requires meeting *at least 4 of 6* criteria: - -[cols="1,3,1"] -|=== -| # | Criterion | Target - -| 1 | Pilot deployments completed | >= 2 -| 2 | Academic papers published or accepted | >= 1 -| 3 | Consulting revenue (cumulative) | >= $50,000 -| 4 | GitHub stars | >= 2,000 -| 5 | External contributors (active in last 90 days) | >= 10 -| 6 | Conference talks given (cumulative) | >= 5 -|=== - -=== Contingency: Phase 2 Stalls - -If fewer than 3 criteria are met at Month 24: - -* *Pivot domain*: If bioinformatics is not responding, try digital twins - (manufacturing) or data mesh (enterprise data teams). Different domains have - different adoption timescales. -* *Strengthen Phase 1*: More community content, more conference talks, more - integration guides. Phase 2 may fail because Phase 1 did not build enough - awareness. -* *Reassess product-market fit*: If no institution is willing to pilot, the - problem may be less acute than hypothesised. Consider whether drift detection - is a "nice to have" versus "must have" for the target audience. - -== Phase 3: Enterprise Expansion (Months 24-36) - -=== Objective - -Launch commercial support licences and expand into enterprise markets where -cross-system data consistency has measurable financial impact ($12.9M/year -per Gartner). - -=== Why Enterprise Third - -1. *Credibility required*: Enterprise buyers require proof of capability -- - case studies, academic publications, and production deployments. Phases 1-2 - build this credibility. - -2. *Procurement cycles*: Enterprise procurement takes 6-12 months. Starting - enterprise sales in Phase 1 would mean no revenue until Month 18-24 anyway. - Building community and institutional credibility during that period is a - better use of time. - -3. *Product maturity*: Enterprise deployment requires production hardening, - security review, monitoring tooling, and operational documentation that are - not needed for community or research use. - -=== Target Verticals - -[cols="2,2,1,2"] -|=== -| Vertical | Use Case | Target ACV | Entry Point - -| *Financial services* -| AML/KYC cross-system entity consistency -| $50,000+ -| Compliance teams at tier-2 banks and fintech - -| *Manufacturing* -| Digital twin synchronisation, IoT data consistency -| $40,000+ -| Engineering teams at mid-size manufacturers - -| *Pharmaceuticals* -| Clinical trial data consistency, FDA compliance -| $60,000+ -| Data management teams at CROs and pharma - -| *Supply chain* -| Traceability, certification expiry detection -| $35,000+ -| Data platform teams at logistics providers -|=== - -=== Milestones - -[cols="1,3,2"] -|=== -| Month | Milestone | Deliverable - -| 25 -| Commercial support launch -| Support tier structure published, pricing page live, billing infrastructure -operational - -| 26 -| Enterprise content -| Case studies from Phase 2 pilots, ROI calculator, compliance guides -(SOC 2 readiness, GDPR cross-system erasure) - -| 27 -| First enterprise prospect -| Qualified enterprise prospect in pipeline (identified through community -adoption or conference contact) - -| 28 -| Container hardening -| Production-grade container images with security scanning, minimal base -images (Chainguard Wolfi), SBOM, SLSA attestation - -| 30 -| First enterprise customer -| Signed commercial support contract - -| 32 -| Managed service pilot -| Managed VeriSimDB instance for 1-2 customers (hosted on Fly.io or equivalent) - -| 34 -| Second and third enterprise customers -| Pipeline conversion from Phase 1-2 community and institutional contacts - -| 36 -| Phase 3 retrospective -| ARR assessment, team scaling decision, managed service viability -|=== - -=== Sales Motion - -==== Developer-Led (Primary) - -The expected enterprise sales motion for VeriSimDB is bottom-up: - -1. A data engineer discovers VeriSimDB through a blog post, conference talk, - or community recommendation -2. The engineer evaluates VeriSimDB in a staging environment, running drift - detection on a subset of production data -3. The engineer demonstrates results to their team lead or VP Engineering, - showing quantified drift across existing systems -4. The team requests an architecture review (consulting engagement, $5K-$15K) -5. Successful evaluation leads to a commercial support licence purchase - -This motion has a *low CAC* ($8,000 estimated) because the developer does the -evaluation and internal selling. Traditional enterprise sales (outbound SDR, -demo, POC, procurement) costs $30,000-$50,000 per customer. - -==== Partner Channel (Secondary, Phase 3b) - -After establishing 5-10 commercial customers: - -* Engage system integrators (Accenture, Deloitte, Thoughtworks) who advise - enterprise clients on data infrastructure -* Create partner training programme (VCL certification) -* Provide referral incentives (10-15% of first-year ACV) - -=== Success Criteria (Phase 3) - -[cols="1,3,1"] -|=== -| # | Criterion | Target - -| 1 | Commercial support customers | >= 5 -| 2 | ARR | >= $200,000 -| 3 | GitHub stars | >= 5,000 -| 4 | External contributors (active in last 90 days) | >= 25 -| 5 | Managed service users (pilot) | >= 2 -| 6 | Published case studies | >= 3 -|=== - -== Technical Readiness by Phase - -Each phase requires specific technical capabilities. These are development -priorities, not marketing milestones. - -[cols="3,1,1,1"] -|=== -| Capability | Phase 1 | Phase 2 | Phase 3 - -| Core octad storage (all 8 modalities) | Required | Required | Required -| Drift detection and coherence scoring | Required | Required | Required -| Self-normalisation (configurable strategies) | Required | Required | Required -| VCL parser and query execution | Required | Required | Required -| Getting-started guide and VCL examples | Required | Required | Required -| Federation (PostgreSQL connector) | Demo quality | Production quality | Production quality -| Federation (ArangoDB, Elasticsearch, Neo4j) | Not required | Demo quality | Production quality -| VCL-UT proof certificates | Documented design | Prototype | Production quality -| Container images (Podman/Docker) | Basic | Hardened | Production-grade -| Monitoring and alerting (Prometheus) | Basic | Comprehensive | SLA-grade -| RBAC and access control | Not required | Basic | Comprehensive -| Backup and restore | Not required | Basic | Automated -| Multi-node deployment | Not required | Not required | Documented + tested -| Managed service infrastructure | Not required | Not required | Pilot -|=== - -== Risk Register - -[cols="1,2,1,3"] -|=== -| Phase | Risk | Likelihood | Mitigation - -| 1 -| Low community adoption -| Medium -| Diversify community targets (GraphRAG, Elixir, Rust, data quality). Create -compelling demo. Engage consistently. - -| 1 -| Competitor adds drift detection -| Low -| Architectural depth (octad + VCL-UT) is difficult to retrofit. First-mover -advantage in category definition. - -| 2 -| No institution willing to pilot -| Medium -| Offer free deployment support. Lower barrier: "let us run drift detection -on your data for free." Publication incentive for academics. - -| 2 -| Academic paper rejected -| Medium -| Submit to multiple venues. Prepare demo paper as backup. Publish preprint -on arXiv regardless. - -| 3 -| Enterprise procurement delay (6-12 months) -| High -| Consulting revenue bridges gap. Developer-led adoption reduces procurement -friction. Start enterprise conversations 6 months before Phase 3. - -| 3 -| Single-founder scaling limit -| High -| First hire (DevRel) at Month 12. Second hire (Rust engineer) at Month 18. -Revenue from Phase 2 consulting funds initial hires. - -| All -| Burnout (single developer, 36-month plan) -| High -| Sustainable pace: no artificial urgency. VeriSimDB is self-funded; there -are no investor milestones. Take breaks. The project advances when it advances. -|=== - -== Summary Timeline - -.... -Month 0 Month 6 Month 12 Month 18 Month 24 Month 30 Month 36 - | | | | | | | - v v v v v v v - Launch GraphRAG Phase 1 First Phase 2 First Phase 3 - repo demo review pilot review enterprise review - Blog 500 stars deployed 2000 stars customer 5000 stars - posts 100 clones Paper $50K ARR $200K ARR target - Talks 5 PRs submitted consulting Managed $400K+ - service -.... diff --git a/verisimdb/docs/business/strategy/economics-analysis.adoc b/verisimdb/docs/business/strategy/economics-analysis.adoc deleted file mode 100644 index 0be0558d..00000000 --- a/verisimdb/docs/business/strategy/economics-analysis.adoc +++ /dev/null @@ -1,441 +0,0 @@ -// SPDX-License-Identifier: CC-BY-SA-4.0 -// Copyright (c) 2026 Jonathan D.A. Jewell (hyperpolymath) -= VeriSimDB -- Economics Analysis -Jonathan D.A. Jewell -:revdate: 2026-02-28 -:revremark: Macro and micro economics analysis for VeriSimDB commercialisation -:toc: macro -:toclevels: 3 -:sectnums: -:icons: font - -[abstract] -Macro-economic market analysis and micro-economic pricing strategy for VeriSimDB. -Covers the data quality tools market, the multimodal database segment, competitive -pricing analysis against Neo4j Enterprise, Pinecone, Weaviate, and others, and -open-core pricing model design. - -toc::[] - -== Macro-Economic Analysis - -=== Market Landscape - -VeriSimDB operates at the intersection of three established and one emerging -market segment. The following analysis uses industry reports from 2024-2026 -to size the addressable opportunity. - -==== Data Quality Tools Market - -[cols="1,3"] -|=== -| Metric | Value - -| Global market size (2024) | $1.7 billion -| Projected market size (2027) | $2.3 billion (MarketsAndMarkets) -| CAGR | 10.5% -| Key drivers | AI adoption (data quality as prerequisite), regulatory compliance -(GDPR, CCPA, FDA 21 CFR Part 11), cloud migration (multi-system architectures) -| Key players | Informatica, Talend, IBM, Collibra, Great Expectations, Monte Carlo -|=== - -VeriSimDB's relevance to this market is through *cross-system consistency* -- -a specific and currently unaddressed category of data quality. Existing tools -focus on schema validation (Great Expectations), pipeline observability (Monte -Carlo), or master data management (Informatica). None detect cross-modal drift -between different representation types of the same entity. - -*Addressable sub-segment:* Approximately 20% of data quality spend is related to -cross-system inconsistency resolution (ETL repair, manual reconciliation, MDM -maintenance). This yields a sub-segment of approximately *$460 million* by 2027. - -==== Graph Database Market - -[cols="1,3"] -|=== -| Metric | Value - -| Global market size (2024) | $3.8 billion -| Projected market size (2028) | $5.6 billion (Grand View Research) -| CAGR | 10.2% -| Key drivers | Knowledge graphs, fraud detection, recommendation engines, -GraphRAG architectures -| Key players | Neo4j, Amazon Neptune, TigerGraph, ArangoDB, JanusGraph -|=== - -VeriSimDB is not a graph database competitor. However, approximately 5% of -graph database customers operate multi-model architectures (graph + vector + -document) where cross-modal consistency is a concern. This yields a -sub-segment of approximately *$280 million* by 2028. - -==== Vector Database Market - -[cols="1,3"] -|=== -| Metric | Value - -| Global market size (2024) | $800 million -| Projected market size (2027) | $1.2 billion (estimated, no consensus report) -| CAGR | 14.9% -| Key drivers | LLM/RAG architectures, semantic search, recommendation systems -| Key players | Pinecone, Weaviate, Qdrant, Milvus, Chroma -|=== - -The vector database market is the most rapidly growing segment and the most -directly relevant to VeriSimDB's Phase 1 strategy (GraphRAG community). Teams -building RAG pipelines combine vector databases with knowledge graphs and -document stores -- precisely the multimodal architecture where drift detection -provides value. - -*Addressable sub-segment:* Approximately 10% of vector database customers -operate hybrid vector + graph architectures. This yields a sub-segment of -approximately *$120 million* by 2027. - -==== Emerging: Multimodal Database Segment - -[cols="1,3"] -|=== -| Metric | Value - -| Estimated market size (2026) | $800 million (analyst estimates, no formal report) -| Projected market size (2030) | $2.5 billion -| CAGR | ~33% (emerging market, high uncertainty) -| Key drivers | AI/ML infrastructure, multi-model consolidation, data quality -| Key players | ArangoDB (multi-model), SurrealDB (multi-model), VeriSimDB -(multimodal with drift detection) -|=== - -This segment is nascent and poorly defined. The distinction between -"multi-model" (multiple query models over a single storage engine) and -"multimodal" (multiple representation types with consistency guarantees) is -not yet recognised by analysts. VeriSimDB has an opportunity to define the -category. - -=== Combined Market Opportunity - -[cols="2,1,1,1"] -|=== -| Segment | TAM | SAM (Addressable Sub-Segment) | SOM (1% Serviceable) - -| Data Quality Tools | $2.3B | $460M | $23M -| Graph Databases | $5.6B | $280M | $14M -| Vector Databases | $1.2B | $120M | $6M -| Multimodal Databases | $800M | $160M | $8M -| *Combined* | | | *$51M* -|=== - -The *$51 million SOM* represents a conservative 1% capture of the addressable -sub-segments. This is achievable with 50-100 enterprise customers at an average -ACV of $40,000-$50,000. - -=== Macro Tailwinds - -Several macro-economic trends favour VeriSimDB's positioning: - -AI adoption driving multi-system architectures:: - Every AI/ML pipeline adds at least one new data store (vector database for - embeddings, graph database for knowledge graphs). This increases the number of - systems holding representations of the same entity, making cross-modal drift - more likely and more costly. - -Regulatory pressure on data consistency:: - GDPR (right to erasure across all systems), FDA 21 CFR Part 11 (electronic - records integrity), SOX (financial data accuracy), and emerging AI regulations - (EU AI Act) all create compliance requirements for cross-system data consistency. - Organisations that cannot demonstrate consistency face fines and audit failures. - -Cloud migration creating data fragmentation:: - Organisations migrating to cloud adopt specialised managed services (RDS, - DynamoDB, Neptune, OpenSearch, Pinecone) rather than consolidating on a single - database. This polyglot persistence pattern is the root cause of cross-modal - drift. - -Open source as default infrastructure:: - The database market has shifted to open-source-first adoption. PostgreSQL, - MongoDB, Elasticsearch, and Redis are all open source or source-available. - VeriSimDB's open-core model aligns with this expectation. - -=== Macro Headwinds - -Market education required:: - "Cross-modal drift detection" is not a recognised category. Potential customers - do not know they have this problem until they see it demonstrated. Category - creation is expensive and time-consuming. - -Consolidation pressure:: - Large database vendors (MongoDB, Snowflake, Databricks) are adding multi-model - capabilities. They may eventually add drift detection as a feature, competing - with VeriSimDB using existing market position. - -Economic uncertainty:: - IT budget constraints during economic downturns delay new infrastructure - adoption. Data quality is often deprioritised relative to revenue-generating - features. - -== Micro-Economic Analysis - -=== Competitive Pricing Analysis - -[cols="2,2,2,3"] -|=== -| Product | Free Tier | Commercial Price | Notes - -| *Neo4j* -| Community Edition (GPLv3) -| Enterprise: from $36,000/year; AuraDB: from $65/month -| Community Edition lacks clustering, role-based access, and monitoring. -Enterprise features are gated. - -| *Pinecone* -| Starter: 1 index, 100K vectors -| Standard: from $70/month; Enterprise: custom pricing (typically $25K+/year) -| Fully managed, cloud-only. No self-hosted option. Pricing scales with vector -count and query volume. - -| *Weaviate* -| Open source (BSD-3) -| Weaviate Cloud: from $25/month; Enterprise: custom pricing -| Open source is full-featured. Cloud service adds managed infrastructure. -Enterprise adds SLA and support. - -| *ArangoDB* -| Community Edition (Apache 2.0) -| ArangoGraph: from $99/month; Enterprise: custom pricing (est. $30K+/year) -| Community Edition is full-featured. Enterprise adds SmartGraphs, satellite -collections, and dedicated support. - -| *SurrealDB* -| Open source (BSL 1.1) -| SurrealDB Cloud: beta (pricing TBD); Enterprise: not yet available -| BSL restricts commercial use (converts to Apache 2.0 after 3 years). -Cloud service is in beta. - -| *Qdrant* -| Open source (Apache 2.0) -| Qdrant Cloud: from $25/month; Enterprise: custom pricing -| Full-featured open source. Cloud adds managed infrastructure. - -| *VeriSimDB* -| Community Edition (PMPL) -| Standard: $25,000/year; Professional: $40,000/year; Enterprise: $75,000/year -| Full-featured community edition. Commercial licence provides support and SLA, -not additional features. -|=== - -=== Pricing Strategy - -==== Principle: No Feature Gating - -VeriSimDB's commercial licence does *not* add features beyond the community -edition. All 8 modalities, drift detection, self-normalisation, VCL, VCL-UT, -and federation ship in the open-source release. - -This is a deliberate strategic choice: - -1. *Maximises adoption*: Developers can evaluate the full product without hitting - artificial limits. There is no "upgrade to Enterprise for clustering" moment - that creates friction. - -2. *Builds trust*: The open-source community is increasingly sceptical of - "open core" models that gate essential features. Full-featured community - editions (Weaviate, ArangoDB Community) earn developer trust. - -3. *Targets operational value*: Enterprise customers pay for SLAs, priority - patches, deployment consulting, and expert access -- not for features. The - operational value of "24-hour security patch response" is worth $40,000/year - to a financial services firm; the feature itself (the patch) is free. - -==== Pricing Tiers - -[cols="1,1,4"] -|=== -| Tier | ACV | Value Proposition - -| *Standard* -| $25,000 -| For teams that need a support safety net. Email support with 48-hour response. -Security patches within 7 days. Quarterly architecture review call. -"Insurance policy" positioning. - -| *Professional* -| $40,000 -| For teams running VeriSimDB in production. Priority support with 24-hour -response. Security patches within 48 hours. Monthly architecture review. -8 hours/quarter of deployment consulting included. -"Operational confidence" positioning. - -| *Enterprise* -| $75,000 -| For organisations with strict compliance requirements. Dedicated support -channel with 4-hour response. Security patches within 24 hours. Weekly -architecture review. Unlimited consulting. On-site training (2 days/year). -"Compliance-grade support" positioning. -|=== - -==== Pricing Rationale - -The $40,000 Professional tier is the target ACV, positioned as follows: - -* *Below* Neo4j Enterprise ($36K+ but typically $50K-$100K with all modules), - making it easy for prospects to justify budget -* *Above* Weaviate Cloud and ArangoDB Cloud managed service costs, reflecting - the additional value of drift detection and formal verification -* *Competitive* with Pinecone Enterprise pricing ($25K+/year), which provides - only vector search (1 modality) versus VeriSimDB's 8 modalities -* *Aligned* with industry benchmarks for open-source database support (Red Hat - OpenShift subscriptions, Canonical Ubuntu Pro, Elastic subscriptions) - -=== Unit Economics - -[cols="2,1,3"] -|=== -| Metric | Value | Derivation - -| *Average Contract Value (ACV)* -| $40,000 -| Professional tier, most common entry point - -| *Customer Acquisition Cost (CAC)* -| $8,000 -| Developer-led adoption reduces traditional enterprise sales cost; estimated -as 3 months of DevRel + content costs per conversion - -| *LTV:CAC Ratio* -| 15:1 -| 3-year average retention at $40K ACV = $120K LTV; $120K / $8K = 15:1. -Target > 3:1 is considered healthy. - -| *Payback Period* -| 2.4 months -| $8K CAC / ($40K ACV / 12 months) = 2.4 months. Target < 12 months. - -| *Gross Margin* -| 85% -| Software + support. COGS is primarily support engineer time (15% of ACV). - -| *Net Revenue Retention (NRR)* -| 115% -| Expansion from Standard to Professional to Enterprise, plus consulting -add-ons. 5% annual churn offset by 20% expansion. - -| *Annual Churn* -| 5% -| Conservative estimate. Open-source databases have lower churn than proprietary -because customers continue using the free edition even if they cancel support. - -| *Customer Lifetime* -| 3 years -| Average enterprise support contract duration. -|=== - -=== Break-Even Analysis - -[cols="1,2,2"] -|=== -| Scenario | Customers Needed | Revenue - -| *Minimum viable* (1 employee + founder) -| 3 Professional contracts -| $120,000 (covers 1 DevRel hire + infrastructure) - -| *Stage 1 team* (3-5 people) -| 8-10 mixed contracts -| $300,000-$400,000 (covers core team compensation) - -| *Self-sustaining* (10-15 people) -| 20-25 mixed contracts -| $800,000-$1,000,000 (covers full org + overhead) -|=== - -=== Revenue Projection - -[cols="1,1,1,1,1"] -|=== -| | Year 1 | Year 2 | Year 3 | Year 4 - -| *Community users* | 500 | 2,000 | 5,000 | 10,000 -| *Standard contracts* | 0 | 2 | 5 | 10 -| *Professional contracts* | 0 | 1 | 5 | 15 -| *Enterprise contracts* | 0 | 0 | 1 | 3 -| *Consulting engagements* | 0 | 3 | 5 | 8 -| *Consulting revenue* | $0 | $45,000 | $75,000 | $120,000 -| *Licence revenue* | $0 | $90,000 | $325,000 | $775,000 -| *Total revenue* | $0 | $135,000 | $400,000 | $895,000 -| *Total costs* | $3,000 | $104,000 | $365,000 | $700,000 -| *Operating profit* | -$3,000 | $31,000 | $35,000 | $195,000 -|=== - -The conservative projection reaches *operating profitability in Year 2* through -consulting revenue, with licence revenue becoming the primary driver by Year 3. - -== Sensitivity Analysis - -=== Best Case - -* Community adoption exceeds expectations (GitHub trending, viral blog post) -* First enterprise customer acquired in Year 1 -* Year 3 revenue: $800,000 (double the conservative projection) - -=== Worst Case - -* Community adoption is slow (< 200 stars in Year 1) -* No enterprise customers until Year 3 -* Year 3 revenue: $100,000 (consulting only) -* Mitigation: VeriSimDB remains a viable open-source project even without - commercial revenue; the founder's academic position provides baseline income - -=== Key Variables - -[cols="2,3,1"] -|=== -| Variable | Impact | Sensitivity - -| Community adoption rate -| Determines pipeline for enterprise conversion -| High - -| Enterprise sales cycle length -| 6-12 months typical; longer delays revenue -| High - -| Competitor response (Neo4j adds drift detection) -| Reduces differentiation but VeriSimDB has -architectural depth advantage (octad + VCL-UT) -| Medium - -| Pricing acceptance ($40K ACV) -| May need to lower to $25K for early adopters -| Medium - -| Churn rate -| 5% assumed; 10% would reduce LTV by 33% -| Medium -|=== - -== Investment Considerations - -=== Self-Funded Path - -VeriSimDB can reach sustainability without external investment: - -* Year 1: Self-funded ($3,000 budget) -* Year 2: First revenue from consulting ($135,000); hire DevRel lead -* Year 3: Licence revenue begins ($400,000); hire core engineering team -* Year 4: Self-sustaining ($895,000 revenue, $195,000 operating profit) - -This path is slower but preserves full ownership and avoids the pressure to -scale prematurely. - -=== External Funding (If Pursued) - -If external funding is sought, the appropriate stage would be: - -* *Pre-seed / Angel* ($100K-$250K, Year 1-2): Accelerate community building, - hire DevRel lead 6 months earlier, attend more conferences -* *Seed* ($500K-$1M, Year 2-3): Hire full Stage 1 team, accelerate enterprise - sales, develop managed service -* *Series A* ($3M-$5M, Year 3-4): Scale to Stage 3 organisation, launch - managed service, expand internationally - -The self-funded path is the current plan. External funding is documented here -for completeness, not as a recommendation. diff --git a/verisimdb/docs/business/strategy/functional-units.adoc b/verisimdb/docs/business/strategy/functional-units.adoc deleted file mode 100644 index 2b6e7f48..00000000 --- a/verisimdb/docs/business/strategy/functional-units.adoc +++ /dev/null @@ -1,362 +0,0 @@ -// SPDX-License-Identifier: CC-BY-SA-4.0 -// Copyright (c) 2026 Jonathan D.A. Jewell (hyperpolymath) -= VeriSimDB -- Functional Units and Organisational Structure -Jonathan D.A. Jewell -:revdate: 2026-02-28 -:revremark: Organisational structure to support VeriSimDB at each growth stage -:toc: macro -:toclevels: 3 -:sectnums: -:icons: font - -[abstract] -Organisational structure required to develop, market, support, and operate -VeriSimDB at each growth stage. The structure evolves from a single-developer -open source project through a small focused team to a sustainable organisation -capable of supporting enterprise customers. - -toc::[] - -== Current State (Solo Developer) - -VeriSimDB is currently developed and maintained by a single developer -(Jonathan D.A. Jewell). All functions -- engineering, documentation, testing, -community engagement, infrastructure, and strategy -- are performed by one -person. - -This is viable for the community-building phase (Phase 1) but will not scale -to support enterprise customers, multiple pilot deployments, or sustained -conference presence. - -The functional unit plan below describes the minimal team needed at each growth -stage, not an aspirational org chart. Each role is justified by a specific -bottleneck that a single developer cannot resolve alone. - -== Stage 1: Core Team (3-5 People, Months 12-18) - -=== Engineering (2-3 People) - -==== Rust Core Engineer (1 person) - -*Justification:* The Rust core (storage engine, drift detection, indexing) is -the most performance-critical and complex subsystem. A dedicated Rust engineer -enables parallel development of modality stores, query optimisation, and -federation connectors while the founder focuses on architecture and community. - -*Responsibilities:* - -* Develop and maintain Rust crates (`verisim-graph`, `verisim-vector`, - `verisim-drift`, `verisim-normalizer`, etc.) -* Performance benchmarking and optimisation -* Federation connector development (PostgreSQL, ArangoDB, Elasticsearch, Neo4j) -* CI/CD pipeline maintenance (Rust builds, cross-compilation) - -*Skills required:* - -* Strong Rust experience (3+ years), including async and FFI -* Database internals knowledge (indexing, query planning, storage engines) -* Experience with at least one graph or vector database - -*Hiring signal:* When the first external contributor submits a non-trivial Rust -PR, the candidate pool includes people who already understand the codebase. - -==== Elixir/OTP Engineer (1 person) - -*Justification:* The Elixir orchestration layer (GenServer-per-entity, -supervision trees, distributed coordination) requires OTP expertise that is -distinct from Rust systems programming. An Elixir engineer enables development -of advanced fault tolerance, distributed consensus, and hot code reloading -capabilities. - -*Responsibilities:* - -* Develop and maintain Elixir orchestration layer (`VeriSim.EntityServer`, - `VeriSim.DriftMonitor`, `VeriSim.QueryRouter`) -* Hypatia integration pipeline (`ScanIngester`, `PatternQuery`, `DispatchBridge`) -* Distributed coordination (multi-node deployment, consensus) -* Monitoring and observability (Prometheus metrics, Telemetry events) - -*Skills required:* - -* Elixir/OTP experience (2+ years), including GenServer, Supervisor, and - distributed Erlang -* Understanding of BEAM concurrency model -* Experience with Phoenix or similar Elixir web frameworks (for HTTP API) - -==== Generalist / Zig-Idris2 Specialist (0-1 person, contract) - -*Justification:* The ABI/FFI layer (Idris2 definitions, Zig bindings, generated -C headers) requires a rare skill combination. This role may be filled by a -contractor or part-time contributor rather than a full-time hire, given the -narrow scope. - -*Responsibilities:* - -* Maintain Idris2 ABI definitions and dependent-type proofs -* Develop and maintain Zig FFI bindings -* Generate and validate C headers -* Ensure ABI/FFI compliance across platforms - -*Skills required:* - -* Idris2 or Agda experience (dependent types) -* Zig experience (C ABI, cross-compilation) -* Understanding of FFI patterns and memory management - -=== Developer Relations (1 person) - -==== DevRel Lead - -*Justification:* The single-biggest bottleneck for Phase 2 (institutional -partnerships) is the founder's time. A DevRel lead handles documentation, -community engagement, conference submissions, and content creation -- freeing -the founder for architecture decisions and enterprise conversations. - -*Responsibilities:* - -* Write and maintain documentation (getting-started guide, tutorials, API - reference) -* Create content (blog posts, conference talks, video tutorials) -* Manage community channels (GitHub issues, forum discussions, social media) -* Submit conference talk proposals and prepare presentations -* Coordinate with academic collaborators on paper submissions - -*Skills required:* - -* Technical writing ability (can explain database concepts to developers) -* Conference speaking experience -* Familiarity with database or data engineering communities -* Ability to write code examples in at least one of Rust, Elixir, or ReScript - -*Hiring signal:* When community engagement (GitHub issues, forum posts) exceeds -what one person can respond to within 48 hours. - -=== Operations (0.5 person, shared with Engineering) - -In Stage 1, operations responsibilities are shared across the engineering team: - -* CI/CD pipeline maintenance (GitHub Actions, SLSA, SBOM) -* Container image builds and registry management -* Infrastructure provisioning (staging environment for demos) -* Security patching and dependency updates - -No dedicated operations hire is needed until enterprise customers require SLA -guarantees. - -== Stage 2: Growth Team (6-10 People, Months 18-30) - -=== Engineering (4-5 People) - -The Stage 1 engineering team expands with: - -==== Query Engine Engineer (1 person) - -*Justification:* VCL and VCL-UT require dedicated attention as query complexity -grows. The query planner (currently basic) needs optimisation for cross-modal -queries spanning all 8 modalities. - -*Responsibilities:* - -* VCL parser and query planner development -* VCL-UT type checker integration (Idris2 bridge) -* Query performance optimisation (cost-based planning, modality pushdown) -* Query language evolution (new syntax, new operators) - -==== Test and Quality Engineer (1 person) - -*Justification:* The test suite (662 tests) needs to scale with the codebase. -Property-based testing, fuzzing, and integration testing across federation -targets require dedicated attention. - -*Responsibilities:* - -* Property-based test development (QuickCheck patterns for Rust + Elixir) -* Fuzz testing (VCL parser, drift detection algorithm) -* Integration testing (federation connectors, container deployment) -* Performance benchmarking (automated regression detection) - -=== Support (1 person) - -==== Support Engineer - -*Justification:* Enterprise pilot customers (Phase 2) require responsive -support. The 48-hour response SLA cannot be maintained by the engineering team -while also developing features. - -*Responsibilities:* - -* Triage and respond to customer support requests -* Reproduce and document bugs from customer environments -* Create knowledge base articles from recurring issues -* Escalate complex issues to engineering with reproduction steps - -*Skills required:* - -* Database administration experience -* Comfortable with Rust and Elixir at a reading level -* Strong written communication -* Experience with support ticketing systems - -=== Marketing (1 person, part-time or contract) - -==== Content Marketing Lead - -*Justification:* Phase 3 (enterprise expansion) requires professional marketing -materials: case studies, ROI calculators, compliance guides, and analyst -briefings. These are distinct from DevRel content and require marketing -expertise. - -*Responsibilities:* - -* Create enterprise-facing content (case studies, white papers, compliance guides) -* Manage website and landing pages -* Coordinate analyst briefings (Gartner, Forrester, IDC) -* Event marketing (conference booths, sponsorships) - -*Note:* This role may be a contractor or part-time hire in Stage 2, becoming -full-time in Stage 3. - -== Stage 3: Sustainable Organisation (10-15 People, Months 30-48) - -=== Engineering (6-8 People) - -The Stage 2 engineering team expands with: - -* *Second Rust engineer* -- dedicated to federation connectors and new modality - backends -* *Security engineer* -- dedicated to security hardening, vulnerability response, - and compliance (SOC 2, ISO 27001) -* *Infrastructure engineer* -- dedicated to managed service deployment, - monitoring, and SRE practices - -=== Developer Relations (2 People) - -* *DevRel Lead* (from Stage 1) -* *Developer Advocate* -- focused on community growth, meetup organising, - open-source contribution mentorship - -=== Support (2 People) - -* *Support Engineer* (from Stage 2) -* *Solutions Engineer* -- pre-sales technical support, architecture reviews, - proof-of-concept deployments for enterprise prospects - -=== Marketing (1-2 People) - -* *Content Marketing Lead* (from Stage 2, now full-time) -* *Events Coordinator* (0-1 person) -- conference logistics, sponsorship - management, analyst relations - -=== Operations (1-2 People) - -* *SRE / Infrastructure Engineer* -- managed service operations, monitoring, - alerting, incident response -* *Finance and Admin* (0.5-1 person, possibly outsourced) -- invoicing, - contracts, compliance documentation - -== Functional Unit Summary - -[cols="2,1,1,1"] -|=== -| Function | Stage 1 | Stage 2 | Stage 3 - -| Engineering | 2-3 | 4-5 | 6-8 -| Developer Relations | 1 | 1 | 2 -| Support | 0 | 1 | 2 -| Marketing | 0 | 0.5-1 | 1-2 -| Operations | 0.5 | 0.5 | 1-2 -| *Total* | *3-5* | *7-8* | *12-16* -|=== - -== Compensation Framework - -Compensation ranges are based on UK/remote market rates for the required skill -levels. VeriSimDB competes for talent against well-funded database startups -(SurrealDB, Qdrant, LanceDB) and established companies (Neo4j, DataStax). - -[cols="2,1,3"] -|=== -| Role | Range (GBP, annual) | Notes - -| Rust Core Engineer | 60,000-90,000 | Senior level required; Rust + database internals is a rare combination -| Elixir/OTP Engineer | 55,000-80,000 | BEAM expertise commands a premium -| Query Engine Engineer | 65,000-95,000 | PL theory + database systems intersection -| DevRel Lead | 50,000-75,000 | Technical writing + speaking + community management -| Support Engineer | 40,000-60,000 | Database admin + communication skills -| Content Marketing | 45,000-65,000 | B2B SaaS marketing experience -|=== - -Early hires may accept below-market compensation in exchange for equity or -significant open-source contribution credit. However, VeriSimDB should not -rely on below-market compensation as a strategy -- it creates retention risk -and signals unsustainability. - -== Hiring Priorities - -Hiring order is determined by the bottleneck that most constrains growth at -each stage: - -[cols="1,2,3"] -|=== -| Priority | Role | Bottleneck Addressed - -| 1 | DevRel Lead | Founder cannot maintain 48h issue response, write documentation, -create content, AND develop the core engine simultaneously - -| 2 | Rust Core Engineer | Federation connector development and query optimisation -are blocked by single-developer bandwidth - -| 3 | Elixir/OTP Engineer | Distributed coordination and Hypatia integration -require dedicated OTP expertise - -| 4 | Support Engineer | Enterprise pilot customers require responsive support -that engineering should not be distracted by - -| 5 | Content Marketing | Enterprise expansion requires professional marketing -materials beyond what DevRel produces -|=== - -== Reporting Structure - -=== Stage 1 (Flat) - -.... -Jonathan D.A. Jewell (Founder, Architecture Lead) -├── Rust Core Engineer -├── Elixir/OTP Engineer -└── DevRel Lead -.... - -=== Stage 2 (Two Pillars) - -.... -Jonathan D.A. Jewell (Founder, CTO) -├── Engineering Lead (Rust Core Engineer, promoted) -│ ├── Elixir/OTP Engineer -│ ├── Query Engine Engineer -│ └── Test & Quality Engineer -├── DevRel Lead -│ └── Content Marketing (contract) -└── Support Engineer -.... - -=== Stage 3 (Functional) - -.... -Jonathan D.A. Jewell (Founder, CEO/CTO) -├── VP Engineering -│ ├── Rust Core Team (2-3) -│ ├── Elixir/OTP Engineer -│ ├── Query Engine Engineer -│ ├── Security Engineer -│ └── Test & Quality Engineer -├── Head of DevRel -│ ├── Developer Advocate -│ └── Content Marketing -├── Head of Customer Success -│ ├── Support Engineer -│ └── Solutions Engineer -└── Operations - ├── SRE / Infrastructure - └── Finance & Admin (outsourced) -.... diff --git a/verisimdb/docs/business/strategy/go-to-market.adoc b/verisimdb/docs/business/strategy/go-to-market.adoc deleted file mode 100644 index 01bb0d13..00000000 --- a/verisimdb/docs/business/strategy/go-to-market.adoc +++ /dev/null @@ -1,447 +0,0 @@ -// SPDX-License-Identifier: CC-BY-SA-4.0 -// Copyright (c) 2026 Jonathan D.A. Jewell (hyperpolymath) -= VeriSimDB -- Go-to-Market Strategy -Jonathan D.A. Jewell -:revdate: 2026-02-28 -:revremark: Three-phase GTM strategy for community-led growth -:toc: macro -:toclevels: 3 -:sectnums: -:icons: font - -[abstract] -Go-to-market strategy for VeriSimDB, structured in three phases: community-led -open source adoption, institutional partnerships in research and biomedicine, -and enterprise expansion in financial services and manufacturing. The strategy -is developer-led and bottom-up, reflecting VeriSimDB's current stage (alpha, -single developer) and the need to build credibility before pursuing enterprise -contracts. - -toc::[] - -== Strategic Positioning - -=== Category Definition - -VeriSimDB creates a new category: *cross-modal consistency infrastructure*. -This is distinct from: - -* *Multi-model databases* (ArangoDB, SurrealDB) -- store multiple models but do - not check consistency between them -* *Data quality tools* (Great Expectations, Monte Carlo) -- validate data within - a single system, not across systems -* *Data integration platforms* (Fivetran, Airbyte) -- move data between systems - but do not monitor ongoing consistency - -The category-creation approach is deliberate. Competing within an established -category (graph databases, vector databases) positions VeriSimDB against -well-funded incumbents on their home turf. Creating a new category positions -VeriSimDB as the reference implementation for a problem that incumbents -do not address. - -=== Positioning Statement - -For *data engineers and platform teams* who struggle with *cross-system entity -inconsistency*, VeriSimDB is a *multimodal database with drift detection* that -*detects and repairs divergent representations automatically*. Unlike *multi-model -databases* (ArangoDB, SurrealDB) that store multiple models without consistency -checking, VeriSimDB *continuously monitors cross-modal coherence and -self-normalises drifted data*. - -== Phase 1: Community-Led Growth (Months 0-12) - -=== Objective - -Establish VeriSimDB as the recognised solution for cross-modal drift detection -within developer communities where the problem is acutely felt. - -=== Target Communities - -[cols="1,2,3,1"] -|=== -| Community | Why They Care | Entry Point | Priority - -| *GraphRAG engineers* -| Vector embeddings drift from knowledge graph state, causing LLM hallucinations -| Blog post: "Detecting Embedding Drift in RAG Pipelines" -| Highest - -| *Elixir/BEAM developers* -| OTP orchestration layer is a natural talking point; community is small but -highly engaged -| ElixirForum post, Code BEAM talk proposal -| High - -| *Rust systems programmers* -| Rust core engine, no C/C++ dependencies, pure-Rust graph backend -| r/rust post, RustConf/EuroRust talk proposal -| High - -| *Data quality practitioners* -| Existing tools (Great Expectations, Evidently) detect statistical drift, not -cross-modal drift -| Comparison blog post, Data Council talk proposal -| Medium - -| *Formal methods enthusiasts* -| VCL-UT dependent-type proofs are novel in a database context -| Idris2 Discourse post, TYPES mailing list -| Medium -|=== - -=== Channels - -==== GitHub - -* Repository: https://github.com/hyperpolymath/verisimdb -* Actions: Comprehensive README with architecture diagram, getting-started guide, - VCL examples, contribution guidelines -* Metrics: Stars, forks, issues, PRs, unique cloners - -==== Blog Posts (Hosted on GitHub Pages) - -Planned content calendar (first 6 months): - -[cols="1,3,2"] -|=== -| Month | Title | Target Audience - -| 1 | "Introducing VeriSimDB: Cross-Modal Drift Detection for Multimodal Data" -| General developer audience - -| 2 | "Detecting Embedding Drift in RAG Pipelines with VeriSimDB" -| GraphRAG engineers - -| 3 | "VeriSimDB vs. Great Expectations vs. Monte Carlo: Different Problems" -| Data quality practitioners - -| 4 | "Building a Fault-Tolerant Database Orchestrator with Elixir/OTP" -| Elixir community - -| 5 | "Proof-Carrying Queries: Dependent Types Meet Database Systems" -| Formal methods, PL researchers - -| 6 | "The Octad Model: Why 8 Modalities?" -| Database architects -|=== - -==== Conference Talks - -Target conferences for Year 1: - -[cols="2,1,1,2"] -|=== -| Conference | Type | When | Angle - -| *Code BEAM* (EU/US) | BEAM/Elixir | H2 2026 | OTP orchestration architecture -| *EuroRust / RustConf* | Rust | H2 2026 | Pure-Rust graph backend, no C/C++ deps -| *FOSDEM* (Database Devroom) | Open source | Feb 2027 | Multimodal DB architecture -| *Data Council* | Data engineering | H1 2027 | Cross-modal drift as data quality problem -| *ElixirConf* (EU/US) | Elixir | H2 2026 | GenServer-per-entity architecture -|=== - -==== Community Engagement - -* Respond to every GitHub issue within 48 hours -* Engage in relevant Reddit/HN/Lobsters threads (with disclosure) -* Contribute to adjacent projects (LangChain, Qdrant) where VeriSimDB adds value -* Publish VCL examples and tutorials in `docs/` - -=== Success Metrics (Month 12) - -[cols="1,1,3"] -|=== -| Metric | Target | Rationale - -| GitHub stars | 500+ | Indicates awareness; top 20% of database repos on GitHub -| Monthly unique cloners | 100+ | Indicates evaluation activity -| Open issues | 20+ | Indicates engagement (bugs + feature requests) -| External PRs | 5+ | Indicates contributor interest -| Blog post views | 10,000+ cumulative | Indicates content reach -| Conference talks given | 3+ | Indicates community credibility -|=== - -== Phase 2: Biomedical and Research Institutions (Months 12-24) - -=== Objective - -Secure 2-3 pilot deployments with research institutions where data quality is -critical and the octad model provides clear value. - -=== Target Institutions - -[cols="2,3,2"] -|=== -| Sector | Why VeriSimDB | Entry Point - -| *Bioinformatics labs* -| Multi-omics data integration requires consistency across genomic, proteomic, -and metabolomic representations -| ELIXIR network, ISMB/BOSC conference - -| *Clinical research* -| Patient records span EHR, pharmacy, lab, and imaging systems with -life-safety implications for inconsistency -| OpenHIE community, AMIA conference - -| *Climate science* -| Cross-sensor calibration verification and multi-archive federation -| Pangeo community, ESA CCI - -| *Digital humanities* -| Knowledge graph quality assurance at scale (Wikidata, ConceptNet) -| Wikidata Workshop, ISWC conference -|=== - -=== Engagement Model - -1. *Academic paper submission* (VLDB, SIGMOD, or EDBT) establishing the - theoretical contribution -2. *Tutorial at domain conference* (ISMB for bioinformatics, AMIA for clinical) - demonstrating VeriSimDB with domain-specific data -3. *Pilot partnership* with 1-2 labs, providing free deployment support in - exchange for a published evaluation -4. *Grant applications* co-authored with institution partners (UKRI, Horizon - Europe, NIH data commons) - -=== Revenue Model (Phase 2) - -Revenue in Phase 2 comes from *consulting engagements*, not software licences: - -* Architecture review and deployment planning: $5,000-$15,000 per engagement -* Custom federation connector development: $10,000-$25,000 per connector -* Training workshops for research teams: $2,000-$5,000 per workshop - -Target: 3-5 consulting engagements, $50,000-$150,000 total revenue. - -=== Academic Publication Strategy - -[cols="1,2,2"] -|=== -| Venue | Paper Type | Angle - -| *VLDB* -| Research paper (12 pages) -| Cross-modal drift detection: theory, algorithm, evaluation - -| *SIGMOD* -| Demo paper (4 pages) -| Live demo: 1000 entities, 50 corrupted, 100% detection - -| *EDBT* -| Vision paper (6 pages) -| The octad model and multimodal database taxonomy - -| *CIDR* -| Systems paper (12 pages) -| Rust + Elixir/OTP architecture for multimodal databases - -| *ICFP/POPL* -| Experience report (12 pages) -| VCL-UT: dependent types in database query languages -|=== - -=== Success Metrics (Month 24) - -[cols="1,1,3"] -|=== -| Metric | Target | Rationale - -| Pilot deployments | 2-3 | Validates product-market fit in research sector -| Published papers | 1-2 | Establishes academic credibility -| Consulting revenue | $100K+ | Validates willingness to pay -| GitHub stars | 2,000+ | Growing community awareness -| External contributors | 10+ | Sustainable OSS community -|=== - -== Phase 3: Enterprise Expansion (Months 24-36) - -=== Objective - -Launch commercial support licences and expand into enterprise markets where -cross-system data consistency has measurable financial impact. - -=== Target Verticals - -[cols="2,3,2,1"] -|=== -| Vertical | Pain Point | VeriSimDB Value | ACV - -| *Financial services* -| AML/KYC compliance requires cross-system entity consistency; regulatory -penalties for inconsistent customer data -| Federation + drift detection across trading, compliance, and risk systems -| $50K+ - -| *Manufacturing* -| Digital twin synchronisation; IoT sensor data consistency across simulation -and physical state -| Temporal + spatial + provenance modalities; drift = simulation divergence -| $40K+ - -| *Pharmaceuticals* -| FDA 21 CFR Part 11 compliance; multi-system clinical trial data consistency -| Provenance + proof certificates for audit trails; drift = regulatory risk -| $60K+ - -| *Supply chain* -| End-to-end traceability; certification expiry detection across partner systems -| Federation + provenance + temporal modalities -| $35K+ -|=== - -=== Go-to-Market Motion - -==== Developer-Led (Bottom-Up) - -The expected enterprise sales motion for VeriSimDB is bottom-up: - -1. A data engineer discovers VeriSimDB through a blog post, conference talk, - or community recommendation -2. The engineer evaluates VeriSimDB in a staging environment, running drift - detection on a subset of production data -3. The engineer demonstrates results to their team lead or VP Engineering, - showing quantified drift across existing systems -4. The team requests an architecture review (consulting engagement, $5K-$15K) -5. Successful evaluation leads to a commercial support licence ($40K+/year) - -This motion has a *low CAC* ($8,000 estimated) because the developer does the -evaluation and internal selling. Traditional enterprise sales (outbound SDR, -demo, POC, procurement) costs $30,000-$50,000 per customer. - -==== Partner Channel (Secondary, Phase 3b) - -After establishing 5-10 commercial customers through developer-led adoption: - -* Engage system integrators (Accenture, Deloitte, Thoughtworks) who advise - enterprise clients on data infrastructure -* Create partner training programme (VCL certification) -* Provide referral incentives (10-15% of first-year ACV) - -=== Commercial Support Licence - -[cols="1,1,3"] -|=== -| Tier | ACV | Includes - -| *Standard* -| $25,000 -| Email support (48h response), security patches (7-day response), quarterly -architecture review call - -| *Professional* -| $40,000 -| Priority email + Slack support (24h response), security patches (48h response), -monthly architecture review, deployment consulting (8 hours/quarter) - -| *Enterprise* -| $75,000 -| Dedicated support channel (4h response), security patches (24h response), -weekly architecture review, unlimited deployment consulting, on-site training -(2 days/year) -|=== - -All tiers include the same software. No feature gating. The commercial licence -provides operational guarantees and expert access, not additional functionality. - -=== Success Metrics (Month 36) - -[cols="1,1,3"] -|=== -| Metric | Target | Rationale - -| Commercial customers | 10-15 | Validates enterprise product-market fit -| ARR | $400K-$600K | Path to sustainability without external funding -| GitHub stars | 5,000+ | Top-tier OSS database project recognition -| External contributors | 25+ | Self-sustaining community -| Published papers | 3-5 | Ongoing academic credibility -| Conference talks | 10+ cumulative | Established speaker in database conferences -|=== - -== Risk Mitigation - -=== Phase 1 Risks - -[cols="2,3,3"] -|=== -| Risk | Impact | Mitigation - -| Low community adoption -| No awareness, no pipeline for Phase 2 -| Focus on GraphRAG community (strongest alignment); create compelling demo; -engage consistently (48h issue response) - -| Competitor adds drift detection -| Reduces differentiation -| 8-modality octad model is architecturally difficult to retrofit; deep moat in -VCL-UT formal verification; first-mover advantage in category definition -|=== - -=== Phase 2 Risks - -[cols="2,3,3"] -|=== -| Risk | Impact | Mitigation - -| Academic paper rejected -| Delays credibility building -| Submit to multiple venues; prepare demo paper as backup; publish preprint on -arXiv regardless of venue acceptance - -| No institution willing to pilot -| Consulting revenue delayed -| Offer free deployment support. Lower barrier: "let us run drift detection -on your data for free." Publication incentive for academics. -|=== - -=== Phase 3 Risks - -[cols="2,3,3"] -|=== -| Risk | Impact | Mitigation - -| Enterprise procurement delay (6-12 months) -| Revenue delayed; cash flow pressure -| Consulting revenue bridges gap. Developer-led adoption reduces procurement -friction. Start enterprise conversations 6 months before Phase 3. - -| Single-founder scaling limit -| Cannot support 10+ enterprise customers alone -| Early revenue funds first hire (DevRel or support engineer); comprehensive -documentation and automation reduce per-customer effort. -|=== - -== Budget Summary - -[cols="2,1,1,1"] -|=== -| Item | Phase 1 | Phase 2 | Phase 3 - -| Infrastructure (hosting, CI/CD) -| $500 -| $2,000 -| $10,000 - -| Conference travel -| $2,000 -| $5,000 -| $10,000 - -| Marketing (domain, design) -| $500 -| $1,000 -| $5,000 - -| Personnel (contract) -| $0 -| $0 -| $50,000 - -| *Total* -| *$3,000* -| *$8,000* -| *$75,000* -|=== - -Phase 1 and 2 are self-funded. Phase 3 personnel costs are covered by -commercial support revenue. diff --git a/verisimdb/docs/cache-sharing-strategy.adoc b/verisimdb/docs/cache-sharing-strategy.adoc deleted file mode 100644 index 62e8fe6a..00000000 --- a/verisimdb/docs/cache-sharing-strategy.adoc +++ /dev/null @@ -1,745 +0,0 @@ -// SPDX-License-Identifier: CC-BY-SA-4.0 - -= Cache Sharing Strategy: Local vs Federated -:toc: left -:toclevels: 3 -:sectnums: - -== Overview - -VeriSimDB's cache is **tiered** (L1, L2, L3) and **selectively shared** depending on deployment mode and security requirements. - -**Key Principle:** Share what's safe and beneficial, keep private what's sensitive or node-specific. - -== Three-Tier Architecture - -[source,text] ----- -┌─────────────────────────────────────────────────────────────┐ -│ L1: Per-Node In-Memory Cache (ETS) │ -│ - NOT shared across nodes │ -│ - Each node has its own L1 │ -│ - Fastest: <1ms access │ -│ - Use: Hot data for THIS node │ -└─────────────────────────────────────────────────────────────┘ - │ - ▼ -┌─────────────────────────────────────────────────────────────┐ -│ L2: Distributed Cache (Shared Across Cluster) │ -│ - SHARED within same organization/trust boundary │ -│ - NOT shared across federation (different orgs) │ -│ - Medium speed: <10ms access │ -│ - Use: Warm data shared within trusted cluster │ -│ - Implementation: Raft-based or Redis cluster │ -└─────────────────────────────────────────────────────────────┘ - │ - ▼ -┌─────────────────────────────────────────────────────────────┐ -│ L3: Persistent Cache (verisim-temporal) │ -│ - CAN be shared (depends on policy) │ -│ - Public data: Shared across federation │ -│ - Private data: Local only │ -│ - Slow: <100ms access │ -│ - Use: Cold data, historical queries │ -└─────────────────────────────────────────────────────────────┘ ----- - -== Multi-Organization Architecture - -This diagram shows how cache layers are organized across multiple organizations in a federated deployment: - -[source,text] ----- -┌──────────────────────────────────────────────────────────────────────────────┐ -│ ORGANIZATION A (University) │ -│ ┌────────────┐ ┌────────────┐ ┌────────────┐ │ -│ │ Node A1 │ │ Node A2 │ │ Node A3 │ │ -│ ├────────────┤ ├────────────┤ ├────────────┤ │ -│ │ L1: ETS │ │ L1: ETS │ │ L1: ETS │ ← Per-node, NOT shared │ -│ │ 1GB cache │ │ 1GB cache │ │ 1GB cache │ │ -│ │ <1ms │ │ <1ms │ │ <1ms │ │ -│ └─────┬──────┘ └─────┬──────┘ └─────┬──────┘ │ -│ │ │ │ │ -│ └────────────────┼────────────────┘ │ -│ ▼ │ -│ ┌──────────────────────────────┐ │ -│ │ L2: Distributed Cache │ ← Shared within Org A only │ -│ │ (Raft or Redis Cluster) │ │ -│ │ <10ms access │ │ -│ │ org:A:* keys only │ │ -│ └──────────────┬───────────────┘ │ -│ │ │ -└──────────────────────────┼───────────────────────────────────────────────────┘ - │ - ▼ - ┌─────────────────────────────────────────┐ - │ L3: verisim-temporal │ - │ (Selective Federation) │ - │ │ - │ ┌────────────┐ ┌──────────────┐ │ - │ │ Private │ │ Public │ │ - │ │ (Org A │ │ (Federated │ │ - │ │ only) │ │ shared) │ │ - │ └────────────┘ └──────────────┘ │ - │ <100ms access │ - └─────────────────────────────────────────┘ - ▲ - │ -┌──────────────────────────┼───────────────────────────────────────────────────┐ -│ │ │ -│ ┌──────────────┴───────────────┐ │ -│ │ L2: Distributed Cache │ ← Shared within Org B only │ -│ │ (Raft or Redis Cluster) │ │ -│ │ <10ms access │ │ -│ │ org:B:* keys only │ │ -│ └──────────────┬───────────────┘ │ -│ ▲ │ -│ ┌────────────────┼────────────────┐ │ -│ │ │ │ │ -│ ┌─────┴──────┐ ┌─────┴──────┐ ┌─────┴──────┐ │ -│ │ Node B1 │ │ Node B2 │ │ Node B3 │ │ -│ ├────────────┤ ├────────────┤ ├────────────┤ │ -│ │ L1: ETS │ │ L1: ETS │ │ L1: ETS │ ← Per-node, NOT shared │ -│ │ 1GB cache │ │ 1GB cache │ │ 1GB cache │ │ -│ │ <1ms │ │ <1ms │ │ <1ms │ │ -│ └────────────┘ └────────────┘ └────────────┘ │ -│ ORGANIZATION B (Research Lab) │ -└──────────────────────────────────────────────────────────────────────────────┘ - -KEY POINTS: -• L1 caches are isolated per-node (each node has its own ETS table) -• L2 caches are isolated per-organization (org:A:* vs org:B:* key prefixes) -• L3 is split: private data stays local, public data is federated -• Organizations CANNOT access each other's L2 caches -• Public data in L3 is shared across federation to avoid duplication ----- - -== Cache Lookup Sequence - -This diagram shows the sequence when a query is executed: - -[source,text] ----- -Client Node A L1 Cache L2 Cache L3 Cache Store - │ │ │ │ │ │ - │ Query │ │ │ │ │ - ├────────────────────►│ │ │ │ │ - │ │ │ │ │ │ - │ │ 1. Check L1 │ │ │ │ - │ ├──────────────────►│ │ │ │ - │ │◄──────────────────┤ │ │ │ - │ │ MISS (<1ms) │ │ │ │ - │ │ │ │ │ │ - │ │ 2. Check L2 │ │ │ - │ ├────────────────────────────────►│ │ │ - │ │◄────────────────────────────────┤ │ │ - │ │ MISS (<10ms) │ │ │ - │ │ │ │ │ │ - │ │ 3. Check L3 │ │ - │ ├──────────────────────────────────────────────►│ │ - │ │◄──────────────────────────────────────────────┤ │ - │ │ MISS (<100ms) │ │ - │ │ │ │ │ │ - │ │ 4. Fetch from store │ - │ ├────────────────────────────────────────────────────────────►│ - │ │◄────────────────────────────────────────────────────────────┤ - │ │ Result (200-800ms) │ - │ │ │ │ │ │ - │ │ 5. Cache in L3 │ │ - │ ├──────────────────────────────────────────────►│ │ - │ │ │ │ │ │ - │ │ 6. Cache in L2 │ │ │ - │ ├────────────────────────────────►│ │ │ - │ │ │ │ │ │ - │ │ 7. Cache in L1 │ │ │ │ - │ ├──────────────────►│ │ │ │ - │ │ │ │ │ │ - │ Result │ │ │ │ │ - │◄────────────────────┤ │ │ │ │ - │ (Total: ~1000ms) │ │ │ │ │ - │ │ │ │ │ │ - │ │ │ │ │ │ - │ SECOND QUERY │ │ │ │ │ - ├────────────────────►│ │ │ │ │ - │ │ 1. Check L1 │ │ │ │ - │ ├──────────────────►│ │ │ │ - │ │◄──────────────────┤ │ │ │ - │ │ HIT! (<1ms) │ │ │ │ - │ Result │ │ │ │ │ - │◄────────────────────┤ │ │ │ │ - │ (Total: ~2ms) │ │ │ │ │ - │ 250x faster! ⚡ │ │ │ │ │ - -NOTES: -• First query: Full miss cascade (L1→L2→L3→Store), total ~1000ms -• Second query: L1 hit, total ~2ms (250x speedup) -• Cache promotion: Data fetched from L3 is promoted to L2 and L1 -• TTL-based expiration: Entries expire based on modality policy ----- - -== Invalidation Cascade Sequence - -This diagram shows how invalidation propagates when drift is detected: - -[source,text] ----- -DriftMonitor Node A Node A-L1 L2 Cluster Node B-L1 Node C-L1 - │ │ │ │ │ │ - │ Drift detected │ │ │ │ │ - │ (octad:abc-123)│ │ │ │ │ - ├────────────────►│ │ │ │ │ - │ │ │ │ │ │ - │ │ 1. Invalidate │ │ │ │ - │ │ local L1 │ │ │ │ - │ ├─────────────────►│ │ │ │ - │ │ │ Delete │ │ │ - │ │ │ octad:abc* │ │ │ - │ │◄─────────────────┤ │ │ │ - │ │ Invalidated │ │ │ │ - │ │ │ │ │ │ - │ │ 2. Broadcast to L2 │ │ │ - │ ├────────────────────────────────►│ │ │ - │ │ │ │ Raft/Redis │ │ - │ │ │ │ invalidate │ │ - │ │ │ │ org:A:*abc*│ │ - │ │◄────────────────────────────────┤ │ │ - │ │ Broadcast sent │ │ │ │ - │ │ │ │ │ │ - │ │ │ │ 3. Notify all nodes │ - │ │ │ ├───────────►│ │ - │ │ │ │ │ Invalidate │ - │ │ │ │ │ local L1 │ - │ │ │ │ │ │ - │ │ │ ├────────────────────────►│ - │ │ │ │ │ │ - │ │ │ │ │ Invalidate - │ │ │ │ │ local L1 - │ │ │ │ │ │ - │ │ 4. Log invalidation │ │ │ - │ │ (temporal) │ │ │ │ - │ ├─────────────────►│ │ │ │ - │ │ │ │ │ │ - │ Confirmation │ │ │ │ │ - │◄────────────────┤ │ │ │ │ - │ │ │ │ │ │ - -PROPAGATION TIMES: -• Step 1 (local L1): <1ms -• Step 2 (L2 broadcast): <10ms -• Step 3 (cluster-wide): <50ms total -• Step 4 (audit log): <100ms - -GUARANTEES: -✓ All nodes in organization receive invalidation within 50ms -✓ Audit trail in verisim-temporal for compliance -✓ Tag-based invalidation (can invalidate by octad, modality, federation) -✓ No stale data served after drift detection ----- - -== Tag-Based Invalidation Patterns - -[source,text] ----- -SCENARIO 1: Octad-Specific Drift - invalidate_by_tag("octad:abc-123") - ↓ - Invalidates: - • All query results for octad abc-123 - • All modalities (GRAPH, VECTOR, DOCUMENT, etc.) - • Across all cache layers (L1, L2, L3) - -SCENARIO 2: Modality Store Update - invalidate_by_tag("modality:GRAPH") - ↓ - Invalidates: - • All GRAPH query results - • All octads that have GRAPH data - • Statistics and execution plans for GRAPH - -SCENARIO 3: Federation-Wide Invalidation - invalidate_by_tag("federation:/universities/*") - ↓ - Invalidates: - • All queries targeting /universities/* federation pattern - • Across all organizations in the federation - • Both L2 (per-org) and L3 (federated) layers - -SCENARIO 4: Organization-Scoped Invalidation - invalidate_by_tag("org:university-a") - ↓ - Invalidates: - • All cache entries for organization "university-a" - • L1 on all nodes in that org - • L2 shared cache for that org - • L3 org-scoped data (NOT public federated data) - -TAG COMPOSITION: - Cache entries can have multiple tags: - tags: [ - "octad:abc-123", - "modality:GRAPH", - "federation:/universities/*", - "org:university-a" - ] - - Invalidation of ANY tag clears the entry. ----- - -== Sharing Matrix - -[cols="1,2,2,2"] -|=== -|Cache Layer |Standalone Mode |Hybrid Mode |Federated Mode - -|**L1 (ETS)** -|Per-process (not shared) -|Per-node (not shared across nodes) -|Per-node (not shared across orgs) - -|**L2 (Distributed)** -|N/A (single node) -|Shared across local nodes -|Shared within org, NOT across orgs - -|**L3 (Persistent)** -|Local disk -|Local disk + selective federation -|Selective: public data shared, private local -|=== - -== L1: Always Local, Never Shared - -**Why NOT shared:** -1. **Performance:** ETS is in-process memory, sharing would require serialization/network -2. **Isolation:** Each node has different query patterns -3. **Security:** Prevents cross-node cache poisoning - -**What's cached in L1:** -- Query results for THIS node's queries -- Parsed ASTs for THIS node -- Execution plans for THIS node -- Hot data recently accessed by THIS node - -**Cache invalidation:** -- Local only (doesn't propagate to other nodes) -- Must be explicitly invalidated on each node - -[source,elixir] ----- -# L1 cache is node-local -Node A: Caches query result for "SELECT GRAPH FROM octad abc-123" -Node B: Doesn't have this cached, must fetch separately - -# Invalidation is per-node -Node A: invalidate("query:result:abc123") ← Only affects Node A -Node B: Still has stale cache ← Must invalidate separately ----- - -== L2: Shared Within Trust Boundary - -**Should L2 be shared?** **YES, within same organization/cluster.** - -**Why shared:** -1. **Efficiency:** Multiple nodes benefit from same cache entry -2. **Consistency:** All nodes in cluster see same data -3. **Cost savings:** Avoid duplicate work across nodes - -**Why NOT shared across organizations:** -1. **Privacy:** University A shouldn't see University B's cached queries -2. **Security:** Cross-org cache poisoning risk -3. **Access control:** Different orgs have different permissions - -=== Implementation Options - -==== Option A: Raft-Based Shared Cache (Recommended for <20 nodes) - -[source,text] ----- -┌──────────────────────────────────────────────────┐ -│ Raft Cache Cluster (Same Org) │ -│ │ -│ Node 1 ◄─────┐ │ -│ │ │ -│ Node 2 ◄─────┼─── Raft Consensus ───► Cache │ -│ │ │ -│ Node 3 ◄─────┘ │ -└──────────────────────────────────────────────────┘ - -# Benefits: -- Strongly consistent -- No external dependencies -- Integrates with existing Raft (KRaft metadata log) - -# Drawbacks: -- More complex than Redis -- Limited to <20 nodes (Raft quorum overhead) ----- - -[source,elixir] ----- -defmodule VeriSim.QueryCache.L2Raft do - @moduledoc """ - L2 cache using Raft consensus for distributed caching. - Only shares within same organization cluster. - """ - - def put(key, value, opts) do - # Write to Raft log (replicated across cluster) - RaftCache.append_entry(%{ - command: :put, - key: key, - value: value, - org_id: get_org_id(), # Scoped to organization - ttl: opts[:ttl] - }) - end - - def get(key) do - # Read from local Raft replica (no consensus needed) - org_id = get_org_id() - - case RaftCache.read_local(key) do - {:ok, entry} when entry.org_id == org_id -> {:ok, entry.value} - _ -> {:error, :not_found} - end - end -end ----- - -==== Option B: Redis Cluster (Recommended for >20 nodes) - -[source,text] ----- -┌──────────────────────────────────────────────────┐ -│ Redis Cluster (Same Org) │ -│ │ -│ Node 1 ◄────┐ │ -│ │ │ -│ Node 2 ◄────┼─── Redis Pub/Sub ───► Cache │ -│ │ │ -│ Node 3 ◄────┘ │ -└──────────────────────────────────────────────────┘ - -# Benefits: -- Mature, battle-tested -- Scales to 100+ nodes -- Built-in pub/sub for invalidation - -# Drawbacks: -- External dependency (Redis server) -- Eventually consistent (not strongly consistent) ----- - -[source,elixir] ----- -defmodule VeriSim.QueryCache.L2Redis do - @moduledoc """ - L2 cache using Redis for distributed caching. - Scoped to organization via key prefixes. - """ - - def put(key, value, opts) do - org_id = get_org_id() - redis_key = "org:#{org_id}:#{key}" - - # Store in Redis with TTL - Redix.command(:redix, ["SETEX", redis_key, opts[:ttl], :erlang.term_to_binary(value)]) - - # Publish invalidation event (other nodes update their L1) - Redix.command(:redix, ["PUBLISH", "cache:invalidate:#{org_id}", key]) - end - - def get(key) do - org_id = get_org_id() - redis_key = "org:#{org_id}:#{key}" - - case Redix.command(:redix, ["GET", redis_key]) do - {:ok, nil} -> {:error, :not_found} - {:ok, binary} -> {:ok, :erlang.binary_to_term(binary)} - error -> error - end - end -end ----- - -=== Invalidation Propagation - -When L2 is shared, invalidations must propagate: - -[source,elixir] ----- -# Node A invalidates cache -Node A: QueryCache.invalidate("query:result:abc123") - ↓ -L2 (Raft/Redis): Broadcasts invalidation - ↓ -Node B, C, D: Receive invalidation → Clear L1 - -# Result: All nodes in cluster stay consistent ----- - -== L3: Selectively Shared Based on Data Sensitivity - -**Should L3 be shared?** **It depends on the data.** - -=== Public Data: Share Across Federation - -**What to share:** -- Published research papers (Document modality) -- Public embeddings (Vector modality) -- Open citation graphs (Graph modality) -- Historical public records (Temporal modality) - -**Benefits:** -1. **Efficiency:** Avoid duplicate storage across federation -2. **Consistency:** Everyone sees same historical data -3. **Cost savings:** Reduce federation-wide storage - -[source,elixir] ----- -# Public data cached in federated L3 -University A: Caches public paper "Climate Change 2025" - ↓ -Stored in federated verisim-temporal with tag: "visibility:public" - ↓ -University B: Queries same paper → L3 hit (no need to re-fetch) - -# Result: 200x faster for University B (no network fetch) ----- - -=== Private Data: Keep Local Only - -**What NOT to share:** -- Unpublished drafts -- Internal datasets -- Personally identifiable information (PII) -- Data under NDA/embargo - -**Implementation:** - -[source,elixir] ----- -defmodule VeriSim.QueryCache.L3 do - def put(key, value, opts) do - visibility = Keyword.get(opts, :visibility, :private) - - case visibility do - :public -> - # Store in federated verisim-temporal (shared) - VeriSimTemporal.store_federated(key, value, opts) - - :private -> - # Store in local verisim-temporal only (NOT shared) - VeriSimTemporal.store_local(key, value, opts) - - :org -> - # Store in organization-scoped temporal (shared within org) - org_id = get_org_id() - VeriSimTemporal.store_scoped(org_id, key, value, opts) - end - end - - def get(key) do - # Try local first (private data) - case VeriSimTemporal.get_local(key) do - {:ok, value} -> {:ok, value} - {:error, :not_found} -> - # Try org-scoped - case VeriSimTemporal.get_org_scoped(get_org_id(), key) do - {:ok, value} -> {:ok, value} - {:error, :not_found} -> - # Try federated (public data) - VeriSimTemporal.get_federated(key) - end - end - end -end ----- - -== Cache Consistency Across Layers - -=== Write-Through Strategy - -**When data is cached, write to all appropriate layers:** - -[source,elixir] ----- -def cache_query_result(key, result, opts) do - visibility = get_visibility(result) - - # Always write to L1 (local) - put_l1(key, result, opts) - - # Write to L2 if in cluster mode - if in_cluster_mode?() do - put_l2(key, result, opts) - end - - # Write to L3 based on visibility - case visibility do - :public -> put_l3_federated(key, result, opts) - :org -> put_l3_org_scoped(key, result, opts) - :private -> put_l3_local(key, result, opts) - end -end ----- - -=== Invalidation Cascade - -**When data is invalidated, cascade through all layers:** - -[source,elixir] ----- -def invalidate_cascade(key) do - # 1. Invalidate L1 (local) - invalidate_l1(key) - - # 2. Invalidate L2 (broadcast to cluster) - if in_cluster_mode?() do - broadcast_invalidation_to_cluster(key) - end - - # 3. Invalidate L3 (based on scope) - invalidate_l3(key) -end - -def broadcast_invalidation_to_cluster(key) do - # Raft-based - RaftCache.append_entry(%{command: :invalidate, key: key}) - - # Or Redis-based - Redix.command(:redix, ["PUBLISH", "cache:invalidate", key]) -end ----- - -== Security Considerations - -=== Cache Isolation by Organization - -[source,elixir] ----- -# Every cache key is prefixed by org_id -def cache_key(query, org_id) do - "org:#{org_id}:query:#{hash(query)}" -end - -# Organization A cannot access Organization B's cache -Org A: get("org:A:query:abc123") → ✓ Allowed -Org A: get("org:B:query:abc123") → ✗ Denied (different org) ----- - -=== Cache Poisoning Prevention - -[source,elixir] ----- -# Validate cache entries before serving -def get_from_cache(key) do - case fetch_from_l2(key) do - {:ok, entry} -> - # Verify integrity (ZKP signature) - if verify_signature(entry) do - {:ok, entry.value} - else - Logger.error("Cache poisoning detected: #{key}") - invalidate(key) - {:error, :invalid_signature} - end - - error -> error - end -end ----- - -=== Access Control - -[source,elixir] ----- -# Check permissions before serving cached data -def get_with_access_check(key, user_id) do - case get_from_cache(key) do - {:ok, cached_result} -> - # Verify user still has access - if AccessControl.can_read?(user_id, cached_result.octad_id) do - {:ok, cached_result} - else - Logger.warn("Access denied to cached data: #{key}") - {:error, :access_denied} - end - - error -> error - end -end ----- - -== Deployment Recommendations - -[cols="1,2,2,2"] -|=== -|Deployment Mode |L1 (Local) |L2 (Distributed) |L3 (Persistent) - -|**Standalone** -|✓ ETS (1GB) -|✗ Not needed -|✓ Local disk - -|**Hybrid (Single Org)** -|✓ ETS per node -|✓ Raft/Redis shared -|✓ Local + selective federation - -|**Federated (Multi-Org)** -|✓ ETS per node -|✓ Raft/Redis (per org) -|✓ Public: federated, Private: local -|=== - -=== Cost-Benefit Analysis - -[cols="1,1,1,2"] -|=== -|Sharing Level |Storage Cost |Network Cost |Benefit - -|**L1: Not shared** -|High (duplicate on each node) -|None -|Fastest access (<1ms) - -|**L2: Shared within org** -|Medium (one copy per org) -|Low (<10ms intra-cluster) -|Avoid duplicate work - -|**L3: Selectively shared** -|Low (one copy federation-wide) -|Medium (<100ms cross-org) -|Avoid duplicate storage -|=== - -== Summary - -**Cache Tiering:** -- ✅ **Yes**, VeriSimDB uses 3-tier caching (L1, L2, L3) - -**Cache Sharing:** -- **L1:** ❌ Never shared (per-node, in-memory) -- **L2:** ✅ Shared within organization/cluster, ❌ NOT across organizations -- **L3:** ✅ Public data shared across federation, ❌ Private data kept local - -**Why this design:** -1. **Performance:** L1 local = fastest -2. **Efficiency:** L2 shared within org = avoid duplicate work -3. **Privacy:** L3 selective = respect data sensitivity -4. **Security:** Org isolation = prevent cross-org cache poisoning - -**Recommendation:** -- Start with **L1 only** (simplest) -- Add **L2 (Raft)** when you have 3-5 nodes in same org -- Add **L2 (Redis)** when you exceed 20 nodes -- Use **L3 federation** only for public data - -== References - -- link:caching-strategy.adoc[Caching Strategy Overview] -- link:vcl-architecture.adoc[VCL Architecture] -- link:challenges-federated.adoc[Federated Deployment Challenges] -- link:../lib/verisim/query_cache.ex[Query Cache Implementation] diff --git a/verisimdb/docs/caching-strategy.adoc b/verisimdb/docs/caching-strategy.adoc deleted file mode 100644 index ec33f9f9..00000000 --- a/verisimdb/docs/caching-strategy.adoc +++ /dev/null @@ -1,594 +0,0 @@ -// SPDX-License-Identifier: CC-BY-SA-4.0 - -= VeriSimDB Caching Strategy -:toc: left -:toclevels: 3 -:sectnums: - -== Overview - -VeriSimDB implements a **multi-layer, drift-aware caching system** that dramatically improves query performance while maintaining data consistency. - -== What Gets Cached? - -[cols="1,2,1,2"] -|=== -|Cache Type |What's Stored |TTL |Invalidation - -|**Query Results** -|Complete query output with metadata -|5-60 min (per modality) -|On drift, on write - -|**Parsed ASTs** -|Parsed VCL abstract syntax trees -|1 hour -|Never (queries don't change) - -|**Execution Plans** -|Optimized query plans with cost estimates -|10 min -|On statistics update - -|**ZKP Proofs** -|Generated zero-knowledge proofs -|30 min -|On contract change - -|**Registry Lookups** -|UUID → store mappings -|10 min -|On registry update - -|**Temporal Versions** -|Historical octad states -|Infinite -|Never (immutable) - -|**Store Metadata** -|Index info, capabilities -|5 min -|On store update -|=== - -== Cache Architecture - -=== Three-Layer Design - -[source,text] ----- -┌─────────────────────────────────────────────────────────┐ -│ L1: In-Memory (ETS) │ -│ - Hot data (<1ms access) │ -│ - 1GB limit (LRU eviction) │ -│ - Query results, ASTs, plans │ -└─────────────────────────────────────────────────────────┘ - │ miss - ▼ -┌─────────────────────────────────────────────────────────┐ -│ L2: Distributed Cache (Across Nodes) │ -│ - Warm data (<10ms access) │ -│ - Shared across federation │ -│ - Raft-based consistency │ -└─────────────────────────────────────────────────────────┘ - │ miss - ▼ -┌─────────────────────────────────────────────────────────┐ -│ L3: Persistent Cache (verisim-temporal) │ -│ - Cold data (<100ms access) │ -│ - Survives restarts │ -│ - Historical queries │ -└─────────────────────────────────────────────────────────┘ ----- - -=== Cache Promotion - -When data is accessed: -1. **L3 hit** → Promote to L2 and L1 -2. **L2 hit** → Promote to L1 -3. **L1 hit** → Update access count (LRU) - -== Cache Policies: Per-Modality - -Different modalities have different caching needs: - -[cols="1,1,1,2"] -|=== -|Modality |Policy |TTL |Rationale - -|**VECTOR** -|Aggressive -|1 hour -|Embeddings rarely change, ANN search expensive - -|**GRAPH** -|Relaxed -|5 min -|Edges can change, but queries are expensive - -|**DOCUMENT** -|Aggressive -|1 hour -|Documents rarely change, full-text search expensive - -|**SEMANTIC** -|Strict -|1 min -|ZKP proofs must be fresh, contracts can change - -|**TENSOR** -|Relaxed -|5 min -|Tensors moderately stable, operations expensive - -|**TEMPORAL** -|Aggressive -|Infinite -|Historical data is immutable -|=== - -=== Policy Definitions - -**Strict:** -- Only cache dependent-type (verified) queries -- Short TTL (1 minute) -- Invalidate immediately on any change - -**Relaxed:** -- Cache both query types -- Medium TTL (5 minutes) -- Invalidate on drift detection - -**Aggressive:** -- Cache everything -- Long TTL (1 hour) -- Only invalidate on explicit writes - -== Drift-Aware Invalidation - -**Key Innovation:** Cache is automatically invalidated when drift is detected. - -=== Invalidation Triggers - -[source,elixir] ----- -# 1. Drift detected for specific octad -defmodule VeriSim.DriftMonitor do - def handle_drift_detected(octad_id, _details) do - # Invalidate all cached queries for this octad - VeriSim.QueryRouter.Cached.invalidate_on_drift(octad_id) - end -end - -# 2. Federation-wide drift -def handle_federation_drift(federation_pattern) do - VeriSim.QueryRouter.Cached.invalidate_federation_cache(federation_pattern) -end - -# 3. Modality store updated -def handle_store_update(modality) do - VeriSim.QueryRouter.Cached.invalidate_modality_cache(modality) -end ----- - -=== Tag-Based Invalidation - -Every cache entry has tags for fine-grained invalidation: - -[source,elixir] ----- -# Cache entry with tags -QueryCache.put(key, result, - tags: [ - "octad:abc-123", - "modality:GRAPH", - "federation:/universities/*" - ] -) - -# Invalidate all queries for octad -QueryCache.invalidate_by_tag("octad:abc-123") - -# Invalidate all graph queries -QueryCache.invalidate_by_tag("modality:GRAPH") - -# Invalidate entire federation -QueryCache.invalidate_by_tag("federation:/universities/*") ----- - -== Usage Examples - -=== Basic Query with Cache - -[source,elixir] ----- -# First execution: Cache MISS -{:ok, result} = VeriSim.QueryRouter.Cached.execute_with_cache( - "SELECT GRAPH FROM HEXAD abc-123", - use_dependent_types: true -) -# Execution time: 500ms - -# Second execution: Cache HIT -{:ok, result} = VeriSim.QueryRouter.Cached.execute_with_cache( - "SELECT GRAPH FROM HEXAD abc-123", - use_dependent_types: true -) -# Execution time: 2ms (250x faster!) ----- - -=== Force Fresh Results - -[source,elixir] ----- -# Bypass cache, force fresh query -{:ok, result} = VeriSim.QueryRouter.Cached.execute_with_cache( - query, - use_dependent_types: true, - force_fresh: true -) ----- - -=== Cache Warming at Startup - -[source,elixir] ----- -# Pre-populate cache with common queries -common_queries = [ - "SELECT GRAPH FROM FEDERATION /universities/* LIMIT 100", - "SELECT VECTOR FROM STORE milvus-1 WHERE ... LIMIT 50", - "SELECT TEMPORAL FROM HEXAD abc-123 AS OF 'yesterday'" -] - -VeriSim.QueryRouter.Cached.warm_cache(common_queries) ----- - -=== Manual Cache Control - -[source,elixir] ----- -# Invalidate specific query -QueryCache.invalidate(query_key) - -# Clear entire cache -QueryCache.clear_all() - -# Get cache statistics -QueryCache.stats() -# => %{ -# hit_rate: 87.5, -# hits: %{l1: 1000, l2: 200, l3: 50}, -# misses: 150, -# l1_entries: 5000, -# memory_mb: 234.5 -# } ----- - -== Performance Impact - -=== Benchmark: Cached vs Uncached - -[cols="1,1,1,1"] -|=== -|Query Type |Uncached |Cached (L1) |Speedup - -|**Simple Graph Query** -|150ms -|2ms -|**75x** - -|**Vector Similarity (HNSW)** -|300ms -|1ms -|**300x** - -|**Multi-Modal (Graph+Vector+Doc)** -|800ms -|5ms -|**160x** - -|**Dependent-Type with ZKP** -|1200ms -|3ms -|**400x** - -|**Temporal Historical Query** -|250ms -|1ms -|**250x** -|=== - -=== Cache Hit Rates (Production) - -Based on typical workloads: - -- **Read-heavy workload:** 90-95% hit rate -- **Mixed workload:** 70-80% hit rate -- **Write-heavy workload:** 40-50% hit rate - -=== Memory Usage - -**L1 Cache (1GB limit):** -- Typical entry: 10-50KB -- Capacity: ~20,000-100,000 entries -- LRU eviction when full - -**Storage Overhead:** -- L1: 0% (in-memory ETS) -- L2: ~5% (distributed across nodes) -- L3: ~10% (persistent in verisim-temporal) - -== Cache Configuration - -=== Default Configuration - -[source,elixir] ----- -%{ - max_memory_mb: 1024, # 1GB L1 cache - default_ttl_seconds: 300, # 5 minutes - enable_l2: true, # Distributed cache - enable_l3: true, # Persistent cache - policy: :relaxed, # Global policy - - modality_policies: %{ - "VECTOR" => :aggressive, - "GRAPH" => :relaxed, - "DOCUMENT" => :aggressive, - "SEMANTIC" => :strict, - "TENSOR" => :relaxed, - "TEMPORAL" => :aggressive - } -} ----- - -=== Runtime Configuration - -[source,elixir] ----- -# Update cache size -VeriSim.QueryCache.set_max_memory(2048) # 2GB - -# Change global policy -VeriSim.QueryCache.set_policy(:aggressive) - -# Override per modality -VeriSim.QueryCache.set_modality_policy("GRAPH", :aggressive) ----- - -== Integration with Query Optimization - -Caching integrates seamlessly with query optimization: - -=== Cached Execution Plans - -[source,text] ----- -Query: SELECT GRAPH, VECTOR FROM ... - -First Execution: - 1. Parse query → AST (20ms) ← CACHED for next time - 2. Generate plan (50ms) ← CACHED for next time - 3. Execute query (500ms) ← CACHED for next time - Total: 570ms - -Second Execution: - 1. Get cached AST (1ms) - 2. Get cached plan (1ms) - 3. Get cached result (1ms) - Total: 3ms (190x faster!) ----- - -=== Cached ZKP Proofs - -[source,elixir] ----- -# First query with ZKP generation -{:ok, result} = execute_with_cache( - "SELECT SEMANTIC FROM HEXAD abc-123 PROOF ACCESS(Contract)", - use_dependent_types: true -) -# ZKP generation: 500ms -# Total: 800ms - -# Second query reuses cached proof -{:ok, result} = execute_with_cache( - "SELECT SEMANTIC FROM HEXAD abc-123 PROOF ACCESS(Contract)", - use_dependent_types: true -) -# ZKP retrieved from cache: 1ms -# Total: 50ms (16x faster) ----- - -== Cache Monitoring - -=== Real-Time Statistics - -[source,elixir] ----- -# Get current cache stats -iex> VeriSim.QueryCache.stats() -%{ - hit_rate: 87.5, # 87.5% of queries served from cache - hits: %{ - l1: 10000, # 10k hits from in-memory - l2: 1500, # 1.5k hits from distributed - l3: 500 # 500 hits from persistent - }, - misses: 1800, # 1.8k cache misses - evictions: 200, # 200 LRU evictions - l1_entries: 45000, # 45k entries in L1 - memory_mb: 876.5, # 876MB used (out of 1024MB) - memory_limit_mb: 1024 -} ----- - -=== Performance Metrics - -[source,elixir] ----- -# Cache access times -L1 access: 0.5-2ms (in-memory) -L2 access: 5-15ms (network + deserialization) -L3 access: 50-150ms (disk + Merkle chain) - -# Compared to uncached: -Graph query: 100-500ms -Vector search: 200-800ms -ZKP generation: 300-1000ms ----- - -=== Logging - -[source,text] ----- -[info] Cache hit (l1): 1.2ms for query:result:abc123def -[info] Cache miss: 0.8ms for query:plan:balanced:xyz789 -[info] Evicted 1000 LRU cache entries (memory limit reached) -[info] Invalidated 234 cache entries with tag: octad:abc-123 -[info] Cache warming complete: 50 queries pre-loaded ----- - -== Best Practices - -=== 1. Use Appropriate Policies - -[source,elixir] ----- -# Read-heavy workload → Aggressive caching -QueryCache.set_modality_policy("DOCUMENT", :aggressive) - -# Compliance/audit queries → Strict caching -QueryCache.set_modality_policy("SEMANTIC", :strict) - -# Development/testing → Disable caching -QueryCache.clear_all() # Clear before each test ----- - -=== 2. Warm Cache After Major Changes - -[source,elixir] ----- -# After data import -def after_bulk_import do - # Invalidate old cache - QueryCache.clear_all() - - # Warm with new common queries - QueryRouter.Cached.warm_cache(get_common_queries()) -end ----- - -=== 3. Monitor Hit Rates - -[source,elixir] ----- -# Alert if hit rate drops below 50% -defmodule CacheMonitor do - use GenServer - - def handle_info(:check_hit_rate, state) do - stats = QueryCache.stats() - - if stats.hit_rate < 50.0 do - Logger.warn("Cache hit rate low: #{stats.hit_rate}%") - # Consider: increasing cache size, adjusting TTLs - end - - schedule_check() - {:noreply, state} - end -end ----- - -=== 4. Use Tags for Organized Invalidation - -[source,elixir] ----- -# Tag by project -QueryCache.put(key, result, tags: ["project:thesis-2025"]) - -# Later: invalidate entire project -QueryCache.invalidate_by_tag("project:thesis-2025") ----- - -== Troubleshooting - -=== Problem: Low Hit Rate - -**Symptoms:** Hit rate < 50%, many misses - -**Causes:** -1. TTL too short (entries expire before reuse) -2. High write rate (frequent invalidations) -3. Query patterns too diverse (no repeated queries) -4. Cache too small (LRU evictions) - -**Solutions:** -1. Increase TTL for stable modalities -2. Use more aggressive policies -3. Increase cache size -4. Pre-warm cache with common queries - -=== Problem: Stale Data - -**Symptoms:** Queries return outdated results - -**Causes:** -1. Drift not detected (DriftMonitor not running) -2. Manual writes bypass invalidation -3. TTL too long - -**Solutions:** -1. Ensure DriftMonitor is active -2. Call `invalidate_on_drift()` after writes -3. Reduce TTL for affected modalities -4. Use `force_fresh: true` for critical queries - -=== Problem: High Memory Usage - -**Symptoms:** L1 cache using >1GB, frequent evictions - -**Causes:** -1. Large query results cached -2. Too many entries -3. Memory limit too low - -**Solutions:** -1. Increase max_memory_mb -2. Use more conservative policies (cache less) -3. Reduce TTLs (entries expire sooner) -4. Use L2/L3 for large results - -== Future Enhancements - -=== Planned Features - -1. **Smart Cache Warming** - - ML-based prediction of common queries - - Automatic pre-loading based on usage patterns - -2. **Cross-Federation Cache Sharing** - - Universities share cached results - - Reduces duplicate work across federation - -3. **Query Result Compression** - - Compress large results in L2/L3 - - Trade CPU for memory/network - -4. **Cache Analytics Dashboard** - - Real-time visualization of hit rates - - Query pattern analysis - - Cost savings estimation - -5. **Adaptive TTLs** - - Automatically adjust TTL based on update frequency - - Longer TTL for stable data, shorter for volatile - -== References - -- link:query-optimization-overview.adoc[Query Optimization Overview] -- link:reversibility-design.adoc[Reversibility Design] -- link:vcl-architecture.adoc[VCL Architecture] -- link:../lib/verisim/query_cache.ex[Query Cache Implementation] -- link:../lib/verisim/query_router_cached.ex[Cached Query Router] diff --git a/verisimdb/docs/challenges-federated.adoc b/verisimdb/docs/challenges-federated.adoc deleted file mode 100644 index 9ac84342..00000000 --- a/verisimdb/docs/challenges-federated.adoc +++ /dev/null @@ -1,926 +0,0 @@ -// SPDX-License-Identifier: CC-BY-SA-4.0 - -= Challenges: Federated Deployment Mode - -== Overview - -Federated deployment runs VeriSimDB as a **tiny coordinator** (<5k LOC ReScript + Elixir) while modality stores are **distributed across independent institutions**. This mode enables cross-institutional collaboration and data sovereignty but introduces complex challenges in consensus, trust, and distributed coordination. - -This document details the **technical challenges** unique to federated deployment, focusing on the coordination layer and cross-organizational boundaries. - -== 1. Consensus and Coordination Challenges - -=== 1.1. KRaft Metadata Consensus Overhead - -**Problem:** Every registry update (Octad registration, policy change) requires **Raft quorum consensus**. - -**Raft Requirements:** -- **Minimum Nodes:** 3 (tolerates 1 failure) -- **Recommended:** 5 (tolerates 2 failures) -- **Quorum:** ⌈N/2⌉ + 1 nodes must acknowledge - -**Latency Impact:** -[source,text] ----- -Single-Node Write (Standalone): 1-5ms (local disk) -Raft Consensus Write (Federated): 50-500ms depending on geography - -Breakdown: - - Leader receives write request: 0ms - - Leader → Follower 1 (US East → US West): 70ms - - Leader → Follower 2 (US East → EU): 100ms - - Followers write to log: 10ms each - - Followers → Leader ACK: 100ms (slowest path) - - Leader commits and responds: 180ms total ----- - -**Throughput Limit:** -- **Theoretical:** ~1000 writes/sec per Raft group -- **Practical:** ~100-500 writes/sec with cross-region latency - -**Consequences:** -- **High Latency:** Registry updates 50-100x slower than standalone -- **Write Bottleneck:** Cannot scale writes beyond single Raft group capacity -- **Complexity:** Raft log compaction, snapshot management required - -**Mitigation:** -- **Batching:** Batch multiple registry updates into single Raft entry -- **Regional Sharding:** Multiple Raft groups per region (but: adds complexity) -- **Read-Heavy Optimization:** Serve reads from any replica (eventual consistency acceptable) -- **Log Compaction:** Snapshot registry state every N entries, truncate old log - -=== 1.2. Split-Brain and Network Partitions - -**Problem:** Network partition can split Raft quorum, causing **availability vs consistency tradeoff**. - -**Partition Scenarios:** - -**Scenario 1: Majority Partition** -[source,text] ----- -Initial: 5 nodes (Leader + 4 Followers) -Partition: [Leader, F1, F2] | [F3, F4] - -Majority partition (3 nodes): ✓ Can commit writes -Minority partition (2 nodes): ✗ Cannot elect leader, read-only - -Result: Majority partition remains available, minority unavailable ----- - -**Scenario 2: No Majority Partition** -[source,text] ----- -Initial: 5 nodes -Partition: [L, F1] | [F2, F3] | [F4] - -No partition has 3 nodes → ✗ No writes possible -All partitions read-only until partition heals - -Result: Total write unavailability ----- - -**Real-World Example:** -- University A (US East) - Leader -- University B (US West) - Follower -- University C (EU) - Follower -- University D (Asia) - Follower -- University E (AU) - Follower - -**If EU ↔ US link fails:** US+Asia+AU (3 nodes) continue, EU isolated - -**Consequences:** -- **Prolonged Partitions:** Can last hours-to-days for undersea cable failures -- **Data Divergence:** Minority partition accepts reads of stale data -- **Manual Intervention:** May require admin intervention to resolve - -**Mitigation:** -- **Monitoring:** Detect partitions quickly via heartbeat monitoring -- **Graceful Degradation:** Serve stale data with **staleness indicator** (e.g., "data as of 2 hours ago") -- **Multi-Region Raft:** Deploy separate Raft groups per region, federate at higher level -- **Automated Failover:** Use **Raft pre-vote** to prevent split elections - -=== 1.3. Leader Election Storms - -**Problem:** Frequent leader failures cause **election cascades**, reducing availability. - -**Election Storm Trigger:** -1. Leader node becomes overloaded (CPU 100%) -2. Heartbeat messages delayed → Followers timeout -3. Followers start elections simultaneously -4. Multiple candidates split votes → No leader elected -5. Repeat until timeout increases or load decreases - -**Duration:** Each election round takes **election timeout** (typically 150-300ms) × **retry attempts** - -**Impact Example:** -[source,text] ----- -Scenario: Leader overload + 3 failed elections - -T=0: Leader misses heartbeat -T=200ms: Follower 1 starts election -T=210ms: Follower 2 starts election (vote split) -T=400ms: Timeout, retry election -T=600ms: Timeout, retry election -T=800ms: Finally elect leader - -Result: 800ms write unavailability ----- - -**Consequences:** -- **Write Downtime:** All writes blocked during election -- **Cascading Failures:** If new leader also overloaded, repeat -- **Client Confusion:** Clients retry writes, amplifying load - -**Mitigation:** -- **Pre-Vote Phase:** Raft extension where candidates check viability before election -- **Leader Stickiness:** Increase heartbeat frequency, decrease election timeout variance -- **Load Shedding:** Leader sheds non-critical load when CPU > 80% -- **Backup Leaders:** Designate "preferred leader" based on network centrality - -== 2. Trust and Security Challenges - -=== 2.1. Multi-Party Signature Verification - -**Problem:** Every federated store must **sign** its data; coordinator must **verify** all signatures. - -**Signature Overhead:** -[source,text] ----- -Operation: Query Octad across 3 federated stores - -1. Orchestrator sends query to Store A, B, C (parallel) -2. Each store: - - Fetches data (10ms) - - Computes signature with private key (5ms) - - Returns {data, signature} -3. Orchestrator verifies signatures: - - Fetch Store A public key from registry (cached: 1ms) - - Verify signature A (10ms) - - Verify signature B (10ms) - - Verify signature C (10ms) -4. Merge results - -Total overhead: 30ms signature verification (vs 0ms standalone) ----- - -**Cryptographic Load:** -- **Ed25519 Verification:** ~10ms per signature (single-core) -- **At Scale:** 100 federated stores × 1000 qps = **100k signature verifications/sec** -- **CPU Requirement:** ~10 cores dedicated to signature verification - -**Consequences:** -- **Latency Tax:** Every query pays 30-50ms signature overhead -- **CPU Intensive:** Orchestrator needs powerful CPUs (not I/O bound like standalone) -- **Key Rotation Complexity:** Must coordinate public key updates across all stores - -**Mitigation:** -- **Batch Verification:** Verify multiple signatures in parallel (GPU acceleration) -- **Caching:** Cache verification results for **immutable data** (e.g., historical versions) -- **Signature Aggregation:** Use BLS signatures to aggregate N signatures into 1 -- **Hardware Acceleration:** Use AES-NI/AVX2 instructions for crypto operations - -=== 2.2. Byzantine Fault Tolerance - -**Problem:** Federated stores may be **malicious** (send incorrect data, violate policies). - -**Threat Model:** -[cols="1,2,2"] -|=== -|Attack |Example |Impact - -|**Data Tampering** -|Store A returns modified Octad content -|Clients receive incorrect data - -|**Policy Violation** -|Store B grants access despite policy denying it -|Unauthorized data access - -|**Availability Attack** -|Store C always times out (DoS) -|Query failures, degraded UX - -|**Sybil Attack** -|Attacker registers 10 fake stores, controls majority -|Registry poisoning -|=== - -**Byzantine Generals Problem:** -- **Classic Raft:** Assumes crash failures only (nodes fail-stop) -- **Federated VeriSimDB:** Must tolerate **byzantine failures** (malicious nodes) -- **BFT Requirement:** Need **3f + 1** nodes to tolerate **f** byzantine nodes - -**Example:** -- To tolerate 1 byzantine store: Need 4 stores minimum -- To tolerate 2 byzantine stores: Need 7 stores minimum - -**Consequences:** -- **Higher Overhead:** More replicas needed than crash-tolerant Raft -- **Complex Verification:** Must cross-check responses from multiple stores -- **Trust Assumptions:** Cannot assume federated stores are honest - -**Mitigation:** -- **BFT Consensus:** Replace Raft with **Tendermint** or **HotStuff** for metadata log -- **Merkle Proofs:** Require stores to provide **Merkle proofs** for all data -- **Reputation System:** Track store reliability, downrank misbehaving stores -- **ZKP Verification:** Use `proven` ZKP to verify **logical correctness** without trusting store - -=== 2.3. Trust Boundary Explosion - -**Problem:** In federated mode, **every store is a trust boundary** → N stores = N boundaries. - -**Trust Boundaries:** -[source,text] ----- -Standalone: 1 boundary (orchestrator → local stores) -Hybrid: 3-5 boundaries (orchestrator → local + 2-3 remote) -Federated: 50-100 boundaries (orchestrator → N remote stores) ----- - -**Security Implications:** -1. **Compromise Surface:** Any one store compromise exposes its hosted Octads -2. **Policy Inconsistency:** 100 stores = 100 different security policies to audit -3. **Credential Management:** Must manage 100 pairs of TLS certificates -4. **Audit Complexity:** Must correlate logs across 100 independent systems - -**Example Breach Scenario:** -[source,text] ----- -Attacker compromises University B's federated store: - - Gains access to all Octads hosted by University B - - Can serve malicious data to orchestrator - - Can exfiltrate query patterns (who queries what) - -Impact: Partial breach (only University B's Octads), but: - - Orchestrator cannot detect tampering without cross-validation - - Clients unaware they received compromised data ----- - -**Mitigation:** -- **Zero-Trust Architecture:** Verify **every** response from **every** store (never trust by default) -- **Defense in Depth:** Multiple layers of verification (signatures + Merkle proofs + ZKPs) -- **Continuous Monitoring:** Real-time anomaly detection for store behavior -- **Incident Response Plan:** Pre-negotiated procedures for store compromise - -== 3. Operational Complexity - -=== 3.1. Heterogeneous Store Management - -**Problem:** Federated stores run **different software versions**, **different configurations**, and **different hardware**. - -**Heterogeneity Matrix:** -[cols="1,2,2,2"] -|=== -|Store |Software |Hardware |Network - -|**University A** -|VeriSim 0.9 (Rust 1.70) -|32-core, 128GB RAM -|10 Gbps - -|**Archive.org** -|Custom fork v0.8 -|16-core, 64GB RAM (shared) -|1 Gbps (rate-limited) - -|**Lab B** -|VeriSim 1.0-beta (Rust 1.75) -|8-core, 32GB RAM -|100 Mbps - -|**Community Node** -|VeriSim 0.7 (outdated) -|4-core, 16GB RAM -|50 Mbps DSL -|=== - -**Consequences:** -- **API Incompatibility:** Old stores may not support new API features -- **Performance Variance:** 200x difference between fastest and slowest store -- **Bug Divergence:** Custom forks may have unique bugs -- **Upgrade Coordination:** Cannot force-upgrade 100 independent stores - -**Mitigation:** -- **Versioned APIs:** Support multiple API versions in orchestrator (v1, v2, v3) -- **Capability Negotiation:** Stores advertise capabilities in registration manifest -- **Graceful Degradation:** If store doesn't support feature, skip that modality -- **Minimum Version Policy:** Require stores to be within N versions of current - -=== 3.2. Cross-Institutional SLAs - -**Problem:** Must coordinate **service level agreements** across independent organizations. - -**SLA Negotiation Challenges:** -[cols="1,2,2"] -|=== -|Aspect |Challenge |Example Conflict - -|**Uptime** -|Different institutions have different availability requirements -|University A: 99.9% SLA, Archive.org: 95% SLA - -|**Latency** -|Geographic distribution causes unavoidable latency variance -|US-EU queries: 100ms, US-Asia: 200ms - -|**Data Retention** -|Legal requirements differ by jurisdiction -|EU GDPR: delete after 30 days, US: retain indefinitely - -|**Maintenance Windows** -|Timezones make coordination difficult -|University A: 2am ET, Lab B: 2am PT (3-hour overlap) - -|**Cost Sharing** -|Who pays for bandwidth, storage, compute? -|Free tier vs paid tier stores -|=== - -**Real-World Scenario:** -[source,text] ----- -Query touches 5 stores: - - 3 stores respond in 50ms (99.9% uptime) - - 1 store responds in 500ms (99% uptime, overloaded) - - 1 store times out (maintenance window) - -Result: Query returns incomplete data with 500ms latency -User Experience: Slow and unreliable ----- - -**Mitigation:** -- **SLA Tiers:** Define Bronze/Silver/Gold tiers with clear expectations -- **Circuit Breakers:** Automatically exclude stores violating SLAs -- **Best-Effort Model:** Accept that federated queries are **inherently best-effort** -- **Contractual Agreements:** Formal MoUs between institutions (but: hard to enforce) - -=== 3.3. Drift Detection Across Organizational Boundaries - -**Problem:** Drift detection in federated mode requires **polling remote stores** → network overhead and coordination challenges. - -**Drift Detection Protocol:** -[source,text] ----- -1. Orchestrator queries all stores hosting Octad 550e8400-... - - Store A (Graph): GET /octad/550e8400/graph/metadata - - Store B (Vector): GET /octad/550e8400/vector/metadata - - Store C (Document): GET /octad/550e8400/document/metadata - -2. Compare metadata (timestamps, hashes, version numbers) - -3. If drift detected: - a. Determine authoritative source (policy-driven) - b. Trigger repair: - - Option 1: Pull from authoritative → Push to drifted stores - - Option 2: Flag for manual review (if conflict) - -4. Record drift event in Temporal log ----- - -**Challenges:** -- **Polling Overhead:** Must poll N stores every T minutes (N×T network requests) -- **Consistency Semantics:** Different stores may have different "freshness" guarantees -- **Repair Latency:** Cross-institutional repair can take hours (requires approval) -- **Policy Conflicts:** Authoritative source may disagree across institutions - -**Example Drift Scenario:** -[source,text] ----- -T=0: Research paper retracted at University A -T=60: University A updates Graph modality (removes citation edges) -T=120: Drift detection discovers Archive.org Document still references paper -T=180: Orchestrator sends repair request to Archive.org -T=240: Archive.org admin reviews request (manual process) -T=480: Archive.org applies update (8 hours later) - -Reality: Cross-institutional drift repair is **not real-time** ----- - -**Mitigation:** -- **Event-Driven Updates:** Stores **push** change notifications instead of polling -- **Asynchronous Repair:** Accept that drift is **eventually** resolved (not immediately) -- **Priority Tiers:** Critical drift (e.g., retractions) gets expedited review -- **Automated Repair Policies:** Pre-negotiated rules for common drift cases - -=== 3.4. Monitoring and Observability Gaps - -**Problem:** Cannot directly access **logs**, **metrics**, or **traces** from federated stores. - -**Observability Challenges:** -[cols="1,2,2"] -|=== -|Data Type |Standalone |Federated - -|**Logs** -|Direct access via journalctl/syslog -|Must request logs from store operator (may refuse) - -|**Metrics** -|Prometheus scrape from localhost -|Must expose /metrics endpoint (may not implement) - -|**Traces** -|OpenTelemetry spans in single system -|Distributed tracing requires cooperation (may not support) - -|**Debugging** -|Can attach debugger, inspect memory -|Cannot access remote store internals -|=== - -**Example Debug Scenario:** -[source,text] ----- -Problem: Queries to Store X are slow (500ms vs expected 50ms) - -Standalone: - - Check CPU, memory, disk I/O on local machine - - Profile code with perf/flamegraph - - Fix in minutes - -Federated: - - Email Store X operator: "Can you check if your system is slow?" - - Wait 24 hours for response - - Operator says: "Looks fine on our end" - - Add synthetic monitoring, discover issue is network routing - - Cannot fix (requires Store X ISP intervention) - -Resolution time: Days to weeks ----- - -**Mitigation:** -- **Federated Observability Contract:** Require stores to export metrics/logs to central collector -- **Synthetic Monitoring:** Run continuous health checks from orchestrator to all stores -- **Anomaly Detection:** Use ML to detect outliers (e.g., Store X consistently 10x slower) -- **SLA Enforcement:** Downrank or exclude stores that don't provide observability data - -== 4. Performance and Scaling Challenges - -=== 4.1. Network Latency Dominates - -**Problem:** In federated mode, **network latency** overwhelms compute time. - -**Latency Breakdown:** -[source,text] ----- -Standalone Query (Vector search): - - Orchestrator → Local Vector store: 0.1ms (localhost) - - HNSW search: 2ms - - Return: 0.1ms - Total: 2.2ms - -Federated Query (Vector search): - - Orchestrator → Remote Store X: 50ms (cross-country) - - HNSW search: 2ms - - Return: 50ms - Total: 102ms (46x slower) - -Cross-Region Query: - - US → EU: 100ms RTT - - Total: 202ms (92x slower) ----- - -**Consequences:** -- **Tail Latency:** p99 latency determined by **slowest store** in query path -- **Cannot Optimize Compute:** Optimizing HNSW from 2ms → 1ms irrelevant when network is 100ms -- **Geographic Penalty:** Queries spanning continents inherently slow - -**Mitigation:** -- **CDN-Style Caching:** Cache hot data near orchestrator (stale acceptable for some use cases) -- **Regional Affinity:** Route queries to geographically close stores when possible -- **Speculative Execution:** Query multiple stores in parallel, use fastest response -- **Accept Latency:** Federated mode trades latency for **data sovereignty** and **institutional independence** - -=== 4.2. Fan-Out Queries and Amplification - -**Problem:** Cross-modal queries **fan out** to multiple stores, amplifying latency and failures. - -**Fan-Out Example:** -[source,text] ----- -Query: "Find papers similar to embedding X that cite paper Y" - -Requires: - 1. Vector search (Store A) - 2. Graph traversal (Store B) - 3. Join results (orchestrator) - -Sequential: 50ms + 50ms = 100ms -Parallel: max(50ms, 50ms) = 50ms - -With N modalities: latency scales with max(Store1, Store2, ..., StoreN) ----- - -**Amplification Factors:** -- **2 modalities:** 2x network calls -- **6 modalities:** 6x network calls (if not optimized) -- **Failure Probability:** P(any store fails) = 1 - ∏(1 - P(store_i fails)) - -**Example:** -- Each store: 99% availability -- Query touches 6 stores -- Query success probability: 0.99^6 = **94%** (6% failure rate!) - -**Consequences:** -- **High Failure Rate:** Multi-store queries less reliable than single-store -- **Partial Results:** Must handle cases where some stores timeout -- **Retry Storms:** Clients retry failed queries → amplify load on stores - -**Mitigation:** -- **Query Planning:** Optimize to minimize store fan-out -- **Hedged Requests:** Send duplicate queries to backup stores if primary slow -- **Partial Result Tolerance:** Return best-effort results with **completeness indicator** -- **Circuit Breakers:** Fail fast if primary store down, avoid waiting for timeout - -=== 4.3. Write Propagation Latency - -**Problem:** Writes in federated mode must **propagate across stores**, introducing delays. - -**Write Propagation Flow:** -[source,text] ----- -T=0: Client writes Octad update to orchestrator -T=1: Orchestrator commits to KRaft metadata log (quorum: 100ms) -T=101: Orchestrator dispatches writes to federated stores: - - Store A (Graph): 50ms to reach, 10ms to persist - - Store B (Vector): 80ms to reach, 5ms to persist - - Store C (Document): 120ms to reach, 20ms to persist -T=241: All stores ACK write completion - -Total write latency: 241ms (vs 5ms standalone) ----- - -**Consistency Implications:** -- **Eventual Consistency:** Reads during T=101 to T=241 see **incomplete data** -- **Read-After-Write Anomaly:** Client writes, immediately reads, sees old data -- **Cross-Store Inconsistency:** Store A updated, Store B still processing - -**Mitigation:** -- **Write Acknowledgment Levels:** - - **Metadata Only:** Return after KRaft commit (fast, but data not yet at stores) - - **One Store:** Return after 1 store ACKs (balanced) - - **All Stores:** Return after all stores ACK (slow, but fully consistent) -- **Client Caching:** Client caches pending writes, serves from cache until propagated -- **Version Tokens:** Return version token with write, clients include token in subsequent reads - -=== 4.4. Bandwidth Costs at Scale - -**Problem:** Federated architecture incurs **significant bandwidth costs** for cross-store communication. - -**Bandwidth Analysis:** -[source,text] ----- -Assumptions: - - 100 federated stores - - 1000 qps average load - - Average query touches 3 stores - - Average response size: 100 KB - -Bandwidth required: - - Queries: 1000 qps × 3 stores × 100 KB = 300 MB/s - - Responses: 300 MB/s - - Total: 600 MB/s = 4.8 Gbps - -Monthly bandwidth: - - 4.8 Gbps × 86400 sec/day × 30 days = 15.5 PB/month - - At $0.10/GB egress: $1.5M/month bandwidth cost ----- - -**Cost Breakdown:** -- **Intra-Region:** Typically free or low-cost ($0.01/GB) -- **Cross-Region (Same Cloud):** $0.02-0.05/GB -- **Cross-Region (Different Clouds):** $0.08-0.15/GB -- **To Internet:** $0.10-0.20/GB - -**Consequences:** -- **Cost Explosion:** Bandwidth becomes **dominant cost** at scale -- **Regional Affinity Critical:** Must minimize cross-region traffic -- **Economic Barrier:** Small institutions cannot afford federated participation - -**Mitigation:** -- **Response Compression:** gzip/brotli compression (70% reduction) -- **Delta Encoding:** Send only changes, not full objects -- **Regional Hubs:** Deploy regional orchestrators to localize traffic -- **Peer-to-Peer:** Stores communicate directly (bypass orchestrator for large transfers) - -== 5. Governance and Policy Challenges - -=== 5.1. Multi-Stakeholder Governance - -**Problem:** Federated mode involves **multiple independent institutions** with conflicting interests. - -**Governance Conflicts:** -[cols="1,2,2"] -|=== -|Issue |Institution A |Institution B - -|**Data Retention** -|Retain indefinitely (research) -|Delete after 7 years (legal requirement) - -|**Access Policies** -|Open access (public good) -|Restricted to .edu domains (institutional policy) - -|**Monetization** -|Free tier only -|Charge $0.01/query (cost recovery) - -|**Modality Support** -|All 6 modalities -|Graph and Document only (limited resources) - -|**Schema Changes** -|Propose new metadata fields -|Resist changes (stability) -|=== - -**Coordination Overhead:** -- **Consensus Required:** Major changes need agreement from **majority of stores** -- **Veto Power:** Any store can refuse to adopt changes -- **Glacial Pace:** Months to approve simple changes - -**Example:** -[source,text] ----- -Proposal: Add "license" field to Octad metadata - -Timeline: - - T=0: Proposal submitted - - T+30d: 10 stores approve, 5 neutral, 3 reject - - T+60d: Negotiation with rejectors (concerns about schema bloat) - - T+90d: Compromise: "license" optional, not required - - T+120d: Implementation begins (phased rollout) - - T+180d: Fully adopted - -Result: 6 months for simple metadata addition ----- - -**Mitigation:** -- **Governance Framework:** Pre-defined voting procedures, quorum requirements -- **Lightweight Changes:** Allow non-breaking changes without full consensus -- **Opt-In Features:** New features optional, not mandatory -- **Advisory Board:** Elected representatives from major institutions - -=== 5.2. Policy Synchronization and Drift - -**Problem:** Access policies **diverge** across federated stores over time. - -**Policy Drift Example:** -[source,text] ----- -Initial: All stores enforce "Allow .edu domains only" - -T+6 months: - - Store A: Still enforcing original policy - - Store B: Updated to "Allow .edu and .gov" - - Store C: Reverted to "Public access" (new administration) - -Query from @gmail.com: - - Store A: ✗ Denied - - Store B: ✗ Denied - - Store C: ✓ Allowed - -Result: Inconsistent authorization (confusing UX) ----- - -**Causes of Drift:** -- **Manual Updates:** Stores manually apply policy changes → prone to lag -- **Versioning:** Stores run different software versions with different policy engines -- **Intentional Divergence:** Store operators override global policies - -**Consequences:** -- **Security Gaps:** Overly permissive stores bypass intended restrictions -- **Compliance Violations:** Some stores may violate institutional regulations -- **User Confusion:** "Works on one store, denied on another" - -**Mitigation:** -- **Policy Registry:** Store policies in **global registry**, orchestrator validates before querying -- **Automated Enforcement:** Orchestrator **enforces policies** before forwarding queries (don't trust stores) -- **Policy Audit:** Regular audits to detect drift, flag violating stores -- **Policy Versioning:** Policies have version numbers, require minimum version - -=== 5.3. Economic Incentives and Cost Sharing - -**Problem:** Federated stores incur costs (compute, storage, bandwidth) → need **economic model** for sustainability. - -**Cost Distribution:** -[cols="1,2,2,2"] -|=== -|Store Type |Costs |Revenue Model |Sustainability - -|**Academic** -|$10k/month (grant-funded) -|Free (public good) -|Depends on grant renewal - -|**Commercial** -|$50k/month (production) -|Charge $0.01/query -|Profitable if >5M queries/month - -|**Community** -|$500/month (volunteer) -|Donations -|Fragile (volunteers burn out) - -|**Archive** -|$5k/month (shared infra) -|Institutional funding -|Stable but slow upgrades -|=== - -**Tragedy of the Commons:** -- Popular stores get **overloaded** (high query volume) -- Unpopular stores **idle** (wasted capacity) -- No mechanism to balance load or compensate popular stores - -**Example:** -[source,text] ----- -Archive.org hosts 80% of Document modality data - - Receives 80% of federated queries - - Bandwidth costs skyrocket - - Archive.org considers de-federating to reduce costs - -Result: Federated ecosystem at risk ----- - -**Mitigation:** -- **Reciprocity Agreements:** Stores agree to host each other's overflows -- **Payment Rails:** Integrate micropayments (e.g., Lightning Network, stablecoins) -- **Query Credits:** Each store gets N free queries/month, pays beyond that -- **Funding Pool:** Central fund (grants, donations) subsidizes high-cost stores - -== 6. Data Sovereignty and Legal Challenges - -=== 6.1. Cross-Jurisdictional Data Flows - -**Problem:** Federated queries may **cross legal jurisdictions** with different data protection laws. - -**Legal Frameworks:** -[cols="1,2,2"] -|=== -|Jurisdiction |Law |Key Requirement - -|**EU** -|GDPR -|Data cannot leave EU without adequacy decision - -|**US** -|CLOUD Act -|US government can compel disclosure - -|**China** -|Data Security Law -|Data localization (critical data stays in China) - -|**Russia** -|Personal Data Law -|Russian citizens' data must be stored in Russia -|=== - -**Compliance Conflict:** -[source,text] ----- -Scenario: EU user queries Octad hosted across: - - Store A (Germany): Subject to GDPR - - Store B (US): Subject to CLOUD Act - - Store C (China): Subject to Data Security Law - -Problem: - - GDPR: Cannot send personal data to China (no adequacy) - - CLOUD Act: US government could compel Store B disclosure - - China DSL: If data critical, cannot leave China - -VeriSimDB must: - 1. Determine if query involves personal data - 2. Check user's jurisdiction - 3. Only query compliant stores - 4. Return partial results if necessary ----- - -**Consequences:** -- **Compliance Overhead:** Must track **data residency** for every Octad -- **Partial Results:** May need to exclude stores for legal reasons -- **Audit Complexity:** Must prove compliance to regulators in multiple jurisdictions - -**Mitigation:** -- **Data Tagging:** Tag Octads with **jurisdiction metadata** (e.g., "EU-only", "US-only") -- **Query Planning:** Route queries only to **legally compliant stores** -- **Encryption:** Use **homomorphic encryption** to process data without exposing plaintext (but: performance cost) -- **Legal Counsel:** Consult lawyers specializing in cross-border data flows - -=== 6.2. Right to be Forgotten (RTBF) - -**Problem:** GDPR **Right to be Forgotten** requires deleting data across **all federated stores**. - -**RTBF Challenge:** -[source,text] ----- -User requests deletion of Octad 550e8400-... - -VeriSimDB must: - 1. Identify all stores hosting any modality of Octad - - Registry lookup: Stores [A, B, C, D, E] - 2. Send deletion request to each store - 3. Wait for confirmation (stores may take days to comply) - 4. Verify deletion (how to prove data gone?) - 5. Update registry to mark Octad deleted - -Complications: - - Store B is offline (maintenance window) - - Store C refuses (claims US CLOUD Act overrides GDPR) - - Store D deleted data but has backup tapes (not yet purged) - - Store E is a mirror (didn't receive deletion request) - -Result: Incomplete deletion → GDPR violation ----- - -**Penalties:** GDPR fines up to **€20M or 4% of global revenue** (whichever higher) - -**Mitigation:** -- **Deletion SLA:** Pre-negotiate **deletion timelines** with all stores (e.g., 30 days) -- **Immutable Flag:** Mark Octad as "deleted" in registry immediately, lazy-delete from stores -- **Cryptographic Deletion:** Encrypt Octads with **per-Octad keys**, delete keys instead of data -- **Backup Coordination:** Require stores to certify backups also purged - -=== 6.3. Sovereignty vs Availability Tradeoff - -**Problem:** Strict data sovereignty requirements **limit query availability**. - -**Example:** -[source,text] ----- -Scenario: French research paper (Octad X) hosted only on French stores - -Global query from US user: - - Option 1: Query French stores (slow, 150ms latency due to geography) - - Option 2: Mirror to US stores (violates French data sovereignty laws) - - Option 3: Deny access (poor UX) - -Trade-off: Cannot optimize for both sovereignty AND performance ----- - -**Tension:** -- **Sovereignty:** Data must stay in jurisdiction → **limits geographic distribution** -- **Performance:** Low latency requires **data near users** → conflicts with sovereignty - -**Mitigation:** -- **Edge Caching:** Cache **anonymized** versions near users (metadata only, no personal data) -- **Federated Learning:** Compute aggregates locally, only share aggregates (not raw data) -- **Legal Agreements:** Bilateral agreements between jurisdictions (expensive, slow) - -== 7. Summary of Challenges - -[cols="1,2,1"] -|=== -|Challenge Category |Key Issues |Severity - -|**Consensus** -|Raft quorum overhead, split-brain, leader election storms -|HIGH - -|**Trust** -|Multi-party signatures, Byzantine faults, trust boundary explosion -|CRITICAL - -|**Operations** -|Heterogeneous stores, cross-institutional SLAs, drift detection, observability gaps -|HIGH - -|**Performance** -|Network latency dominates, fan-out amplification, write propagation delay, bandwidth costs -|HIGH - -|**Governance** -|Multi-stakeholder conflicts, policy drift, economic incentives -|MEDIUM - -|**Legal** -|Cross-jurisdictional flows, RTBF compliance, sovereignty vs availability -|CRITICAL -|=== - -== 8. When to Choose Federated Despite Challenges - -Federated deployment is appropriate when: - -1. **Data Sovereignty:** Institutional data MUST remain under institutional control (non-negotiable) -2. **Cross-Institutional Collaboration:** Multiple universities/labs need to share knowledge without centralization -3. **Compliance Requirements:** Legal/regulatory requirements prevent data centralization -4. **Open Science:** Commitment to decentralized, community-governed knowledge infrastructure -5. **Scale Beyond Single Institution:** Data volume exceeds what any single institution can host -6. **Resilience:** Need to survive single-institution failures (geographic redundancy) - -For simpler use cases, **standalone** or **hybrid** modes are recommended. - -== See Also - -- link:deployment-modes.adoc[Deployment Modes Overview] -- link:challenges-standalone.adoc[Challenges: Standalone Mode] -- link:challenges-hybrid.adoc[Challenges: Hybrid Mode] -- link:technical-specification-kraft-metadata-log.adoc[KRaft Metadata Log Specification] -- link:zkp-and-sanctify-integration.adoc[Zero-Trust Security with ZKP] -- link:../WHITEPAPER.md[VeriSimDB White Paper] diff --git a/verisimdb/docs/challenges-hybrid.adoc b/verisimdb/docs/challenges-hybrid.adoc deleted file mode 100644 index 3d3ea90f..00000000 --- a/verisimdb/docs/challenges-hybrid.adoc +++ /dev/null @@ -1,782 +0,0 @@ -// SPDX-License-Identifier: CC-BY-SA-4.0 - -= Challenges: Hybrid Deployment Mode - -== Overview - -Hybrid deployment runs **some modalities locally** (fast access, low latency) while **federating others** (cost savings, institutional collaboration). This mode offers the best balance of performance and federation but introduces unique challenges at the **boundary between local and remote stores**. - -This document details the **technical challenges** specific to hybrid deployment, focusing on consistency, configuration, and cross-boundary coordination. - -== 1. Architectural Challenges - -=== 1.1. Hot/Cold Data Boundary Design - -**Problem:** Deciding **which modalities** stay local vs federate requires predicting access patterns. - -**Design Questions:** -1. Which modalities are "hot" (high query rate) vs "cold" (archive)? -2. Do access patterns change over time (seasonality, research trends)? -3. What if a "cold" modality suddenly becomes hot (e.g., regulatory audit)? - -**Example Decision Matrix:** -[cols="1,2,2,1"] -|=== -|Modality |Access Pattern |Decision |Rationale - -|**Vector** -|10k qps (RAG pipeline) -|**Local** -|Sub-ms latency critical - -|**Graph** -|500 qps (citation lookup) -|**Local** -|Medium load, queries cheap - -|**Document** -|50 qps (full-text search) -|**Remote** -|Archive.org has corpus already - -|**Temporal** -|10 qps (version history) -|**Remote** -|Cold data, institutional archive - -|**Tensor** -|100 qps (model inference) -|**Local** -|GPU-bound, needs local acceleration - -|**Semantic** -|5 qps (proof validation) -|**Remote** -|Validation logic centralized -|=== - -**Risk:** Incorrect boundary placement leads to: -- **Over-provisioning local:** Wasting resources on cold data -- **Over-federating:** Network latency kills performance - -**Mitigation:** -- **Instrumentation:** Deploy with full telemetry, collect 30 days of query logs before deciding -- **Dynamic Reconfiguration:** Design orchestrator to support **live migration** of modalities (hot → local, cold → remote) -- **Hybrid-Hybrid:** Allow **same modality split** across local/remote (e.g., recent Temporal versions local, old versions remote) - -=== 1.2. Network Topology Complexity - -**Problem:** Hybrid mode requires orchestrator to route queries across **multiple network hops**. - -**Network Paths:** -[source,text] ----- -Client → Orchestrator → Local Modality (1 hop) ✓ Fast -Client → Orchestrator → Remote Store → Modality (2+ hops) ✗ Slow -Client → Orchestrator → Remote Store → Remote Store → ... (N hops) ✗✗ Cascade failure risk ----- - -**Latency Breakdown Example:** -|=== -|Query Type |Local Path |Remote Path |Slowdown Factor - -|Vector search -|2ms -|50ms (network) + 2ms = 52ms -|**26x** - -|Graph SPARQL -|10ms -|50ms + 10ms = 60ms -|**6x** - -|Cross-modal join (Vector + Graph) -|12ms (parallel) -|52ms + 60ms = 112ms (sequential) -|**9x** -|=== - -**Consequences:** -- **Unpredictable Latency:** Queries touching mixed local/remote modalities have variable latency -- **Timeout Tuning:** Must set different timeouts per modality (local: 100ms, remote: 5s) -- **Cascading Failures:** Remote store timeout delays local query completion - -**Mitigation:** -- **Parallel Dispatch:** Query local and remote modalities **concurrently**, combine results -- **Timeouts Per Modality:** Configure `VeriSim.QueryRouter` with modality-specific timeouts -- **Circuit Breakers:** Implement Hystrix-style circuit breakers to fail fast when remote stores degrade - -=== 1.3. Partial Consistency Models - -**Problem:** Local modalities offer **strong consistency**, remote modalities offer **eventual consistency** → hybrid has **mixed consistency**. - -**Consistency Matrix:** -[cols="1,2,2"] -|=== -|Modality Location |Consistency Model |Guarantee - -|**Local (standalone)** -|Strong (ACID per modality) -|Read-your-writes within same modality - -|**Remote (federated)** -|Eventual (drift-tolerant) -|Stale reads possible, bounded by drift detection interval - -|**Hybrid (mixed)** -|**Undefined** -|Cross-modal queries may observe inconsistent state -|=== - -**Example Inconsistency Scenario:** -[source,text] ----- -T=0: Client writes Octad update (Document + Vector) -T=1: Local Vector modality updates immediately -T=5: Remote Document modality still processing (network delay) -T=6: Client queries: Vector reflects new data, Document shows old data -Result: Inconsistent cross-modal view ----- - -**Consequences:** -- **Anomalous Query Results:** Joins across local/remote modalities return incomplete data -- **Debugging Difficulty:** Inconsistencies appear non-deterministically depending on network timing -- **Semantic Violations:** Business logic assumptions (e.g., "Vector embedding always matches Document text") break - -**Mitigation:** -- **Causal Consistency:** Implement **vector clocks** or **hybrid logical clocks** to track causality across local/remote boundary -- **Read Repair:** On query, check if remote modality is stale, trigger **on-demand sync** before returning results -- **Timestamp Bounds:** Return query results with **timestamp range** indicating freshness (e.g., "data as of T=5 ± 30s") - -=== 1.4. Configuration Explosion - -**Problem:** Hybrid mode requires **per-modality configuration** specifying local vs remote. - -**Configuration File Complexity:** -[source,elixir] ----- -# config/hybrid.exs -config :verisim, - deployment_mode: :hybrid, - local_modalities: [ - graph: %{ - crate: "verisim-graph-rs", - resources: %{ram: "32GB", cpu: "4 cores"} - }, - vector: %{ - crate: "verisim-vector-rs", - resources: %{ram: "64GB", cpu: "8 cores"} - }, - tensor: %{ - crate: "verisim-tensor-rs", - resources: %{ram: "16GB", gpu: "4x A100"} - } - ], - federated_modalities: [ - document: %{ - store_endpoint: "https://archive.org/verisim", - backup_endpoint: "https://mirror.archive.org/verisim", - timeout: 5000, - retry_policy: %{max_attempts: 3, backoff: :exponential} - }, - temporal: %{ - store_endpoint: "https://uni-b.edu/verisim", - backup_endpoint: nil, - timeout: 10_000, - retry_policy: %{max_attempts: 1, backoff: :none} - }, - semantic: %{ - store_endpoint: "https://zkp-service.org/verisim", - backup_endpoint: "https://zkp-backup.org/verisim", - timeout: 2000, - retry_policy: %{max_attempts: 5, backoff: :constant} - } - ], - drift_detection: %{ - local_interval: 60_000, # 1 minute for local - remote_interval: 300_000, # 5 minutes for remote - cross_boundary_check: true - } ----- - -**Operational Burden:** -- **Environment-Specific:** Dev/staging/prod likely have **different** hybrid configurations -- **Drift Risk:** Config files diverge across environments → "works in staging, fails in prod" -- **Testing Complexity:** Must test all combinations of local/remote modality placement - -**Mitigation:** -- **Config Validation:** Add Elixir schema validation to catch errors at compile time -- **Infrastructure as Code:** Use Terraform/Pulumi to generate hybrid configs from **single source of truth** -- **Config Testing:** Add integration tests that verify all declared modalities are reachable - -== 2. Performance Challenges - -=== 2.1. Mixed Latency Profiles - -**Problem:** Queries that touch both local and remote modalities have **unpredictable latency**. - -**Latency Distribution:** -[source,text] ----- -Pure Local Query: - p50: 5ms, p99: 20ms, p999: 50ms ✓ Predictable - -Pure Remote Query: - p50: 80ms, p99: 500ms, p999: 5000ms ✗ High tail latency - -Hybrid Query (Local Vector + Remote Document): - p50: 85ms, p99: 520ms, p999: 5050ms ✗✗ Dominated by remote ----- - -**User Impact:** -- **Inconsistent UX:** Some queries feel instant, others sluggish -- **Timeout Tuning Difficulty:** Single timeout value fails (too short → false timeouts, too long → hangs) -- **SLA Violation Risk:** Cannot commit to p99 < 100ms when remote stores involved - -**Mitigation:** -- **Query Planning:** Analyze query to determine if local-only path exists, use it -- **Adaptive Timeouts:** Adjust timeout dynamically based on query plan (local-only: 100ms, hybrid: 5s) -- **Caching:** Cache remote query results locally (invalidate on drift detection event) - -=== 2.2. Cross-Boundary Drift Detection Overhead - -**Problem:** Detecting drift across local/remote boundary requires **network round-trips**. - -**Drift Detection Protocol (Simplified):** -[source,text] ----- -1. Orchestrator queries local Vector modality → embedding_local -2. Orchestrator queries remote Document modality → embedding_remote (via HTTP) -3. Compute cosine_similarity(embedding_local, embedding_remote) -4. If similarity < threshold → trigger drift alert ----- - -**Overhead:** -- **Network Cost:** Each drift check = 1 HTTP request to remote store -- **Frequency Tradeoff:** Check often (fresh data) vs rarely (low overhead) -- **False Positives:** Network latency variance causes spurious drift alerts - -**Quantitative Impact:** -- Drift checks every 60s × 1000 Octads = **16.6 qps to remote stores** (just for monitoring) -- At scale (100k Octads), drift monitoring alone consumes **1.6k qps** - -**Mitigation:** -- **Sampling:** Only check drift for **subset** of Octads per interval (e.g., 1% sample) -- **Push-Based Updates:** Remote stores **push** drift events to orchestrator (avoid polling) -- **Bloom Filters:** Use probabilistic data structures to quickly identify unchanged Octads - -=== 2.3. Write Propagation Delay - -**Problem:** Writes to local modalities propagate **instantly**, writes to remote modalities are **eventually consistent**. - -**Example Timeline:** -[source,text] ----- -T=0: Client writes Octad (updates Vector + Document) -T=1: Local Vector modality persists immediately -T=50: Network delay to remote Document store -T=100: Remote Document store persists -T=150: Drift detection discovers propagation lag -T=200: Drift repair job enqueued (if policy requires sync) ----- - -**Consequences:** -- **Read-After-Write Inconsistency:** Client reads back Octad immediately, sees old Document content -- **Cross-Modal Queries Fail:** Join between Vector (updated) and Document (stale) returns partial results -- **Audit Log Ambiguity:** Temporal log shows write at T=0, but Document not updated until T=100 - -**Mitigation:** -- **Write-Through Cache:** Orchestrator caches pending writes, serves from cache until remote ack -- **Versioning:** Tag each modality update with **vector clock**, expose version in query API -- **Async Writes with Callback:** Return write receipt immediately, notify client when remote persisted - -=== 2.4. Hot Modality Cannot Scale Independently - -**Problem:** If local Vector modality is hot, cannot add **Vector-only replicas** without duplicating other local modalities. - -**Scaling Constraint:** -- Standalone mode limitation: scaling hot modality requires scaling entire node -- Hybrid mode partially helps: can **move cold modalities to remote**, leaving local resources for hot modality -- But: still cannot add **multiple local Vector instances** without multiple orchestrator nodes - -**Workaround:** -- **Read Replicas:** Deploy read-only Vector replicas behind load balancer (but: replication lag!) -- **Cache Layer:** Add Redis cache in front of Vector modality (but: cache invalidation!) -- **Migrate to Full Federation:** Split Vector modality across multiple federated stores (but: loses hybrid simplicity!) - -== 3. Operational Challenges - -=== 3.1. Monitoring Across Boundaries - -**Problem:** Must monitor **both local and remote** modalities with different tools/access patterns. - -**Monitoring Requirements:** -[cols="1,2,2"] -|=== -|Aspect |Local Modalities |Remote Modalities - -|**Metrics Collection** -|Direct Prometheus scrape from localhost -|Must scrape remote store's `/metrics` (may not expose) - -|**Logging** -|Local syslog/journald -|Remote store logs (may not have access) - -|**Tracing** -|Direct OpenTelemetry instrumentation -|Requires remote store to export spans (may not support) - -|**Alerting** -|Orchestrator can detect crashes immediately -|Must rely on **health checks** (delayed detection) -|=== - -**Operational Complexity:** -- **Dual Dashboards:** Need separate Grafana dashboards for local vs remote -- **Alerting Gaps:** May not detect remote store issues until queries time out -- **Root Cause Analysis:** Difficult to determine if issue is local, network, or remote - -**Mitigation:** -- **Unified Observability:** Require remote stores to export metrics/logs/traces to **central collector** -- **Synthetic Monitoring:** Run continuous health checks from orchestrator to remote stores -- **Federated Prometheus:** Use Prometheus federation to scrape remote store metrics - -=== 3.2. Backup Strategy Divergence - -**Problem:** Local and remote modalities require **different backup approaches**. - -**Backup Strategies:** -[cols="1,2,2"] -|=== -|Modality Location |Backup Method |Responsibility - -|**Local** -|Snapshot to S3 every 1 hour -|Orchestrator operator - -|**Remote** -|Archive.org's internal backup (unknown schedule) -|Remote store operator (out of your control) - -|**Hybrid** -|**Inconsistent:** Local backed up hourly, remote maybe daily? -|**Split responsibility** → coordination overhead -|=== - -**Risk Scenarios:** -1. **Partial Restore:** Restore local modalities from backup, but remote modalities have newer data → drift -2. **Remote Store Disaster:** Archive.org loses data, your backups don't include Document modality -3. **RPO Mismatch:** Local modalities have 1-hour RPO, remote have 24-hour RPO → uneven data loss - -**Mitigation:** -- **SLA Requirements:** Contractually require remote stores to meet **minimum backup SLAs** -- **Local Caching:** Cache remote modality data locally (stale copy better than no copy) -- **Disaster Recovery Testing:** Quarterly DR drills that simulate remote store failure - -=== 3.3. Dependency on External Services - -**Problem:** Hybrid mode **depends on remote stores** being available and performant. - -**Failure Modes:** -1. **Remote Store Downtime:** Archive.org goes down → Document queries fail -2. **Network Partition:** Internet link fails → remote modalities unreachable -3. **Remote Store Overload:** Archive.org rate-limits your requests → queries timeout -4. **Remote Store Deprecation:** Archive.org shuts down VeriSimDB API → forced migration - -**Blast Radius:** -- Partial system outage (remote modalities unavailable) -- Degraded query results (cross-modal queries incomplete) -- Reputation damage (users perceive VeriSimDB as unreliable) - -**Mitigation:** -- **Fallback Stores:** Configure **backup remote stores** for each modality -- **Circuit Breakers:** Automatically fail over to backup if primary remote store unhealthy -- **Local Fallback Cache:** Serve stale data from local cache during remote outage (with staleness warning) -- **Contractual SLAs:** Negotiate **uptime SLAs** with remote store operators (99.9%+) - -=== 3.4. Cross-Boundary Security Policies - -**Problem:** Local modalities enforce **local policies**, remote modalities enforce **their own policies** → coordination required. - -**Policy Conflict Example:** -[source,text] ----- -Octad 550e8400-... access policy: - Local Vector: "Allow all authenticated users" - Remote Document: "Allow only .edu domains" - -Query from user@gmail.com: - ✓ Allowed to read Vector embedding - ✗ Denied access to Document content - Result: Partial query result (confusing UX) ----- - -**Consequences:** -- **Inconsistent Access Control:** User can access some modalities but not others -- **Policy Drift:** Local and remote policies diverge over time (no sync mechanism) -- **Audit Complexity:** Must trace authorization across multiple policy engines - -**Mitigation:** -- **Unified Policy Registry:** Store **global policy hash** in ReScript registry, both local and remote enforce same policy -- **Policy Synchronization:** Orchestrator **pushes policy updates** to remote stores when changed -- **Fail-Secure Mode:** If policies conflict, deny access (prefer security over availability) - -== 4. Scaling Challenges - -=== 4.1. Asymmetric Scaling - -**Problem:** Local and remote modalities scale **independently** → resource imbalance. - -**Scaling Scenarios:** -[cols="1,2,2,2"] -|=== -|Scenario |Local Modalities |Remote Modalities |Impact - -|**Query Spike** -|Hit vertical limit (maxed out) -|Scale horizontally (remote adds nodes) -|**Asymmetry:** Local becomes bottleneck - -|**Data Growth** -|Disk fills up (must upgrade) -|Remote scales seamlessly -|**Cost:** Forced hardware upgrade for local - -|**Regional Expansion** -|Cannot replicate (single node) -|Remote has CDN-like distribution -|**Latency:** Remote users slow despite remote stores nearby -|=== - -**Operational Complexity:** -- Must **separately** plan capacity for local vs remote -- Cannot leverage remote scalability for local workloads -- May need to **migrate local → remote** dynamically as load increases - -**Mitigation:** -- **Autoscaling Local:** Use VM autoscaling to add local capacity (requires migration to cloud) -- **Overflow to Remote:** When local capacity exhausted, temporarily route queries to remote (degraded latency) -- **Hybrid-to-Federated Migration:** Prepare to move all modalities to federated if scaling pain exceeds hybrid benefits - -=== 4.2. Hot Data Migration - -**Problem:** If "cold" remote modality becomes "hot", must **migrate to local** → data transfer cost and downtime. - -**Migration Steps:** -1. **Identify Hot Data:** Query logs show Remote Document now 500 qps (was 50 qps) -2. **Provision Local Capacity:** Add 2TB disk for Tantivy index -3. **Export from Remote:** Request data dump from Archive.org (may take days) -4. **Import to Local:** Load data into local Document modality (hours) -5. **Cutover:** Update orchestrator config to route Document queries locally -6. **Deregister Remote:** Remove Archive.org Document endpoint from registry - -**Downtime:** Potentially 6-48 hours depending on data volume and remote store cooperation - -**Cost:** -- **Data Transfer:** Egress fees from remote store (can be $$$) -- **Dual Storage:** Temporarily storing data both locally and remotely during migration -- **Engineering Time:** Migration is manual, error-prone process - -**Mitigation:** -- **Pre-Emptive Caching:** Cache hot data locally **before** fully migrating (reduces cutover time) -- **Gradual Migration:** Migrate data in **chunks** (e.g., 10% per day) while dual-running -- **Remote Store API:** Negotiate **incremental export API** with remote stores (avoid full dumps) - -=== 4.3. Network Bandwidth Bottleneck - -**Problem:** High query volume to remote modalities saturates **network link**. - -**Bandwidth Analysis:** -[source,text] ----- -Assumption: - - 500 qps to remote Document modality - - Avg response size: 100 KB - - Bandwidth required: 500 qps × 100 KB = 50 MB/s = 400 Mbps - -Constraint: - - Orchestrator has 1 Gbps uplink - - Other services share link (Elixir orchestrator, monitoring, etc.) - - Effective available: ~600 Mbps - -Risk: At 750 qps, network saturates → queries timeout ----- - -**Consequences:** -- **Query Timeouts:** Network congestion causes packet loss -- **Cascading Failures:** Retries amplify bandwidth usage -- **Noisy Neighbor:** Remote queries starve local traffic (monitoring, SSH, etc.) - -**Mitigation:** -- **Response Compression:** Enable gzip/brotli for HTTP responses (reduce size by ~70%) -- **CDN/Caching:** Place CDN (e.g., Cloudflare) in front of remote stores (cache hot data near orchestrator) -- **QoS:** Configure Quality of Service to prioritize critical traffic over bulk transfers -- **Upgrade Link:** Negotiate 10 Gbps uplink if remote query volume critical - -== 5. Drift Detection Challenges - -=== 5.1. Cross-Boundary Drift Semantics - -**Problem:** Drift between **local modalities** has different semantics than drift **across local/remote boundary**. - -**Drift Types:** -[cols="1,2,2"] -|=== -|Drift Type |Example |Severity - -|**Local-Local Drift** -|Local Vector embedding diverges from local Graph edges -|**High** (same system, should stay in sync) - -|**Local-Remote Drift** -|Local Vector embedding diverges from remote Document text -|**Medium** (expected due to eventual consistency) - -|**Remote-Remote Drift** -|Remote Document (Archive.org) diverges from remote Temporal (Uni-B) -|**Low** (different operators, expected lag) -|=== - -**Challenge:** Drift detection policy must distinguish between expected (remote lag) and anomalous (local desync) drift. - -**Mitigation:** -- **Separate Thresholds:** Use different drift thresholds for local-local vs local-remote -- **Time-Windowed Detection:** Only alert on remote drift if lag exceeds **expected propagation time** -- **Manual Review Required:** Flag cross-boundary drift for human review (don't auto-repair blindly) - -=== 5.2. Network Overhead of Continuous Monitoring - -**Problem:** Drift detection requires **continuous sampling** of remote modalities → network cost. - -**Monitoring Load:** -[source,text] ----- -Scenario: 10,000 Octads, drift check every 5 minutes - -Checks per hour: 10,000 × (60 / 5) = 120,000 checks/hour -If each check = 1 HTTP request of 1 KB: - Network: 120,000 × 1 KB = 120 MB/hour = 2.8 GB/day - -At scale (1M Octads): 280 GB/day just for drift monitoring ----- - -**Consequence:** Drift monitoring becomes **major cost driver** (bandwidth, remote store load) - -**Mitigation:** -- **Adaptive Sampling:** Only check drift for Octads recently accessed (others assumed stable) -- **Event-Driven Updates:** Remote stores **push** change notifications instead of pull-based polling -- **Batching:** Batch multiple drift checks into single HTTP request (reduce overhead) - -=== 5.3. Repair Policy Coordination - -**Problem:** Repairing drift across local/remote boundary requires **coordination** with remote store operator. - -**Repair Workflow (Complex):** -[source,text] ----- -1. Orchestrator detects drift: Local Vector ≠ Remote Document -2. Determine repair direction: - - Option A: Update local to match remote (download from remote) - - Option B: Update remote to match local (push to remote) - - Option C: Manual review required (conflict) -3. If Option B: - - Orchestrator generates signed update request - - Sends to remote store's /update endpoint - - Remote store validates signature, applies update - - Remote store returns ack - - Orchestrator confirms drift resolved -4. If Option C: - - Orchestrator creates "drift alert" Octad - - Sends notification to domain custodians - - Waits for manual resolution - - Custodian signs resolution decision - - Orchestrator applies resolution ----- - -**Complexity:** -- **Multi-Party Coordination:** Requires cooperation from remote store (may not respond) -- **Policy Conflicts:** Local policy may say "auto-repair", remote policy may say "manual review" -- **Repair Latency:** Manual review can take hours-to-days - -**Mitigation:** -- **Pre-Negotiated Policies:** Establish **repair SLAs** with remote stores before federation -- **Automated Escalation:** If remote store doesn't respond within N hours, escalate to manual review -- **Repair Auditing:** Log all repair actions to immutable Temporal log (accountability) - -== 6. Security Challenges - -=== 6.1. Expanded Attack Surface - -**Problem:** Hybrid mode exposes orchestrator to **both local and internet-based attacks**. - -**Attack Vectors:** -[cols="1,2,2"] -|=== -|Vector |Local (Standalone) |Hybrid - -|**Container Escape** -|✓ Risk exists -|✓ Same risk - -|**Remote Store Compromise** -|✗ N/A -|✓ **New risk:** Remote store breached → sends malicious responses - -|**Man-in-the-Middle** -|✗ N/A (localhost only) -|✓ **New risk:** HTTPS connection to remote store intercepted - -|**DDoS** -|✗ Limited (internal network) -|✓ **New risk:** Remote store DDoSed → cascading failure -|=== - -**Mitigation:** -- **mTLS:** Use mutual TLS for orchestrator ↔ remote store communication (prevent MITM) -- **Response Validation:** Verify signatures on **all** remote responses (detect tampering) -- **Rate Limiting:** Limit queries to remote stores (prevent amplification attacks) -- **Zero-Trust:** Treat remote stores as untrusted (validate all data before use) - -=== 6.2. Policy Enforcement at Boundary - -**Problem:** Orchestrator must enforce **global access policies** while remote stores enforce **local policies**. - -**Policy Layering:** -[source,text] ----- -User Query: "Read Octad 550e8400-..." - -Layer 1 (Orchestrator): Check global policy - - Is user authenticated? ✓ - - Does user have read permission for Octad? ✓ - - Forward to remote Document store - -Layer 2 (Remote Store): Check local policy - - Is orchestrator authorized? ✓ - - Does user meet remote's domain restrictions (.edu)? ✗ - - Return 403 Forbidden - -Result: Query denied at Layer 2 despite passing Layer 1 ----- - -**Consequences:** -- **Inconsistent UX:** Users confused why some modalities accessible, others not -- **Debugging Difficulty:** Must check **two** policy engines to understand denial -- **Security Gaps:** If Layer 2 fails open, global policy bypassed - -**Mitigation:** -- **Policy Synchronization:** Push orchestrator's global policy to remote stores (enforce consistently) -- **Pre-Flight Checks:** Orchestrator queries remote store's policy **before** forwarding request -- **Fail-Secure:** If policies conflict, deny access (prefer security over availability) - -== 7. Cost Challenges - -=== 7.1. Dual Infrastructure Costs - -**Problem:** Hybrid mode requires **both local and remote** infrastructure → double the cost? - -**Cost Breakdown:** -[cols="1,2,2,2"] -|=== -|Component |Standalone |Hybrid |Increase - -|**Local Hardware** -|$10k/month (all modalities) -|$6k/month (3 modalities) -|**-40%** ✓ - -|**Remote Store Fees** -|$0 -|$3k/month (3 modalities) -|**+$3k** ✗ - -|**Network Egress** -|$0 (localhost) -|$1k/month (to remote) -|**+$1k** ✗ - -|**Total** -|$10k/month -|$10k/month -|**Break-even** -|=== - -**Reality:** Hybrid is **not cheaper** unless: -- Remote stores offer significantly cheaper storage (e.g., Archive.org free tier) -- Local hardware would require expensive upgrade to support all modalities -- Network egress is negligible (e.g., colocated datacenters) - -**Mitigation:** -- **Cost Analysis:** Model costs before migrating to hybrid (may not save money) -- **Negotiate Pricing:** Bulk discounts from remote stores for sustained usage -- **Regional Optimization:** Use remote stores in same datacenter (reduce egress) - -=== 7.2. Operational Overhead - -**Problem:** Hybrid mode requires **more operational expertise** than standalone. - -**Skills Required:** -- **Standalone:** Understand 6 modality stores + Elixir orchestrator (7 components) -- **Hybrid:** All above + Network troubleshooting + Remote store APIs + Policy coordination (10+ components) - -**Hiring Impact:** May need to hire **DevOps specialist** for hybrid management (+$150k/year salary) - -**Mitigation:** -- **Managed Services:** Use managed VeriSimDB-as-a-Service (offload operational burden) -- **Documentation:** Invest in runbooks and troubleshooting guides -- **Monitoring:** Heavy monitoring investment to reduce MTTR (Mean Time To Repair) - -== 8. Summary of Challenges - -[cols="1,2,1"] -|=== -|Challenge Category |Key Issues |Severity - -|**Architecture** -|Hot/cold boundary design, network topology, partial consistency, config explosion -|HIGH - -|**Performance** -|Mixed latency, drift detection overhead, write propagation delay -|HIGH - -|**Operations** -|Cross-boundary monitoring, backup divergence, external dependencies, policy coordination -|MEDIUM - -|**Scaling** -|Asymmetric scaling, hot data migration, network bandwidth bottleneck -|MEDIUM - -|**Drift** -|Cross-boundary semantics, monitoring overhead, repair coordination -|HIGH - -|**Security** -|Expanded attack surface, dual policy enforcement -|MEDIUM - -|**Cost** -|Dual infrastructure, operational overhead -|LOW (situational) -|=== - -== 9. When to Choose Hybrid Despite Challenges - -Hybrid deployment is appropriate when: - -1. **Clear Hot/Cold Split:** 80% queries to 2-3 "hot" modalities, rest are cold -2. **Cost Pressure:** Local hardware expensive, remote storage cheap (e.g., archival) -3. **Institutional Requirements:** Must federate certain modalities (e.g., Document with university archives) while keeping others local (e.g., Vector for performance) -4. **Growth Path:** Starting standalone, planning eventual full federation (hybrid is stepping stone) -5. **Team Capacity:** DevOps team skilled in distributed systems (can manage complexity) - -For simpler deployments, **standalone** may be better. For full-scale federation, skip hybrid and go **directly to federated** mode. - -== See Also - -- link:deployment-modes.adoc[Deployment Modes Overview] -- link:challenges-standalone.adoc[Challenges: Standalone Mode] -- link:challenges-federated.adoc[Challenges: Federated Mode] -- link:../WHITEPAPER.md[VeriSimDB White Paper] diff --git a/verisimdb/docs/challenges-standalone.adoc b/verisimdb/docs/challenges-standalone.adoc deleted file mode 100644 index 8344eff6..00000000 --- a/verisimdb/docs/challenges-standalone.adoc +++ /dev/null @@ -1,439 +0,0 @@ -// SPDX-License-Identifier: CC-BY-SA-4.0 - -= Challenges: Standalone Deployment Mode - -== Overview - -Standalone deployment runs VeriSimDB as a **traditional database** with all six modalities (Graph, Vector, Tensor, Semantic, Document, Temporal) hosted locally on the same machine or cluster. While this offers simplicity and performance, it introduces specific architectural, operational, and scaling challenges. - -This document details the **technical challenges** unique to standalone deployment and provides mitigation strategies where applicable. - -== 1. Architectural Challenges - -=== 1.1. Vertical Scaling Limits - -**Problem:** All modalities compete for shared resources (CPU, RAM, disk I/O). - -[cols="1,2,2"] -|=== -|Resource |Constraint |Impact - -|**RAM** -|Oxigraph (Graph) + HNSW (Vector) + ndarray (Tensor) all hold data structures in memory -|OOM kills under heavy concurrent load - -|**Disk I/O** -|Tantivy (Document) writes inverted indices while Temporal writes Merkle-tree snapshots -|I/O contention degrades latency - -|**CPU** -|Tensor operations (Burn inference) saturate cores during batch jobs -|Query starvation for other modalities -|=== - -**Mitigation:** -- Use **cgroup limits** to reserve resources per modality (e.g., `systemd` slices) -- Run modality stores in separate **containers** with CPU/memory quotas -- Profile with `perf` to identify hotspots and optimize Rust code paths - -=== 1.2. Single Point of Failure - -**Problem:** Loss of the standalone node means loss of all six modalities simultaneously. - -**Risk Scenarios:** -1. Hardware failure → No Graph, Vector, Tensor, Semantic, Document, OR Temporal access -2. Kernel panic → All GenServers crash (Elixir orchestrator down) -3. Disk corruption → All modalities potentially compromised - -**Mitigation:** -- **Replication:** Run multiple standalone nodes with **async replication** (see `verisim-temporal` Merkle-tree sync) -- **Backups:** Snapshot all modality stores to S3/B2 every N hours -- **High Availability:** Deploy behind HAProxy with health checks per modality - -**Trade-off:** Replication adds complexity approaching hybrid/federated mode challenges. - -=== 1.3. Modality Coupling - -**Problem:** All six modalities share the same **Elixir orchestration layer**, creating tight coupling. - -**Consequences:** -- **Upgrade Hell:** Updating `verisim-graph-rs` (Oxigraph) requires restarting the entire stack -- **Blast Radius:** A bug in `verisim-tensor-rs` can crash the GenServer supervision tree, taking down Document/Temporal -- **Dependency Conflicts:** Rust crate version mismatches (e.g., `tokio` version across modalities) force global coordination - -**Mitigation:** -- Use **modular Elixir releases** (`mix release`) with per-modality umbrella apps -- Implement **graceful degradation**: if Vector modality crashes, Graph/Document continue serving -- Adopt **interface-driven design**: all modality crates expose identical C-ABI functions - -== 2. Performance Challenges - -=== 2.1. Interference Between Modalities - -**Problem:** Workloads from different modalities compete unpredictably. - -**Example Scenario:** -[source,text] ----- -T=0: Client A starts HNSW vector search (high CPU) -T=5: Client B starts Tantivy full-text query (high disk read) -T=10: Client C requests Oxigraph SPARQL (high RAM) -Result: All three queries degrade due to resource contention ----- - -**Quantitative Impact:** -- Vector search latency: 5ms → 50ms (10x) -- Document queries: 20ms → 200ms (10x) -- Graph queries: OOM → killed - -**Mitigation:** -- **Query Scheduling:** Implement priority queues in `VeriSim.QueryRouter` -- **Async I/O:** Use `tokio` async runtime in Rust crates to yield during I/O waits -- **Read Replicas:** Add read-only replicas for heavy query workloads - -=== 2.2. Write Amplification - -**Problem:** A single Octad update can trigger writes across **all six modalities**. - -**Example:** -1. User updates Octad `550e8400-...` with new document text -2. **Document modality**: Tantivy reindexes (writes segments) -3. **Vector modality**: Recomputes embedding, updates HNSW index (rebuilds layers) -4. **Temporal modality**: Appends new Merkle-tree node (writes snapshot) -5. **Semantic modality**: Validates CBOR proof (writes metadata) -6. **Graph modality**: May add citation edges (writes RDF triples) - -**Write Amplification Factor:** ~6x per logical update - -**Consequences:** -- **SSD Wear:** Premature SSD exhaustion (TBW limits) -- **Latency Spikes:** Write storms block reads -- **Replication Lag:** If using async replication, lag increases proportionally - -**Mitigation:** -- **Lazy Updates:** Only write to modalities that actually changed -- **Batching:** Buffer updates in memory, flush every N seconds -- **Write Coalescing:** Merge multiple updates to same Octad before persisting - -=== 2.3. Query Routing Overhead - -**Problem:** `VeriSim.QueryRouter` (Elixir GenServer) becomes bottleneck. - -**Bottleneck Analysis:** -- All queries pass through single GenServer process -- Erlang message passing adds ~1-5µs per routing decision -- Under high load (>10k qps), GenServer mailbox saturates - -**Symptoms:** -- Increasing `gen_server:call` timeouts -- Queries queued despite modality stores idle -- Orchestrator CPU at 100% (single core) - -**Mitigation:** -- **Pooling:** Use `poolboy` to spawn N parallel routers (shard by Octad UUID) -- **Direct Dispatch:** For hot paths, bypass GenServer and call Rust NIFs directly -- **Load Balancing:** Use `pg` (process groups) to distribute across BEAM scheduler cores - -== 3. Operational Challenges - -=== 3.1. Backup and Recovery Complexity - -**Problem:** Six modalities require **six separate backup strategies**. - -[cols="1,2,2"] -|=== -|Modality |Backup Method |Recovery Complexity - -|**Graph** -|Oxigraph RDF dump → `.nt` file -|Restore requires full re-parse (slow for large graphs) - -|**Vector** -|HNSW index serialization → binary blob -|Cannot restore incrementally; all-or-nothing - -|**Tensor** -|`ndarray` → HDF5 or `.npy` files -|Must re-validate tensor shapes on restore - -|**Semantic** -|CBOR blobs → JSON export -|Schema validation may fail if CBOR spec evolved - -|**Document** -|Tantivy segments → tarball -|Index rebuild required if segments corrupted - -|**Temporal** -|Merkle-tree nodes → recursive snapshot -|Tree integrity check expensive on large histories -|=== - -**Operational Burden:** -- **Inconsistent Snapshots:** Modalities backed up at different times → cross-modal inconsistency -- **Restore Testing:** Must test restore for all 6 modalities (6x testing effort) -- **Storage Costs:** 6 separate backups consume significant space - -**Mitigation:** -- **Coordinated Snapshots:** Use `VeriSim.DriftMonitor` to trigger simultaneous backups -- **Differential Backups:** Only back up changed modality stores (requires version tracking) -- **Restore Automation:** Script full-stack restore and test in CI/CD - -=== 3.2. Monitoring and Observability - -**Problem:** Need to monitor **six independent data stores** plus orchestration layer. - -**Metrics to Track (per modality):** -- Query latency (p50, p99, p999) -- Write throughput (ops/sec) -- Error rates (by error type) -- Resource usage (CPU, RAM, disk I/O) -- Cache hit rates (where applicable) - -**Operational Complexity:** -- **7 dashboards** (6 modalities + orchestrator) → alert fatigue -- **Correlation Difficulty:** Determining which modality caused a slowdown requires cross-dashboard analysis -- **Logging Volume:** Combined logs from all modalities can reach GB/day - -**Mitigation:** -- **Unified Metrics:** Export all modalities to Prometheus with consistent labels (`modality=graph`, etc.) -- **Distributed Tracing:** Instrument with OpenTelemetry to trace queries across modalities -- **Anomaly Detection:** Use ML-based tools (e.g., Prometheus Alertmanager + `aiops`) to detect correlated failures - -=== 3.3. Dependency Management Hell - -**Problem:** Six Rust crates with divergent dependency trees. - -**Conflict Example:** -[source,toml] ----- -# verisim-graph-rs/Cargo.toml -tokio = { version = "1.35", features = ["rt-multi-thread"] } - -# verisim-vector-rs/Cargo.toml -tokio = { version = "1.40", features = ["rt-multi-thread", "io-util"] } - -# Conflict: Both crates linked into same Elixir NIF → panic ----- - -**Consequences:** -- **Build Failures:** Cargo refuses to resolve incompatible versions -- **Runtime Panics:** Mismatched async runtimes cause deadlocks -- **Security Vulnerabilities:** Stuck on old dependency versions due to conflicts - -**Mitigation:** -- **Workspace Unification:** Use Cargo workspace to enforce single `tokio` version across all crates -- **Feature Flags:** Abstract over incompatible features with conditional compilation -- **Forking:** Fork problematic dependencies and vendor them (last resort) - -== 4. Scaling Challenges - -=== 4.1. Data Growth Without Horizontal Scaling - -**Problem:** Standalone mode lacks **horizontal partitioning** (sharding). - -**Growth Scenarios:** -1. **Graph grows to 1B triples** → Oxigraph in-memory structures exceed RAM -2. **Vector index reaches 100M embeddings** → HNSW construction time becomes prohibitive (hours) -3. **Document corpus hits 10TB** → Tantivy segment count explodes, query latency degrades - -**Scaling Options (All Suboptimal):** -- **Vertical Scaling:** Buy bigger machine → diminishing returns, CapEx explosion -- **Manual Sharding:** Split Octads across multiple standalone nodes → lose single-system image -- **Migrate to Hybrid/Federated:** Requires architectural overhaul - -**Constraint:** Standalone architecture fundamentally incompatible with horizontal scaling. - -=== 4.2. Hot Modality Bottlenecks - -**Problem:** Some modalities become "hot" (high query rate) while others remain idle. - -**Example:** -- **Vector modality:** 95% of queries (embedding search for RAG pipelines) -- **Temporal modality:** 1% of queries (historical lookups) -- **Result:** Vector store saturated, Temporal store wasted capacity - -**Consequences:** -- **Inefficient Resource Use:** Paying for 6 modalities, using 1-2 heavily -- **Cannot Scale Hot Path:** Standalone mode cannot add more Vector replicas without duplicating all modalities - -**Mitigation:** -- **Read Replicas (Partial):** Add read-only Vector replicas behind load balancer -- **Caching Layer:** Add Redis cache for hot Vector queries -- **Profile-Driven Optimization:** Use query logs to identify hot modalities and optimize those code paths - -=== 4.3. Write Scaling Limits - -**Problem:** All writes funnel through single orchestrator GenServer. - -**Theoretical Limit:** -- Elixir GenServer: ~50k messages/sec per process (microbenchmark) -- Each write requires ~5 GenServer calls (validation, routing, 6 modality dispatches, logging) -- **Practical Write Throughput:** ~10k writes/sec - -**Symptom Under Load:** -- GenServer mailbox depth increases linearly -- Timeout errors on `GenServer.call(..., 5000)` despite modality stores idle -- Writes queued for seconds/minutes - -**Mitigation:** -- **Pooling:** Use `poolboy` to spawn N parallel `EntityServer` processes -- **Sharding:** Shard Octads by UUID across multiple orchestrator nodes (requires migration to distributed Elixir) -- **Async Writes:** Use `GenServer.cast` for non-critical writes (lose ordering guarantees) - -== 5. Security Challenges - -=== 5.1. Single Security Boundary - -**Problem:** Compromise of standalone node exposes **all six modalities** simultaneously. - -**Attack Scenarios:** -1. **Container Escape:** Attacker gains root on host → accesses all modality stores -2. **Elixir VM Exploit:** Remote code execution in Erlang VM → full database compromise -3. **Rust NIF Bug:** Buffer overflow in modality crate → arbitrary code execution with DB privileges - -**Blast Radius:** 100% of data (all modalities, all Octads) - -**Mitigation:** -- **Defense in Depth:** - - Run modalities in **separate rootless containers** (`svalinn/vordr`) - - Use **seccomp profiles** to restrict syscalls per container - - Enable **SELinux/AppArmor** mandatory access control -- **Least Privilege:** Modality stores run as non-root user with minimal file permissions -- **Network Isolation:** Use **separate network namespaces** for modalities (Unix sockets instead of TCP) - -=== 5.2. Audit Log Single Point of Failure - -**Problem:** `verisim-temporal` (audit log) shares fate with other modalities. - -**Risk:** -- If standalone node compromised, attacker can **modify or delete audit logs** -- **No external verification** of log integrity - -**Mitigation:** -- **Immutable Logs:** Use **append-only filesystems** (e.g., `chattr +a` on Linux) -- **External Replication:** Stream audit logs to **remote syslog server** or **blockchain** (tamper-evident) -- **Cryptographic Chaining:** Each log entry includes hash of previous entry (Merkle chain) - -== 6. Migration Challenges - -=== 6.1. Path to Federation - -**Problem:** If standalone deployment outgrows single machine, migration to federated mode is **non-trivial**. - -**Migration Steps Required:** -1. **Split Modalities:** Decide which modalities move to which federated stores -2. **Data Export:** Dump each modality's data (see Backup section for complexity) -3. **Registry Setup:** Deploy ReScript registry with quorum (requires 3+ nodes for Raft) -4. **Policy Rewrite:** Translate single-machine access policies to distributed policies -5. **Re-import Data:** Load data into federated stores, register with registry -6. **Cutover:** Update orchestrator config, test cross-store queries - -**Downtime:** Potentially hours-to-days depending on data volume - -**Risk:** Data inconsistency during migration if updates occur mid-process - -**Mitigation:** -- **Dual-Write Phase:** Write to both standalone and federated stores during transition -- **Shadow Mode:** Run federated cluster in parallel, compare results before cutover -- **Rollback Plan:** Keep standalone node operational until federated cluster validated - -=== 6.2. Hybrid Mode Complexity - -**Problem:** Partial migration to hybrid (some modalities local, some remote) increases operational complexity. - -**Challenges:** -- **Network Latency Variance:** Local modalities <1ms, remote modalities 10-100ms → unpredictable query times -- **Cross-Modal Drift:** Drift detection now spans local/remote boundary → more complex reconciliation logic -- **Configuration Explosion:** Must track which modalities are local vs remote per deployment - -**Recommendation:** Only migrate to hybrid if **clear performance/cost benefit** (e.g., archive old Temporal data remotely, keep hot Vector data local) - -== 7. Cost Challenges - -=== 7.1. Over-Provisioning - -**Problem:** Must provision hardware for **peak load across all six modalities**. - -**Example:** -- Vector modality needs 64GB RAM for HNSW index -- Graph modality needs 32GB RAM for RDF triples -- Tensor modality needs 4x GPUs for inference -- **Total:** 96GB RAM + 4 GPUs, even if modalities rarely peak simultaneously - -**Cost Impact:** Paying for capacity that sits idle 80% of the time - -**Alternative:** Federated/hybrid mode allows **independent scaling** per modality (pay-as-you-grow) - -=== 7.2. Maintenance Burden - -**Problem:** Single admin team must understand **six different data stores**. - -**Skill Requirements:** -- Oxigraph (RDF/SPARQL) -- HNSW (vector indexing algorithms) -- ndarray/Burn (tensor operations, ML model serving) -- CBOR (semantic proofs, schema validation) -- Tantivy (full-text search, inverted indices) -- Merkle trees (temporal versioning, cryptographic proofs) - -**Hiring Difficulty:** Finding engineers with all six skill sets is rare/expensive - -**Mitigation:** -- **Specialization:** Split team into modality-specific sub-teams (requires larger org) -- **Documentation:** Invest heavily in runbooks and operational guides -- **Managed Service:** Consider offering VeriSimDB-as-a-Service (centralize expertise) - -== 8. Summary of Challenges - -[cols="1,2,1"] -|=== -|Challenge Category |Key Issues |Severity - -|**Architecture** -|Vertical scaling limits, SPOF, tight coupling -|HIGH - -|**Performance** -|Resource interference, write amplification, routing bottleneck -|HIGH - -|**Operations** -|Backup complexity, monitoring overhead, dependency hell -|MEDIUM - -|**Scaling** -|No horizontal scaling, hot modality bottlenecks, write limits -|HIGH - -|**Security** -|Single security boundary, audit log SPOF -|MEDIUM - -|**Migration** -|Complex path to federation/hybrid -|MEDIUM - -|**Cost** -|Over-provisioning, specialized maintenance burden -|MEDIUM -|=== - -== 9. When to Choose Standalone Despite Challenges - -Standalone deployment is appropriate when: - -1. **Data Volume:** < 1TB total across all modalities -2. **Query Rate:** < 1k qps sustained -3. **Team Size:** < 5 engineers (federated coordination overhead not justified) -4. **Latency Requirements:** Sub-millisecond response times critical -5. **Simplicity Premium:** Operational simplicity outweighs scaling concerns - -For larger deployments, consider **hybrid** or **federated** mode to mitigate these challenges. - -== See Also - -- link:deployment-modes.adoc[Deployment Modes Overview] -- link:challenges-hybrid.adoc[Challenges: Hybrid Mode] -- link:challenges-federated.adoc[Challenges: Federated Mode] -- link:../WHITEPAPER.md[VeriSimDB White Paper] diff --git a/verisimdb/docs/consultation-dependent-types-zkp.adoc b/verisimdb/docs/consultation-dependent-types-zkp.adoc deleted file mode 100644 index 0a338780..00000000 --- a/verisimdb/docs/consultation-dependent-types-zkp.adoc +++ /dev/null @@ -1,1166 +0,0 @@ -= Consultation Paper: Dependent Type System with Zero-Knowledge Proofs -:author: VeriSimDB Design Team -:date: 2026-01-22 -:toc: left -:toclevels: 4 -:sectnums: -:source-highlighter: rouge - -// SPDX-License-Identifier: CC-BY-SA-4.0 - -[abstract] -== Abstract - -This consultation paper explores the **second most challenging technical problem in VeriSimDB**: integrating dependent types with zero-knowledge proofs to enable **formally verified queries with privacy preservation**. We analyze the dual execution paths (dependent-type PROOF vs slipstream), type-level contracts, soundness guarantees, and ZKP circuit generation. - -**Status**: Open for consultation + -**Decision Required By**: v1.0 implementation phase + -**Stakeholders**: Core team, type theorists, cryptographers, compliance officers - ---- - -== 1. Problem Statement - -=== 1.1 The Core Challenge - -VeriSimDB must support **two fundamentally different query execution models**: - -[cols="1,2,2,2",options="header"] -|=== -|Path |Type System |Execution |Output - -|**PROOF** + -(verified) -|Dependent types with refinements -|Type-checked, proof-generating -|`(Result, Proof[φ])` - -|**Slipstream** + -(unverified) -|Simple types (Int, String, UUID) -|Direct execution, no verification -|`Result` -|=== - -**The question**: How do we: - -1. **Type-check dependent types** in a language that compiles to WASM (ReScript)? -2. **Generate ZKP circuits** from type-level predicates? -3. **Verify proofs** without re-executing queries? -4. **Maintain soundness** (well-typed queries don't produce invalid proofs)? -5. **Optimize** (type erasure for slipstream path)? - -**Example query that exposes the challenge**: - -```vcl --- Dependent-type query: Prove that ALL results satisfy predicate -SELECT GRAPH, DOCUMENT -FROM verisim:semantic -WHERE octad.types INCLUDES "Paper" - AND VERIFIED(octad.document.peer_reviewed = true) -PROOF EXISTENCE { - ∀ r ∈ Result. r.document.peer_reviewed = true -} -LIMIT 100; - --- What ZKP circuit should be generated? --- How do we prove ∀ r without revealing r? --- How do we type-check VERIFIED()? -``` - -=== 1.2 Why This Matters - -**Wrong implementation consequences**: - -1. **Unsound type system**: Well-typed queries produce invalid proofs → compliance violations -2. **Performance disaster**: Type-checking adds 10× latency → no one uses PROOF path -3. **Privacy leak**: ZKP circuits reveal more than necessary → GDPR violations -4. **Complexity explosion**: Dependent types infect entire codebase → maintenance nightmare - -**Scale**: At 1000 queries/sec: -- Type-checking cost: ~10ms per query (dependent types) vs ~1ms (simple types) -- ZKP proof generation: ~100ms per query (circuit complexity) -- Proof verification: ~10ms per proof (by third party) - -This is the **technical credibility challenge**: Can we actually deliver on the promise of "formally verified queries"? - ---- - -== 2. Type System Deep Dive - -=== 2.1 Type Language - -==== 2.1.1 Base Types - -```ocaml -(* Simple types (slipstream path) *) -type simple_type = - | TUnit - | TBool - | TInt - | TFloat - | TString - | TUuid - | TTimestamp - | TVector of int (* Vector[n] *) - | TTensor of int list (* Tensor[d₁, d₂, ...] *) - | TOption of simple_type (* Option[T] *) - | TList of simple_type (* List[T] *) - | TRecord of (string * simple_type) list (* {field: Type, ...} *) -``` - -```ocaml -(* Dependent types (PROOF path) *) -type dependent_type = - | TSimple of simple_type - | TRefinement of { - base: dependent_type; - var: string; - predicate: expr; (* {x : τ | φ(x)} *) - } - | TPi of { (* Π x : τ₁. τ₂(x) - dependent function *) - var: string; - domain: dependent_type; - codomain: dependent_type; - } - | TSigma of { (* Σ x : τ₁. τ₂(x) - dependent pair *) - var: string; - first: dependent_type; - second: dependent_type; - } - | TProof of expr (* Proof[φ] - proof type *) - | TModalResult of { (* QueryResult[M, φ] *) - modalities: modality list; - contract: expr option; - } - -(* Predicates (refinement type expressions) *) -and expr = - | EVar of string - | EBool of bool - | EInt of int - | EString of string - | EBinOp of binop * expr * expr (* x < 10, x = "foo" *) - | EUnOp of unop * expr (* !x, -x *) - | EForall of string * dependent_type * expr (* ∀ x : τ. φ(x) *) - | EExists of string * dependent_type * expr (* ∃ x : τ. φ(x) *) - | EImplies of expr * expr (* φ₁ ⇒ φ₂ *) - | EAnd of expr * expr - | EOr of expr * expr - | EField of expr * string (* x.field *) - | EApp of expr * expr (* f(x) *) -``` - -==== 2.1.2 Example Types - -**Example 1: Positive integer** - -```ocaml -{x : Int | x > 0} -``` - -In VCL: -```vcl -VERIFIED(citation_count > 0) -- Type: {n : Int | n > 0} -``` - -**Example 2: Non-empty list** - -```ocaml -{xs : List[String] | length(xs) > 0} -``` - -In VCL: -```vcl -VERIFIED(authors != []) -- Type: {as : List[String] | length(as) > 0} -``` - -**Example 3: Proved query result** - -```ocaml -Σ r : QueryResult[GRAPH, DOCUMENT]. - Proof[∀ h ∈ r. h.document.peer_reviewed = true] -``` - -In VCL: -```vcl -PROOF EXISTENCE { - ∀ r ∈ Result. r.document.peer_reviewed = true -} --- Type: ProvedResult[GRAPH, DOCUMENT] -``` - -=== 2.2 Typing Rules - -==== 2.2.1 Simple Type Rules (Slipstream) - -**Rule T-Var (Variable)**: -``` -Γ(x) = τ -───────────── -Γ ⊢ x : τ -``` - -**Rule T-Query (Simple query)**: -``` -Γ ⊢ modalities : M -Γ ⊢ source : Source -Γ ⊢ condition : Octad → Bool -──────────────────────────────────────── -Γ ⊢ SELECT M FROM source WHERE condition : QueryResult[M] -``` - -==== 2.2.2 Dependent Type Rules (PROOF path) - -**Rule T-Refinement (Refinement type intro)**: -``` -Γ ⊢ e : τ -Γ ⊢ φ[e/x] : Bool -φ[e/x] = true -──────────────────────────────────────── -Γ ⊢ e : {x : τ | φ(x)} -``` - -**Rule T-ProvedQuery (Query with proof)**: -``` -Γ ⊢ modalities : M -Γ ⊢ source : Source -Γ ⊢ condition : Octad → Bool -Γ ⊢ proof-spec : ProofSpec[φ] -Γ ⊢ φ : QueryResult[M] → Bool -──────────────────────────────────────── -Γ ⊢ SELECT M ... PROOF proof-spec : ProvedResult[M, φ] - -Where: -ProvedResult[M, φ] = Σ r : QueryResult[M]. Proof[φ(r)] -``` - -**Rule T-Verified (VERIFIED predicate)**: -``` -Γ ⊢ expr : Bool -Γ ⊢ expr verifiable (* Can be checked via ZKP *) -──────────────────────────────────────── -Γ ⊢ VERIFIED(expr) : {b : Bool | b = true} -``` - -=== 2.3 Bidirectional Type Checking - -**Problem**: Type inference for dependent types is **undecidable** in general. - -**Solution**: Bidirectional type checking (synthesis + checking modes): - -```ocaml -(* Synthesis mode: infer type from expression *) -val synthesize : context -> expr -> dependent_type option - -(* Checking mode: check expression against expected type *) -val check : context -> expr -> dependent_type -> bool -``` - -**Algorithm**: - -```ocaml -let rec synthesize ctx = function - | EVar x -> - (* Look up variable in context *) - Context.lookup ctx x - - | EInt n -> - (* Integer literal synthesizes to Int *) - Some TInt - - | EBinOp (Lt, e1, e2) -> - (* Both operands must be Int *) - let* t1 = synthesize ctx e1 in - let* t2 = synthesize ctx e2 in - if equal_type t1 TInt && equal_type t2 TInt then - Some TBool - else - None - - | EVERIFIED expr -> - (* VERIFIED(φ) synthesizes to refinement type *) - let* _ = check ctx expr TBool in - if is_verifiable expr then - Some (TRefinement { - base = TBool; - var = "b"; - predicate = expr; - }) - else - None - -let rec check ctx expr expected_type = - match synthesize ctx expr with - | Some actual_type -> - is_subtype actual_type expected_type - | None -> - false -``` - -**Example**: - -```vcl --- Query -SELECT GRAPH -FROM verisim:semantic -WHERE octad.graph.citation_count > 10 - AND VERIFIED(octad.graph.peer_reviewed = true) - --- Type checking steps: --- 1. Synthesize: octad.graph.citation_count → Int --- 2. Check: citation_count > 10 → Bool --- 3. Synthesize: VERIFIED(...) → {b : Bool | b = true} --- 4. Check: overall condition → {b : Bool | b = true} --- 5. Result type: QueryResult[GRAPH] -``` - -=== 2.4 Subtyping - -**Rule Sub-Refine (Refinement subtyping)**: -``` -φ₁(x) ⇒ φ₂(x) (* Implication *) -──────────────────────────────────────── -{x : τ | φ₁(x)} <: {x : τ | φ₂(x)} -``` - -**Example**: -```ocaml -{n : Int | n > 10} <: {n : Int | n > 0} - -(* Because: n > 10 implies n > 0 *) -``` - -**Implication checking** (via SMT solver): - -```ocaml -let implies (phi1 : expr) (phi2 : expr) : bool = - (* Use Z3 SMT solver to check: φ₁ ⇒ φ₂ *) - let ctx = Z3.mk_context [] in - let solver = Z3.Solver.mk_solver ctx None in - - (* Encode φ₁ and φ₂ as Z3 expressions *) - let z3_phi1 = encode_expr ctx phi1 in - let z3_phi2 = encode_expr ctx phi2 in - - (* Check: φ₁ ∧ ¬φ₂ unsatisfiable? *) - Z3.Solver.add solver [z3_phi1; Z3.Boolean.mk_not ctx z3_phi2]; - - match Z3.Solver.check solver [] with - | Z3.Solver.UNSATISFIABLE -> true (* φ₁ ⇒ φ₂ *) - | _ -> false -``` - ---- - -== 3. Zero-Knowledge Proof Integration - -=== 3.1 Proof Contracts - -**Four types of proofs** in VCL: - -[cols="1,2,2",options="header"] -|=== -|Contract Type |Predicate |Use Case - -|**EXISTENCE** |`∃ r ∈ Result. φ(r)` |Prove at least one result satisfies φ - -|**UNIVERSAL** |`∀ r ∈ Result. φ(r)` |Prove all results satisfy φ - -|**ACCESS** |`∀ r ∈ Result. authorized(user, r)` |Prove user authorized for all results - -|**INTEGRITY** |`∀ r ∈ Result. hash(r) = expected` |Prove data integrity (no tampering) -|=== - -**VCL examples**: - -```vcl --- EXISTENCE proof: At least one peer-reviewed paper exists -SELECT * -FROM verisim:semantic -WHERE octad.types INCLUDES "Paper" -PROOF EXISTENCE { - ∃ r ∈ Result. r.document.peer_reviewed = true -}; - --- UNIVERSAL proof: All results are peer-reviewed -SELECT * -FROM verisim:semantic -WHERE octad.types INCLUDES "Paper" -PROOF UNIVERSAL { - ∀ r ∈ Result. r.document.peer_reviewed = true -}; - --- ACCESS proof: User authorized for all results -SELECT * -FROM verisim:restricted -WHERE octad.types INCLUDES "MedicalRecord" -PROOF ACCESS { - ∀ r ∈ Result. has_permission(user, r, "read") -}; -``` - -=== 3.2 ZKP Circuit Generation - -**Problem**: How to convert type-level predicate to ZKP circuit? - -**Architecture**: - -``` -┌─────────────────┐ -│ VCL Query │ -│ (with PROOF) │ -└────────┬────────┘ - │ - v -┌─────────────────┐ -│ Type Checker │ ← Synthesize proof obligation -└────────┬────────┘ - │ - v -┌─────────────────┐ -│ Circuit Builder │ ← Convert predicate to arithmetic circuit -└────────┬────────┘ - │ - v -┌─────────────────┐ -│ proven Library │ ← Generate SNARK proof -└────────┬────────┘ - │ - v -┌─────────────────┐ -│ (Result, Proof) │ -└─────────────────┘ -``` - -**Circuit builder pseudo-code**: - -```rust -pub struct CircuitBuilder { - constraints: Vec, - variables: HashMap, -} - -impl CircuitBuilder { - pub fn build_from_predicate(&mut self, pred: Expr) -> Result { - match pred { - // ∀ r ∈ Result. φ(r) - Expr::Forall(var, ty, body) => { - // For each result in query result set: - // Add constraint: φ(r) = true - let circuit_var = self.allocate_variable(&var, ty)?; - let body_circuit = self.build_from_predicate(body)?; - self.add_constraint(Constraint::ForAll(circuit_var, body_circuit)); - Ok(self.finalize()?) - } - - // r.field = value - Expr::BinOp(Eq, lhs, rhs) => { - let lhs_var = self.build_from_predicate(lhs)?; - let rhs_var = self.build_from_predicate(rhs)?; - self.add_constraint(Constraint::Eq(lhs_var, rhs_var)); - Ok(self.finalize()?) - } - - // r.field > value - Expr::BinOp(Gt, lhs, rhs) => { - let lhs_var = self.build_from_predicate(lhs)?; - let rhs_var = self.build_from_predicate(rhs)?; - self.add_constraint(Constraint::Gt(lhs_var, rhs_var)); - Ok(self.finalize()?) - } - - _ => Err(CircuitError::UnsupportedPredicate(pred)), - } - } -} -``` - -**Example circuit**: - -```vcl -PROOF UNIVERSAL { - ∀ r ∈ Result. r.document.peer_reviewed = true -} - --- Generated circuit (pseudocode): --- Input: result_set = [r₁, r₂, ..., rₙ] --- Witness: peer_reviewed = [true, true, ..., true] --- Constraint: ∀ i. peer_reviewed[i] = 1 (boolean true) --- Output: proof π that all results are peer-reviewed -``` - -**Circuit complexity**: -- **Universal (∀)**: O(n) constraints (n = result set size) -- **Existence (∃)**: O(1) constraints (prove one witness) -- **Equality (=)**: O(1) constraint per equality -- **Comparison (>, <)**: O(log k) constraints (k = bit width) - -=== 3.3 Proof Generation - -**Using proven library** (Rust): - -```rust -use proven::{Circuit, Proof, ProvingKey, VerifyingKey}; - -pub async fn generate_proof( - query: &VCLQuery, - result: &QueryResult, - contract: &ProofContract, -) -> Result { - // 1. Build circuit from proof contract - let circuit = CircuitBuilder::new() - .build_from_predicate(contract.predicate)?; - - // 2. Generate proving key (one-time setup per circuit) - let proving_key = ProvingKey::generate(&circuit)?; - - // 3. Compute witness (private inputs) - let witness = compute_witness(result, contract)?; - - // 4. Generate SNARK proof - let proof = proven::prove( - &proving_key, - &circuit, - &witness, - )?; - - Ok(proof) -} - -fn compute_witness( - result: &QueryResult, - contract: &ProofContract, -) -> Result { - let mut witness = Witness::new(); - - match contract.contract_type { - ContractType::Universal => { - // For ∀ r ∈ Result. φ(r): - // Witness = [φ(r₁), φ(r₂), ..., φ(rₙ)] - for octad in &result.octads { - let satisfied = evaluate_predicate(&contract.predicate, octad)?; - witness.add_boolean(satisfied); - } - } - - ContractType::Existence => { - // For ∃ r ∈ Result. φ(r): - // Witness = (index i, φ(rᵢ) = true) - let index = result.octads.iter() - .position(|h| evaluate_predicate(&contract.predicate, h).unwrap_or(false)) - .ok_or(WitnessError::NoExistentialWitness)?; - witness.add_integer(index as i64); - } - - _ => return Err(WitnessError::UnsupportedContractType), - } - - Ok(witness) -} -``` - -=== 3.4 Proof Verification - -**Verifier side** (anyone can verify without re-executing query): - -```rust -pub fn verify_proof( - proof: &Proof, - contract: &ProofContract, - verifying_key: &VerifyingKey, -) -> Result { - // Check proof validity - let valid = proven::verify( - &verifying_key, - &proof, - )?; - - if !valid { - return Ok(false); - } - - // Check contract type matches - if proof.contract_type != contract.contract_type { - return Err(VerificationError::ContractMismatch); - } - - Ok(true) -} -``` - -**Key properties**: - -1. **Soundness**: Valid proof ⇒ predicate actually true (cryptographic assumption) -2. **Zero-knowledge**: Proof reveals nothing except "predicate satisfied" -3. **Succinctness**: Proof size O(1) regardless of result set size -4. **Fast verification**: O(1) time regardless of query complexity - ---- - -== 4. Soundness & Safety - -=== 4.1 Type Safety Theorems - -**Theorem 1 (Progress)**: If `Γ ⊢ e : τ` and `e` is closed, then either: -- `e` is a value, or -- `e → e'` for some `e'` - -**Proof sketch**: By induction on typing derivation. - -**Theorem 2 (Preservation)**: If `Γ ⊢ e : τ` and `e → e'`, then `Γ ⊢ e' : τ`. - -**Proof sketch**: By induction on evaluation derivation. - -**Corollary (Type Safety)**: Well-typed expressions don't get stuck. - -**Theorem 3 (Proof Soundness)**: If `Γ ⊢ query : ProvedResult[M, φ]`, then: -- Executing `query` produces `(result, proof)` -- `verify(proof, φ) = true ⇒ φ(result) = true` - -**Proof sketch**: -1. Type checking ensures proof obligation φ is well-formed -2. Circuit builder correctly encodes φ as arithmetic constraints -3. proven library soundness: valid proof ⇒ constraints satisfied -4. Constraints satisfied ⇒ φ(result) = true - -=== 4.2 Attack Scenarios - -**Attack 1: Malicious query (invalid proof claim)** - -```vcl --- Attacker tries to prove false statement -SELECT * -FROM verisim:semantic -WHERE octad.types INCLUDES "Paper" -PROOF UNIVERSAL { - ∀ r ∈ Result. r.document.h_index > 100 -- FALSE for many results -}; -``` - -**Defense**: Proof generation fails (cannot construct valid witness) - -```rust -// In generate_proof(): -let satisfied = evaluate_predicate(&contract.predicate, octad)?; -if !satisfied { - return Err(ProofError::PredicateViolation { - octad_id: octad.id, - predicate: contract.predicate.clone(), - }); -} -``` - -**Attack 2: Proof forgery** - -Attacker tries to construct fake proof without executing query. - -**Defense**: Cryptographic soundness of SNARK (computational assumption) - -**Attack 3: Type confusion** - -```vcl --- Attacker tries to confuse type checker -SELECT GRAPH -FROM verisim:semantic -WHERE VERIFIED(octad.graph.node_count > "foo") -- Type error -PROOF EXISTENCE { ... }; -``` - -**Defense**: Type checker rejects query before execution - -```ocaml -(* In type_check_condition: *) -let t1 = synthesize ctx (EField (EVar "octad", "graph", "node_count")) in -let t2 = synthesize ctx (EString "foo") in -if not (equal_type t1 TInt && equal_type t2 TInt) then - raise (TypeError "Cannot compare Int with String") -``` - ---- - -== 5. Performance Analysis - -=== 5.1 Type Checking Cost - -**Measurement** (on 100 sample queries): - -[cols="1,1,1,1",options="header"] -|=== -|Query Type |Type Check Time |Circuit Build Time |Proof Gen Time - -|**Simple** (slipstream) |1ms |N/A |N/A - -|**Refinement** (VERIFIED) |8ms |5ms |N/A - -|**PROOF EXISTENCE** |12ms |15ms |80ms - -|**PROOF UNIVERSAL** |15ms |50ms |300ms (n=100 results) -|=== - -**Bottlenecks**: - -1. **SMT solver** (implication checking): 5-10ms per subtyping check -2. **Circuit building** (universal): O(n) constraints, 0.5ms per result -3. **Proof generation**: 100-500ms (depends on circuit size) - -**Optimization strategies**: - -```rust -// 1. Cache type-checking results -static TYPE_CACHE: Lazy>> = ...; - -pub fn type_check_cached(query: &VCLQuery) -> Result { - let hash = compute_hash(query); - - if let Some(ty) = TYPE_CACHE.lock().unwrap().get(&hash) { - return Ok(ty.clone()); // Cache hit: 0.1ms - } - - let ty = type_check_uncached(query)?; - TYPE_CACHE.lock().unwrap().insert(hash, ty.clone()); - Ok(ty) -} - -// 2. Incremental type checking (only check changed parts) -pub fn type_check_incremental( - query: &VCLQuery, - prev_query: &VCLQuery, - prev_type: &DependentType, -) -> Result { - // Diff queries, only re-check changed AST nodes - let diff = diff_queries(query, prev_query); - if diff.is_empty() { - return Ok(prev_type.clone()); // No changes: 0ms - } - // ... check only diff ... -} - -// 3. Parallel proof generation (batch queries) -pub async fn generate_proofs_batch( - queries: Vec<(VCLQuery, QueryResult)>, -) -> Result, ProofError> { - use tokio::task::spawn; - - let handles: Vec<_> = queries.into_iter() - .map(|(query, result)| spawn(async move { - generate_proof(&query, &result, query.proof_contract.as_ref().unwrap()).await - })) - .collect(); - - // Wait for all proofs in parallel - let proofs = futures::future::join_all(handles).await; - // ... -} -``` - -**Results** (after optimization): - -[cols="1,1,1",options="header"] -|=== -|Query Type |Before |After - -|**Simple** |1ms |1ms (no change) - -|**PROOF EXISTENCE** |107ms |35ms (3× faster) - -|**PROOF UNIVERSAL** |365ms |120ms (3× faster) -|=== - -=== 5.2 Type Erasure - -**Idea**: Erase dependent types to simple types after type-checking. - -```ocaml -let rec erase_type = function - | TSimple t -> t - | TRefinement { base; _ } -> erase_type base (* Drop predicate *) - | TPi { codomain; _ } -> erase_type codomain (* Drop dependency *) - | TSigma { first; second; _ } -> - TRecord [("fst", erase_type first); ("snd", erase_type second)] - | TProof _ -> TUnit (* Proofs erased to unit *) - | TModalResult { modalities; _ } -> TQueryResult modalities -``` - -**Example**: - -```ocaml -(* Before erasure *) -{x : Int | x > 0} - -(* After erasure *) -Int -``` - -**Benefit**: Slipstream execution uses erased types (faster, no proof generation). - -```rust -pub fn execute_query_erased(query: &VCLQuery) -> Result { - // Type check with dependent types - let full_type = type_check(query)?; - - // Erase to simple types - let erased_type = erase_type(full_type); - - // Execute query with erased types (no proof generation) - let result = execute_with_simple_types(query, erased_type)?; - - Ok(result) -} -``` - ---- - -== 6. Implementation Challenges - -=== 6.1 Challenge 1: ReScript Type Checker - -**Problem**: ReScript doesn't natively support dependent types. - -**Options**: - -1. **Implement type checker in ReScript** (manually) - - Pros: Type-safe, compiles to WASM, portable - - Cons: No SMT solver in WASM, must call out to native - -2. **Implement type checker in Rust** (via FFI) - - Pros: Z3 bindings available, fast - - Cons: Must cross FFI boundary, not portable to WASM - -3. **Implement type checker in OCaml** (separate service) - - Pros: Strong type system heritage, SMT bindings - - Cons: Another runtime, another language - -**Recommendation**: **Option 2 (Rust with Z3)** - -```rust -// Rust type checker with Z3 -use z3::{Config, Context, Solver}; - -pub struct TypeChecker { - z3_ctx: Context, - solver: Solver<'static>, - type_ctx: TypeContext, -} - -// Expose to ReScript via FFI -#[no_mangle] -pub extern "C" fn type_check_query( - query_json: *const c_char, -) -> *mut c_char { - let query_str = unsafe { CStr::from_ptr(query_json).to_str().unwrap() }; - let query: VCLQuery = serde_json::from_str(query_str).unwrap(); - - let type_checker = TypeChecker::new(); - match type_checker.check(&query) { - Ok(ty) => { - let result = serde_json::to_string(&ty).unwrap(); - CString::new(result).unwrap().into_raw() - } - Err(e) => { - let error = format!("{{\"error\": \"{}\"}}", e); - CString::new(error).unwrap().into_raw() - } - } -} -``` - -**ReScript bindings**: - -```rescript -// VCLTypeChecker.res -@module("./native/type_checker.node") -external typeCheckQuery: string => string = "type_check_query" - -let checkQuery = (query: VCLQuery.t): result => { - let queryJson = query->VCLQuery.toJson->Js.Json.stringify - let resultJson = typeCheckQuery(queryJson) - - switch resultJson->Js.Json.parseExn { - | exception _ => Error(TypeError.ParseFailed) - | json => - switch json->Js.Json.decodeObject { - | Some(obj) if obj->Js.Dict.get("error")->Belt.Option.isSome => - Error(TypeError.fromJson(json)) - | Some(_) => Ok(DependentType.fromJson(json)) - | None => Error(TypeError.InvalidResponse) - } - } -} -``` - -=== 6.2 Challenge 2: Circuit Complexity - -**Problem**: Universal quantification over large result sets creates huge circuits. - -**Example**: - -```vcl --- 1000 results, each with 10-field record -PROOF UNIVERSAL { - ∀ r ∈ Result. r.document.word_count > 1000 -} - --- Circuit constraints: 1000 × 10 = 10,000 constraints --- Proof generation time: ~5 seconds -``` - -**Solutions**: - -1. **Merkle tree proofs** (batch results) -``` -Instead of proving φ(r₁) ∧ φ(r₂) ∧ ... ∧ φ(r₁₀₀₀), -prove Merkle root of [φ(r₁), φ(r₂), ..., φ(r₁₀₀₀)] -``` - -2. **Recursive proofs** (compose smaller proofs) -``` -Prove φ(r₁...r₅₀₀) → π₁ -Prove φ(r₅₀₁...r₁₀₀₀) → π₂ -Prove "π₁ valid ∧ π₂ valid" → π_final -``` - -3. **Sampling proofs** (probabilistic guarantees) -``` -Instead of proving ∀ r ∈ Result. φ(r), -prove φ(sample(Result, 100)) with high probability -``` - -**Tradeoffs**: - -[cols="1,2,2,2",options="header"] -|=== -|Approach |Proof Size |Proof Time |Security - -|**Naive** (prove all) |O(n) |O(n) |Perfect - -|**Merkle tree** |O(log n) |O(n) |Perfect - -|**Recursive** |O(1) |O(n log n) |Perfect - -|**Sampling** |O(1) |O(k) (k=sample size) |Probabilistic -|=== - -=== 6.3 Challenge 3: Verifiable Predicates - -**Problem**: Not all predicates are efficiently verifiable via ZKP. - -**Verifiable** (arithmetic circuits): -- Equality: `x = y` -- Comparison: `x < y`, `x > y` -- Boolean logic: `φ ∧ ψ`, `φ ∨ ψ` -- Arithmetic: `x + y`, `x × y` - -**Not efficiently verifiable**: -- String operations: `substring(s, 0, 10) = "Machine"` -- Cryptographic hashes: `sha256(x) = y` (expensive) -- Modular arithmetic: `x mod p = 0` -- Floating-point: `x / y > 0.5` (requires fixed-point encoding) - -**Strategy**: Restrict `VERIFIED()` to verifiable predicates only. - -```rescript -// VCLTypeChecker.res -let isVerifiable = (expr: VCLExpr.t): bool => { - switch expr { - | EBinOp(Eq, _, _) => true - | EBinOp(Lt | Gt | Lte | Gte, _, _) => true - | EBinOp(And | Or, e1, e2) => isVerifiable(e1) && isVerifiable(e2) - | EUnOp(Not, e) => isVerifiable(e) - | EField(_, _) => true - | EInt(_) | EBool(_) => true - - // Non-verifiable predicates - | EFuncCall("substring", _) => false - | EFuncCall("sha256", _) => false - | _ => false - } -} -``` - ---- - -== 7. Open Questions & Consultation - -=== 7.1 Critical Questions - -1. **Type checker language**: Rust (with Z3) or OCaml (better type theory support)? - -2. **Circuit optimization**: Naive, Merkle tree, recursive, or sampling? - -3. **Verifiable predicate restrictions**: Should we support string operations (expensive) or only arithmetic? - -4. **Proof caching**: Should proofs be cached? For how long? (Proofs become invalid if data changes) - -5. **Type erasure boundary**: When to erase types (after type-check or after proof generation)? - -=== 7.2 Consultation Questions - -==== For Type Theorists: - -1. Is our dependent type system sound? (Any counterexamples?) -2. Should we support liquid types (automatic predicate inference) in v3? -3. Is bidirectional type checking the right approach, or should we use constraint solving? - -==== For Cryptographers: - -1. Is SNARK the right proof system, or should we use STARK (transparent, post-quantum)? -2. Should we support zk-STARK proofs for specific contracts (e.g., Merkle tree)? -3. What's the security parameter for proven library? (128-bit, 256-bit?) - -==== For Compliance Officers: - -1. Do ZKP proofs satisfy GDPR/HIPAA requirements for data minimization? -2. What audit trail is needed for proof verification? (Who verified? When?) -3. Should proofs be stored long-term? (For compliance audits) - ---- - -== 8. Recommendation & Next Steps - -=== 8.1 Recommended Architecture - -**⭐ RUST TYPE CHECKER WITH Z3 + MERKLE TREE PROOFS ⭐** - -[cols="1,2,2",options="header"] -|=== -|Component |Implementation |Rationale - -|**Type Checker** |Rust (with z3-sys bindings) |Fast, SMT solver available, portable - -|**Circuit Builder** |Rust (proven library) |SNARK generation, type-safe - -|**Proof Optimization** |Merkle tree (for n > 100) |O(log n) proof size, perfect security - -|**Verifiable Predicates** |Arithmetic only (v1) + -String ops (v2, expensive) |Start simple, add complexity later - -|**Type Erasure** |After type-check, before exec |Slipstream fast, PROOF verified -|=== - -=== 8.2 Implementation Roadmap - -**v1.0 (Dependent types + basic ZKP)**: -1. Rust type checker with simple types + refinement types -2. EXISTENCE and UNIVERSAL proof contracts -3. Naive circuit generation (no Merkle optimization) -4. Type erasure for slipstream path - -**v2.0 (Optimized circuits)**: -1. Merkle tree proof optimization -2. String operation support (via SHA-256 circuit) -3. Proof caching (5-minute TTL) - -**v3.0 (Advanced types)**: -1. Liquid types (automatic refinement inference) -2. Effect types (track side effects: read, write, network) -3. Recursive proofs (for very large result sets) - -=== 8.3 Success Metrics - -| Metric | Target | Measurement | -|--------|--------|-------------| -| Type check latency | <15ms (p99) | Time from query to type | -| Proof gen latency | <200ms (p99) | Time from result to proof | -| Proof verification | <10ms | Time to verify proof | -| Proof size | <1 KB (for n=100) | Bytes per proof | -| Type checker memory | <50 MB | RSS after 1000 queries | -| Soundness | 100% | Valid proof ⇒ predicate true | - ---- - -== 9. References - -1. Dependent types: Benjamin Pierce, "Types and Programming Languages" -2. Refinement types: Ranjit Jhala, "Liquid Types" -3. ZK-SNARKs: Zcash Sapling protocol: https://z.cash/technology/zksnarks/ -4. proven library: https://github.com/proven-network/proven -5. Z3 SMT solver: https://github.com/Z3Prover/z3 -6. Bidirectional type checking: Pierce & Turner, "Local Type Inference" - ---- - -== Appendix A: Full Type System - -=== A.1 Syntax - -```ocaml -(* Types *) -τ ::= Int | Bool | String | Uuid | Unit - | τ₁ → τ₂ (* Function *) - | τ₁ × τ₂ (* Product *) - | τ₁ + τ₂ (* Sum *) - | {x : τ | φ(x)} (* Refinement *) - | Π x : τ₁. τ₂(x) (* Dependent function *) - | Σ x : τ₁. τ₂(x) (* Dependent pair *) - | Proof[φ] (* Proof type *) - -(* Expressions *) -e ::= x | n | true | false | () - | e₁ e₂ (* Application *) - | λx. e (* Abstraction *) - | (e₁, e₂) (* Pair *) - | e.1 | e.2 (* Projection *) - | inl e | inr e (* Injection *) - | case e of inl x => e₁ | inr y => e₂ (* Case *) - -(* Predicates *) -φ ::= true | false - | e₁ = e₂ | e₁ < e₂ | e₁ > e₂ - | φ₁ ∧ φ₂ | φ₁ ∨ φ₂ | ¬φ - | ∀ x : τ. φ(x) | ∃ x : τ. φ(x) - | φ₁ ⇒ φ₂ -``` - -=== A.2 Typing Rules (Complete) - -See docs/vcl-type-system.adoc for full 30+ typing rules. - ---- - -== Appendix B: Example Proofs - -=== B.1 EXISTENCE Proof - -**Query**: -```vcl -SELECT * -FROM verisim:semantic -WHERE octad.types INCLUDES "Paper" -PROOF EXISTENCE { - ∃ r ∈ Result. r.document.peer_reviewed = true -}; -``` - -**Generated circuit** (pseudocode): -``` -Input: result_ids = [id₁, id₂, ..., idₙ] -Witness: (index = 5, peer_reviewed[5] = true) -Constraint: peer_reviewed[index] = 1 -Output: proof π -``` - -**Proof size**: ~256 bytes (independent of n) - -=== B.2 UNIVERSAL Proof - -**Query**: -```vcl -SELECT * -FROM verisim:restricted -WHERE octad.types INCLUDES "MedicalRecord" -PROOF UNIVERSAL { - ∀ r ∈ Result. has_permission(user_id, r, "read") -}; -``` - -**Generated circuit** (pseudocode): -``` -Input: result_ids = [id₁, id₂, ..., idₙ] - user_id = "user:alice" -Witness: permissions = [1, 1, 1, ..., 1] (* All true *) -Constraint: ∀ i ∈ [1..n]. permissions[i] = 1 -Output: proof π -``` - -**Proof size**: ~512 bytes + O(log n) for Merkle tree diff --git a/verisimdb/docs/consultation-normalization-strategy.adoc b/verisimdb/docs/consultation-normalization-strategy.adoc deleted file mode 100644 index 20a22029..00000000 --- a/verisimdb/docs/consultation-normalization-strategy.adoc +++ /dev/null @@ -1,920 +0,0 @@ -= Consultation Paper: Normalization Cascade Strategy -:author: VeriSimDB Design Team -:date: 2026-01-22 -:toc: left -:toclevels: 4 -:sectnums: -:source-highlighter: rouge - -// SPDX-License-Identifier: CC-BY-SA-4.0 - -[abstract] -== Abstract - -This consultation paper explores the **most challenging architectural decision in VeriSimDB**: how to manage normalization across five levels of the system (L0: intra-modality → L4: cross-lineage) while balancing consistency, performance, and federation scalability. We present deep technical analysis of push vs pull strategies, Byzantine fault tolerance, and adaptive threshold learning. - -**Status**: Open for consultation + -**Decision Required By**: v1.0 implementation phase + -**Stakeholders**: Core team, federated store operators, compliance officers - ---- - -== 1. Problem Statement - -=== 1.1 The Core Challenge - -VeriSimDB's multimodal federation creates **five distinct levels of normalization**, each with different consistency requirements: - -[cols="1,2,2,2",options="header"] -|=== -|Level |Scope |Example Drift |Consistency Requirement - -|**L0** |Intra-modality + -(within single store) -|Graph edge exists but vector embedding missing -|**Strong** - should not occur - -|**L1** |Cross-modality + -(across modalities in octad) -|Title mismatch: graph says "ML Paper" vs document says "Machine Learning Paper" -|**Eventually consistent** - cosmetic - -|**L2** |Cross-octad + -(between related octads) -|Author octad links to paper octad that doesn't exist yet -|**Eventually consistent** - temporal ordering - -|**L3** |Cross-store + -(federated stores) -|University A has version 1.2, University B has version 1.1 -|**Quorum-based** - Byzantine tolerance needed - -|**L4** |Cross-lineage + -(external systems) -|DOI resolver returns different metadata than our cache -|**Best-effort** - we don't control external systems -|=== - -**The question**: For each level, should normalization be: - -- **Pushed** (proactive propagation: detect drift → notify all affected parties → repair immediately)? -- **Pulled** (reactive repair: detect drift on query → repair locally → cache result)? -- **Hybrid** (push critical issues, pull optimizations)? - -=== 1.2 Why This Matters - -**Wrong choice consequences**: - -1. **Too much push**: Thundering herd problem, network storms, federated stores overwhelmed with repair messages -2. **Too much pull**: Stale data, compliance violations (GDPR right-to-forget not propagated), users see inconsistent results -3. **No strategy**: Drift accumulates until system unusable - -**Scale**: At 100 federated stores with 1M octads each: -- Push strategy: ~100 billion messages/day (if 1% daily drift) -- Pull strategy: ~1 billion query-time repairs/day (if 10 queries/octad/day) - -This is not a theoretical problem—it's the **make-or-break architectural decision** for federation scalability. - ---- - -== 2. Deep Technical Analysis - -=== 2.1 Push Strategy (Eventual Consistency Model) - -==== 2.1.1 How It Works - -When drift detected at any level: - -``` -┌─────────────┐ -│ Detect Drift│ -└──────┬──────┘ - │ - v -┌─────────────────────┐ -│ Classify Drift Type │ -└──────┬──────────────┘ - │ - v -┌─────────────────────────┐ -│ Broadcast Repair Message│ ← To ALL affected stores -└──────┬──────────────────┘ - │ - v -┌─────────────────────────┐ -│ Wait for ACK (optional) │ -└──────┬──────────────────┘ - │ - v -┌─────────────────────────┐ -│ Mark as Consistent │ -└─────────────────────────┘ -``` - -**Elixir GenServer pseudo-code**: - -```elixir -defmodule VeriSim.DriftMonitor do - use GenServer - - @doc "Detected drift at L1 (cross-modality)" - def handle_cast({:drift_detected, :l1_cross_modal, drift_details}, state) do - # Classify drift severity - severity = classify_drift(drift_details) - - case severity do - :critical -> - # Push immediately to all affected octads - affected_octads = find_affected_octads(drift_details) - Enum.each(affected_octads, fn octad_id -> - VeriSim.Octad.repair(octad_id, drift_details) - end) - - # Broadcast to federation if cross-store impact - if affects_federation?(drift_details) do - VeriSim.Federation.broadcast_repair(drift_details) - end - - :optimization -> - # Queue for batch processing (don't push) - {:noreply, enqueue_for_batch(state, drift_details)} - - :cosmetic -> - # Log only, no action - Logger.info("Cosmetic drift detected: #{inspect(drift_details)}") - {:noreply, state} - end - end -end -``` - -==== 2.1.2 Advantages - -1. **Strong consistency guarantees**: All stores converge to same state eventually -2. **Compliance-friendly**: GDPR "right to forget" propagates immediately -3. **Predictable behavior**: Users see consistent data across federation -4. **Audit trail**: Push messages create immutable log of all repairs - -==== 2.1.3 Disadvantages - -1. **Network overhead**: Every drift generates N messages (N = number of affected stores) -2. **Latency**: Mutations block until repair messages sent (unless async) -3. **Thundering herd**: If one store detects drift, all stores receive repair messages simultaneously -4. **Byzantine attacks**: Malicious store can flood federation with fake repair messages -5. **Cascading failures**: If repair fails at one store, should transaction abort everywhere? - -==== 2.1.4 Real-World Measurements - -From Kafka's migration to KRaft (similar problem domain): - -- **Metadata propagation latency**: 50ms median, 500ms p99 for 10-node cluster -- **Network amplification**: 1 metadata change → 10× messages (gossip protocol) -- **Failure mode**: Controller node failure causes 5-10 second unavailability - -**Extrapolated to VeriSimDB**: -- At 100 federated stores, 1% daily drift rate (1M octads): - - **10,000 drift events/day** - - **1,000,000 repair messages/day** (100× amplification) - - **~12 messages/second** sustained load - - **Peak load**: 1000+ messages/second (if drift batches correlate) - -=== 2.2 Pull Strategy (Lazy Consistency Model) - -==== 2.2.1 How It Works - -``` -┌──────────┐ -│ Query │ -└────┬─────┘ - │ - v -┌─────────────────┐ -│ Detect Drift? │ ← During query execution -└────┬────────────┘ - │ Yes - v -┌─────────────────┐ -│ Repair Locally │ ← Only for this query -└────┬────────────┘ - │ - v -┌─────────────────┐ -│ Cache Result │ ← TTL-based cache -└────┬────────────┘ - │ - v -┌─────────────────┐ -│ Return to User │ -└─────────────────┘ -``` - -**VCL query-time repair**: - -```vcl --- User's query -SELECT GRAPH, DOCUMENT -FROM verisim:semantic -WHERE octad.types INCLUDES "Paper" -LIMIT 100; - --- VeriSimDB internally detects: graph title ≠ document title --- Applies repair strategy: LATEST_WINS --- Returns repaired results to user --- Caches repair for TTL duration (adaptive: 5 min → 1 hour) -``` - -**Rust store pseudo-code**: - -```rust -impl OctadStore { - pub async fn query_with_repair( - &self, - query: &VCLQuery, - repair_policy: RepairPolicy, - ) -> Result { - // Execute query - let raw_results = self.execute_query(query).await?; - - // Detect drift in results - let drift_detected = self.detect_drift(&raw_results)?; - - if !drift_detected.is_empty() { - // Repair on-the-fly - let repaired = match repair_policy { - RepairPolicy::LatestWins => self.repair_latest_wins(raw_results, drift_detected), - RepairPolicy::Quorum => self.repair_quorum(raw_results, drift_detected).await?, - RepairPolicy::Manual => return Err(VCLError::ManualRepairRequired(drift_detected)), - }; - - // Cache repaired results - self.cache.insert(query.cache_key(), repaired.clone(), adaptive_ttl()); - - Ok(repaired) - } else { - Ok(raw_results) - } - } -} -``` - -==== 2.2.2 Advantages - -1. **Low network overhead**: No proactive propagation, repair only when needed -2. **No thundering herd**: Repairs happen independently per query -3. **Graceful degradation**: If federation unavailable, local store still works -4. **Adaptive**: Can increase cache TTL if drift rare, decrease if frequent -5. **Byzantine-resistant**: Malicious store can't flood network (no broadcast) - -==== 2.2.3 Disadvantages - -1. **Stale data**: Users may see inconsistent results until query triggers repair -2. **Query latency**: First query after drift pays repair cost (subsequent queries cached) -3. **Compliance risk**: GDPR "right to forget" not immediate (depends on query frequency) -4. **Cache invalidation complexity**: When to invalidate? TTL-based risks stale cache -5. **Repair inconsistency**: Two simultaneous queries might repair differently - -==== 2.2.4 Real-World Measurements - -From DNS caching (similar problem domain): - -- **Cache hit rate**: 80-95% for popular domains (TTL: 5 min - 24 hours) -- **Stale data window**: TTL duration (median: 1 hour) -- **Repair latency**: First query pays full cost, subsequent queries ~1ms (cached) - -**Extrapolated to VeriSimDB**: -- At 100 federated stores, 1% daily drift rate (1M octads), 10 queries/octad/day: - - **100,000 drift instances/day** - - **1,000,000 query-time repairs/day** (if 10 queries hit each drift) - - **Cache hit rate**: 90% (if TTL = 1 hour, query distribution even) - - **Actual repairs**: 100,000/day (10% cache miss) - - **Network messages**: 100,000/day (vs 1,000,000 for push strategy) - -=== 2.3 Hybrid Strategy (Recommended) - -==== 2.3.1 Core Principle - -**Push critical safety issues, pull optimization issues.** - -``` -┌─────────────┐ -│ Detect Drift│ -└──────┬──────┘ - │ - v -┌─────────────────┐ -│ Classify Drift │ -└──────┬──────────┘ - │ - ├─ Critical? ──> PUSH (retraction, integrity, access) - │ - └─ Optimization? ──> PULL (title mismatch, cache) -``` - -**Classification table**: - -[cols="1,1,2,2",options="header"] -|=== -|Drift Type |Severity |Strategy |Rationale - -|**Retraction** + -(paper retracted) -|Critical -|**PUSH** immediately -|Compliance: must propagate retraction to all cites - -|**Integrity violation** + -(hash mismatch) -|Critical -|**PUSH** immediately -|Security: data corruption must be flagged - -|**Access revocation** + -(permission removed) -|Critical -|**PUSH** immediately -|Privacy: unauthorized access must stop immediately - -|**Schema version** + -(v1.1 vs v1.2) -|High -|**ASYNC PUSH** (batched) -|Federation: stores should converge but not urgent - -|**Title mismatch** + -("ML Paper" vs "Machine Learning Paper") -|Low -|**PULL** on query -|Cosmetic: user sees consistent result, no propagation needed - -|**Derived data stale** + -(citation count cached) -|Low -|**PULL** on query -|Optimization: recompute on demand - -|**External sync** + -(DOI metadata changed) -|Advisory -|**PULL** periodically (cron) -|Best-effort: we don't control external system -|=== - -==== 2.3.2 Implementation Strategy - -**Elixir adaptive thresholds**: - -```elixir -defmodule VeriSim.AdaptiveLearner do - @moduledoc """ - Learn optimal push/pull thresholds based on observed metrics. - """ - - def adjust_thresholds(state) do - # Observe metrics - push_frequency = Metrics.get(:push_messages_per_hour) - pull_latency = Metrics.get(:pull_repair_latency_p99) - drift_frequency = Metrics.get(:drift_events_per_hour) - - # Adjust push threshold (when does "high severity" become "critical"?) - new_push_threshold = cond do - push_frequency > 1000 -> - # Too many push messages, raise threshold (push less) - state.push_threshold + 0.1 - - pull_latency > 500 -> - # Pull repairs too slow, lower threshold (push more) - state.push_threshold - 0.1 - - true -> - state.push_threshold - end - - # Adjust cache TTL (how long to cache pull repairs?) - new_cache_ttl = cond do - drift_frequency < 10 -> - # Drift rare, increase TTL (cache longer) - min(state.cache_ttl * 1.5, 3600) - - drift_frequency > 100 -> - # Drift frequent, decrease TTL (cache shorter) - max(state.cache_ttl * 0.7, 60) - - true -> - state.cache_ttl - end - - %{state | - push_threshold: clamp(new_push_threshold, 0.0, 1.0), - cache_ttl: new_cache_ttl - } - end -end -``` - -==== 2.3.3 Advantages - -1. **Best of both worlds**: Safety of push, scalability of pull -2. **Adaptive**: System learns optimal thresholds from workload -3. **Predictable critical path**: Compliance issues (retraction, access) always pushed -4. **Graceful degradation**: If push fails, fall back to pull - -==== 2.3.4 Disadvantages - -1. **Implementation complexity**: Must classify every drift type -2. **Classification errors**: Misclassifying critical as optimization = compliance violation -3. **Threshold tuning**: Initial thresholds may be wrong, takes time to converge -4. **Hybrid confusion**: Users/operators must understand two consistency models - ---- - -== 3. Byzantine Fault Tolerance - -=== 3.1 Threat Model - -In federated VeriSimDB, we must assume **Byzantine faults**: - -1. **Malicious store**: Intentionally sends fake drift messages ("paper retracted" when it's not) -2. **Compromised store**: Attacker controls a store, floods network with repair messages -3. **Network partition**: Store can't reach quorum, must decide: abort or continue? - -**Not in threat model** (out of scope): -- Nation-state attacks -- Compromise of all stores simultaneously -- Side-channel attacks (timing, power analysis) - -=== 3.2 Quorum-Based Consensus - -**For L3 (cross-store) normalization**, use quorum voting: - -``` -┌─────────────────────┐ -│ Store A detects │ -│ drift: title changed│ -└──────┬──────────────┘ - │ - v -┌────────────────────────────┐ -│ Query federation: │ -│ "What is correct title?" │ -└──────┬─────────────────────┘ - │ - v -┌────────────────────────────┐ -│ Collect responses: │ -│ Store A: "ML Paper" │ -│ Store B: "ML Paper" │ -│ Store C: "Machine Learning"│ -│ Store D: "ML Paper" │ -│ Store E: [no response] │ -└──────┬─────────────────────┘ - │ - v -┌────────────────────────────┐ -│ Quorum (3/5): "ML Paper" │ -│ → Repair to quorum value │ -└────────────────────────────┘ -``` - -**Quorum size**: `2f + 1` for `f` Byzantine faults - -- **f=1** (tolerate 1 malicious store): Quorum = 3 -- **f=2** (tolerate 2 malicious stores): Quorum = 5 - -**Rust implementation**: - -```rust -pub struct FederationQuorum { - stores: Vec, - timeout: Duration, -} - -impl FederationQuorum { - pub async fn repair_with_quorum( - &self, - octad_id: OctadId, - field: &str, - ) -> Result { - // Query all stores in parallel - let responses = self.query_all_stores(octad_id, field).await?; - - // Count responses - let mut counts: HashMap = HashMap::new(); - for response in responses { - *counts.entry(response.value).or_insert(0) += 1; - } - - // Find quorum value (> 50% of stores) - let total_stores = self.stores.len(); - let quorum_size = (total_stores / 2) + 1; - - for (value, count) in counts { - if count >= quorum_size { - return Ok(value); - } - } - - // No quorum reached - Err(QuorumError::NoConsensus { - octad_id, - field: field.to_string(), - responses, - }) - } -} -``` - -=== 3.3 Signature Verification - -**All repair messages must be signed** (using sactify-php): - -```elixir -defmodule VeriSim.Federation.RepairMessage do - defstruct [ - :octad_id, - :drift_type, - :repair_action, - :timestamp, - :source_store_id, - :signature # Ed25519 signature from source store - ] - - def verify(message) do - # Get source store's public key from registry - public_key = VeriSim.Registry.get_store_pubkey(message.source_store_id) - - # Verify signature - payload = encode_payload(message) - case :crypto.verify(:eddsa, :sha512, payload, message.signature, [public_key, :ed25519]) do - true -> {:ok, message} - false -> {:error, :invalid_signature} - end - end -end -``` - -**Attack mitigation**: -- **Fake repair messages**: Rejected (invalid signature) -- **Replay attacks**: Timestamp checked, messages expire after 5 minutes -- **Store impersonation**: Public keys registered in immutable registry - ---- - -== 4. Performance Modeling - -=== 4.1 Push Strategy Cost Model - -Variables: -- `N` = number of federated stores -- `H` = number of octads -- `D` = daily drift rate (as fraction, e.g., 0.01 = 1%) -- `M` = message size (bytes) - -**Daily network cost**: -``` -Messages_per_day = H × D × N -Bytes_per_day = Messages_per_day × M -``` - -**Example (conservative)**: -- N = 100 stores -- H = 1,000,000 octads -- D = 0.01 (1% daily drift) -- M = 1024 bytes (1 KB message) - -``` -Messages_per_day = 1,000,000 × 0.01 × 100 = 1,000,000 -Bytes_per_day = 1,000,000 × 1024 = 1 GB/day -Network_cost = $0.01/GB × 1 GB × 100 stores = $1/day -``` - -**Scaling**: -- At N = 1,000 stores: $10/day = $3,650/year -- At N = 10,000 stores: $100/day = $36,500/year - -**Break-even analysis**: -- Pull strategy cost: Query latency × number of queries -- Push strategy cost: Network bandwidth × number of stores -- Break-even when query latency savings > network costs - -=== 4.2 Pull Strategy Cost Model - -Variables: -- `Q` = queries per octad per day -- `L_repair` = repair latency (ms) -- `L_cache` = cache hit latency (ms) -- `C` = cache hit rate (as fraction, e.g., 0.9 = 90%) - -**Daily query latency cost**: -``` -Repairs_per_day = H × D × Q -Cache_hits = Repairs_per_day × C -Cache_misses = Repairs_per_day × (1 - C) - -Latency_cost = Cache_misses × L_repair + Cache_hits × L_cache -``` - -**Example (conservative)**: -- H = 1,000,000 octads -- D = 0.01 (1% daily drift) -- Q = 10 queries/octad/day -- L_repair = 100 ms -- L_cache = 1 ms -- C = 0.9 (90% cache hit) - -``` -Repairs_per_day = 1,000,000 × 0.01 × 10 = 100,000 -Cache_hits = 100,000 × 0.9 = 90,000 -Cache_misses = 100,000 × 0.1 = 10,000 - -Latency_cost = 10,000 × 100ms + 90,000 × 1ms = 1,090,000 ms = ~18 minutes total -``` - -**User experience**: -- 10,000 users experience 100ms latency (cache miss) -- 90,000 users experience 1ms latency (cache hit) -- Average latency: 10.9ms per query - -=== 4.3 Hybrid Strategy Cost Model - -**Assumptions**: -- 5% of drift is critical (pushed) -- 95% of drift is optimization (pulled) - -**Push component**: -``` -Push_messages = H × D × 0.05 × N = 50,000 messages/day -Push_bytes = 50,000 × 1024 = 51 MB/day -Push_cost = $0.01/GB × 0.051 GB × 100 stores = $0.05/day -``` - -**Pull component**: -``` -Pull_repairs = H × D × 0.95 × Q = 95,000 repairs/day -Pull_latency = 9,500 × 100ms + 85,500 × 1ms = 1,035,500 ms = ~17 minutes -``` - -**Total cost**: $0.05/day + 17 minutes latency (vs $1/day push-only or 18 minutes pull-only) - -**Conclusion**: Hybrid strategy reduces network cost by 95% vs push-only, with only 5% increase in latency vs pull-only. - ---- - -== 5. Open Questions & Consultation Points - -=== 5.1 Critical Questions - -1. **What constitutes "critical" drift?** - - Current classification: retraction, integrity violation, access revocation - - Should schema version mismatches be critical? (impacts type safety) - - Should citation count mismatches be critical? (impacts research metrics) - -2. **What quorum size for Byzantine tolerance?** - - f=1 (quorum=3): Fast, but only tolerates 1 malicious store - - f=2 (quorum=5): Slower, but tolerates 2 malicious stores - - f=3 (quorum=7): Very slow, but very robust - -3. **What cache TTL for pulled repairs?** - - Short TTL (5 min): Fresh data, high query load - - Long TTL (1 hour): Stale data, low query load - - Adaptive TTL: Starts at 5 min, increases to 1 hour if hit rate > 90% - -4. **Should L4 (external systems) be in scope at all?** - - Option A: External sync is advisory only (current recommendation) - - Option B: Treat external systems as untrusted federation members (push critical issues) - - Option C: No external sync (users must manually import) - -=== 5.2 Consultation Questions for Stakeholders - -==== For Core Team: - -1. Is the classification of critical vs optimization drift correct? Are we missing any critical types? -2. Should adaptive learning (miniKanren in v3) be allowed to change push/pull classification, or must it be static? -3. What happens if quorum fails (network partition)? Abort query or return stale data with warning? - -==== For Federated Store Operators: - -1. What network bandwidth constraints do you have? (Determines max push frequency) -2. What compliance requirements do you have? (GDPR, HIPAA, FAIR) (Determines minimum push coverage) -3. Are you willing to participate in quorum voting? (Latency cost for your queries) - -==== For Compliance Officers: - -1. Is eventual consistency acceptable for non-critical drift? (e.g., title mismatches) -2. What maximum delay is acceptable for critical drift propagation? (GDPR: 72 hours, our target: 5 seconds) -3. Must all repair actions be auditable? (Impacts whether pull repairs must be logged) - ---- - -== 6. Recommendation & Next Steps - -=== 6.1 Recommended Strategy - -**⭐ HYBRID PUSH/PULL WITH ADAPTIVE THRESHOLDS ⭐** - -[cols="1,1,2",options="header"] -|=== -|Level |Strategy |Implementation - -|**L0** (Intra-modality) -|PULL -|Store-internal consistency on insert - -|**L1** (Cross-modality) -|HYBRID -|Push: retractions, integrity violations + -Pull: title mismatches, derived data - -|**L2** (Cross-octad) -|PULL -|Validate at query time, cache results - -|**L3** (Cross-store) -|HYBRID -|Push: retractions, access revocation (signed, quorum-verified) + -Pull: schema version, optimization - -|**L4** (External) -|ADVISORY -|Periodic sync (cron), best-effort only -|=== - -**Rationale**: -1. **Scales**: 95% reduction in network cost vs push-only -2. **Safe**: Critical issues (compliance, security) always pushed -3. **Fast**: 95% of repairs cached, <10ms latency -4. **Adaptive**: System learns optimal thresholds from workload - -=== 6.2 Implementation Roadmap - -**v1.0 (Elixir feedback loops)**: -1. Static classification (hardcoded critical types) -2. Fixed cache TTL (15 minutes) -3. Quorum size f=1 (tolerates 1 Byzantine fault) -4. Signature verification (sactify-php) - -**v2.0 (Adaptive thresholds)**: -1. Elixir GenServer learns cache TTL from hit rate -2. Elixir GenServer adjusts push threshold from network metrics -3. Metrics dashboard (Grafana) for operators - -**v3.0 (miniKanren rule synthesis)**: -1. Learn classification rules from error examples -2. Synthesize new repair strategies -3. Optimize quorum selection (which stores to query first?) - -=== 6.3 Success Metrics - -| Metric | Target | Measurement | -|--------|--------|-------------| -| Network cost | <$10/day for 100 stores | Bytes sent per day | -| Query latency (cache hit) | <10ms | p99 latency | -| Query latency (cache miss) | <100ms | p99 latency | -| Critical drift propagation | <5 seconds | Time from detection to all stores repaired | -| Cache hit rate | >90% | Hits / (Hits + Misses) | -| Byzantine tolerance | Tolerate 1 malicious store | Quorum voting success rate | - ---- - -== 7. References - -1. Apache Kafka KRaft design: https://cwiki.apache.org/confluence/display/KAFKA/KIP-500 -2. Raft consensus algorithm: https://raft.github.io/ -3. Byzantine Paxos: Castro & Liskov, "Practical Byzantine Fault Tolerance" -4. DNS caching best practices: RFC 2181 -5. GDPR compliance timelines: Article 12(3), 72-hour breach notification - ---- - -== Appendix A: Classification Examples - -=== A.1 Critical Drift (Push Immediately) - -```elixir -# Example 1: Paper retracted -%DriftEvent{ - type: :retraction, - severity: :critical, - octad_id: "octad:550e8400-e29b-41d4-a716-446655440000", - field: "semantic.status", - old_value: "published", - new_value: "retracted", - reason: "Fabricated data, journal retraction notice", - detected_at: ~U[2026-01-22 14:30:00Z], - action: :push_immediately -} - -# Push to all stores that cite this paper -# Push to all octads that reference this paper -# Log in immutable temporal store -``` - -```elixir -# Example 2: Hash mismatch (integrity violation) -%DriftEvent{ - type: :integrity_violation, - severity: :critical, - octad_id: "octad:6ba7b810-9dad-11d1-80b4-00c04fd430c8", - field: "document.content_hash", - old_value: "sha256:abc123...", - new_value: "sha256:def456...", - reason: "Document content changed but hash not updated", - detected_at: ~U[2026-01-22 14:35:00Z], - action: :push_immediately -} - -# Push to all federation members -# Quarantine octad until resolved -``` - -=== A.2 Optimization Drift (Pull on Query) - -```elixir -# Example 3: Title mismatch (cosmetic) -%DriftEvent{ - type: :title_mismatch, - severity: :low, - octad_id: "octad:7c9e6679-7425-40de-944b-e07fc1f90ae7", - field: "graph.title vs document.title", - old_value: "ML Paper", - new_value: "Machine Learning Paper", - reason: "Different abbreviation styles", - detected_at: ~U[2026-01-22 14:40:00Z], - action: :pull_on_query -} - -# Don't push (cosmetic difference) -# Repair when user queries this octad -# Cache repaired result for 1 hour -``` - -```elixir -# Example 4: Stale derived data -%DriftEvent{ - type: :derived_stale, - severity: :low, - octad_id: "octad:8d8a8b4c-5d3e-4f5a-9b6c-7e8f9a0b1c2d", - field: "semantic.citation_count", - old_value: 42, - new_value: 45, - reason: "New citations added, cached count not updated", - detected_at: ~U[2026-01-22 14:45:00Z], - action: :pull_on_query -} - -# Don't push (expensive to recompute for all octads) -# Recompute when user queries citation count -# Cache for 1 day (citations change slowly) -``` - ---- - -== Appendix B: Quorum Voting Examples - -=== B.1 Successful Quorum (3/5 agreement) - -``` -Store responses for octad:550e8400, field "title": - -Store A: "Machine Learning Paper" (200ms latency) -Store B: "Machine Learning Paper" (150ms latency) -Store C: "ML Paper" (180ms latency) -Store D: "Machine Learning Paper" (220ms latency) -Store E: [timeout] (1000ms timeout) - -Quorum calculation: - "Machine Learning Paper": 3 votes (A, B, D) - "ML Paper": 1 vote (C) - No response: 1 (E) - -Quorum size: 5 stores × 0.5 + 1 = 3 votes required -Result: QUORUM REACHED → "Machine Learning Paper" -Action: Repair local store to quorum value -``` - -=== B.2 Failed Quorum (no majority) - -``` -Store responses for octad:6ba7b810, field "status": - -Store A: "published" (200ms latency) -Store B: "retracted" (150ms latency) -Store C: "published" (180ms latency) -Store D: "retracted" (220ms latency) -Store E: [network partition] - -Quorum calculation: - "published": 2 votes (A, C) - "retracted": 2 votes (B, D) - No response: 1 (E) - -Quorum size: 5 stores × 0.5 + 1 = 3 votes required -Result: NO QUORUM (tie) -Action: MANUAL INTERVENTION REQUIRED -``` - -**Resolution strategies for failed quorum**: - -1. **Retry with more stores**: Query 10 stores instead of 5 -2. **Increase timeout**: Maybe Store E would respond with 2000ms timeout -3. **Manual override**: Operator investigates and sets correct value -4. **Pessimistic default**: Assume most restrictive value ("retracted" safer than "published") diff --git a/verisimdb/docs/deployment-modes.adoc b/verisimdb/docs/deployment-modes.adoc deleted file mode 100644 index 986e9fe4..00000000 --- a/verisimdb/docs/deployment-modes.adoc +++ /dev/null @@ -1,260 +0,0 @@ -// SPDX-License-Identifier: CC-BY-SA-4.0 - -= VeriSimDB Deployment Modes - -VeriSimDB is **both** a database and a federation coordinator—the architecture supports multiple deployment modes to suit different use cases. - -== Three Deployment Modes - -=== Mode 1: Standalone Database (Traditional) - -Deploy VeriSimDB as a **complete database system** with all six modalities local. - -[cols="1,3"] -|=== -|Component |Function - -|Registry -|Local in-memory namespace - -|Modality Stores -|All 6 modalities running on the same machine/cluster - -|Use Case -|Single institution, traditional database replacement - -|Pros -|Simple deployment, no federation overhead, full ACID guarantees - -|Cons -|No cross-institutional sharing, scaling limited to vertical/sharding -|=== - -[source,bash] ----- -# Start all-in-one VeriSimDB node -cargo run -p verisim-api # Rust stores -iex -S mix # Elixir orchestration ----- - -**Example:** University research lab wants a multimodal database for their internal projects. - -=== Mode 2: Federated Coordinator (Tiny Core) - -Deploy only the **ReScript registry + Elixir orchestrator** while modality stores run elsewhere. - -[cols="1,3"] -|=== -|Component |Function - -|Registry -|Central namespace, KRaft quorum for consensus - -|Modality Stores -|External: University A (graph), Lab B (vector), Archive C (document) - -|Use Case -|Cross-institutional federation, data sovereignty - -|Pros -|Minimal footprint (<5k LOC), institutions keep data local, drift detection across boundaries - -|Cons -|Network latency, requires federation protocol adoption -|=== - -[source,bash] ----- -# Deploy only coordination layer -cd registry && npm run build # ReScript → WASM -cd ../elixir-orchestration && iex -S mix - -# Remote stores register via API: -curl -X POST http://registry:8080/api/v1/stores \ - -d '{"store_id": "...", "endpoint": "https://uni-a.edu/verisim", ...}' ----- - -**Example:** Open Science consortium where each university keeps their data but shares namespace. - -=== Mode 3: Hybrid (Recommended for Most Users) - -Run **some modalities locally** and **federate others**. - -[cols="1,3"] -|=== -|Component |Function - -|Registry -|Local or quorum - -|Local Modalities -|Graph, Vector (performance-critical) - -|Federated Modalities -|Document (archive.org), Temporal (institution B) - -|Use Case -|Performance + federation balance - -|Pros -|Hot data local, cold data federated, flexible scaling - -|Cons -|More complex configuration -|=== - -[source,bash] ----- -# Start local stores for graph + vector -cargo run -p verisim-api --features graph,vector - -# Configure remote federation for document + temporal -# via orchestrator config ----- - -**Example:** AI research lab needs fast vector search locally but shares documentation with broader community. - -== Is VeriSimDB a Database? - -**Yes**, VeriSimDB is a database in all three modes: - -1. **Standalone mode** → Traditional multimodal database -2. **Federated mode** → Distributed database with controlled consistency -3. **Hybrid mode** → Partially local, partially federated database - -The key architectural innovation is that it **also** enables federation without forcing you into that model. - -== Comparison to Traditional Architectures - -[cols="1,2,2,2"] -|=== -|Architecture |Data Location |Consistency Model |Best For - -|**PostgreSQL** -|Single instance or replicas -|Strong ACID -|Transactional apps - -|**Cassandra** -|Distributed replicas -|Eventual consistency -|High availability - -|**IPFS** -|Content-addressed blocks -|Immutable -|Static content - -|**VeriSimDB (Standalone)** -|All local -|Strong per-modality -|Multimodal apps - -|**VeriSimDB (Federated)** -|Across institutions -|Drift-tolerant -|Cross-institutional collaboration - -|**VeriSimDB (Hybrid)** -|Partially local/remote -|Configurable -|Performance + federation -|=== - -== Technical: How Mode Selection Works - -The deployment mode is determined by: - -1. **Registry Configuration** - Single node vs quorum cluster -2. **Store Registration** - Which modalities are local vs remote -3. **Orchestrator Config** - Drift detection scope and repair policies - -[source,elixir] ----- -# config/config.exs -config :verisim, - deployment_mode: :hybrid, - local_modalities: [:graph, :vector], - federated_modalities: [ - document: "https://archive.org/verisim", - temporal: "https://uni-b.edu/verisim" - ] ----- - -== Migration Paths - -=== Start Standalone → Federate Later - -1. Deploy Mode 1 (standalone) -2. As you grow, register external stores -3. Gradually migrate modalities to federation -4. Orchestrator handles transition transparently - -=== Start Federated → Add Local Stores - -1. Deploy Mode 2 (coordinator only) -2. Add local modality stores for performance -3. Configure hybrid routing rules -4. Hot data stays local, cold data federated - -== Operational Considerations - -[cols="1,2,2"] -|=== -|Aspect |Standalone |Federated - -|**Latency** -|Sub-millisecond -|Network-dependent (10-100ms typical) - -|**Availability** -|Single point of failure (unless clustered) -|Resilient (quorum consensus) - -|**Data Sovereignty** -|Full control -|Institutions keep ownership - -|**Scaling** -|Vertical + sharding -|Horizontal federation - -|**Complexity** -|Low (one system) -|Medium (coordination protocol) -|=== - -== Choosing Your Mode - -[cols="1,3"] -|=== -|Use Case |Recommended Mode - -|Small team, fast iteration -|**Standalone** - -|Cross-institutional research -|**Federated** - -|Performance-critical AI workloads with external docs -|**Hybrid** - -|Building a product -|**Standalone** (start), **Hybrid** (scale) - -|Open Science consortium -|**Federated** - -|Enterprise with multiple departments -|**Hybrid** (departments = federated stores) -|=== - -== Summary - -VeriSimDB's architecture supports three deployment modes: - -* **Standalone** - Traditional database, all modalities local -* **Federated** - Tiny coordinator, stores distributed across institutions -* **Hybrid** - Best of both worlds, performance + federation - -The "tiny core" refers to the **federated coordination layer**, but VeriSimDB **is a database** in all three modes. The choice depends on your use case, not the technology's capabilities. diff --git a/verisimdb/docs/design/DESIGN-2026-02-27-level-data-model.md b/verisimdb/docs/design/DESIGN-2026-02-27-level-data-model.md deleted file mode 100644 index 197c6ccb..00000000 --- a/verisimdb/docs/design/DESIGN-2026-02-27-level-data-model.md +++ /dev/null @@ -1,138 +0,0 @@ -# Design Document: IDApTIK Level Architect — Canonical Level Data Model -# SPDX-License-Identifier: CC-BY-SA-4.0 - -**Date:** 2026-02-27 -**Author:** Jonathan D.A. Jewell (hyperpolymath) -**Repo:** `idaptik/idaptik-level-architect` -**Session agent:** Claude Opus 4.6 - ---- - -## Summary - -Define the canonical level data model for the IDApTIK Level Architect in Idris2 -with dependent type proofs. This replaces the main game's `LevelConfig.res` as -the source of truth — ReScript types will derive from the Idris2 ABI definitions. - -## Motivation - -The level architect needs to: - -1. Represent every aspect of a level (devices, guards, dogs, drones, wiring, - zones, items, missions, physical grid, defence flags) -2. Prove correctness invariants at compile time (referential integrity of device - IPs, guard zone existence, zone spatial ordering, PBX consistency) -3. Export levels to `LevelConfig.res` format for the main game (round-trip - fidelity) -4. Store levels in VerisimDB for persistence, version history, and simulation - result tracking - -## Architectural Decisions - -### AD-1: Idris2 ABI as canonical source of truth - -**Decision:** All level types are defined in Idris2 with dependent type proofs. -The ReScript types in the main game and level architect derive from these. - -**Rationale:** Idris2's dependent types can prove invariants (referential -integrity, spatial ordering, positive values) that ReScript cannot express. A -broken level config is caught at the ABI layer before it reaches the game. - -### AD-2: VerisimDB for level architect only - -**Decision:** The level architect uses VerisimDB for: -- Level persistence (save/load work-in-progress designs) -- Version history (track edits, branch level variants, diff versions) -- Simulation results (guard patrol traces, detection events, timing data) -- Validation logs (Idris2 proof outcomes per level version) - -The main game does NOT use a database. It receives static exported -`LevelConfig.res` files. - -**Rationale:** The game is a single-player Tauri desktop app. It needs fast, -offline, read-only access to level data. A database adds runtime complexity -with no benefit. Save games, player progress, and inventory state are all -handled with local files. - -VerisimDB's federated mode provides an escape hatch: if a game-side database -is ever needed (e.g., level marketplace, multiplayer lobby), federation can -bridge to a lightweight secondary without redesigning the architect. - -### AD-3: Module-per-domain architecture - -**Decision:** 13 new Idris2 modules, one per game domain, plus a composition -root (`Level.idr`) and a validation module (`Validation.idr`). - -**Rationale:** -- Clean compilation units (change one domain, rebuild one module) -- Independent proof development per domain -- Clear 1:1 mapping to main game ReScript source files -- No circular dependencies (types flow upward from leaves to root) - -### AD-4: So-based witnesses over DecEq - -**Decision:** Use `So`-based witnesses via `decSo` on `(==)` for IPv4 -comparison in proofs, rather than full `DecEq` instances. - -**Rationale:** Full `DecEq` for `Bits8` requires proving `Not (a = b)` for the -negative case, which is awkward without `believe_me`. `So (a == b)` is -sufficient for our `InRegistry` and referential integrity proofs, and can be -constructed safely via `decSo` (which is in `Data.So` in the standard library). - -### AD-5: JSON serialization via Zig FFI - -**Decision:** Levels are serialized as JSON. The Zig FFI layer handles -parsing/emitting JSON. The Idris2 layer guarantees the parsed data is -well-formed. - -**Rationale:** `LevelConfig.res` is JSON-shaped (ReScript records compile to JS -objects). Round-trip verification is easiest with text format. Binary would add -complexity without benefit at this scale (levels are <100KB). - -## Module Dependency Graph - -``` - Level.idr - / | \ - / | \ - Mission Validation Physical - | / | \ | - Inventory / | \ Wiring - | / | \ | - Guards Dogs Drones Assassin - \ | / / - \ | / / - Devices - | - Network - | - Primitives - | - Types.idr (existing) -``` - -## Cross-Domain Proofs - -| Proof | Type | Guarantees | -|-------|------|-----------| -| `InRegistry` | `IPv4 -> DeviceRegistry -> Type` | An IP address exists in the device list | -| `GuardsInZones` | `List GuardPlacement -> List ZoneTransition -> Type` | Every guard's zone appears in zone transitions | -| `DefenceTargetsValid` | `List DeviceDefenceConfig -> DeviceRegistry -> Type` | failoverTarget/cascadeTrap/mirrorTarget IPs exist | -| `ZonesOrdered` | `List ZoneTransition -> Type` | Zone x-coordinates are monotonically increasing | -| `PBXConsistent` | `Bool -> Maybe IPv4 -> Type` | PBX IP is set iff hasPBX is True | -| `ValidatedLevel` | Record | Bundles LevelData with all proof terms | - -## Banned Patterns - -- `believe_me` — compiles silently, hides unsoundness -- `assert_total` — bypasses totality checker -- `assert_smaller` — bypasses termination checker -- `unsafePerformIO` — side effects in pure code - -## Verification Criteria - -1. `idris2 --build idaptik.ipkg` compiles with zero errors, zero warnings -2. Zero instances of banned patterns -3. Every field in `LevelConfig.res` has a corresponding field in `LevelConfig` -4. `ValidatedLevel` bundles all 5 proof types -5. All functions under `%default total` diff --git a/verisimdb/docs/design/DESIGN-2026-02-27-strategic-improvements.adoc b/verisimdb/docs/design/DESIGN-2026-02-27-strategic-improvements.adoc deleted file mode 100644 index 293bb29e..00000000 --- a/verisimdb/docs/design/DESIGN-2026-02-27-strategic-improvements.adoc +++ /dev/null @@ -1,750 +0,0 @@ -// SPDX-License-Identifier: CC-BY-SA-4.0 -// Design Document: VerisimDB Strategic Improvements -// Author: Jonathan D.A. Jewell -// Date: 2026-02-27 -// Repository: verisimdb - -= VerisimDB Strategic Improvements -Jonathan D.A. Jewell -2026-02-27 -:toc: left -:toclevels: 3 -:icons: font -:source-highlighter: rouge -:sectnums: - -== Purpose - -An honest assessment of where VerisimDB stands, what makes it genuinely novel, -what is weak, and what specific work would make it a database that people outside -hyperpolymath would actually want to use. - -This is not a roadmap with deadlines. It is a prioritised list of improvements -that can be picked up in any future session. - -== Honest Assessment - -=== What VerisimDB Has That Nothing Else Does - -No existing database -- Virtuoso, ArangoDB, SurrealDB, Tigris, Neo4j, none of -them -- tracks whether different representations of the same entity are -consistent with each other. - -VerisimDB's *drift detection* thesis: you store an entity, it gets a graph -representation, a document representation, a vector embedding, a temporal -snapshot, and VerisimDB _notices_ when those diverge and can reconcile them -automatically. This is genuinely novel. - -=== What Is Actually Strong - -[cols="1,3"] -|=== -| Area | Status - -| **Drift detection** -| Real implementation. `DriftMonitor` sweeps on schedule, `compute_modality_drift/3` - uses cosine distance between modality embeddings, Rust core computes per-entity - drift scores. - -| **Self-normalisation** -| Real implementation. Regeneration strategies inspect octad data, select the - authoritative modality, and derive drifted modality content from it. - -| **HNSW vector search** -| Real implementation (~670 lines Rust). Genuine approximate nearest-neighbour, - not brute-force. - -| **Tantivy full-text** -| Real implementation. Inverted index, ranked results, phrase queries. - -| **Temporal versioning** -| Real implementation. Native version history per octad, not audit table hacks. - -| **ETS caching** -| Real implementation. Read-through cache with TTL in the Elixir layer, write - invalidation. - -| **Federation** -| Real implementation. HTTP fanout to peers, query decomposition, result merging, - drift policy filtering. - -| **Benchmarks** -| All 6 modality stores benchmarked via Criterion. Baselines established. - -| **Zero believe_me** -| All Idris2 ABI code is clean. No unsafe casts. -|=== - -=== What Is Actually Weak - -[cols="1,1,3"] -|=== -| Problem | Severity | Detail - -| **VCL is incomplete** -| HIGH -| VCL parses and executes basic queries, but the language is not fully specified - and lacks documentation. Without a usable query language, VerisimDB is an HTTP - CRUD API with caching. _This is the single biggest gap._ - -| **VCL-UT not wired** -| MEDIUM -| Lean type definitions exist but the checker is not invoked at runtime. PROOF - clauses are syntactically supported but proof certificates are not formally - verified. See KNOWN-ISSUES #21. - -| **No production users** -| HIGH -| Only used by hyperpolymath repos. No adversarial workloads, no unexpected - query patterns, no real-world edge cases discovered. - -| **Modality stores are wrappers** -| LOW -| The stores wrap Oxigraph, Tantivy, HNSW, ndarray, CBOR. This is fine - architecturally -- the novel contribution is the octad coordination and drift - layer, not the individual stores. But it means the "8 modalities" pitch is - weaker than it sounds; the real pitch is drift detection. (The tensor - modality is an exception — active research is uncovering novel applications - with significant future potential.) - -| **Single developer** -| MEDIUM -| Bus factor of 1. No adversarial review of design decisions. - -| **oxrocksdb-sys build pain** -| MEDIUM -| Oxigraph pulls ~400K lines of C++ via RocksDB. 15-minute container builds, - requires clang-19. See KNOWN-ISSUES #24. - -| **No introspection by design** -| NONE -| This is correct -- no `/_introspect` endpoint, no `SHOW COLLECTIONS`. The - Elixir layer knows the schema from source code. Listed here to confirm it is - intentional, not an oversight. -|=== - -== Competitive Position - -=== VerisimDB vs Virtuoso - -[cols="<3,^2,^2",options="header"] -|=== -| Capability -| Virtuoso -| VerisimDB - -| **Core model** -| RDF triple store + relational SQL -| 8-modal octad (graph, vector, tensor, semantic, document, temporal, provenance, spatial) - -| **Graph** -| RDF graphs (SPARQL 1.1), billions of triples proven at scale -| Oxigraph RDF + property graph, smaller scale - -| **Relational / SQL** -| Full SQL support (ODBC/JDBC/ADO.NET) -| None -- not a goal - -| **Full-text search** -| Built-in (decent, older tech) -| Tantivy (Rust Lucene equivalent, modern) - -| **Vector / similarity search** -| Not built-in -| HNSW embeddings, native - -| **Tensor / numeric** -| Not built-in -| ndarray/Burn, native - -| **Semantic / proofs** -| RDF inference, OWL reasoning (mature) -| CBOR proof blobs (novel, less mature) - -| **Temporal / versioning** -| Graph versioning via named graphs (manual) -| Native temporal modality, automatic - -| **Drift detection** -| None -| Core feature -- cross-modal consistency - -| **Self-normalisation** -| None -| Core feature -- auto-repair diverged modalities - -| **Query language** -| SPARQL + SQL (both mature, standardised) -| VCL (custom, incomplete) - -| **Linked Data / LOD** -| First-class citizen -- LOD cloud backbone -| Not a focus - -| **SPARQL endpoint** -| Production-grade, federation support -| Via Oxigraph (basic) - -| **Wire protocols** -| HTTP (8890) + ODBC/JDBC (1111) -| HTTP only (single port) - -| **Maturity** -| ~25 years, powers DBpedia, LOD cloud -| ~1 year, powers hyperpolymath repos - -| **Community** -| Established (OpenLink, academic users) -| Solo developer - -| **Documentation** -| Extensive (if sometimes dated) -| Sparse (VCL docs incomplete) - -| **Operational complexity** -| Moderate-high (multi-port, config-heavy) -| High (6 stores + Rust + Elixir) - -| **License** -| GPLv2 (open source edition) / commercial -| PMPL-1.0-or-later - -| **Performance at scale** -| Proven (billions of triples, production workloads) -| Unproven at scale - -| **"Find things similar to X"** -| SPARQL text patterns only -| Vector embeddings -- semantic similarity - -| **"Show me X as it was last week"** -| Manual (named graph snapshots) -| Single temporal query - -| **"Are these modalities consistent?"** -| Not a concept -| Core concept (drift scores) -|=== - -=== Where VerisimDB Should NOT Compete - -Do not try to out-query ArangoDB at graph traversal or out-SPARQL Virtuoso at -linked data. They have had decades. VerisimDB is not a general-purpose -database and should not pretend to be one. - -=== Where VerisimDB Wins - -The moment you say "ArangoDB doesn't know that your graph and your document -disagree -- VerisimDB does" -- that is the differentiator. Nobody else is -playing in that space. - -VerisimDB is closer to "database + data observability" than "another database." -The real competitors are data quality tools like Great Expectations or Monte -Carlo, not database engines. - -== The Three Differentiators (Layered Strategy) - -VerisimDB has three genuinely novel capabilities. They serve different -audiences and should be presented in this order: - -=== Layer 1: Drift Detection (Door-Opener) - -The easiest to explain, the broadest audience. - -_"Your Postgres and your Elasticsearch disagree about the same customer. -VerisimDB detects that."_ - -Every enterprise data team has this problem. Nobody else solves it at the -entity level across heterogeneous systems. Data quality tools like Great -Expectations validate individual tables; VerisimDB validates cross-system -entity consistency. - -=== Layer 2: Heterogeneous Database Federation (Enterprise Value) - -The same customer exists in five different databases. VerisimDB sits above -all of them as a consistency layer, monitoring for drift without requiring -data to move into VerisimDB itself. - -_"VerisimDB watches your existing databases and tells you when they disagree. -No migration required."_ - -The current federation implementation works between VerisimDB peers (resolved -in KNOWN-ISSUES #2, #3, #19). The strategic evolution is federation over -heterogeneous databases — ArangoDB, Postgres, Elasticsearch, Neo4j — with -VerisimDB as the coordination layer. The IDApTIK database bridge (ArangoDB + -VerisimDB) is the first working example of this pattern. - -=== Layer 3: VCL-UT — Formally Verified Queries (Technical Moat) - -The deepest technical capability. Small audience (formal methods, regulated -industries, safety-critical systems) but impossible to replicate quickly. - -_"VerisimDB returns your query results with a machine-verifiable proof -certificate. The Lean type checker proves the result satisfies your -constraints."_ - -No database in existence does this. VCL-UT is the reason a competitor -cannot clone VerisimDB in a weekend after seeing the drift detection demo. -It takes years of type theory knowledge to even attempt it. - -**Who cares about VCL-UT:** - -- Regulated industries (finance, medical devices, aerospace) where you must - _prove_ a query result is correct, not just assert it -- Cross-organisational trust scenarios — verify results without trusting - the database operator -- Safety-critical systems (Ada/SPARK territory) -- Academic formal methods community - -**VCL-UT is not the lead pitch.** It is the moat. Drift detection gets -people in the door. Federation keeps them. VCL-UT ensures nobody can -compete. - -== Competitor Analysis: Data Quality / Observability Space - -VerisimDB overlaps more with data quality tools than with traditional databases. -Here is how it compares: - -[cols="<2,<3,<3",options="header"] -|=== -| Tool -| What It Does -| How It Differs from VerisimDB - -| **Great Expectations** -| Python library. Assertions on data ("expect this column to have values - between 0 and 100"). Validates data in pipelines. -| Single tables/dataframes. No cross-system consistency. No storage — it - is a testing framework, not a database. - -| **Monte Carlo** -| Commercial SaaS. ML-based anomaly detection across data warehouses. - Detects freshness issues, volume changes, schema drift. -| Watches one system (Snowflake/BigQuery/Redshift). Statistical anomalies, - not semantic inconsistencies between representations. - -| **Soda** -| Open-source data quality checks. YAML-based rules, integrates with - Airflow/dbt. -| Similar to Great Expectations with different UX. Single-system, - pipeline-level. - -| **Atlan** -| Data catalog + quality. Metadata management, lineage tracking. -| Catalog, not enforcement. Tells you _what_ data exists, not whether - representations agree. - -| **Bigeye** -| Automated data quality monitoring. Thresholds, alerts on anomalies. -| Statistical monitoring (row counts, null rates, distribution shifts). - No cross-system entity-level consistency. - -| **Anomalo** -| ML-based data quality. Learns "normal" patterns, alerts on deviations. -| Smart anomaly detection but still single-system, table-level. - -| **Elementary** -| dbt-native data observability. Monitors dbt model quality. -| Tied to dbt. If you don't use dbt, it doesn't exist for you. - -| **Datafold** -| Data diff tool. Compares table states before/after migrations. -| Diff, not continuous monitoring. Useful but narrow. -|=== - -**The gap none of them fill:** - -Every one of these tools operates on _one system at a time_ at the _table or -column level_. None of them ask: "is entity X in Postgres consistent with -entity X in Elasticsearch?" - -Great Expectations says "this column should have no nulls." VerisimDB says -"this entity's graph representation has 3 edges but its document says 4 -relationships — drift score 0.31." - -**The integration gap (honest):** - -All of those tools integrate with the existing ecosystem. -`pip install great-expectations`, point it at your Snowflake, done. -VerisimDB currently requires data to be stored _in_ VerisimDB. If federation -evolves to watch external databases (Layer 2 above), VerisimDB gets the -ecosystem integration of Great Expectations with the cross-system consistency -checking that none of them offer. - -== From Octad to Octad: Two New Modalities - -The current 6 modalities cover data representation well. Two additions -complete the model. - -=== The Octad - -[cols="1,2,1",options="header"] -|=== -| Modality | Question It Answers | Status - -| Graph -| "How is this related to that?" -| Implemented - -| Vector -| "What's similar to this?" -| Implemented - -| Tensor -| "What patterns exist in this high-dimensional data?" The tensor modality's - multi-dimensional representation capabilities open possibilities for novel - applications currently under active research. Early results suggest - compelling use cases that go well beyond traditional numeric storage — - details will be shared as the work matures. -| Implemented - -| Semantic -| "Is this valid? Prove it." -| Implemented - -| Document -| "Search for this text." -| Implemented - -| Temporal -| "What did this look like before?" -| Implemented - -| **Provenance** -| **"Where did this come from?"** -| Planned (CRITICAL) - -| **Spatial** -| **"Where is this? What's nearby?"** -| Planned -|=== - -=== Provenance / Lineage — 7th Modality (CRITICAL) - -Temporal tracks _versions_ (what changed). Provenance tracks _origins_ -(where it came from, how it was transformed, who touched it). These are -fundamentally different questions. - -==== Why This Is Critical - -Provenance / lineage compounds with federation to create something no other -tool offers: - -_"This entity was created in Postgres, transformed by a Spark pipeline, -cached in Redis, and indexed in Elasticsearch. Here is the full chain of -custody. Here is where the drift was introduced — step 3 of the pipeline -dropped a field."_ - -This is the difference between "your data is inconsistent" (drift detection -alone) and "your data is inconsistent _and here is exactly where it went -wrong_" (drift + provenance). - -==== Compliance and Regulatory Value - -[cols="1,3"] -|=== -| Regulation | What Provenance Enables - -| **GDPR Article 22** -| Right to explanation. Prove how a data-driven decision was made by - tracing the data lineage from source to query result. - -| **SOX / MiFID II** -| Financial audit trail. Prove the report's source data was not tampered - with by showing the full transformation chain. - -| **IEC 62304** -| Medical device data integrity. Trace every input to a clinical decision - back to its origin system. - -| **FAIR Principles** -| Provenance is required for Reusability (the R in FAIR). VerisimDB already - implements FAIR; provenance completes it. -|=== - -==== What the 7th Modality Stores - -Each entity gains a provenance record: - -- **Origin system** — which database or service created this entity -- **Transformation chain** — ordered list of pipeline steps that modified it -- **Actor trail** — which user, service account, or automation touched it -- **Timestamp chain** — when each transformation occurred -- **Integrity hash** — cryptographic hash at each step (detect tampering) -- **Federation source** — which federated peer contributed this data - -==== Implementation Approach - -The provenance modality is structurally simpler than graph or vector: - -- Storage: append-only log per entity (similar to temporal, but tracking - _origins_ not _versions_) -- Query: `SELECT octads WHERE provenance.origin = 'postgres'` or - `WHERE provenance.chain CONTAINS 'spark-etl-step-3'` -- Drift integration: when drift is detected, provenance can pinpoint - _which transformation step_ introduced the inconsistency -- VCL extension: `PROVENANCE` clause alongside existing `PROOF` clause - -==== Priority - -**CRITICAL.** Implement after VCL (Priority 1) and the drift demo -(Priority 2), but before or alongside heterogeneous federation (Priority 5). -Provenance and federation are mutually reinforcing — federation without -provenance tells you data disagrees; federation _with_ provenance tells you -why. - -=== Spatial / Geospatial — 8th Modality - -Spatial drift is a real consistency problem: "the map says this device is in -zone A but the database record says zone B." Every IoT, logistics, and -fleet management system has this. - -==== What the 8th Modality Stores - -- **Coordinates** — latitude/longitude, 3D points, projected coordinates -- **Geometries** — polygons, lines, multipoints (zone boundaries, patrol routes) -- **Spatial index** — R-tree for efficient proximity and containment queries -- **CRS metadata** — coordinate reference system (WGS84, UTM, etc.) - -==== Spatial Queries in VCL - -- `WHERE spatial.within(polygon)` — containment -- `WHERE spatial.distance(point) < 5000` — proximity (metres) -- `WHERE spatial.intersects(geometry)` — overlap - -==== Implementation - -Rust crates `geo` and `rstar` (R-tree) provide the core. Same architectural -slot as the other modality stores — `verisim-spatial` crate, HTTP API -endpoints, Elixir orchestration integration. - -==== Priority - -After provenance. Spatial is valuable but not thesis-critical. Provenance -compounds with drift detection and federation; spatial is additive. - -== Priority Improvements - -Ordered by impact. Each item is self-contained and can be picked up in any -future session. - -=== Priority 1: Finish VCL (HIGH) - -**Why:** A database without a usable query language is a key-value store with -extra steps. This is the single highest-impact improvement. - -**Scope:** Not VCL-UT. Just VCL. The query language that end users type. - -**What "finished" means:** - -1. **Language specification** -- a document (adoc or djot) that defines VCL - syntax and semantics exhaustively. Every valid query form documented with - examples. Every error case specified. - -2. **Cross-modal queries work end-to-end** -- a query like - `SELECT octads WHERE drift_score > 0.3 AND modality.temporal.version > 5` - must parse, plan, execute, and return correct results. - -3. **WHERE clause routing is correct** -- fulltext conditions go to Tantivy, - vector conditions go to HNSW, graph patterns go to Oxigraph, drift - conditions go to the drift engine. The resolved issues (#12, #13) fixed - the stubs; this is about ensuring coverage of all clause combinations. - -4. **Aggregates work** -- `COUNT`, `AVG`, `MIN`, `MAX`, `SUM` over octad - fields. `GROUP BY` modality or custom fields. - -5. **EXPLAIN is accurate** -- the resolved EXPLAIN (#18) produces real cost - estimates; verify they correlate with actual execution time. - -6. **Error messages are helpful** -- parse errors include line/column, - semantic errors explain what went wrong (e.g., "modality 'graphh' does - not exist, did you mean 'graph'?"). - -**Deliverables:** - -- `docs/VCL-SPEC.adoc` -- language specification -- Updated parser/executor to cover all specified forms -- 20+ integration tests proving the spec is implemented - -=== Priority 2: Drift Detection Demo (HIGH) - -**Why:** This is the differentiator. It needs to be undeniable. - -**Scope:** Build a reproducible demonstration. - -**What the demo does:** - -1. Create 1000 octad entities with consistent data across all modalities -2. Deliberately corrupt 50 of them -- alter the document text without - updating the vector embedding, change graph edges without updating the - document, modify temporal history to create inconsistencies -3. Run drift detection -- show VerisimDB identifying all 50 corrupted - entities with drift scores -4. Run self-normalisation -- show VerisimDB repairing the inconsistencies - by regenerating drifted modalities from the authoritative one -5. Verify -- re-run drift detection, show all 1000 entities are now - consistent - -**Deliverables:** - -- `demos/drift-detection/` directory with a runnable script -- `demos/drift-detection/README.adoc` explaining what it does and what to look for -- Output showing before/after drift scores -- This is the thing you show people. This is the pitch. - -=== Priority 3: One External User (HIGH) - -**Why:** Using your own database for your own projects proves it works for you. -One external user proves it works for someone else. That external user will -find assumptions you did not know you made. - -**How:** - -- IDApTIK (Joshua) is a start but you control both sides -- Find one other project -- academic, open-source, or a friend's -- that has - multi-representation data and would benefit from drift detection -- The demo from Priority 2 is the sales pitch -- Do not wait for VCL to be perfect; the REST API is sufficient for a first - user - -=== Priority 4: Positioning and Documentation (MEDIUM) - -**Why:** "6-modal database" sounds like marketing. "Database that catches -data rot" is a value proposition. - -**What to change:** - -1. **README** -- lead with drift detection, not the 6 modalities. The - modalities are implementation; drift detection is the product. - -2. **Tagline** -- not "6-core multimodal database with self-normalization" - but "the database that knows when your data disagrees with itself." - -3. **Comparison page** -- show the Virtuoso/ArangoDB comparison table from - this document. Be honest about where VerisimDB loses (maturity, query - language, scale) and where it wins (drift, temporal, vector). - -4. **Tutorial** -- a 10-minute "store a octad, search it, check its drift, - corrupt it, watch it heal" walkthrough. - -=== Priority 5: Heterogeneous Federation (MEDIUM-HIGH) - -**Why:** This is Layer 2 of the differentiator stack. VerisimDB watching -_other_ databases for consistency — not just other VerisimDB instances — is -the enterprise value proposition. - -**Scope:** Extend the existing federation system to support non-VerisimDB -peers via HTTP adapters. - -**What "done" means:** - -1. **Adapter interface** — a trait/behaviour that maps an external database's - API to VerisimDB's octad model. Each adapter knows how to read an entity - from its source and extract modality-relevant data. - -2. **ArangoDB adapter** — the IDApTIK bridge (`ArangoClient`) is the first - working example. Generalise it into a federation adapter that VerisimDB - can monitor for drift. - -3. **PostgreSQL adapter** — HTTP via PostgREST or direct SQL via Postgrex. - Read an entity from Postgres, extract document/graph/semantic data, compare - against VerisimDB's representation. - -4. **Drift detection across adapters** — "entity X in ArangoDB has 3 edges - but entity X in VerisimDB's graph modality has 4 — drift score 0.28." - -**Deliverables:** - -- `lib/verisim/federation/adapter.ex` — adapter behaviour -- `lib/verisim/federation/adapters/arango.ex` — ArangoDB adapter -- `lib/verisim/federation/adapters/postgres.ex` — PostgreSQL adapter -- Integration test with ArangoDB + VerisimDB running - -=== Priority 6: Replace oxrocksdb-sys (MEDIUM) - -**Why:** 15-minute container builds and a clang-19 dependency are painful. - -**Options (from KNOWN-ISSUES #24):** - -- `fjall` -- LSM-tree, same architecture as RocksDB, pure Rust -- `redb` -- B-tree, simpler, pure Rust -- ETS backend for the graph store (Elixir-side) - -The `GraphStore` trait abstraction already exists, so this is a backend swap. - -=== Priority 7: Wire VCL-UT (HIGH — but after VCL) - -**Why:** VCL-UT is the technical moat — the reason nobody can clone VerisimDB -quickly. Formally verified query results via Lean type checking is genuinely -unique in the database world. But VCL must work reliably first (Priority 1). -No point verifying proofs in a query language that cannot execute reliably. - -**Why it matters more than it might seem:** - -VCL-UT is not a nice-to-have. It is the capability that separates VerisimDB -from every other database and every data quality tool. Great Expectations can -assert "this column has no nulls." VerisimDB with VCL-UT can return a -_machine-verifiable proof certificate_ that the query result satisfies -dependent type constraints. This is the kind of guarantee required in: - -- Financial regulation (MiFID II, SOX compliance — prove the report is correct) -- Medical devices (IEC 62304 — formally verified data integrity) -- Aerospace (DO-178C — evidence-based assurance) -- Cross-institutional research (prove results without exposing raw data) - -No other database can do this. Not Virtuoso, not ArangoDB, not Postgres. -The Lean type definitions already exist. The proof obligation generation is -designed. The gap is wiring the Lean checker into the execution path. - -**Defer until:** VCL spec is complete and integration-tested (Priority 1). -Then this becomes the highest priority. - -**See:** KNOWN-ISSUES #21, `docs/vcl-vs-vcl-dt.adoc` for the phased roadmap. - -== What NOT to Do - -These are tempting but would dilute focus: - -1. **Do not add SQL support.** VerisimDB is not a relational database. Adding - SQL invites comparison with Postgres/MariaDB where VerisimDB loses. - -2. **Do not add SPARQL support.** Virtuoso and Oxigraph already do this. - VerisimDB's Oxigraph store speaks SPARQL internally; exposing it externally - makes VerisimDB look like a worse Virtuoso. - -3. **Do not add a web admin UI.** Admin UIs are maintenance burdens. The REPL - works. The REST API works. The Elixir orchestration layer is the admin - interface. - -4. **Do not chase benchmarks against other databases.** VerisimDB will not - beat ArangoDB at graph traversal speed or Tantivy at raw text search. The - benchmark that matters is: "how fast does VerisimDB detect that these 50 - entities are inconsistent?" That is a benchmark nobody else can run. - -5. **Do not add introspection endpoints.** No `/_introspect`, no - `SHOW COLLECTIONS`, no `DESCRIBE HEXAD`. The schema is in the code. - Introspection is attack surface with no upside for a system where the - orchestration layer is the sole client. - -== Summary - -VerisimDB is not crap. It has a genuine thesis that no other database -addresses. But the thesis is buried under incomplete query language support -and zero external validation. - -The path forward (priority order): - -1. **Finish VCL** so the database is _queryable_ -2. **Drift detection demo** so the differentiator is _undeniable_ -3. **Provenance / lineage modality** so cross-system data origins are _traceable_ (CRITICAL) -4. **One external user** so the assumptions are _tested_ -5. **Heterogeneous federation** so enterprise adoption is _practical_ -6. **Reposition** around drift detection so the value proposition is _clear_ -7. **Wire VCL-UT** so the technical moat is _operational_ (after VCL is solid) - -The layered pitch: - -- **Drift detection** gets people in the door (broad audience, easy to explain) -- **Federation** keeps them (watches existing databases, no migration) -- **Provenance** makes them depend on it (audit trail, GDPR, compliance) -- **VCL-UT** ensures nobody can compete (years of type theory to replicate) diff --git a/verisimdb/docs/design/DESIGN-2026-02-27-vcl-dt-assessment.adoc b/verisimdb/docs/design/DESIGN-2026-02-27-vcl-dt-assessment.adoc deleted file mode 100644 index ddcebacc..00000000 --- a/verisimdb/docs/design/DESIGN-2026-02-27-vcl-dt-assessment.adoc +++ /dev/null @@ -1,572 +0,0 @@ -// SPDX-License-Identifier: CC-BY-SA-4.0 -// Copyright (c) 2026 Jonathan D.A. Jewell (hyperpolymath) -= VCL-UT (Dependent Type Path) Wiring Assessment -:author: Jonathan D.A. Jewell (hyperpolymath) -:date: 2026-02-27 -:repo: verisimdb -:toc: -:toclevels: 3 - -== Executive Summary - -VCL-UT has *substantially more real implementation than initially expected*. -The ReScript type checker and Rust proof/semantic layer are genuinely functional -with real algorithms, tests, and meaningful code. However, the three layers -(ReScript, Rust, Elixir) are **not wired together end-to-end**. Each layer -works in isolation but the integration seams are broken or stubbed. - -[cols="1,1,1,1"] -|=== -| Layer | Implementation | Integration | Assessment - -| ReScript VCL -| ~85% real code -| 0% connected to Elixir executor -| *Impressive standalone work, needs integration* - -| Rust Proof/Semantic -| ~90% real code -| ~20% reachable from Elixir -| *Strong foundation, partially exposed via API* - -| Elixir Orchestration -| ~60% real code -| ~15% calls Rust proof APIs -| *Proof verification mostly stubbed* - -| Lean/Idris2 Type Checker -| 0% — does not exist -| N/A -| *No formal type checker present* -|=== - -**Honest overall VCL-UT completion: ~35%** (individual components exist; -end-to-end pipeline does not). - -== Layer-by-Layer Assessment - -=== 1. ReScript VCL Layer (`src/vcl/`) - -==== What Is Actually Implemented (Real Working Code) - -[cols="1,3"] -|=== -| File | Status - -| `VCLTypes.res` (296 lines) -| *FULLY IMPLEMENTED.* Pi types, Sigma types, ProofType, ProvedResultType, - QueryResultType. Conversion functions between AST types and type-level types. - No stubs. - -| `VCLBidir.res` (841 lines) -| *SUBSTANTIALLY IMPLEMENTED.* Real bidirectional type inference with - `synthesizeQuery` (9-step pipeline) and `checkQuery`. Handles proof clause - branching: Slipstream path returns `QueryResultType`, dependent-type path - calls `checkMultiProof` and returns `ProvedResultType`. Multi-proof - composition checks contract registry compatibility and mutual composability. - Mutation type checking (`synthesizeMutation`) covers INSERT/UPDATE/DELETE - with optional proofs. - -| `VCLParser.res` (1154 lines) -| *FULLY IMPLEMENTED.* Complete parser combinator library. Handles PROOF - clause parsing (`proofClause`), dependent type mode detection - (`parseDependentType`), Slipstream rejection of PROOF (`parseSlipstream`). - Full mutation parsing. 48 tests in `VCLParser_test.res`. - -| `VCLContext.res` (228 lines) -| *FULLY IMPLEMENTED.* Typing environment with bindings, contract registry, - modality field registries. Default context provides known fields per - modality. Contract composition checking. - -| `VCLSubtyping.res` (248 lines) -| *5 of 6 rules implemented.* Rule 6 (Refinement subsumption) explicitly - DEFERRED with comment "needs SMT solver". ProvedResultType correctly - subtypes plain QueryResultType (can forget proof). - -| `VCLProofObligation.res` (252 lines) -| *IMPLEMENTED with hardcoded circuit names.* Generates structured - `composedProofPlan` with composition strategies (Independent, Sequential, - Nested). Circuit names are hardcoded strings like "existence-proof-v1" — - these do NOT correspond to registered circuits in the Rust circuit registry. - -| `VCLCircuit.res` (72 lines) -| *TYPE DEFINITIONS ONLY.* Circuit DSL types (gates, wires, constraints). - `parseCustomProof` is trivial — just converts params to a Dict. - `serializeCircuitDef` delegates to `Js.Json.stringifyAny`. No actual - circuit compilation or verification logic on the ReScript side. - -| `VCLError.res` (447 lines) -| *FULLY IMPLEMENTED.* Comprehensive error hierarchy covering Phase 1 - (dependent type errors), Phase 2 (cross-modal), Phase 3 (write path). - -| `VCLExplain.res` (427 lines) -| *FULLY IMPLEMENTED.* Query plan visualization including proof obligation - nodes. Hardcoded cost estimates per modality (reasonable heuristics). - -| `VCLTypeChecker.res` (355 lines) -| *FULLY IMPLEMENTED.* Thin facade over VCLBidir. Backward-compatible API. -|=== - -==== What Is Stubbed/Incomplete - -1. **VCLSubtyping Rule 6** — Refinement subsumption needs an SMT solver - (e.g., Z3 or CVC5). This is required for proving that a refined type - (with proof) is a subtype of another refined type. - -2. **VCLCircuit.res** — The circuit DSL is type definitions only. No - serialization to R1CS or compilation logic on the ReScript side. The Rust - `circuit_compiler.rs` has this functionality but it is not reachable from - ReScript. - -3. **Hardcoded circuit names** in VCLProofObligation.res — The function - `proofKindToCircuit` returns static strings that are not validated against - the Rust circuit registry. - -==== What Is Missing - -1. **No integration with Elixir executor.** The ReScript type checker runs - entirely client-side (in the browser or via Deno/Node). The Elixir VCL - executor comment says "Type-check if PROOF clause present (VCLTypeChecker - -> VCLBidir)" but this invocation never happens. The Elixir executor - parses VCL independently using its own built-in parser. - -2. **No proof generation triggered from type checker.** The type checker - produces a `ProvedResultType` and `composedProofPlan`, but these are not - sent anywhere for actual proof generation. - -=== 2. Rust Proof/Semantic Layer (`rust-core/verisim-semantic/`) - -==== What Is Actually Implemented (Real Working Code) - -[cols="1,3"] -|=== -| File | Status - -| `lib.rs` (395 lines) -| *FULLY IMPLEMENTED.* SemanticStore trait, ProofBlob with CBOR - serialization/deserialization, InMemorySemanticStore with type registration, - annotation validation, proof storage. `ProofBlob::verify()` attempts ZKP - `VerifiableProofData` decoding with graceful legacy fallback. - -| `zkp.rs` (368 lines) -| *FULLY IMPLEMENTED.* Real SHA-256 hash commitments, Merkle tree - construction (`build_merkle_tree`), Merkle proof generation - (`generate_merkle_proof`) and verification (`verify_merkle_proof`). - `VerifiableProofData` enum with Commitment, Reveal, MerkleInclusion, - ContentIntegrity variants. Constant-time byte comparison. 12 tests. - -| `zkp_bridge.rs` (629 lines) -| *SUBSTANTIALLY IMPLEMENTED but NOT true ZK-SNARKs.* Privacy-level routing - (Public/Private/ZeroKnowledge). ZeroKnowledge path uses blinded Merkle - proofs with committed witnesses — this is a real cryptographic commitment - scheme but NOT actual zero-knowledge proofs. Comments acknowledge this: - "Full ZK-SNARK via sanctify is designed but not yet compiled in" (line 49) - and "Full ZK verification would invoke a ZK-SNARK verifier here" (line - 337). Nonce generation is deterministic (NOT cryptographically secure). - 11 tests. - -| `circuit_registry.rs` (314 lines) -| *FULLY IMPLEMENTED.* In-memory circuit registry with RwLock. R1CS - constraint system (`R1CSConstraint { a, b, c }` where A * B = C). - `CompiledCircuit::verify` evaluates constraints against witness/public - inputs. Full CRUD API. Tests pass. - -| `circuit_compiler.rs` (297 lines) -| *FULLY IMPLEMENTED.* Compiles circuit definitions into R1CS constraints. - Handles AND, OR, XOR, NOT, LinearCombination gate types. Generates - verification keys from SHA-256 Merkle commitment over constraints. - -| `verification_keys.rs` (247 lines) -| *FULLY IMPLEMENTED.* Key management with rotation support (active + - previous key). Federation key export/import with instance-prefixed naming. - Expiry checking. Tests pass. - -| `proven_bridge.rs` (278 lines) -| *FULLY IMPLEMENTED for certificate parsing.* Parses JSON/CBOR certificates - from the `proven` library. ProverKind enum (Z3, Lean, Coq, Agda, Idris2, - Custom). Certificate verification: checks version, signature (SHA-256 of - prover+statement+proof_term), timestamp. Converts certificates to - ProofBlob. BUT: the `proven` library itself is NOT integrated — this - bridge only consumes pre-existing certificates. - -| `sanctify_bridge.rs` (414 lines) -| *FULLY IMPLEMENTED.* Parses sanctify-php security reports, security issue - types, severity levels, contract validation, proof blob conversion. - Tangential to VCL-UT but functional. -|=== - -==== What Is Stubbed/Incomplete - -1. **No actual ZK-SNARK backend.** The `zkp_bridge.rs` uses hash - commitments and Merkle proofs as an approximation. True zero-knowledge - proofs (Groth16, PLONK, etc.) are not implemented. - -2. **Deterministic nonce generation** in `zkp_bridge.rs` (line 349-350): - `"In production, replace with a CSPRNG."` This means ZKP proofs are - currently NOT cryptographically secure. - -3. **No connection to external proof assistants.** The `proven_bridge.rs` - parses certificates that would be generated by the `proven` library (an - Idris2-based ZKP system), but `proven` is not integrated or invoked. The - bridge is a one-way importer for externally-generated certificates. - -==== What Is Exposed via API - -The Rust API (`verisim-api/src/lib.rs`) exposes three proof endpoints: - -- `POST /proofs/generate` — calls `zkp_api::generate_zkp()` (hash - commitments/Merkle proofs, not ZK-SNARKs) -- `POST /proofs/verify` — calls `zkp_api::verify_zkp()` -- `POST /proofs/generate-with-circuit` — calls - `zkp_api::generate_zkp_with_circuit()` (circuit registry integration) - -These endpoints are REAL and functional. The Elixir executor calls the first -two for `:zkp` and `:proven` proof types. - -==== What Is NOT Exposed - -The Rust VCL endpoint (`vcl.rs`) is a **Slipstream-only parser**. It handles -SELECT, SEARCH, INSERT, DELETE, SHOW, COUNT, EXPLAIN but has *zero PROOF -clause handling*. Any VCL query with a PROOF clause sent to the Rust VCL -endpoint would fail or be silently ignored. - -=== 3. Elixir Orchestration Layer (`elixir-orchestration/lib/verisim/query/`) - -==== What Is Actually Implemented - -[cols="1,3"] -|=== -| File | Status - -| `vcl_bridge.ex` (672 lines) -| *IMPLEMENTED with degraded fallback.* GenServer managing a Deno/Node - subprocess that runs compiled ReScript VCL parser. Falls back to built-in - Elixir parser when subprocess unavailable. `parse_dependent/2` exists but - the fallback parser only collects PROOF tokens as raw strings, not - structured proof specs. - -| `vcl_executor.ex` (~1000 lines) -| *PARTIALLY IMPLEMENTED.* Full SELECT/INSERT/UPDATE/DELETE execution, - cross-modal condition evaluation (DRIFT, CONSISTENT, EXISTS/NOT EXISTS - between modalities), pagination, ORDER BY, GROUP BY, aggregates. BUT: proof - verification is **largely stubbed** (see below). - -| `query_router.ex` (187 lines) -| *IMPLEMENTED.* Routes queries by type (:text, :vector, :graph, :semantic, - :temporal, :multi) to RustClient. No VCL-UT-specific logic. -|=== - -==== The Critical Proof Verification Gap - -The `verify_single_proof/1` function in `vcl_executor.ex` (lines 514-592) -reveals the core disconnect: - -[source,elixir] ----- -# What each proof type ACTUALLY does: - -:existence -> :ok # Stub — always passes -:citation -> validate_contract_exists() # Text search for contract name -:access -> :ok # Stub — always passes -:integrity -> text search for contract # Superficial check -:provenance -> :ok # Stub — always passes -:custom -> validate_contract_exists() # Text search for contract name -:zkp -> RustClient.post("/proofs/generate", ...) # REAL call to Rust -:proven -> RustClient.post("/proofs/generate", ...) # REAL call to Rust -:sanctify -> validate_contract_exists() # Text search for contract name ----- - -Out of 9 proof types, **only 2 actually call the Rust ZKP API**. - -Furthermore, `validate_contract_exists/1` (line 626-638) has a critical -safety bypass: - -[source,elixir] ----- -{:error, _} -> - # If search fails (e.g., Rust core unavailable), allow proof to pass - # with a warning — this prevents query failures during development - Logger.warning("Cannot verify contract '#{contract_name}': semantic store unreachable") - :ok # <-- SILENTLY PASSES when Rust is down ----- - -This means even the contract existence checks are unreliable — if the Rust -core is down, ALL proof verification silently passes. - -==== What The Elixir Executor Does NOT Do - -1. Does **not invoke the ReScript bidirectional type checker** -2. Does **not type-check VCL-UT queries** (only the ReScript layer does this, - client-side) -3. Does **not generate actual proof obligations** from PROOF clauses (only the - ReScript `VCLProofObligation.res` does this, client-side) -4. Does **not verify circuit constraints** via the Rust circuit registry -5. Does **not call the circuit compilation endpoint** (`/proofs/generate-with-circuit`) -6. Does **not verify proven certificates** (calls `/proofs/generate` instead, - which generates new hash commitments rather than verifying existing - certificates) - -=== 4. Lean/Idris2 Type Checker - -**Does not exist.** The Idris2 files found in `src/abi/` (Types.idr, -Layout.idr, Foreign.idr) are standard ABI definitions for VeriSimDB's FFI -layer, not VCL type checkers. No `.lean` files exist anywhere in the -repository. - -The CLAUDE.md Known Issues section acknowledges: "VCL-UT not connected to VCL -PROOF runtime (Lean type checker not invoked)." This is accurate — no Lean -type checker has been written. - -== Integration Gap Analysis - -=== Disconnected Seams - -[cols="1,1,1"] -|=== -| From | To | Status - -| ReScript type checker -| Elixir executor -| *BROKEN.* Executor never invokes type checker. - -| ReScript proof obligations -| Rust ZKP bridge -| *BROKEN.* Obligations are generated client-side but never sent to Rust. - -| ReScript circuit DSL -| Rust circuit compiler -| *BROKEN.* Circuit definitions cannot reach the compiler. - -| Elixir executor proof verify -| Rust circuit registry -| *BROKEN.* Executor never calls `/proofs/generate-with-circuit`. - -| Elixir executor -| Rust VCL endpoint -| *PARTIAL.* The Rust VCL endpoint has no PROOF handling. - -| proven library (external) -| Rust proven_bridge -| *BROKEN.* proven is not integrated; bridge only imports certificates. - -| Rust ZKP bridge -| Real ZK-SNARK backend -| *MISSING.* Only hash commitments, no real ZK-SNARKs. -|=== - -=== Data Flow: What Should Happen vs What Actually Happens - -**Intended VCL-UT data flow:** -[source] ----- -Client sends VCL with PROOF clause - -> Elixir executor receives query - -> ReScript type checker validates types + generates proof obligations - -> Proof obligations sent to Rust ZKP bridge - -> Rust generates real ZK proofs (SNARKs) - -> Proofs stored in semantic store - -> Results returned with proof certificate (ProvedResultType) ----- - -**What actually happens today:** -[source] ----- -Client sends VCL with PROOF clause - -> Elixir executor receives query (via fallback parser) - -> PROOF tokens collected as raw strings - -> verify_single_proof() called per spec - -> Most proof types return :ok (stub) or do text search - -> :zkp/:proven types call Rust /proofs/generate - -> Rust generates hash commitments (NOT ZK-SNARKs) - -> Query results returned WITHOUT proof certificate - -> ReScript type checker never invoked - -> ProvedResultType never constructed server-side ----- - -== Scope of Remaining Work - -=== Priority 1: Wire ReScript Type Checker to Elixir Executor (HIGH) - -**Effort: 3-5 days** - -The ReScript type checker is the most complete component but is entirely -disconnected. Wire it: - -1. Ensure `vcl_bridge.ex` reliably spawns the ReScript parser process (or - compile ReScript to a standalone binary via Deno). -2. Add a `typecheck/2` function to `vcl_bridge.ex` that sends the parsed AST - to the ReScript type checker and receives back: - - Inferred types - - Proof obligations (from `VCLProofObligation.res`) - - Composition strategy -3. In `vcl_executor.ex`, after parsing a query with PROOF clause, call the - type checker before executing. -4. Use the returned proof obligations to drive actual proof generation - (Priority 2). - -=== Priority 2: Connect Proof Obligations to Rust ZKP Bridge (HIGH) - -**Effort: 2-3 days** - -Replace the stubbed `verify_single_proof/1` with real proof generation: - -1. For each proof obligation from the type checker, construct a proper - `ZkpBridgeRequest` and call the Rust API. -2. Use `/proofs/generate-with-circuit` for Custom proof types (requires - circuit name resolution). -3. Store generated proofs in the semantic store via `POST /proofs/store` - (this endpoint may need to be added to the Rust API). -4. Return proofs alongside query results. - -=== Priority 3: Implement Real Proof Verification in Executor (MEDIUM) - -**Effort: 2-3 days** - -1. Replace `:existence`, `:access`, `:provenance` stubs with actual checks - (query the octad store, check access policies, verify provenance chain). -2. Remove the silent-pass-on-error in `validate_contract_exists/1`. -3. For `:proven` type, call the proven bridge's certificate verification - rather than generating new hash commitments. -4. For `:sanctify` type, call the sanctify bridge's contract validation. - -=== Priority 4: VCLSubtyping Rule 6 — Refinement Subsumption (MEDIUM) - -**Effort: 3-5 days (significant)** - -Requires integrating an SMT solver (Z3 or CVC5) to prove refinement subtyping -relationships. Options: - -1. **Z3 via WASM** — Run Z3 in the browser/Deno alongside the ReScript type - checker. Heaviest but most powerful. -2. **Z3 via Rust binding** — Add `z3-sys` crate to `verisim-semantic` and - expose an endpoint. Requires linking to Z3's C library. -3. **Simplified refinement checking** — Implement a subset of refinement - subsumption without a full SMT solver (handles common cases). - -=== Priority 5: Real ZK-SNARK Backend (LOW — significant scope) - -**Effort: 2-4 weeks** - -The current hash commitment + Merkle proof approach is NOT zero-knowledge. -To get real ZK-SNARKs: - -1. Integrate a Rust ZK library (`bellman`, `arkworks`, or `halo2`). -2. Implement Groth16 or PLONK proving/verification. -3. Replace the blinded Merkle proof approximation in `zkp_bridge.rs`. -4. Fix the deterministic nonce generation with a CSPRNG. -5. Update circuit compiler to emit actual ZK circuits. - -=== Priority 6: Lean/Idris2 Formal Type Checker (LOW — aspirational) - -**Effort: 4-8 weeks** - -The CLAUDE.md mentions a Lean type checker but none exists. Building one would -mean: - -1. Define VCL-UT's type system in Lean 4 or Idris2. -2. Implement bidirectional type checking with dependent types. -3. Generate proof obligations as Lean/Idris2 terms. -4. Verify proof obligations against the type system. -5. Bridge to Rust/Elixir for runtime integration. - -This is the most ambitious item and may not be necessary if the ReScript type -checker is sufficient for practical use. - -== Honest Completion Percentages - -[cols="1,1,1,1"] -|=== -| Component | Code Written | Code Working | End-to-End Wired - -| ReScript VCL Parser -| 95% -| 95% -| 0% - -| ReScript Type Checker (VCLBidir) -| 85% -| 85% -| 0% - -| ReScript Proof Obligations -| 80% -| 80% -| 0% - -| Rust ZKP Core (zkp.rs) -| 90% -| 90% -| 20% - -| Rust ZKP Bridge (privacy levels) -| 85% -| 85% -| 20% - -| Rust Circuit Registry + Compiler -| 90% -| 90% -| 0% - -| Rust Proven Bridge -| 90% -| 90% -| 0% - -| Rust Verification Keys -| 90% -| 90% -| 0% - -| Elixir VCL Bridge -| 70% -| 70% -| 40% - -| Elixir VCL Executor (proof path) -| 30% -| 30% -| 15% - -| Real ZK-SNARK Backend -| 0% -| 0% -| 0% - -| Lean/Idris2 Type Checker -| 0% -| 0% -| 0% - -| *Overall VCL-UT Pipeline* -| *~60%* -| *~55%* -| *~10%* -|=== - -NOTE: "Code Written" and "Code Working" are high because real algorithms are -implemented and tests pass. "End-to-End Wired" is abysmal because the layers -do not talk to each other for the VCL-UT path. - -== Key Takeaways - -1. **The individual components are much better than expected.** The ReScript - type checker is a real bidirectional type checker with dependent types. The - Rust proof layer has real cryptographic primitives. Neither is stub code. - -2. **The integration is much worse than expected.** The three layers operate - as isolated islands. The Elixir executor's proof verification is the weakest - link — it's largely hardcoded `:ok` returns. - -3. **The shortest path to a working VCL-UT pipeline** is wiring the ReScript - type checker to the Elixir executor (Priority 1) and connecting proof - obligations to the Rust ZKP bridge (Priority 2). This would give ~60% - end-to-end functionality in roughly 1-2 weeks of focused work. - -4. **Real ZK-SNARKs are a separate, larger project.** The current hash - commitment approach is functional for demonstration and development but - provides no actual zero-knowledge properties. This is the honest truth. - -5. **The Lean type checker is aspirational.** The ReScript type checker is - capable enough for practical VCL-UT use. A Lean version would provide - formal verification guarantees but is not blocking functionality. diff --git a/verisimdb/docs/design/DESIGN-2026-02-28-panll-interop-telemetry.md b/verisimdb/docs/design/DESIGN-2026-02-28-panll-interop-telemetry.md deleted file mode 100644 index 643ad8d9..00000000 --- a/verisimdb/docs/design/DESIGN-2026-02-28-panll-interop-telemetry.md +++ /dev/null @@ -1,178 +0,0 @@ -# Design Document: PanLL Database Interop + Telemetry - -**Date:** 2026-02-28 -**Repo:** nextgen-databases/verisimdb + panll -**Session:** PanLL database module protocol, opt-in telemetry, product development insights - -## Problem Statement - -VeriSimDB's PanLL integration is currently hardcoded — `verisimdbState` lives directly -in PanLL's Model, with VeriSimDB-specific messages and Tauri commands. This means: - -1. **No plugin architecture** for other databases (QuandleDB, LithoGlyph) -2. **No telemetry** — we have `:telemetry` events defined but no aggregation, export, or UI -3. **No product development feedback loop** — we don't know how the system is used -4. **Each playground is standalone** — VCL Playground, future KQL/GQL playgrounds are islands - -## Design Goals - -1. **Database Module Protocol** — standard interface for any database to plug into PanLL -2. **Opt-In Telemetry** — transparent, privacy-first product development metrics -3. **User-Facing Insights** — telemetry is also useful to database operators -4. **Playground Gallery** — PanLL manages query language playgrounds uniformly - -## Architecture - -### Database Module Protocol - -Each database module implements a standard protocol for PanLL integration: - -``` -┌───────────────────────────────────────────────────────┐ -│ PanLL Database Module Protocol │ -│ │ -│ capabilities(): → list │ -│ health(): → result │ -│ query(string): → result │ -│ telemetry(): → telemetrySnapshot │ -│ playground(): → playgroundConfig │ -│ │ -│ Pane-L Mapping: grammar, type system, constraints │ -│ Pane-N Mapping: query engine, inference, reasoning │ -│ Pane-W Mapping: results, drift, telemetry dashboard │ -└───────────────────────────────────────────────────────┘ -``` - -**Capabilities** (what a database can do): -- `QueryExecution` — run queries in its native language -- `DriftDetection` — detect cross-modal consistency issues -- `ProofGeneration` — generate verifiable proof certificates -- `Normalisation` — self-repair drifted modalities -- `Federation` — federate queries across backends -- `Telemetry` — expose operational metrics -- `Playground` — provide an interactive query editor - -### Telemetry Architecture - -``` - ┌──────────────┐ - │ VeriSimDB │ - │ Telemetry │ - │ (Elixir) │ - └──────┬───────┘ - │ :telemetry events - ┌──────▼───────┐ - │ Collector │ - │ (ETS-based) │ - └──────┬───────┘ - │ periodic flush - ┌──────▼───────┐ - │ Reporter │ - │ (JSON export│ - │ + HTTP API) │ - └──────┬───────┘ - │ - ┌─────────────┼─────────────┐ - │ │ │ - ┌──────▼───┐ ┌─────▼────┐ ┌─────▼────┐ - │ PanLL UI │ │ JSON │ │ Product │ - │ Dashboard│ │ Export │ │ Insights │ - └──────────┘ └──────────┘ └──────────┘ -``` - -**Privacy Guarantees:** -- **Opt-in only** — telemetry disabled by default, explicit `VERISIM_TELEMETRY=true` -- **No PII** — never captures query content, entity data, or user identifiers -- **Aggregate only** — counts, distributions, rates — never individual records -- **Local first** — all data stays on the machine unless user explicitly exports -- **User-visible** — telemetry dashboard shows exactly what is collected - -### Telemetry Events - -| Event | What it measures | Why it helps | -|-------|-----------------|-------------| -| `verisim.query.modality_usage` | Which modalities are queried most | Prioritise optimisation work | -| `verisim.query.pattern` | Query shape (SELECT, SEARCH, INSERT, etc.) | Understand usage patterns | -| `verisim.query.duration_distribution` | Query latency percentiles | Performance regression detection | -| `verisim.drift.frequency` | How often drift is detected | Gauge system stability | -| `verisim.drift.modality_breakdown` | Which modalities drift most | Focus normalisation work | -| `verisim.normalise.success_rate` | How often normalisation succeeds | Quality metric | -| `verisim.federation.peer_health` | Federated peer availability | Operational monitoring | -| `verisim.proof.type_usage` | Which proof types are used | VCL-UT adoption metric | -| `verisim.entity.modality_coverage` | Average modalities per entity | Data completeness metric | -| `verisim.system.uptime` | Server uptime | Reliability metric | - -### Product Insights Derived from Telemetry - -The reporter aggregates raw telemetry into actionable insights: - -1. **Modality Heatmap** — which of the 8 modalities see real use vs. are ignored -2. **Query Pattern Distribution** — what percentage is read vs. write vs. drift vs. proof -3. **Performance Trends** — is p95 latency improving or degrading over time? -4. **Drift Frequency** — are entities staying consistent or constantly re-normalising? -5. **Federation Health** — are peer backends reliable or flaky? -6. **VCL-UT Adoption** — are users using dependent type proofs? - -### PanLL Integration Points - -**Existing (already built):** -- `Model.res` — `verisimdbState` with drift scores, proof obligations -- `Msg.res` — `verisimdbMsg` with health, query, drift, normalise, entity detail -- `TauriCmd.res` — 7 VeriSimDB Tauri commands -- `PaneW.res` — Database tools panel, drift heatmap, VCL query area - -**New (this session):** -- `Model.res` — add `telemetryState` to verisimdbState -- `Msg.res` — add `FetchTelemetry`, `TelemetryLoaded` messages -- `TauriCmd.res` — add `getTelemetry` Tauri command -- `PaneW.res` — add telemetry dashboard panel -- `DatabaseModule.res` (NEW) — protocol types for generic database modules -- `DatabaseRegistry.res` (NEW) — manages registered database modules - -## Implementation Plan - -### Phase 1: VeriSimDB Telemetry (Elixir-side) - -1. Extend `lib/verisim/telemetry.ex` with event emission across modules -2. Create `lib/verisim/telemetry/collector.ex` — ETS-based metric aggregation -3. Create `lib/verisim/telemetry/reporter.ex` — JSON export + HTTP endpoint -4. Create `lib/verisim/telemetry/product_insights.ex` — derived insights -5. Add telemetry endpoint to verisim-api (GET /api/v1/telemetry) - -### Phase 2: PanLL Database Module Protocol - -1. Create `src/modules/DatabaseModule.res` — type definitions -2. Create `src/modules/DatabaseRegistry.res` — module registry -3. Extend `Model.res` with telemetry state -4. Extend `Msg.res` with telemetry messages -5. Add telemetry Tauri command to `TauriCmd.res` - -### Phase 3: PanLL Telemetry Dashboard - -1. Extend `PaneW.res` with telemetry visualization panel -2. Modality usage heatmap (reuses drift heatmap pattern) -3. Query pattern distribution bar chart -4. Performance trend indicators - -### Phase 4: Playground Gallery (Future) - -1. Abstract VCL Playground as a PanLL module -2. Define playground protocol (editor, linter, formatter, executor) -3. Register VCL, future KQL, future GQL playgrounds -4. PanLL manages playground lifecycle - -## Files Changed/Created - -### VeriSimDB (elixir-orchestration) -- `lib/verisim/telemetry.ex` — extend with new events -- `lib/verisim/telemetry/collector.ex` — NEW: ETS-based aggregation -- `lib/verisim/telemetry/reporter.ex` — NEW: JSON export + insights -- `test/verisim/telemetry_test.exs` — NEW: telemetry tests - -### PanLL -- `src/modules/DatabaseModule.res` — NEW: protocol types -- `src/modules/DatabaseRegistry.res` — NEW: module registry -- `src/Model.res` — extend verisimdbState with telemetry -- `src/Msg.res` — add telemetry messages -- `src/commands/TauriCmd.res` — add telemetry command -- `src/components/PaneW.res` — add telemetry panel diff --git a/verisimdb/docs/drift-handling.adoc b/verisimdb/docs/drift-handling.adoc deleted file mode 100644 index 6fbee35a..00000000 --- a/verisimdb/docs/drift-handling.adoc +++ /dev/null @@ -1,839 +0,0 @@ -// SPDX-License-Identifier: CC-BY-SA-4.0 - -= Drift Handling in VeriSimDB -:toc: -:toc-placement!: - -Comprehensive drift detection, repair, and query integration for federated knowledge stores. - -toc::[] - -== Overview - -**Drift** occurs when the same Octad exists in multiple modality stores with inconsistent representations. VeriSimDB provides: - -1. **Drift Detection** - Automatic monitoring across modalities -2. **Drift Repair** - Reconciliation strategies to restore consistency -3. **Drift-Aware Queries** - VCL extensions for querying drifted data -4. **Drift Tolerance** - Configurable policies for acceptable inconsistency - -== Types of Drift - -=== Cross-Modal Drift - -**Definition:** Representations of the same Octad differ across modalities. - -**Example:** - -``` -Octad: 550e8400-e29b-41d4-a716-446655440000 - -verisim:graph → title: "Machine Learning Paper" -verisim:document → title: "ML Paper" ❌ DRIFT -verisim:vector → embedding: [0.1, 0.2, ...] ✓ OK -``` - -**Causes:** - -- Partial update (only graph updated, document not refreshed) -- Network partition during write -- Concurrent updates from different clients - -=== Temporal Drift - -**Definition:** Data changes over time without proper versioning. - -**Example:** - -``` -t0: octad.title = "Draft Paper" -t1: octad.title = "Published Paper" (document updated) -t1: octad.graph still has "Draft Paper" ❌ DRIFT -``` - -**Causes:** - -- Asynchronous replication lag -- Cache staleness -- Delayed batch updates - -=== Federation Drift - -**Definition:** Same Octad stored at multiple organizations with divergent state. - -**Example:** - -``` -Octad: 550e8400-e29b-41d4-a716-446655440000 - -University A: retraction_status = "active" -University B: retraction_status = "retracted" ❌ DRIFT -``` - -**Causes:** - -- Network partitions -- Conflicting updates -- Malicious tampering (Byzantine drift) - -=== Semantic Drift - -**Definition:** Type annotations or contracts become inconsistent. - -**Example:** - -``` -verisim:semantic → types: ["https://schema.org/Paper", "https://schema.org/Article"] -verisim:graph → types: ["https://schema.org/Paper"] ❌ DRIFT -``` - -**Causes:** - -- Schema evolution -- Type inference errors -- Manual type corrections - -== Drift Detection - -=== Automatic Detection - -**Continuous monitoring** (every 5 minutes by default): - -```elixir -# lib/verisim/drift_monitor.ex -defmodule VeriSim.DriftMonitor do - @doc """ - Check all octads for cross-modal drift. - """ - def detect_all_drift do - octad_ids = list_all_octad_ids() - - octad_ids - |> Stream.chunk_every(100) - |> Enum.each(fn chunk -> - chunk - |> Task.async_stream(&detect_octad_drift/1, max_concurrency: 10) - |> Stream.filter(fn {:ok, result} -> result.has_drift end) - |> Enum.each(&log_drift/1) - end) - end - - defp detect_octad_drift(octad_id) do - # Fetch from all modalities - graph_rep = VeriSim.Graph.get(octad_id) - vector_rep = VeriSim.Vector.get(octad_id) - document_rep = VeriSim.Document.get(octad_id) - semantic_rep = VeriSim.Semantic.get(octad_id) - temporal_rep = VeriSim.Temporal.get(octad_id) - - # Check for inconsistencies - drifts = [] - - drifts = if graph_rep.title != document_rep.title do - [{:title_mismatch, graph_rep.title, document_rep.title} | drifts] - else - drifts - end - - drifts = if graph_rep.updated_at != vector_rep.updated_at do - [{:timestamp_mismatch, graph_rep.updated_at, vector_rep.updated_at} | drifts] - else - drifts - end - - %{octad_id: octad_id, has_drift: length(drifts) > 0, drifts: drifts} - end -end -``` - -=== On-Demand Detection (VCL) - -**Explicit drift check query:** - -```vcl -DRIFT DETECT -FROM verisim:graph -WHERE octad.id = @id; -``` - -**Response:** - -```json -{ - "octad_id": "550e8400-e29b-41d4-a716-446655440000", - "has_drift": true, - "drifts": [ - { - "type": "title_mismatch", - "modalities": ["graph", "document"], - "values": { - "graph": "Machine Learning Paper", - "document": "ML Paper" - }, - "detected_at": "2025-01-15T10:30:00Z", - "severity": "medium" - } - ] -} -``` - -**Batch drift detection:** - -```vcl -DRIFT DETECT -FROM verisim:graph -WHERE octad.types INCLUDES "https://schema.org/Paper" -LIMIT 100; -``` - -== Drift Repair - -=== Repair Strategies - -**1. Latest Wins** - -Most recent update across all modalities becomes canonical. - -```vcl -DRIFT REPAIR -FROM verisim:graph -WHERE octad.id = @id -USING STRATEGY latest_wins; -``` - -**Implementation:** - -```rust -// rust-core/verisim-drift/src/repair.rs - -pub fn repair_latest_wins(octad_id: &Uuid) -> Result { - // Fetch from all modalities with timestamps - let reps = fetch_all_representations(octad_id)?; - - // Find representation with latest updated_at - let latest = reps.iter() - .max_by_key(|r| r.updated_at) - .ok_or(Error::NoRepresentations)?; - - // Propagate latest to all modalities - for modality in &["graph", "vector", "document", "semantic", "temporal"] { - if modality != &latest.modality { - write_representation(modality, &latest)?; - } - } - - Ok(RepairResult { - octad_id: *octad_id, - canonical_source: latest.modality.clone(), - repaired_modalities: reps.len() - 1, - }) -} -``` - -**2. Quorum Consensus** - -Value that appears in majority of modalities wins. - -```vcl -DRIFT REPAIR -FROM verisim:graph -WHERE octad.id = @id -USING STRATEGY quorum; -``` - -**Implementation:** - -```rust -pub fn repair_quorum(octad_id: &Uuid) -> Result { - let reps = fetch_all_representations(octad_id)?; - - // Group by value, count occurrences - let mut value_counts: HashMap = HashMap::new(); - - for rep in &reps { - *value_counts.entry(rep.title.clone()).or_insert(0) += 1; - } - - // Find value with most votes - let (canonical_value, _count) = value_counts.iter() - .max_by_key(|(_, count)| *count) - .ok_or(Error::NoConsensus)?; - - // Propagate canonical value - for rep in &reps { - if &rep.title != canonical_value { - update_representation(&rep.modality, octad_id, canonical_value)?; - } - } - - Ok(RepairResult { ... }) -} -``` - -**3. Manual Resolution** - -Present conflict to user for manual decision. - -```vcl -DRIFT REPAIR -FROM verisim:graph -WHERE octad.id = @id -USING STRATEGY manual; -``` - -**Response:** - -```json -{ - "octad_id": "550e8400-e29b-41d4-a716-446655440000", - "conflicts": [ - { - "field": "title", - "options": [ - { - "value": "Machine Learning Paper", - "modalities": ["graph", "semantic"], - "updated_at": "2025-01-15T10:00:00Z" - }, - { - "value": "ML Paper", - "modalities": ["document", "vector"], - "updated_at": "2025-01-15T10:05:00Z" - } - ] - } - ], - "resolution_token": "abc123..." -} -``` - -**User resolves manually:** - -```http -POST /api/v1/drift/resolve HTTP/1.1 - -{ - "resolution_token": "abc123...", - "selected_value": "Machine Learning Paper" -} -``` - -**4. Merge** - -Combine values intelligently (domain-specific). - -```vcl -DRIFT REPAIR -FROM verisim:graph -WHERE octad.id = @id -USING STRATEGY merge; -``` - -**Example: Merging tags** - -``` -graph: tags = ["machine-learning", "neural-networks"] -document: tags = ["deep-learning", "neural-networks"] - -merged: tags = ["machine-learning", "neural-networks", "deep-learning"] -``` - -=== Automatic Repair - -**Triggered automatically** when drift exceeds tolerance threshold: - -```elixir -# config.exs -config :verisim, - drift_tolerance: 0.05, # 5% of octads can have drift - auto_repair: true, - repair_strategy: :latest_wins -``` - -**Repair workflow:** - -``` -1. Drift detection runs every 5 minutes -2. If drift_rate > threshold → trigger repair -3. Apply repair_strategy to all drifted octads -4. Log repair actions to verisim-temporal -5. Verify repair success (re-check drift) -``` - -=== Drift Repair Verification - -**After repair, verify consistency:** - -```vcl -DRIFT VERIFY -FROM verisim:graph -WHERE octad.id = @id; -``` - -**Response:** - -```json -{ - "octad_id": "550e8400-e29b-41d4-a716-446655440000", - "is_consistent": true, - "checked_modalities": ["graph", "vector", "document", "semantic", "temporal"], - "last_repair": "2025-01-15T10:30:00Z", - "verified_at": "2025-01-15T10:31:00Z" -} -``` - -== Drift-Aware Queries - -=== Query with Drift Tolerance - -**Accept some drift** (performance optimization): - -```vcl -FROM verisim:graph -WITH DRIFT TOLERANCE 0.1 -- Allow 10% drift -WHERE octad.types INCLUDES "https://schema.org/Paper" -LIMIT 100; -``` - -**Behavior:** - -- Query returns results even if some octads have drift -- Drift warnings included in response metadata -- Faster than strict consistency checks - -**Response:** - -```json -{ - "data": [...], - "metadata": { - "total_results": 100, - "drifted_results": 8, - "drift_rate": 0.08, - "drift_tolerance": 0.1, - "warning": "Some results may have inconsistent representations" - } -} -``` - -=== Query with Strict Consistency - -**Require perfect consistency** (default): - -```vcl -FROM verisim:graph -WITH DRIFT TOLERANCE 0.0 -- No drift allowed -WHERE octad.types INCLUDES "https://schema.org/Paper" -LIMIT 100; -``` - -**Behavior:** - -- Query fails if ANY result has drift -- Suggests running `DRIFT REPAIR` first -- Slowest (checks all modalities) - -=== Query Before/After Drift Repair - -**Time-travel query** using temporal modality: - -```vcl --- Query BEFORE repair (as of specific time) -FROM verisim:temporal -WHERE octad.id = @id -AS OF TIMESTAMP '2025-01-15T09:00:00Z'; - --- Query AFTER repair (latest) -FROM verisim:graph -WHERE octad.id = @id; -``` - -**Use case:** Audit trail, rollback analysis, drift impact assessment - -=== Drift-Tolerant Federation - -**Query across federated stores with mixed drift:** - -```vcl -FROM verisim:federation -WHERE octad.types INCLUDES "https://schema.org/Paper" -WITH DRIFT TOLERANCE 0.2 -- Tolerate 20% drift across federation -MIN QUORUM 3; -- At least 3 stores must respond -``` - -**Behavior:** - -- Query sent to all federated stores -- Accept partial results if quorum met -- Drift warnings per store -- Aggregated drift rate in response - -== Drift Policies - -=== Policy Configuration - -**Per-modality drift policies:** - -```elixir -# config.exs -config :verisim, - drift_policies: %{ - graph: %{ - tolerance: 0.05, - auto_repair: true, - strategy: :latest_wins, - check_interval_ms: 300_000 # 5 minutes - }, - vector: %{ - tolerance: 0.1, # Embeddings can have more drift - auto_repair: false, # Manual review for embeddings - strategy: :manual - }, - document: %{ - tolerance: 0.0, # Text must be consistent - auto_repair: true, - strategy: :latest_wins - }, - semantic: %{ - tolerance: 0.0, # Types must be consistent - auto_repair: true, - strategy: :quorum - }, - temporal: %{ - tolerance: 0.0, # History is immutable - auto_repair: false, - strategy: :none - } - } -``` - -=== Drift Severity Levels - -**Classify drift by impact:** - -| Severity | Description | Action | -|----------|-------------|--------| -| **LOW** | Formatting differences (e.g., "ML Paper" vs "ML Paper ") | Log only | -| **MEDIUM** | Content differences (e.g., title mismatch) | Alert + auto-repair | -| **HIGH** | Semantic differences (e.g., retraction status) | Alert + manual review | -| **CRITICAL** | Security differences (e.g., access control mismatch) | Block queries + escalate | - -**Severity detection:** - -```rust -pub fn classify_drift_severity(drift: &Drift) -> DriftSeverity { - match drift.field { - "title" | "body" if drift.edit_distance() < 5 => DriftSeverity::Low, - "title" | "body" => DriftSeverity::Medium, - "retraction_status" | "access_control" => DriftSeverity::Critical, - "types" | "contracts" => DriftSeverity::High, - _ => DriftSeverity::Medium, - } -} -``` - -== Federation Drift - -=== Cross-Organization Drift - -**Challenge:** Universities A and B have divergent state for the same Octad. - -**Detection:** - -```vcl -DRIFT DETECT FEDERATION -FROM verisim:federation -WHERE octad.id = @id -STORES [@university_a, @university_b]; -``` - -**Response:** - -```json -{ - "octad_id": "550e8400-e29b-41d4-a716-446655440000", - "federation_drift": true, - "stores": { - "university_a": { - "title": "Retracted Paper", - "retraction_status": "retracted", - "updated_at": "2025-01-15T10:00:00Z" - }, - "university_b": { - "title": "Active Paper", - "retraction_status": "active", - "updated_at": "2025-01-14T09:00:00Z" - } - }, - "conflict_type": "retraction_dispute", - "resolution": "governance_vote" -} -``` - -=== Byzantine Drift - -**Definition:** Malicious or faulty node provides inconsistent data. - -**Detection:** - -1. **Quorum-based verification** - Majority vote determines truth -2. **ZKP validation** - Verify cryptographic proofs -3. **Temporal consistency** - Check version history - -**Example:** - -``` -Stores A, B, C, D all report: title = "Active Paper" -Store E reports: title = "HACKED!!!" ❌ BYZANTINE - -Action: Exclude Store E from quorum, investigate tampering -``` - -**VCL query with Byzantine tolerance:** - -```vcl -FROM verisim:federation -WHERE octad.id = @id -WITH BYZANTINE TOLERANCE 1 -- Tolerate 1 faulty node -MIN QUORUM 3; -``` - -=== Federated Repair Coordination - -**Multi-party repair protocol:** - -1. **Leader election** - One store coordinates repair -2. **Consensus phase** - All stores vote on canonical value -3. **Propagation phase** - Leader broadcasts canonical value -4. **Verification phase** - All stores confirm consistency - -**VCL:** - -```vcl -DRIFT REPAIR FEDERATION -FROM verisim:federation -WHERE octad.id = @id -USING STRATEGY consensus -COORDINATOR @university_a; -``` - -== Drift History and Audit - -=== Query Drift History - -**View all drift events for a Octad:** - -```vcl -DRIFT HISTORY -FROM verisim:temporal -WHERE octad.id = @id -ORDER BY detected_at DESC -LIMIT 10; -``` - -**Response:** - -```json -{ - "octad_id": "550e8400-e29b-41d4-a716-446655440000", - "drift_events": [ - { - "detected_at": "2025-01-15T10:30:00Z", - "repaired_at": "2025-01-15T10:31:00Z", - "drift_type": "title_mismatch", - "affected_modalities": ["graph", "document"], - "repair_strategy": "latest_wins", - "canonical_value": "Machine Learning Paper" - }, - { - "detected_at": "2025-01-14T15:00:00Z", - "repaired_at": "2025-01-14T15:02:00Z", - "drift_type": "timestamp_mismatch", - "affected_modalities": ["vector", "semantic"], - "repair_strategy": "latest_wins", - "canonical_value": "2025-01-14T14:58:00Z" - } - ] -} -``` - -=== Drift Metrics Dashboard - -**System-wide drift statistics:** - -```vcl -DRIFT METRICS -FROM verisim:system -WHERE period = 'last_24_hours'; -``` - -**Response:** - -```json -{ - "period": "2025-01-14T10:00:00Z to 2025-01-15T10:00:00Z", - "total_octads": 100000, - "drifted_octads": 850, - "drift_rate": 0.0085, - "drift_rate_threshold": 0.05, - "status": "healthy", - "drift_by_type": { - "title_mismatch": 400, - "timestamp_mismatch": 350, - "type_mismatch": 100 - }, - "drift_by_modality": { - "graph": 300, - "vector": 200, - "document": 250, - "semantic": 100 - }, - "repairs_performed": 820, - "repairs_pending": 30, - "repair_success_rate": 0.96 -} -``` - -== Drift Handling Best Practices - -=== Prevention - -**1. Atomic Writes** - -Write to all modalities in a transaction: - -```rust -pub fn create_octad_atomic(octad: &Octad) -> Result { - let tx = begin_transaction()?; - - tx.write_graph(&octad)?; - tx.write_vector(&octad)?; - tx.write_document(&octad)?; - tx.write_semantic(&octad)?; - tx.write_temporal(&octad)?; - - tx.commit()?; - Ok(octad.id) -} -``` - -**2. Cache Invalidation** - -Invalidate caches immediately after writes: - -```elixir -def update_octad(octad) do - # Update all modalities - VeriSim.Graph.update(octad) - VeriSim.Vector.update(octad) - VeriSim.Document.update(octad) - - # Invalidate all caches - VeriSim.QueryCache.invalidate(octad.id) -end -``` - -**3. Version Vectors** - -Use version vectors to track causality: - -```json -{ - "octad_id": "550e8400-...", - "version_vector": { - "graph": 5, - "vector": 5, - "document": 4, // Out of sync! - "semantic": 5, - "temporal": 5 - } -} -``` - -=== Detection - -**1. Continuous Monitoring** - -Check for drift every 5 minutes (configurable): - -```elixir -# Scheduled via Quantum or similar -defmodule VeriSim.Scheduler do - def schedule_drift_detection do - every(5, :minutes, fn -> - VeriSim.DriftMonitor.detect_all_drift() - end) - end -end -``` - -**2. Write-Time Verification** - -Check for drift immediately after writes: - -```rust -pub fn update_octad_with_verification(octad: &Octad) -> Result<(), Error> { - write_all_modalities(octad)?; - - // Immediate drift check - let drift = detect_drift(&octad.id)?; - - if drift.has_drift { - return Err(Error::DriftDetectedAfterWrite(drift)); - } - - Ok(()) -} -``` - -=== Repair - -**1. Gradual Repair** - -Don't repair everything at once (avoid system overload): - -```elixir -def repair_drifted_octads_gradually do - drifted = list_drifted_octads(limit: 100) - - drifted - |> Enum.chunk_every(10) - |> Enum.each(fn chunk -> - Enum.each(chunk, &repair_octad/1) - Process.sleep(1000) # Rate limit - end) -end -``` - -**2. Repair During Off-Peak Hours** - -Schedule intensive repairs for low-traffic periods: - -```elixir -# Repair at 3 AM UTC -defmodule VeriSim.Scheduler do - def schedule_intensive_repair do - at("03:00", fn -> - VeriSim.DriftMonitor.repair_all_drift(strategy: :latest_wins) - end) - end -end -``` - -== Summary - -VeriSimDB provides comprehensive drift handling through: - -1. **Detection** - Automatic monitoring + on-demand VCL queries -2. **Repair** - Multiple strategies (latest_wins, quorum, manual, merge) -3. **Query Integration** - Drift-aware VCL with tolerance controls -4. **Federation** - Cross-org drift detection and Byzantine tolerance -5. **Audit** - Complete drift history in verisim-temporal - -**Key Design Principles:** - -- **Prefer prevention over repair** - Atomic writes, version vectors -- **Make drift visible** - Don't hide inconsistencies from users -- **Provide escape hatches** - Manual resolution when automation fails -- **Federation-first** - Drift is expected, not exceptional -- **Safety over speed** - Default to strict consistency, opt-in to tolerance diff --git a/verisimdb/docs/error-handling-strategy.adoc b/verisimdb/docs/error-handling-strategy.adoc deleted file mode 100644 index e90b5b62..00000000 --- a/verisimdb/docs/error-handling-strategy.adoc +++ /dev/null @@ -1,1259 +0,0 @@ -// SPDX-License-Identifier: CC-BY-SA-4.0 - -= VCL Error Handling Strategy -:toc: left -:toclevels: 3 -:sectnums: - -== Overview - -VeriSimDB's error handling provides **comprehensive, actionable feedback** across all failure modes while maintaining **type safety** and **recoverability**. - -**Key Principles:** - -1. **Structured Errors** - Type-safe error representations (ReScript variants) -2. **Helpful Messages** - Context-rich error descriptions with hints -3. **Graceful Degradation** - Partial results when possible, fallback strategies -4. **Audit Trail** - All errors logged to verisim-temporal for compliance -5. **Recovery Strategies** - Automatic retry, circuit breakers, compensating transactions - -== Error Taxonomy - -=== Five Error Categories - -[cols="1,2,1,2"] -|=== -|Category |When It Happens |Recoverable? |Example - -|**Parse Errors** -|Invalid VCL syntax -|No (client fix) -|`SELECT GRPH FROM ...` (typo) - -|**Type Errors** -|Dependent-type verification fails -|No (contract violation) -|ZKP proof generation failed - -|**Runtime Errors** -|Execution failures -|Sometimes -|Store unavailable, timeout - -|**Modality Errors** -|Modality-specific failures -|Sometimes -|Dimension mismatch (vector), cycle detected (graph) - -|**Federation Errors** -|Cross-org coordination failures -|Often (partial results) -|Remote store unreachable, consensus timeout -|=== - -== Parse Errors - -**Cause:** Invalid VCL syntax detected by parser - -**Characteristics:** - -- Caught early (before execution) -- Always non-recoverable (client must fix) -- Include line/column information -- Provide helpful hints - -=== Example: Unexpected Token - -[source,vcl] ----- -SELECT GRAPH, VECTRO FROM octad abc-123 - ^^^^^^ - typo here ----- - -**Error Message:** - -[source,text] ----- -Parse Error at 1:14-1:20: Expected 'VECTOR', 'TENSOR', 'SEMANTIC', 'DOCUMENT', 'TEMPORAL', found 'VECTRO' - Hint: Did you mean 'VECTOR'? ----- - -**ReScript Error Type:** - -[source,rescript] ----- -ParseError({ - kind: UnexpectedToken({ - expected: ["VECTOR", "TENSOR", "SEMANTIC", "DOCUMENT", "TEMPORAL"], - found: "VECTRO" - }), - span: { - start: {line: 1, column: 14, offset: 14}, - end_: {line: 1, column: 20, offset: 20} - }, - source: "SELECT GRAPH, VECTRO FROM octad abc-123", - hint: Some("Did you mean 'VECTOR'?") -}) ----- - -=== Common Parse Errors - -[cols="1,2,2"] -|=== -|Error |Query |Hint - -|**Missing FROM** -|`SELECT GRAPH WHERE ...` -|Every query needs a FROM clause - -|**Invalid Modality** -|`SELECT GRPH FROM ...` -|Valid: GRAPH, VECTOR, TENSOR, SEMANTIC, DOCUMENT, TEMPORAL - -|**Invalid Drift Policy** -|`WITH DRIFT IGNORE` -|Valid: STRICT, REPAIR, TOLERATE, LATEST - -|**Unterminated String** -|`SELECT GRAPH FROM octad "abc-123` -|Missing closing quote - -|**Invalid Proof Type** -|`PROOF VALIDITY(...)` -|Valid: EXISTENCE, CITATION, ACCESS, INTEGRITY, PROVENANCE -|=== - -== Type Errors - -**Cause:** Dependent-type verification fails (ZKP proof cannot be generated or verified) - -**Characteristics:** - -- Only occur in **dependent-type path** (not slipstream) -- Non-recoverable (contract violated or not found) -- Include contract name and reason -- Logged for compliance auditing - -=== Example: Contract Violation - -[source,vcl] ----- -SELECT SEMANTIC FROM octad abc-123 - PROOF ACCESS(EthicsApprovalContract) ----- - -**Error Message:** - -[source,text] ----- -Type Error [octad: abc-123, modality: SEMANTIC]: Contract 'EthicsApprovalContract' violated: ethics approval expired on 2025-11-01 - Context: Query requires valid ethics approval for medical data access ----- - -**ReScript Error Type:** - -[source,rescript] ----- -TypeError({ - kind: ContractViolation({ - contract: "EthicsApprovalContract", - reason: "ethics approval expired on 2025-11-01" - }), - octad_id: Some("abc-123"), - modality: Some("SEMANTIC"), - context: "Query requires valid ethics approval for medical data access" -}) ----- - -=== Common Type Errors - -[cols="1,2"] -|=== -|Error |Meaning - -|**ContractNotFound** -|The specified contract doesn't exist in the registry - -|**ContractViolation** -|Data doesn't satisfy contract requirements - -|**ProofGenerationFailed** -|ZKP proof couldn't be generated (insufficient witness) - -|**ProofVerificationFailed** -|ZKP proof verification failed (data tampered or invalid) - -|**TypeMismatch** -|Expected type doesn't match actual type - -|**CircularDependency** -|Contracts have circular dependencies (A depends on B depends on A) -|=== - -== Runtime Errors - -**Cause:** Failures during query execution - -**Characteristics:** - -- May be recoverable (retry, fallback) -- Include timestamp and query ID for debugging -- Circuit breaker pattern for failing stores -- Partial results may be available - -=== Example: Store Unavailable - -[source,vcl] ----- -SELECT GRAPH FROM store milvus-1 WHERE ... ----- - -**Error Message:** - -[source,text] ----- -Runtime Error [recoverable]: Store 'milvus-1' unavailable: connection timeout after 5000ms ----- - -**Recovery Strategy:** - -[source,elixir] ----- -case execute_with_retry(query, max_retries: 3) do - {:ok, result} -> {:ok, result} - {:error, {:store_unavailable, store_id}} -> - # Fallback to replica - execute_on_replica(query, store_id) -end ----- - -=== Recoverable vs Non-Recoverable - -[cols="1,1,2"] -|=== -|Error |Recoverable? |Strategy - -|**StoreUnavailable** -|✅ Yes -|Retry with backoff, fallback to replica - -|**QueryTimeout** -|✅ Yes -|Increase timeout, use cached result if available - -|**DriftDetected** -|✅ Yes -|Trigger repair, re-execute after repair - -|**PermissionDenied** -|❌ No -|Fail immediately, audit log - -|**ResourceExhausted** -|⚠️ Sometimes -|Clear cache, wait for resources, fail if persistent - -|**InvalidOctadId** -|❌ No -|Fail immediately, client error - -|**NetworkError** -|✅ Yes -|Retry with exponential backoff - -|**InternalError** -|❌ No -|Fail, log for investigation -|=== - -== Modality-Specific Errors - -Each modality has its own error types due to domain-specific constraints. - -=== Graph Errors - -[cols="1,2,1"] -|=== -|Error |Example |Recovery - -|**MalformedRDF** -|`` -|Sanitize input, validate before insert - -|**CycleDetected** -|`A → B → C → A` -|Use DAG constraint, limit traversal depth - -|**TraversalDepthExceeded** -|Depth > 10 -|Increase limit or optimize query -|=== - -=== Vector Errors - -[cols="1,2,1"] -|=== -|Error |Example |Recovery - -|**DimensionMismatch** -|Query: 768-dim, Stored: 1536-dim -|Re-embed with correct model - -|**InvalidDistanceMetric** -|`DISTANCE "euclidian"` (typo) -|Use valid metric: euclidean, cosine, dot - -|**ANNIndexUnavailable** -|HNSW index not built yet -|Wait for indexing, use brute-force search -|=== - -=== Semantic Errors - -[cols="1,2,1"] -|=== -|Error |Example |Recovery - -|**ZKPVerificationFailed** -|Proof doesn't match data -|Re-generate proof, investigate tampering - -|**WitnessGenerationFailed** -|Insufficient data for proof -|Enrich data, relax contract requirements - -|**ContractExpired** -|Contract valid until 2025-11-01 -|Renew contract, use new contract version -|=== - -=== Temporal Errors - -[cols="1,2,1"] -|=== -|Error |Example |Recovery - -|**VersionNotFound** -|`AS OF '2024-01-01'` (no data) -|Query earlier/later version, check date format - -|**MerkleVerificationFailed** -|Merkle proof invalid -|Investigate corruption, use backup - -|**TemporalConflict** -|Concurrent writes to same version -|Use conflict resolution policy -|=== - -== Federation Errors - -**Cause:** Failures in cross-organization coordination - -**Characteristics:** - -- Often result in **partial results** (some stores succeeded) -- May indicate Byzantine faults -- Require consensus timeout handling -- Access control across org boundaries - -=== Example: Partial Results - -[source,vcl] ----- -SELECT GRAPH FROM FEDERATION /universities/* LIMIT 100 ----- - -**Error Message:** - -[source,text] ----- -Federation Error [/universities/*]: Partial results: succeeded=[oxford, cambridge, mit], failed=[stanford, berkeley] ----- - -**Recovery Strategy:** - -[source,elixir] ----- -case execute_federated(query) do - {:ok, results} -> {:ok, results} - {:partial, results, failed_stores} -> - Logger.warn("Partial results from federation: failed=#{inspect(failed_stores)}") - - # Return what we have with warning - {:ok, %{ - data: results, - partial: true, - failed_stores: failed_stores, - warning: "Not all stores responded" - }} -end ----- - -=== Federation Error Handling Strategies - -[cols="1,2"] -|=== -|Error |Strategy - -|**RemoteStoreUnreachable** -|Timeout with backoff, exclude from quorum, retry later - -|**PartialResults** -|Return what succeeded + warning, log for monitoring - -|**CrossOrgAccessDenied** -|Fail immediately, audit log, notify org admin - -|**ByzantineFaultDetected** -|Isolate suspicious nodes, trigger investigation, use trusted subset - -|**ConsensusTimeout** -|Use partial quorum if safe, fail if critical (e.g., writes) - -|**FederationPolicyViolation** -|Fail, audit log, review federation agreements -|=== - -== Error Recovery Strategies - -=== 1. Retry with Exponential Backoff - -For transient failures (network errors, store unavailable): - -[source,elixir] ----- -defmodule VeriSim.ErrorRecovery do - def retry_with_backoff(func, opts \\ []) do - max_retries = Keyword.get(opts, :max_retries, 3) - base_delay_ms = Keyword.get(opts, :base_delay_ms, 100) - - retry(func, max_retries, base_delay_ms, 0) - end - - defp retry(func, max_retries, base_delay_ms, attempt) when attempt < max_retries do - case func.() do - {:ok, result} -> {:ok, result} - {:error, error} when is_recoverable?(error) -> - delay_ms = base_delay_ms * :math.pow(2, attempt) |> round() - Logger.debug("Retry attempt #{attempt + 1}/#{max_retries} after #{delay_ms}ms") - Process.sleep(delay_ms) - retry(func, max_retries, base_delay_ms, attempt + 1) - - {:error, error} -> - {:error, error} - end - end - - defp retry(_func, _max_retries, _base_delay_ms, _attempt) do - {:error, :max_retries_exceeded} - end - - defp is_recoverable?({:store_unavailable, _}), do: true - defp is_recoverable?({:network_error, _}), do: true - defp is_recoverable?({:timeout, _}), do: true - defp is_recoverable?(_), do: false -end ----- - -=== 2. Circuit Breaker Pattern - -Prevent cascading failures by stopping requests to failing stores: - -[source,elixir] ----- -defmodule VeriSim.CircuitBreaker do - use GenServer - - # States: :closed (normal), :open (failing), :half_open (testing recovery) - defstruct [ - state: :closed, - failure_count: 0, - failure_threshold: 5, - timeout_ms: 60_000, # 1 minute - last_failure: nil - ] - - def call_with_breaker(store_id, func) do - case get_state(store_id) do - :open -> - {:error, {:circuit_open, "Store #{store_id} circuit breaker is open"}} - - :half_open -> - case func.() do - {:ok, result} -> - record_success(store_id) - {:ok, result} - - {:error, _} = error -> - record_failure(store_id) - error - end - - :closed -> - case func.() do - {:ok, result} -> - {:ok, result} - - {:error, _} = error -> - record_failure(store_id) - error - end - end - end -end ----- - -=== 3. Fallback to Cached Results - -When fresh execution fails, use cached results if available: - -[source,elixir] ----- -def execute_with_fallback(query) do - case VeriSim.QueryRouter.execute(query) do - {:ok, result} -> {:ok, result} - - {:error, {:store_unavailable, _}} -> - # Try cache even if expired - case VeriSim.QueryCache.get(query, allow_stale: true) do - {:ok, cached} -> - Logger.warn("Using stale cached result due to store unavailability") - {:ok, %{cached | stale: true, warning: "Data may be outdated"}} - - {:error, :not_found} -> - {:error, :no_fallback_available} - end - end -end ----- - -=== 4. Partial Results Handling - -For federation queries, return what succeeded: - -[source,elixir] ----- -def execute_federated_with_partial(query) do - stores = get_stores_for_federation(query.federation_pattern) - - results = stores - |> Task.async_stream(fn store -> - execute_on_store(store, query) - end, timeout: query.timeout_ms) - |> Enum.to_list() - - succeeded = results - |> Enum.filter(fn - {:ok, {:ok, _}} -> true - _ -> false - end) - |> Enum.map(fn {:ok, {:ok, result}} -> result end) - - failed = results - |> Enum.filter(fn - {:ok, {:error, _}} -> true - {:exit, _} -> true - _ -> false - end) - - cond do - Enum.empty?(failed) -> - # All succeeded - {:ok, Enum.flat_map(succeeded, & &1.data)} - - Enum.empty?(succeeded) -> - # All failed - {:error, :all_stores_failed} - - length(succeeded) >= query.min_quorum -> - # Enough succeeded for partial results - {:partial, Enum.flat_map(succeeded, & &1.data), failed} - - true -> - # Not enough succeeded - {:error, {:insufficient_quorum, length(succeeded), query.min_quorum}} - end -end ----- - -=== 5. Compensating Transactions - -For mutations that partially fail, roll back changes: - -[source,elixir] ----- -def execute_mutation_with_compensation(mutation) do - # Track what we've done for potential rollback - saga = Saga.new(mutation.id) - - try do - # Step 1: Update GRAPH modality - {:ok, graph_result} = update_graph(mutation) - Saga.add_step(saga, :graph, graph_result, fn -> rollback_graph(graph_result) end) - - # Step 2: Update VECTOR modality - {:ok, vector_result} = update_vector(mutation) - Saga.add_step(saga, :vector, vector_result, fn -> rollback_vector(vector_result) end) - - # Step 3: Update TEMPORAL log - {:ok, temporal_result} = append_to_temporal(mutation) - - Saga.commit(saga) - {:ok, %{graph: graph_result, vector: vector_result, temporal: temporal_result}} - rescue - error -> - Logger.error("Mutation failed, rolling back: #{inspect(error)}") - Saga.rollback(saga) - {:error, {:mutation_failed, error}} - end -end ----- - -== Error Audit Trail - -**All errors are logged to verisim-temporal for compliance:** - -[source,elixir] ----- -defmodule VeriSim.ErrorLogger do - def log_error(error, context) do - entry = %{ - timestamp: DateTime.utc_now(), - error_code: VCLError.get_error_code(error), - error_message: VCLError.format(error), - query_id: context.query_id, - user_id: context.user_id, - octad_ids: context.octad_ids, - recoverable: VCLError.is_recoverable?(error), - recovery_attempted: context.recovery_attempted, - recovery_successful: context.recovery_successful - } - - VeriSim.Temporal.append_audit_log("errors", entry) - end -end ----- - -**Audit queries:** - -[source,vcl] ----- --- Find all errors for a user -SELECT TEMPORAL FROM audit_log -WHERE event_type = 'error' - AND user_id = 'alice' - AS OF 'last 7 days' - --- Find all non-recoverable errors -SELECT TEMPORAL FROM audit_log -WHERE event_type = 'error' - AND recoverable = false - AS OF 'last 30 days' - --- Find stores with high error rates -SELECT TEMPORAL FROM audit_log -WHERE event_type = 'error' - AND error_code LIKE 'VCL_STORE_%' -GROUP BY store_id -ORDER BY COUNT(*) DESC ----- - -== Error Response Format - -=== HTTP API Error Response - -[source,json] ----- -{ - "error": { - "code": "VCL_PARSE_ERROR", - "message": "Parse Error at 1:14-1:20: Expected 'VECTOR', found 'VECTRO'", - "hint": "Did you mean 'VECTOR'?", - "recoverable": false, - "details": { - "line": 1, - "column": 14, - "span": { - "start": 14, - "end": 20 - }, - "query": "SELECT GRAPH, VECTRO FROM octad abc-123" - } - }, - "request_id": "req_abc123", - "timestamp": "2026-01-22T15:30:00Z" -} ----- - -=== Elixir Error Tuples - -[source,elixir] ----- -# Success -{:ok, result} - -# Recoverable error -{:error, {:store_unavailable, "milvus-1"}} - -# Non-recoverable error -{:error, {:permission_denied, "User cannot access octad abc-123"}} - -# Partial success -{:partial, results, failed_stores} ----- - -=== ReScript Result Type - -[source,rescript] ----- -type queryResult<'a> = - | Ok('a) - | Error(vclError) - | Partial({data: 'a, failed: array}) ----- - -== Best Practices - -=== 1. Always Handle Errors Explicitly - -[source,elixir] ----- -# ❌ Bad: Ignoring errors -result = execute_query(query) -process(result) - -# ✅ Good: Explicit error handling -case execute_query(query) do - {:ok, result} -> process(result) - {:error, error} -> handle_error(error) -end ----- - -=== 2. Use Pattern Matching for Error Types - -[source,elixir] ----- -case execute_query(query) do - {:ok, result} -> - {:ok, result} - - {:error, {:store_unavailable, store_id}} -> - execute_on_replica(query, store_id) - - {:error, {:timeout, _}} -> - execute_with_cache_fallback(query) - - {:error, {:permission_denied, _}} -> - {:error, :unauthorized} - - {:error, error} -> - Logger.error("Unexpected error: #{inspect(error)}") - {:error, :internal_error} -end ----- - -=== 3. Log Errors with Context - -[source,elixir] ----- -Logger.error("Query execution failed", - error: inspect(error), - query_id: query.id, - user_id: query.user_id, - octad_ids: query.octad_ids, - duration_ms: duration, - retry_count: retry_count -) ----- - -=== 4. Fail Fast for Non-Recoverable Errors - -[source,elixir] ----- -case VCLParser.parse(query_string) do - {:ok, ast} -> execute(ast) - {:error, parse_error} -> - # Don't retry parse errors - {:error, parse_error} -end ----- - -=== 5. Set Reasonable Timeouts - -[source,elixir] ----- -# Per-modality timeouts -@timeout_config %{ - "GRAPH" => 5_000, # 5 seconds - "VECTOR" => 10_000, # 10 seconds (ANN search) - "SEMANTIC" => 30_000, # 30 seconds (ZKP generation) - "TEMPORAL" => 5_000 # 5 seconds -} ----- - -== Error Monitoring - -=== Metrics to Track - -[cols="1,2,1"] -|=== -|Metric |Description |Alert Threshold - -|**Error Rate** -|Errors / Total Requests -|> 5% - -|**Store Availability** -|% of successful store requests -|< 95% - -|**Retry Success Rate** -|% of retries that succeed -|< 50% - -|**Circuit Breaker Opens** -|Count of circuit breaker activations -|> 3 per hour - -|**Partial Result Rate** -|% of federation queries with partial results -|> 10% - -|**Error Recovery Time** -|Time to recover from errors -|> 5 minutes -|=== - -=== Alerting Rules - -[source,elixir] ----- -defmodule VeriSim.ErrorMonitor do - def check_error_thresholds do - stats = get_error_stats(last: 5.minutes) - - # Alert on high error rate - if stats.error_rate > 0.05 do - alert(:high_error_rate, "Error rate: #{stats.error_rate * 100}%") - end - - # Alert on store unavailability - Enum.each(stats.store_availability, fn {store_id, availability} -> - if availability < 0.95 do - alert(:store_unavailable, "Store #{store_id} availability: #{availability * 100}%") - end - end) - end -end ----- - -== Verbosity Levels - -VCL supports **four verbosity levels** for error and diagnostic output: - -[cols="1,2,3"] -|=== -|Level |Use Case |Output - -|**`SILENT`** -|Production batch jobs, CI/CD -|No output except fatal errors - -|**`NORMAL`** (default) -|Interactive queries, REPL -|Errors + warnings + query results - -|**`VERBOSE`** -|Debugging, development -|Errors + warnings + execution plan + timing - -|**`DEBUG`** -|System debugging, troubleshooting -|Full trace: parse tree, type checking, store calls, drift detection -|=== - -=== Configuration - -**Command-line:** -[source,bash] ----- -verisim query --verbosity=verbose query.vcl -verisim query --verbosity=silent query.vcl # CI/CD mode ----- - -**VCL pragma:** -[source,vcl] ----- -SET VERBOSITY VERBOSE; - -SELECT GRAPH, DOCUMENT -FROM verisim:semantic -WHERE octad.types INCLUDES "Paper" -LIMIT 10; ----- - -**Environment variable:** -[source,bash] ----- -export VERISIM_VERBOSITY=debug -verisim query query.vcl ----- - -=== Verbosity Output Examples - -==== SILENT Mode - -[source,text] ----- -# No output unless fatal error -# Query succeeded: 0 bytes output ----- - -==== NORMAL Mode - -[source,text] ----- -Warning: Store 'university-archive' cache expired, refreshing... -Query returned 10 results in 234ms ----- - -==== VERBOSE Mode - -[source,text] ----- -[PLAN] Query plan selected: federated-quorum (3 stores) -[EXEC] Querying store 1/3: university-archive -[EXEC] Querying store 2/3: research-lab -[EXEC] Querying store 3/3: corporate-db -[DRIFT] Detected title mismatch: "ML Paper" vs "Machine Learning Paper" -[DRIFT] Repair strategy: quorum (2/3 agree on "Machine Learning Paper") -[RESULT] Query returned 10 results in 1.2s -[TIMING] Parse: 5ms, Type check: 12ms, Execute: 1183ms ----- - -==== DEBUG Mode - -[source,text] ----- -[PARSE] Tokens: [SELECT, GRAPH, COMMA, DOCUMENT, FROM, ...] -[PARSE] AST: - Query { - modalities: [GRAPH, DOCUMENT], - source: Source::Semantic("verisim:semantic"), - condition: Some(Condition::Includes(...)) - } -[TYPE] Checking modalities: GRAPH, DOCUMENT -[TYPE] Octad type: {h : Octad | Graph(h) ≠ None ∧ Document(h) ≠ None} -[TYPE] Condition type: Octad → Bool -[TYPE] Query type: QueryResult[{GRAPH, DOCUMENT}] -[EXEC] HTTP GET https://university-archive.edu/verisim/query -[EXEC] Response: 200 OK, 3 octads -[EXEC] HTTP GET https://research-lab.org/verisim/query -[EXEC] Response: 200 OK, 5 octads -[EXEC] HTTP GET https://corporate-db.com/verisim/query -[EXEC] Response: 504 Gateway Timeout (retry 1/3) -[EXEC] Response: 200 OK, 2 octads -[DRIFT] Store 1 title: "ML Paper" -[DRIFT] Store 2 title: "Machine Learning Paper" -[DRIFT] Store 3 title: "Machine Learning Paper" -[DRIFT] Quorum result: "Machine Learning Paper" (2/3) -[RESULT] Merged 10 octads with drift repair -[TIMING] Total: 1.2s (Parse: 5ms, Type: 12ms, Execute: 1183ms) ----- - -== Friendly Notices and Warnings - -Beyond errors, VCL provides **helpful notices** for non-error situations: - -=== Notice Types - -[cols="1,2,3"] -|=== -|Type |When Shown |Example - -|**Info** -|Informational messages -|"Cache hit for query XYZ (saved 2.3s)" - -|**Warning** -|Potential issues, not errors -|"Query uses deprecated syntax (VERSION 0.9)" - -|**Hint** -|Performance/style suggestions -|"Consider adding LIMIT clause for large result sets" - -|**Deprecation** -|Features scheduled for removal -|"FULLTEXT CONTAINS is deprecated, use FULLTEXT MATCHES instead (removal: v2.0)" -|=== - -=== Example Notices - -==== Cache Hit (Info) - -[source,text] ----- -ℹ Cache hit for query SHA256:a1b2c3... (saved 1.8s) -Query returned 10 results in 23ms ----- - -==== Deprecated Syntax (Warning) - -[source,vcl] ----- -VERSION 0.9; -SELECT * FROM HEXAD abc-123; ----- - -[source,text] ----- -⚠ Warning: VERSION 0.9 is deprecated (current: 1.0, removal: v2.0) - Migration guide: https://verisimdb.org/migration/v0.9-to-v1.0 ----- - -==== Missing LIMIT (Hint) - -[source,vcl] ----- -SELECT GRAPH, DOCUMENT -FROM FEDERATION /universities/* -WHERE octad.types INCLUDES "Paper"; ----- - -[source,text] ----- -💡 Hint: Federated query without LIMIT may return large result sets - Consider adding: LIMIT 100 ----- - -==== Slow Query (Warning) - -[source,text] ----- -⚠ Warning: Query took 5.2s (threshold: 1s) - Consider adding modality-specific filters to reduce result set ----- - -=== Configuration - -**Suppress warnings:** -[source,bash] ----- -verisim query --no-warnings query.vcl ----- - -**Suppress hints:** -[source,bash] ----- -verisim query --no-hints query.vcl ----- - -**Strict mode (warnings as errors):** -[source,bash] ----- -verisim query --strict query.vcl -# Exit code 1 on any warning ----- - -== Learning Hooks (v3 - miniKanren Integration) - -**Status:** STUB - Planned for v3.0 - -VCL error handling will integrate with **miniKanren** (link:minikanren-integration-v3.adoc[see roadmap]) to: - -1. **Learn error patterns** - Synthesize predicates from error examples -2. **Suggest fixes** - Generate correction rules from observed errors -3. **Optimize recovery** - Infer optimal retry/fallback strategies - -=== Learning Hook Architecture (v3) - -[source,text] ----- -┌─────────────────────────────────────────────────────────┐ -│ Error Handling Pipeline │ -│ ├── Error Detection (current v1) │ -│ ├── Error Logging (current v1) │ -│ ├── Error Recovery (current v1) │ -│ └── Error Learning (v3 - miniKanren) ◄─ NEW │ -│ ↓ │ -├─────────────────────────────────────────────────────────┤ -│ miniKanren Learning Engine │ -│ ├── Pattern Synthesis (learn from examples) │ -│ ├── Fix Generation (suggest corrections) │ -│ └── Strategy Optimization (infer recovery policies) │ -└─────────────────────────────────────────────────────────┘ ----- - -=== Use Case 1: Pattern Synthesis (v3) - -**Problem:** Detect common error patterns across queries - -**miniKanren approach:** - -[source,scheme] ----- -;; Learn error predicate from positive/negative examples -(defrel (error-patterno examples pattern) - (fresh (error-examples success-examples) - (partition-errorso examples error-examples success-examples) - (pattern-matches-allo pattern error-examples) - (pattern-rejects-allo pattern success-examples) - (pattern-is-minimalo pattern))) - -;; Example: Learn that missing modality causes errors -(run 1 (pattern) - (error-patterno - '((query "SELECT FROM octad abc" . error) - (query "SELECT GRAPH FROM octad abc" . ok) - (query "SELECT FROM store xyz" . error)) - pattern)) -;; Output: (pattern (missing-modality-in-select)) ----- - -=== Use Case 2: Fix Generation (v3) - -**Problem:** Suggest corrections for common syntax errors - -**miniKanren approach:** - -[source,scheme] ----- -;; Relational specification of fixes -(defrel (fix-suggestiono error-query fixed-query) - (conde - ;; Fix 1: Add missing modality - [(missing-modalityo error-query) - (add-default-modalityo error-query fixed-query)] - - ;; Fix 2: Correct typo - [(typoo error-query field) - (correct-typoo error-query field fixed-query)] - - ;; Fix 3: Add missing LIMIT - [(unbounded-queryo error-query) - (add-limito error-query 100 fixed-query)])) - -;; Query: What fix for this error? -(run* (fixed) - (fix-suggestiono - "SELECT GRPH FROM octad abc" ;; Typo: GRPH - fixed)) -;; Output: ("SELECT GRAPH FROM octad abc") ----- - -=== Use Case 3: Strategy Optimization (v3) - -**Problem:** Given error history, infer optimal retry/fallback strategy - -**miniKanren approach:** - -[source,scheme] ----- -;; Learn retry strategy from error history -(defrel (retry-strategyo error-history strategy) - (fresh (error-type frequency transient?) - (classify-errorso error-history error-type frequency) - (transient-erroro? error-type transient?) - (conde - ;; High-frequency transient: exponential backoff - [(>o frequency 0.3) - (== transient? #t) - (== strategy 'exponential-backoff)] - - ;; Low-frequency transient: simple retry - [( {:ok, ast} - {:error, parse_error} -> - # v3: Ask miniKanren for fix suggestions - suggested_fixes = MiniKanren.suggest_fixes(parse_error, query_string) - - {:error, parse_error, suggestions: suggested_fixes} - end - end -end ----- - -**Output example (v3):** - -[source,text] ----- -Parse Error at 1:8: Expected 'GRAPH', 'VECTOR', ..., found 'GRPH' - Hint: Did you mean 'GRAPH'? - Suggested fix (confidence: 95%): - SELECT GRAPH FROM HEXAD abc-123 - ^^^^^ - [miniKanren learned from 47 similar errors] ----- - -== Summary - -**Error Handling Guarantees:** - -✅ **Structured Errors** - Type-safe error representation across all languages (ReScript, Elixir) - -✅ **Helpful Messages** - Context-rich errors with hints for common mistakes - -✅ **Graceful Degradation** - Partial results, cache fallback, retry with backoff - -✅ **Audit Trail** - All errors logged to verisim-temporal for compliance - -✅ **Recovery Strategies** - Circuit breakers, compensating transactions, quorum-based consensus - -✅ **Monitoring** - Real-time error metrics, alerting on thresholds - -**Error Recovery Time:** - -- Transient errors: < 1 second (retry with backoff) -- Store unavailable: < 5 seconds (replica fallback) -- Federation timeout: < 30 seconds (partial results) -- Circuit breaker recovery: < 60 seconds (half-open test) - -== References - -- link:vcl-architecture.adoc[VCL Architecture] -- link:caching-strategy.adoc[Caching Strategy] -- link:reversibility-design.adoc[Reversibility Design] -- link:../src/vcl/VCLError.res[VCL Error Types] -- link:../lib/verisim/error_recovery.ex[Error Recovery Implementation] diff --git a/verisimdb/docs/federation-readiness.adoc b/verisimdb/docs/federation-readiness.adoc deleted file mode 100644 index 9f215873..00000000 --- a/verisimdb/docs/federation-readiness.adoc +++ /dev/null @@ -1,167 +0,0 @@ -// SPDX-License-Identifier: CC-BY-SA-4.0 -// Copyright (c) 2026 Jonathan D.A. Jewell (hyperpolymath) - -= Federation Readiness Assessment -:toc: left -:toclevels: 3 -:sectnums: - -== What Federation Means in VeriSimDB - -Federation is the ability for multiple VeriSimDB instances to coordinate queries across their respective octad stores. Each instance maintains its own six modality stores (Graph, Vector, Tensor, Semantic, Document, Temporal) and its own set of octad entities. A federated query spans multiple instances, retrieving and correlating data from peers that may be geographically distributed, operated by different organizations, or running different versions of VeriSimDB. - -Federation in VeriSimDB is distinct from database replication. Replicated databases maintain identical copies of the same data. Federated VeriSimDB instances hold _different_ octad entities that may reference each other through graph edges, share semantic vocabularies, or have vector embeddings in the same latent space. The federation layer coordinates cross-instance queries while each instance retains sovereignty over its own data. - -The VCL query language supports federation natively: - -[source,vcl] ----- -SELECT GRAPH, SEMANTIC -FROM FEDERATION cross_org_research -WHERE (h)-[:CITES]->(target) -DRIFT POLICY strict_consistency ----- - -This query targets a named federation (`cross_org_research`), requests graph and semantic modalities, traverses citation edges that may span instances, and enforces a drift policy governing how cross-instance consistency is handled. - -== What Works Today - -=== API Routes for Peer Registration - -The `verisim-api` Rust crate includes API endpoints for federation peer management: - -* `POST /api/federation/peers` -- Register a new peer instance with its endpoint URL and capabilities. -* `GET /api/federation/peers` -- List all registered peers. -* `DELETE /api/federation/peers/:id` -- Deregister a peer. - -These routes accept and return JSON. Peer metadata includes the instance URL, supported modalities, version, and last-seen timestamp. Peer registration works and persists across restarts. - -=== VCL Parser Accepts FEDERATION Queries - -The VCL parser (ReScript) correctly parses `FROM FEDERATION ` clauses and `DRIFT POLICY ` directives. These produce valid AST nodes that flow through the query router. The parser also handles federation-scoped `WHERE` clauses and multi-modality `SELECT` statements against federation targets. - -=== Hypatia FileExecutor Handles FEDERATION Queries - -Hypatia's `FileExecutor` module can execute `FEDERATION` queries against local flat files (the verisimdb-data git-backed store). It performs cross-store matching by loading scan results from multiple repositories and correlating them by entity identifiers. This is a file-based simulation of federation, not true distributed query execution, but it validates the query structure and result format. - -== What VeriSimDB Loses Without VCL - -If VCL were removed or not used, VeriSimDB would lose the following federation-critical capabilities: - -* **No structured query language for multi-store coordination.** Without VCL, federated queries would require hand-written API calls to each peer, manual result merging, and no unified query plan. VCL provides a single declarative interface that the federation layer decomposes into per-peer sub-queries. - -* **No drift policy enforcement on federated queries.** Drift policies (`DRIFT POLICY strict_consistency`, `DRIFT POLICY eventual`, etc.) are expressed in VCL and enforced by the query router. Without VCL, there is no standard mechanism to declare or enforce consistency requirements across federated peers. - -* **No cross-modal consistency checks across instances.** VCL's `PROOF CONSISTENCY` clause can request verification that graph, vector, and semantic modalities agree across federation peers. Without VCL-UT, there is no way to express or verify cross-instance consistency at query time. - -* **No unified result format.** VCL queries return results in a consistent format regardless of whether the data came from one instance or twenty. Without VCL, each peer's API response format would need to be handled individually. - -== Current Gaps - -=== Normalizer Regeneration Strategies Are Stubs - -The `verisim-normalizer` crate's regeneration strategies currently return hardcoded `[regenerated]` placeholder values. When drift is detected and normalization is triggered, the normalizer identifies the authoritative modality and calls the appropriate regeneration strategy, but that strategy does not actually regenerate the drifted modality's data. Each of the six modalities needs a real regeneration implementation: - -* **Graph**: Re-derive edges from document content and semantic annotations. -* **Vector**: Re-embed from document text using the configured embedding model. -* **Tensor**: Re-compute tensor representation from source data. -* **Semantic**: Re-extract type annotations and RDF triples from document and graph. -* **Document**: Re-generate searchable text from graph, semantic, and temporal data. -* **Temporal**: Re-build version chain from modification history. - -=== Federation Executor Always Returns Empty - -The federation executor in the Elixir orchestration layer (`VeriSim.FederationExecutor`) receives a parsed federation query and should decompose it into per-peer sub-queries, dispatch them, and merge results. Currently it always returns `{:ok, []}` regardless of the query or registered peers. This is a placeholder implementation. - -=== Federation Resolver Peer Queries Unimplemented - -Peers can be registered and listed, but the resolver cannot actually _query_ a registered peer. The `VeriSim.FederationResolver` module exists but its `query_peer/3` function is unimplemented. Calling it returns `{:error, :not_implemented}`. - -=== No Consensus Protocol for Cross-Instance Writes - -Federation currently has no mechanism for coordinated writes across instances. If entity A on instance 1 references entity B on instance 2 through a graph edge, and entity B is modified, there is no protocol to notify instance 1 or maintain referential integrity. The Raft consensus design (documented in link:challenges-federated.adoc[Challenges: Federated Deployment]) is specified but not implemented. - -=== Drift Auto-Trigger Missing - -Drift detection works: the `verisim-drift` crate computes drift scores across modalities, and the Elixir `DriftMonitor` can evaluate them against thresholds. However, drift detection and repair must be triggered manually (via API call or Elixir function invocation). There is no automatic trigger that runs on a schedule or fires on write events. In a federation context, this means cross-instance drift can accumulate undetected until someone manually checks. - -== Readiness Matrix - -[cols="2,1,3",options="header"] -|=== -|Component |Readiness |Notes - -|**API (peer registration)** -|Ready -|Endpoints work, persistence confirmed. Peer metadata stored and retrievable. - -|**VCL Parser (FEDERATION syntax)** -|Ready -|Parses `FROM FEDERATION`, `DRIFT POLICY`, and federation-scoped queries correctly. Produces valid AST. - -|**Federation Executor** -|Stub -|Always returns `{:ok, []}`. No query decomposition, dispatch, or result merging implemented. - -|**Federation Resolver (peer queries)** -|Stub -|Cannot query registered peers. `query_peer/3` returns `{:error, :not_implemented}`. - -|**Normalizer (regeneration)** -|Stub -|Returns hardcoded `[regenerated]`. Six modality-specific strategies needed. - -|**Consensus Protocol** -|Design only -|Raft-based design documented in challenges-federated.adoc. No implementation exists. - -|**Drift Detection** -|Working -|Computes drift scores, evaluates thresholds. Functional but manual-trigger only. - -|**Drift Auto-Trigger** -|Missing -|No scheduled or event-driven drift detection. Manual only. - -|**Cross-Instance Consistency** -|Missing -|No mechanism to verify or enforce consistency across federation peers. - -|**VCL-UT PROOF over Federation** -|Missing -|PROOF clauses parse but do not generate real proofs, even for local queries. Federation adds no additional proof capability. -|=== - -== Roadmap to Production Federation - -=== Phase 1: Local Federation Simulation - -**Goal:** Validate the full query lifecycle against multiple local stores without network coordination. - -* Implement `FederationExecutor.execute/2` to decompose queries and dispatch to local store partitions. -* Implement `FederationResolver.query_peer/3` for local-only peers (same instance, different store namespaces). -* Add integration tests that create two logical "instances" in a single process and run cross-instance queries. -* Wire drift auto-trigger to run on a configurable interval (GenServer timer in Elixir). -* Implement at least one real normalizer regeneration strategy (Document is the most straightforward). - -=== Phase 2: Network Federation - -**Goal:** Federated queries across actual network boundaries. - -* Implement HTTP-based peer query protocol. `FederationResolver.query_peer/3` makes HTTP calls to peer endpoints. -* Add authentication and authorization for peer-to-peer communication (mTLS or signed requests). -* Implement query result merging in `FederationExecutor` with configurable merge strategies (union, intersection, ranked). -* Add timeout and fallback handling for unreachable peers. -* Implement remaining normalizer regeneration strategies for all six modalities. -* Add federation-scoped drift detection: compute drift across instances, not just within one. - -=== Phase 3: Verified Federation - -**Goal:** VCL-UT proofs work across federation boundaries. - -* Wire Lean type checker into the VCL-UT execution path (prerequisite: VCL-UT Phase 3 from link:vcl-vs-vcl-dt.adoc[VCL Slipstream vs VCL-UT]). -* Implement proof witness collection across peers: each peer contributes its portion of the proof. -* Implement Raft consensus for coordinated writes (cross-instance referential integrity). -* Implement `PROOF CONSISTENCY` verification across federation peers (cross-instance modal agreement). -* Add sanctify ZKP integration for privacy-preserving proofs in multi-tenant federations. -* Performance optimization: proof caching, incremental proof updates, parallel peer queries. diff --git a/verisimdb/docs/getting-started.adoc b/verisimdb/docs/getting-started.adoc deleted file mode 100644 index c182b90e..00000000 --- a/verisimdb/docs/getting-started.adoc +++ /dev/null @@ -1,583 +0,0 @@ -// SPDX-License-Identifier: CC-BY-SA-4.0 -= VeriSimDB Getting Started Guide -Jonathan D.A. Jewell -:toc: left -:toclevels: 3 -:icons: font -:source-highlighter: rouge - -== What Is VeriSimDB? - -VeriSimDB is a *cross-system entity consistency engine* with drift detection -and self-normalisation. It treats every entity as an **octad** -- eight -simultaneous representations (modalities) that are continuously monitored -for consistency: - -[cols="1,3"] -|=== -| Modality | What It Stores - -| *Graph* | RDF triples, property graph edges, relationships -| *Vector* | Embedding vectors for similarity search (HNSW) -| *Tensor* | Multi-dimensional numerical arrays (ndarray/Burn) -| *Semantic* | Type annotations, ontology terms, proof blobs (CBOR) -| *Document* | Full-text searchable content (Tantivy) -| *Temporal* | Version history, time-series data -| *Provenance* | Origin tracking, transformation chains, actor trails -| *Spatial* | Geospatial coordinates, geometries (R-tree) -|=== - -When any modality drifts from the others -- an embedding no longer matches -its source text, a graph edge references a deleted document, a provenance -chain's hash integrity breaks -- VeriSimDB detects it and can automatically -repair it. - -=== Architecture Overview - -[source] ----- -┌──────────────────────────────────────────────────────┐ -│ Elixir/OTP Orchestration Layer │ -│ ├── EntityServer (GenServer per entity) │ -│ ├── DriftMonitor (continuous consistency check) │ -│ ├── QueryRouter (distributes VCL queries) │ -│ ├── SchemaRegistry (type system coordination) │ -│ └── Federation (heterogeneous peer queries) │ -│ ↕ HTTP / NIF │ -├──────────────────────────────────────────────────────┤ -│ Rust Core Engine │ -│ ├── 8 modality stores (graph, vector, tensor...) │ -│ ├── Drift detection (per-modality scoring) │ -│ ├── Normaliser (5 regeneration strategies) │ -│ ├── WAL (write-ahead log for durability) │ -│ └── HTTP API (Axum, TLS, IPv6, Prometheus) │ -└──────────────────────────────────────────────────────┘ ----- - -== Prerequisites - -=== System Requirements - -* **Rust** 1.80+ (stable) -* **Elixir** 1.17+ with OTP 27+ -* **Podman** (for container deployment) or Docker - -=== Optional Dependencies - -* **Deno** 2.0+ (for VCL parser bridge -- falls back to built-in parser) -* **pgvector**, **PostGIS** (if federating with PostgreSQL) - -== Installation - -=== From Source - -[source,bash] ----- -# Clone the repository -git clone https://gitlab.com/hyperpolymath/verisimdb.git -cd verisimdb - -# Build the Rust core -cargo build --release - -# Set up the Elixir orchestration layer -cd elixir-orchestration -mix deps.get -mix compile - -# Start the Rust API server -cd .. && cargo run --release -p verisim-api & - -# Start the Elixir orchestration -cd elixir-orchestration && mix run --no-halt ----- - -=== Container Deployment - -[source,bash] ----- -# Build the container image (in-memory, default) -podman build -t verisimdb:latest -f container/Containerfile . - -# Run VeriSimDB (in-memory -- data lost on restart) -podman run -d \ - --name verisimdb \ - -p 8080:8080 \ - verisimdb:latest - -# Verify it's running -curl http://localhost:8080/api/v1/health ----- - -=== Persistent Storage - -To persist data across restarts, build with the `persistent` feature: - -[source,bash] ----- -# Build with persistent storage (redb graph + file-backed Tantivy + WAL) -podman build -t verisimdb:persistent \ - --build-arg FEATURES=persistent \ - -f container/Containerfile . - -# Run with a named volume for data persistence -podman run -d \ - --name verisimdb \ - -p 8080:8080 \ - -v verisimdb-data:/data \ - verisimdb:persistent ----- - -Persistent mode stores data at `VERISIM_PERSISTENCE_DIR` (defaults to `/data` -in the container, `/var/lib/verisimdb` outside containers): - -* `graph.redb` -- redb B-tree database for graph triples (pure Rust, ACID) -* `documents/` -- Tantivy full-text index (mmap-backed) -* `wal/` -- Write-ahead log for crash recovery - -Other modalities (vector, tensor, semantic, temporal, provenance, spatial) -remain in-memory. Graph and document persistence covers the two most -query-intensive modalities. - -=== Verified Container Deployment (stapeln) - -For supply-chain-verified deployment using the -link:https://github.com/hyperpolymath/stapeln[stapeln] container ecosystem: - -[source,bash] ----- -# Build and sign as a .ctp (Cerro Torre Package) bundle -cd container && ./ct-build.sh persistent --push - -# Deploy the full stack with selur-compose -selur-compose verify # Verify all .ctp signatures -selur-compose up --detach # Start: rust-core + elixir + svalinn -selur-compose ps # Check status -selur-compose logs -f rust-core # Stream logs ----- - -The `compose.toml` orchestrates three services: - -* **rust-core** -- Modality stores, drift detection, HTTP/gRPC API (port 8080) -* **elixir-orchestration** -- Entity servers, drift monitor, federation (port 4000) -* **svalinn** -- Edge gateway with JWT auth, rate limiting, policy enforcement (port 443) - -The `.gatekeeper.yaml` policy controls authentication, rate limits, and -trust requirements. All `.ctp` bundles are cryptographically signed via -cerro-torre (Ed25519) and verified before deployment. - -== Quick Start: Your First Octad - -A **octad** (historical name; now an octad) is an entity with up to 8 -modality representations. Let's create one. - -=== Step 1: Create an Entity - -[source,bash] ----- -curl -X POST http://localhost:8080/api/v1/octads \ - -H "Content-Type: application/json" \ - -d '{ - "title": "Douglas Adams", - "body": "English author, best known for The Hitchhiker'\''s Guide to the Galaxy", - "types": ["https://schema.org/Person", "https://schema.org/Author"], - "embedding": [0.12, -0.34, 0.56, 0.78, -0.91, 0.23, -0.45, 0.67], - "relationships": [ - {"predicate": "wrote", "object": "hitchhikers-guide"}, - {"predicate": "bornIn", "object": "cambridge-uk"} - ] - }' ----- - -Response: - -[source,json] ----- -{ - "id": "a1b2c3d4-e5f6-4a7b-8c9d-0e1f2a3b4c5d", - "status": "created", - "modalities": { - "document": true, - "vector": true, - "graph": true, - "semantic": true, - "tensor": false, - "temporal": true, - "provenance": true, - "spatial": false - } -} ----- - -=== Step 2: Retrieve the Entity - -[source,bash] ----- -curl http://localhost:8080/api/v1/octads/a1b2c3d4-e5f6-4a7b-8c9d-0e1f2a3b4c5d ----- - -=== Step 3: Check Drift Status - -[source,bash] ----- -curl http://localhost:8080/api/v1/drift/entity/a1b2c3d4-e5f6-4a7b-8c9d-0e1f2a3b4c5d ----- - -Response: - -[source,json] ----- -{ - "entity_id": "a1b2c3d4-e5f6-4a7b-8c9d-0e1f2a3b4c5d", - "graph": 0.0, - "vector": 0.0, - "tensor": 0.0, - "semantic": 0.0, - "document": 0.0, - "temporal": 0.0, - "provenance": 0.0, - "spatial": 0.0, - "overall": 0.0 -} ----- - -All zeros -- no drift detected. The entity is consistent across all -populated modalities. - -=== Step 4: Simulate Drift - -Now deliberately break consistency by updating the embedding without -updating the document: - -[source,bash] ----- -# Update ONLY the vector embedding (not the document text) -curl -X PUT http://localhost:8080/api/v1/octads/a1b2c3d4-e5f6-4a7b-8c9d-0e1f2a3b4c5d \ - -H "Content-Type: application/json" \ - -d '{ - "embedding": [0.99, 0.99, 0.99, 0.99, 0.99, 0.99, 0.99, 0.99] - }' - -# Check drift again -curl http://localhost:8080/api/v1/drift/entity/a1b2c3d4-e5f6-4a7b-8c9d-0e1f2a3b4c5d ----- - -Response: - -[source,json] ----- -{ - "entity_id": "a1b2c3d4-e5f6-4a7b-8c9d-0e1f2a3b4c5d", - "vector": 0.87, - "document": 0.0, - "overall": 0.43, - "...": "..." -} ----- - -The vector modality is now 87% drifted from the document. VeriSimDB -detected that the embedding no longer corresponds to the text. - -=== Step 5: Trigger Self-Normalisation - -[source,bash] ----- -curl -X POST http://localhost:8080/api/v1/normalizer/trigger/a1b2c3d4-e5f6-4a7b-8c9d-0e1f2a3b4c5d ----- - -VeriSimDB identifies the document as the authoritative modality and -regenerates the vector embedding to match. Re-checking drift will -show all scores back near zero. - -== Querying with VCL - -VeriSimDB uses its own query language, *VCL* (VeriSim Consonance Language), -designed for cross-modal queries. VCL is not SQL -- it natively -understands modalities and drift. - -=== Basic Queries - -[source,sql] ----- --- Get all modalities for an entity -SELECT * FROM HEXAD 'a1b2c3d4-e5f6-4a7b-8c9d-0e1f2a3b4c5d' - --- Get specific modalities -SELECT GRAPH.*, DOCUMENT.*, VECTOR.* FROM HEXAD 'entity-id' - --- Full-text search across documents -SELECT DOCUMENT.* FROM FEDERATION /* WHERE DOCUMENT CONTAINS 'galaxy' - --- Vector similarity search -SELECT VECTOR.* FROM FEDERATION /* WHERE VECTOR SIMILAR TO [0.1, 0.2, ...] - --- Drift-aware query -SELECT * FROM FEDERATION /* WHERE DRIFT(VECTOR, DOCUMENT) > 0.3 ----- - -=== Cross-Modal Conditions - -[source,sql] ----- --- Find entities where graph structure contradicts document -SELECT GRAPH.*, DOCUMENT.* -FROM FEDERATION /* -WHERE DRIFT(GRAPH, DOCUMENT) > 0.5 - --- Modality existence checks -SELECT * FROM FEDERATION /* -WHERE PROVENANCE EXISTS AND SPATIAL EXISTS - --- Cross-modal field comparison -SELECT * FROM FEDERATION /* -WHERE DOCUMENT.severity > GRAPH.importance ----- - -=== VCL-UT (Formally Verified Queries) - -VCL-UT extends VCL with proof certificates. Every result includes a -cryptographic proof that the data satisfies stated constraints: - -[source,sql] ----- -SELECT DOCUMENT.*, GRAPH.* -FROM HEXAD 'entity-id' -PROOF EXISTENCE(entity-id), PROVENANCE(entity-id) ----- - -NOTE: VCL-UT proof generation is under active development. The syntax -parses correctly and proof obligations are extracted, but end-to-end -verifiable certificates require the Lean type checker integration -(planned for Q2 2026). - -== Federation - -VeriSimDB can federate queries across heterogeneous databases. A single -VCL query can transparently route to PostgreSQL, ArangoDB, Elasticsearch, -and other VeriSimDB instances. - -=== Registering Peers - -[source,elixir] ----- -# Register another VeriSimDB instance -VeriSim.Federation.Resolver.register_peer( - "verisim-prod", - "http://verisim-prod:8080/api/v1", - [:graph, :vector, :document] -) - -# Register an ArangoDB instance -VeriSim.Federation.Resolver.register_peer("arango-archive", %{ - endpoint: "http://arango.internal:8529", - adapter_type: :arangodb, - adapter_config: %{ - database: "research_archive", - collection: "entities", - graph_name: "entity_graph", - auth: {:basic, "readonly", "password"} - }, - modalities: [:graph, :document, :semantic, :spatial] -}) - -# Register a PostgreSQL instance with pgvector -VeriSim.Federation.Resolver.register_peer("pg-analytics", %{ - endpoint: "http://postgrest.internal:3000", - adapter_type: :postgresql, - adapter_config: %{ - database: "analytics", - table: "entities", - schema: "public", - extensions: [:pgvector, :postgis], - auth: {:basic, "reader", "password"} - }, - modalities: [:document, :vector, :spatial, :temporal] -}) - -# Register an Elasticsearch cluster -VeriSim.Federation.Resolver.register_peer("es-search", %{ - endpoint: "http://elastic.internal:9200", - adapter_type: :elasticsearch, - adapter_config: %{ - index: "entities", - version: 8, - auth: {:basic, "elastic", "password"} - }, - modalities: [:document, :vector] -}) ----- - -=== Federated Queries - -[source,sql] ----- --- Query all registered peers -SELECT DOCUMENT.* FROM FEDERATION /* WHERE DOCUMENT CONTAINS 'climate data' - --- Query only production peers -SELECT * FROM FEDERATION /prod/* WHERE VECTOR SIMILAR TO [0.1, ...] - --- Strict drift policy: exclude untrusted peers -SELECT * FROM FEDERATION /* DRIFT POLICY STRICT - --- Repair policy: trigger normalisation on drifted peers -SELECT * FROM FEDERATION /* DRIFT POLICY REPAIR ----- - -=== Adapter Capabilities - -[cols="1,1,1,1,1"] -|=== -| Modality | VeriSimDB | ArangoDB | PostgreSQL | Elasticsearch - -| Graph | Native | AQL traversal | Recursive CTE | -- -| Vector | HNSW | -- | pgvector | dense_vector kNN -| Tensor | ndarray/Burn | -- | -- | -- -| Semantic | CBOR proofs | Document fields | JSONB | Nested objects -| Document | Tantivy FTS | Fulltext index | tsvector/GIN | Full-text -| Temporal | Native | Date fields | tstzrange | date_range -| Provenance | Hash-chain | Edge collections | Audit table | -- -| Spatial | R-tree | GeoJSON index | PostGIS | geo_shape -|=== - -== NIF Transport (Same-Node Deployment) - -For Elixir applications running on the same node as the Rust core, -the NIF transport bypasses HTTP entirely for 10-100x lower latency: - -[source,bash] ----- -# Enable NIF transport -export VERISIM_TRANSPORT=nif # Direct NIF calls -# or -export VERISIM_TRANSPORT=auto # NIF if available, HTTP fallback ----- - -The NIF bridge is loaded automatically from `priv/native/libverisim_nif.so`. -All operations use the same API -- only the transport layer changes. - -== Drift Detection Deep Dive - -=== How Drift Is Measured - -Each modality pair has a drift score from 0.0 (perfectly consistent) to -1.0 (completely diverged): - -* *semantic_vector_drift*: Embedding doesn't match semantic content -* *graph_document_drift*: Graph structure doesn't match document text -* *temporal_consistency_drift*: Version history has gaps or contradictions -* *tensor_drift*: Tensor representation diverged from source data -* *schema_drift*: Type constraint violations in semantic modality -* *quality_drift*: Overall data quality score - -=== Configuring Thresholds - -[source,elixir] ----- -# In config/config.exs -config :verisim, :drift, - detection_interval_ms: 30_000, # Check every 30 seconds - warning_threshold: 0.3, # Log warning above 0.3 - critical_threshold: 0.7, # Trigger normalisation above 0.7 - auto_normalise: true # Automatically repair drifted entities ----- - -=== Self-Normalisation Strategies - -When drift exceeds the critical threshold, the normaliser selects a -repair strategy: - -[cols="1,3"] -|=== -| Strategy | Description - -| *Authoritative Source* | Identify the most trusted modality and regenerate others from it -| *Consensus* | Use majority agreement across modalities to determine correct state -| *Temporal Rewind* | Roll back to the last consistent version -| *Selective Repair* | Only regenerate the specific drifted modality -| *Full Rebuild* | Regenerate all modalities from the primary source of truth -|=== - -== API Reference - -=== REST Endpoints - -[cols="1,1,3"] -|=== -| Method | Endpoint | Description - -| `GET` | `/api/v1/health` | Health check -| `POST` | `/api/v1/octads` | Create a new entity -| `GET` | `/api/v1/octads/:id` | Retrieve entity by ID -| `PUT` | `/api/v1/octads/:id` | Update entity -| `DELETE` | `/api/v1/octads/:id` | Delete entity -| `GET` | `/api/v1/search/text?q=...&limit=10` | Full-text search -| `POST` | `/api/v1/search/vector` | Vector similarity search -| `GET` | `/api/v1/search/related/:id` | Graph traversal -| `POST` | `/api/v1/spatial/search/radius` | Spatial radius search -| `POST` | `/api/v1/spatial/search/bounds` | Spatial bounding box -| `POST` | `/api/v1/spatial/search/nearest` | k-nearest spatial -| `GET` | `/api/v1/drift/entity/:id` | Entity drift scores -| `GET` | `/api/v1/drift/status` | Overall drift status -| `POST` | `/api/v1/normalizer/trigger/:id` | Trigger normalisation -| `GET` | `/api/v1/normalizer/status` | Normaliser status -| `GET` | `/api/v1/provenance/:id` | Provenance chain -| `GET` | `/api/v1/provenance/:id/verify` | Verify provenance integrity -| `POST` | `/api/v1/federation/register` | Register federation peer -| `POST` | `/api/v1/federation/query` | Execute federated query -| `GET` | `/api/v1/federation/peers` | List federation peers -|=== - -== Running the Test Suite - -[source,bash] ----- -# Rust tests (510+ tests) -cd rust-core && cargo test - -# Elixir tests (152+ tests, excluding integration) -cd elixir-orchestration && mix test --exclude integration - -# Integration tests (requires Rust core running) -cd elixir-orchestration && mix test ----- - -== Troubleshooting - -=== Rust core won't build - -The default build requires only a Rust toolchain (no C/C++ compiler, no protoc, -no external system libraries). If you enable the optional `oxigraph-backend` -feature flag, you will need clang and cmake: - -[source,bash] ----- -# Only needed for --features oxigraph-backend (NOT the default) -sudo dnf install clang cmake # Fedora -sudo apt install clang cmake # Debian/Ubuntu ----- - -=== Elixir can't connect to Rust core - -[source,bash] ----- -# Check the Rust API is running -curl http://localhost:8080/api/v1/health - -# Set the correct URL in Elixir config -export VERISIM_RUST_CORE_URL=http://localhost:8080/api/v1 ----- - -=== VCL parser falls back to built-in - -This is normal. The external Deno-based VCL parser provides richer -error messages but the built-in Elixir parser handles all standard -queries. The fallback warning can be safely ignored. - -== Next Steps - -* Read the link:VCL-SPEC.adoc[VCL Specification] for the full query language -* See link:vcl-examples.adoc[VCL Examples] for more query patterns -* Review link:federation-readiness.adoc[Federation Readiness] for deployment planning -* Explore link:drift-handling.adoc[Drift Handling] for advanced configuration -* Check link:deployment-modes.adoc[Deployment Modes] for production setup diff --git a/verisimdb/docs/minikanren-integration-v3.adoc b/verisimdb/docs/minikanren-integration-v3.adoc deleted file mode 100644 index 9c2f63b8..00000000 --- a/verisimdb/docs/minikanren-integration-v3.adoc +++ /dev/null @@ -1,329 +0,0 @@ -// SPDX-License-Identifier: CC-BY-SA-4.0 - -= miniKanren Integration (v3 Roadmap) -:toc: -:toc-placement!: - -**Status:** STUB - Planned for VeriSimDB v3.0 - -Proposed integration of miniKanren for advanced constraint-based query optimization and learning. - -toc::[] - -== Overview - -**Current (v1.x):** Adaptive learning uses Elixir feedback loops (`adaptive_learner.ex`) - -**Future (v3.x):** Enhance with miniKanren for: -- Constraint-based query plan optimization -- Rule-based normalization policy synthesis -- Declarative drift repair strategies -- Symbolic reasoning over ZKP contracts - -== Why miniKanren? - -=== Advantages Over Mozart/Oz - -- **Lightweight** - Embeddable, no separate runtime -- **Minimal footprint** - ~500 LOC core implementation -- **Multiple host languages** - Available for Scheme, Racket, Clojure, JavaScript, Python -- **Relational programming** - Natural fit for query optimization - -=== Why Not Now (v1.x)? - -- **Elixir feedback loops sufficient** for current learning needs -- **Tiny core philosophy** - Avoid premature complexity -- **Prove value first** - Demonstrate adaptive learning works before adding logic programming - -== Proposed Architecture (v3) - -``` -┌─────────────────────────────────────────────────────────────┐ -│ Elixir Orchestration Layer │ -│ ├── VeriSim.EntityServer (GenServer per Octad) │ -│ ├── VeriSim.DriftMonitor (drift detection coordinator) │ -│ ├── VeriSim.QueryRouter (distributes queries) │ -│ ├── VeriSim.AdaptiveLearner (v1: feedback loops) │ -│ └── VeriSim.MiniKanren (v3: constraint solving) ◄─ NEW │ -│ ↓ FFI/NIFs │ -├─────────────────────────────────────────────────────────────┤ -│ miniKanren Logic Engine (Rust or AffineScript binding) │ -│ ├── Query plan constraint solver │ -│ ├── Normalization rule synthesizer │ -│ └── Drift repair strategy generator │ -└─────────────────────────────────────────────────────────────┘ -``` - -== Use Cases for miniKanren (v3) - -=== 1. Query Plan Optimization - -**Problem:** Given query constraints and store capabilities, find optimal execution plan. - -**miniKanren approach:** - -```scheme -;; Define relations for query planning -(defrel (query-plano query plan) - (fresh (modalities predicates stores) - (query-modalitieso query modalities) - (query-predicateso query predicates) - (available-storeso stores) - (plan-satisfieso plan modalities predicates stores) - (plan-minimizes-costo plan))) - -;; Ask miniKanren: What plans satisfy this query? -(run* (plan) - (query-plano - '(FROM verisim:graph WHERE octad.types INCLUDES "Paper") - plan)) -``` - -**Output:** Multiple valid plans, ranked by estimated cost. - -=== 2. Normalization Policy Synthesis - -**Problem:** Given observed deviance patterns, synthesize detection rules. - -**miniKanren approach:** - -```scheme -;; Learn rule from positive/negative examples -(defrel (normalization-ruleo examples rule) - (fresh (positive-examples negative-examples) - (partition-exampleso examples positive-examples negative-examples) - (rule-accepts-allo rule positive-examples) - (rule-rejects-allo rule negative-examples) - (rule-is-minimalo rule))) - -;; Synthesize rule from training data -(run 1 (rule) - (normalization-ruleo - '((cache-ttl 300 . ok) - (cache-ttl 3600 . deviance) - (cache-ttl 60 . ok)) - rule)) -``` - -**Output:** `(rule (lambda (ttl) (and (>= ttl 60) (<= ttl 600))))` - -=== 3. Drift Repair Strategy Selection - -**Problem:** Given drift type and context, choose best repair strategy. - -**miniKanren approach:** - -```scheme -;; Relational specification of repair strategies -(defrel (repair-strategyo drift-type context strategy) - (conde - ;; latest_wins when temporal ordering clear - [(== drift-type 'timestamp-mismatch) - (fresh (timestamps) - (all-have-timestampso context timestamps) - (== strategy 'latest-wins))] - - ;; quorum when values diverge - [(== drift-type 'title-mismatch) - (multiple-valueso context) - (== strategy 'quorum)] - - ;; manual for critical fields - [(critical-fieldo drift-type) - (== strategy 'manual)])) - -;; Query: What strategy for this drift? -(run* (strategy) - (repair-strategyo - 'title-mismatch - '{:values ["ML Paper" "Machine Learning Paper"] - :modalities [graph document]} - strategy)) -``` - -**Output:** `(quorum manual)` - multiple valid strategies - -=== 4. ZKP Contract Satisfaction - -**Problem:** Verify if a octad satisfies contract constraints symbolically. - -**miniKanren approach:** - -```scheme -;; Relational specification of ZKP contracts -(defrel (contract-satisfiedo octad contract) - (fresh (contract-type constraints) - (contract-typeo contract contract-type) - (contract-constraintso contract constraints) - (all-constraints-satisfiedo octad constraints))) - -;; Check satisfaction without executing ZKP -(run* (result) - (contract-satisfiedo - '{:types ["Paper"] :citations []} - 'CitationContract - result)) -``` - -**Output:** `(#f)` - octad does NOT satisfy contract (missing citations) - -== Implementation Options (v3) - -=== Option A: Rust miniKanren Binding - -**Crate:** `rust-kanren` or `microkanren-rs` - -```rust -// lib/verisim_minikanren/src/lib.rs - -use microkanren::*; - -#[rustler::nif] -fn solve_query_plan(query_ast: String) -> Result, String> { - let state = State::new(); - - // Encode query as miniKanren goal - let goal = parse_query_to_goal(&query_ast)?; - - // Run miniKanren search - let results = state.run(goal, 10); // Top 10 plans - - Ok(results.into_iter().map(|r| r.to_string()).collect()) -} -``` - -**Integration:** Elixir calls Rust NIF → miniKanren → returns results - -=== Option B: ReScript miniKanren Binding - -**Package:** `rescript-minikanren` (port from JavaScript implementation) - -```rescript -// src/bindings/MiniKanren.res - -module MiniKanren = { - type goal - type state - - @module("minikanren") external run: (state, goal, int) => array = "run" - @module("minikanren") external fresh: (state => goal) => goal = "fresh" - @module("minikanren") external eq: (string, string) => goal = "==" -} - -let solveQueryPlan = (queryAst: string): array => { - let state = MiniKanren.State.new() - let goal = parseQueryToGoal(queryAst) - MiniKanren.run(state, goal, 10) -} -``` - -**Integration:** Elixir calls ReScript (compiled to JS) → Deno → miniKanren → returns results - -=== Option C: Scheme miniKanren (via Guile) - -**Use existing STATE.scm infrastructure:** - -```scheme -;; lib/verisim_minikanren.scm - -(use-modules (minikanren)) - -(define (solve-query-plan query-ast) - (run 10 (plan) - (query-plano query-ast plan))) - -;; Called from Elixir via Ports -``` - -**Integration:** Elixir spawns Guile process → miniKanren → returns results via stdin/stdout - -== Performance Considerations (v3) - -=== When to Use miniKanren - -**Good for:** -- Small search spaces (< 1000 possibilities) -- Symbolic reasoning (type checking, contract verification) -- Policy synthesis (learn rules from examples) -- Offline optimization (query plan analysis) - -**Not good for:** -- Large search spaces (> 10,000 possibilities) - use heuristics instead -- Real-time critical path (miniKanren search is slow) -- Numeric optimization (use gradient descent, not logic search) - -=== Hybrid Approach (Recommended) - -```elixir -defmodule VeriSim.QueryPlanner do - @doc """ - v1: Use adaptive learner (fast feedback loops) - v3: Use miniKanren for hard cases (constraint solving) - """ - def plan_query(query) do - if simple_query?(query) do - # Fast path: Use learned plans from adaptive_learner - VeriSim.AdaptiveLearner.get_policy(:query_plan) - |> apply_to_query(query) - else - # Slow path: Use miniKanren for complex constraints - VeriSim.MiniKanren.solve_query_plan(query) - end - end -end -``` - -== Migration Path (v1 → v3) - -=== Phase 1: v1.x (Current) - -- ✅ Adaptive learner with Elixir feedback loops -- ✅ Heuristic query planning -- ✅ Rule-based drift repair - -=== Phase 2: v2.x (Intermediate) - -- Collect query plan training data -- Measure adaptive learner limitations -- Prototype miniKanren integration in non-critical paths - -=== Phase 3: v3.x (Full Integration) - -- Replace heuristic query planner with miniKanren solver -- Synthesize normalization rules from examples -- Symbolic ZKP contract verification -- Declarative drift repair strategies - -== Success Criteria (v3) - -Before adding miniKanren, we need: - -1. **Demonstrated Need** - - Adaptive learner hits optimization ceiling - - Query planning requires constraint solving (not just heuristics) - - Rule synthesis requested by users - -2. **Performance Acceptable** - - miniKanren query planning < 100ms for typical queries - - Acceptable latency impact on critical path - -3. **Tiny Core Maintained** - - miniKanren binding < 1,000 LOC - - Total coordination logic still < 5,000 LOC - -== References - -- https://minikanren.org/[miniKanren official site] -- http://minikanren.org/workshop/2023/[miniKanren Workshop 2023] -- https://github.com/webyrd/faster-minikanren[Faster miniKanren implementation] -- https://www.youtube.com/watch?v=eQL48qYDwp4[The Reasoned Schemer (book/video)] - -== Summary - -miniKanren integration is **planned for v3** as an enhancement to v1's adaptive learner. - -**Current approach (v1):** Elixir feedback loops - fast, simple, good enough - -**Future approach (v3):** miniKanren for hard problems - constraint solving, rule synthesis, symbolic reasoning - -**Decision:** Prove adaptive learning works before adding logic programming. diff --git a/verisimdb/docs/normalization-cascade.adoc b/verisimdb/docs/normalization-cascade.adoc deleted file mode 100644 index 8c7acb73..00000000 --- a/verisimdb/docs/normalization-cascade.adoc +++ /dev/null @@ -1,762 +0,0 @@ -// SPDX-License-Identifier: CC-BY-SA-4.0 - -= Normalization Cascade Architecture -:toc: left -:toclevels: 4 -:sectnums: -:stem: latexmath - -== Overview - -This document analyzes the **normalization cascade** in VeriSimDB - the hierarchical levels at which consistency is maintained, and whether normalization should be **pushed** (proactive propagation) or **pulled** (on-demand correction). - -**Core Question:** Is VeriSimDB responsible for: -1. **Database-only normalization** - Ensuring its own internal consistency? -2. **System-wide normalization** - Actively propagating changes across federated nodes? - -== ⭐ RECOMMENDATION: Hybrid Push/Pull with Database-Internal Focus - -[IMPORTANT] -==== -**Decision:** VeriSimDB normalizes what it controls, advises on what it doesn't. - -**Core Philosophy:** Push is expensive - reserve for safety-critical issues. Pull scales better for optimization. -==== - -=== Strategy by Level (The Answer) - -[cols="1,2,1,3"] -|=== -|Level |Scope |Strategy |Rationale - -|**L0: Intra-Modality** -|Within single modality store -|✅ **PULL** -|Stores normalize on insert. VeriSimDB delegates to modality stores. - -|**L1: Cross-Modality** -|Across modalities in octad -|✅ **HYBRID** -|**PUSH:** Critical (integrity violations, retractions) + -**PULL:** Optimization (title mismatches, cosmetic drift) + -*This is VeriSimDB's unique value - multi-modal consistency* - -|**L2: Cross-Octad** -|Between related octads -|✅ **PULL** -|Validate citations at query time. Most citations are valid, scales better. - -|**L3: Cross-Store** -|Federated VeriSimDB instances -|✅ **HYBRID** -|**ASYNC PUSH:** Critical (retractions, access changes) + -**PULL:** Version drift, schema updates + -*Shared responsibility, no single authority* - -|**L4: Cross-Lineage** -|External systems (arXiv, PubMed, etc.) -|⚠️ **ADVISORY** -|Best-effort validation, cache results, warn but don't enforce. + -*VeriSimDB can't control external systems* -|=== - -=== What Gets Pushed vs Pulled - -**PUSH (Proactive Propagation):** - -- ❗ **Retractions** - Paper retracted → notify all citing papers -- ❗ **Integrity violations** - Hash mismatch detected → immediate alert -- ❗ **Access revocations** - User loses permission → propagate immediately -- ⚠️ **Major version updates** - Schema change → background sync (async push) - -**PULL (On-Demand Repair):** - -- ✏️ **Title mismatches** - "ML Paper" vs "Machine Learning Paper" → repair at query time -- ✏️ **Citation counts** - Derived data → recompute on demand -- ✏️ **Cache invalidation** - Stale cache → refresh when accessed -- ✏️ **External validation** - Check DOI exists → best-effort at query time - -=== Why This Works - -✅ **Scales:** Most drift is cosmetic, doesn't need immediate propagation - -✅ **Clear boundaries:** VeriSimDB controls L0-L2 (internal), collaborates on L3 (federation), advises on L4 (external) - -✅ **User control:** Conservative/Balanced/Aggressive modes let users tune aggressiveness - -✅ **Learns:** Adaptive learner can promote frequent drift from PULL→PUSH based on observations - -✅ **Safety-first:** Critical issues are pushed immediately, optimizations are pulled lazily - -=== Configuration Example - -```elixir -# config/config.exs -config :verisim, VeriSim.NormalizationCoordinator, - # Conservative: minimal push (retractions only) - # Balanced: push critical + important (DEFAULT) - # Aggressive: push nearly everything - mode: :balanced, - - # What triggers immediate push? - push_critical: [ - :retraction, # Paper retracted - :integrity_violation, # Hash mismatch - :access_revocation # Permission removed - ], - - # What triggers async push (background job)? - push_important: [ - :schema_version, # Schema updated - :major_update # Significant change - ], - - # Everything else: PULL (query-time repair) - - # External systems - validate_external: true, # Best-effort validation - external_timeout: 5_000, # Don't wait > 5s - external_cache_ttl: :timer.hours(24) -``` - -=== The Answer to "Is This VeriSimDB's Job?" - -[cols="1,2,1"] -|=== -|Level |VeriSimDB's Job? |Why? - -|L0-L1 -|✅ **Yes** (core responsibility) -|Multi-modal consistency is VeriSimDB's unique value - -|L2 -|✅ **Yes** (internal graphs) -|VeriSimDB can query its own graph store - -|L3 -|✅ **Yes** (federation) -|VeriSimDB instances collaborate via KRaft - -|L4 -|⚠️ **Partial** (advisory) -|External systems unreliable, can't enforce -|=== - -**Bottom line:** VeriSimDB normalizes L0-L3 (what it controls), advises on L4 (what it doesn't). - -== Normalization Cascade Levels - -We identify **five levels** of normalization scope: - -[cols="1,2,3,2"] -|=== -|Level |Scope |What Normalizes |Responsibility - -|**L0: Intra-Modality** -|Within a single modality store -|Vector dimensions, tensor shapes, RDF consistency -|**Modality store** (verisim-graph, verisim-vector, etc.) - -|**L1: Cross-Modality** -|Across modalities within a octad -|Title in GRAPH matches DOCUMENT, embedding dimension matches metadata -|**VeriSimDB core** (EntityServer, DriftMonitor) - -|**L2: Cross-Octad** -|Between octads in same store -|Citation targets exist, provenance chains valid -|**VeriSimDB core** (via graph queries) - -|**L3: Cross-Store** -|Between federated stores -|Replicated octads match across institutions -|**Federation layer** (KRaft registry + drift policies) - -|**L4: Cross-Lineage** -|Through citation/provenance chains -|Retractions propagate, versioning cascades -|**System-wide** (VeriSimDB + external services) -|=== - -=== Level Details - -==== L0: Intra-Modality Normalization - -**Scope:** Within a single modality store (e.g., verisim-vector) - -**Examples:** -- Vector normalization (L2 norm = 1.0) -- Tensor shape validation (declared rank matches data) -- RDF triple deduplication -- Document encoding consistency (UTF-8) - -**Responsibility:** **Modality store itself** - -**Rationale:** Each store knows its own invariants. VeriSimDB core shouldn't micromanage. - -**Implementation:** -```rust -// rust-core/verisim-vector/src/lib.rs -impl VectorStore { - pub fn insert(&mut self, octad_id: UUID, vector: Vec) -> Result<()> { - // L0 normalization: ensure vector is normalized - let norm = vector.iter().map(|x| x * x).sum::().sqrt(); - let normalized = vector.iter().map(|x| x / norm).collect(); - - self.storage.insert(octad_id, normalized) - } -} -``` - -==== L1: Cross-Modality Normalization - -**Scope:** Across modalities within a single octad - -**Examples:** -- Title mismatch: GRAPH says "ML Paper", DOCUMENT says "Machine Learning Paper" -- Embedding dimension: SEMANTIC metadata says 768-dim, VECTOR has 512-dim -- Timestamp conflict: TEMPORAL shows updated 2025-01-20, GRAPH shows 2025-01-22 - -**Responsibility:** **VeriSimDB core (DriftMonitor)** - -**Rationale:** Only VeriSimDB sees all modalities for a octad. This is its core value-add. - -**Implementation:** -```elixir -# lib/verisim/drift_monitor.ex -defmodule VeriSim.DriftMonitor do - def detect_cross_modal_drift(octad_id) do - modalities = fetch_all_modalities(octad_id) - - # L1 normalization: check title consistency - titles = [ - modalities.graph["title"], - modalities.document["title"], - modalities.semantic["title"] - ] - - if length(Enum.uniq(titles)) > 1 do - {:drift_detected, :title_mismatch, titles} - else - :consistent - end - end -end -``` - -==== L2: Cross-Octad Normalization - -**Scope:** Between related octads in the same store - -**Examples:** -- Citation target exists: Paper A cites Paper B, ensure B exists -- Provenance chain valid: Octad derives from Parent, ensure Parent accessible -- Circular reference detection: A → B → C → A (invalid) - -**Responsibility:** **VeriSimDB core (via graph queries)** - -**Rationale:** VeriSimDB can query its own graph store to validate relationships. - -**Implementation:** -```elixir -# lib/verisim/citation_validator.ex -defmodule VeriSim.CitationValidator do - def validate_citations(octad_id) do - # L2 normalization: ensure citation targets exist - citations = get_citations(octad_id) - - Enum.map(citations, fn cited_id -> - case EntityServer.exists?(cited_id) do - true -> {:ok, cited_id} - false -> {:error, :citation_target_missing, cited_id} - end - end) - end -end -``` - -==== L3: Cross-Store Normalization - -**Scope:** Between federated stores (different institutions) - -**Examples:** -- Replicated octad mismatch: University A has version 3, University B has version 5 -- Schema drift: Store A uses new field "keywords", Store B doesn't -- Trust boundary: Store A accepts update from Store B, need verification - -**Responsibility:** **Federation layer (shared between VeriSimDB instances)** - -**Rationale:** No single VeriSimDB instance has authority. Requires distributed consensus. - -**Implementation:** -```elixir -# lib/verisim/federation_normalizer.ex -defmodule VeriSim.FederationNormalizer do - def normalize_across_stores(octad_id, stores) do - # L3 normalization: quorum-based repair - versions = Enum.map(stores, fn store -> - fetch_version(store, octad_id) - end) - - case apply_drift_policy(:quorum, versions) do - {:ok, canonical_version} -> - # Push canonical version to minority stores - push_to_stores(canonical_version, stores) - {:error, :no_quorum} -> - # Escalate to manual resolution - {:manual_intervention_required, versions} - end - end -end -``` - -==== L4: Cross-Lineage Normalization - -**Scope:** Through citation/provenance chains (potentially cross-system) - -**Examples:** -- Paper retraction: Paper A is retracted → notify all papers citing A -- Data correction: Dataset B is corrected → cascade to all analyses using B -- Version update: Library C releases v2 → notify dependent projects - -**Responsibility:** **System-wide (VeriSimDB + external notification services)** - -**Rationale:** Lineage may span systems VeriSimDB doesn't control (external publishers, repositories). - -**Implementation:** -```elixir -# lib/verisim/lineage_propagator.ex -defmodule VeriSim.LineagePropagator do - def propagate_retraction(retracted_octad_id) do - # L4 normalization: notify citing octads - citing_octads = get_all_citations_to(retracted_octad_id) - - # Within VeriSimDB: update status - Enum.each(citing_octads, fn citing_id -> - add_warning(citing_id, "Cites retracted work: #{retracted_octad_id}") - end) - - # System-wide: notify external systems - external_citations = get_external_citations(retracted_octad_id) - Enum.each(external_citations, fn ext -> - notify_external_system(ext.system_url, :retraction_notice, retracted_octad_id) - end) - end -end -``` - -== Push vs Pull: Architectural Tradeoffs - -=== Push Model (Proactive Propagation) - -**Definition:** Changes are **actively propagated** to dependent entities. - -**Mechanism:** Database triggers, event sourcing, message queues - -**Example:** -``` -Octad A updated → DriftMonitor detects change → Pushes update to Octad B (cites A) -``` - -==== Push Model: Pros - -✅ **Eventual consistency guaranteed** - All nodes converge without query intervention - -✅ **Proactive detection** - Drift is caught immediately, not at query time - -✅ **Audit trail** - Every propagation is logged (good for compliance) - -✅ **Predictable state** - System reaches consistency without external stimulus - -==== Push Model: Cons - -❌ **Cascade storms** - One change triggers avalanche of updates - -❌ **Network overhead** - Constant background traffic for federation - -❌ **Ordering issues** - Concurrent pushes may conflict - -❌ **Complexity** - Need distributed transaction coordination - -❌ **Wasted work** - Propagate changes to data that's never queried - -=== Pull Model (On-Demand Correction) - -**Definition:** Normalization happens **only when data is queried**. - -**Mechanism:** Lazy evaluation, cache invalidation, query-time drift detection - -**Example:** -``` -Query Octad A → Detect drift with Octad B → Repair on-the-fly → Return normalized result -``` - -==== Pull Model: Pros - -✅ **No background work** - Zero overhead when data isn't accessed - -✅ **Scales better** - Only normalize hot data - -✅ **Simpler** - No distributed coordination for propagation - -✅ **Flexible repair** - Can choose strategy based on query context - -==== Pull Model: Cons - -❌ **Query latency** - First query after drift pays normalization cost - -❌ **Stale data** - Drift exists until queried - -❌ **Redundant work** - Multiple queries may re-normalize same data - -❌ **No proactive guarantees** - Cold data stays drifted indefinitely - -=== Hybrid Model (Tiered Push/Pull) - -**Definition:** Push critical changes, pull for optimizations. - -**Tiers:** - -[cols="1,2,2,1"] -|=== -|Tier |Examples |Strategy |Rationale - -|**Critical (Push)** -|Retraction, integrity violation, access revocation -|Immediate propagation -|Safety-critical, must be consistent - -|**Important (Async Push)** -|Major version update, schema change -|Background job (within 1 hour) -|Should converge, but not blocking - -|**Optimization (Pull)** -|Cache TTL, minor title mismatch -|Query-time repair -|Cosmetic, low impact - -|**Informational (Pull)** -|Citation count, popularity metrics -|Never pushed, computed on demand -|Derived data, not canonical -|=== - -==== Hybrid Implementation - -```elixir -# lib/verisim/normalization_coordinator.ex -defmodule VeriSim.NormalizationCoordinator do - def handle_drift(drift_type, octad_id, details) do - case classify_drift(drift_type) do - :critical -> - # PUSH: Immediate propagation - push_to_all_dependents(octad_id, details) - - :important -> - # ASYNC PUSH: Background job - Oban.insert(NormalizationJob.new(%{octad_id: octad_id, details: details})) - - :optimization -> - # PULL: Mark as drifted, repair at query time - mark_drifted(octad_id, drift_type) - - :informational -> - # PULL: Do nothing, recompute on demand - :ok - end - end - - defp classify_drift(:retraction), do: :critical - defp classify_drift(:integrity_violation), do: :critical - defp classify_drift(:schema_version), do: :important - defp classify_drift(:title_mismatch), do: :optimization - defp classify_drift(:citation_count), do: :informational -end -``` - -== Database-Internal vs System-Wide - -=== Option A: Database-Only Normalization - -**Philosophy:** VeriSimDB ensures its own consistency. External systems handle their own. - -**Scope:** -- ✅ L0: Intra-modality (modality stores) -- ✅ L1: Cross-modality (VeriSimDB core) -- ✅ L2: Cross-octad (VeriSimDB graph queries) -- ⚠️ L3: Cross-store (only if both stores are VeriSimDB instances) -- ❌ L4: Cross-lineage (not VeriSimDB's job) - -**Pros:** -- ✅ Clear responsibility boundary -- ✅ VeriSimDB doesn't need to understand external systems -- ✅ Simpler implementation - -**Cons:** -- ❌ Federation is limited (only VeriSimDB ↔ VeriSimDB) -- ❌ External citations can't be validated -- ❌ System-wide retractions require manual coordination - -**Example:** -``` -VeriSimDB detects Paper A cites Paper B (external arXiv) -→ VeriSimDB stores citation, but does NOT validate B exists -→ User must manually check arXiv -``` - -=== Option B: System-Wide Normalization - -**Philosophy:** VeriSimDB actively maintains consistency across all linked systems. - -**Scope:** -- ✅ L0-L4: All levels, including external systems - -**Pros:** -- ✅ End-to-end consistency guarantees -- ✅ Can validate external citations (via APIs) -- ✅ Retractions propagate everywhere - -**Cons:** -- ❌ VeriSimDB becomes a "master coordinator" (single point of failure) -- ❌ Must understand every external system (arXiv, PubMed, ORCID, etc.) -- ❌ External systems may not support callbacks -- ❌ Huge complexity explosion - -**Example:** -``` -VeriSimDB detects Paper A cites Paper B (external arXiv) -→ VeriSimDB calls arXiv API to validate B exists -→ VeriSimDB subscribes to arXiv retraction feed -→ If B is retracted, VeriSimDB pushes warning to A -``` - -=== Option C: Hybrid Boundary (RECOMMENDED) - -**Philosophy:** VeriSimDB normalizes internally and within its federation. External systems are **advisory**. - -**Scope:** -- ✅ L0-L2: Full normalization (VeriSimDB internal) -- ✅ L3: Full normalization (VeriSimDB federation) -- ⚠️ L4: **Advisory only** (external systems) - -**External Citation Handling:** - -1. **Store external citations** (URL, DOI, identifier) -2. **Validate at query time** (optional, best-effort) -3. **Cache validation results** (with TTL) -4. **Don't enforce** (warn if unreachable, but allow) - -**Implementation:** -```elixir -# lib/verisim/external_validator.ex -defmodule VeriSim.ExternalValidator do - def validate_external_citation(doi) do - # Try to validate, but don't fail if we can't - case CrossrefAPI.resolve(doi) do - {:ok, metadata} -> - # Cache result - Cache.put("doi:#{doi}", metadata, ttl: :timer.hours(24)) - {:ok, :validated, metadata} - - {:error, :not_found} -> - # Warn but don't block - {:warning, :external_not_found, doi} - - {:error, :network_error} -> - # Assume valid if we can't reach external system - {:ok, :assumed_valid, "Network unavailable"} - end - end -end -``` - -== Recommendation: Hybrid Push/Pull + Hybrid Boundary - -=== Proposed Architecture - -**L0 (Intra-Modality):** -- **Strategy:** PULL (modality stores normalize on insert/update) -- **Responsibility:** Modality store - -**L1 (Cross-Modality):** -- **Strategy:** HYBRID - - Critical: PUSH (integrity violations) - - Optimization: PULL (title mismatches) -- **Responsibility:** VeriSimDB DriftMonitor - -**L2 (Cross-Octad):** -- **Strategy:** PULL (validate citations at query time) -- **Responsibility:** VeriSimDB graph queries - -**L3 (Cross-Store):** -- **Strategy:** HYBRID - - Critical: ASYNC PUSH (retractions, access changes) - - Optimization: PULL (version drift) -- **Responsibility:** Federation layer (KRaft + drift policies) - -**L4 (Cross-Lineage):** -- **Strategy:** ADVISORY (best-effort external validation) -- **Responsibility:** VeriSimDB + user responsibility - -=== Decision Matrix - -[cols="1,2,2,2"] -|=== -|Level |Strategy |Push or Pull? |VeriSimDB's Job? - -|L0 -|On-insert validation -|PULL -|✅ Yes (delegated to stores) - -|L1 -|Critical: push, Opt: pull -|HYBRID -|✅ Yes (core responsibility) - -|L2 -|Query-time validation -|PULL -|✅ Yes (internal graph) - -|L3 -|Critical: async push, Opt: pull -|HYBRID -|✅ Yes (federation layer) - -|L4 -|Best-effort advisory -|PULL -|⚠️ Partial (warn only) -|=== - -=== Rationale - -1. **Push is expensive** - Reserve for critical safety issues (retraction, access) -2. **Pull scales better** - Most drift is cosmetic (title variations) -3. **External systems are unreliable** - Don't block on them -4. **VeriSimDB controls its boundary** - Internal + federation are its domain -5. **Users own external validation** - VeriSimDB helps but doesn't enforce - -== Configuration - -Users can configure normalization aggressiveness: - -```elixir -# config/config.exs -config :verisim, VeriSim.NormalizationCoordinator, - # Conservative: minimal push, mostly pull - mode: :conservative, # :conservative | :balanced | :aggressive - - # What drift types trigger push? - push_critical: [:retraction, :integrity_violation, :access_revocation], - push_important: [:schema_version, :major_update], - - # External validation - validate_external: true, # Best-effort validation - external_timeout: 5_000, # Don't wait > 5s - - # Cross-store normalization - federation_sync_interval: :timer.hours(1), # Background sync - drift_policy: :repair # :strict | :repair | :tolerate | :latest -``` - -=== Conservative Mode - -- **Push:** Only retractions and security issues -- **Pull:** Everything else -- **Best for:** High-volume systems, performance-critical - -=== Balanced Mode (Default) - -- **Push:** Critical + important drift -- **Pull:** Optimizations -- **Best for:** Most deployments - -=== Aggressive Mode - -- **Push:** Nearly everything -- **Pull:** Only informational metrics -- **Best for:** High-assurance systems (medical, legal) - -== Integration with Adaptive Learner (v1) and miniKanren (v3) - -=== Adaptive Learner (v1) - -The normalization cascade can **learn optimal strategies** via feedback loops: - -```elixir -# lib/verisim/adaptive_learner.ex -defmodule VeriSim.AdaptiveLearner do - def learn_normalization_policy(state) do - # Observe: what drift types occur most? - drift_frequency = analyze_drift_frequency(state.observations) - - # Decide: should we push this drift type? - Enum.map(drift_frequency, fn {drift_type, freq} -> - if freq > 0.1 do - # High frequency → switch to PUSH - recommend_push(drift_type) - else - # Low frequency → keep as PULL - recommend_pull(drift_type) - end - end) - end -end -``` - -=== miniKanren (v3) - -In v3, miniKanren can **synthesize normalization rules** from examples: - -```scheme -;; Learn: when should we push vs pull? -(defrel (normalization-strategyo drift-examples strategy) - (fresh (frequency severity user-impact) - (analyze-exampleso drift-examples frequency severity user-impact) - (conde - ;; Rule 1: High frequency + high severity → PUSH - [(>o frequency 0.2) - (== severity 'high) - (== strategy 'push)] - - ;; Rule 2: Low frequency → PULL - [( -// Date: 2026-02-28 - -= PanLL Module Audit -Jonathan D.A. Jewell -2026-02-28 -:toc: left -:toclevels: 3 -:icons: font -:sectnums: - -== Purpose - -Systematically evaluate which of the ~265 hyperpolymath repos map to PanLL -modules using the three-pane model: - -* **Pane-L** (Left): Constraints — types, proofs, policies, invariants -* **Pane-N** (Centre): Agent reasoning — planning, analysis, decision-making -* **Pane-W** (Right): Results — output, visualisation, artefacts - -PanLL is the unified mission control / accessibility layer for the entire -hyperpolymath ecosystem — a "multi-monitor trader desk" with neurosymbolic -intelligence behind each panel. - -== Evaluation Criteria - -[cols="1,3"] -|=== -| Rating | Definition - -| **Strong fit** -| Clear 3-pane mapping, immediate value. The repo naturally decomposes into - constraints (Pane-L), reasoning (Pane-N), and output (Pane-W). Could be - integrated as a PanLL module in 1-2 sessions. - -| **Moderate fit** -| Partial mapping — maybe 2 of 3 panes are clear, or the fit is for a - future evolution of the repo. Worth tracking but not immediate. - -| **Not applicable** -| Libraries, utilities, archived repos, or repos that are consumed by other - modules rather than being modules themselves. -|=== - -== Already Mapped (8 modules) - -These repos are already identified as PanLL modules: - -[cols="1,2,3"] -|=== -| Module | Repo | Pane Mapping - -| **VeriSimDB** -| `nextgen-databases/verisimdb` -| L: VCL-UT proofs, drift thresholds. N: Query planning, drift analysis. W: Entity data, drift heatmap. - -| **QuandleDB** -| `nextgen-databases/quandledb` -| L: Quandle algebra constraints. N: Query optimisation. W: Query results, algebraic visualisation. - -| **LithoGlyph** -| `nextgen-databases/lithoglyph` -| L: Schema types, formal specs. N: Query decomposition. W: Glyph rendering, results. - -| **reposystem** -| `reposystem/` -| L: RSR compliance rules. N: Template selection, scaffolding logic. W: Generated repo structure. - -| **gitbot-fleet** -| `gitbot-fleet/` -| L: Bot policies, confidence thresholds. N: Issue analysis, PR review. W: Bot actions, reports. - -| **IDApTIK** -| `idaptik/` -| L: Dependent type level validation. N: Level design reasoning. W: Game UI, level preview. - -| **panic-attack** -| `panic-attack/` -| L: Security rules, vulnerability patterns. N: Static analysis, threat modelling. W: Scan reports, findings. - -| **hypatia** -| `hypatia/` -| L: CI/CD policies, scan thresholds. N: Pattern correlation, dispatch planning. W: Dashboard, alerts. -|=== - -== Strong Fit (18 candidates) - -[cols="1,2,3"] -|=== -| Repo | Monorepo | Pane Mapping - -| **airie** -| standalone -| L: Network topology constraints, routing policies. N: Simulation engine, traffic analysis. W: Network visualisation, latency maps. - -| **proven** -| standalone -| L: Formal proof obligations, type safety rules. N: Proof search, verification. W: Proof certificates, audit trail. - -| **stapeln** -| `ambientops` -| L: Container policies, supply chain rules. N: Build orchestration, layer optimisation. W: Container images, SBOM. - -| **echidna** -| standalone -| L: Code safety rules (believe_me bans). N: AST analysis, pattern matching. W: Lint reports, fix suggestions. - -| **Eclexia** -| `nextgen-languages/eclexia` -| L: Grammar rules, type system. N: Parse/typecheck/evaluate. W: REPL output, AST visualisation. - -| **AffineScript** -| `nextgen-languages/affinescript` -| L: Affine type constraints, linearity. N: Type inference, borrow checking. W: Compiled output, type errors. - -| **Anvomidav** -| `nextgen-languages/anvomidav` -| L: Language grammar, type rules. N: Compilation, optimisation. W: Compiled artefacts. - -| **protocol-squisher** -| standalone -| L: Protocol specs, compression constraints. N: Compression strategy selection. W: Compressed output, stats. - -| **squisher-corpus** -| standalone -| L: Corpus validation rules. N: Benchmark analysis, comparison. W: Benchmark results, charts. - -| **PanLL itself** -| standalone -| L: Module registry, pane layout constraints. N: Module coordination, routing. W: The dashboard UI itself. - -| **ambientops/volumod** -| `ambientops` -| L: Volume/mount policies. N: Storage analysis, allocation. W: Volume status, usage reports. - -| **ambientops/displace** -| `ambientops` -| L: Deployment constraints. N: Deployment planning, rollback logic. W: Deployment status, logs. - -| **claim-forge** -| `reposystem` -| L: Claim validation rules. N: Claim analysis, evidence gathering. W: Claim reports, certificates. - -| **scaffoldia** -| `reposystem` -| L: Template constraints, RSR rules. N: Template selection, customisation. W: Generated project. - -| **bitfuckit** -| `reposystem` -| L: Bitbucket API constraints. N: Repo analysis, migration planning. W: Migration reports, API responses. - -| **Axiom.jl** -| `developer-ecosystem/julia-ecosystem` -| L: Mathematical axioms, type constraints. N: Computation, proof search. W: Results, proofs. - -| **filesoup/fslint** -| `filesoup` -| L: Lint rules, file quality policies. N: File analysis, deduplication. W: Lint reports, cleanup actions. - -| **social-media-tools** -| `social-media-tools` -| L: Platform API constraints, rate limits. N: Content scheduling, analytics. W: Post status, engagement metrics. -|=== - -== Moderate Fit (22 candidates) - -[cols="1,2,3"] -|=== -| Repo | Reason for Moderate | Notes - -| **polystack/poly-*components** -| `polystack` -| Partial — the individual components are libraries, but the stack as a whole has a 3-pane mapping. Could be a meta-module. - -| **developer-ecosystem/deno-ecosystem** -| `developer-ecosystem` -| Deno project management tools have some 3-pane potential (L: package policies, N: dependency analysis, W: project status). But many are thin wrappers. - -| **developer-ecosystem/zig-ecosystem** -| `developer-ecosystem` -| Build system tools could map to PanLL for cross-compilation management. - -| **developer-ecosystem/rescript-ecosystem** -| `developer-ecosystem` -| ReScript tooling ecosystem — compiler/formatter integration could be a PanLL dev-tools module. - -| **developer-ecosystem/idris2-ecosystem** -| `developer-ecosystem` -| Idris2 package management. Moderate fit for type theory module. - -| **developer-ecosystem/elixir-ecosystem** -| `developer-ecosystem` -| OTP project management. Could integrate with VeriSimDB's Elixir layer. - -| **nickel-augmentation** -| `nickel-augmentation` -| Nickel config management has strong L-pane (schemas) but weaker N/W panes. - -| **ipv6-tools** -| `ipv6-tools` -| Network tools — L: RFC compliance. N: Address analysis. W: Connectivity reports. But narrow scope. - -| **standards/axel-protocol** -| `standards` -| Protocol definition — strong L-pane, but the standard itself isn't runtime software. - -| **standards/a2ml** -| `standards` -| Markup language spec — more of a consumed standard than an interactive module. - -| **standards/rhodium-standard-repositories** -| `standards` -| RSR spec — consumed by reposystem tools, not interactive itself. - -| **nextgen-languages/wokelang** -| `nextgen-languages` -| Language design — could be Eclexia-like module but less developed. - -| **nextgen-languages/ephapax** -| `nextgen-languages` -| Concatenative language — moderate fit for language design module. - -| **nextgen-languages/betlang** -| `nextgen-languages` -| Domain-specific language — narrow scope. - -| **games/phantom-metal-taste** -| `games` -| Game project — could be IDApTIK-like but different genre. - -| **games/dicti0nary-attack** -| `games` -| Word game — lightweight, less module potential. - -| **games/candy-crash** -| `games` -| Game — similar to above. - -| **asdf-tool-plugins (74 repos)** -| `asdf-tool-plugins` -| Plugin ecosystem — could be a toolchain management module collectively. L: version constraints. N: version resolution. W: installed tools. - -| **patallm-gallery/echomesh** -| `patallm-gallery` -| Network mesh — moderate fit for distributed systems module. - -| **patallm-gallery/elegant-state** -| `patallm-gallery` -| State management — could integrate with PanLL's own state. - -| **patallm-gallery/claude-integrations** -| `patallm-gallery` -| AI integrations — moderate fit, depends on direction. - -| **0-ai-gatekeeper-protocol** -| standalone -| AI manifest protocol — strong L-pane (rules), but it's a consumed spec not an interactive module. -|=== - -== Not Applicable (215+ repos) - -The remaining ~215 repos fall into categories that don't map to PanLL modules: - -[cols="1,1,2"] -|=== -| Category | Count | Reason - -| **hyperpolymath-archive** -| 113 -| Archived repos — historical, not active - -| **asdf-tool-plugins (individual)** -| 74 -| Each individual plugin is too thin — they work collectively (see moderate fit above) - -| **Library dependencies** -| ~15 -| Consumed by other modules, not interactive themselves (e.g., language-bridges, cadre-router) - -| **Mirror/practice repos** -| ~8 -| Practice mirrors, forks, not original work - -| **Config/infrastructure** -| ~5 -| .git-private-farm, farm manifests, CI config -|=== - -== Recommended Next Steps - -1. **Immediate**: Register the 8 already-mapped modules in PanLL's DatabaseRegistry (VeriSimDB, QuandleDB, LithoGlyph) and create a parallel LanguageRegistry (Eclexia, AffineScript) and ToolRegistry (panic-attack, hypatia, gitbot-fleet) - -2. **Short-term**: Implement the top 5 strong-fit candidates as PanLL modules: airie, proven, stapeln, echidna, Eclexia - -3. **Medium-term**: Evaluate whether the moderate-fit repos would benefit from a lighter-weight "PanLL widget" concept (single-pane display) vs full 3-pane modules - -4. **Long-term**: Dedicated session to systematically implement the remaining strong-fit candidates - -== Summary - -[cols="1,1"] -|=== -| Category | Count - -| Already mapped -| 8 - -| Strong fit -| 18 - -| Moderate fit -| 22 - -| Not applicable -| 215+ - -| **Total evaluated** -| **~265** -|=== - -The hyperpolymath ecosystem has **26 repos** (8 + 18) that naturally fit the PanLL -three-pane model. With the 22 moderate-fit candidates, up to **48 PanLL modules** -could eventually populate the "multi-monitor trader desk" vision. diff --git a/verisimdb/docs/papers/verisimdb-federated-consistency.adoc b/verisimdb/docs/papers/verisimdb-federated-consistency.adoc deleted file mode 100644 index 93e33163..00000000 --- a/verisimdb/docs/papers/verisimdb-federated-consistency.adoc +++ /dev/null @@ -1,840 +0,0 @@ -// SPDX-License-Identifier: CC-BY-SA-4.0 -// Copyright (c) 2026 Jonathan D.A. Jewell (hyperpolymath) - -= VeriSimDB: Cross-Modal Drift Detection and Self-Normalisation in Heterogeneous Database Federations -:author: Jonathan D.A. Jewell -:email: j.d.a.jewell@open.ac.uk -:affiliation: The Open University, Milton Keynes, United Kingdom -:revnumber: 1.0 -:revdate: 2026-02-28 -:toc: left -:toclevels: 3 -:sectnums: -:stem: latexmath -:icons: font -:source-highlighter: rouge -:keywords: multimodal databases, drift detection, self-normalisation, formal verification, federation, data consistency - -// ============================================================================ -// ABSTRACT -// ============================================================================ - -[abstract] --- -Modern data systems increasingly store entities across multiple representation types -- graphs for relationships, vectors for similarity, documents for full-text search, time-series for temporal analysis -- in separate, specialised databases. -This polyglot persistence pattern creates consistency challenges that existing solutions cannot detect: when one representation of an entity is updated, others may silently diverge, producing _cross-modal drift_ invisible to any single store. -We present VeriSimDB, a multimodal database engine that maintains entities simultaneously across eight modalities (the "octad": Graph, Vector, Tensor, Semantic, Document, Temporal, Provenance, Spatial) and employs cross-modal drift detection to identify and repair representational divergence. -Our approach introduces a drift scoring algorithm based on cosine distance between modality embeddings, with adaptive thresholds that account for natural representational variance. -We describe VCL-UT, a query language extension that supports proof-carrying queries with 11 formally verified proof types, enabling clients to request cryptographic certificates attesting to the consistency, provenance, and integrity of query results. -We evaluate VeriSimDB on a suite of 510 Rust and 152 Elixir tests encompassing synthetic drift injection, federation coordination, and proof verification scenarios. -Results show that VeriSimDB detects semantic-vector drift with precision exceeding 0.95 and recall of 0.91, and repairs 94% of detected divergences automatically through its self-normalisation engine. -Federation overhead for cross-instance queries remains below 12ms per additional peer for modality-specific queries. --- - - -// ============================================================================ -// 1. INTRODUCTION -// ============================================================================ - -== Introduction - -The volume and heterogeneity of data managed by modern organisations has driven a fragmentation of storage systems along representational lines. -Graph databases such as Neo4j and Amazon Neptune store relational structures. -Vector stores including Pinecone, Milvus, and Weaviate manage embedding spaces for similarity search. -Document databases like Elasticsearch and Apache Solr handle full-text indexing. -Time-series databases (InfluxDB, TimescaleDB) track temporal evolution. -Each system excels at its specific representational paradigm, and the prevailing architectural advice -- often termed _polyglot persistence_ <> -- recommends selecting the best tool for each access pattern. - -This advice has a fundamental flaw: it assumes that the different representations of the same real-world entity can be managed independently. -In practice, entities are not cleanly partitioned. -A scientific paper exists simultaneously as a graph node (with citation edges), a vector (with semantic embedding), a document (with searchable content), a temporal entity (with revision history), a provenance record (with authorship chain), and a spatial entity (with affiliated institutional coordinates). -When a researcher retracts the paper, the graph store must remove citation edges, the vector store must invalidate the embedding, the document store must update the searchable text, and the provenance store must record the retraction event. -If any one of these updates fails or is delayed, the entity's representations diverge -- a condition we term _cross-modal drift_. - -Existing multi-model databases such as ArangoDB <>, SurrealDB, and Tigris offer multiple access patterns within a single system, but they do not detect or repair cross-modal inconsistencies. -They provide syntax for querying graphs, documents, and key-value pairs, but treat these as independent access patterns over the same underlying storage rather than as distinct, synchronised representations that must agree. -No existing system detects that a graph edge contradicts the document content, or that a vector embedding no longer reflects the semantic annotations. - -This paper makes four contributions: - -1. **The Octad Model** -- a formal entity model in which each entity exists simultaneously across eight modalities with well-defined interaction semantics and lifecycle events. - -2. **Cross-Modal Drift Detection** -- an algorithm that computes drift scores between modality pairs using cosine distance, Jaccard overlap, and temporal consistency measures, with adaptive thresholds that learn from historical drift patterns. - -3. **Self-Normalisation** -- an authority-ranked repair engine that identifies the most authoritative modality for a drifted entity and regenerates divergent representations, with five configurable strategies ranging from fully automatic to human-in-the-loop. - -4. **VCL-UT** -- a dependent-type extension to the VeriSim Consonance Language (VCL) that supports 11 proof types, enabling clients to request and verify formal certificates of data consistency, integrity, and provenance. - - -// ============================================================================ -// 2. BACKGROUND -// ============================================================================ - -== Background - -=== Multimodal and Multi-Model Databases - -The term _multi-model database_ refers to systems that support more than one data model within a single engine. -ArangoDB <> supports graph, document, and key-value access over the same storage layer. -OrientDB and CosmosDB offer similar capabilities. -SurrealDB adds vector search and graph traversal to a document core. -These systems reduce operational complexity by consolidating storage, but they share a critical limitation: they provide no mechanism to verify that the graph view, document view, and vector view of the same entity are mutually consistent. - -The _multimodal database_ concept, as we define it, is distinct from multi-model in a critical respect: each modality is a first-class, independently queryable representation with its own storage semantics, and the system actively monitors cross-modality consistency. -Where multi-model databases offer _syntactic_ polyglotism (multiple query languages over one store), multimodal databases offer _semantic_ polyglotism (multiple representation types with enforced agreement). - -=== Data Quality and Drift Detection - -The concept of _data drift_ originates in the machine learning literature, where it describes the phenomenon of production data distributions diverging from training data distributions <>. -Tools such as Evidently AI, Great Expectations, and Whylabs detect drift in tabular feature distributions, but they operate on single-modality datasets (typically feature vectors or structured tables) and do not address cross-modal consistency. - -In the database literature, consistency has traditionally been framed in terms of ACID transactions or eventual consistency models <>. -These models address the question of whether multiple _copies_ of the same data agree, not whether multiple _representations_ of the same entity agree. -A system can be perfectly consistent in the ACID sense (all replicas agree on the current state) while exhibiting severe cross-modal drift (the graph and document views contradict each other). - -=== Formal Verification in Database Systems - -Proof-carrying data has been explored in the context of authenticated data structures <>, where Merkle trees and hash chains provide integrity certificates. -CertikOS and seL4 demonstrate that formal verification can be applied to systems software, but database systems have largely resisted formal methods due to the complexity of query optimisation and storage management. - -VCL-UT draws inspiration from proof-carrying code <> and dependent type systems in languages such as Idris 2 <> and Lean 4 <>, adapting these ideas to the database query context. - -=== Marr's Three Levels of Analysis - -VeriSimDB's design philosophy is grounded in David Marr's three levels of analysis <>, adapted from computational neuroscience to database systems: - -1. **Computational level** -- _What problem are we solving?_ Maintaining cross-modal consistency across eight representations of the same entity, detecting drift before it causes data quality issues, and providing unified querying across all modalities. - -2. **Algorithmic level** -- _How do we solve it?_ Octad entities with one identifier and eight synchronised stores; drift detection with configurable thresholds; self-normalisation triggered by drift events; OTP supervision for fault tolerance. - -3. **Implementational level** -- _How is it built?_ Rust for performance-critical modality stores; Elixir/OTP for distributed coordination; HTTP API for inter-layer communication; Prometheus metrics for observability. - -This separation ensures that architectural decisions are traceable to the problem they solve, not to implementation convenience. - - -// ============================================================================ -// 3. THE OCTAD MODEL -// ============================================================================ - -== The Octad Model - -=== Design Rationale - -The octad comprises eight modalities, each addressing a distinct representational need. -The choice of eight is not arbitrary -- it emerges from an analysis of the irreducible representation types required by real-world data systems. -Earlier versions of VeriSimDB used a octad (six modalities), but operational experience demonstrated that provenance and spatial data require first-class treatment rather than being subsumed into other modalities. - -.The eight modalities of the VeriSimDB octad -[cols="1,2,2,1",options="header"] -|=== -| Modality | Representation | Storage Backend | Primary Query - -| Graph -| RDF triples and property graph edges -| Oxigraph (via redb pure-Rust backend) -| SPARQL patterns - -| Vector -| Dense float embeddings (f32) -| HNSW index (custom implementation) -| Approximate nearest neighbour - -| Tensor -| Multi-dimensional numeric arrays -| ndarray / Burn -| Slice and projection - -| Semantic -| Type annotations and CBOR proof blobs -| In-memory registry + CBOR serialisation -| Type matching and contract verification - -| Document -| Full-text searchable content -| Tantivy -| BM25 full-text search - -| Temporal -| Version history and time-series -| WAL-backed append-only log -| Time-range and version queries - -| Provenance -| Origin tracking, transformation chain, actor trail -| Hash-chain store -| Lineage traversal, actor search - -| Spatial -| Geospatial coordinates, geometries -| R-tree index -| Radius, bounding box, nearest-neighbour -|=== - -=== Entity Lifecycle - -An octad entity is created by supplying data to one or more modalities. -VeriSimDB assigns a unique Octad identifier (UUID v7, time-sortable) and initialises all eight modality slots. -Unpopulated modalities are marked as `absent` rather than null, distinguishing "no data provided" from "data was deleted." - -The lifecycle comprises four states: - -1. **Creation** -- Entity is initialised. Each provided modality is indexed. An initial consistency check runs to detect contradictions in the seed data. - -2. **Steady state** -- Entity exists across its populated modalities. The DriftMonitor periodically sweeps all entities, computing pairwise drift scores. Updates to any modality trigger an immediate drift check against related modalities. - -3. **Drift detected** -- One or more modality pairs exceed the configured drift threshold. The entity is flagged and placed on the normalisation queue. Depending on the drift policy, queries may still return the entity (with a drift annotation) or may block until normalisation completes. - -4. **Normalisation** -- The self-normalisation engine identifies the authoritative modality, regenerates the drifted modalities, validates the result, and atomically commits all changes. - -=== Modality Interactions and Dependencies - -Not all modality pairs have meaningful drift relationships. -VeriSimDB defines a _drift interaction matrix_ that specifies which pairs are monitored: - -- **semantic-vector**: The embedding should reflect the semantic types. Drift here indicates the embedding model disagrees with the type annotations. -- **graph-document**: Graph edges should correspond to entities mentioned in the document. Missing or extra edges indicate drift. -- **temporal-provenance**: The version history should be consistent with the provenance record. A version without a corresponding provenance entry indicates drift. -- **spatial-graph**: Spatial relationships (proximity, containment) should align with graph edges (same-zone, adjacent-to). Divergence indicates layout inconsistency. - -The drift interaction matrix is configurable per deployment, allowing operators to focus monitoring on the modality pairs most relevant to their data. - -=== Octad Identifier and Cross-Modal Linking - -The Octad ID serves as the universal join key across all eight modality stores. -Each store maintains its own index structures, but cross-modal queries are resolved by Octad ID lookup rather than by full-table scans. -The ID is stored as a 128-bit UUID in each modality's index, enabling constant-time cross-modal joins. - - -// ============================================================================ -// 4. DRIFT DETECTION -// ============================================================================ - -== Drift Detection - -=== The Drift Scoring Algorithm - -VeriSimDB's drift detection operates on pairs of modalities for a given entity. -For each monitored pair stem:[(M_i, M_j)], the system computes a drift score stem:[d_{ij} \in [0, 1\]] where 0 indicates perfect agreement and 1 indicates maximum divergence. - -The scoring functions are modality-pair-specific: - -**Semantic-vector drift.** Given an entity's embedding stem:[\mathbf{v} \in \mathbb{R}^n] and its semantic type annotations stem:[T = \{t_1, \ldots, t_k\}], each type stem:[t_i] has a pre-computed type embedding stem:[\mathbf{e}_i]. -The drift score is: - -[stem] -++++ -d_{\text{sem-vec}} = 1 - \frac{1}{k} \sum_{i=1}^{k} \frac{\mathbf{v} \cdot \mathbf{e}_i}{\|\mathbf{v}\| \|\mathbf{e}_i\|} -++++ - -This is one minus the average cosine similarity between the entity embedding and its type embeddings. -A score near 0 indicates the embedding faithfully represents the semantic types; a score near 1 indicates the embedding has diverged from what the types predict. - -**Graph-document drift.** Given entities mentioned in the document text stem:[E_{\text{doc}}] and entities connected by graph edges stem:[E_{\text{graph}}], the drift score is the complement of the Jaccard coefficient: - -[stem] -++++ -d_{\text{graph-doc}} = 1 - \frac{|E_{\text{doc}} \cap E_{\text{graph}}|}{|E_{\text{doc}} \cup E_{\text{graph}}|} -++++ - -**Temporal-consistency drift.** For each version in the temporal log, the system verifies that a corresponding provenance entry exists and that timestamps are monotonically ordered. -The score is the fraction of versions that fail these checks. - -**Tensor drift.** Tensor drift measures whether the tensor representation is consistent with the source modalities from which it was derived. -This is computed as the normalised Frobenius distance between the current tensor and a re-derived tensor: - -[stem] -++++ -d_{\text{tensor}} = \frac{\| T_{\text{current}} - T_{\text{rederived}} \|_F}{\max(\| T_{\text{current}} \|_F, \| T_{\text{rederived}} \|_F)} -++++ - -**Schema drift.** Schema drift detects violations of type constraints declared in the semantic modality. -The score is the fraction of declared constraints that are violated by the current data across all modalities. - -**Quality drift.** An aggregate score computed as a weighted mean of all other drift types, providing a single summary metric per entity. - -=== Per-Modality Drift Types - -VeriSimDB classifies drift into six named types, each corresponding to a specific modality interaction: - -.Drift type classification -[cols="2,3,1",options="header"] -|=== -| Drift Type | Detection Mechanism | Default Threshold (Warning / Critical) - -| `semantic_vector_drift` -| Cosine distance between entity embedding and type embeddings -| 0.3 / 0.7 - -| `graph_document_drift` -| Jaccard complement between document entities and graph edges -| 0.4 / 0.8 - -| `temporal_consistency_drift` -| Fraction of versions missing provenance entries -| 0.2 / 0.6 - -| `tensor_drift` -| Normalised Frobenius distance to re-derived tensor -| 0.35 / 0.75 - -| `schema_drift` -| Fraction of violated type constraints -| 0.1 / 0.5 - -| `quality_drift` -| Weighted aggregate of all other drift scores -| 0.25 / 0.65 -|=== - -=== Adaptive Thresholds - -Static thresholds are insufficient for heterogeneous deployments where natural representational variance differs between entity types. -For example, academic papers may exhibit low semantic-vector drift (embeddings closely track types), while social media posts may exhibit higher natural variance due to informal language and rapidly shifting topics. - -VeriSimDB implements adaptive thresholds using an exponentially weighted moving average (EWMA) of historical drift scores per entity type: - -[stem] -++++ -\theta_t = \alpha \cdot d_t + (1 - \alpha) \cdot \theta_{t-1} -++++ - -where stem:[\alpha = 0.1] by default. -The warning threshold is set at stem:[\theta_t + 2\sigma] and the critical threshold at stem:[\theta_t + 3\sigma], where stem:[\sigma] is the standard deviation of recent drift scores. -This allows the system to learn what "normal" drift looks like for each entity type and only alert when drift exceeds the expected range. - -=== DriftMonitor Architecture - -The DriftMonitor is implemented as an Elixir GenServer that coordinates drift detection across all entities. -It operates on a configurable sweep interval (default: 60 seconds) and maintains per-entity drift scores in ETS (Erlang Term Storage) for O(1) lookup. - -On each sweep: - -1. The DriftMonitor retrieves the list of entities modified since the last sweep. -2. For each modified entity, it calls the Rust `DriftCalculator` via HTTP to compute pairwise drift scores. -3. Scores exceeding warning thresholds are logged with telemetry events. -4. Scores exceeding critical thresholds trigger normalisation by placing the entity on the pending normalisation queue. -5. The DriftMonitor dispatches up to `max_concurrent_normalizations` (default: 10) repair operations concurrently. - -The Rust `DriftCalculator` performs the numerically intensive drift computation (cosine similarity, Jaccard coefficients, Frobenius norms) while the Elixir layer handles scheduling, concurrency, and fault tolerance through OTP supervision trees. - - -// ============================================================================ -// 5. SELF-NORMALISATION -// ============================================================================ - -== Self-Normalisation - -=== Five Normalisation Strategies - -When drift exceeds the critical threshold, VeriSimDB's normaliser selects a repair strategy. -Five strategies are available, ordered from most automatic to most manual: - -1. **FromAuthoritative** -- Regenerate the drifted modality from the highest-authority consistent modality. VeriSimDB maintains a default authority ranking: Document > Semantic > Provenance > Graph > Vector > Tensor > Spatial > Temporal. The rationale is that human-written content (Document) is the most authoritative source of truth, followed by explicitly declared type annotations (Semantic) and lineage records (Provenance). - -2. **Merge** -- Combine data from all non-drifted modalities, weighted by authority rank, to produce a consensus representation. This strategy is useful when no single modality is clearly authoritative and the drift involves minor discrepancies. - -3. **VectorRegeneration** -- A specialised strategy for semantic-vector drift that re-embeds the entity's document content through the configured embedding model. This is the most common normalisation operation, as embeddings are the most frequent source of drift. - -4. **FullReconciliation** -- Performs a complete cross-modal reconciliation, re-deriving each modality from the authoritative source and validating all pairwise consistency checks before committing. This is the most expensive strategy and is reserved for entities with multi-modality drift. - -5. **UserResolve** -- Flags the entity for manual resolution and adds it to a resolution queue accessible via the API. This strategy is invoked when automated repair confidence is below a configurable threshold (default: 0.6) or when the drift type is flagged as requiring human judgement. - -=== Authoritative Modality Selection - -The authority ranking is configurable per deployment and per entity type. -The default ranking (Document > Semantic > Provenance > Graph > Vector > Tensor > Spatial > Temporal) reflects the principle that human-curated data is more authoritative than computed representations. -However, in domains such as scientific computing where tensor data is the primary artefact, operators may promote Tensor to the top of the ranking. - -Selection proceeds as follows: - -1. Identify all non-drifted modalities (those with drift scores below the warning threshold for all pairs involving them). -2. From these, select the modality with the highest authority rank. -3. If no modality is fully non-drifted (cascading drift), select the modality with the lowest aggregate drift score. - -=== Atomic Cross-Modality Updates - -Normalisation must update multiple modality stores atomically to avoid introducing new drift during repair. -VeriSimDB achieves this through a two-phase approach: - -1. **Preparation phase**: The normaliser computes all regenerated modality values and validates their mutual consistency _before_ writing any of them. -2. **Commit phase**: All modality updates are applied within a single Elixir process, holding write locks on all affected modality stores for the duration of the commit. The write-ahead log (WAL) records all changes, enabling rollback if any individual store write fails. - -=== Verification After Normalisation - -After committing the normalised data, VeriSimDB immediately re-runs drift detection on the repaired entity. -If any drift scores remain above the warning threshold after normalisation, the entity is re-queued for a second normalisation pass with a more aggressive strategy. -If the second pass also fails, the entity is escalated to the UserResolve strategy. - -Normalisation results are recorded as telemetry events including the entity ID, drift type, strategy used, duration, and success status, enabling operators to monitor normalisation health via Prometheus dashboards. - - -// ============================================================================ -// 6. VCL AND VCL-UT -// ============================================================================ - -== VCL and VCL-UT - -=== VCL Syntax and Semantics - -The VeriSim Consonance Language (VCL) is the native query interface for VeriSimDB. -Its grammar is specified in ISO/IEC 14977 EBNF notation <> and supports `SELECT`, `INSERT`, `UPDATE`, and `DELETE` operations across octad entities. - -A VCL query comprises: - -[source,ebnf] ----- -query = select_clause, from_clause, [where_clause], - [group_by_clause], [having_clause], - [proof_clause], [order_by_clause], - [limit_clause], [offset_clause] ; ----- - -The `SELECT` clause specifies which modalities to retrieve, with optional per-modality projections: - -[source,vcl] ----- -SELECT GRAPH((?s)-[:CITES]->(?target)), DOCUMENT(title, abstract) -FROM STORE research_papers -WHERE CONTAINS(h.content, "drift detection") -LIMIT 50 ----- - -The `FROM` clause supports three source types: `HEXAD ` for direct entity access, `STORE ` for collection queries, and `FEDERATION ` for cross-instance queries. - -VCL is not SQL. -Key differences include: no table schemas (entities are schema-flexible with type annotations in the Semantic modality), modality-scoped projections rather than column selections, graph pattern matching in `WHERE` clauses, and cross-modal conditions such as `DRIFT_SCORE(semantic, vector) > 0.5`. - -=== The PROOF Clause and 11 Proof Types - -VCL-UT (VCL with Dependent Types) is activated by the presence of a `PROOF` clause. -When a query includes `PROOF`, the query router directs it through the formal verification pipeline rather than the standard (slipstream) execution path. - -The `PROOF` clause takes the form `PROOF ()` and supports 11 proof types: - -.VCL-UT proof types -[cols="1,3",options="header"] -|=== -| Proof Type | Guarantee - -| `EXISTENCE` -| The referenced entity exists in the target store and has data in the requested modalities at query time. - -| `INTEGRITY` -| The returned data matches the stored data exactly. The proof certificate includes a Merkle root or hash chain. - -| `CONSISTENCY` -| All populated modalities of the entity are mutually consistent (no cross-modal drift detected). - -| `PROVENANCE` -| The complete lineage chain of the entity is intact, from creation through all transformations. - -| `FRESHNESS` -| The data was retrieved within a specified time window (e.g., within the last 60 seconds). - -| `ACCESS` -| The querying principal has the required permissions for all requested modalities. - -| `CITATION` -| The entity's citation chain is verifiable and no cited entities have been retracted or deleted. - -| `CUSTOM` -| A user-defined proof contract that specifies arbitrary verification logic. - -| `ZKP` -| A zero-knowledge proof that the query result satisfies a predicate without revealing the underlying data. - -| `PROVEN` -| The result has been verified by the external `proven` library, producing a certificate-based JSON/CBOR attestation. - -| `SANCTIFY` -| A composite proof combining INTEGRITY, PROVENANCE, and CONSISTENCY into a single audit-grade certificate. -|=== - -Multiple proofs can be composed in a single query: - -[source,vcl] ----- -SELECT * -FROM HEXAD 550e8400-e29b-41d4-a716-446655440000 -PROOF INTEGRITY(DataIntegrityContract) - AND PROVENANCE(FullLineageContract) ----- - -=== Proof Obligation Generation and Certificate Verification - -When a VCL-UT query is executed, the proof pipeline generates a _proof obligation_ for each requested proof type. -The obligation specifies: - -* The entity (or entities) to verify. -* The modalities involved. -* The contract parameters (e.g., freshness window, access principal). - -The obligation is dispatched to the appropriate verifier (Rust-side for INTEGRITY and ZKP, Elixir-side for EXISTENCE and ACCESS, external `proven` library for PROVEN). -Upon successful verification, a proof certificate is generated in CBOR format and attached to the query result. -The certificate contains the proof type, the verified predicate, a timestamp, and a digital signature. - -Clients can independently verify proof certificates without querying VeriSimDB, enabling offline audit and compliance workflows. - -=== Composition Strategies - -VCL-UT supports three composition strategies for multi-proof queries: - -1. **Conjunction (AND)** -- All proofs must succeed for the query to return results. Failure of any proof aborts the query with a proof failure diagnostic. - -2. **Sequential** -- Proofs are evaluated in order. The result of one proof is available to subsequent proofs (e.g., EXISTENCE before INTEGRITY). - -3. **Independent** -- Each proof is evaluated independently. The query result includes a per-proof status, allowing clients to inspect which proofs succeeded and which failed. - - -// ============================================================================ -// 7. FEDERATION -// ============================================================================ - -== Federation - -=== Adapter-Based Federation Architecture - -VeriSimDB federation enables cross-instance queries over heterogeneous databases. -The federation layer is adapter-based: each external database type implements a `FederationAdapter` trait that translates VCL queries into the native query language of the target system and maps results back into octad entities. - -Four adapters are currently defined: - -1. **VeriSimDB-native** -- Peer VeriSimDB instances communicating via the HTTP API. This is the primary federation mode, supporting all VCL features including drift policies and proof clauses. - -2. **PostgreSQL** -- Queries relational tables via SQL, mapping rows to Document and Semantic modalities. Graph modalities require explicit relationship tables. - -3. **ArangoDB** -- Leverages ArangoDB's native graph and document capabilities, mapping AQL query results to Graph and Document modalities. - -4. **Elasticsearch** -- Maps full-text queries to the Document modality and structured filters to the Semantic modality. - -=== Query Decomposition and Result Merging - -When a federated query targets multiple peers, the QueryRouter decomposes it into per-peer sub-queries: - -1. The federation registry is consulted to determine which peers hold relevant data. -2. Each sub-query is scoped to the modalities available at the target peer. -3. Sub-queries are dispatched concurrently via the Elixir Task module. -4. Results are merged by Octad ID, with per-modality data assembled from whichever peer provided it. - -When multiple peers provide data for the same modality of the same entity, conflict resolution is governed by the query's drift policy. - -=== Drift Policy Filtering Across Federated Peers - -The `DRIFT POLICY` clause in federated queries specifies how cross-instance drift is handled: - -- **STRICT** -- Fail the query if any entity exhibits drift across peers. Suitable for compliance-critical queries. -- **REPAIR** -- Automatically trigger normalisation on drifted entities before returning results. Increases latency but guarantees consistency. -- **TOLERATE** -- Return data despite detected drift, annotating each result with its drift scores. Suitable for analytics and exploration. -- **LATEST** -- Use the most recent version from any peer, as determined by the temporal modality. Suitable for eventual-consistency workloads. - - -// ============================================================================ -// 8. IMPLEMENTATION -// ============================================================================ - -== Implementation - -=== Rust Core - -The Rust core comprises 18 crates organised as a Cargo workspace. -Performance-critical components -- modality stores, drift computation, API serving, and write-ahead logging -- are implemented in Rust to minimise latency and maximise throughput. - -Key crates include: - -- **verisim-graph**: Property graph and RDF triple store built on Oxigraph with a pure-Rust redb backend (eliminating the C++ dependency of RocksDB). -- **verisim-vector**: HNSW (Hierarchical Navigable Small World) index for approximate nearest-neighbour search over f32 embeddings. -- **verisim-document**: Full-text search engine built on Tantivy, a Rust implementation of Apache Lucene's indexing algorithms. -- **verisim-drift**: The `DriftCalculator` struct implementing all drift scoring functions (cosine similarity, Jaccard coefficients, Frobenius norms). -- **verisim-normalizer**: The self-normalisation engine with five strategies, authority ranking, conflict resolution, and audit logging. -- **verisim-provenance**: Hash-chain-based lineage tracking with actor search and integrity verification. -- **verisim-spatial**: R-tree index for geospatial queries (radius, bounding box, k-nearest). -- **verisim-octad**: The octad entity abstraction, builder pattern, and cross-modal linking logic. -- **verisim-api**: Actix-web HTTP server exposing RESTful endpoints for all operations plus federation peer management. -- **verisim-wal**: Write-ahead log for crash recovery and atomic multi-modality commits. - -=== Elixir/OTP Orchestration - -The Elixir layer coordinates distributed operations using OTP's battle-tested supervision and concurrency primitives: - -- **VeriSim.EntityServer**: A GenServer per entity, managing the entity's lifecycle and serialising concurrent updates. Entity servers are distributed across the cluster using Horde (a distributed DynamicSupervisor). -- **VeriSim.DriftMonitor**: A GenServer that sweeps entities on a configurable interval, dispatching drift computation to the Rust core and queuing normalisation tasks. -- **VeriSim.QueryRouter**: Routes VCL queries to the appropriate execution pipeline (slipstream vs VCL-UT) and handles federation decomposition. -- **VeriSim.SchemaRegistry**: Manages entity type schemas and proof contracts, coordinating with the Semantic modality store. -- **VeriSim.Consensus**: A KRaft-inspired (Kafka Raft) consensus implementation for metadata replication in clustered deployments. - -=== ReScript Components - -The VCL parser is implemented in ReScript, compiling to JavaScript and running under Deno. -This layer handles: - -- VCL tokenisation and parsing to AST. -- Federation registry management (mapping peer names to endpoints and capabilities). -- PanLL database module protocol integration (registering VeriSimDB as a panel-capable database in the broader ecosystem). - -=== Test Suite - -The test suite comprises 510 Rust tests and 152 Elixir tests, totalling 662 tests with 0 failures. -Test categories include: - -- **Unit tests**: Per-crate tests for each modality store, drift calculator, normaliser strategy, and VCL parser. -- **Integration tests**: Cross-crate tests verifying drift detection triggers normalisation, federation queries span peers, and proof certificates verify correctly. -- **Property tests**: Proptest-based generative testing for the drift calculator, verifying that drift scores are always in [0, 1] and that normalisation reduces drift scores. - -=== Container Deployment - -VeriSimDB is deployed as OCI containers built with Podman, using Chainguard wolfi-base images for minimal attack surface. -Two deployment profiles are supported: - -- **In-memory**: All modality stores use in-memory backends. Suitable for development and testing. -- **Persistent**: Graph (redb), Document (Tantivy on filesystem), WAL, and provenance (file-backed hash chain). Suitable for production. - -The `compose.toml` definition (for selur-compose or Podman Compose) specifies a three-service stack: verisim-api (Rust core), verisim-otp (Elixir orchestration), and svalinn (edge gateway with authentication and rate limiting). - - -// ============================================================================ -// 9. EVALUATION -// ============================================================================ - -== Evaluation - -=== Test Results Summary - -All 662 tests (510 Rust, 152 Elixir) pass with 0 failures across all supported platforms (Linux x86_64, Linux aarch64). -The Rust test suite executes in 14.3 seconds; the Elixir test suite in 8.7 seconds. - -=== Drift Detection Precision and Recall - -We evaluated drift detection on a synthetic dataset of 10,000 octad entities with injected drift. -Drift was injected by randomly modifying modality representations without updating related modalities, simulating partial updates and network partitions. - -.Drift detection performance on synthetic dataset -[cols="2,1,1,1",options="header"] -|=== -| Drift Type | Precision | Recall | F1 - -| semantic_vector_drift -| 0.96 -| 0.91 -| 0.93 - -| graph_document_drift -| 0.94 -| 0.89 -| 0.91 - -| temporal_consistency_drift -| 0.99 -| 0.97 -| 0.98 - -| tensor_drift -| 0.92 -| 0.88 -| 0.90 - -| schema_drift -| 0.98 -| 0.95 -| 0.96 - -| quality_drift (aggregate) -| 0.95 -| 0.92 -| 0.93 -|=== - -Temporal consistency drift achieves the highest precision and recall because its detection mechanism (checking for missing provenance entries) is exact rather than approximate. -Tensor drift has the lowest scores because re-deriving tensors introduces floating-point rounding differences that can be misclassified as drift. - -=== Self-Normalisation Success Rate - -Of 3,847 drift events detected in the synthetic evaluation, 3,617 (94.0%) were successfully normalised automatically: - -- **FromAuthoritative**: 2,891 events (79.9% of successful normalisations) -- **VectorRegeneration**: 412 events (11.4%) -- **Merge**: 198 events (5.5%) -- **FullReconciliation**: 116 events (3.2%) -- **UserResolve (escalated)**: 230 events (6.0% of total -- not auto-resolved) - -The mean normalisation latency was 23ms for single-modality repairs and 87ms for FullReconciliation. - -=== Query Latency Characteristics - -.P50 and P99 query latencies by modality (single-entity, single-modality) -[cols="2,1,1",options="header"] -|=== -| Query Type | P50 (ms) | P99 (ms) - -| Graph (SPARQL pattern) -| 2.1 -| 8.4 - -| Vector (k-NN, k=10) -| 1.8 -| 6.2 - -| Document (full-text) -| 3.4 -| 12.1 - -| Temporal (version range) -| 1.2 -| 4.3 - -| Cross-modal (3 modalities) -| 5.7 -| 18.9 - -| VCL-UT with INTEGRITY proof -| 8.3 -| 27.4 -|=== - -VCL-UT queries incur approximately 2-3x the latency of slipstream queries due to proof generation and certificate construction. - -=== Federation Overhead - -Federation overhead was measured by comparing single-instance queries against federated queries spanning 2, 5, and 10 peers: - -.Federation latency overhead per additional peer -[cols="2,1,1",options="header"] -|=== -| Configuration | Median Overhead / Peer (ms) | P99 Overhead / Peer (ms) - -| 2 peers -| 8.2 -| 22.1 - -| 5 peers -| 10.7 -| 31.4 - -| 10 peers -| 11.9 -| 38.2 -|=== - -Overhead scales sub-linearly due to concurrent sub-query dispatch. -The primary cost is network round-trip time to each peer, not query execution at the peer. - - -// ============================================================================ -// 10. RELATED WORK -// ============================================================================ - -== Related Work - -**Graph databases.** -Neo4j <> is the leading graph database, supporting Cypher queries over property graphs. -It excels at relationship traversal but provides no vector, tensor, or document modalities. -Amazon Neptune supports RDF and property graphs but similarly lacks cross-modal consistency checking. -Neither system detects drift between graph structure and other representations of the same entity. - -**Vector databases.** -Pinecone <>, Milvus <>, and Weaviate offer high-performance approximate nearest-neighbour search. -These systems are single-modality: they store and query vectors but do not maintain relationships, documents, or provenance for the entities those vectors represent. -Drift between an embedding and its source data is invisible to these systems. - -**Multi-model databases.** -ArangoDB <> and SurrealDB support graph, document, and key-value access patterns within a single engine. -CosmosDB provides multiple API frontends (SQL, MongoDB, Gremlin, Cassandra) over a common storage layer. -These systems offer syntactic convenience but do not monitor cross-model consistency. -An entity whose graph edges contradict its document content will not trigger any alert. - -**Metadata management.** -Apache Atlas <> and DataHub <> provide metadata cataloguing, lineage tracking, and governance. -They address the "what data exists and where did it come from" problem but do not perform runtime drift detection or automated repair. -Their lineage tracking is metadata-level (recording ETL pipeline provenance), not data-level (verifying that representations agree). - -**ML drift detection.** -Evidently AI <>, Great Expectations <>, and Whylabs provide statistical drift detection for machine learning pipelines. -They detect distributional drift in feature columns (covariate shift, concept drift, data quality rules) but operate on single-modality tabular data and do not address cross-representation consistency. - -**Authenticated data structures.** -Merkle trees and authenticated skip lists <> provide tamper-evidence for data stored in untrusted environments. -VeriSimDB's INTEGRITY proof type draws on this work, extending it from single-store integrity to cross-modal integrity. - -**Formal methods in databases.** -Amazon's use of TLA+ for DynamoDB <> demonstrates that formal methods can improve database reliability, but TLA+ is used for protocol verification, not query-level proofs. -VCL-UT is, to our knowledge, the first query language to support proof-carrying results with dependent-type semantics. - - -// ============================================================================ -// 11. FUTURE WORK -// ============================================================================ - -== Future Work - -**Lean 4 integration for VCL-UT type checking.** -The current VCL-UT type checker is implemented in Elixir with a bridge to the Rust ZKP module. -We plan to integrate Lean 4 as an external proof assistant, enabling VCL-UT contracts to be expressed as Lean theorems and verified by Lean's kernel. -This would provide the highest level of formal assurance for proof certificates. - -**GraphRAG application integration.** -Retrieval-Augmented Generation (RAG) systems that combine graph traversal with vector retrieval are an ideal use case for VeriSimDB's multimodal architecture. -A GraphRAG adapter would expose VeriSimDB's combined graph-vector-document modalities as a unified retrieval backend for large language model applications. - -**Performance optimisation at scale.** -The current evaluation is limited to 10,000 entities. -We plan to conduct benchmarks at 1 million and 10 million entity scales, profiling bottlenecks in drift detection sweep time, normalisation throughput, and federation query planning. - -**Formal verification of self-normalisation correctness.** -We aim to prove that the self-normalisation engine is _convergent_: that repeated normalisation of a drifted entity always reduces drift scores to below the warning threshold in bounded time. -This property is currently validated empirically but not formally proven. - -**Dynamic Raft membership.** -The current KRaft consensus implementation supports static membership. -Dynamic membership changes (adding and removing nodes without cluster restart) are planned for the next release, enabling elastic scaling of clustered deployments. - -**Cross-drift lineage tracking.** -When drift in entity A causes cascading drift in entities B and C (through graph edges or provenance chains), the current system detects each drift independently. -Future work will track drift _lineage_, enabling operators to identify and repair root-cause drift rather than treating symptoms. - - -// ============================================================================ -// 12. CONCLUSION -// ============================================================================ - -== Conclusion - -VeriSimDB addresses a gap in the database landscape that has grown in proportion to the adoption of polyglot persistence. -As organisations store entities across graphs, vectors, documents, time-series, and spatial databases, the consistency of those representations has become a silent reliability risk. -No individual database can detect that its view of an entity contradicts another database's view. - -The octad model provides a principled foundation for multimodal entity management, with eight modalities covering the irreducible representation types required by modern data systems. -Cross-modal drift detection transforms consistency from an invisible hope into a measurable, actionable metric. -Self-normalisation automates the most tedious and error-prone aspect of multi-system data management: keeping representations in agreement. - -VCL-UT extends this consistency guarantee to the query interface, enabling clients to not merely _hope_ that their data is consistent but to _prove_ it with cryptographic certificates. -The 11 proof types cover a range of assurance levels from basic existence checking to zero-knowledge proofs, allowing clients to select the appropriate level of verification for their use case. - -Federation extends these guarantees across organisational boundaries, enabling multi-institutional data sharing with configurable drift policies that balance consistency against latency and autonomy. - -VeriSimDB is open source under the Palimpsest License (PMPL-1.0-or-later), with the Rust core, Elixir orchestration layer, and VCL parser available at https://github.com/hyperpolymath/verisimdb. - - -// ============================================================================ -// REFERENCES -// ============================================================================ - -[bibliography] -== References - -- [[[sadalage2012]]] Sadalage, P.J. and Fowler, M. (2012). _NoSQL Distilled: A Brief Guide to the Emerging World of Polyglot Persistence_. Addison-Wesley Professional. -- [[[arangodb2023]]] ArangoDB GmbH. (2023). "ArangoDB: The Multi-Model Database." https://www.arangodb.com -- [[[lu2019]]] Lu, J., Liu, A., Dong, F., Gu, F., Gama, J. and Zhang, G. (2019). "Learning under Concept Drift: A Review." _IEEE Transactions on Knowledge and Data Engineering_, 31(12), pp. 2346-2363. -- [[[vogels2009]]] Vogels, W. (2009). "Eventually Consistent." _Communications of the ACM_, 52(1), pp. 40-44. -- [[[tamassia2003]]] Tamassia, R. (2003). "Authenticated Data Structures." In _Proceedings of the 11th European Symposium on Algorithms (ESA)_, pp. 2-5. Springer. -- [[[necula1997]]] Necula, G.C. (1997). "Proof-Carrying Code." In _Proceedings of the 24th ACM SIGPLAN-SIGACT Symposium on Principles of Programming Languages (POPL)_, pp. 106-119. -- [[[brady2021]]] Brady, E. (2021). _Idris 2: Quantitative Type Theory in Practice_. In _Proceedings of the 35th European Conference on Object-Oriented Programming (ECOOP)_. -- [[[demoura2021]]] de Moura, L. and Ullrich, S. (2021). "The Lean 4 Theorem Prover and Programming Language." In _Proceedings of the 28th International Conference on Automated Deduction (CADE)_, pp. 625-635. -- [[[marr1982]]] Marr, D. (1982). _Vision: A Computational Investigation into the Human Representation and Processing of Visual Information_. W.H. Freeman and Company. -- [[[iso14977]]] ISO/IEC 14977:1996. _Information technology -- Syntactic metalanguage -- Extended BNF_. -- [[[neo4j2023]]] Neo4j, Inc. (2023). "Neo4j Graph Database." https://neo4j.com -- [[[pinecone2023]]] Pinecone Systems, Inc. (2023). "Pinecone: The Vector Database for Machine Learning." https://www.pinecone.io -- [[[wang2021]]] Wang, J., Yi, X., Guo, R., Jin, H., Xu, P., Li, S., Wang, X., Guo, X., Li, C., Xu, X. et al. (2021). "Milvus: A Purpose-Built Vector Data Management System." In _Proceedings of the 2021 ACM SIGMOD International Conference on Management of Data_, pp. 2614-2627. -- [[[atlas2023]]] Apache Software Foundation. (2023). "Apache Atlas: Data Governance and Metadata Framework." https://atlas.apache.org -- [[[datahub2022]]] Acryl Data, Inc. (2022). "DataHub: A Generalized Metadata Search & Discovery Tool." https://datahubproject.io -- [[[evidentlyai2023]]] Evidently AI. (2023). "Evidently: Open-Source ML Monitoring." https://www.evidentlyai.com -- [[[ge2023]]] Great Expectations. (2023). "Great Expectations: Data Quality Framework." https://greatexpectations.io -- [[[newcombe2015]]] Newcombe, C., Rath, T., Zhang, F., Muehlfeld, B., Brooker, M. and Deardeuff, M. (2015). "How Amazon Web Services Uses Formal Methods." _Communications of the ACM_, 58(4), pp. 66-73. diff --git a/verisimdb/docs/papers/verisimdb-idaptik-case-study.adoc b/verisimdb/docs/papers/verisimdb-idaptik-case-study.adoc deleted file mode 100644 index 52a263da..00000000 --- a/verisimdb/docs/papers/verisimdb-idaptik-case-study.adoc +++ /dev/null @@ -1,737 +0,0 @@ -// SPDX-License-Identifier: CC-BY-SA-4.0 -// Copyright (c) 2026 Jonathan D.A. Jewell (hyperpolymath) - -= Applying Multimodal Database Technology to Game Level Architecture: A VeriSimDB Case Study -:author: Jonathan D.A. Jewell -:email: j.d.a.jewell@open.ac.uk -:affiliation: The Open University, Milton Keynes, United Kingdom -:revnumber: 1.0 -:revdate: 2026-02-28 -:toc: left -:toclevels: 3 -:sectnums: -:stem: latexmath -:icons: font -:source-highlighter: rouge -:keywords: multimodal databases, game level design, dependent types, formal verification, Idris2, VeriSimDB - -// ============================================================================ -// ABSTRACT -// ============================================================================ - -[abstract] --- -Game level data is inherently multimodal: spatial layouts define geometry, entity -relationships form a graph, temporal state tracks version history, and provenance -records the designer's decision trail. Traditional approaches store this data in -flat configuration files or single-model databases, sacrificing cross-modal -consistency guarantees. We present a case study of IDApTIK, a spy infiltration -game whose Level Architect component uses VeriSimDB -- a multimodal database -with cross-modal drift detection -- as its persistence layer. The level data model -is defined in Idris2 with dependent-type proofs that guarantee referential -integrity, spatial ordering, and defence configuration consistency at compile time. -The Zig FFI layer provides C-compatible bindings with zero runtime overhead. We -describe the octad mapping of game entities to VeriSimDB's 8 modalities, the -14-module Idris2 ABI architecture, and the cross-domain proof system. The result -is a level persistence system where invalid levels cannot be stored -- not by -runtime validation, but by type-level impossibility. We discuss lessons learned -about the overhead of the octad model for game data, the value of formal -verification in game design tooling, and guidance on when VeriSimDB's multimodal -approach is warranted versus simpler alternatives. --- - -// ============================================================================ -// 1. EXECUTIVE SUMMARY -// ============================================================================ - -== Executive Summary - -IDApTIK is a spy infiltration game in which players navigate a building -defended by guards, dogs, drones, electronic devices, and an assassin. The -game's Level Architect is a design tool that creates, validates, and iterates -on level configurations. Level data spans multiple representation types -- -spatial coordinates for zone layouts, graph relationships between guards and -patrol zones, temporal version history of design iterations, and provenance -records of designer decisions. - -This paper describes the integration of VeriSimDB as the persistence layer for -the IDApTIK Level Architect. VeriSimDB's octad model (8 simultaneous data -modalities) provides a natural fit for game level data. Each level entity is -stored across Graph, Vector, Semantic, Document, Temporal, Provenance, and -Spatial modalities simultaneously, with cross-modal drift detection ensuring -that spatial layouts, entity relationships, and design metadata remain -consistent across all representations. - -The level data model is defined in Idris2, a dependently-typed programming -language, providing compile-time guarantees that go beyond what conventional -type systems can express. Five cross-domain proofs verify invariants such as -referential integrity of device IP addresses, guard placement within valid -zones, monotonic zone ordering, and PBX configuration consistency. These proofs -carry zero runtime overhead through Idris2's quantity erasure mechanism. - -The Zig FFI layer translates between the Idris2 ABI definitions and C-compatible -bindings consumed by VeriSimDB's Rust core. JSON serialisation preserves -round-trip fidelity with the game's ReScript frontend. - -=== Key Results - -* *14 Idris2 modules* defining the complete level data model -* *5 cross-domain proofs* verified at compile time with zero `believe_me` usage -* *Zero runtime overhead* from proof terms (quantity 0 erasure) -* *7 of 8 octad modalities* mapped to game level concepts -* *Type-safe persistence* -- invalid levels cannot be stored in VeriSimDB - -// ============================================================================ -// 2. THE CHALLENGE -// ============================================================================ - -== The Challenge - -=== Game Level Data Is Inherently Multimodal - -A game level is not a single data structure. It is a collection of interrelated -representations that must remain consistent with each other. Consider the data -required to fully describe an IDApTIK level: - -Spatial layout:: - Zone coordinates, guard positions, device placements, physical grid dimensions, - room boundaries, and corridor geometry. This data is inherently spatial -- it has - coordinates, extents, and proximity relationships. - -Entity relationships:: - Guards patrol specific zones. Devices wire to other devices. Drones follow - defined routes through zones. The assassin targets specific areas. These - relationships form a graph -- a directed, typed graph with domain-specific edge - semantics. - -Temporal state:: - Level designs evolve through iterations. A designer creates version 1, playtests, - adjusts guard positions, adds a new device, playtests again. Each version must be - preserved for comparison, rollback, and branching. This is temporal data -- - time-versioned state with branching semantics. - -Provenance:: - Why was this guard placed here? Which playtest session revealed the blind spot? - Who approved the final level layout? Provenance data tracks the designer's - decision trail -- the origin and justification for every element. - -Similarity:: - When designing a new level, a designer wants to find existing levels that are - similar in difficulty, layout complexity, or enemy density. This requires - vector similarity search over level embeddings. - -Semantic meaning:: - Device types, guard behaviours, zone classifications, and mission objectives - carry semantic meaning defined by the game's type system. This meaning must be - machine-readable and verifiable. - -Searchable descriptions:: - Levels have names, briefing texts, mission descriptions, and designer notes. - These must be full-text searchable. - -=== The Consistency Problem - -When level data is stored in a single flat configuration file (as is conventional -in game development), cross-modal consistency is trivially maintained -- there is -only one representation. However, this approach sacrifices queryability, version -history, similarity search, and provenance tracking. - -When level data is distributed across multiple specialised stores -- a graph -database for relationships, a spatial index for geometry, a document store for -descriptions -- the representations can diverge. A guard might be placed in a -zone that no longer exists in the spatial layout. A device might reference an IP -address that was removed from the device registry. A level description might -claim "3 guards" when the graph shows 5 guard-zone edges. - -These inconsistencies are typically caught by runtime validation -- if they are -caught at all. Runtime validation is reactive: it detects errors after they -occur, not before they can be introduced. It provides no formal guarantee that -the validation is complete. - -=== The Formal Verification Gap - -Game development rarely employs formal verification. Level data models are -typically defined in dynamically-typed languages (JSON schemas, Lua tables, -Unity ScriptableObjects) or statically-typed but not dependently-typed languages -(C# classes, Rust structs). These type systems can express "a guard has a zone -ID" but cannot express "a guard's zone ID refers to a zone that exists in the -level's zone list." - -The gap between what the type system can express and what the domain requires -leads to a class of bugs that are invisible to the compiler: - -* Dangling references (device IP not in registry) -* Ordering violations (zones not spatially sorted) -* Configuration inconsistencies (PBX IP set but PBX disabled) -* Cross-entity invariant violations (defence target not in device list) - -// ============================================================================ -// 3. SOLUTION ARCHITECTURE -// ============================================================================ - -== Solution Architecture - -=== VeriSimDB as Persistence Layer - -VeriSimDB provides a natural mapping for game level data through its octad model. -The Level Architect stores each level as a VeriSimDB entity with representations -across 7 of the 8 available modalities: - -[cols="1,2,4"] -|=== -| Modality | Level Data | Example - -| *Graph* -| Entity relationships -| Guards patrol zones (guard -> patrols -> zone). Devices wire to zones -(device -> wired_to -> zone). Drones follow routes (drone -> route -> [zone_1, -zone_2, ...]). Assassin targets areas (assassin -> targets -> zone). - -| *Vector* -| Level similarity embeddings -| Numeric representation of level characteristics (difficulty, enemy density, -device complexity, zone count, spatial compactness) enabling similarity search -across the level catalogue. "Find levels similar to this one in difficulty." - -| *Semantic* -| Idris2 ABI type annotations -| Dependent-type proofs as semantic metadata. `InRegistry guard_ip device_list` -proves referential integrity. `ZonesOrdered zone_transitions` proves spatial -monotonicity. Stored as CBOR proof blobs. - -| *Document* -| Searchable level descriptions -| Level name, mission briefing text, designer notes, walkthrough hints. Full-text -searchable via Tantivy. "Find all levels mentioning 'laser grid' in the briefing." - -| *Temporal* -| Version history -| Every save creates a temporal snapshot. Designers can diff versions, branch -level variants (easy/hard), and roll back to any previous state. Bitemporal -queries: "What did the level look like at design time T when queried at time Q?" - -| *Provenance* -| Designer decision trail -| Which designer placed each guard. Which playtest session triggered each change. -Which validation run approved the final layout. Actor trail, hash-chain integrity. - -| *Spatial* -| Zone coordinates, positions -| Zone boundaries (x-start, x-end, y-start, y-end), guard positions within zones, -device placements, physical grid coordinates. Spatial queries: "Find all devices -within 5 grid units of this guard." -|=== - -The 8th modality, *Tensor*, is not currently mapped. Tensor representation could -be used for ML-based level balancing (feature matrices of enemy placement patterns), -but this is future work. - -=== Architectural Decisions - -==== AD-1: Level Architect Only - -VeriSimDB is used by the Level Architect only -- not by the game itself. The main -game receives static, exported `LevelConfig.res` files. This decision is deliberate: - -* The game is a single-player Tauri desktop app requiring fast, offline, read-only - access to level data -* A database adds runtime complexity with no benefit for the game -* VeriSimDB's federated mode provides an escape hatch if a game-side database is - ever needed (e.g., level marketplace, multiplayer lobby) - -==== AD-2: Idris2 ABI as Canonical Source of Truth - -All level types are defined in Idris2 with dependent-type proofs. The ReScript -types in the main game and level architect derive from these. The Idris2 ABI is -the single source of truth. - -==== AD-3: JSON Serialisation via Zig FFI - -Levels are serialised as JSON. The Zig FFI layer handles parsing and emitting. -Level files are small (< 100 KB), making text format preferable to binary for -round-trip verification and human readability. - -=== Cross-Modal Drift Detection - -VeriSimDB's drift detection provides a unique benefit for level design: it catches -inconsistencies between representations automatically. Examples: - -Spatial-Graph drift:: - A guard's graph edge says it patrols Zone C, but Zone C's spatial coordinates - were deleted in the last edit. Drift score increases, flagging the inconsistency - before export. - -Document-Graph drift:: - The level briefing says "two guards patrol the east wing," but the graph shows - three guard-zone edges for the east wing. Document-graph drift detection flags - the stale description. - -Temporal-Semantic drift:: - The current version's type proofs pass, but the previous version's proofs are - invalidated by a schema change. Temporal-semantic drift detection identifies - which historical versions need re-validation. - -// ============================================================================ -// 4. IMPLEMENTATION -// ============================================================================ - -== Implementation - -=== Idris2 ABI: 14 Modules - -The level data model is implemented across 14 Idris2 modules, organised in a -dependency graph from primitive types at the leaves to the composed level at -the root. - -[cols="1,3,2"] -|=== -| Module | Purpose | Key Types - -| *Primitives* -| Base types shared across all modules -| `IPv4`, `DeviceIP`, `ZoneId`, `GridCoord`, `Percentage` - -| *Types* -| Core game enumerations -| `DifficultyLevel`, `MissionType`, `WeaponClass`, `DeviceCategory` - -| *Devices* -| Electronic devices in the building -| `Device`, `DeviceRegistry`, `DeviceDefenceConfig` - -| *Zones* -| Building zones with spatial extent -| `ZoneTransition`, `ZoneDifficulty`, `ZoneConfig` - -| *Inventory* -| Player starting equipment -| `InventoryItem`, `InventorySlot`, `StartingLoadout` - -| *Guards* -| Human guards with patrol assignments -| `GuardPlacement`, `GuardBehaviour`, `PatrolRoute` - -| *Dogs* -| Guard dogs with territory definitions -| `DogPlacement`, `DogBehaviour`, `Territory` - -| *Drones* -| Aerial drones with flight paths -| `DronePlacement`, `DroneRoute`, `DetectionCone` - -| *Assassin* -| The assassin enemy (singular per level) -| `AssassinConfig`, `TargetPriority`, `StalkPattern` - -| *Mission* -| Mission objectives and win conditions -| `MissionObjective`, `WinCondition`, `TimerConfig` - -| *Wiring* -| Device-to-device and device-to-zone wiring -| `WireConnection`, `CircuitConfig`, `PBXConfig` - -| *Physical* -| Physical grid and room layout -| `PhysicalGrid`, `RoomBoundary`, `CorridorSegment` - -| *Level* -| Composed level (all modules combined) -| `LevelData`, `LevelConfig`, `LevelMetadata` - -| *Validation* -| Cross-domain proof orchestration -| `ValidatedLevel`, proof witness types -|=== - -==== Module Dependency Graph - -.... - Level.idr - / | \ - / | \ - Mission Validation Physical - | / | \ | - Inventory / | \ Wiring - | / | \ | - Guards Dogs Drones Assassin - \ | / / - \ | / / - Devices - | - Network - | - Primitives - | - Types.idr -.... - -All module names are bare (e.g., `module Primitives`, not `module Abi.Primitives`) -because Idris2's `sourcedir = "src/abi"` in the `.ipkg` file sets the root. -Modules are listed comma-separated on a single line in the `.ipkg` file due to -an Idris2 parser requirement. - -=== Cross-Domain Proofs - -Five proof types enforce cross-domain invariants at compile time: - -[cols="2,2,4"] -|=== -| Proof | Signature | Guarantee - -| *InRegistry* -| `IPv4 -> DeviceRegistry -> Type` -| An IP address referenced by a guard, drone, or wiring configuration actually -exists in the level's device list. Prevents dangling device references. - -| *GuardsInZones* -| `List GuardPlacement -> List ZoneTransition -> Type` -| Every guard's assigned zone appears in the level's zone transition list. -Prevents guards from being assigned to nonexistent zones. - -| *DefenceTargetsValid* -| `List DeviceDefenceConfig -> DeviceRegistry -> Type` -| Every `failoverTarget`, `cascadeTrap`, and `mirrorTarget` IP in the defence -configuration refers to a device in the registry. Prevents defence systems -from targeting nonexistent devices. - -| *ZonesOrdered* -| `List ZoneTransition -> Type` -| Zone x-coordinates are monotonically increasing (left to right). Prevents -overlapping or disordered zone layouts that would break spatial queries. - -| *PBXConsistent* -| `Bool -> Maybe IPv4 -> Type` -| The PBX IP address is set if and only if `hasPBX` is `True`. Prevents -configuration states where a PBX IP exists but PBX is disabled, or PBX is -enabled but no IP is configured. -|=== - -==== Proof Construction Strategy - -Proofs are constructed using `So`-based witnesses via `decSo` on boolean -equality, rather than full `DecEq` instances. This approach avoids the need to -prove `Not (a = b)` for the negative case of `DecEq` on `Bits8`, which would -require `believe_me` -- a banned pattern. - -[source,idris] ----- --- So-based witness: safe, no believe_me -InRegistry : IPv4 -> DeviceRegistry -> Type -InRegistry ip reg = So (elem ip (registeredIPs reg)) - --- Constructed via decSo (from Data.So in stdlib) -checkInRegistry : (ip : IPv4) -> (reg : DeviceRegistry) -> Dec (InRegistry ip reg) -checkInRegistry ip reg = decSo (elem ip (registeredIPs reg)) ----- - -=== ValidatedLevel Record - -The `ValidatedLevel` record bundles a `LevelData` value with all 5 proof terms. -The proofs are erased at runtime (quantity 0) -- they carry zero overhead in the -compiled output. - -[source,idris] ----- -record ValidatedLevel where - constructor MkValidatedLevel - levelData : LevelData - 0 devicesValid : InRegistry (allDeviceIPs levelData) (deviceRegistry levelData) - 0 guardsValid : GuardsInZones (guards levelData) (zones levelData) - 0 defenceValid : DefenceTargetsValid (defences levelData) (deviceRegistry levelData) - 0 zonesOrdered : ZonesOrdered (zones levelData) - 0 pbxConsistent : PBXConsistent (hasPBX levelData) (pbxIP levelData) ----- - -A `ValidatedLevel` can only be constructed if all 5 proofs are satisfied. The -`0` quantity annotation means the proof values are erased during compilation -- -they exist only at type-checking time and contribute zero bytes to the runtime -representation. - -=== Zig FFI Layer - -The Zig FFI provides C-compatible bindings between the Idris2 ABI and VeriSimDB's -Rust core: - -[cols="1,4"] -|=== -| Component | Responsibility - -| *JSON parser* -| Parses level JSON into C structs compatible with the Idris2 ABI types. -Validates structural integrity before passing to the Idris2 layer. - -| *JSON emitter* -| Serialises validated level data back to JSON for export to the game's -`LevelConfig.res` format. Preserves field ordering for diff-friendly output. - -| *C header generation* -| Auto-generated C headers from Idris2 ABI definitions bridge the Idris2 and -Zig type systems. Located in `generated/abi/`. - -| *Memory management* -| Arena allocator for level data parsing. No dynamic allocation in the -serialisation path. Deterministic cleanup. - -| *Cross-compilation* -| Zig's built-in cross-compilation supports targeting all platforms the game -runs on (Linux, macOS, Windows via Tauri) from a single build. -|=== - -// ============================================================================ -// 5. RESULTS -// ============================================================================ - -== Results - -=== Type-Safe Level Validation - -The integration achieves compile-time guarantees that are impossible in -conventional game development workflows: - -Invalid device references:: - A level where a guard references device IP `192.168.1.50` but the device - registry only contains `192.168.1.1` through `192.168.1.10` will fail to - compile. The `InRegistry` proof cannot be constructed. - -Orphaned guards:: - A level where a guard is assigned to "Zone F" but the zone list only contains - Zones A through E will fail to compile. The `GuardsInZones` proof cannot be - constructed. - -Misordered zones:: - A level where Zone B's x-coordinate is less than Zone A's will fail to compile. - The `ZonesOrdered` proof cannot be constructed. - -These are not runtime errors caught by validation logic. They are compile-time -impossibilities enforced by the type system. - -=== Zero Runtime Overhead - -All proof terms use quantity 0 erasure. The compiled `ValidatedLevel` is identical -in memory layout to a plain `LevelData` -- the proofs are erased entirely. There -is no performance penalty for formal verification. - -Measurement confirms this: serialisation and deserialisation benchmarks show no -measurable difference between `LevelData` and `ValidatedLevel` in the Zig FFI -layer, because the Zig layer never sees the proof terms. - -=== Cross-Modal Consistency - -VeriSimDB's drift detection provides ongoing consistency monitoring beyond what -compile-time proofs alone can achieve: - -* *Compile-time proofs* verify that a level is internally consistent at the moment - of creation or modification -* *Drift detection* verifies that the level's multiple representations in VeriSimDB - remain consistent over time -- that the graph, spatial, document, and semantic - representations have not diverged due to partial updates, schema evolution, or - external modifications - -The two mechanisms are complementary: proofs guarantee initial consistency, drift -detection maintains it. - -=== Zero `believe_me` Usage - -The entire 14-module ABI contains zero instances of `believe_me`, `assert_total`, -or `assert_smaller`. Every proof is constructive. Every function is total. This -is significant because `believe_me` is an axiom that asserts an arbitrary -proposition -- its presence undermines the value of every other proof in the -module. Its absence means the proofs are genuine. - -For context, a survey of Idris2 projects on GitHub found that `believe_me` usage -ranges from 2% to 15% of proof obligations in typical projects. The IDApTIK ABI -achieves 0% through careful proof strategy (So-based witnesses rather than DecEq). - -// ============================================================================ -// 6. LESSONS LEARNED -// ============================================================================ - -== Lessons Learned - -=== When to Use VeriSimDB vs. Simpler Alternatives - -VeriSimDB's octad model is not always the right choice. The decision depends on -how many distinct representation types the domain naturally requires: - -[cols="1,2,3"] -|=== -| Modalities Needed | Recommendation | Rationale - -| 1-2 -| Use a specialised database (PostgreSQL, Neo4j, etc.) -| The octad model adds conceptual overhead without benefit when data is -naturally single-modal. - -| 3-4 -| Consider a multi-model database (ArangoDB, SurrealDB) or VeriSimDB -| Multi-model databases are simpler if you do not need drift detection. VeriSimDB -is warranted if cross-modal consistency matters. - -| 5+ -| VeriSimDB is strongly indicated -| No existing database supports 5+ simultaneous modalities with consistency -checking. The octad model earns its overhead. -|=== - -Game level data naturally spans 5-7 modalities (spatial, graph, document, -temporal, provenance, semantic, optionally vector). This places it firmly in the -range where VeriSimDB's approach is justified. - -However, a simpler game -- one with flat levels, no version history, and no -design tools -- should use flat configuration files. VeriSimDB is a tool for -level *design*, not level *runtime*. - -=== Overhead of the Octad Model - -The octad model introduces overhead in three areas: - -Conceptual overhead:: - Developers must think about their data in terms of modalities. For game - developers accustomed to flat config files, this requires a mental shift. The - mapping exercise (Section 3.1) is essential but non-trivial. - -Storage overhead:: - Storing a level across 7 modalities uses more disk space than a single JSON - file. For IDApTIK (levels < 100 KB), this is negligible. For games with - millions of procedurally generated levels, it could be significant. - -Query complexity:: - Cross-modal queries (e.g., "find all guards in zones near a specific device, - sorted by level similarity") require understanding VCL's modality-spanning - syntax. Single-modality queries are straightforward. - -The overhead is justified when the benefits -- drift detection, version history, -similarity search, provenance tracking, and formal verification -- outweigh the -learning curve. For a level design tool used by a small team, the benefits are -clear. For a runtime game engine serving millions of players, the trade-off -is different. - -=== Value of Formal Verification in Game Design - -Formal verification is rarely applied to game data. The IDApTIK case study -suggests it is valuable in specific contexts: - -*Where it helps:* - -* Level design tools where an invalid configuration causes hard-to-diagnose - runtime bugs (e.g., a guard patrolling a nonexistent zone causes a null - reference crash during gameplay) -* Modding communities where user-created content must be validated before - distribution -* Competitive games where level fairness is a design requirement (provable - zone ordering guarantees symmetry properties) - -*Where it does not help:* - -* Rapid prototyping where level configurations change every few minutes -* Procedural generation where levels are created algorithmically (the generator - should enforce invariants, not the data model) -* Games with simple, flat level structures (a 2D platformer with tile maps - does not benefit from dependent types) - -=== Idris2 Practical Considerations - -Several Idris2 quirks affected the implementation: - -Module naming:: - Module names must be bare (not qualified) when `sourcedir` is set in the - `.ipkg` file. This differs from Haskell convention. - -Module listing:: - The `.ipkg` file requires modules to be listed comma-separated on a single - line. Multi-line module lists cause parser errors in Idris2 0.8.0. - -Pattern matching:: - Idris2 0.8.0 has a known bug with dependent type pattern matching that - prevented `Layout.idr` and `Foreign.idr` from being included in the build. - These modules are excluded pending a compiler fix. - -Proof construction:: - `DecEq` for `Bits8` requires `believe_me` for the negative case. Using - `So`-based witnesses avoids this entirely but requires a different proof - style that may be unfamiliar to Idris2 developers. - -// ============================================================================ -// 7. CONCLUSION -// ============================================================================ - -== Conclusion - -The IDApTIK Level Architect integration demonstrates that VeriSimDB's multimodal -architecture provides genuine value for game level data management -- a domain -that is inherently multimodal but traditionally treated as single-modal. - -The key contributions of this case study are: - -1. *A concrete mapping* from game level concepts to VeriSimDB's octad modalities, - demonstrating that 7 of 8 modalities have natural game-domain interpretations - -2. *A dependently-typed ABI* in Idris2 that proves cross-domain invariants at - compile time with zero runtime overhead, eliminating entire classes of level - configuration bugs - -3. *A complementary consistency model* where compile-time proofs guarantee initial - correctness and VeriSimDB's drift detection maintains correctness over time - -4. *Honest guidance* on when the octad model is warranted versus when simpler - alternatives are preferable - -The broader implication is that multimodal databases are not limited to enterprise -data integration or scientific computing. Any domain with inherently multimodal -data -- and game design is one such domain -- can benefit from cross-modal -consistency guarantees. The question is not whether the data is multimodal -(it usually is), but whether the consistency between modalities matters enough -to warrant the overhead of tracking it. - -For IDApTIK, where an invalid level configuration can crash the game or create an -unfair player experience, the answer is yes. - -// ============================================================================ -// REFERENCES -// ============================================================================ - -== References - -[bibliography] -* [[[brady2013]]] Edwin Brady, _Idris, a General Purpose Dependently Typed Programming Language: Design and Implementation_, Journal of Functional Programming, 23(5), 2013. -* [[[verisimdb2026]]] Jonathan D.A. Jewell, _VeriSimDB: Cross-Modal Drift Detection and Self-Normalisation in Heterogeneous Database Federations_, Technical Report, 2026. -* [[[tauri2025]]] Tauri Contributors, _Tauri 2.0: Build Smaller, Faster, and More Secure Desktop and Mobile Applications_, https://tauri.app/, 2025. -* [[[rescript2025]]] ReScript Contributors, _ReScript: Fast, Simple, Fully Typed JavaScript from the Future_, https://rescript-lang.org/, 2025. -* [[[zig2025]]] Andrew Kelley et al., _Zig Programming Language_, https://ziglang.org/, 2025. -* [[[oxigraph2024]]] Thomas Tanon, _Oxigraph: A SPARQL-Compliant Graph Database Written in Rust_, https://oxigraph.org/, 2024. - -// ============================================================================ -// APPENDIX -// ============================================================================ - -[appendix] -== Octad Modality Mapping Reference - -[cols="1,1,1,3"] -|=== -| Modality | Used | IDApTIK Entity | Notes - -| Graph | Yes | Guard-zone edges, device wiring, drone routes | Primary relationship model -| Vector | Yes | Level similarity embeddings | Used for "find similar levels" -| Tensor | No | (Future: ML balancing features) | Not yet mapped -| Semantic | Yes | Idris2 proof blobs (CBOR) | Type-level annotations -| Document | Yes | Level names, briefings, notes | Full-text searchable -| Temporal | Yes | Version history, design iterations | Bitemporal queries -| Provenance | Yes | Designer decisions, playtest origins | Hash-chain integrity -| Spatial | Yes | Zone coordinates, guard positions | R-tree indexed -|=== - -[appendix] -== Cross-Domain Proof Summary - -[cols="2,3,2"] -|=== -| Proof | What It Prevents | Construction Method - -| InRegistry | Dangling device IP references | So + decSo on elem -| GuardsInZones | Guards in nonexistent zones | So + decSo on elem -| DefenceTargetsValid | Defence targeting nonexistent devices | So + decSo on elem -| ZonesOrdered | Disordered zone x-coordinates | So + decSo on (<=) -| PBXConsistent | PBX IP/enabled flag mismatch | Pattern matching on (Bool, Maybe) -|=== diff --git a/verisimdb/docs/query-optimization-overview.adoc b/verisimdb/docs/query-optimization-overview.adoc deleted file mode 100644 index f1f16521..00000000 --- a/verisimdb/docs/query-optimization-overview.adoc +++ /dev/null @@ -1,483 +0,0 @@ -// SPDX-License-Identifier: CC-BY-SA-4.0 - -= VCL Query Optimization: Complete Guide -:toc: left -:toclevels: 3 -:sectnums: - -== Overview - -VeriSimDB provides comprehensive query optimization through three integrated systems: - -1. **Query Planning** - Cost-based optimization with selectivity estimation -2. **Bidirectional Propagation** - Forward (predicate pushdown) + backward (store capabilities) -3. **Tunable Modes** - Conservative/Balanced/Aggressive with adaptive learning -4. **EXPLAIN Support** - Visual query plans with performance hints -5. **Reversibility** - Time-travel queries and undo operations - -== Quick Start: EXPLAIN - -To understand how your query will execute: - -[source,vcl] ----- -EXPLAIN -SELECT GRAPH, VECTOR -FROM FEDERATION /universities/* -WHERE (h)-[:CITES]->(target) - AND h.embedding SIMILAR TO [0.1, 0.2, 0.3] WITHIN 0.9 -LIMIT 10 ----- - -**Output:** -[source,text] ----- -╔════════════════════════════════════════════════════════════════╗ -║ VCL QUERY EXECUTION PLAN ║ -╚════════════════════════════════════════════════════════════════╝ - -Strategy: Sequential Pipeline (operations run in series) -Optimization Mode: Balanced -Bidirectional Optimization: Enabled -Estimated Total Cost: 180ms - -───────────────────────────────────────────────────────────────── - -Step 1: Query (VECTOR) - Cost: 80ms - Selectivity: 5.0% of data - Optimization: Using HNSW index - Pushed predicates: - - LIMIT 10 - - WITHIN 0.9 - -Step 2: Query (GRAPH) - Cost: 60ms - Selectivity: 2.0% of data - Optimization: Using edge_type_index - Pushed predicates: - - Filter by Step 1 UUIDs - -Step 3: Query (DOCUMENT) - Cost: 40ms - Selectivity: 1.0% of data - -───────────────────────────────────────────────────────────────── - -Cost Breakdown by Modality: - VECTOR: 80ms (44%) - GRAPH: 60ms (33%) - DOCUMENT: 40ms (22%) - -Performance Hints: - ✓ Query plan is optimal ----- - -== Query Planner Modes - -=== Three Optimization Modes - -[cols="1,2,2,2"] -|=== -|Mode |Selectivity Estimate |Cost Estimate |Best For - -|**Conservative** -|2x (assume more results) -|1.5x (add safety buffer) -|Production, compliance, first-time queries - -|**Balanced** -|1x (use historical averages) -|1x (realistic estimates) -|Most workloads, general use - -|**Aggressive** -|0.5x (assume fewer results) -|0.8x (optimistic) -|Development, exploratory queries, known-selective queries -|=== - -=== Configuration API - -**Global Mode:** -[source,elixir] ----- -# Set for all queries -VeriSim.QueryPlannerConfig.set_global_mode(:balanced) ----- - -**Per-Modality Override:** -[source,elixir] ----- -# Vector searches are predictable → aggressive -VeriSim.QueryPlannerConfig.set_modality_mode("VECTOR", :aggressive) - -# Graph traversal is unpredictable → conservative -VeriSim.QueryPlannerConfig.set_modality_mode("GRAPH", :conservative) - -# Semantic ZKP verification is expensive → conservative -VeriSim.QueryPlannerConfig.set_modality_mode("SEMANTIC", :conservative) ----- - -**Adaptive Tuning:** -[source,elixir] ----- -# Enable automatic mode adjustment based on query performance -VeriSim.QueryPlannerConfig.set_adaptive(true) ----- - -When enabled, the system: -1. Tracks actual query costs vs. estimates -2. If consistently wrong (>30% error), adjusts mode -3. Logs mode changes for transparency - -=== Default Configuration - -[source,elixir] ----- -%{ - global_mode: :balanced, - modality_overrides: %{ - "VECTOR" => :aggressive, # Predictable (HNSW) - "GRAPH" => :conservative, # Unpredictable (traversal) - "SEMANTIC" => :conservative # Expensive (ZKP) - }, - statistics_weight: 0.7, # 70% historical, 30% estimates - enable_adaptive: true -} ----- - -== Bidirectional Optimization - -=== Forward Propagation (Top-Down) - -**Push optimizations from query to stores:** - -1. **Predicate Pushdown** - [source,vcl] - ---- - -- Query - SELECT VECTOR, DOCUMENT - WHERE h.embedding SIMILAR TO [0.1, 0.2] - AND FULLTEXT CONTAINS "machine learning" - LIMIT 10 - - -- Optimization: Push predicates to stores - Milvus: "SIMILAR TO [0.1, 0.2] LIMIT 20" (2x buffer) - Tantivy: "CONTAINS 'machine learning'" - ---- - -2. **Projection Pushdown** - [source,vcl] - ---- - -- Only request needed fields - SELECT GRAPH(nodes, edges) -- Don't fetch properties - ---- - -3. **LIMIT Pushdown** - [source,vcl] - ---- - -- Tell stores to limit results early - LIMIT 10 → Each store returns ~20 (buffer for joins) - ---- - -=== Backward Propagation (Bottom-Up) - -**Pull store capabilities into query plan:** - -1. **Index Detection** - - Query Oxigraph: "Do you have edge_type_index?" - - If yes → Reorder to use indexed operation first - -2. **Cache Awareness** - - Query verisim-temporal: "Is version cached?" - - If yes → Execute early (instant) - -3. **GPU Availability** - - Query Burn: "GPU available?" - - If yes → Prioritize tensor operations - -4. **Partition Hints** - - Query Milvus: "Which partitions match this filter?" - - If subset → Only query relevant partitions - -=== Example Optimization Flow - -**Initial Query:** -[source,vcl] ----- -SELECT GRAPH, VECTOR, DOCUMENT -FROM FEDERATION /universities/* -WHERE (h)-[:CITES]->(target) - AND h.embedding SIMILAR TO [0.1, 0.2, 0.3] - AND FULLTEXT CONTAINS "quantum computing" -LIMIT 10 ----- - -**Step 1: Forward Propagation** -- Push LIMIT 10 → stores return 20 (buffer) -- Push VECTOR predicate to Milvus -- Push GRAPH predicate to Oxigraph -- Push DOCUMENT predicate to Tantivy - -**Step 2: Backward Propagation** -- Milvus reports: "HNSW index available, estimated 50 results" -- Oxigraph reports: "No edge_type_index, estimated 10,000 results" -- Tantivy reports: "'quantum computing' is rare, estimated 100 results" - -**Step 3: Reorder Based on Hints** -1. Execute Tantivy first (most selective: 100 results) -2. Filter Milvus by Tantivy UUIDs (50 → 20 results) -3. Filter Oxigraph by Milvus UUIDs (10,000 → 10 results) - -**Result:** 180ms instead of 2000ms (10x faster!) - -== Reversibility: Time-Travel Queries - -=== What You Get For Free - -VeriSimDB's architecture provides **80% reversibility with zero overhead:** - -[cols="1,2,1"] -|=== -|Component |Reversibility Support |Overhead - -|**Temporal Modality** -|✅ Full (Merkle trees) -|0% (already designed for versioning) - -|**Semantic Modality** -|✅ Full (ZKP proofs immutable) -|0% (proofs are versioned) - -|**Metadata Registry** -|✅ Full (KRaft append-only log) -|0% (log truncation optional) -|=== - -=== Time-Travel Query Syntax - -[source,vcl] ----- --- Query historical state -SELECT * -FROM HEXAD abc-123 -WHERE AS OF '2026-01-22T12:00:00Z' - --- Compare current vs historical -SELECT * -FROM HEXAD abc-123 -WHERE BETWEEN '2026-01-22T00:00:00Z' AND '2026-01-22T23:59:59Z' - --- Get specific version -SELECT * -FROM HEXAD abc-123 -WHERE VERSION v3-final ----- - -=== Undo Operations - -[source,elixir] ----- -# Create checkpoint before risky operation -checkpoint = VeriSim.Reversibility.create_checkpoint(octad_id) - -# Try operation -case apply_operation(octad_id) do - {:ok, result} -> - # Success → commit - VeriSim.Reversibility.commit_checkpoint(checkpoint) - {:ok, result} - - {:error, reason} -> - # Failure → undo - VeriSim.Reversibility.revert_to_checkpoint(checkpoint) - {:error, reason} -end ----- - -=== Drift Repair Validation - -**Before (No Reversibility):** -[source,text] ----- -1. Detect drift -2. Apply repair -3. Hope it works ❌ ----- - -**After (With Reversibility):** -[source,text] ----- -1. Detect drift -2. Checkpoint current state -3. Apply repair -4. Validate repair -5. If bad → REVERT ✅ -6. If good → COMMIT ✅ ----- - -[source,elixir] ----- -defmodule VeriSim.DriftMonitor do - def repair_with_validation(octad_id, policy) do - checkpoint = Reversibility.create_checkpoint(octad_id) - - case apply_repair(octad_id, policy) do - {:ok, new_state} -> - if validate_consistency(new_state) do - Reversibility.commit_checkpoint(checkpoint) - {:ok, new_state} - else - Reversibility.revert_to_checkpoint(checkpoint) - {:error, :repair_failed_validation} - end - - {:error, reason} -> - Reversibility.revert_to_checkpoint(checkpoint) - {:error, reason} - end - end -end ----- - -== Performance Best Practices - -=== 1. Use EXPLAIN for Slow Queries - -[source,vcl] ----- --- If query is slow, explain it -EXPLAIN SELECT ... - --- Look for hints like: --- "First step has low selectivity" --- "Operation not using indexes" --- "Query might benefit from parallel execution" ----- - -=== 2. Tune Per-Modality - -[source,elixir] ----- -# Vector searches are fast and predictable -QueryPlannerConfig.set_modality_mode("VECTOR", :aggressive) - -# Graph traversal can explode -QueryPlannerConfig.set_modality_mode("GRAPH", :conservative) ----- - -=== 3. Enable Adaptive Tuning - -[source,elixir] ----- -# Let system learn from your workload -QueryPlannerConfig.set_adaptive(true) ----- - -=== 4. Use Time-Travel for Debugging - -[source,vcl] ----- --- Query was working yesterday, broken today? --- Compare states -SELECT * FROM HEXAD abc-123 WHERE AS OF 'yesterday' ----- - -=== 5. Add Indexes to Stores - -[source,bash] ----- -# Oxigraph: Create edge type index -curl -X POST /oxigraph/indexes -d '{"type": "edge_type"}' - -# Milvus: Use HNSW for high recall -curl -X POST /milvus/indexes -d '{"type": "HNSW", "M": 16}' - -# Tantivy: Index frequently queried fields -curl -X POST /tantivy/schema -d '{"indexed_fields": ["title", "author"]}' ----- - -== Integration Example - -**Complete workflow showing all features:** - -[source,vcl] ----- --- 1. Check query plan -EXPLAIN -SELECT GRAPH, VECTOR -FROM FEDERATION /universities/* -WHERE (h)-[:CITES]->(target) - AND h.embedding SIMILAR TO [0.1, 0.2, 0.3] -LIMIT 10 - --- 2. If slow, tune to aggressive --- (via API: QueryPlannerConfig.set_global_mode(:aggressive)) - --- 3. Create checkpoint before executing --- (via API: checkpoint = Reversibility.create_checkpoint()) - --- 4. Execute query -SELECT GRAPH, VECTOR -FROM FEDERATION /universities/* -WHERE (h)-[:CITES]->(target) - AND h.embedding SIMILAR TO [0.1, 0.2, 0.3] -LIMIT 10 - --- 5. If results wrong, revert --- (via API: Reversibility.revert_to_checkpoint(checkpoint)) - --- 6. Or compare with historical state -SELECT GRAPH, VECTOR -FROM FEDERATION /universities/* -WHERE (h)-[:CITES]->(target) - AND h.embedding SIMILAR TO [0.1, 0.2, 0.3] - AND AS OF '2026-01-22T12:00:00Z' -LIMIT 10 ----- - -== Cost Summary - -|=== -|Feature |Storage Cost |CPU Cost |Benefit - -|**Query Planning** -|0% -|2% -|10x faster queries - -|**Bidirectional Optimization** -|0% -|3% -|Automatic index usage - -|**Tunable Modes** -|0% -|0% -|Flexibility - -|**EXPLAIN** -|0% -|0% -|Debugging - -|**Reversibility (Temporal/Semantic)** -|0% -|0% -|Time-travel, undo - -|**Reversibility (Full)** -|20% -|10% -|Drift validation, A/B testing -|=== - -== References - -- link:vcl-architecture.adoc[VCL Architecture] - Dual-path router design -- link:vcl-examples.adoc[VCL Examples] - 35+ example queries -- link:reversibility-design.adoc[Reversibility Design] - Full technical specification -- link:../lib/verisim/query_planner_config.ex[Query Planner Config] - Source code -- link:../lib/verisim/query_planner_bidirectional.ex[Bidirectional Optimization] - Source code -- link:../src/vcl/VCLExplain.res[EXPLAIN Implementation] - Source code diff --git a/verisimdb/docs/rescript-registry-types.adoc b/verisimdb/docs/rescript-registry-types.adoc deleted file mode 100644 index c1ad4ee3..00000000 --- a/verisimdb/docs/rescript-registry-types.adoc +++ /dev/null @@ -1,71 +0,0 @@ -// SPDX-License-Identifier: CC-BY-SA-4.0 - -= VeriSimDB: ReScript Registry Type Definitions - -These definitions form the backbone of the "Tiny Core". They ensure that every Octad resolved by the WASM proxy is type-safe and consistent with the KRaft controller's state. - -== 1. Core Octad Types - -[source,rescript] ----- -type uuid = string // 128-bit UUID represented as hex -type did = string // Decentralized Identifier for the owner - -type modalityType = - | Graph - | Vector - | Tensor - | Semantic - | Document - | Temporal - -type storeMeta = { - endpoint: string, - supportedModalities: array, - policyHash: string, -} - -type octad = { - id: uuid, - owner: did, - modalities: Belt.Map.String.t, - policyHash: string, - lastModified: float, -} ----- - -== 2. KRaft Metadata Log Types - -To support the KRaft-style quorum, the registry must understand log indices and terms. - -[source,rescript] ----- -type term = int -type index = int - -type logEntry = { - term: term, - index: index, - command: - | RegisterOctad(octad) - | UpdatePolicy(uuid, string) - | RevokeStore(string) -} - -type registryState = { - lastIncludedIndex: index, - lastIncludedTerm: term, - octads: Belt.Map.String.t, -} ----- - -== 3. The Resolution Logic - -This is the primary function invoked by the WASM proxy during a lookup. - -[source,rescript] ----- -let resolveOctad = (registry: registryState, id: uuid): option => { - Belt.Map.String.get(registry.octads, id) -} ----- diff --git a/verisimdb/docs/reversibility-design.adoc b/verisimdb/docs/reversibility-design.adoc deleted file mode 100644 index de25e8b2..00000000 --- a/verisimdb/docs/reversibility-design.adoc +++ /dev/null @@ -1,416 +0,0 @@ -// SPDX-License-Identifier: CC-BY-SA-4.0 - -= Reversibility in VeriSimDB: Design Analysis -:toc: left -:toclevels: 3 -:sectnums: - -== Overview - -This document analyzes the feasibility, costs, and benefits of making VeriSimDB operations **reversible** - i.e., the ability to undo/redo any operation at any point in time. - -== What is Reversibility? - -**Reversibility** means every operation has an **inverse operation** that can restore the previous state: - -[source,text] ----- -State A → [Operation O] → State B → [Inverse O⁻¹] → State A ----- - -**Examples:** -- `INSERT octad X` → `DELETE octad X` -- `UPDATE octad X field=Y` → `UPDATE octad X field=` -- `QUERY returns 100 results` → `RESTORE query state (cached)` - -== Why Reversibility Matters - -=== Use Cases - -1. **Undo Accidental Changes** - - User deletes important octad → Undo - - Drift repair applies wrong fix → Undo - -2. **Time-Travel Queries** - - "What did the database look like 3 hours ago?" - - "Show me the state before the retraction was applied" - -3. **A/B Testing** - - Branch 1: Apply update - - Branch 2: Don't apply update - - Compare results, revert losing branch - -4. **Debugging** - - Replay sequence of operations - - Step backward through query execution - -5. **Audit Compliance** - - Prove database state at specific time - - Reconstruct history for legal requirements - -6. **Drift Repair Validation** - - Apply repair → Test → If bad, revert - -== Good News: VeriSimDB is Already 80% There! - -VeriSimDB's architecture **naturally supports reversibility** because: - -=== 1. verisim-temporal Uses Merkle Trees - -**Already immutable and append-only:** - -[source,text] ----- -Time 0: [State A] → Hash: 0xABC123 -Time 1: [State B] → Hash: 0xDEF456, Parent: 0xABC123 -Time 2: [State C] → Hash: 0x789GHI, Parent: 0xDEF456 - -To revert to Time 1: - 1. Look up hash 0xDEF456 - 2. Reconstruct state from Merkle chain - 3. Apply as current state ----- - -**Cost:** Already implemented! Zero additional cost for temporal history. - -=== 2. KRaft Metadata Log is Append-Only - -The registry doesn't delete entries - it appends new versions: - -[source,text] ----- -Log Entry 1: REGISTER octad abc-123 → store-1 -Log Entry 2: UPDATE octad abc-123 → store-2 (moved) -Log Entry 3: DELETE octad abc-123 (tombstone) - -To revert to Entry 2: - - Replay log up to Entry 2 - - Ignore Entry 3 ----- - -**Cost:** Already implemented! Log truncation is optional. - -=== 3. Modality Stores Support Versioning - -Each modality can be made reversible: - -|=== -| Modality | Reversibility Support | Implementation Cost - -| **Temporal** -| ✅ Native (Merkle trees) -| **FREE** - Already designed for versioning - -| **Graph** (Oxigraph) -| ⚠️ Partial (RDF versioning) -| **LOW** - Use named graphs per version - -| **Vector** (Milvus) -| ❌ Not native -| **MEDIUM** - Wrap with versioned layer - -| **Document** (Tantivy) -| ⚠️ Partial (index snapshots) -| **LOW** - Snapshot-based revert - -| **Tensor** (Burn) -| ⚠️ Depends on backend -| **LOW** - File-based snapshots - -| **Semantic** (verisim-semantic) -| ✅ Native (ZKP proofs are immutable) -| **FREE** - Proofs are already versioned -|=== - -== Implementation Strategy - -=== Level 1: Metadata Reversibility (FREE) - -**What:** Revert registry and metadata changes only. - -**How:** -1. Replay KRaft log to target timestamp -2. Reconstruct registry state -3. Ignore later entries - -**Limitations:** Modality data not reverted (just pointers). - -**Use Case:** Undo registry mistakes (wrong store mapping). - -=== Level 2: Temporal + Semantic Reversibility (CHEAP) - -**What:** Full reversibility for Temporal and Semantic modalities. - -**How:** -1. Use verisim-temporal Merkle chain to retrieve old state -2. ZKP proofs are immutable, just replay - -**Cost:** ~5% storage overhead (Merkle chain hashes). - -**Use Case:** Compliance audits, provenance verification. - -=== Level 3: Full Multi-Modal Reversibility (MODERATE COST) - -**What:** All six modalities support revert. - -**Implementation:** - -==== Option A: Snapshot-Based (Simple) - -[source,rust] ----- -// Take periodic snapshots -pub struct ModalitySnapshot { - timestamp: i64, - modality: Modality, - data_hash: [u8; 32], - delta_from_previous: Option, -} - -impl ModalitySnapshot { - pub fn revert_to(&self, target_time: i64) -> Result { - // Find snapshot before target_time - let snapshot = self.find_snapshot_before(target_time)?; - - // Apply deltas forward from snapshot to target_time - snapshot.apply_deltas_until(target_time) - } -} ----- - -**Storage Cost:** 10-20% overhead (depends on snapshot frequency). - -**Revert Speed:** Fast (O(log n) if deltas are small). - -==== Option B: Copy-on-Write (Elegant) - -[source,rust] ----- -// Use persistent data structures (like Git) -pub struct CowOctad { - id: Uuid, - versions: BTreeMap>, -} - -impl CowOctad { - pub fn update(&mut self, new_state: OctadState) { - // Share unchanged data, clone only modified parts - let new_version = Arc::new(new_state); - self.versions.insert(Utc::now(), new_version); - } - - pub fn revert_to(&self, timestamp: Timestamp) -> Option> { - self.versions.range(..=timestamp).next_back().map(|(_, state)| state.clone()) - } -} ----- - -**Storage Cost:** 15-30% overhead (shared data reduces duplication). - -**Revert Speed:** Instant (O(1) lookup). - -==== Option C: Event Sourcing (Most Powerful) - -[source,rust] ----- -// Store operations, not states -pub enum OctadEvent { - Created { id: Uuid, data: OctadData }, - Updated { id: Uuid, field: String, old: Value, new: Value }, - Deleted { id: Uuid }, -} - -pub struct EventLog { - events: Vec, -} - -impl EventLog { - pub fn reconstruct_at(&self, timestamp: i64) -> OctadState { - self.events - .iter() - .take_while(|e| e.timestamp <= timestamp) - .fold(OctadState::default(), |state, event| { - state.apply(event) - }) - } - - pub fn revert(&self, steps: usize) -> OctadState { - // Replay all events except last N - self.reconstruct_from_events(&self.events[..self.events.len() - steps]) - } -} ----- - -**Storage Cost:** 20-40% overhead (every operation stored). - -**Revert Speed:** Slow for full reconstruction (O(n)), fast for recent reverts. - -**Benefit:** Enables **query replay** (see below). - -== Recommended Approach: Hybrid - -**Phase 1: Metadata + Temporal + Semantic (Immediate)** -- Use existing Merkle trees and KRaft log -- **Cost:** FREE (already implemented) -- **Benefit:** Compliance, audit trails, metadata undo - -**Phase 2: Document + Graph Snapshots (3 months)** -- Snapshot-based reversibility for Tantivy and Oxigraph -- Snapshots every 1 hour (configurable) -- **Cost:** 10% storage overhead -- **Benefit:** Undo accidental deletes, time-travel queries - -**Phase 3: Vector + Tensor Copy-on-Write (6 months)** -- CoW for Milvus embeddings and Burn tensors -- Share unchanged data between versions -- **Cost:** 20% storage overhead -- **Benefit:** A/B testing, drift repair validation - -**Phase 4: Full Event Sourcing (Optional, 12 months)** -- Event log for all operations -- Enable query replay and advanced debugging -- **Cost:** 30% storage overhead -- **Benefit:** Full reversibility, query optimization via replay - -== Cost-Benefit Analysis - -|=== -| Capability | Storage Cost | CPU Cost | Revert Speed | Benefit - -| **Metadata Only** -| 0% -| 0% -| Instant -| Registry undo - -| **+ Temporal/Semantic** -| 5% -| 2% -| Instant -| Compliance, provenance - -| **+ Snapshots (Doc/Graph)** -| 15% -| 5% -| Fast (seconds) -| Time-travel queries - -| **+ CoW (Vector/Tensor)** -| 25% -| 10% -| Instant -| A/B testing, validation - -| **+ Event Sourcing** -| 40% -| 20% -| Varies -| Full replay, debugging -|=== - -== Query Reversibility (Bonus Feature) - -**Event sourcing enables query replay:** - -[source,vcl] ----- --- Original query (3 hours ago) -SELECT GRAPH, VECTOR -FROM FEDERATION /universities/* -WHERE FULLTEXT CONTAINS "machine learning" -LIMIT 100 - --- Replay query at historical timestamp -SELECT GRAPH, VECTOR -FROM FEDERATION /universities/* -WHERE FULLTEXT CONTAINS "machine learning" - AND AS OF '2026-01-22T12:00:00Z' -LIMIT 100 ----- - -**Use Case:** "Show me what this query returned yesterday" (for debugging drift). - -== Drift Repair with Reversibility - -**Current Drift Repair:** -1. Detect drift -2. Apply repair -3. Hope it works - -**With Reversibility:** -1. Detect drift -2. Snapshot current state -3. Apply repair -4. Validate repair -5. **If bad: Revert to snapshot** -6. If good: Commit repair - -**Code Example:** - -[source,elixir] ----- -defmodule VeriSim.DriftMonitor do - def repair_with_validation(octad_id, repair_policy) do - # Create reversible checkpoint - checkpoint = Reversibility.create_checkpoint(octad_id) - - # Apply repair - case apply_repair(octad_id, repair_policy) do - {:ok, new_state} -> - # Validate repair - if validate_consistency(new_state) do - # Good repair, commit - Reversibility.commit_checkpoint(checkpoint) - {:ok, new_state} - else - # Bad repair, revert - Reversibility.revert_to_checkpoint(checkpoint) - {:error, :repair_failed_validation} - end - - {:error, reason} -> - # Repair failed, revert - Reversibility.revert_to_checkpoint(checkpoint) - {:error, reason} - end - end -end ----- - -== Conclusion - -**Is Reversibility Hard?** No - VeriSimDB already has 80% of the infrastructure. - -**Is It Expensive?** Moderate - 15-25% storage overhead for full reversibility. - -**Is It Useful?** **EXTREMELY** - Enables: -- Compliance audits -- Time-travel queries -- Drift repair validation -- A/B testing -- Debugging -- Undo operations - -**Recommendation:** Implement in phases: -1. **Now:** Use existing Merkle trees (FREE) -2. **3 months:** Add snapshots (10% cost, high value) -3. **6 months:** Add CoW (20% cost, very high value) -4. **Optional:** Event sourcing (40% cost, specialized use cases) - -== Next Steps - -1. Enable reversibility for Temporal + Semantic (already done!) -2. Add snapshot infrastructure for Document/Graph -3. Design CoW layer for Vector/Tensor -4. Create VCL syntax for time-travel queries: - ```vcl - SELECT * FROM HEXAD abc-123 AS OF '2026-01-22T12:00:00Z' - REVERT TO '2026-01-22T12:00:00Z' - ``` - -== References - -- link:../WHITEPAPER.md[VeriSimDB White Paper] - Temporal modality design -- link:technical-specification-kraft-metadata-log.adoc[KRaft Log] - Append-only log -- link:vcl-architecture.adoc[VCL Architecture] - Query execution -- https://martin.kleppmann.com/2015/05/27/logs-for-data-infrastructure.html[Designing Data-Intensive Applications] diff --git a/verisimdb/docs/safety-and-fault-tolerance.adoc b/verisimdb/docs/safety-and-fault-tolerance.adoc deleted file mode 100644 index fcd4894b..00000000 --- a/verisimdb/docs/safety-and-fault-tolerance.adoc +++ /dev/null @@ -1,1161 +0,0 @@ -// SPDX-License-Identifier: CC-BY-SA-4.0 - -= Safety and Fault Tolerance -:toc: left -:toclevels: 3 -:sectnums: - -== Overview - -VeriSimDB implements **defense in depth** across all safety dimensions: from kernel-level protections to socio-technical governance. This document catalogs safety guarantees, fault tolerance mechanisms, and failure recovery strategies. - -**Threat Model:** - -- Byzantine failures (malicious nodes) -- Crash failures (nodes die unexpectedly) -- Network partitions (split brain scenarios) -- Data corruption (bit flips, disk failures) -- Resource exhaustion (memory, disk, CPU) -- Supply chain attacks (compromised dependencies) -- Human errors (operational mistakes) - -== Fault Tolerance - -=== BEAM/OTP Supervisor Trees - -**Guarantee:** Process crashes don't cascade to entire system. - -[source,elixir] ----- -defmodule VeriSim.Application do - use Application - - def start(_type, _args) do - children = [ - # Registry for octad servers - {Registry, keys: :unique, name: VeriSim.EntityRegistry}, - - # DynamicSupervisor for octad entities - {DynamicSupervisor, name: VeriSim.EntitySupervisor, strategy: :one_for_one}, - - # Query router (restart if crashes) - {VeriSim.QueryRouter, []}, - - # Cache (restart if crashes, lose cache but not data) - {VeriSim.QueryCache, []}, - - # Drift monitor (restart if crashes, resume monitoring) - {VeriSim.DriftMonitor, []}, - ] - - opts = [strategy: :one_for_one, name: VeriSim.Supervisor] - Supervisor.start_link(children, opts) - end -end ----- - -**Strategy:** - -- `:one_for_one` - If one child crashes, restart only that child -- `:one_for_all` - If one child crashes, restart all children (for tightly coupled processes) -- `:rest_for_one` - If one child crashes, restart it and all children started after it - -**Crash Isolation:** - -[source,text] ----- -┌─────────────────────────────────────────────────────────┐ -│ VeriSim.Supervisor (one_for_one) │ -│ │ -│ ┌──────────────┐ ┌──────────────┐ ┌──────────────┐ │ -│ │ QueryRouter │ │ QueryCache │ │ DriftMonitor │ │ -│ └──────┬───────┘ └──────┬───────┘ └──────┬───────┘ │ -│ │ crashes │ │ │ -│ ✗ │ │ │ -│ │ restarted │ │ │ -│ ✓ │ unaffected │ │ -│ ✓ ✓ │ -└─────────────────────────────────────────────────────────┘ ----- - -=== Self-Healing - -**Guarantee:** System automatically recovers from transient failures. - -==== 1. Automatic Process Restart - -[source,elixir] ----- -defmodule VeriSim.EntityServer do - use GenServer, restart: :transient # Restart only if abnormal exit - - def init(octad_id) do - # Restore state from verisim-temporal - case VeriSim.Temporal.load_octad_state(octad_id) do - {:ok, state} -> - Logger.info("Restored octad #{octad_id} from temporal log") - {:ok, state} - - {:error, :not_found} -> - # First time initialization - {:ok, %{octad_id: octad_id, data: %{}}} - end - end -end ----- - -==== 2. Circuit Breaker Recovery - -[source,elixir] ----- -# Circuit breaker automatically transitions from :open → :half_open → :closed -defmodule VeriSim.CircuitBreaker do - def handle_info(:check_circuit, state) when state.state == :open do - if time_since_last_failure(state) > state.timeout_ms do - Logger.info("Circuit breaker transitioning to half_open for #{state.store_id}") - {:noreply, %{state | state: :half_open}} - else - schedule_check() - {:noreply, state} - end - end -end ----- - -==== 3. Cache Warming After Restart - -[source,elixir] ----- -defmodule VeriSim.QueryCache do - def init(_opts) do - # Schedule cache warming after startup - Process.send_after(self(), :warm_cache, 5_000) - - state = %{...} - {:ok, state} - end - - def handle_info(:warm_cache, state) do - Logger.info("Warming cache with common queries") - - common_queries = get_common_queries() - VeriSim.QueryRouter.Cached.warm_cache(common_queries) - - {:noreply, state} - end -end ----- - -==== 4. Drift Repair - -[source,elixir] ----- -# Automatic drift repair when detected -defmodule VeriSim.DriftMonitor do - def handle_drift_detected(octad_id, drift_details) do - Logger.warn("Drift detected for octad #{octad_id}: #{inspect(drift_details)}") - - case get_repair_policy(octad_id) do - :auto_repair -> - Logger.info("Auto-repairing octad #{octad_id}") - VeriSim.DriftRepair.repair(octad_id, drift_details) - - :manual -> - Logger.warn("Manual repair required for octad #{octad_id}") - notify_admin(octad_id, drift_details) - - :tolerate -> - Logger.info("Tolerating drift for octad #{octad_id}") - :ok - end - end -end ----- - -== Memory Safety - -**Guarantee:** Rust provides memory safety without garbage collection. - -=== Rust Ownership System - -[source,rust] ----- -// SAFE: Ownership prevents use-after-free -fn process_octad(octad: Octad) { - let graph = octad.graph; // Ownership transferred - analyze_graph(graph); - // octad.graph cannot be accessed here (ownership moved) -} - -// SAFE: Borrowing allows temporary access -fn analyze_octad(octad: &Octad) { - println!("Octad ID: {}", octad.id); - // octad remains valid after function returns -} - -// SAFE: Lifetimes ensure references don't outlive data -fn get_octad_id<'a>(octad: &'a Octad) -> &'a str { - &octad.id // Lifetime 'a ensures this reference is valid -} ----- - -=== No Buffer Overflows - -[source,rust] ----- -// SAFE: Bounds checking on array access -let embeddings: Vec = vec![0.1, 0.2, 0.3]; -let value = embeddings.get(10); // Returns None (safe) - -// UNSAFE code is explicitly marked and isolated -unsafe fn read_raw_memory(ptr: *const u8, len: usize) -> &'static [u8] { - std::slice::from_raw_parts(ptr, len) -} ----- - -=== Elixir Immutability - -[source,elixir] ----- -# Immutable data structures prevent accidental mutation -octad = %{id: "abc-123", data: %{title: "Original"}} - -# This creates a NEW map, doesn't mutate original -updated = Map.put(octad, :data, %{title: "Updated"}) - -# Original remains unchanged -IO.inspect(octad.data.title) # => "Original" -IO.inspect(updated.data.title) # => "Updated" ----- - -== Kernel Safety - -**Guarantee:** OS-level isolation between processes. - -=== Process Isolation - -[source,text] ----- -┌──────────────────────────────────────────────────────┐ -│ Linux Kernel │ -│ │ -│ ┌─────────────┐ ┌─────────────┐ ┌─────────────┐ │ -│ │ Process A │ │ Process B │ │ Process C │ │ -│ │ (Elixir) │ │ (Rust API) │ │ (Postgres) │ │ -│ │ │ │ │ │ │ │ -│ │ Memory: │ │ Memory: │ │ Memory: │ │ -│ │ 0x1000-2000 │ │ 0x3000-4000 │ │ 0x5000-6000 │ │ -│ │ │ │ │ │ │ │ -│ │ Cannot │ │ Cannot │ │ Cannot │ │ -│ │ access B/C │ │ access A/C │ │ access A/B │ │ -│ └─────────────┘ └─────────────┘ └─────────────┘ │ -└──────────────────────────────────────────────────────┘ - -Isolation via: -- Virtual memory (MMU) -- System calls (controlled kernel entry) -- File descriptors (capability-based access) ----- - -=== Resource Limits (cgroups) - -[source,bash] ----- -# Limit memory per process -systemctl set-property verisim-api.service MemoryMax=2G - -# Limit CPU shares -systemctl set-property verisim-api.service CPUQuota=200% - -# Limit open file descriptors -ulimit -n 4096 ----- - -=== Seccomp (System Call Filtering) - -[source,rust] ----- -// Restrict which syscalls Rust processes can make -use seccomp::*; - -fn setup_seccomp() -> Result<(), seccomp::Error> { - let mut filter = SeccompFilter::new()?; - - // Allow only essential syscalls - filter.allow(libc::SYS_read)?; - filter.allow(libc::SYS_write)?; - filter.allow(libc::SYS_open)?; - filter.allow(libc::SYS_close)?; - - // Block dangerous syscalls - filter.deny(libc::SYS_fork)?; - filter.deny(libc::SYS_execve)?; - - filter.load()?; - Ok(()) -} ----- - -== Platform Safety - -**Guarantee:** Cross-platform compatibility without platform-specific vulnerabilities. - -=== Supported Platforms - -[cols="1,1,2"] -|=== -|Platform |Status |Notes - -|**Linux x86_64** -|✅ Primary -|Reference platform, full testing - -|**macOS x86_64** -|✅ Supported -|Developer platform - -|**macOS ARM64** -|✅ Supported -|Apple Silicon - -|**Windows x86_64** -|⚠️ Experimental -|WSL2 recommended - -|**FreeBSD** -|🔬 Research -|Rust support good, BEAM needs testing -|=== - -=== Platform-Specific Issues - -[source,elixir] ----- -defmodule VeriSim.Platform do - def get_temp_dir do - case :os.type() do - {:unix, :darwin} -> "/tmp" - {:unix, :linux} -> "/tmp" - {:unix, :freebsd} -> "/tmp" - {:win32, _} -> System.get_env("TEMP") || "C:\\Temp" - end - end - - def get_data_dir do - case :os.type() do - {:unix, _} -> "~/.local/share/verisim" - {:win32, _} -> Path.join(System.get_env("APPDATA"), "verisim") - end - end -end ----- - -=== Endianness Handling - -[source,rust] ----- -use byteorder::{LittleEndian, ReadBytesExt, WriteBytesExt}; - -// Always use explicit endianness for network/disk I/O -fn serialize_vector(embedding: &[f32]) -> Vec { - let mut buf = Vec::new(); - - for &value in embedding { - buf.write_f32::(value).unwrap(); - } - - buf -} ----- - -== Supply Chain Safety - -**Guarantee:** Dependencies are audited, pinned, and reproducible. - -=== Dependency Auditing - -[source,bash] ----- -# Rust: cargo-audit checks for known vulnerabilities -cargo audit - -# Elixir: mix audit checks hex.pm packages -mix audit - -# Continuous monitoring in CI ----- - -=== Dependency Pinning - -**Cargo.lock (Rust):** -[source,toml] ----- -# Exact versions locked -[[package]] -name = "serde" -version = "1.0.195" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "63261df402c67811e9ac6def069e4786148c4563f4b50fd4bf30aa370d626b02" ----- - -**mix.lock (Elixir):** -[source,elixir] ----- -%{ - "phoenix": {:hex, :phoenix, "1.7.10", "02189140a61b2ce85bb633a9b6fd02dff705a5f1596869547aeb2e2b8a61511", [:mix], [...], "hexpm", "..."}, -} ----- - -=== Vendor Dependencies - -[source,bash] ----- -# Rust: Vendor all dependencies locally -cargo vendor - -# Commit vendor/ directory to repo for air-gapped builds -git add vendor/ -git commit -m "Vendor Rust dependencies" ----- - -=== Reproducible Builds - -[source,dockerfile] ----- -# Containerfile with pinned base images -FROM rust:1.75.0-bookworm AS rust-builder - -# Pin Elixir/Erlang versions -FROM hexpm/elixir:1.17.3-erlang-27.1.2-debian-bookworm-20241016-slim - -# Verify checksums -RUN echo "sha256:abc123..." | sha256sum -c - ----- - -=== SBOM Generation - -[source,bash] ----- -# Generate Software Bill of Materials -cargo sbom > sbom-rust.json -mix sbom > sbom-elixir.json - -# Publish SBOM for transparency ----- - -== System Safety - -**Guarantee:** Resilience to hardware failures, disk corruption, power loss. - -=== Disk Corruption Detection - -[source,rust] ----- -// CRC32 checksums for all writes -use crc32fast::Hasher; - -fn write_with_checksum(data: &[u8], file: &mut File) -> io::Result<()> { - let mut hasher = Hasher::new(); - hasher.update(data); - let checksum = hasher.finalize(); - - // Write: [length: u32][data: bytes][checksum: u32] - file.write_u32::(data.len() as u32)?; - file.write_all(data)?; - file.write_u32::(checksum)?; - file.sync_all()?; // fsync - - Ok(()) -} - -fn read_with_checksum(file: &mut File) -> io::Result> { - let len = file.read_u32::()? as usize; - let mut data = vec![0u8; len]; - file.read_exact(&mut data)?; - let stored_checksum = file.read_u32::()?; - - let mut hasher = Hasher::new(); - hasher.update(&data); - let computed_checksum = hasher.finalize(); - - if computed_checksum != stored_checksum { - return Err(io::Error::new(io::ErrorKind::InvalidData, "Checksum mismatch")); - } - - Ok(data) -} ----- - -=== Crash-Safe Writes (fsync) - -[source,rust] ----- -// Write-ahead logging with fsync -fn append_to_wal(entry: &LogEntry) -> io::Result<()> { - let mut file = OpenOptions::new() - .append(true) - .create(true) - .open("wal.log")?; - - let serialized = bincode::serialize(entry)?; - - file.write_all(&serialized)?; - file.sync_all()?; // Guarantee durability before returning - - Ok(()) -} ----- - -=== Redundant Storage (RAID-like) - -[source,elixir] ----- -# Write to multiple stores for redundancy -defmodule VeriSim.ReplicatedStore do - def write_with_replication(key, value, replicas: 3) do - stores = select_replica_stores(3) - - results = stores - |> Task.async_stream(fn store -> - write_to_store(store, key, value) - end, timeout: 5_000) - |> Enum.to_list() - - succeeded = Enum.count(results, fn {:ok, {:ok, _}} -> true; _ -> false end) - - if succeeded >= 2 do # Quorum: 2/3 - {:ok, :replicated} - else - {:error, :replication_failed} - end - end -end ----- - -=== Power Loss Recovery - -[source,rust] ----- -// Atomic rename for crash-safe file updates -use std::fs; - -fn atomic_write(path: &Path, data: &[u8]) -> io::Result<()> { - let temp_path = path.with_extension("tmp"); - - // 1. Write to temp file - fs::write(&temp_path, data)?; - - // 2. Sync temp file to disk - let file = File::open(&temp_path)?; - file.sync_all()?; - - // 3. Atomic rename (POSIX guarantees) - fs::rename(&temp_path, path)?; - - // 4. Sync directory entry - let dir = path.parent().unwrap(); - let dir_file = File::open(dir)?; - dir_file.sync_all()?; - - Ok(()) -} ----- - -== Runtime Safety - -**Guarantee:** Controlled resource usage, graceful degradation under load. - -=== BEAM Scheduler - -[source,text] ----- -┌──────────────────────────────────────────────────────┐ -│ BEAM VM (Erlang Runtime) │ -│ │ -│ ┌──────────┐ ┌──────────┐ ┌──────────┐ │ -│ │ Scheduler│ │ Scheduler│ │ Scheduler│ │ -│ │ 1 │ │ 2 │ │ 3 │ │ -│ └────┬─────┘ └────┬─────┘ └────┬─────┘ │ -│ │ │ │ │ -│ ┌───┴─────────────┴─────────────┴───┐ │ -│ │ Process Queue │ │ -│ │ [P1] [P2] [P3] ... [Pn] │ │ -│ └────────────────────────────────────┘ │ -│ │ -│ Each process gets fair CPU time (reduction count) │ -│ Preemptive multitasking (no process starves) │ -└──────────────────────────────────────────────────────┘ ----- - -**Reduction-based scheduling:** -- Each process gets 2000 reductions -- Function calls, message sends count as reductions -- Scheduler preempts after 2000 reductions - -=== Rust Panic Handling - -[source,rust] ----- -use std::panic; - -fn safe_execute_query(query: &str) -> Result { - // Catch panics to prevent process crash - let result = panic::catch_unwind(|| { - execute_query_internal(query) - }); - - match result { - Ok(Ok(res)) => Ok(res), - Ok(Err(e)) => Err(e), - Err(panic_payload) => { - eprintln!("Query panicked: {:?}", panic_payload); - Err(QueryError::InternalError("Query execution panicked".into())) - } - } -} ----- - -=== Resource Limits - -[source,elixir] ----- -defmodule VeriSim.QueryRouter do - @max_query_time_ms 60_000 # 60 seconds - @max_memory_mb 512 - - def execute_with_limits(query) do - task = Task.async(fn -> - # Monitor memory usage - Process.flag(:max_heap_size, %{ - size: @max_memory_mb * 1024 * 1024 div :erlang.system_info(:wordsize), - kill: true, - error_logger: true - }) - - execute_query(query) - end) - - case Task.yield(task, @max_query_time_ms) || Task.shutdown(task) do - {:ok, result} -> result - nil -> {:error, {:timeout, @max_query_time_ms}} - end - end -end ----- - -=== Back Pressure - -[source,elixir] ----- -defmodule VeriSim.RateLimiter do - use GenServer - - # Token bucket algorithm - def init(opts) do - state = %{ - tokens: opts[:max_tokens], - max_tokens: opts[:max_tokens], - refill_rate: opts[:refill_rate], # tokens per second - } - - schedule_refill() - {:ok, state} - end - - def handle_call(:acquire, _from, state) do - if state.tokens > 0 do - {:reply, :ok, %{state | tokens: state.tokens - 1}} - else - {:reply, {:error, :rate_limit_exceeded}, state} - end - end - - defp handle_info(:refill, state) do - new_tokens = min(state.tokens + state.refill_rate, state.max_tokens) - schedule_refill() - {:noreply, %{state | tokens: new_tokens}} - end -end ----- - -== Execution Safety - -**Guarantee:** Least privilege, sandboxing, capability-based security. - -=== Capability-Based Security - -[source,elixir] ----- -defmodule VeriSim.Capability do - @type capability :: :read | :write | :execute | :admin - - def check_capability(user, octad_id, required_cap) do - user_caps = get_capabilities(user, octad_id) - - if required_cap in user_caps do - :ok - else - {:error, {:insufficient_capability, required_cap}} - end - end -end - -# Usage -case Capability.check_capability(user, octad_id, :write) do - :ok -> update_octad(octad_id, new_data) - {:error, reason} -> {:error, reason} -end ----- - -=== Sandboxed Query Execution - -[source,rust] ----- -// Execute queries in isolated environment -use nix::unistd; -use nix::sched::{clone, CloneFlags}; - -fn execute_sandboxed(query: &str) -> Result { - // Create new PID namespace (isolated process tree) - // Create new NET namespace (no network access unless explicitly granted) - // Create new MOUNT namespace (isolated filesystem view) - - let flags = CloneFlags::CLONE_NEWPID - | CloneFlags::CLONE_NEWNET - | CloneFlags::CLONE_NEWNS; - - // Child process executes query in sandbox - match unsafe { clone(Box::new(|| { - drop_privileges()?; - execute_query_internal(query) - }), flags) } { - Ok(pid) => wait_for_child(pid), - Err(e) => Err(QueryError::SandboxError(e.to_string())) - } -} - -fn drop_privileges() -> Result<(), nix::Error> { - // Drop to non-root user - unistd::setuid(unistd::Uid::from_raw(1000))?; - unistd::setgid(unistd::Gid::from_raw(1000))?; - Ok(()) -} ----- - -=== SELinux/AppArmor Policies - -[source,text] ----- -# AppArmor profile for verisim-api -/usr/local/bin/verisim-api { - # Read-only access to config - /etc/verisim/** r, - - # Read-write access to data directory - /var/lib/verisim/** rw, - - # Network access - network inet stream, - network inet6 stream, - - # Deny everything else - deny /** wx, -} ----- - -== Driver Safety - -**Guarantee:** Safe interaction with external systems (databases, networks, devices). - -=== Database Connection Pooling - -[source,elixir] ----- -# Connection pools prevent resource exhaustion -config :verisim, VeriSim.Repo, - pool_size: 10, # Max 10 concurrent connections - queue_target: 50, # Target queue time (ms) - queue_interval: 1000, # Check queue every second - timeout: 15_000, # Query timeout - connect_timeout: 5_000 # Connection timeout ----- - -=== Network Timeouts - -[source,rust] ----- -use reqwest::Client; -use std::time::Duration; - -fn create_http_client() -> Client { - Client::builder() - .timeout(Duration::from_secs(10)) - .connect_timeout(Duration::from_secs(5)) - .pool_max_idle_per_host(10) - .build() - .unwrap() -} ----- - -=== Device Access Control - -[source,elixir] ----- -defmodule VeriSim.DiskAccess do - @allowed_paths [ - "/var/lib/verisim", - "/tmp/verisim" - ] - - def read_file(path) do - if is_safe_path?(path) do - File.read(path) - else - {:error, {:access_denied, path}} - end - end - - defp is_safe_path?(path) do - canonical = Path.expand(path) - - Enum.any?(@allowed_paths, fn allowed -> - String.starts_with?(canonical, allowed) - end) - end -end ----- - -== Process Safety - -**Guarantee:** Process isolation, crash containment, message passing without shared memory. - -=== BEAM Process Model - -[source,text] ----- -┌──────────────────────────────────────────────────────┐ -│ Process A Process B Process C │ -│ ┌──────────┐ ┌──────────┐ ┌──────────┐ │ -│ │ Mailbox │ │ Mailbox │ │ Mailbox │ │ -│ │ [msg1] │ │ [msg3] │ │ │ │ -│ │ [msg2] │ │ │ │ │ │ -│ └────┬─────┘ └────┬─────┘ └────┬─────┘ │ -│ │ │ │ │ -│ ┌────▼─────┐ ┌────▼─────┐ ┌────▼─────┐ │ -│ │ Heap │ │ Heap │ │ Heap │ │ -│ │ (private)│ │ (private)│ │ (private)│ │ -│ └──────────┘ └──────────┘ └──────────┘ │ -│ │ -│ NO SHARED MEMORY - messages copied between heaps │ -└──────────────────────────────────────────────────────┘ ----- - -**Guarantees:** - -- Process crashes don't affect other processes -- No data races (no shared memory) -- Message passing is safe (deep copy) -- Each process has independent garbage collection - -=== Link and Monitor - -[source,elixir] ----- -# Link: Bidirectional crash propagation (for tightly coupled processes) -defmodule Parent do - def start_child do - pid = spawn_link(Child, :run, []) # If child crashes, parent crashes too - pid - end -end - -# Monitor: Unidirectional notification (supervisor pattern) -defmodule Supervisor do - def start_child do - {pid, ref} = spawn_monitor(Child, :run, []) - - receive do - {:DOWN, ^ref, :process, ^pid, reason} -> - Logger.warn("Child process #{inspect(pid)} died: #{inspect(reason)}") - restart_child() - end - end -end ----- - -=== Crash Reports - -[source,elixir] ----- -defmodule VeriSim.CrashLogger do - require Logger - - def log_crash(pid, reason, stacktrace) do - crash_report = %{ - timestamp: DateTime.utc_now(), - pid: inspect(pid), - reason: inspect(reason), - stacktrace: Exception.format_stacktrace(stacktrace), - memory: :erlang.process_info(pid, :memory), - message_queue_len: :erlang.process_info(pid, :message_queue_len) - } - - Logger.error("Process crash: #{inspect(crash_report)}") - - # Write to verisim-temporal for audit trail - VeriSim.Temporal.append_audit_log("crashes", crash_report) - end -end ----- - -== Program Safety - -**Guarantee:** Type safety, formal verification (ZKP), contract-based programming. - -=== ReScript Type Safety - -[source,rescript] ----- -// Exhaustive pattern matching (compiler enforces) -let handleResult = (result: result) => { - switch result { - | Ok(data) => processData(data) - | Error(ParseError(e)) => logParseError(e) - | Error(TypeError(e)) => logTypeError(e) - | Error(RuntimeError(e)) => logRuntimeError(e) - // Compiler error if any case is missing! - } -} - -// Phantom types for compile-time guarantees -type verified<'a> -type unverified<'a> - -let verify: query => result, error> -let execute: query => result - -// Cannot execute unverified queries (type error!) -let unsafeQuery: query = parseQuery("...") -// execute(unsafeQuery) // TYPE ERROR: expected query - -let safeQuery: query = verify(unsafeQuery)->Belt.Result.getExn -execute(safeQuery) // OK ----- - -=== Contract Programming (Elixir Dialyzer) - -[source,elixir] ----- -defmodule VeriSim.Octad do - @type octad_id :: String.t() - @type modality :: :graph | :vector | :tensor | :semantic | :document | :temporal - - @spec get_octad(octad_id()) :: {:ok, map()} | {:error, atom()} - def get_octad(octad_id) when is_binary(octad_id) do - # Implementation - end - - # Dialyzer catches type mismatches at compile time - # get_octad(123) # WARNING: function expects binary, got integer -end ----- - -=== ZKP Formal Verification - -[source,text] ----- -Contract: CitationContract - -Precondition: - ∀ citation ∈ octad.citations, - ∃ cited_octad ∈ registry, - octad.timestamp > cited_octad.timestamp - -Postcondition: - proof ← generate_zkp(octad, CitationContract) - verify_zkp(proof, public_inputs) = true - -Invariant: - Citations cannot reference future octads ----- - -== Socio-Technical Safety - -**Guarantee:** Human factors considered, operational safety, governance. - -=== Human Error Prevention - -[cols="1,2,2"] -|=== -|Hazard |Mitigation |Enforcement - -|**Accidental Data Deletion** -|Soft delete + retention policy -|Reversibility with time-travel queries - -|**Incorrect Access Control** -|Principle of least privilege -|Capability-based security, audit logs - -|**Deployment Errors** -|Blue-green deployments -|Rollback scripts, canary releases - -|**Configuration Mistakes** -|Config validation + defaults -|Nickel type system, schema validation - -|**Operational Burnout** -|On-call rotation limits -|Governance policy (max 1 week on-call) -|=== - -=== Operational Safety - -[source,text] ----- -┌────────────────────────────────────────────────────┐ -│ OPERATIONAL SAFETY LAYERS │ -│ │ -│ 1. Monitoring & Alerting │ -│ ├─ Prometheus metrics │ -│ ├─ Grafana dashboards │ -│ └─ PagerDuty integration │ -│ │ -│ 2. Incident Response │ -│ ├─ Runbooks (step-by-step guides) │ -│ ├─ Post-mortem process │ -│ └─ Blameless culture │ -│ │ -│ 3. Change Management │ -│ ├─ Change advisory board │ -│ ├─ Approval workflow │ -│ └─ Rollback procedures │ -│ │ -│ 4. Documentation │ -│ ├─ Architecture decision records (ADRs) │ -│ ├─ Deployment guides │ -│ └─ Troubleshooting guides │ -└────────────────────────────────────────────────────┘ ----- - -=== Governance - -[source,adoc] ----- -= VeriSimDB Governance - -== Decision-Making Process - -1. **Proposal** - Anyone can propose changes (RFC process) -2. **Discussion** - Community feedback (minimum 1 week) -3. **Vote** - Steering committee votes (simple majority) -4. **Implementation** - Approved proposals merged -5. **Review** - Post-deployment review (1 month) - -== Steering Committee - -- 5 members (elected annually) -- Responsibilities: - * Approve major architectural changes - * Resolve conflicts - * Oversee security advisories - -== Security Advisory Process - -1. **Report** - security@verisimdb.org (private) -2. **Triage** - Assess severity (CVSS score) -3. **Patch** - Develop fix (private repo) -4. **Coordinated Disclosure** - 90 days or when patched -5. **Post-Mortem** - Public analysis after disclosure ----- - -=== Ethics & Privacy - -[cols="1,2"] -|=== -|Principle |Implementation - -|**Data Minimization** -|Only collect necessary data, ZKP for verification without disclosure - -|**Purpose Limitation** -|Data used only for stated purpose, no secondary use without consent - -|**Transparency** -|Audit logs public (privacy-preserving), open-source codebase - -|**User Control** -|Users can export, delete, or restrict their data - -|**Security by Design** -|End-to-end encryption, defense in depth -|=== - -== Safety Guarantees Summary - -[cols="1,1,2"] -|=== -|Safety Dimension |Status |Key Mechanism - -|**Fault Tolerance** -|✅ High -|BEAM supervisors, automatic restart - -|**Self-Healing** -|✅ High -|Circuit breakers, drift repair, cache warming - -|**Memory Safety** -|✅ Guaranteed -|Rust ownership, Elixir immutability - -|**Kernel Safety** -|✅ Strong -|Process isolation, cgroups, seccomp - -|**Platform Safety** -|✅ Good -|Cross-platform testing, explicit endianness - -|**Supply Chain** -|✅ Strong -|Dependency auditing, pinning, SBOM - -|**System Safety** -|✅ High -|CRC checksums, fsync, redundant storage - -|**Runtime Safety** -|✅ High -|Resource limits, back pressure, preemptive scheduling - -|**Execution Safety** -|✅ Strong -|Capabilities, sandboxing, least privilege - -|**Driver Safety** -|✅ Good -|Connection pools, timeouts, access control - -|**Process Safety** -|✅ Guaranteed -|No shared memory, message passing, crash isolation - -|**Program Safety** -|✅ Strong -|Type safety, contracts, ZKP formal verification - -|**Socio-Technical** -|✅ Evolving -|Governance, ethics, operational safety culture -|=== - -== References - -- link:error-handling-strategy.adoc[Error Handling Strategy] -- link:reversibility-design.adoc[Reversibility Design] -- link:vcl-architecture.adoc[VCL Architecture] -- link:../lib/verisim/supervisor.ex[Supervisor Tree Implementation] -- https://www.erlang.org/doc/design_principles/des_princ.html[OTP Design Principles] -- https://doc.rust-lang.org/book/ch04-01-what-is-ownership.html[Rust Ownership System] diff --git a/verisimdb/docs/safety-theory-applied.adoc b/verisimdb/docs/safety-theory-applied.adoc deleted file mode 100644 index 29d8b76b..00000000 --- a/verisimdb/docs/safety-theory-applied.adoc +++ /dev/null @@ -1,1336 +0,0 @@ -// SPDX-License-Identifier: CC-BY-SA-4.0 - -= Safety Theory Applied to VeriSimDB -:toc: left -:toclevels: 3 -:sectnums: - -== Overview - -VeriSimDB's design explicitly considers **safety science research** from James Reason, Jop Groeneweg, Jens Rasmussen, and others. This document analyzes how their theories apply to database systems and how VeriSimDB addresses identified hazards. - -**Key Insight:** Safety is not just about technical defenses, but about recognizing **incompatible goals** and understanding that **defenses themselves can become threats**. - -== James Reason: Swiss Cheese Model - -=== The Model - -[source,text] ----- -Hazard → [Layer 1] → [Layer 2] → [Layer 3] → [Layer 4] → Accident - 🧀 🧀 🧀 🧀 - Holes align across all layers = failure ----- - -**Layers in VeriSimDB:** - -1. **Design** - Architecture prevents classes of errors (type safety, ownership) -2. **Implementation** - Code reviews, testing, formal verification -3. **Runtime** - Supervisors, circuit breakers, drift detection -4. **Operational** - Monitoring, alerts, runbooks, incident response - -**Reason's Distinction:** - -- **Active Failures** - Direct causes (bug in code, operator mistake) -- **Latent Conditions** - Systemic issues (poor design, inadequate testing, production pressure) - -=== Incompatible Goals in VeriSimDB - -**Goal Conflict 1: Performance vs Correctness** - -[cols="1,1,2"] -|=== -|Goal A |Goal B |Conflict - -|**Query Speed** -Fast response times -|**Dependent-Type Verification** -ZKP proof generation (expensive) -|Users skip verification path when performance matters - -|**Cache Hit Rate** -Serve from cache (< 1ms) -|**Data Freshness** -Always query authoritative store -|Stale data served under load - -|**Slipstream Path** -No verification overhead -|**Byzantine Fault Tolerance** -Verify all data sources -|Malicious data accepted when verification skipped -|=== - -**VeriSimDB Mitigation:** - -[source,elixir] ----- -defmodule VeriSim.IncompatibleGoals do - @moduledoc """ - Explicit recognition of goal conflicts with policy enforcement. - """ - - @doc """ - Enforce minimum verification for sensitive operations. - """ - def enforce_verification_policy(query, context) do - # High-risk queries MUST use dependent-type path - if is_high_risk?(query) and not context.use_dependent_types do - {:error, {:verification_required, - "This query involves sensitive data and requires dependent-type verification"}} - else - :ok - end - end - - defp is_high_risk?(query) do - # Medical data, financial data, PII - has_sensitive_octads?(query) or - has_cross_org_access?(query) or - has_write_operations?(query) - end -end ----- - -**Goal Conflict 2: Availability vs Consistency** - -[cols="1,1,2"] -|=== -|Goal A |Goal B |Conflict - -|**Always Available** -Return results even if some stores fail -|**Strong Consistency** -All stores must agree -|Partial/divergent results under partition - -|**Federation Scale** -100+ organizations -|**Consensus Overhead** -Raft quorum requires majority -|Cannot achieve consensus at scale - -|**Drift Tolerance** -Allow temporary inconsistency -|**Referential Integrity** -Citations must reference valid octads -|Broken citations during drift windows -|=== - -**VeriSimDB Mitigation:** - -[source,elixir] ----- -defmodule VeriSim.CAP do - @moduledoc """ - Explicit CAP theorem trade-offs based on operation criticality. - """ - - def execute_with_policy(query, policy) do - case policy do - :cp_mode -> - # Consistency + Partition tolerance (sacrifice availability) - execute_with_strong_consistency(query) - - :ap_mode -> - # Availability + Partition tolerance (sacrifice consistency) - execute_with_eventual_consistency(query) - - :auto -> - # Choose based on query semantics - if requires_strong_consistency?(query) do - execute_with_strong_consistency(query) - else - execute_with_eventual_consistency(query) - end - end - end - - defp requires_strong_consistency?(query) do - # Writes, financial transactions, medical records - query.mutation? or - query.octads |> Enum.any?(&critical_data?/1) - end -end ----- - -**Goal Conflict 3: Developer Velocity vs Security** - -[cols="1,1,2"] -|=== -|Goal A |Goal B |Conflict - -|**Ship Features Fast** -Short development cycles -|**Thorough Security Review** -Slow, careful analysis -|Security corners cut under pressure - -|**Experiment Freely** -Try new modalities/stores -|**Audit Trail** -Every change logged -|Experimental code bypasses logging - -|**Flexible Federation** -Easy to join/leave -|**Byzantine Fault Prevention** -Strict vetting of participants -|Malicious nodes join federation -|=== - -**VeriSimDB Mitigation:** - -[source,adoc] ----- -= Development Policy (Governance) - -## Feature Development - -1. **Security Review Gate** - All features touching sensitive data require security team sign-off - - NOT optional, even under deadline pressure - - Estimated 2-5 days review time (budget this upfront) - -2. **Experimental Flag** - New features behind feature flags - - `config :verisim, experimental_features: [:new_modality]` - - Disabled by default in production - - Audit log captures all experimental feature usage - -3. **Federation Vetting** - New federation members require: - - Technical review (2 weeks) - - Legal agreement (data governance) - - Trial period (3 months, revocable) - - Background check for operators - -**Why This Works:** -- Makes incompatible goals visible (not hidden) -- Slows down risky operations (by design) -- Provides escape valve (experimental flags) without compromising production ----- - -=== Defenses as Threats (Reason's Paradox) - -**Threat 1: Circuit Breakers Mask Problems** - -[source,text] ----- -Problem: Store is slow/unreliable -Defense: Circuit breaker opens, fail fast -Threat: Operators don't notice degraded service - (circuit breaker "working as intended") -Result: Underlying issue never fixed ----- - -**Mitigation:** - -[source,elixir] ----- -defmodule VeriSim.CircuitBreakerMonitor do - @moduledoc """ - Alert when circuit breakers open too frequently. - Defense working = problem hidden. - """ - - def check_circuit_health do - breakers = CircuitBreaker.get_all_stats() - - # Alert if ANY breaker opened in last hour - recently_opened = Enum.filter(breakers, fn {_store, stats} -> - stats.last_open_time && - DateTime.diff(DateTime.utc_now(), stats.last_open_time, :minute) < 60 - end) - - if length(recently_opened) > 0 do - alert(:circuit_breaker_opened, """ - Circuit breakers opened recently: - #{inspect(recently_opened)} - - This indicates underlying store problems. - Circuit breaker is MASKING the issue. - Investigate root cause immediately. - """) - end - end -end ----- - -**Threat 2: Caching Serves Incorrect Data** - -[source,text] ----- -Problem: Cache serves stale data -Defense: Drift detection + invalidation -Threat: Drift detection has false negatives - (doesn't catch all inconsistencies) -Result: Stale cache incorrectly trusted ----- - -**Mitigation:** - -[source,elixir] ----- -defmodule VeriSim.CacheAudit do - @moduledoc """ - Randomly audit cached results against authoritative store. - """ - - def audit_random_cache_entries(sample_rate \\ 0.01) do - # 1% of cache hits trigger audit - if :rand.uniform() < sample_rate do - cached_entries = QueryCache.get_random_sample(100) - - Enum.each(cached_entries, fn entry -> - # Re-execute query without cache - fresh_result = QueryRouter.execute(entry.query, force_fresh: true) - - unless results_match?(entry.cached_result, fresh_result) do - alert(:cache_mismatch, """ - Cache audit failed! - Query: #{entry.query} - Cached: #{inspect(entry.cached_result)} - Fresh: #{inspect(fresh_result)} - - This indicates drift detection missed an inconsistency. - """) - - # Invalidate this cache entry - QueryCache.invalidate(entry.key) - end - end) - end - end -end ----- - -**Threat 3: Automatic Retry Amplifies Cascading Failures** - -[source,text] ----- -Problem: Store is overloaded -Defense: Retry with backoff -Threat: All clients retry simultaneously - (thundering herd) -Result: Store goes from degraded → down ----- - -**Mitigation:** - -[source,elixir] ----- -defmodule VeriSim.ErrorRecovery do - @doc """ - Add jitter to prevent thundering herd. - """ - defp calculate_backoff(attempt, base_delay_ms, max_delay_ms) do - delay = base_delay_ms * :math.pow(2, attempt) |> round() - - # Jitter: ±25% randomization - jitter = delay * (0.75 + :rand.uniform() * 0.5) |> round() - - min(jitter, max_delay_ms) - end - - @doc """ - Global retry budget: limit concurrent retries. - """ - def retry_with_budget(func, opts \\ []) do - if can_acquire_retry_token?() do - try do - retry_with_backoff(func, opts) - after - release_retry_token() - end - else - # Retry budget exhausted, fail fast - {:error, {:retry_budget_exhausted, "System under load, retry later"}} - end - end - - defp can_acquire_retry_token? do - # Token bucket: max 100 concurrent retries system-wide - RetryBucket.try_acquire() - end -end ----- - -**Threat 4: ZKP Overhead Incentivizes Skipping Verification** - -[source,text] ----- -Problem: ZKP proof generation is expensive (500ms) -Defense: Dependent-type path with verification -Threat: Users switch to slipstream path to avoid overhead -Result: Verification bypassed entirely ----- - -**Mitigation:** - -[source,elixir] ----- -defmodule VeriSim.VerificationPolicy do - @doc """ - Mandatory verification for sensitive operations. - Cannot be overridden even by administrator. - """ - def enforce_mandatory_verification(query, context) do - if requires_mandatory_verification?(query) do - if context.use_dependent_types do - :ok - else - {:error, {:mandatory_verification, - """ - This query involves sensitive data and REQUIRES dependent-type verification. - The slipstream path is NOT available for this operation. - - Rationale: Medical data, financial transactions, and PII require - cryptographic proofs for audit and compliance. - """}} - end - else - :ok - end - end - - # Audit: Track verification opt-outs - def log_verification_decision(query, context) do - if not context.use_dependent_types and context.available_dependent_types do - Temporal.append_audit_log("verification_optout", %{ - query_id: query.id, - user_id: context.user_id, - reason: "User chose slipstream path despite dependent-type availability", - octad_ids: extract_octad_ids(query) - }) - end - end -end ----- - -== Jop Groeneweg: Tripod Theory - -=== 11 Basic Risk Factors (BRFs) - -Groeneweg identified 11 organizational factors that create preconditions for accidents: - -[cols="1,2,2"] -|=== -|BRF |How It Manifests in Databases |VeriSimDB Mitigation - -|**1. Hardware** -|Disk failures, network outages, memory corruption -|Redundant storage, CRC checksums, ECC memory - -|**2. Design** -|SQL injection, buffer overflows, race conditions -|Type safety (ReScript), ownership (Rust), immutability (Elixir) - -|**3. Procedures** -|Unclear deployment steps, missing rollback procedures -|Runbooks, ADRs, documented incident response - -|**4. Error-Enforcing Conditions** -|Time pressure, understaffing, 24/7 on-call -|On-call rotation limits (max 1 week), blameless post-mortems - -|**5. Housekeeping** -|Stale dependencies, obsolete code, technical debt -|Dependency auditing, automated cleanup, refactoring sprints - -|**6. Incompatible Goals** -|Performance vs correctness, availability vs consistency -|**Explicit policy enforcement** (see above) - -|**7. Communication** -|Undocumented assumptions, siloed teams -|ADRs, architecture docs, cross-team reviews - -|**8. Organization** -|Unclear responsibility, no owner for security -|Governance model, steering committee, RACI matrix - -|**9. Training** -|Operators don't understand ZKP, drift, federation -|Onboarding docs, runbooks, chaos engineering exercises - -|**10. Defenses** -|Over-reliance on automation, complacency -|**Random audits** (cache correctness, circuit breaker health) - -|**11. Incompatible Materials/Equipment** -|Mixing trusted and untrusted data sources -|Org-scoped caches, cross-org access control -|=== - -=== Tripod Delta: Preconditions → Defenses → Accident - -[source,text] ----- -┌──────────────────────────────────────────────────────┐ -│ BASIC RISK FACTORS (Latent Conditions) │ -│ • Time pressure (ship features fast) │ -│ • Incompatible goals (performance vs correctness) │ -│ • Inadequate training (operators don't know ZKP) │ -└───────────────────┬──────────────────────────────────┘ - │ - ▼ -┌──────────────────────────────────────────────────────┐ -│ PRECONDITIONS (Active in Workplace) │ -│ • Slipstream path chosen for speed │ -│ • Cache hit rate prioritized over freshness │ -│ • Circuit breaker masks slow store │ -└───────────────────┬──────────────────────────────────┘ - │ - ▼ -┌──────────────────────────────────────────────────────┐ -│ DEFENSES (Should Prevent Accident) │ -│ • Drift detection (passive) │ -│ • ZKP verification (skipped!) │ -│ • Access control (bypassed via cache) │ -└───────────────────┬──────────────────────────────────┘ - │ DEFENSES FAIL - ▼ -┌──────────────────────────────────────────────────────┐ -│ ACCIDENT │ -│ • Stale data served from cache │ -│ • User receives incorrect medical diagnosis │ -│ • Lawsuit, regulatory fine, loss of trust │ -└──────────────────────────────────────────────────────┘ ----- - -**VeriSimDB Strategy: Break the Chain** - -1. **Eliminate BRFs** - Address incompatible goals explicitly -2. **Detect Preconditions** - Audit when slipstream chosen for sensitive data -3. **Active Defenses** - Random cache audits, not just passive drift detection -4. **Redundant Defenses** - Multiple layers (Swiss cheese) - -== Jens Rasmussen: Dynamic Safety Model - -=== The Boundary Model - -[source,text] ----- - ┌─────────────────────────┐ - │ SAFE OPERATING SPACE │ - │ │ - Economically │ ⬤ ← System │ Unacceptable - Infeasible │ │ Workload - │ │ - └─────────────────────────┘ - │ - Safety Boundary - (Accident happens) - -System migrates toward boundaries under pressure: -• Management pressure → Economic boundary (under-provision) -• Workload pressure → Unacceptable workload (burnout) -• Least effort → Safety boundary (skip verification) ----- - -**Rasmussen's Insight:** Systems naturally drift toward boundaries due to: - -1. **Economic Pressure** - Cost reduction, faster time-to-market -2. **Workload Pressure** - Do more with less, 24/7 operations -3. **Gradient of Least Effort** - Shortcuts save time, rarely punished - -=== Boundaries in VeriSimDB - -**Economic Boundary:** - -[source,text] ----- -Pressure: "Why provision 3 stores? Use 1 to save cost!" -Drift: Single store, no redundancy -Accident: Store fails → complete outage ----- - -**Mitigation:** - -[source,elixir] ----- -defmodule VeriSim.CapacityPlanning do - @doc """ - Enforce minimum redundancy for production. - """ - def validate_deployment(config) do - if config.environment == :production do - # Require minimum 3 stores per modality - Enum.each(config.modalities, fn {modality, stores} -> - if length(stores) < 3 do - raise """ - PRODUCTION SAFETY VIOLATION: - Modality #{modality} has only #{length(stores)} stores. - Minimum 3 required for fault tolerance. - - This is NOT negotiable. Under-provisioning to save cost - will result in complete outage when ANY store fails. - """ - end - end) - end - end -end ----- - -**Workload Boundary:** - -[source,text] ----- -Pressure: "We need 24/7 on-call coverage but only 2 engineers" -Drift: Engineers burned out, skip careful verification -Accident: Operator mistake during incident, data corruption ----- - -**Mitigation:** - -[source,adoc] ----- -= On-Call Policy (Governance) - -## Rotation Limits - -1. **Maximum on-call duration:** 1 week -2. **Minimum team size for 24/7 coverage:** 4 engineers -3. **Mandatory break after on-call:** 3 days off-rotation - -## Escalation - -If staffing inadequate for 24/7: -1. Reduce coverage to business hours (8am-6pm) -2. Hire contractors for interim coverage -3. Use external monitoring service (PagerDuty, Opsgenie) - -**DO NOT:** -- Extend on-call shifts beyond 1 week -- Expect engineers to "cover for each other" indefinitely -- Sacrifice safety for 24/7 availability ----- - -**Safety Boundary (Least Effort Gradient):** - -[source,text] ----- -Pressure: "ZKP verification is slow, just use slipstream" -Drift: Operators routinely skip verification -Accident: Malicious data accepted, Byzantine fault undetected ----- - -**Mitigation:** - -[source,elixir] ----- -defmodule VeriSim.SafetyBoundary do - @doc """ - Make the safe path the EASIEST path (not the hard one). - """ - - # DEFAULT: Dependent-type path (safe) - def execute_query(query, opts \\ []) do - # use_dependent_types: true by default - use_dependent_types = Keyword.get(opts, :use_dependent_types, true) - - # Cache results to amortize ZKP cost - if use_dependent_types do - VeriSim.QueryRouter.Cached.execute_with_cache(query, opts) - else - # Slipstream requires explicit opt-in + reason - reason = Keyword.fetch!(opts, :slipstream_reason) - log_slipstream_optout(query, reason) - - VeriSim.QueryRouter.execute_slipstream(query, opts) - end - end - - defp log_slipstream_optout(query, reason) do - Logger.warn("Slipstream path chosen: #{reason}") - - Temporal.append_audit_log("slipstream_usage", %{ - query: query, - reason: reason, - timestamp: DateTime.utc_now() - }) - end -end ----- - -**Key Insight:** Make the **safe path the default**, make the **unsafe path require justification**. - -=== Rasmussen's SRK Framework - -**Skill-based** - Automatic, unconscious (muscle memory) -**Rule-based** - Following procedures (if X then Y) -**Knowledge-based** - Problem-solving, novel situations - -**Application to VeriSimDB Operators:** - -[cols="1,2,2"] -|=== -|Level |Operator Behavior |Error Mode - -|**Skill-based** -|Routine deployments, restart services -|Slip: Type wrong command, forget step - -|**Rule-based** -|Follow runbook for known incident -|Mistake: Misapply rule, wrong context - -|**Knowledge-based** -|Novel failure, no runbook -|Knowledge gap: Wrong mental model -|=== - -**VeriSimDB Mitigation:** - -[source,adoc] ----- -= Runbook Design Principles - -## Skill-based Operations (Routine) - -1. **Automate** - Use scripts, not manual commands -2. **Checklist** - Pre-flight checklist before deploy -3. **Muscle memory** - Practice in staging first - -## Rule-based Operations (Known Incidents) - -1. **Decision trees** - Clear if/then flowcharts -2. **Verification steps** - "Check X before proceeding to Y" -3. **Rollback procedure** - Always provide undo path - -## Knowledge-based Operations (Novel Failures) - -1. **Escalation path** - When to call expert -2. **Diagnostics** - How to gather information -3. **Safe mode** - Fall back to minimal functionality -4. **Post-incident review** - Convert novel → rule-based - -**Example Runbook: Store Unreachable** - -SKILL-BASED: -``` -1. Check monitoring dashboard -2. Run: `systemctl status verisim-store` -3. Check logs: `journalctl -u verisim-store -n 100` -``` - -RULE-BASED: -``` -IF store_status == "stopped" THEN - 1. Restart: `systemctl start verisim-store` - 2. Verify: `curl http://store:8080/health` - 3. IF health_check == "ok" THEN done - 4. ELSE escalate - -IF store_status == "degraded" THEN - 1. Check disk space: `df -h /var/lib/verisim` - 2. IF disk_full THEN clear old logs - 3. ELSE escalate -``` - -KNOWLEDGE-BASED: -``` -IF none of the above work: - 1. Escalate to on-call engineer - 2. Gather: logs, metrics, recent changes - 3. Consider: network partition, hardware failure, Byzantine fault - 4. Safe fallback: Isolate store, use replicas -``` ----- - -== Diane Vaughan: Normalization of Deviance - -**The Challenger Disaster Lesson:** - -NASA gradually accepted O-ring erosion as normal because: -1. Flights succeeded despite erosion (luck masked danger) -2. Small violations accumulated over time -3. Engineers normalized the deviance socially -4. New engineers didn't know original safety standards - -**Definition:** **Normalization of Deviance** is the gradual process where unacceptable practices become acceptable through repetition without immediate consequences. - -=== Deviance Normalization in Databases - -[source,text] ----- -NORMAL → DEVIANT → NORMALIZED DEVIANCE → CATASTROPHE - -Week 1: "Cache should be fresh" (NORMAL) -Week 2: "Cache hit rate is low, let's increase TTL to 5 min" -Week 3: "5 min worked fine, let's try 1 hour" -Week 4: "1 hour is great! No issues so far" (DEVIANCE NORMALIZED) -Week 8: User receives 3-hour-old medical diagnosis → HARM - -Why it happens: -• No immediate negative feedback (luck) -• Gradual boundary expansion (boiling frog) -• Social reinforcement ("everyone does it") -• Production pressure (performance metrics reward it) ----- - -**Example 1: Cache TTL Creep** - -[source,text] ----- -Timeline of Normalization: - -Month 1: Cache TTL = 5 minutes (by design) - "Data must be reasonably fresh" - -Month 2: Performance team increases TTL to 15 minutes - "No user complaints, looks good!" - -Month 3: TTL increased to 1 hour - "Still no issues, cache hit rate excellent!" - -Month 6: TTL increased to 6 hours - "System is so fast now!" - [DEVIANCE NORMALIZED] - -Month 7: Medical database serves 8-hour-old diagnosis - Patient receives outdated treatment protocol - HARM - -What went wrong: -• Each increase seemed small/safe -• No immediate consequences (lag between cause and effect) -• Organizational memory: New engineers don't know original 5-min rationale -• Metrics rewarded high cache hit rate, not data freshness ----- - -**VeriSimDB Protection:** - -[source,elixir] ----- -defmodule VeriSim.DevianceMonitor do - @moduledoc """ - Detect normalization of deviance via configuration drift. - """ - - @original_config %{ - cache_ttl_seconds: 300, # 5 minutes (design intent) - verification_required: true, - drift_tolerance: :strict - } - - def check_for_deviance do - current_config = get_current_config() - - # Flag ANY deviation from original design - deviations = Map.keys(@original_config) - |> Enum.filter(fn key -> - @original_config[key] != current_config[key] - end) - |> Enum.map(fn key -> - %{ - setting: key, - original: @original_config[key], - current: current_config[key], - changed_at: get_config_change_date(key), - changed_by: get_config_change_author(key) - } - end) - - if length(deviations) > 0 do - alert(:deviance_detected, """ - NORMALIZATION OF DEVIANCE ALERT - - Current configuration deviates from original design: - #{inspect(deviations, pretty: true)} - - WHY THIS MATTERS: - Each change seemed small and safe at the time. - But accumulated deviations indicate boundary erosion. - - REQUIRED ACTION: - 1. Review each deviation's rationale (ADR required) - 2. Measure actual risk vs perceived risk - 3. Restore original config OR document why deviation is safe - - Remember: Challenger disaster happened because small O-ring - erosions were normalized over 24 flights. - """) - end - end -end ----- - -**Example 2: Verification Bypass Creep** - -[source,text] ----- -Timeline of Normalization: - -Month 1: 100% of writes use dependent-type path (by design) - "All mutations must be verified via ZKP" - -Month 2: Dev team adds "skip verification" flag for testing - "Just for local dev, will remove before production" - -Month 3: Flag makes it to production (behind feature flag) - "Only used in emergencies" - -Month 4: Emergency occurs, flag used - "It worked! No issues!" - -Month 5: Flag used routinely to speed up batch imports - "Verification is too slow for bulk operations" - [DEVIANCE NORMALIZED] - -Month 6: Malicious data inserted via unverified bulk import - Byzantine fault undetected - HARM - -What went wrong: -• "Emergency" exception became routine -• Speed prioritized over security (incompatible goals) -• No immediate consequences from skipping verification -• New engineers don't know why verification was mandatory ----- - -**VeriSimDB Protection:** - -[source,elixir] ----- -defmodule VeriSim.VerificationAudit do - @moduledoc """ - Track verification bypass usage and alert on normalization. - """ - - def audit_verification_patterns do - # Get last 30 days of verification decisions - usage = Temporal.query_audit_log(""" - SELECT - DATE(timestamp) as date, - COUNT(*) as total_queries, - SUM(CASE WHEN use_dependent_types = false THEN 1 ELSE 0 END) as slipstream_count - FROM query_log - WHERE timestamp > NOW() - INTERVAL '30 days' - GROUP BY DATE(timestamp) - ORDER BY date - """) - - # Detect upward trend in slipstream usage - trend = calculate_trend(usage, :slipstream_count) - - if trend.direction == :increasing and trend.slope > 0.05 do - alert(:deviance_normalizing, """ - NORMALIZATION OF DEVIANCE: Verification Bypass Increasing - - Slipstream path (unverified queries) usage is trending UP: - #{format_trend(trend)} - - This indicates verification is being bypassed more frequently. - - TYPICAL PATTERN (Challenger-style): - Week 1: "Just this once for performance testing" - Week 2: "It worked fine, let's use it again" - Week 4: "Everyone's doing it now" - Week 8: Malicious data inserted, Byzantine fault - - REQUIRED ACTION: - 1. Interview teams: Why is verification being skipped? - 2. Address root cause (is ZKP too slow? improve performance) - 3. DO NOT normalize the bypass as acceptable - """) - end - end -end ----- - -=== Practical Drift (Scott Snook) - -**Definition:** **Practical Drift** is the slow, steady uncoupling of practice from written procedure due to: - -1. **Local Rationality** - Actions make sense in local context -2. **Decentralization** - No single person sees the full picture -3. **Production Pressure** - Get the job done trumps follow the rules - -**Snook's Friendly Fire Analysis:** - -- Procedure: Pilots must confirm target ID before firing -- Reality: Pilots skip confirmation when time-pressured -- Local Rationality: "We're in a combat zone, it's obviously hostile" -- Drift: Confirmation step routinely skipped -- Result: Friendly fire (two US helicopters shot down) - -**Application to VeriSimDB:** - -[source,text] ----- -PROCEDURE vs PRACTICE DRIFT - -Written Procedure: - "All cross-org queries must use federation quorum (3/5 stores)" - -Actual Practice (Month 1): - 100% compliance - -Actual Practice (Month 3): - Developer: "Quorum is slow, let's use 1/5 for read-only queries" - [LOCAL RATIONALITY: Makes sense for this specific query] - -Actual Practice (Month 6): - Team: "Everyone uses 1/5 now, it's the new normal" - [DECENTRALIZATION: No one sees they ALL drifted] - -Actual Practice (Month 12): - New engineer: "The docs say 3/5 but everyone uses 1/5" - [ORGANIZATIONAL MEMORY LOSS: Procedure obsolete] - -Result: - Byzantine fault undetected (malicious store serves bad data) - No other stores consulted to detect inconsistency - HARM ----- - -**VeriSimDB Protection:** - -[source,elixir] ----- -defmodule VeriSim.ProcedureDriftDetector do - @moduledoc """ - Detect drift between documented procedure and actual practice. - """ - - @documented_procedures %{ - federation_quorum: 3, # Documented: 3/5 stores required - cache_ttl_medical: 60, # Documented: 1 minute for medical data - verification_rate: 1.0 # Documented: 100% of writes verified - } - - def detect_drift do - actual_practice = measure_actual_practice() - - drifts = Map.keys(@documented_procedures) - |> Enum.filter(fn key -> - documented = @documented_procedures[key] - actual = actual_practice[key] - - # Significant drift: >10% deviation - abs(actual - documented) / documented > 0.1 - end) - |> Enum.map(fn key -> - %{ - procedure: key, - documented: @documented_procedures[key], - actual: actual_practice[key], - drift_percent: calculate_drift_percent(key, actual_practice[key]) - } - end) - - if length(drifts) > 0 do - alert(:procedure_drift, """ - PRACTICAL DRIFT DETECTED - - Written procedures vs actual practice: - #{inspect(drifts, pretty: true)} - - WHAT THIS MEANS: - Your documentation describes one thing, but teams do another. - - This is "practical drift" (Snook): - • Each deviation made sense locally ("faster this way") - • Decentralized: No one saw EVERYONE was drifting - • New members learn from practice, not documentation - - REQUIRED ACTION: - 1. Determine correct approach (update docs OR enforce procedure) - 2. If docs are wrong: Update them (ADR explaining why) - 3. If practice is wrong: Re-train teams, add enforcement - 4. DO NOT let drift continue silently - """) - end - end - - defp measure_actual_practice do - # Sample actual system behavior - %{ - federation_quorum: measure_avg_quorum_size(), - cache_ttl_medical: measure_avg_cache_ttl_for_medical_data(), - verification_rate: measure_verification_usage_rate() - } - end -end ----- - -=== Organizational Memory Loss - -**Problem:** Original design rationale forgotten over time. - -[source,text] ----- -Year 1: "Cache TTL is 5 minutes because medical data must be fresh" - [Original engineer documents this] - -Year 2: Original engineer leaves - -Year 3: New engineer: "Why is TTL only 5 minutes? That's slow!" - [No memory of rationale] - Changes TTL to 1 hour - -Year 4: Another engineer: "Why was it ever 5 minutes? 1 hour seems normal" - [Deviance fully normalized] ----- - -**VeriSimDB Protection: Architecture Decision Records (ADRs)** - -[source,adoc] ----- -= ADR-023: Cache TTL for Medical Data - -Date: 2025-01-15 -Status: Accepted -Supersedes: ADR-015 - -## Context - -Medical diagnosis data must be fresh to prevent harm: -- Treatments change rapidly (new studies, drug recalls) -- Stale data can lead to incorrect medical decisions -- Regulatory requirement (HIPAA): data freshness guarantees - -## Decision - -Cache TTL for medical data (octads tagged "medical:*") is **5 minutes**. - -This is NOT negotiable. Do NOT increase this TTL without: -1. New ADR with medical review board approval -2. Risk assessment by safety team -3. Legal review (regulatory compliance) - -## Consequences - -Positive: -- Patients receive up-to-date medical information -- Regulatory compliance maintained -- Reduced liability - -Negative: -- Lower cache hit rate for medical queries -- Higher load on medical data stores - -## Why 5 Minutes Specifically? - -Medical review board determined: -- Most treatment updates occur hourly -- 5-minute lag is acceptable risk -- 1-hour lag is UNACCEPTABLE (could miss critical updates) - -## Rationale for Future Maintainers - -If you're reading this and thinking "5 minutes is too short": - -STOP. Consider: -1. Challenger disaster: Small O-ring erosions normalized -2. Each TTL increase seems harmless until it isn't -3. Medical data requires freshness, not just performance - -If you have a legitimate reason to change this: -- Write a new ADR (supersede this one) -- Get medical review board approval -- Document your rationale for the NEXT maintainer ----- - -== Other Relevant Researchers - -=== Nancy Leveson: STAMP/STPA - -**Systems-Theoretic Accident Model (STAMP):** - -- Accidents caused by inadequate **control** of interactions -- Not just component failures, but **emergent properties** -- Focus on constraints, not just events - -**Application to VeriSimDB:** - -[source,text] ----- -Control Structure: - -┌─────────────────────────────────────────────┐ -│ Governance (Steering Committee) │ -│ Control: Policy, standards, oversight │ -└─────────────────┬───────────────────────────┘ - │ Policy feedback - ▼ -┌─────────────────────────────────────────────┐ -│ Development Team │ -│ Control: Code review, testing, deploy │ -└─────────────────┬───────────────────────────┘ - │ Metrics, alerts - ▼ -┌─────────────────────────────────────────────┐ -│ VeriSimDB (System) │ -│ Control: Supervisors, circuit breakers │ -└─────────────────┬───────────────────────────┘ - │ State, logs - ▼ -┌─────────────────────────────────────────────┐ -│ Modality Stores (Controlled Process) │ -│ Control: Rate limits, timeouts │ -└─────────────────────────────────────────────┘ - -Hazard: Data corruption -Constraint: Writes must be verified via ZKP -Control Action: Enforce dependent-type path for writes ----- - -=== Erik Hollnagel: Safety-II (Resilience Engineering) - -**Safety-I:** Things go wrong (accident prevention) -**Safety-II:** Things go right (performance variability) - -**Application:** - -Instead of asking "Why did the circuit breaker fail?", ask: - -- "How do operators successfully manage load spikes?" (learn from success) -- "What workarounds do users employ when stores are slow?" (performance variability) -- "How does the system gracefully degrade?" (resilience) - -**VeriSimDB Resilience:** - -[source,elixir] ----- -defmodule VeriSim.ResilienceMonitor do - @moduledoc """ - Study how the system succeeds under stress. - """ - - def analyze_successful_degradation do - # Find incidents where system degraded but didn't fail - incidents = Temporal.query_audit_log(""" - SELECT * FROM audit_log - WHERE event_type = 'degraded_service' - AND duration > 5 minutes - AND NOT EXISTS ( - SELECT * FROM audit_log AS outages - WHERE outages.event_type = 'complete_outage' - AND outages.timestamp BETWEEN incident.start AND incident.end - ) - """) - - # Analyze what kept the system running - Enum.each(incidents, fn incident -> - resilience_factors = %{ - cache_hit_rate: incident.cache_stats.hit_rate, - partial_results_used: incident.partial_results_count > 0, - circuit_breakers_active: incident.circuit_breakers_open > 0, - load_shedding_active: incident.rate_limit_exceeded > 0 - } - - Logger.info(""" - Successful degradation (not failure): - #{inspect(resilience_factors)} - - System remained available despite: - - #{incident.failed_stores} stores unavailable - - #{incident.latency_p99}ms p99 latency - - Learn from this success! - """) - end) - end -end ----- - -== Summary: Safety Theory in Practice - -[cols="1,2,2"] -|=== -|Theory |Key Insight |VeriSimDB Implementation - -|**Reason (Swiss Cheese)** -|Multiple defense layers, latent conditions -|Design, implementation, runtime, operational layers - -|**Reason (Incompatible Goals)** -|Safety vs performance trade-offs -|Explicit policy enforcement, mandatory verification - -|**Reason (Defenses as Threats)** -|Safety mechanisms can mask problems -|Random audits, circuit breaker alerts - -|**Vaughan (Normalization of Deviance)** -|Small violations accumulate, become normal -|Configuration drift monitoring, ADRs preserve rationale - -|**Snook (Practical Drift)** -|Practice uncouples from procedure -|Measure actual behavior vs documented procedure - -|**Groeneweg (Tripod)** -|11 organizational factors create preconditions -|Address BRFs (training, procedures, housekeeping) - -|**Rasmussen (Boundaries)** -|Systems drift toward unsafe boundaries -|Make safe path easiest, enforce minimums - -|**Rasmussen (SRK)** -|Different error modes at different cognitive levels -|Runbooks for skill/rule/knowledge levels - -|**Leveson (STAMP)** -|Accidents from inadequate control -|Control structure with feedback loops - -|**Hollnagel (Safety-II)** -|Learn from success, not just failure -|Analyze successful degradation, resilience factors -|=== - -=== The Normalization Cascade (Most Dangerous) - -[source,text] ----- -1. Incompatible Goals (Reason) - ↓ - Pressure to prioritize performance over safety - -2. Practical Drift (Snook) - ↓ - Teams skip safety steps "just this once" - (locally rational, no immediate consequences) - -3. Normalization of Deviance (Vaughan) - ↓ - Skipping safety becomes routine practice - (organizational amnesia, social reinforcement) - -4. Boundary Migration (Rasmussen) - ↓ - System operates closer to safety boundary - (gradient of least effort) - -5. ACCIDENT - ↓ - Eventually the lack of immediate consequences - runs out (luck exhausted) ----- - -**VeriSimDB Protection Against The Cascade:** - -1. **Make goals compatible** - Fast AND safe (via caching) -2. **Monitor drift** - Alert when practice ≠ procedure -3. **Preserve memory** - ADRs explain "why" for future maintainers -4. **Enforce boundaries** - Hard limits, not soft guidelines -5. **Learn from success** - What kept system safe under pressure? - -**Core Principle:** - -Safety is not just **avoiding bad outcomes** (Safety-I), but **understanding and amplifying successful adaptations** (Safety-II) while **recognizing that defenses themselves can become threats** (Reason) and **incompatible goals create systemic pressure toward unsafe boundaries** (Rasmussen). - -== References - -- Reason, J. (1997). Managing the Risks of Organizational Accidents -- Reason, J. (1990). Human Error -- Groeneweg, J. (1992). Controlling the Controllable: The Management of Safety -- Rasmussen, J. (1997). Risk Management in a Dynamic Society: A Modelling Problem -- Rasmussen, J. (1983). Skills, Rules, and Knowledge (SRK Framework) -- Leveson, N. (2011). Engineering a Safer World (STAMP/STPA) -- Hollnagel, E. (2014). Safety-I and Safety-II (Resilience Engineering) -- link:safety-and-fault-tolerance.adoc[Safety & Fault Tolerance] - Technical implementation -- link:error-handling-strategy.adoc[Error Handling Strategy] - Recovery mechanisms diff --git a/verisimdb/docs/security-lessons.lgt b/verisimdb/docs/security-lessons.lgt deleted file mode 100644 index c3ebc8da..00000000 --- a/verisimdb/docs/security-lessons.lgt +++ /dev/null @@ -1,227 +0,0 @@ -% SPDX-License-Identifier: MPL-2.0 -% Security Lessons - Logtalk Rules for VeriSimDB Security Fixes -% Intended for incorporation into reposystem and gitbot-fleet repos - -:- object(security_lessons). - - :- info([ - version is 1.0, - author is 'VeriSimDB Security Team', - date is 2026-01-22, - comment is 'Lessons learned from Dependabot and OpenSSF Scorecard remediation' - ]). - - %% Dependency Management Rules - - :- public(should_update_dependency/3). - :- mode(should_update_dependency(+atom, +atom, +atom), zero_or_one). - :- info(should_update_dependency/3, [ - comment is 'Determine if a dependency should be updated based on vulnerability', - argnames is ['Package', 'CurrentVersion', 'TargetVersion'] - ]). - - % Rule: Update protobuf from 2.x to 3.x to fix recursion CVE - should_update_dependency(protobuf, CurrentVersion, '3.7.2') :- - atom_concat('2.', _, CurrentVersion), - !. % Recursion crash vulnerability - - % Rule: Update prometheus when it depends on vulnerable protobuf - should_update_dependency(prometheus, '0.13.4', '0.14.0') :- - transitive_depends_on(prometheus, protobuf, '2.28.0'), - !. - - % Rule: Update tantivy to fix lru Stacked Borrows issue - should_update_dependency(tantivy, CurrentVersion, '0.25.0') :- - version_less_than(CurrentVersion, '0.25.0'), - transitive_depends_on(tantivy, lru, LruVersion), - vulnerable_lru(LruVersion), - !. - - :- public(vulnerable_lru/1). - vulnerable_lru('0.12.5'). % Stacked Borrows violation - - :- public(transitive_depends_on/3). - :- mode(transitive_depends_on(+atom, +atom, +atom), zero_or_one). - transitive_depends_on(Parent, Child, ChildVersion) :- - % Placeholder: would query Cargo.lock/package-lock.json - % In gitbot-fleet, implement via cargo tree or npm ls - fail. - - %% Workflow Security Rules - - :- public(workflow_needs_fixing/2). - :- mode(workflow_needs_fixing(+atom, -list), one). - :- info(workflow_needs_fixing/2, [ - comment is 'Identify security issues in GitHub Actions workflows', - argnames is ['WorkflowPath', 'Issues'] - ]). - - workflow_needs_fixing(Path, Issues) :- - findall(Issue, workflow_issue(Path, Issue), Issues). - - :- private(workflow_issue/2). - - % Missing SPDX header - workflow_issue(Path, missing_spdx_header) :- - read_first_line(Path, Line), - \+ atom_concat('# SPDX-License-Identifier:', _, Line). - - % Unpinned GitHub Actions - workflow_issue(Path, unpinned_action(Action, Line)) :- - read_workflow_line(Path, LineNum, Line), - atom_concat('uses: ', Rest, Line), - atom_concat(Action, '@', Rest), - \+ (atom_concat(_, '@', Rest), atom_length(Rest, Len), Len > 40), % SHA is 40 chars - Line. - - % Missing permissions declaration - workflow_issue(Path, missing_permissions) :- - \+ workflow_has_permissions(Path). - - :- public(fix_unpinned_action/3). - :- mode(fix_unpinned_action(+atom, +atom, -atom), one). - :- info(fix_unpinned_action/3, [ - comment is 'Convert version tag to SHA-pinned reference', - argnames is ['Action', 'Version', 'SHA'] - ]). - - % Known SHA mappings (January 2026) - fix_unpinned_action('actions/checkout', 'v6.0.1', 'b4ffde65f46336ab88eb53be808477a3936bae11'). - fix_unpinned_action('actions/checkout', 'v4', 'b4ffde65f46336ab88eb53be808477a3936bae11'). - fix_unpinned_action('actions/configure-pages', 'v5', '983d7736d9b0ae728b81ab479565c72886d7745b'). - fix_unpinned_action('actions/upload-pages-artifact', 'v4', '56afc609e74202658d3ffba0e8f6dda462b719fa'). - fix_unpinned_action('actions/deploy-pages', 'v4', 'd6db90164ac5ed86f2b6aed7e0febac5b3c0c03e'). - fix_unpinned_action('actions/jekyll-build-pages', 'v1', '44a6e6beabd48582f863aeeb6cb2151cc1716697'). - fix_unpinned_action('ruby/setup-ruby', 'v1.207.0', '708024e6c902387ab41de36e1669e43b5ee7085e'). - fix_unpinned_action('dtolnay/rust-toolchain', 'stable', '6d9817901c499d6b02debbb57edb38d33daa680b'). - - %% Branch Protection Rules - - :- public(branch_protection_config/2). - :- mode(branch_protection_config(+atom, -term), one). - :- info(branch_protection_config/2, [ - comment is 'Required branch protection settings for OpenSSF Scorecard', - argnames is ['Branch', 'Config'] - ]). - - branch_protection_config(main, config( - required_approving_review_count(1), - enforce_admins(false), - required_status_checks(null), - restrictions(null), - allow_force_pushes(false), - allow_deletions(false), - allow_fork_syncing(true) - )). - - %% OpenSSF Scorecard Compliance Rules - - :- public(scorecard_check_status/3). - :- mode(scorecard_check_status(+atom, +atom, -atom), one). - :- info(scorecard_check_status/3, [ - comment is 'Determine if OpenSSF Scorecard check passes', - argnames is ['CheckID', 'RepoPath', 'Status'] - ]). - - % Branch-Protection check - scorecard_check_status('Branch-Protection', Repo, pass) :- - branch_protected(Repo, main), - !. - scorecard_check_status('Branch-Protection', _, fail). - - % Code-Review check - scorecard_check_status('Code-Review', Repo, pass) :- - branch_protected(Repo, main), - required_reviews(Repo, main, Count), - Count >= 1, - !. - scorecard_check_status('Code-Review', _, fail). - - % Pinned-Dependencies check - scorecard_check_status('Pinned-Dependencies', Repo, pass) :- - forall(workflow_file(Repo, Path), all_actions_pinned(Path)), - !. - scorecard_check_status('Pinned-Dependencies', _, fail). - - % Vulnerabilities check - scorecard_check_status('Vulnerabilities', Repo, pass) :- - \+ has_dependabot_alerts(Repo), - !. - scorecard_check_status('Vulnerabilities', _, fail). - - %% Automation Helpers - - :- public(generate_fix_pr/3). - :- mode(generate_fix_pr(+atom, +list, -atom), one). - :- info(generate_fix_pr/3, [ - comment is 'Generate PR branch with automated security fixes', - argnames is ['Repo', 'Issues', 'BranchName'] - ]). - - generate_fix_pr(Repo, Issues, BranchName) :- - atom_concat('security-fixes-', Timestamp, BranchName), - get_time(Timestamp), - create_branch(Repo, BranchName), - forall(member(Issue, Issues), apply_fix(Repo, BranchName, Issue)), - commit_fixes(Repo, BranchName, 'chore(security): automated Scorecard fixes'), - create_pull_request(Repo, BranchName, main, 'Security: OpenSSF Scorecard Fixes'). - - %% Integration Points for gitbot-fleet - - :- public(scan_repo_for_issues/2). - :- mode(scan_repo_for_issues(+atom, -list), one). - :- info(scan_repo_for_issues/2, [ - comment is 'Scan repository and return all security issues', - argnames is ['RepoPath', 'Issues'] - ]). - - scan_repo_for_issues(Repo, AllIssues) :- - findall(dependency_issue(Pkg, Old, New), - (cargo_dependency(Repo, Pkg, Old), - should_update_dependency(Pkg, Old, New)), - DepIssues), - findall(workflow_issue(Path, Issue), - (workflow_file(Repo, Path), - workflow_issue(Path, Issue)), - WorkflowIssues), - findall(scorecard_fail(CheckID), - scorecard_check_status(CheckID, Repo, fail), - ScorecardIssues), - append([DepIssues, WorkflowIssues, ScorecardIssues], AllIssues). - - %% Priority Rules - - :- public(issue_priority/2). - :- mode(issue_priority(+term, -atom), one). - - issue_priority(dependency_issue(_, Version, _), critical) :- - atom_concat('2.', _, Version), % Major version jump (e.g., protobuf 2.x → 3.x) - !. - - issue_priority(workflow_issue(_, unpinned_action(_, _)), high). - issue_priority(workflow_issue(_, missing_permissions), high). - issue_priority(workflow_issue(_, missing_spdx_header), medium). - - issue_priority(scorecard_fail('Branch-Protection'), high). - issue_priority(scorecard_fail('Code-Review'), high). - issue_priority(scorecard_fail('Vulnerabilities'), critical). - issue_priority(scorecard_fail('Pinned-Dependencies'), medium). - issue_priority(scorecard_fail(_), low). - - %% Lesson Summary - - :- public(security_lesson/2). - :- mode(security_lesson(+atom, -atom), one). - - security_lesson(transitive_dependencies, - 'Update parent crates when transitive dependencies are vulnerable'). - security_lesson(major_version_jumps, - 'Major version updates (2.x→3.x) often fix critical CVEs'). - security_lesson(sha_pinning, - 'Always SHA-pin GitHub Actions to prevent tag manipulation'). - security_lesson(branch_protection, - 'Require at least 1 review for Code-Review scorecard compliance'). - security_lesson(automation, - 'Batch fix similar issues across repos to save time'). - -:- end_object. diff --git a/verisimdb/docs/snapshotting-and-truncation-logic.adoc b/verisimdb/docs/snapshotting-and-truncation-logic.adoc deleted file mode 100644 index b1fc793e..00000000 --- a/verisimdb/docs/snapshotting-and-truncation-logic.adoc +++ /dev/null @@ -1,77 +0,0 @@ -// SPDX-License-Identifier: CC-BY-SA-4.0 - -= VeriSimDB: Snapshotting & Truncation Logic - -To maintain the "Tiny Core" mandate, VeriSimDB uses a KRaft-inspired snapshotting mechanism. This prevents the append-only metadata log from growing indefinitely. - -== 1. The Trigger Mechanism - -The Elixir Orchestrator monitors the log size. Once the log reaches a configurable threshold (e.g., 10,000 entries or 10MB), a snapshot is initiated. - -[source,elixir] ----- -defmodule VeriSim.Registry.Snapshotter do - @doc """ - Initiates a snapshot of the current FSM state. - """ - def take_snapshot(state_machine_pid) do - # 1. Capture the current state from the Raft FSM - state = Raft.get_state(state_machine_pid) - - # 2. Serialize to CBOR (compact binary format) - snapshot_data = CBOR.encode(state) - - # 3. Calculate the metadata for the snapshot - meta = %{ - last_included_index: Raft.get_last_index(state_machine_pid), - last_included_term: Raft.get_last_term(state_machine_pid) - } - - # 4. Save to disk/store and truncate the log - Raft.commit_snapshot(state_machine_pid, meta, snapshot_data) - end -end ----- - -== 2. Log Truncation - -Once the snapshot is successfully written and acknowledged by the Quorum: - -* All log entries preceding `last_included_index` are deleted. -* Memory is freed. -* Any new node joining the cluster downloads the snapshot first, then catches up on the few remaining log entries. - -== 3. Structural ASCII Overview - -=== System Architecture - ----- -[ Client (WASM SDK) ] - | - | (1) Signed Request (sactify-php) - v -[ WASM Proxy Node ] <---- (2) Trust Window Lookup ----> [ Elixir Orchestrator ] - | | - | (3) Multi-Modality Fetch | (4) Drift Check - | | - +------------+-------------+-------------+ | - | | | | v -[ Store A ] [ Store B ] [ Store C ] [ Store D ] <--- [ Registry Quorum ] - (Graph) (Vector) (Tensor) (Semantic) (KRaft Consistency) ----- - -=== Sequence Flow: Registration - ----- -Client WASM Proxy Controller Quorum Store (Rust) - | | | | - |---(Sign/Auth)-->| | | - | |---(Propose)------->| | - | | |--[ Consensusing ]--| - | | | | - | |<--(Commit/Ack)-----| | - | | | | - | |---(Distribute)------------------------->| - | | | | - |<---(201 Created)| | | ----- diff --git a/verisimdb/docs/technical-specification-kraft-metadata-log.adoc b/verisimdb/docs/technical-specification-kraft-metadata-log.adoc deleted file mode 100644 index 412fc5c0..00000000 --- a/verisimdb/docs/technical-specification-kraft-metadata-log.adoc +++ /dev/null @@ -1,33 +0,0 @@ -// SPDX-License-Identifier: CC-BY-SA-4.0 - -= Technical Specification: The KRaft Metadata Log - -In VeriSimDB, the Registry is not a passive database but an active Replicated State Machine (RSM). Following the KRaft pattern, all metadata changes are treated as a sequence of events. - -== 1. The Append-Only Ledger - -Every change to the global namespace—whether a new Octad registration or a policy update—is appended to a log. - -*Consensus*:: A change is only considered "Committed" once it has been replicated to a majority of the Registry Quorum. - -*Order*:: The log index ensures a total ordering of events, preventing "Split-Brain" scenarios in the federation. - -== 2. Log Truncation & Snapshots - -To maintain a "Tiny Core" footprint, the log is not infinite. - -When the log exceeds the `THRESHOLD_LIMIT`, the current `registryState` (as defined in the ReScript types) is serialized into a Snapshot. - -The log is then truncated, keeping only the snapshot and any entries created after it. - -== 3. Pull-Based Follower Sync - -Stores or passive registry nodes that have been offline do not wait for a push. They perform a Catch-up Pull: - -. Request the current Snapshot from the Leader. -. Apply the snapshot to their local state. -. Replay the remaining log entries from the `lastIncludedIndex`. - -== 4. The "Trust Window" Integration - -The KRaft log also stores the short-lived symmetric keys used for Trust Windows. By replicating these keys across the quorum, we ensure that a client can fail-over from one registry node to another without needing a "Heavy Handshake" re-authentication. diff --git a/verisimdb/docs/vcl-architecture.adoc b/verisimdb/docs/vcl-architecture.adoc deleted file mode 100644 index 2d561adc..00000000 --- a/verisimdb/docs/vcl-architecture.adoc +++ /dev/null @@ -1,707 +0,0 @@ -// SPDX-License-Identifier: CC-BY-SA-4.0 - -= VCL Architecture: Dual-Path Query Router -:toc: left -:toclevels: 3 -:sectnums: - -== Overview - -This document describes the **corrected** VCL (VeriSim Consonance Language) architecture, aligned with link:../design-decisions.adoc[Architecture Design Decisions] and the link:../WHITEPAPER.md[VeriSimDB White Paper]. - -**Key Corrections from Initial Design:** - -* ❌ **NOT** using CockroachDB for query metadata -* ✅ **Using** ReScript registry + Raft consensus (Elixir) -* ❌ **NOT** using Fluree for audit trails -* ✅ **Using** `verisim-temporal` (Merkle trees) for immutable logs -* ✅ **Integrated** drift detection/repair via Elixir `DriftMonitor` - -== 1. Core Architecture: Dual-Path Router - -VCL provides two execution paths for queries: - -[source,text] ----- - ┌─────────────────┐ - │ VCL Query │ - │ (raw string) │ - └────────┬────────┘ - │ - ┌─────────────▼────────────────┐ - │ Has PROOF clause? │ - └──────┬──────────────┬────────┘ - │ │ - YES ◄────┘ └────► NO - │ │ - ┌──────▼──────────┐ ┌─────────▼──────────┐ - │ Dependent-Type │ │ Slipstream │ - │ Path │ │ Path │ - │ (Verified) │ │ (Unverified) │ - └──────┬──────────┘ └─────────┬──────────┘ - │ │ - │ │ - ┌──────▼──────────────────────┬────────▼─────────┐ - │ │ - │ VeriSim Orchestrator (Elixir) │ - │ │ - └──────┬─────────────────────────────┬────────────┘ - │ │ - ┌──────▼────────┐ ┌───────▼───────┐ - │ Oxigraph │ │ Milvus │ - │ (Graph) │ │ (Vector) │ - └───────────────┘ └───────────────┘ - │ │ - ┌──────▼────────┐ ┌───────▼───────┐ - │ Tantivy │ │ Burn │ - │ (Document) │ │ (Tensor) │ - └───────────────┘ └───────────────┘ ----- - -== 2. Component Architecture - -=== 2.1. ReScript Core - -**Location:** `verisimdb/src/vcl/` - -**Components:** - -[source,text] ----- -src/vcl/ -├── VCLParser.res # Parses raw query to AST -├── VCLTypeChecker.res # Type checks with proven-library -├── VCLRouter.res # Main routing logic (dual-path) -├── VCLResponse.res # Response formatting -└── bindings/ - ├── ProvenLibrary.res # FFI to proven-library (Rust) - └── ElixirRouter.res # FFI to Elixir orchestrator ----- - -**Registry Integration:** - -[source,rescript] -// VCLRouter.res -module VCL = { - type queryMetadata = { - queryId: string, - timestamp: float, - usedDependentTypes: bool, - proof: option, - octadIds: array, - } - - let handleQuery = (rawQuery: string, useDependentTypes: bool) => { - // Log to verisim-temporal (NOT Fluree) - let queryId = TemporalLog.appendQuery(rawQuery) - - if (useDependentTypes) { - // Dependent-Type Path - rawQuery - -> VCLParser.parseToTypedAST - -> TypeChecker.validateWithProvenLibrary - -> DriftMonitor.checkConsistency // Check for drift - -> FederatedExecutor.executeWithZKP - -> TemporalLog.appendResult(queryId) // Merkle tree audit trail - -> Response.withProof - } else { - // Slipstream Path - rawQuery - -> VCLParser.parseToUntypedAST - -> FederatedExecutor.executeUnverified - -> TemporalLog.appendResult(queryId) // Still log (no proof) - -> Response.withoutProof - } - } - - // Store metadata in ReScript registry (NOT CockroachDB) - let storeQueryMetadata = (metadata: queryMetadata) => { - Registry.updateOctadMetadata(metadata.octadIds, { - lastQueried: metadata.timestamp, - queryCount: Registry.incrementCounter, - }) - } -} ----- - -=== 2.2. Elixir Orchestration - -**Location:** `verisimdb/lib/verisim/` - -**Components:** - -[source,text] ----- -lib/verisim/ -├── query_router.ex # Main query routing GenServer -├── drift_monitor.ex # Drift detection/repair -├── federated_executor.ex # Distributes queries to stores -├── raft_metadata.ex # Raft consensus for metadata -└── temporal_log.ex # Audit trail (Merkle trees) ----- - -**Query Router (Corrected):** - -[source,elixir] -# lib/verisim/query_router.ex -defmodule VeriSim.QueryRouter do - use GenServer - - @doc """ - Routes queries to appropriate modality stores. - Integrates with DriftMonitor for consistency checks. - """ - - def handle_query(%{typed: true} = query) do - # Dependent-Type Path - with {:ok, proof} <- ProvenLibrary.generate_zkp(query.ast), - {:ok, _} <- DriftMonitor.check_consistency(query.octad_ids), - {:ok, results} <- distribute_to_stores(query), - :ok <- TemporalLog.append(query.id, results, proof) do - {:ok, %{data: results, proof: proof, verified: true}} - else - {:error, :drift_detected} = err -> - case query.drift_policy do - :strict -> err - :repair -> - DriftMonitor.repair_and_retry(query) - :tolerate -> - # Continue despite drift - {:ok, results} = distribute_to_stores(query) - {:ok, %{data: results, proof: nil, verified: false, warning: "Drift detected"}} - end - error -> error - end - end - - def handle_query(%{typed: false} = query) do - # Slipstream Path - {:ok, results} = distribute_to_stores(query) - :ok = TemporalLog.append(query.id, results, nil) # Log without proof - {:ok, %{data: results, proof: nil, verified: false}} - end - - defp distribute_to_stores(query) do - # Route to Oxigraph, Milvus, Tantivy, etc. based on modalities - query.modalities - |> Enum.map(&route_to_store(&1, query)) - |> Task.async_stream(&execute_subquery/1, max_concurrency: 6) - |> Enum.reduce({:ok, []}, &aggregate_results/2) - end - - defp route_to_store("GRAPH", query), do: {:oxigraph, query.graph_ast} - defp route_to_store("VECTOR", query), do: {:milvus, query.vector_ast} - defp route_to_store("DOCUMENT", query), do: {:tantivy, query.document_ast} - defp route_to_store("TENSOR", query), do: {:burn, query.tensor_ast} - defp route_to_store("SEMANTIC", query), do: {:verisim_semantic, query.semantic_ast} - defp route_to_store("TEMPORAL", query), do: {:verisim_temporal, query.temporal_ast} -end ----- - -=== 2.3. Metadata Storage (Corrected) - -**NOT using CockroachDB.** Instead: - -**Registry (ReScript + Raft):** - -[source,elixir] -# lib/verisim/raft_metadata.ex -defmodule VeriSim.RaftMetadata do - @moduledoc """ - Raft consensus for registry metadata. - Stores UUID → store mappings, NOT query results. - - Per design-decisions.adoc: - - Use ReScript + Raft for <20 nodes - - Only migrate to CockroachDB if >20 nodes AND need SQL queries - """ - - use GenServer - - # Raft log entries - @type log_entry :: %{ - term: integer(), - index: integer(), - command: :register_octad | :update_metadata | :register_node, - data: map() - } - - def register_octad(octad_id, store_id, modalities) do - # Append to Raft log (consensus required) - entry = %{ - command: :register_octad, - data: %{octad_id: octad_id, store_id: store_id, modalities: modalities} - } - GenServer.call(__MODULE__, {:append_entry, entry}) - end - - def lookup_octad(octad_id) do - # Read from replicated state (no consensus needed) - GenServer.call(__MODULE__, {:lookup, octad_id}) - end -end ----- - -**Migration Path:** - -[source,text] ----- -Phase 1: ReScript + Raft (in-memory, <20 nodes) ← START HERE - ↓ -Phase 2: ReScript + etcd (20-50 nodes) - ↓ -Phase 3: CockroachDB (50+ nodes, complex queries) ----- - -=== 2.4. Audit Trail (Corrected) - -**NOT using Fluree.** Instead: - -**verisim-temporal (Merkle Trees):** - -[source,rust] -// verisim-temporal/src/audit_log.rs (Rust crate) -use blake3::Hasher; -use serde::{Deserialize, Serialize}; - -#[derive(Serialize, Deserialize)] -pub struct QueryLogEntry { - pub query_id: String, - pub timestamp: i64, - pub query_text: String, - pub result_hash: [u8; 32], // Blake3 hash of results - pub proof: Option>, // ZKP bytes (if dependent-type) - pub parent_hash: [u8; 32], // Previous entry in Merkle chain -} - -impl QueryLogEntry { - pub fn compute_hash(&self) -> [u8; 32] { - let mut hasher = Hasher::new(); - hasher.update(self.query_id.as_bytes()); - hasher.update(&self.timestamp.to_le_bytes()); - hasher.update(self.query_text.as_bytes()); - hasher.update(&self.result_hash); - hasher.update(&self.parent_hash); - *hasher.finalize().as_bytes() - } - - pub fn verify_chain(&self, previous: &QueryLogEntry) -> bool { - self.parent_hash == previous.compute_hash() - } -} - -// Elixir NIF interface -#[rustler::nif] -fn append_query_log( - query_id: String, - query_text: String, - result_hash: Vec, - proof: Option>, - parent_hash: Vec, -) -> Result, String> { - let entry = QueryLogEntry { - query_id, - timestamp: chrono::Utc::now().timestamp(), - query_text, - result_hash: result_hash.try_into().map_err(|_| "Invalid hash")?, - proof, - parent_hash: parent_hash.try_into().map_err(|_| "Invalid parent hash")?, - }; - Ok(entry.compute_hash().to_vec()) -} ----- - -**Elixir Integration:** - -[source,elixir] -# lib/verisim/temporal_log.ex -defmodule VeriSim.TemporalLog do - @moduledoc """ - Immutable audit trail using Merkle trees. - Implemented in verisim-temporal Rust crate (NOT Fluree). - """ - - use Rustler, otp_app: :verisimdb, crate: "verisim_temporal" - - # NIF functions (implemented in Rust) - def append_query_log(_query_id, _query_text, _result_hash, _proof, _parent_hash), - do: :erlang.nif_error(:nif_not_loaded) - - def verify_chain(_start_hash, _end_hash, _entries), - do: :erlang.nif_error(:nif_not_loaded) - - def append(query_id, results, proof) do - result_hash = :crypto.hash(:blake3, :erlang.term_to_binary(results)) - parent_hash = get_latest_hash() - - case append_query_log( - query_id, - get_query_text(query_id), - result_hash, - proof, - parent_hash - ) do - {:ok, new_hash} -> - # Store in append-only log file - persist_to_disk(query_id, new_hash) - :ok - {:error, reason} -> - {:error, reason} - end - end -end ----- - -=== 2.5. Drift Detection Integration - -**Elixir DriftMonitor:** - -[source,elixir] -# lib/verisim/drift_monitor.ex -defmodule VeriSim.DriftMonitor do - @moduledoc """ - Detects and repairs drift across federated nodes. - Integrates with VCL query path. - """ - - use GenServer - - def check_consistency(octad_ids) do - octad_ids - |> Enum.map(&fetch_versions_from_nodes/1) - |> detect_drift() - |> case do - {:ok, :consistent} -> {:ok, :consistent} - {:error, :drift_detected, details} -> {:error, :drift_detected, details} - end - end - - def repair_and_retry(query) do - # Statistical drift repair per whitepaper Section 4 - case apply_repair_policy(query.octad_ids, query.drift_policy) do - {:ok, :repaired} -> - # Retry query after repair - VeriSim.QueryRouter.handle_query(query) - {:error, reason} -> - {:error, reason} - end - end - - defp fetch_versions_from_nodes(octad_id) do - # Query all nodes in federation for versions - nodes = Registry.get_nodes_for_octad(octad_id) - Enum.map(nodes, fn node -> - {node, TemporalLog.get_latest_version(node, octad_id)} - end) - end - - defp detect_drift(versions_by_node) do - # Check for Merkle tree divergence - hashes = Enum.map(versions_by_node, fn {_node, version} -> version.hash end) - - if Enum.uniq(hashes) |> length() == 1 do - {:ok, :consistent} - else - {:error, :drift_detected, %{divergent_hashes: hashes}} - end - end -end ----- - -== 3. API Endpoint Design - -=== 3.1. WASM Proxy Endpoint - -**ReScript WASM module behind Cloudflare SDP:** - -[source,rescript] -// src/vcl/VCLEndpoint.res -module Endpoint = { - @post("/api/v1/query") - let queryEndpoint = (req: {body: string, headers: dict}) => { - let useDependentTypes = Dict.get(req.headers, "X-VCL-Safe") == Some("true") - - // Parse and route query - let result = VCL.handleQuery(req.body, useDependentTypes) - - // Return with appropriate headers - { - status: 200, - headers: { - "Content-Type": "application/json", - "X-VCL-Verified": result.verified ? "true" : "false", - "X-Query-ID": result.queryId, - }, - body: JSON.stringify(result), - } - } -} ----- - -=== 3.2. Client SDK - -**ReScript Client:** - -[source,rescript] -// Client SDK for VCL -module VCLClient = { - type queryOptions = { - useDependentTypes: bool, - driftPolicy: option<[#Strict | #Repair | #Tolerate | #Latest]>, - timeout: option, - } - - let query = async (vcl: string, options: queryOptions) => { - let headers = options.useDependentTypes - ? {"X-VCL-Safe": "true", "Content-Type": "text/vcl"} - : {"Content-Type": "text/vcl"} - - let response = await fetch("/api/v1/query", { - method: "POST", - headers: headers, - body: vcl, - }) - - let result = await response->Response.json - - if (result.verified) { - // Dependent-type response - verify proof - let proofValid = await ProvenLibrary.verify(result.proof, result.data) - if (!proofValid) { - raise(ProofVerificationError("ZKP verification failed")) - } - } else { - Console.warn("Unverified query result (slipstream path)") - } - - result - } -} - -// Usage example -let searchPapers = async () => { - let vcl = ` - SELECT GRAPH, VECTOR - FROM FEDERATION /universities/* - WHERE (h)-[:CITES]->(target) - AND h.embedding SIMILAR TO [0.1, 0.2, 0.3] - PROOF CITATION(CitationContract) - LIMIT 100 - ` - - let result = await VCLClient.query(vcl, { - useDependentTypes: true, - driftPolicy: Some(#Repair), - timeout: Some(30000), - }) - - Console.log(result.data) -} ----- - -== 4. Security and Zero Trust - -=== 4.1. Dependent-Type Path Security - -**Guarantees:** - -1. **ZKP Verification** - All queries require valid proofs from `proven-library` -2. **Drift Detection** - Consistency checked before execution -3. **Audit Trail** - Immutable Merkle chain in `verisim-temporal` -4. **Access Control** - Enforced by semantic contracts - -**Example Flow:** - -[source,text] ----- -1. Client sends query with PROOF clause -2. ReScript parser creates typed AST -3. TypeChecker validates with proven-library - ├─ Generates ZKP for query constraints - └─ Validates modality compatibility -4. DriftMonitor checks consistency - ├─ If STRICT: Fail on any drift - ├─ If REPAIR: Auto-repair and continue - └─ If TOLERATE: Continue despite drift -5. FederatedExecutor distributes to stores -6. Results aggregated with proof attached -7. TemporalLog appends to Merkle chain -8. Response returned with ZKP ----- - -=== 4.2. Slipstream Path Security - -**Limitations:** - -1. **No ZKP** - Results not formally verified -2. **No Drift Check** - May return inconsistent data -3. **Rate Limited** - Enforced by Cloudflare SDP -4. **Scoped to Public Data** - Semantic contracts still enforced - -**Use Cases:** - -- Exploratory queries -- Performance-critical applications -- Internal queries with trusted data -- Real-time recommendations - -== 5. Performance Comparison - -|=== -| Metric | Slipstream | Dependent-Type | Notes - -| **Parse Time** -| 5-10ms -| 20-50ms -| Typed AST more complex - -| **Type Check** -| N/A -| 50-100ms -| proven-library validation - -| **Drift Check** -| N/A -| 20-200ms -| Depends on federation size - -| **ZKP Generation** -| N/A -| 100-500ms -| Per whitepaper Section 5 - -| **Query Execution** -| 50-500ms -| 50-500ms -| Same (federated execution) - -| **Total Latency** -| 55-510ms -| 240-1350ms -| 4-5x slower for verification - -| **Audit Trail** -| Logged -| Logged + Proof -| Both paths logged - -| **Guarantees** -| None -| Formal -| Trade-off -|=== - -== 6. Implementation Roadmap - -=== Phase 1: Grammar & Slipstream Parser (Week 1-2) - -- [ ] Finalize VCL grammar (link:vcl-grammar.ebnf[]) -- [ ] Implement slipstream parser in ReScript -- [ ] Add basic AST types -- [ ] Test with simple queries - -=== Phase 2: Type Checker (Week 3-4) - -- [ ] Design typed AST with dependent types -- [ ] Integrate `proven-library` for ZKP -- [ ] Add modality compatibility rules -- [ ] Test dependent-type queries - -=== Phase 3: Elixir Federation (Week 5-6) - -- [ ] Extend `VeriSim.QueryRouter` for dual-path -- [ ] Integrate `DriftMonitor` checks -- [ ] Implement result aggregation -- [ ] Test federated queries - -=== Phase 4: WASM Endpoint (Week 7-8) - -- [ ] Compile VCL to WASM -- [ ] Deploy behind Cloudflare SDP -- [ ] Add rate limiting -- [ ] Complete audit logging - -== 7. References - -- link:vcl-grammar.ebnf[VCL Grammar (EBNF)] -- link:vcl-examples.adoc[VCL Examples] -- link:../WHITEPAPER.md[VeriSimDB White Paper] -- link:../design-decisions.adoc[Architecture Design Decisions] -- link:technical-specification-kraft-metadata-log.adoc[KRaft Metadata Log] -- link:challenges-federated.adoc[Federated Deployment Challenges] - -== Appendix A: Migration from Initial Design - -|=== -| Initial Design | Corrected Design | Justification - -| CockroachDB for metadata -| ReScript + Raft -| Simpler, sufficient for <20 nodes (link:../design-decisions.adoc:247[]) - -| Fluree for audit trails -| verisim-temporal (Merkle trees) -| Lighter weight, no blockchain overhead (link:../design-decisions.adoc:140[]) - -| No drift integration -| DriftMonitor checks -| Core feature per whitepaper Section 4 - -| SQL queries on metadata -| Key-value registry lookups -| Registry is simple UUID → store mapping - -| CockroachDB geo-partitioning -| Raft quorum -| Only migrate to CockroachDB if >20 nodes (link:../design-decisions.adoc:224[]) -|=== - -== Appendix B: Technology Stack - -|=== -| Component | Technology | Why - -| **Parser** -| ReScript + parser combinators -| Type safety, WASM target - -| **Registry** -| ReScript + Raft (Elixir) -| Whitepaper Section 2.1 - -| **Orchestration** -| Elixir GenServers -| Fault tolerance, distribution - -| **ZKP** -| proven-library (Rust) -| Your existing ZKP library - -| **Audit Trail** -| verisim-temporal (Rust) -| Merkle trees, immutability - -| **Graph Store** -| Oxigraph (Rust) -| SPARQL, RDF compliance - -| **Vector Store** -| Milvus -| HNSW, GPU support - -| **Document Store** -| Tantivy (Rust) -| Full-text, inverted indices - -| **Tensor Store** -| Burn (Rust) -| ndarray, WASM target - -| **Semantic Store** -| verisim-semantic (Rust) -| CBOR + proven integration - -| **Temporal Store** -| verisim-temporal (Rust) -| Version trees, drift detection -|=== diff --git a/verisimdb/docs/vcl-examples.adoc b/verisimdb/docs/vcl-examples.adoc deleted file mode 100644 index 9e7a2baf..00000000 --- a/verisimdb/docs/vcl-examples.adoc +++ /dev/null @@ -1,1077 +0,0 @@ -// SPDX-License-Identifier: CC-BY-SA-4.0 - -= VeriSim Consonance Language (VCL) Examples -:toc: left -:toclevels: 3 -:sectnums: - -== Overview - -This document provides **63 comprehensive examples** of VCL queries and mutations for both **dependent-type** (formally verified) and **slipstream** (fast, unverified) paths, including backwards compatibility (VERSION), drift handling (DRIFT REPAIR), SQL-compatible extensions (ORDER BY, GROUP BY, HAVING, aggregates, column projections), multi-proof composition, cross-modal correlation conditions, and write path mutations (INSERT/UPDATE/DELETE). - -== Basic Queries - -=== Example 1: Simple Octad Retrieval - -**Slipstream** (no proof): -[source,vcl] ----- -SELECT * -FROM HEXAD 550e8400-e29b-41d4-a716-446655440000 ----- - -**Dependent-Type** (with proof): -[source,vcl] ----- -SELECT * -FROM HEXAD 550e8400-e29b-41d4-a716-446655440000 -PROOF EXISTENCE(ExistenceContract) ----- - -=== Example 2: Modality-Specific Selection - -**Graph Only:** -[source,vcl] ----- -SELECT GRAPH -FROM HEXAD 550e8400-e29b-41d4-a716-446655440000 -WHERE (h)-[:CITES]->(target) ----- - -**Vector Only:** -[source,vcl] ----- -SELECT VECTOR -FROM HEXAD 550e8400-e29b-41d4-a716-446655440000 -WHERE h.embedding SIMILAR TO [0.1, 0.2, 0.3, 0.4] ----- - -**Multiple Modalities:** -[source,vcl] ----- -SELECT GRAPH, VECTOR, SEMANTIC -FROM HEXAD 550e8400-e29b-41d4-a716-446655440000 -PROOF INTEGRITY(DataIntegrityContract) ----- - -== Graph Queries (Oxigraph) - -=== Example 3: Citation Chain Traversal - -**Slipstream:** -[source,vcl] ----- -SELECT GRAPH -FROM FEDERATION /universities/* -WHERE (h)-[:CITES*1..5]->(paper) - AND paper.title LIKE "%climate change%" -LIMIT 100 ----- - -**Dependent-Type:** -[source,vcl] ----- -SELECT GRAPH -FROM FEDERATION /universities/* WITH DRIFT STRICT -WHERE (h)-[:CITES*1..5]->(paper) - AND paper.title LIKE "%climate change%" - AND SATISFIES CitationContract -PROOF CITATION(CitationChainContract) -LIMIT 100 ----- - -=== Example 4: Bidirectional Graph Traversal - -[source,vcl] ----- -SELECT GRAPH -FROM HEXAD abc12345-0000-0000-0000-000000000000 -WHERE (h)-[:CO_AUTHOR]-(colleague) - AND (colleague)-[:AFFILIATED_WITH]->(institution) - AND institution.country == "USA" ----- - -=== Example 5: Property Graph Query - -[source,vcl] ----- -SELECT GRAPH -FROM STORE oxigraph-node-1 -WHERE (author)-[:WROTE]->(paper) - AND paper.citations > 100 - AND paper.year >= 2020 -PROOF ACCESS(ReadAccessContract) ----- - -== Vector Queries (Milvus) - -=== Example 6: Similarity Search (ANN) - -**Slipstream:** -[source,vcl] ----- -SELECT VECTOR -FROM FEDERATION /embeddings/* -WHERE h.embedding SIMILAR TO [0.12, 0.45, 0.78, ...] WITHIN 0.8 -LIMIT 10 ----- - -**Dependent-Type:** -[source,vcl] ----- -SELECT VECTOR -FROM FEDERATION /embeddings/* WITH DRIFT REPAIR -WHERE h.embedding SIMILAR TO [0.12, 0.45, 0.78, ...] - WITHIN 0.9 - USING COSINE - AND SATISFIES EmbeddingQualityContract(min_dimension=768) -PROOF INTEGRITY(VectorIntegrityContract) -LIMIT 20 ----- - -=== Example 7: K-Nearest Neighbors - -[source,vcl] ----- -SELECT VECTOR -FROM HEXAD query-embedding-uuid -WHERE h.embedding NEAREST 15 USING DOT_PRODUCT - AND h.metadata.domain == "neuroscience" ----- - -=== Example 8: Filtered Vector Search - -[source,vcl] ----- -SELECT VECTOR, DOCUMENT -FROM FEDERATION /research-papers/* -WHERE h.embedding SIMILAR TO [0.5, 0.3, 0.2, ...] - AND FULLTEXT CONTAINS "deep learning" - AND h.year >= 2023 -LIMIT 50 ----- - -== Semantic Queries (verisim-semantic + proven) - -=== Example 9: ZKP-Verified Access - -**Dependent-Type:** -[source,vcl] ----- -SELECT SEMANTIC, DOCUMENT -FROM HEXAD confidential-data-uuid -WHERE SATISFIES AccessControlContract(role=researcher, institution=MIT) - AND HAS PROOF ACCESS -PROOF ACCESS(InstitutionalAccessContract) WITH verifier=ethics-board ----- - -=== Example 10: Provenance Verification - -**Dependent-Type:** -[source,vcl] ----- -SELECT SEMANTIC, TEMPORAL -FROM FEDERATION /clinical-trials/* -WHERE SATISFIES ProvenanceContract - AND VERIFIED BY fda-validator - AND AS OF 2025-01-01T00:00:00Z -PROOF PROVENANCE(ClinicalTrialContract) ----- - -=== Example 11: Multi-Contract Validation - -**Dependent-Type:** -[source,vcl] ----- -SELECT * -FROM HEXAD research-dataset-uuid -WHERE SATISFIES GDPRContract(anonymized=true) - AND SATISFIES FAIRContract(findable=true, accessible=true) - AND SATISFIES LicenseContract(license=CC-BY-4.0) -PROOF CUSTOM(ComplianceContract) WITH auditor=legal-team ----- - -== Document Queries (Tantivy) - -=== Example 12: Full-Text Search - -**Slipstream:** -[source,vcl] ----- -SELECT DOCUMENT -FROM FEDERATION /archives/* -WHERE FULLTEXT CONTAINS "quantum computing" -LIMIT 1000 ----- - -**Dependent-Type:** -[source,vcl] ----- -SELECT DOCUMENT -FROM FEDERATION /archives/* WITH DRIFT TOLERATE -WHERE FULLTEXT CONTAINS "quantum computing" - AND SATISFIES CitationContract -PROOF EXISTENCE(ArchiveContract) -LIMIT 100 ----- - -=== Example 13: Regex Pattern Matching - -[source,vcl] ----- -SELECT DOCUMENT -FROM STORE tantivy-node-3 -WHERE FULLTEXT MATCHES /\b(AI|ML|DL)\b.*ethics/ - AND FIELD author == "Jane Doe" - AND FIELD year >= 2020 ----- - -=== Example 14: Structured Field Queries - -[source,vcl] ----- -SELECT DOCUMENT, SEMANTIC -FROM FEDERATION /journals/* -WHERE FIELD doi LIKE "10.1234/%" - AND FIELD impact_factor >= 5.0 - AND SATISFIES PeerReviewContract ----- - -== Tensor Queries (Burn/ndarray) - -=== Example 15: Tensor Shape Filtering - -[source,vcl] ----- -SELECT TENSOR -FROM HEXAD tensor-data-uuid -WHERE tensor.data SHAPE == [256, 256, 3] - AND tensor.dtype == "float32" ----- - -=== Example 16: Tensor Operations - -[source,vcl] ----- -SELECT TENSOR -FROM STORE burn-node-2 -WHERE tensor.data RANK == 3 - AND tensor.statistics.mean > 0.5 - AND tensor.statistics.std < 0.2 ----- - -== Temporal Queries (verisim-temporal) - -=== Example 17: Point-in-Time Query - -**Dependent-Type:** -[source,vcl] ----- -SELECT TEMPORAL, DOCUMENT -FROM HEXAD legal-document-uuid -WHERE AS OF 2024-06-15T10:30:00Z -PROOF PROVENANCE(TemporalContract) ----- - -=== Example 18: Version Range Query - -[source,vcl] ----- -SELECT TEMPORAL, GRAPH -FROM FEDERATION /contracts/* -WHERE BETWEEN 2024-01-01T00:00:00Z AND 2024-12-31T23:59:59Z - AND MODIFIED BY legal-team ----- - -=== Example 19: Specific Version - -**Dependent-Type:** -[source,vcl] ----- -SELECT * -FROM HEXAD retracted-paper-uuid -WHERE VERSION v3-final - AND SATISFIES RetractionContract(reason=error-in-methodology) -PROOF PROVENANCE(RetractionContract) ----- - -== Multi-Modal Queries - -=== Example 20: Neurosymbolic AI Query - -**Dependent-Type:** -[source,vcl] ----- --- Find papers similar to embedding that cite a specific work -SELECT GRAPH, VECTOR, DOCUMENT -FROM FEDERATION /research-db/* WITH DRIFT REPAIR -WHERE h.embedding SIMILAR TO [0.1, 0.2, ...] WITHIN 0.85 - AND (h)-[:CITES]->(target_paper) - AND target_paper.uuid == "123e4567-e89b-12d3-a456-426614174000" - AND FULLTEXT CONTAINS "neural-symbolic integration" - AND SATISFIES CitationContract -PROOF CITATION(NeurosymbolicContract) -LIMIT 50 ----- - -=== Example 21: Cross-Modal Consistency Check - -**Dependent-Type:** -[source,vcl] ----- --- Verify document content matches vector embedding -SELECT DOCUMENT, VECTOR, SEMANTIC -FROM HEXAD suspicious-entry-uuid -WHERE SATISFIES ConsistencyContract( - vector_matches_text=true, - semantic_alignment=high - ) -PROOF INTEGRITY(CrossModalContract) ----- - -=== Example 22: Federated Provenance Audit - -**Dependent-Type:** -[source,vcl] ----- --- Audit trail across multiple institutions -SELECT TEMPORAL, SEMANTIC, GRAPH -FROM FEDERATION /universities/* WITH DRIFT STRICT -WHERE AS OF 2025-01-01T00:00:00Z - AND (h)-[:DERIVED_FROM]->(source) - AND VERIFIED BY institutional-review-board - AND SATISFIES ProvenanceContract - AND SATISFIES EthicsContract(irb_approved=true) -PROOF PROVENANCE(MultiInstitutionalContract) WITH auditor=nsf ----- - -== Drift Handling Queries - -=== Example 23: Strict Drift (Fail on Inconsistency) - -**Dependent-Type:** -[source,vcl] ----- -SELECT * -FROM FEDERATION /legal-records/* WITH DRIFT STRICT -WHERE AS OF 2024-01-01T00:00:00Z -PROOF INTEGRITY(LegalContract) ----- - -**Behavior:** Query fails if any node has drifted from the specified timestamp. - -=== Example 24: Auto-Repair Drift - -**Dependent-Type:** -[source,vcl] ----- -SELECT GRAPH, DOCUMENT -FROM FEDERATION /collaborative-wiki/* WITH DRIFT REPAIR -WHERE (h)-[:LINKS_TO]->(target) - AND FULLTEXT CONTAINS "VeriSimDB" -PROOF EXISTENCE(WikiContract) ----- - -**Behavior:** Elixir `DriftMonitor` detects and repairs inconsistencies before returning results. - -=== Example 25: Tolerate Drift (Return Inconsistent Data) - -**Slipstream:** -[source,vcl] ----- -SELECT * -FROM FEDERATION /mirrors/* WITH DRIFT TOLERATE -WHERE FULLTEXT CONTAINS "archived content" -LIMIT 1000 ----- - -**Behavior:** Returns data even if nodes are inconsistent. Useful for exploratory queries. - -=== Example 26: Latest Version (Ignore Drift) - -**Slipstream:** -[source,vcl] ----- -SELECT DOCUMENT -FROM FEDERATION /news-feeds/* WITH DRIFT LATEST -WHERE FULLTEXT CONTAINS "breaking news" -LIMIT 10 ----- - -**Behavior:** Uses the most recent version from each node, ignoring temporal consistency. - -== Performance Optimization Examples - -=== Example 27: Pagination for Large Result Sets - -[source,vcl] ----- -SELECT DOCUMENT -FROM FEDERATION /archives/* -WHERE FULLTEXT CONTAINS "machine learning" -LIMIT 100 -OFFSET 500 ----- - -=== Example 28: Store-Specific Query (Bypass Federation) - -[source,vcl] ----- --- Query specific store for lowest latency -SELECT VECTOR -FROM STORE milvus-us-east-1 -WHERE h.embedding NEAREST 20 USING COSINE ----- - -=== Example 29: Minimal Modality Selection - -[source,vcl] ----- --- Only request needed modalities to reduce data transfer -SELECT GRAPH(nodes, edges) -FROM HEXAD abc-123 -WHERE (h)-[:CITES]->(target) ----- - -== Error Handling Examples - -=== Example 30: Proof Verification Failure - -**Query:** -[source,vcl] ----- -SELECT SEMANTIC -FROM HEXAD protected-data-uuid -WHERE SATISFIES AccessControlContract(role=guest) -PROOF ACCESS(GuestAccessContract) ----- - -**Expected Error:** -[source,json] ----- -{ - "error": "ProofVerificationFailed", - "message": "ZKP verification failed for AccessControlContract", - "details": { - "contract": "GuestAccessContract", - "reason": "Insufficient permissions (requires role=researcher)" - } -} ----- - -=== Example 31: Drift Detection Failure - -**Query:** -[source,vcl] ----- -SELECT * -FROM FEDERATION /distributed-ledger/* WITH DRIFT STRICT -WHERE AS OF 2024-01-01T00:00:00Z ----- - -**Expected Error:** -[source,json] ----- -{ - "error": "DriftDetected", - "message": "Inconsistency detected across federated nodes", - "details": { - "divergent_nodes": ["node-2", "node-5"], - "drift_type": "TemporalInconsistency", - "suggestion": "Use DRIFT REPAIR or DRIFT TOLERATE" - } -} ----- - -=== Example 32: Malformed Query - -**Query:** -[source,vcl] ----- -SELECT GRAPH -FROM HEXAD not-a-valid-uuid -WHERE (h)-[:INVALID_SYNTAX ----- - -**Expected Error:** -[source,json] ----- -{ - "error": "ParseError", - "message": "Failed to parse VCL query", - "details": { - "line": 3, - "column": 27, - "expected": "]->" - } -} ----- - -== Advanced Use Cases - -=== Example 33: Recursive Citation Network - -**Dependent-Type:** -[source,vcl] ----- --- Find all papers in transitive citation closure -SELECT GRAPH -FROM FEDERATION /research-db/* WITH DRIFT REPAIR -WHERE (seed_paper)-[:CITES*1..10]->(h) - AND seed_paper.uuid == "root-paper-uuid" - AND h.year >= 2020 - AND SATISFIES CitationContract -PROOF CITATION(TransitiveClosureContract) -LIMIT 1000 ----- - -=== Example 34: Multi-Institutional Data Aggregation - -**Dependent-Type:** -[source,vcl] ----- --- Aggregate COVID-19 data from multiple hospitals -SELECT SEMANTIC, TEMPORAL, DOCUMENT -FROM FEDERATION /hospitals/* WITH DRIFT STRICT -WHERE AS OF 2024-12-31T23:59:59Z - AND SATISFIES HIPAAContract(anonymized=true, aggregated=true) - AND SATISFIES IRBContract(protocol=COVID-2024-001) - AND FIELD diagnosis LIKE "%COVID%" -PROOF PROVENANCE(MultiInstitutionalHealthContract) - WITH auditor=cdc ----- - -=== Example 35: Real-Time Embeddings Query - -**Slipstream:** -[source,vcl] ----- --- Fast similarity search for recommendation system -SELECT VECTOR -FROM STORE milvus-cache -WHERE h.embedding SIMILAR TO [0.5, 0.3, ...] WITHIN 0.75 - AND h.metadata.category == "electronics" - AND h.metadata.in_stock == true -LIMIT 20 ----- - -== Comparison: Dependent-Type vs Slipstream - -=== Same Query, Two Paths - -**Slipstream (Fast, Unverified):** -[source,vcl] ----- -SELECT GRAPH, VECTOR -FROM FEDERATION /papers/* -WHERE (h)-[:CITES]->(target) - AND h.embedding SIMILAR TO [0.1, 0.2, 0.3] -LIMIT 100 ----- - -**Dependent-Type (Slow, Verified):** -[source,vcl] ----- -SELECT GRAPH, VECTOR -FROM FEDERATION /papers/* WITH DRIFT STRICT -WHERE (h)-[:CITES]->(target) - AND h.embedding SIMILAR TO [0.1, 0.2, 0.3] - AND SATISFIES CitationContract - AND SATISFIES EmbeddingQualityContract(dimension=768) -PROOF CITATION(VerifiedCitationContract) -LIMIT 100 ----- - -**Performance Comparison:** - -|=== -| Metric | Slipstream | Dependent-Type - -| Latency -| ~50ms -| ~300ms (includes ZKP generation) - -| Guarantees -| None -| Formal verification via ZKP - -| Use Case -| Exploratory analysis -| Audit-ready queries - -| Audit Trail -| Logged to verisim-temporal -| Logged with proof to verisim-temporal -|=== - -== Summary - -VCL provides two execution paths: - -1. **Dependent-Type Path** (`PROOF` clause present): - - Formal verification via `proven-library` - - Drift detection/repair via Elixir `DriftMonitor` - - ZKP generation for audit trails - - Slower but audit-ready - -2. **Slipstream Path** (no `PROOF` clause): - - Fast, unverified execution - - No proof generation overhead - - Suitable for exploratory queries - - Still logged for audit (but without proofs) - -**Recommendation:** Use dependent-type path for: -- Compliance queries (GDPR, HIPAA, FAIR) -- Multi-institutional data sharing -- Citation verification -- Provenance audits - -Use slipstream path for: -- Real-time recommendations -- Exploratory data analysis -- Performance-critical applications -- Internal queries with trusted data - -== Backwards Compatibility Queries - -=== Example 36: Version Pinning - -**Query with explicit VCL version:** -[source,vcl] ----- -VERSION 1.0; - -SELECT GRAPH, DOCUMENT -FROM verisim:semantic -WHERE octad.types INCLUDES "https://schema.org/Person" -LIMIT 10; ----- - -**Query using deprecated syntax (with warning):** -[source,vcl] ----- -VERSION 0.9; -- Deprecated syntax accepted - -FROM verisim:semantic -SELECT GRAPH, DOCUMENT -WHERE octad.types INCLUDES "https://schema.org/Person" -LIMIT 10; ----- - -=== Example 37: Federation Version Negotiation - -**Query with version compatibility check:** -[source,vcl] ----- -VERSION 1.0; - -SELECT * -FROM FEDERATION /universities/* -WITH DRIFT LATEST -- Use latest schema version -WHERE octad.types INCLUDES "Paper" -LIMIT 100; ----- - -== Drift Handling Queries - -=== Example 38: Drift Detection - -**Detect cross-modal drift:** -[source,vcl] ----- -SELECT DRIFT -FROM verisim:all_modalities -WHERE octad.id = @octad_id; - --- Returns: { --- "drifts": [ --- {"type": "title_mismatch", "graph": "ML Paper", "document": "Machine Learning Paper"} --- ] --- } ----- - -=== Example 39: Drift Repair - Latest Wins - -**Repair using most recent version:** -[source,vcl] ----- -DRIFT REPAIR -FROM verisim:graph -WHERE octad.id = @id -USING STRATEGY latest_wins; ----- - -=== Example 40: Drift Repair - Quorum - -**Repair using majority voting:** -[source,vcl] ----- -DRIFT REPAIR -FROM verisim:federation /universities/* -WHERE octad.id = @id -USING STRATEGY quorum; ----- - -=== Example 41: Drift-Aware Query - -**Query with drift tolerance:** -[source,vcl] ----- -SELECT GRAPH, DOCUMENT -FROM FEDERATION /universities/* -WITH DRIFT TOLERATE -- Return data despite drift -WHERE octad.types INCLUDES "Paper" -LIMIT 100; ----- - -=== Example 42: Byzantine-Resistant Query - -**Query with drift repair:** -[source,vcl] ----- -SELECT GRAPH, DOCUMENT -FROM FEDERATION /universities/* -WITH DRIFT REPAIR -- Auto-repair detected drift -WHERE octad.types INCLUDES "Paper" -LIMIT 100; ----- - -== SQL-Compatible Extensions - -=== Example 43: Column Selection within Modality - -**Select specific fields from a modality (slipstream):** -[source,vcl] ----- -SELECT DOCUMENT.name, DOCUMENT.severity -FROM HEXAD 550e8400-e29b-41d4-a716-446655440000 ----- - -**Mixed modalities and column projections:** -[source,vcl] ----- -SELECT GRAPH, DOCUMENT.name, DOCUMENT.severity -FROM FEDERATION /universities/* -LIMIT 50 ----- - -=== Example 44: ORDER BY — Sort Results - -**Sort by severity descending:** -[source,vcl] ----- -SELECT DOCUMENT -FROM STORE archive-1 -WHERE FULLTEXT CONTAINS "security" -ORDER BY DOCUMENT.severity DESC -LIMIT 50 ----- - -**Multi-field sort:** -[source,vcl] ----- -SELECT DOCUMENT.name, DOCUMENT.severity -FROM FEDERATION /archives/* -ORDER BY DOCUMENT.severity DESC, DOCUMENT.name ASC -LIMIT 100 ----- - -=== Example 45: COUNT(*), SUM, AVG — Aggregate Functions - -**Count all matching octads:** -[source,vcl] ----- -SELECT COUNT(*) -FROM FEDERATION /archives/* -WHERE FULLTEXT CONTAINS "vulnerability" ----- - -**Aggregate with field reference:** -[source,vcl] ----- -SELECT AVG(DOCUMENT.severity), MIN(DOCUMENT.severity), MAX(DOCUMENT.severity) -FROM FEDERATION /scans/* ----- - -**Sum and count together:** -[source,vcl] ----- -SELECT COUNT(*), SUM(DOCUMENT.severity) -FROM STORE tantivy-node-1 -WHERE FIELD severity > 3 ----- - -=== Example 46: GROUP BY / HAVING — Grouping with Filters - -**Group results by field:** -[source,vcl] ----- -SELECT DOCUMENT.name, COUNT(*), AVG(DOCUMENT.severity) -FROM FEDERATION /universities/* -GROUP BY DOCUMENT.name -LIMIT 100 ----- - -**GROUP BY with HAVING filter:** -[source,vcl] ----- -SELECT DOCUMENT.name, COUNT(*) -FROM FEDERATION /archives/* -GROUP BY DOCUMENT.name -HAVING FIELD count > 3 -ORDER BY DOCUMENT.name ASC -LIMIT 50 ----- - -=== Example 47: Full SQL-Compatible Query - -**Combining all SQL-compatible features with VCL's model:** -[source,vcl] ----- -SELECT DOCUMENT.name, DOCUMENT.severity, COUNT(*), SUM(DOCUMENT.severity), AVG(DOCUMENT.severity) -FROM FEDERATION /universities/* WITH DRIFT REPAIR -WHERE FIELD severity > 3 -GROUP BY DOCUMENT.name, DOCUMENT.severity -HAVING FIELD total > 10 -ORDER BY DOCUMENT.severity DESC, DOCUMENT.name ASC -LIMIT 100 -OFFSET 20 ----- - -=== Example 48: SQL-Compatible with Dependent-Type Verification - -**Combining aggregates with formal verification (VCL-UT path):** -[source,vcl] ----- -SELECT DOCUMENT.name, COUNT(*) -FROM FEDERATION /universities/* WITH DRIFT STRICT -GROUP BY DOCUMENT.name -PROOF INTEGRITY(DataIntegrityContract) -ORDER BY DOCUMENT.name ASC -LIMIT 50 ----- - -NOTE: In the dependent-type path, aggregation results are also covered by the proof -guarantee — the ZKP witnesses include the aggregate computation, ensuring verified totals. - -== Multi-Proof Composition (v2.0) - -=== Example 49: Dual Proof — Existence and Integrity - -**Dependent-Type:** -[source,vcl] ----- -SELECT * -FROM HEXAD 550e8400-e29b-41d4-a716-446655440000 -PROOF EXISTENCE(ExistenceContract) AND INTEGRITY(DataIntegrityContract) ----- - -Both proofs must pass before results are returned. The ZKP witnesses for each proof -are generated independently and composed into a single proof bundle. - -=== Example 50: Triple Proof — Access, Provenance, and Integrity - -**Dependent-Type:** -[source,vcl] ----- -SELECT SEMANTIC, TEMPORAL, DOCUMENT -FROM FEDERATION /hospitals/* WITH DRIFT STRICT -WHERE SATISFIES HIPAAContract(anonymized=true) -PROOF ACCESS(InstitutionalAccessContract) AND PROVENANCE(ClinicalTrialContract) AND INTEGRITY(DataIntegrityContract) -LIMIT 100 ----- - -Multi-proof composition validates that the user has access, the data lineage -is verifiable, and the data has not been tampered with — all three conditions -must hold simultaneously. - -=== Example 51: Citation and Custom Proof - -**Dependent-Type:** -[source,vcl] ----- -SELECT GRAPH, DOCUMENT -FROM FEDERATION /journals/* WITH DRIFT REPAIR -WHERE (h)-[:CITES]->(target) - AND FULLTEXT CONTAINS "reproducibility crisis" -PROOF CITATION(CitationChainContract) AND CUSTOM(ReproducibilityContract) -LIMIT 50 ----- - -== Cross-Modal Correlation Queries (v2.0) - -=== Example 52: Cross-Modal Field Compare - -**Compare fields across modalities:** -[source,vcl] ----- -SELECT DOCUMENT, GRAPH -FROM HEXAD 550e8400-e29b-41d4-a716-446655440000 -WHERE DOCUMENT.severity > GRAPH.centrality ----- - -Evaluates after both modalities are fetched, comparing the document severity -field against the graph centrality field on the same octad. - -=== Example 53: Drift Between Modalities - -**Detect embedding drift:** -[source,vcl] ----- -SELECT VECTOR, DOCUMENT -FROM FEDERATION /research-papers/* -WHERE DRIFT(VECTOR, DOCUMENT) > 0.3 -LIMIT 50 ----- - -Returns octads where the vector embedding has drifted more than 0.3 from the -document representation. Useful for identifying stale embeddings. - -=== Example 54: Consistency Check - -**Verify cross-modal consistency:** -[source,vcl] ----- -SELECT VECTOR, SEMANTIC -FROM HEXAD suspicious-entry-uuid -WHERE CONSISTENT(VECTOR, SEMANTIC) USING COSINE -PROOF INTEGRITY(CrossModalContract) ----- - -Checks that the vector representation and semantic annotation are consistent -using cosine similarity as the metric. - -=== Example 55: Modality Existence Check - -**Filter by available modalities:** -[source,vcl] ----- -SELECT * -FROM FEDERATION /archives/* -WHERE VECTOR EXISTS - AND TENSOR NOT EXISTS -LIMIT 100 ----- - -Returns octads that have vector embeddings but lack tensor data. Useful for -identifying octads that need tensor representations generated. - -=== Example 56: Combined Cross-Modal and Standard Conditions - -**Mix cross-modal with single-modality conditions:** -[source,vcl] ----- -SELECT GRAPH, VECTOR, DOCUMENT -FROM FEDERATION /research-db/* WITH DRIFT REPAIR -WHERE FULLTEXT CONTAINS "neural networks" - AND DRIFT(VECTOR, DOCUMENT) < 0.2 - AND VECTOR EXISTS - AND GRAPH EXISTS -PROOF CITATION(NeuralNetworkContract) AND INTEGRITY(DataIntegrityContract) -LIMIT 50 ----- - -Pushdown conditions (FULLTEXT CONTAINS) are sent to individual stores, -while cross-modal conditions (DRIFT, EXISTS) are evaluated post-fetch -on full octad data. - -== Write Path: INSERT / UPDATE / DELETE (v2.0) - -=== Example 57: Insert a New Octad - -**Insert with document and vector data:** -[source,vcl] ----- -INSERT HEXAD WITH - DOCUMENT(title = "New Research Paper", author = "Jane Doe", severity = 3), - VECTOR([0.12, 0.45, 0.78, 0.23, 0.91]) ----- - -Creates a new octad with document and vector modalities populated. -The octad ID is auto-generated and returned in the result. - -=== Example 58: Insert with Graph Relationship - -**Insert with cross-modal data:** -[source,vcl] ----- -INSERT HEXAD WITH - DOCUMENT(title = "Follow-Up Study", year = 2026), - VECTOR([0.5, 0.3, 0.2, 0.8]), - GRAPH(CITES, 550e8400-e29b-41d4-a716-446655440000) ----- - -Creates a octad with document content, embedding, and a graph edge -(CITES relationship to an existing octad). - -=== Example 59: Verified Insert with Proof - -**Insert with integrity verification:** -[source,vcl] ----- -INSERT HEXAD WITH - DOCUMENT(title = "Clinical Trial Result", status = "verified"), - SEMANTIC(ClinicalTrialContract), - TEMPORAL(2026-02-13T10:30:00Z) -PROOF INTEGRITY(WriteContract) AND PROVENANCE(AuditTrailContract) ----- - -The ZKP proofs are generated for the write operation itself, ensuring the -insert is verified and auditable. - -=== Example 60: Update Octad Fields - -**Update specific fields:** -[source,vcl] ----- -UPDATE HEXAD 550e8400-e29b-41d4-a716-446655440000 -SET DOCUMENT.title = "Corrected Title", DOCUMENT.severity = 5 ----- - -Updates the document modality fields without affecting other modalities. -Drift detection will fire if the change creates cross-modal inconsistency. - -=== Example 61: Verified Update with Proof - -**Update with access proof:** -[source,vcl] ----- -UPDATE HEXAD 550e8400-e29b-41d4-a716-446655440000 -SET DOCUMENT.status = "retracted", DOCUMENT.retraction_reason = "Data fabrication" -PROOF ACCESS(EditAccessContract) AND PROVENANCE(RetractionContract) ----- - -The update requires proof of edit access and creates a provenance record -of the retraction in the audit trail. - -=== Example 62: Delete a Octad - -**Simple deletion:** -[source,vcl] ----- -DELETE HEXAD 550e8400-e29b-41d4-a716-446655440000 ----- - -Removes the octad and all its modality data. This operation is logged -in verisim-temporal for auditability. - -=== Example 63: Verified Delete with Proof - -**Delete with access and provenance proof:** -[source,vcl] ----- -DELETE HEXAD 550e8400-e29b-41d4-a716-446655440000 -PROOF ACCESS(AdminAccessContract) AND PROVENANCE(DeletionAuditContract) ----- - -Requires proof of admin access before deletion. The provenance proof -ensures the deletion is recorded in the audit chain. - -== References - -- link:vcl-grammar.ebnf[VCL Grammar (EBNF)] -- link:backwards-compatibility.adoc[Backwards Compatibility Strategy] -- link:drift-handling.adoc[Drift Handling Documentation] -- link:../WHITEPAPER.md[VeriSimDB White Paper] -- link:technical-specification-kraft-metadata-log.adoc[KRaft Metadata Log Specification] -- link:../design-decisions.adoc[Architecture Design Decisions] diff --git a/verisimdb/docs/vcl-formal-semantics.adoc b/verisimdb/docs/vcl-formal-semantics.adoc deleted file mode 100644 index 2e7e0a98..00000000 --- a/verisimdb/docs/vcl-formal-semantics.adoc +++ /dev/null @@ -1,654 +0,0 @@ -// SPDX-License-Identifier: CC-BY-SA-4.0 - -= VCL Formal Semantics -:toc: left -:toclevels: 4 -:sectnums: -:stem: latexmath - -== Overview - -This document specifies the **formal operational semantics** and **static type system** for VeriSim Consonance Language (VCL). VCL has two execution paths with different semantic guarantees: - -1. **Dependent-Type Path** - Queries with `PROOF` clause, formally verified via ZKP -2. **Slipstream Path** - Queries without `PROOF` clause, fast but unverified - -== Notation - -We use standard mathematical notation for formal semantics: - -[cols="1,3"] -|=== -|Notation |Meaning - -|stem:[\Gamma] -|Type environment (context) - -|stem:[\Gamma \vdash e : \tau] -|Expression `e` has type stem:[\tau] in context stem:[\Gamma] - -|stem:[e \Rightarrow v] -|Expression `e` evaluates to value `v` (big-step semantics) - -|stem:[e \rightarrow e'] -|Expression `e` reduces to `e'` (small-step semantics) - -|stem:[\llbracket e \rrbracket_{\rho}] -|Denotational semantics of `e` in environment stem:[\rho] - -|stem:[\pi : \phi] -|Proof term stem:[\pi] for proposition stem:[\phi] - -|stem:[\{x : \tau \mid \phi(x)\}] -|Refinement type: values of type stem:[\tau] satisfying predicate stem:[\phi] -|=== - -== Syntax (Abstract) - -We define the abstract syntax of VCL (concrete syntax in link:vcl-grammar.ebnf[vcl-grammar.ebnf]): - ----- -Query ::= SELECT Modalities FROM Source [WHERE Condition] - [PROOF ProofSpec] [LIMIT Int] [OFFSET Int] - -Modalities ::= * | Modality, ... -Modality ::= GRAPH | VECTOR | TENSOR | SEMANTIC | DOCUMENT | TEMPORAL - -Source ::= HEXAD UUID - | FEDERATION Pattern [WITH DRIFT DriftMode] - | STORE ID - -Condition ::= SimpleCond | CompoundCond | (Condition) -SimpleCond ::= GraphCond | VectorCond | TensorCond | ... -CompoundCond::= Condition AND Condition - | Condition OR Condition - | NOT Condition - -ProofSpec ::= ProofType(Contract) [WITH Params] -ProofType ::= EXISTENCE | CITATION | ACCESS | INTEGRITY | PROVENANCE | CUSTOM ----- - -== Type System (Static Semantics) - -=== Base Types - -VCL has the following base types: - ----- -τ ::= UUID -- 128-bit identifier - | String -- UTF-8 string - | Int -- Integer - | Float -- IEEE 754 double - | Bool -- Boolean - | Vector[n] -- n-dimensional vector - | Tensor[d₁, ..., dₙ] -- n-dimensional tensor - | Timestamp -- ISO 8601 datetime - | RDFTriple -- (subject, predicate, object) - | OctadRef -- Reference to a Octad - | Modality -- GRAPH | VECTOR | TENSOR | ... - | Proof[φ] -- Proof of proposition φ ----- - -=== Modal Type System - -Each modality has an associated type projection: - -[stem] -++++ -\begin{aligned} -\text{Graph}(h) &: \text{Set}(\text{RDFTriple}) \\ -\text{Vector}(h) &: \text{Vector}[n] \\ -\text{Tensor}(h) &: \text{Tensor}[d_1, \ldots, d_n] \\ -\text{Semantic}(h) &: \text{Set}(\text{TypeAnnotation}) \\ -\text{Document}(h) &: \text{String} \\ -\text{Temporal}(h) &: \text{List}(\text{Version}) -\end{aligned} -++++ - -=== Octad Type - -A **Octad** is a dependent product over modalities: - -[stem] -++++ -\text{Octad} = \{h : \text{UUID} \mid \exists m_1, \ldots, m_k \in \text{Modalities}. \text{Valid}(h, m_1) \land \cdots \land \text{Valid}(h, m_k)\} -++++ - -Where stem:[\text{Valid}(h, m)] asserts that octad `h` has a valid representation in modality `m`. - -=== Refinement Types for Queries - -==== Query Result Type - -A query result has a refinement type that depends on the modalities selected: - -[stem] -++++ -\text{QueryResult}[\mathcal{M}] = \text{List}(\text{Octad}_\mathcal{M}) -++++ - -Where stem:[\mathcal{M} \subseteq \{\text{GRAPH}, \text{VECTOR}, \ldots\}] and: - -[stem] -++++ -\text{Octad}_\mathcal{M} = \{h : \text{Octad} \mid \forall m \in \mathcal{M}. \text{Valid}(h, m)\} -++++ - -==== Dependent-Type Query Result - -For queries with `PROOF` clause, the result type includes a proof term: - -[stem] -++++ -\text{ProvedResult}[\mathcal{M}, \phi] = \{(r, \pi) : \text{QueryResult}[\mathcal{M}] \times \text{Proof}[\phi] \mid \pi : \phi(r)\} -++++ - -Where stem:[\phi] is the proposition specified by the `PROOF` clause. - -=== Typing Rules - -==== T-Query (Slipstream Path) - -[stem] -++++ -\frac{ - \Gamma \vdash \text{modalities} : \mathcal{M} \quad - \Gamma \vdash \text{source} : \text{Source} \quad - \Gamma \vdash \text{condition} : \text{Octad} \to \text{Bool} -}{ - \Gamma \vdash \texttt{SELECT } \mathcal{M} \texttt{ FROM } \text{source} \texttt{ WHERE } \text{condition} : \text{QueryResult}[\mathcal{M}] -} -++++ - -**Interpretation:** A slipstream query returns a list of octads with selected modalities, **without proof guarantees**. - -==== T-ProvedQuery (Dependent-Type Path) - -[stem] -++++ -\frac{ - \Gamma \vdash \text{modalities} : \mathcal{M} \quad - \Gamma \vdash \text{source} : \text{Source} \quad - \Gamma \vdash \text{condition} : \text{Octad} \to \text{Bool} \quad - \Gamma \vdash \text{proof-spec} : \text{ProofType} \times \text{Contract} -}{ - \Gamma \vdash \texttt{SELECT } \mathcal{M} \texttt{ ... PROOF } \text{proof-spec} : \text{ProvedResult}[\mathcal{M}, \phi_{\text{proof-spec}}] -} -++++ - -**Interpretation:** A dependent-type query returns results **with proof** that the contract stem:[\phi_{\text{proof-spec}}] holds. - -==== T-Condition (Graph Condition Example) - -[stem] -++++ -\frac{ - \Gamma \vdash \text{pattern} : \text{SPARQLPattern} -}{ - \Gamma \vdash \texttt{WHERE } \text{pattern} : \text{Octad} \to \text{Bool} -} -++++ - -==== T-VectorSimilarity - -[stem] -++++ -\frac{ - \Gamma \vdash v : \text{Vector}[n] \quad - \Gamma \vdash \text{threshold} : \text{Float} -}{ - \Gamma \vdash \texttt{WHERE h.embedding SIMILAR TO } v \texttt{ WITHIN } \text{threshold} : \text{Octad} \to \text{Bool} -} -++++ - -==== T-ProofSpec (EXISTENCE) - -[stem] -++++ -\frac{ - \Gamma \vdash \text{contract} : \text{ExistenceContract} -}{ - \Gamma \vdash \texttt{PROOF EXISTENCE(contract)} : \text{ProofSpec}[\exists h. \text{Valid}(h)] -} -++++ - -==== T-ProofSpec (CITATION) - -[stem] -++++ -\frac{ - \Gamma \vdash \text{contract} : \text{CitationContract} -}{ - \Gamma \vdash \texttt{PROOF CITATION(contract)} : \text{ProofSpec}[\forall h. \text{CitationValid}(h)] -} -++++ - -== Operational Semantics (Dynamic Semantics) - -=== Evaluation Environments - -We define two evaluation environments: - -1. **Store Environment** stem:[\Sigma]: Maps UUIDs to Octad data -2. **Proof Environment** stem:[\Pi]: Maps contracts to proof functions - -[stem] -++++ -\begin{aligned} -\Sigma &: \text{UUID} \rightharpoonup \text{OctadData} \\ -\Pi &: \text{Contract} \rightharpoonup (\text{OctadData} \to \text{Proof}) -\end{aligned} -++++ - -=== Big-Step Semantics (Slipstream Path) - -==== E-SlipstreamQuery - -[stem] -++++ -\frac{ - \llbracket \text{source} \rrbracket_\Sigma = \{h_1, \ldots, h_n\} \quad - \forall i. \llbracket \text{condition}(h_i) \rrbracket_\Sigma = \text{true} \Rightarrow h_i \in R \quad - \forall i. R_i = \text{project}_\mathcal{M}(h_i) -}{ - \llbracket \texttt{SELECT } \mathcal{M} \texttt{ FROM } \text{source} \texttt{ WHERE } \text{condition} \rrbracket_\Sigma = [R_1, \ldots, R_k] -} -++++ - -**Interpretation:** -1. Retrieve candidate octads from source -2. Filter by condition -3. Project to selected modalities -4. **No proof generation** - -==== E-ProjectModality - -[stem] -++++ -\frac{ - \Sigma(h) = \text{octad-data} \quad - \text{modality} \in \mathcal{M} -}{ - \text{project}_\mathcal{M}(h)[\text{modality}] = \text{extract}(\text{octad-data}, \text{modality}) -} -++++ - -=== Big-Step Semantics (Dependent-Type Path) - -==== E-ProvedQuery - -[stem] -++++ -\frac{ - \llbracket \text{source} \rrbracket_\Sigma = \{h_1, \ldots, h_n\} \quad - \forall i. \llbracket \text{condition}(h_i) \rrbracket_\Sigma = \text{true} \Rightarrow h_i \in R \quad - \forall i. R_i = \text{project}_\mathcal{M}(h_i) \quad - \pi = \text{prove}_\Pi(\text{proof-spec}, [R_1, \ldots, R_k]) \quad - \pi : \phi_{\text{proof-spec}}([R_1, \ldots, R_k]) -}{ - \llbracket \texttt{SELECT } \mathcal{M} \texttt{ ... PROOF } \text{proof-spec} \rrbracket_{\Sigma,\Pi} = ([R_1, \ldots, R_k], \pi) -} -++++ - -**Interpretation:** -1. Same filtering as slipstream path -2. **Generate proof** stem:[\pi] that contract holds -3. Return results **with proof term** - -==== E-Prove (EXISTENCE) - -[stem] -++++ -\frac{ - \text{contract} = \text{ExistenceContract} \quad - \forall h \in R. \Sigma(h) \neq \bot -}{ - \text{prove}_\Pi(\texttt{EXISTENCE(contract)}, R) = \pi_{\exists} -} -++++ - -Where stem:[\pi_{\exists}] is a ZKP proving existence without revealing data. - -==== E-Prove (CITATION) - -[stem] -++++ -\frac{ - \text{contract} = \text{CitationContract} \quad - \forall h \in R. \text{validateCitationChain}_\Sigma(h) = \text{true} -}{ - \text{prove}_\Pi(\texttt{CITATION(contract)}, R) = \pi_{\text{cite}} -} -++++ - -Where stem:[\pi_{\text{cite}}] is a ZKP proving citation chain validity. - -=== Small-Step Semantics (Drift Handling) - -For federated queries with drift, we model evaluation as a state machine: - -[stem] -++++ -\langle Q, \Sigma_1 \parallel \cdots \parallel \Sigma_n, \text{DRIFT-MODE} \rangle \rightarrow^* \langle R, \Pi \rangle -++++ - -==== S-DriftDetect - -[stem] -++++ -\frac{ - \Sigma_i(h)[\text{modality}_1] \neq \Sigma_j(h)[\text{modality}_1] \quad - i \neq j -}{ - \langle h, \Sigma_1 \parallel \cdots \parallel \Sigma_n \rangle \xrightarrow{\text{drift}} \langle h, \text{INCONSISTENT} \rangle -} -++++ - -==== S-DriftRepair-LatestWins - -[stem] -++++ -\frac{ - \text{DRIFT-MODE} = \texttt{LATEST} \quad - \Sigma_k(h)[\text{timestamp}] = \max_i \Sigma_i(h)[\text{timestamp}] -}{ - \langle h, \text{INCONSISTENT} \rangle \xrightarrow{\text{repair}} \langle h, \Sigma_k(h) \rangle -} -++++ - -==== S-DriftRepair-Quorum - -[stem] -++++ -\frac{ - \text{DRIFT-MODE} = \texttt{REPAIR} \quad - \exists v. |\{i \mid \Sigma_i(h)[\text{field}] = v\}| > n/2 -}{ - \langle h, \text{INCONSISTENT} \rangle \xrightarrow{\text{quorum}} \langle h, v \rangle -} -++++ - -== Proof Obligations (Dependent-Type Path) - -=== ZKP Contract Specifications - -For each proof type, we specify the **proposition** stem:[\phi] that must be proven: - -==== EXISTENCE Contract - -[stem] -++++ -\phi_{\text{EXISTENCE}}(h) \equiv \exists \text{data} \in \Sigma. \Sigma(h) = \text{data} \land \text{accessible}(h) -++++ - -**Informal:** The octad exists in at least one store and is accessible to the querier. - -==== CITATION Contract - -[stem] -++++ -\phi_{\text{CITATION}}(h) \equiv \forall h' \in \text{citations}(h). \text{ValidCitation}(h, h') \land \phi_{\text{EXISTENCE}}(h') -++++ - -**Informal:** All citations are valid and cited octads exist. - -==== ACCESS Contract - -[stem] -++++ -\phi_{\text{ACCESS}}(h, u) \equiv \text{hasPermission}(u, h, \text{READ}) -++++ - -**Informal:** User `u` has read permission for octad `h`. - -==== INTEGRITY Contract - -[stem] -++++ -\phi_{\text{INTEGRITY}}(h) \equiv \forall t \in \text{versions}(h). \text{Hash}(\text{data}(h, t)) = \text{CommittedHash}(h, t) -++++ - -**Informal:** Data has not been tampered with since commitment. - -==== PROVENANCE Contract - -[stem] -++++ -\phi_{\text{PROVENANCE}}(h) \equiv \exists \text{lineage}. \text{ValidLineage}(h, \text{lineage}) \land \forall h' \in \text{lineage}. \phi_{\text{INTEGRITY}}(h') -++++ - -**Informal:** Full lineage is traceable and all ancestors maintain integrity. - -=== Proof Generation - -For dependent-type queries, VCL delegates to **proven-library** (external ZKP framework): - -[stem] -++++ -\text{prove}_\Pi(\text{proof-spec}, R) = \text{proven.generate}(\phi_{\text{proof-spec}}, R, \text{witness}) -++++ - -Where: -- stem:[\phi_{\text{proof-spec}}] is the proposition from the contract -- `R` is the query result (public input) -- `witness` is private data used for proof construction - -== Soundness Theorem - -=== Slipstream Path - -**Theorem (Slipstream Correctness):** - -If stem:[\Gamma \vdash Q : \text{QueryResult}[\mathcal{M}]] and stem:[\llbracket Q \rrbracket_\Sigma = R], then: - -[stem] -++++ -\forall h \in R. \exists \text{data} \in \Sigma. h = \text{project}_\mathcal{M}(\text{data}) -++++ - -**Informal:** All results are valid projections from the store (but no additional guarantees). - -=== Dependent-Type Path - -**Theorem (Dependent-Type Soundness):** - -If stem:[\Gamma \vdash Q : \text{ProvedResult}[\mathcal{M}, \phi]] and stem:[\llbracket Q \rrbracket_{\Sigma,\Pi} = (R, \pi)], then: - -[stem] -++++ -\pi : \phi(R) \land \text{verify}(\pi, \phi(R), \text{public-params}) = \text{true} -++++ - -**Informal:** The proof term stem:[\pi] is a valid proof of proposition stem:[\phi] applied to results `R`, and it verifies correctly. - -**Proof sketch:** By induction on the derivation of stem:[\Gamma \vdash Q : \text{ProvedResult}[\mathcal{M}, \phi]]. The critical step uses the soundness of the underlying ZKP system (proven-library). - -== Semantic Equivalence - -=== Query Equivalence - -Two queries stem:[Q_1] and stem:[Q_2] are **semantically equivalent** (stem:[Q_1 \equiv Q_2]) if: - -[stem] -++++ -\forall \Sigma. \llbracket Q_1 \rrbracket_\Sigma = \llbracket Q_2 \rrbracket_\Sigma -++++ - -**Example equivalences:** - -1. **Commutativity of AND:** -+ -[stem] -++++ -\texttt{WHERE } C_1 \texttt{ AND } C_2 \equiv \texttt{WHERE } C_2 \texttt{ AND } C_1 -++++ - -2. **Idempotence of modality selection:** -+ -[stem] -++++ -\texttt{SELECT GRAPH, GRAPH} \equiv \texttt{SELECT GRAPH} -++++ - -3. **Filter distribution:** -+ -[stem] -++++ -\texttt{WHERE } C_1 \texttt{ AND } (C_2 \texttt{ OR } C_3) \equiv (\texttt{WHERE } C_1 \texttt{ AND } C_2) \texttt{ OR } (\texttt{WHERE } C_1 \texttt{ AND } C_3) -++++ - -=== Proof Equivalence - -Two proofs stem:[\pi_1] and stem:[\pi_2] are **equivalent** if they prove the same proposition: - -[stem] -++++ -\pi_1 \sim \pi_2 \iff \pi_1 : \phi \land \pi_2 : \phi -++++ - -**Note:** ZKPs may have different witnesses but prove the same statement. - -== Drift Semantics - -=== Consistency Levels - -We define a **consistency lattice** for federated queries: - -[stem] -++++ -\begin{aligned} -\text{STRICT} &\sqsubset \text{LATEST} \\ -\text{STRICT} &\sqsubset \text{REPAIR} \\ -\text{LATEST} &\sqsubset \text{TOLERATE} \\ -\text{REPAIR} &\sqsubset \text{TOLERATE} -\end{aligned} -++++ - -Where stem:[A \sqsubset B] means "A provides stronger guarantees than B". - -=== Drift-Aware Semantics - -For federated queries with drift mode stem:[D]: - -[stem] -++++ -\llbracket Q \rrbracket_{\Sigma_1 \parallel \cdots \parallel \Sigma_n, D} = \text{merge}_D(\llbracket Q \rrbracket_{\Sigma_1}, \ldots, \llbracket Q \rrbracket_{\Sigma_n}) -++++ - -Where stem:[\text{merge}_D] implements the drift policy: - -==== STRICT - -[stem] -++++ -\text{merge}_{\text{STRICT}}(R_1, \ldots, R_n) = -\begin{cases} -R_1 & \text{if } R_1 = \cdots = R_n \\ -\bot & \text{otherwise} -\end{cases} -++++ - -==== LATEST - -[stem] -++++ -\text{merge}_{\text{LATEST}}(R_1, \ldots, R_n) = R_i \text{ where } \text{timestamp}(R_i) = \max_j \text{timestamp}(R_j) -++++ - -==== REPAIR - -[stem] -++++ -\text{merge}_{\text{REPAIR}}(R_1, \ldots, R_n) = \text{quorum}(R_1, \ldots, R_n) -++++ - -==== TOLERATE - -[stem] -++++ -\text{merge}_{\text{TOLERATE}}(R_1, \ldots, R_n) = R_1 \cup \cdots \cup R_n \text{ (with drift annotations)} -++++ - -== Denotational Semantics (Alternative View) - -We can also give a **denotational semantics** by interpreting queries as functions on stores: - -[stem] -++++ -\llbracket \cdot \rrbracket : \text{Query} \to (\text{Store} \to \mathcal{P}(\text{Octad})) -++++ - -Where stem:[\mathcal{P}(\text{Octad})] is the powerset of octads. - -=== Basic Queries - -[stem] -++++ -\llbracket \texttt{SELECT } \mathcal{M} \texttt{ FROM HEXAD } u \rrbracket(\Sigma) = -\begin{cases} -\{\text{project}_\mathcal{M}(\Sigma(u))\} & \text{if } u \in \text{dom}(\Sigma) \\ -\emptyset & \text{otherwise} -\end{cases} -++++ - -=== Filtered Queries - -[stem] -++++ -\llbracket \texttt{SELECT } \mathcal{M} \texttt{ ... WHERE } C \rrbracket(\Sigma) = \{h \in \llbracket \texttt{SELECT } \mathcal{M} \texttt{ ...} \rrbracket(\Sigma) \mid \llbracket C \rrbracket(h) = \text{true}\} -++++ - -=== Compositional Semantics - -Conditions compose via logical operators: - -[stem] -++++ -\begin{aligned} -\llbracket C_1 \texttt{ AND } C_2 \rrbracket(h) &= \llbracket C_1 \rrbracket(h) \land \llbracket C_2 \rrbracket(h) \\ -\llbracket C_1 \texttt{ OR } C_2 \rrbracket(h) &= \llbracket C_1 \rrbracket(h) \lor \llbracket C_2 \rrbracket(h) \\ -\llbracket \texttt{NOT } C \rrbracket(h) &= \neg \llbracket C \rrbracket(h) -\end{aligned} -++++ - -== Future Extensions - -=== Dependent Types with Refinements - -Future versions may support **refinement predicates** in query syntax: - -[source,vcl] ----- -SELECT GRAPH, VECTOR -FROM verisim:semantic -WHERE octad : {h : Octad | citationCount(h) > 10} -LIMIT 100; ----- - -Formal type: - -[stem] -++++ -\{h : \text{Octad} \mid \text{citationCount}(h) > 10\} -++++ - -=== Liquid Types - -Integration with **liquid types** for automatic predicate inference: - -[stem] -++++ -\Gamma \vdash Q : \text{QueryResult}[\mathcal{M}]\{\nu : \text{List}(\text{Octad}) \mid \phi(\nu)\} -++++ - -Where stem:[\phi] is automatically inferred from the query structure. - -== References - -- link:vcl-grammar.ebnf[VCL Grammar (ISO EBNF)] -- link:vcl-type-system.adoc[VCL Type System Specification] -- link:backwards-compatibility.adoc[Backwards Compatibility (VERSION semantics)] -- link:drift-handling.adoc[Drift Handling (DRIFT semantics)] -- Pierce, B. C. (2002). *Types and Programming Languages*. MIT Press. -- Chlipala, A. (2013). *Certified Programming with Dependent Types*. MIT Press. -- https://crypto.stanford.edu/~dabo/cryptobook/[Boneh & Shoup. *A Graduate Course in Applied Cryptography*] (ZKP foundations) diff --git a/verisimdb/docs/vcl-grammar.ebnf b/verisimdb/docs/vcl-grammar.ebnf deleted file mode 100644 index 0142a336..00000000 --- a/verisimdb/docs/vcl-grammar.ebnf +++ /dev/null @@ -1,316 +0,0 @@ -(* SPDX-License-Identifier: MPL-2.0 *) -(* VeriSim Consonance Language (VCL) Grammar *) -(* Format: Extended Backus-Naur Form (EBNF) *) -(* Version: 3.0 — Octad (8 modalities), provenance conditions, spatial conditions *) -(* Date: 2026-02-27 *) - -(* ============================================================================ - 1. TOP-LEVEL STRUCTURE - ============================================================================ *) -statement = query | mutation ; - -query = select_clause, - from_clause, - [where_clause], - [group_by_clause], - [having_clause], - [proof_clause], - [order_by_clause], - [limit_clause], - [offset_clause] ; - -(* ============================================================================ - 2. SELECT CLAUSE - Modality Selection, Column Projections & Aggregates - ============================================================================ *) -select_clause = 'SELECT', select_item_list ; - -select_item_list = select_item, { ',', select_item } ; - -(* A select item can be an aggregate, a column projection, or a full modality *) -select_item = aggregate_expr | field_ref | modality_spec ; - -(* Column projection: MODALITY.field_name *) -field_ref = modality_name, '.', identifier ; - -(* Aggregate functions *) -aggregate_expr = count_all | aggregate_field ; -count_all = 'COUNT', '(', '*', ')' ; -aggregate_field = aggregate_func, '(', field_ref, ')' ; -aggregate_func = 'COUNT' | 'SUM' | 'AVG' | 'MIN' | 'MAX' ; - -(* Bare modality selection *) -modality_spec = 'GRAPH', [graph_projection] | 'VECTOR', [vector_projection] | 'TENSOR', [tensor_projection] | 'SEMANTIC', [semantic_projection] | 'DOCUMENT', [document_projection] | 'TEMPORAL', [temporal_projection] | 'PROVENANCE', [provenance_projection] | 'SPATIAL', [spatial_projection] | '*' (* All available modalities *) ; - -modality_name = 'GRAPH' | 'VECTOR' | 'TENSOR' | 'SEMANTIC' | 'DOCUMENT' | 'TEMPORAL' | 'PROVENANCE' | 'SPATIAL' ; - -(* Projection specifics per modality *) -graph_projection = '(', sparql_pattern, ')' ; -vector_projection = '(', vector_fields, ')' ; -tensor_projection = '(', tensor_slice, ')' ; -semantic_projection = '(', contract_names, ')' ; -document_projection = '(', document_fields, ')' ; -temporal_projection = '(', version_spec, ')' ; -provenance_projection = '(', provenance_fields, ')' ; -spatial_projection = '(', spatial_fields, ')' ; - -(* ============================================================================ - 3. FROM CLAUSE - Data Sources - ============================================================================ *) -from_clause = 'FROM', source_spec ; - -source_spec = hexad_source | federation_source | store_source ; - -(* Direct hexad reference *) -hexad_source = 'HEXAD', uuid ; - -(* Federation pattern (multiple nodes) *) -federation_source = 'FEDERATION', node_pattern, [drift_policy] ; - -node_pattern = glob_pattern (* e.g., '/universities/*' *) | node_list (* e.g., '[node1, node2, node3]' *) ; - -drift_policy = 'WITH', 'DRIFT', drift_mode ; -drift_mode = 'STRICT' (* Fail on any drift *) | 'REPAIR' (* Auto-repair detected drift *) | 'TOLERATE' (* Return data despite drift *) | 'LATEST' (* Use most recent version *) ; - -(* Specific store reference *) -store_source = 'STORE', store_id ; - -(* ============================================================================ - 4. WHERE CLAUSE - Filtering Conditions, - ============================================================================ *) -where_clause = 'WHERE', condition ; - -condition = simple_condition | compound_condition | '(', condition, ')' ; - -simple_condition = graph_condition | vector_condition | tensor_condition | semantic_condition | document_condition | temporal_condition | provenance_condition | spatial_condition | cross_modal_condition ; - -(* 4.9. Cross-Modal Conditions — relationships BETWEEN modalities *) -cross_modal_condition = cross_modal_field_compare | drift_condition | consistency_condition | exists_condition | not_exists_condition ; - -(* Compare fields across modalities: WHERE DOCUMENT.severity > GRAPH.centrality *) -cross_modal_field_compare = field_ref, comparison_op, field_ref ; - -(* Drift between modalities: WHERE DRIFT(VECTOR, DOCUMENT) > 0.3 *) -drift_condition = 'DRIFT', '(', modality_name, ',', modality_name, ')', comparison_op, float ; - -(* Consistency check: WHERE CONSISTENT(VECTOR, SEMANTIC) USING COSINE *) -consistency_condition = 'CONSISTENT', '(', modality_name, ',', modality_name, ')', 'USING', metric_name ; -metric_name = 'COSINE' | 'EUCLIDEAN' | 'DOT_PRODUCT' | 'JACCARD' ; - -(* Modality existence: WHERE VECTOR EXISTS *) -exists_condition = modality_name, 'EXISTS' ; - -(* Modality absence: WHERE TENSOR NOT EXISTS *) -not_exists_condition = modality_name, 'NOT', 'EXISTS' ; - -compound_condition = condition, 'AND', condition | condition, 'OR', condition | 'NOT', condition ; - -(* 4.1. Graph Conditions (SPARQL-like) *) -graph_condition = sparql_pattern | path_pattern ; - -sparql_pattern = '(', node_var, ')', edge_pattern, '(', node_var, ')' ; -edge_pattern = '-[', edge_type, ']->' | '-[', edge_type, ']-' | '<-[', edge_type, ']-' ; -node_var = identifier | ('?', identifier) ; -edge_type = ':', identifier ; - -path_pattern = node_var, path_quantifier, node_var ; -path_quantifier = '-[', edge_type, ('*' | '+' | '{', integer, ',', integer, '}'), ']->' ; - -(* 4.2. Vector Conditions (Similarity Search) *) -vector_condition = vector_field, 'SIMILAR', 'TO', vector_literal, [similarity_threshold] | vector_field, 'NEAREST', integer, [metric_type] ; - -vector_field = identifier, '.', 'embedding' ; -vector_literal = '[', float, { ',', float }, ']' ; -similarity_threshold = 'WITHIN', float ; -metric_type = 'USING', ('COSINE' | 'EUCLIDEAN' | 'DOT_PRODUCT') ; - -(* 4.3. Tensor Conditions (Multi-dimensional) *) -tensor_condition = tensor_field, tensor_op, tensor_literal ; -tensor_op = '==' | '>' | '<' | '>=' | '<=' | 'SHAPE' | 'RANK' ; -tensor_literal = array_literal | scalar_literal ; - -(* 4.4. Semantic Conditions (ZKP Contracts) *) -semantic_condition = 'SATISFIES', contract_name, [contract_params] | 'HAS', 'PROOF', proof_type | 'VERIFIED', 'BY', verifier_id ; - -contract_name = identifier ; -contract_params = '(', param_list, ')' ; -param_list = identifier, '=', literal, { ',', identifier, '=', literal } ; - -(* 4.5. Document Conditions (Full-text Search) *) -document_condition = 'FULLTEXT', 'CONTAINS', string_literal | 'FULLTEXT', 'MATCHES', regex_literal | 'FIELD', identifier, comparison_op, literal ; - -comparison_op = '==' | '!=' | '>' | '<' | '>=' | '<=' | 'LIKE' ; - -(* 4.6. Temporal Conditions (Versioning) *) -temporal_condition = 'AS', 'OF', timestamp | 'BETWEEN', timestamp, 'AND', timestamp | 'VERSION', version_id | 'MODIFIED', 'BY', actor_id ; - -(* 4.7. Provenance Conditions (Lineage Tracking) *) -provenance_condition = provenance_actor | provenance_origin | provenance_chain | provenance_event ; - -(* Filter by actor: WHERE PROVENANCE.actor = 'system' *) -provenance_actor = 'PROVENANCE', '.', 'actor', comparison_op, string_literal ; - -(* Filter by origin: WHERE PROVENANCE.origin = 'import' *) -provenance_origin = 'PROVENANCE', '.', 'origin', comparison_op, string_literal ; - -(* Filter by chain integrity: WHERE PROVENANCE.chain_valid = true *) -provenance_chain = 'PROVENANCE', '.', 'chain_valid', '=', boolean - | 'PROVENANCE', '.', 'chain_length', comparison_op, integer ; - -(* Filter by event type: WHERE PROVENANCE.event_type = 'Modified' *) -provenance_event = 'PROVENANCE', '.', 'event_type', comparison_op, string_literal ; - -(* Provenance-specific fields *) -provenance_fields = identifier, { ',', identifier } ; - -(* 4.8. Spatial Conditions (Geospatial Queries) *) -spatial_condition = spatial_radius | spatial_bounds | spatial_nearest | spatial_field ; - -(* Radius search: WHERE WITHIN RADIUS(51.5074, -0.1278, 100.0) *) -spatial_radius = 'WITHIN', 'RADIUS', '(', float, ',', float, ',', float, ')' ; - -(* Bounding box: WHERE WITHIN BOUNDS(51.0, -1.0, 52.0, 0.5) *) -spatial_bounds = 'WITHIN', 'BOUNDS', '(', float, ',', float, ',', float, ',', float, ')' ; - -(* K-nearest: WHERE NEAREST(51.5074, -0.1278, 10) *) -spatial_nearest = 'NEAREST', '(', float, ',', float, ',', integer, ')' ; - -(* Field access: WHERE SPATIAL.latitude > 51.0 *) -spatial_field = 'SPATIAL', '.', identifier, comparison_op, literal ; - -(* Spatial-specific fields *) -spatial_fields = identifier, { ',', identifier } ; - -(* ============================================================================ - 5. PROOF CLAUSE - Dependent-Type Safety (Multi-Proof Composition) - ============================================================================ *) -proof_clause = 'PROOF', proof_spec_list ; - -proof_spec_list = proof_spec, { 'AND', proof_spec } ; - -proof_spec = proof_type, '(', contract_name, ')', [proof_params] ; - -proof_type = 'EXISTENCE' (* Hexad exists and is accessible *) | 'CITATION' (* Citation chain is valid *) | 'ACCESS' (* User has access rights *) | 'INTEGRITY' (* Data has not been tampered with *) | 'PROVENANCE' (* Lineage is verifiable *) | 'CUSTOM' (* Custom ZKP contract *) ; - -proof_params = 'WITH', param_list ; - -(* ============================================================================ - 6. GROUP BY / HAVING CLAUSES - Aggregation - ============================================================================ *) -group_by_clause = 'GROUP', 'BY', field_ref_list ; -field_ref_list = field_ref, { ',', field_ref } ; - -having_clause = 'HAVING', condition ; - -(* ============================================================================ - 7. ORDER BY CLAUSE - Sorting - ============================================================================ *) -order_by_clause = 'ORDER', 'BY', order_by_list ; -order_by_list = order_by_item, { ',', order_by_item } ; -order_by_item = field_ref, [sort_direction] ; -sort_direction = 'ASC' | 'DESC' ; - -(* ============================================================================ - 8. PAGINATION CLAUSES - ============================================================================ *) -limit_clause = 'LIMIT', integer ; -offset_clause = 'OFFSET', integer ; - -(* ============================================================================ - 9. MUTATIONS - Write Path (INSERT / UPDATE / DELETE) - ============================================================================ *) -mutation = insert_mutation | update_mutation | delete_mutation ; - -(* INSERT: Create a new hexad with modality data *) -insert_mutation = 'INSERT', 'HEXAD', 'WITH', modality_data_list, [proof_clause] ; - -modality_data_list = modality_data, { ',', modality_data } ; - -modality_data = document_data | vector_data | graph_data | tensor_data | semantic_data | temporal_data | provenance_data | spatial_data ; - -document_data = 'DOCUMENT', '(', field_assignment_list, ')' ; -vector_data = 'VECTOR', '(', vector_literal, ')' ; -graph_data = 'GRAPH', '(', identifier, ',', identifier, ')' ; (* edge_type, target_hexad_id *) -tensor_data = 'TENSOR', '(', array_literal, ')' ; -semantic_data = 'SEMANTIC', '(', identifier, ')' ; (* contract name *) -temporal_data = 'TEMPORAL', '(', timestamp, ')' ; -provenance_data = 'PROVENANCE', '(', field_assignment_list, ')' ; (* actor, event_type, description, source *) -spatial_data = 'SPATIAL', '(', field_assignment_list, ')' ; (* latitude, longitude, altitude, geometry_type, srid *) - -field_assignment_list = field_assignment, { ',', field_assignment } ; -field_assignment = identifier, '=', literal ; - -(* UPDATE: Modify fields of an existing hexad *) -update_mutation = 'UPDATE', 'HEXAD', uuid, 'SET', set_list, [proof_clause] ; - -set_list = set_assignment, { ',', set_assignment } ; -set_assignment = field_ref, '=', literal ; - -(* DELETE: Remove a hexad *) -delete_mutation = 'DELETE', 'HEXAD', uuid, [proof_clause] ; - -(* ============================================================================ - 10. LEXICAL ELEMENTS - ============================================================================ *) -uuid = hex_digit, hex_digit, hex_digit, hex_digit, hex_digit, hex_digit, hex_digit, hex_digit, '-', - hex_digit, hex_digit, hex_digit, hex_digit, '-', - hex_digit, hex_digit, hex_digit, hex_digit, '-', - hex_digit, hex_digit, hex_digit, hex_digit, '-', - hex_digit, hex_digit, hex_digit, hex_digit, hex_digit, hex_digit, hex_digit, hex_digit, hex_digit, hex_digit, hex_digit, hex_digit ; -hex_digit = '0' | '1' | '2' | '3' | '4' | '5' | '6' | '7' | '8' | '9' - | 'a' | 'b' | 'c' | 'd' | 'e' | 'f' - | 'A' | 'B' | 'C' | 'D' | 'E' | 'F' ; - -identifier = ? letter or underscore ?, { ? letter, digit, or underscore ? } ; -store_id = identifier ; -verifier_id = identifier ; -actor_id = identifier ; -version_id = identifier ; - -integer = ? digit ?, { ? digit ? } ; -float = ? digit ?, { ? digit ? }, '.', ? digit ?, { ? digit ? }, [('e' | 'E'), ['+' | '-'], ? digit ?, { ? digit ? }] ; - -string_literal = "'", { ? any character except ' and \ ? | "\'", | "\\" }, "'" ; -regex_literal = '/', { ? any character except / and \ ? | '\/' | '\\' }, '/' ; - -timestamp = iso8601_datetime ; -iso8601_datetime = year, '-', month, '-', day, 'T', hour, ':', minute, ':', second, ['.', fraction], timezone ; - -glob_pattern = '/', path_segment, { '/', path_segment }, ('/*' | '/**') ; -path_segment = ? alphanumeric, underscore, or hyphen ?, { ? alphanumeric, underscore, or hyphen ? } ; - -array_literal = '[', literal, { ',', literal }, ']' ; -scalar_literal = integer | float | string_literal | boolean ; -boolean = 'true' | 'false' ; -literal = scalar_literal | array_literal ; - -(* ============================================================================ - 11. COMMENTS - ============================================================================ *) -comment = '--', { ? any character except newline ? }, '\n' (* Line comment *) - | '/*', ? any characters ?, '*/' (* Block comment *) ; - -(* ============================================================================ - 12. RESERVED KEYWORDS - ============================================================================ *) -(* Keywords are case-insensitive in VCL *) -keywords = 'SELECT' | 'FROM' | 'WHERE' | 'PROOF' | 'LIMIT' | 'OFFSET' - | 'GRAPH' | 'VECTOR' | 'TENSOR' | 'SEMANTIC' | 'DOCUMENT' | 'TEMPORAL' | 'PROVENANCE' | 'SPATIAL' - | 'HEXAD' | 'FEDERATION' | 'STORE' - | 'WITH' | 'DRIFT' | 'STRICT' | 'REPAIR' | 'TOLERATE' | 'LATEST' - | 'AND' | 'OR' | 'NOT' - | 'SIMILAR' | 'TO' | 'WITHIN' | 'NEAREST' | 'USING' - | 'SATISFIES' | 'HAS' | 'VERIFIED' | 'BY' - | 'FULLTEXT' | 'CONTAINS' | 'MATCHES' | 'FIELD' | 'LIKE' - | 'AS' | 'OF' | 'BETWEEN' | 'VERSION' | 'MODIFIED' - | 'EXISTENCE' | 'CITATION' | 'ACCESS' | 'INTEGRITY' | 'CUSTOM' - | 'COSINE' | 'EUCLIDEAN' | 'DOT_PRODUCT' | 'JACCARD' | 'SHAPE' | 'RANK' - | 'RADIUS' | 'BOUNDS' - | 'ORDER' | 'GROUP' | 'HAVING' | 'ASC' | 'DESC' - | 'COUNT' | 'SUM' | 'AVG' | 'MIN' | 'MAX' - | 'INSERT' | 'UPDATE' | 'DELETE' | 'SET' - | 'EXISTS' | 'CONSISTENT' - | 'true' | 'false' ; - -(* ============================================================================ - END OF GRAMMAR - ============================================================================ *) diff --git a/verisimdb/docs/vcl-type-system.adoc b/verisimdb/docs/vcl-type-system.adoc deleted file mode 100644 index 4b4a2a43..00000000 --- a/verisimdb/docs/vcl-type-system.adoc +++ /dev/null @@ -1,922 +0,0 @@ -// SPDX-License-Identifier: CC-BY-SA-4.0 - -= VCL Type System Specification -:toc: left -:toclevels: 4 -:sectnums: -:stem: latexmath - -== Overview - -This document specifies the **static type system** for VCL, including: - -1. Type syntax and semantics -2. Typing rules for all VCL constructs -3. Type checking algorithm -4. Subtyping and type equivalence -5. Differences between dependent-type and slipstream paths - -This complements link:vcl-formal-semantics.adoc[VCL Formal Semantics] which covers operational semantics. - -== Type Language - -=== Type Syntax - -The VCL type language stem:[\tau] is defined by: - -[stem] -++++ -\begin{aligned} -\tau ::= &\ \text{UUID} \mid \text{String} \mid \text{Int} \mid \text{Float} \mid \text{Bool} \\ - | &\ \text{Vector}[n] \mid \text{Tensor}[d_1, \ldots, d_k] \\ - | &\ \text{Timestamp} \mid \text{RDFTriple} \\ - | &\ \text{Octad} \mid \text{OctadRef} \\ - | &\ \text{Modality} \mid \text{ModalitySet} \\ - | &\ \text{List}(\tau) \mid \text{Option}(\tau) \\ - | &\ \tau_1 \times \tau_2 \quad \text{(product)} \\ - | &\ \tau_1 \to \tau_2 \quad \text{(function)} \\ - | &\ \{x : \tau \mid \phi(x)\} \quad \text{(refinement)} \\ - | &\ \Pi x : \tau_1. \tau_2(x) \quad \text{(dependent product)} \\ - | &\ \Sigma x : \tau_1. \tau_2(x) \quad \text{(dependent sum)} \\ - | &\ \text{Proof}[\phi] \quad \text{(proof type)} -\end{aligned} -++++ - -=== Type Constructors - -==== Product Types - -Represent pairs/tuples: - -[stem] -++++ -\tau_1 \times \tau_2 = \{(v_1, v_2) \mid v_1 : \tau_1 \land v_2 : \tau_2\} -++++ - -**Example:** `(UUID, Timestamp)` - a octad ID paired with a timestamp - -==== Function Types - -[stem] -++++ -\tau_1 \to \tau_2 = \{f \mid \forall v : \tau_1. f(v) : \tau_2\} -++++ - -**Example:** `Octad → Bool` - predicate on octads (used in WHERE clauses) - -==== List Types - -[stem] -++++ -\text{List}(\tau) = \{[], [v_1], [v_1, v_2], \ldots \mid v_i : \tau\} -++++ - -**Example:** `List(Octad)` - query results - -==== Option Types - -[stem] -++++ -\text{Option}(\tau) = \{\text{None}\} \cup \{\text{Some}(v) \mid v : \tau\} -++++ - -**Example:** `Option(Vector[n])` - optional embedding - -==== Refinement Types - -[stem] -++++ -\{x : \tau \mid \phi(x)\} = \{v : \tau \mid \phi(v) = \text{true}\} -++++ - -**Example:** `{n : Int | n > 0}` - positive integers - -==== Dependent Product (Pi Types) - -[stem] -++++ -\Pi x : \tau_1. \tau_2(x) = \{f \mid \forall v : \tau_1. f(v) : \tau_2(v)\} -++++ - -**Example:** Vector dimension-indexed types: - -[stem] -++++ -\Pi n : \text{Nat}. \text{Vector}[n] \to \text{Float} -++++ - -A function that takes vectors of any dimension and returns a float (e.g., norm). - -==== Dependent Sum (Sigma Types) - -[stem] -++++ -\Sigma x : \tau_1. \tau_2(x) = \{(v, w) \mid v : \tau_1 \land w : \tau_2(v)\} -++++ - -**Example:** Dimensioned vectors: - -[stem] -++++ -\Sigma n : \text{Nat}. \text{Vector}[n] -++++ - -A pair of a dimension `n` and a vector of that dimension. - -=== Modality Types - -Each modality has an associated type: - -[cols="1,2"] -|=== -|Modality |Type Signature - -|`GRAPH` -|stem:[\text{Octad} \to \text{Set}(\text{RDFTriple})] - -|`VECTOR` -|stem:[\Pi n : \text{Nat}. \text{Octad} \to \text{Option}(\text{Vector}[n])] - -|`TENSOR` -|stem:[\Pi d_1, \ldots, d_k : \text{Nat}. \text{Octad} \to \text{Option}(\text{Tensor}[d_1, \ldots, d_k])] - -|`SEMANTIC` -|stem:[\text{Octad} \to \text{Set}(\text{TypeAnnotation})] - -|`DOCUMENT` -|stem:[\text{Octad} \to \text{Option}(\text{String})] - -|`TEMPORAL` -|stem:[\text{Octad} \to \text{List}(\text{Version})] -|=== - -**Note:** Vector and tensor modalities have **dependent types** - the return type depends on the dimension parameters. - -== Type Environments - -=== Context (Type Environment) - -A **type environment** stem:[\Gamma] is a finite map from variables to types: - -[stem] -++++ -\Gamma ::= \emptyset \mid \Gamma, x : \tau -++++ - -**Operations:** - -- stem:[\Gamma(x)] - lookup type of `x` in stem:[\Gamma] -- stem:[\Gamma, x : \tau] - extend stem:[\Gamma] with binding `x : τ` -- stem:[\text{dom}(\Gamma)] - domain of stem:[\Gamma] (set of variables) - -=== Well-Formed Environments - -[stem] -++++ -\frac{ - \Gamma \vdash \tau : \text{Type} \quad - x \notin \text{dom}(\Gamma) -}{ - \Gamma, x : \tau \text{ well-formed} -} -++++ - -== Typing Judgments - -We use the following typing judgments: - -[cols="1,3"] -|=== -|Judgment |Meaning - -|stem:[\Gamma \vdash e : \tau] -|Expression `e` has type stem:[\tau] in context stem:[\Gamma] - -|stem:[\Gamma \vdash Q : \tau] -|Query `Q` has type stem:[\tau] in context stem:[\Gamma] - -|stem:[\Gamma \vdash C : \text{Octad} \to \text{Bool}] -|Condition `C` is a predicate on octads - -|stem:[\Gamma \vdash \tau : \text{Type}] -|stem:[\tau] is a well-formed type - -|stem:[\Gamma \vdash \tau_1 \leq \tau_2] -|stem:[\tau_1] is a subtype of stem:[\tau_2] -|=== - -== Core Typing Rules - -=== Variables - -[stem] -++++ -\frac{ - x : \tau \in \Gamma -}{ - \Gamma \vdash x : \tau -} \text{(T-Var)} -++++ - -=== Literals - -[stem] -++++ -\frac{ - n \in \mathbb{Z} -}{ - \Gamma \vdash n : \text{Int} -} \text{(T-Int)} -\quad -\frac{ - f \in \mathbb{R} -}{ - \Gamma \vdash f : \text{Float} -} \text{(T-Float)} -++++ - -[stem] -++++ -\frac{ - s \in \text{UTF-8} -}{ - \Gamma \vdash s : \text{String} -} \text{(T-String)} -\quad -\frac{ - b \in \{\text{true}, \text{false}\} -}{ - \Gamma \vdash b : \text{Bool} -} \text{(T-Bool)} -++++ - -=== UUID Literals - -[stem] -++++ -\frac{ - u \text{ matches UUID format} -}{ - \Gamma \vdash u : \text{UUID} -} \text{(T-UUID)} -++++ - -=== Vector Literals - -[stem] -++++ -\frac{ - \Gamma \vdash v_1 : \text{Float} \quad \cdots \quad \Gamma \vdash v_n : \text{Float} -}{ - \Gamma \vdash [v_1, \ldots, v_n] : \text{Vector}[n] -} \text{(T-VectorLit)} -++++ - -=== List Construction - -[stem] -++++ -\frac{ - \Gamma \vdash v_1 : \tau \quad \cdots \quad \Gamma \vdash v_n : \tau -}{ - \Gamma \vdash [v_1, \ldots, v_n] : \text{List}(\tau) -} \text{(T-ListLit)} -++++ - -== Query Typing Rules - -=== Slipstream Query (No Proof) - -[stem] -++++ -\frac{ - \Gamma \vdash \mathcal{M} : \text{ModalitySet} \quad - \Gamma \vdash S : \text{Source} \quad - \Gamma \vdash C : \text{Octad} \to \text{Bool} -}{ - \Gamma \vdash \texttt{SELECT } \mathcal{M} \texttt{ FROM } S \texttt{ WHERE } C : \text{QueryResult}[\mathcal{M}] -} \text{(T-SlipstreamQuery)} -++++ - -Where: - -[stem] -++++ -\text{QueryResult}[\mathcal{M}] = \text{List}(\text{Octad}_\mathcal{M}) -++++ - -And: - -[stem] -++++ -\text{Octad}_\mathcal{M} = \{h : \text{Octad} \mid \forall m \in \mathcal{M}. m(h) \neq \text{None}\} -++++ - -=== Dependent-Type Query (With Proof) - -[stem] -++++ -\frac{ - \Gamma \vdash \mathcal{M} : \text{ModalitySet} \quad - \Gamma \vdash S : \text{Source} \quad - \Gamma \vdash C : \text{Octad} \to \text{Bool} \quad - \Gamma \vdash P : \text{ProofSpec}[\phi] -}{ - \Gamma \vdash \texttt{SELECT } \mathcal{M} \texttt{ ... PROOF } P : \text{ProvedResult}[\mathcal{M}, \phi] -} \text{(T-ProvedQuery)} -++++ - -Where: - -[stem] -++++ -\text{ProvedResult}[\mathcal{M}, \phi] = \Sigma r : \text{QueryResult}[\mathcal{M}]. \text{Proof}[\phi(r)] -++++ - -**Interpretation:** A dependent sum - a pair of query results `r` and a proof that stem:[\phi(r)] holds. - -=== Modality Set Typing - -[stem] -++++ -\frac{ - \forall m \in \{m_1, \ldots, m_k\}. m \in \{\texttt{GRAPH}, \texttt{VECTOR}, \ldots\} -}{ - \Gamma \vdash \{m_1, \ldots, m_k\} : \text{ModalitySet} -} \text{(T-ModalitySet)} -++++ - -[stem] -++++ -\frac{ -}{ - \Gamma \vdash * : \text{ModalitySet} -} \text{(T-AllModalities)} -++++ - -=== Source Typing - -==== Octad Source - -[stem] -++++ -\frac{ - \Gamma \vdash u : \text{UUID} -}{ - \Gamma \vdash \texttt{HEXAD } u : \text{Source} -} \text{(T-OctadSource)} -++++ - -==== Federation Source - -[stem] -++++ -\frac{ - \Gamma \vdash p : \text{Pattern} \quad - \Gamma \vdash d : \text{DriftMode} -}{ - \Gamma \vdash \texttt{FEDERATION } p \texttt{ WITH DRIFT } d : \text{Source} -} \text{(T-FederationSource)} -++++ - -==== Store Source - -[stem] -++++ -\frac{ - \Gamma \vdash \text{id} : \text{String} -}{ - \Gamma \vdash \texttt{STORE } \text{id} : \text{Source} -} \text{(T-StoreSource)} -++++ - -== Condition Typing Rules - -=== Simple Conditions - -==== Graph Condition (SPARQL Pattern) - -[stem] -++++ -\frac{ - \Gamma \vdash p : \text{SPARQLPattern} -}{ - \Gamma \vdash p : \text{Octad} \to \text{Bool} -} \text{(T-GraphCond)} -++++ - -==== Vector Similarity - -[stem] -++++ -\frac{ - \Gamma \vdash v : \text{Vector}[n] \quad - \Gamma \vdash t : \text{Float} -}{ - \Gamma \vdash \texttt{h.embedding SIMILAR TO } v \texttt{ WITHIN } t : \text{Octad} \to \text{Bool} -} \text{(T-VectorSimilarity)} -++++ - -==== Vector Nearest-K - -[stem] -++++ -\frac{ - \Gamma \vdash k : \text{Int} \quad - k > 0 \quad - \Gamma \vdash m : \text{MetricType} -}{ - \Gamma \vdash \texttt{h.embedding NEAREST } k \texttt{ USING } m : \text{Octad} \to \text{Bool} -} \text{(T-VectorNearest)} -++++ - -==== Tensor Condition - -[stem] -++++ -\frac{ - \Gamma \vdash \text{field} : \text{Octad} \to \text{Tensor}[\ldots] \quad - \Gamma \vdash \text{op} : \text{TensorOp} \quad - \Gamma \vdash \text{literal} : \text{Tensor}[\ldots] -}{ - \Gamma \vdash \text{field op literal} : \text{Octad} \to \text{Bool} -} \text{(T-TensorCond)} -++++ - -==== Semantic Condition (Contract Satisfaction) - -[stem] -++++ -\frac{ - \Gamma \vdash c : \text{Contract} \quad - \Gamma \vdash p : \text{Params} -}{ - \Gamma \vdash \texttt{SATISFIES } c(p) : \text{Octad} \to \text{Bool} -} \text{(T-SemanticCond)} -++++ - -==== Temporal Condition (As Of) - -[stem] -++++ -\frac{ - \Gamma \vdash t : \text{Timestamp} -}{ - \Gamma \vdash \texttt{AS OF } t : \text{Octad} \to \text{Bool} -} \text{(T-TemporalAsOf)} -++++ - -=== Compound Conditions - -==== Conjunction - -[stem] -++++ -\frac{ - \Gamma \vdash C_1 : \text{Octad} \to \text{Bool} \quad - \Gamma \vdash C_2 : \text{Octad} \to \text{Bool} -}{ - \Gamma \vdash C_1 \texttt{ AND } C_2 : \text{Octad} \to \text{Bool} -} \text{(T-And)} -++++ - -==== Disjunction - -[stem] -++++ -\frac{ - \Gamma \vdash C_1 : \text{Octad} \to \text{Bool} \quad - \Gamma \vdash C_2 : \text{Octad} \to \text{Bool} -}{ - \Gamma \vdash C_1 \texttt{ OR } C_2 : \text{Octad} \to \text{Bool} -} \text{(T-Or)} -++++ - -==== Negation - -[stem] -++++ -\frac{ - \Gamma \vdash C : \text{Octad} \to \text{Bool} -}{ - \Gamma \vdash \texttt{NOT } C : \text{Octad} \to \text{Bool} -} \text{(T-Not)} -++++ - -== Proof Specification Typing - -=== Existence Proof - -[stem] -++++ -\frac{ - \Gamma \vdash c : \text{ExistenceContract} -}{ - \Gamma \vdash \texttt{PROOF EXISTENCE}(c) : \text{ProofSpec}[\exists h. \text{Valid}(h)] -} \text{(T-ProofExistence)} -++++ - -=== Citation Proof - -[stem] -++++ -\frac{ - \Gamma \vdash c : \text{CitationContract} -}{ - \Gamma \vdash \texttt{PROOF CITATION}(c) : \text{ProofSpec}[\forall h. \text{CitationValid}(h)] -} \text{(T-ProofCitation)} -++++ - -=== Access Proof - -[stem] -++++ -\frac{ - \Gamma \vdash c : \text{AccessContract} \quad - \Gamma \vdash u : \text{User} -}{ - \Gamma \vdash \texttt{PROOF ACCESS}(c) : \text{ProofSpec}[\forall h. \text{hasPermission}(u, h)] -} \text{(T-ProofAccess)} -++++ - -=== Integrity Proof - -[stem] -++++ -\frac{ - \Gamma \vdash c : \text{IntegrityContract} -}{ - \Gamma \vdash \texttt{PROOF INTEGRITY}(c) : \text{ProofSpec}[\forall h. \text{Untampered}(h)] -} \text{(T-ProofIntegrity)} -++++ - -=== Provenance Proof - -[stem] -++++ -\frac{ - \Gamma \vdash c : \text{ProvenanceContract} -}{ - \Gamma \vdash \texttt{PROOF PROVENANCE}(c) : \text{ProofSpec}[\forall h. \text{ValidLineage}(h)] -} \text{(T-ProofProvenance)} -++++ - -== Subtyping - -=== Subtyping Relation - -stem:[\tau_1 \leq \tau_2] means "stem:[\tau_1] is a subtype of stem:[\tau_2]" (values of type stem:[\tau_1] can be used where stem:[\tau_2] is expected). - -==== Reflexivity - -[stem] -++++ -\frac{ -}{ - \tau \leq \tau -} \text{(Sub-Refl)} -++++ - -==== Transitivity - -[stem] -++++ -\frac{ - \tau_1 \leq \tau_2 \quad - \tau_2 \leq \tau_3 -}{ - \tau_1 \leq \tau_3 -} \text{(Sub-Trans)} -++++ - -==== Refinement Subsumption - -[stem] -++++ -\frac{ - \forall v. \phi_1(v) \Rightarrow \phi_2(v) -}{ - \{x : \tau \mid \phi_1(x)\} \leq \{x : \tau \mid \phi_2(x)\} -} \text{(Sub-Refine)} -++++ - -**Intuition:** If stem:[\phi_1] is stronger (more restrictive) than stem:[\phi_2], then the refined type with stem:[\phi_1] is a subtype. - -==== List Covariance - -[stem] -++++ -\frac{ - \tau_1 \leq \tau_2 -}{ - \text{List}(\tau_1) \leq \text{List}(\tau_2) -} \text{(Sub-List)} -++++ - -==== Function Contravariance (Input) and Covariance (Output) - -[stem] -++++ -\frac{ - \tau_1' \leq \tau_1 \quad - \tau_2 \leq \tau_2' -}{ - \tau_1 \to \tau_2 \leq \tau_1' \to \tau_2' -} \text{(Sub-Arrow)} -++++ - -**Note:** Input is contravariant, output is covariant. - -==== Modality Subsumption - -[stem] -++++ -\frac{ - \mathcal{M}_1 \subseteq \mathcal{M}_2 -}{ - \text{Octad}_{\mathcal{M}_2} \leq \text{Octad}_{\mathcal{M}_1} -} \text{(Sub-Octad)} -++++ - -**Intuition:** A octad with more modalities can be used where fewer are required (contravariant). - -=== Subsumption in Typing - -[stem] -++++ -\frac{ - \Gamma \vdash e : \tau_1 \quad - \tau_1 \leq \tau_2 -}{ - \Gamma \vdash e : \tau_2 -} \text{(T-Sub)} -++++ - -== Type Equivalence - -=== Definitional Equality - -Two types stem:[\tau_1] and stem:[\tau_2] are **definitionally equal** (stem:[\tau_1 \equiv \tau_2]) if they are syntactically identical up to alpha-renaming. - -**Examples:** - -1. stem:[\text{Vector}[10] \equiv \text{Vector}[10]] -2. stem:[\Pi x : \tau. \sigma \equiv \Pi y : \tau. \sigma[y/x]] - -=== Semantic Equality - -Two types are **semantically equal** if they denote the same set of values: - -[stem] -++++ -\tau_1 \simeq \tau_2 \iff \forall v. (v : \tau_1 \iff v : \tau_2) -++++ - -**Example:** - -[stem] -++++ -\{x : \text{Int} \mid x \geq 0\} \simeq \{x : \text{Int} \mid x > -1\} -++++ - -== Type Checking Algorithm - -=== Bidirectional Type Checking - -VCL uses **bidirectional type checking** with two modes: - -1. **Synthesis (stem:[\Uparrow])**: Infer type of an expression -2. **Checking (stem:[\Downarrow])**: Check expression against expected type - -==== Synthesis Rules - -[stem] -++++ -\Gamma \vdash e \Uparrow \tau -++++ - -"In context stem:[\Gamma], synthesize that expression `e` has type stem:[\tau]" - -**Examples:** - -[stem] -++++ -\frac{ - x : \tau \in \Gamma -}{ - \Gamma \vdash x \Uparrow \tau -} \text{(Synth-Var)} -++++ - -[stem] -++++ -\frac{ - n \in \mathbb{Z} -}{ - \Gamma \vdash n \Uparrow \text{Int} -} \text{(Synth-Int)} -++++ - -==== Checking Rules - -[stem] -++++ -\Gamma \vdash e \Downarrow \tau -++++ - -"In context stem:[\Gamma], check that expression `e` has type stem:[\tau]" - -**Example:** - -[stem] -++++ -\frac{ - \Gamma \vdash e \Uparrow \tau' \quad - \tau' \leq \tau -}{ - \Gamma \vdash e \Downarrow \tau -} \text{(Check-Sub)} -++++ - -=== Type Checking Queries - -==== Slipstream Query - -**Input:** Query `Q`, context stem:[\Gamma] -**Output:** Type stem:[\tau] or type error - -[source,haskell] ----- -check_slipstream_query Γ (SELECT M FROM S WHERE C) = - let τ_M = check_modalities Γ M in - let τ_S = check_source Γ S in - let τ_C = check_condition Γ C in - if τ_C ≡ (Octad → Bool) then - QueryResult[M] - else - TypeError "Condition must be predicate on Octad" ----- - -==== Dependent-Type Query - -**Input:** Query `Q` with `PROOF` clause, context stem:[\Gamma] -**Output:** Proved result type or type error - -[source,haskell] ----- -check_proved_query Γ (SELECT M ... PROOF P) = - let τ_base = check_slipstream_query Γ (SELECT M ...) in - let (ProofSpec[φ]) = check_proof_spec Γ P in - ProvedResult[M, φ] ----- - -=== Type Inference - -Some VCL constructs support **type inference**: - -==== Vector Dimension Inference - -From vector literal: - -[source,vcl] ----- -WHERE h.embedding SIMILAR TO [0.1, 0.2, 0.3] ----- - -Inferred: `Vector[3]` - -==== Refinement Predicate Inference - -From conditions: - -[source,vcl] ----- -WHERE h.embedding SIMILAR TO v WITHIN 0.8 ----- - -Inferred: `{h : Octad | similarity(h.embedding, v) ≥ 0.8}` - -== Type Safety Properties - -=== Progress - -**Theorem (Progress):** - -If stem:[\Gamma \vdash Q : \tau] and `Q` is a closed query (no free variables), then either: - -1. `Q` is a value, or -2. stem:[\exists Q'. Q \rightarrow Q'] (Q can take a step) - -=== Preservation - -**Theorem (Preservation):** - -If stem:[\Gamma \vdash Q : \tau] and stem:[Q \rightarrow Q'], then stem:[\Gamma \vdash Q' : \tau]. - -**Informal:** Type is preserved during evaluation. - -=== Soundness - -**Theorem (Type Soundness):** - -If stem:[\Gamma \vdash Q : \tau] and stem:[Q \Rightarrow^* v], then stem:[\Gamma \vdash v : \tau]. - -**Informal:** Well-typed queries don't "go wrong" - they either diverge or produce a value of the expected type. - -== Dependent Types vs Slipstream Paths - -=== Type Differences - -[cols="2,3,3"] -|=== -|Aspect |Slipstream Path |Dependent-Type Path - -|**Result Type** -|stem:[\text{List}(\text{Octad})] -|stem:[\Sigma r : \text{List}(\text{Octad}). \text{Proof}[\phi(r)]] - -|**Guarantees** -|Best-effort filtering -|Formally verified via ZKP - -|**Performance** -|Fast (no proof overhead) -|Slower (proof generation) - -|**Use Case** -|Exploratory queries -|Compliance, audits - -|**Type Complexity** -|Simple (no dependent types) -|Complex (refinements, dependent products) -|=== - -=== Type Erasure - -For runtime optimization, **dependent types can be erased** from slipstream queries: - -[stem] -++++ -\text{erase}(\Pi x : \tau_1. \tau_2) = \text{erase}(\tau_1) \to \text{erase}(\tau_2) -++++ - -[stem] -++++ -\text{erase}(\Sigma x : \tau_1. \tau_2) = \text{erase}(\tau_1) \times \text{erase}(\tau_2) -++++ - -[stem] -++++ -\text{erase}(\{x : \tau \mid \phi\}) = \text{erase}(\tau) -++++ - -[stem] -++++ -\text{erase}(\text{Proof}[\phi]) = \text{Unit} -++++ - -**Dependent-type path** performs full type checking with refinements. -**Slipstream path** can use erased types for performance. - -== Type System Extensions - -=== Row Polymorphism (Future) - -Support for flexible octad projections: - -[stem] -++++ -\text{Octad}\{\text{graph} : \text{Graph}, \text{vector} : \text{Vector} \mid r\} -++++ - -Where `r` is a row variable representing "possibly more modalities". - -=== Effect Types (Future) - -Track side effects in query execution: - -[stem] -++++ -\text{Query} : \text{OctadSet} \xrightarrow{\{\text{IO}, \text{Network}\}} \text{List}(\text{Octad}) -++++ - -Indicating the query performs IO and network operations. - -=== Linear Types (Future) - -Ensure ZKP witnesses are used exactly once: - -[stem] -++++ -\text{Witness} : \text{LinearType} -++++ - -== References - -- link:vcl-formal-semantics.adoc[VCL Formal Semantics] -- link:vcl-grammar.ebnf[VCL Grammar] -- Pierce, B. C. (2002). *Types and Programming Languages*. MIT Press. -- Pierce, B. C. (ed.) (2005). *Advanced Topics in Types and Programming Languages*. MIT Press. -- Chlipala, A. (2013). *Certified Programming with Dependent Types*. MIT Press. -- Norell, U. (2007). *Towards a practical programming language based on dependent type theory*. PhD thesis, Chalmers. diff --git a/verisimdb/docs/vcl-vs-sql.adoc b/verisimdb/docs/vcl-vs-sql.adoc deleted file mode 100644 index 22df1e7b..00000000 --- a/verisimdb/docs/vcl-vs-sql.adoc +++ /dev/null @@ -1,240 +0,0 @@ -// SPDX-License-Identifier: CC-BY-SA-4.0 -// Copyright (c) 2026 Jonathan D.A. Jewell (hyperpolymath) - -= VCL vs SQL: A Comparative Guide -:toc: left -:toclevels: 3 -:sectnums: - -== Introduction - -VCL (VeriSim Consonance Language) is the query language for VeriSimDB, a 6-core multimodal database. While VCL borrows familiar keywords from SQL (`SELECT`, `FROM`, `WHERE`, `LIMIT`, `OFFSET`), it serves a fundamentally different purpose. - -VCL is **read-only**. It does not support `INSERT`, `UPDATE`, or `DELETE` statements. All mutations go through the Octad API (Rust core or Elixir orchestration layer). This is a deliberate design choice: VeriSimDB treats writes as coordinated multi-modal operations that must maintain consistency across all six modality stores (Graph, Vector, Tensor, Semantic, Document, Temporal). A simple `INSERT` statement cannot express the cross-modal invariants that a Octad write requires. - -VCL is **multimodal**. Where SQL selects columns from relational tables, VCL selects _modalities_ from _octad stores_ or _federations_. A single VCL query can retrieve graph edges, vector embeddings, tensor slices, semantic annotations, document text, and temporal versions in one pass. - -VCL is **federation-aware**. Queries can target a local `STORE`, a specific `HEXAD` entity, or a `FEDERATION` of distributed VeriSimDB instances, with configurable drift policies governing cross-instance consistency. - -== Concept Comparison - -[cols="2,3,3",options="header"] -|=== -|Concept |VCL |SQL - -|**Column/Field Selection** -|`SELECT` with _modalities_: `GRAPH`, `VECTOR`, `TENSOR`, `SEMANTIC`, `DOCUMENT`, `TEMPORAL`, or `*` (all six). Modalities can be combined: `SELECT GRAPH, VECTOR, SEMANTIC`. -|`SELECT` with _column names_ or `*` for all columns. - -|**Data Source** -|`FROM STORE `, `FROM HEXAD `, or `FROM FEDERATION `. A STORE is a collection of octad entities. A HEXAD is a single entity addressed by UUID. A FEDERATION spans multiple VeriSimDB instances. -|`FROM ` or `FROM
AS `. Tables are flat relational structures. - -|**Filtering** -|`WHERE` supports field conditions (`h.field = value`) and full-text predicates: `CONTAINS(h.content, "term")` for substring matching, `MATCHES(h.content, "pattern")` for pattern matching. Graph traversal uses Cypher-inspired syntax: `(h)-[:EDGE_TYPE]->(target)`. Vector similarity uses `SIMILAR TO [...]`. -|`WHERE` supports standard comparison operators, `LIKE`, `IN`, `BETWEEN`, `IS NULL`, boolean combinators (`AND`, `OR`, `NOT`), and subqueries. - -|**Joins** -|**No JOINs.** Graph modality replaces relational joins entirely. Relationships are first-class: `(h)-[:CITES]->(target)` traverses edges without explicit join syntax. Cross-modal correlation is implicit within a octad. -|`JOIN`, `LEFT JOIN`, `RIGHT JOIN`, `FULL OUTER JOIN`, `CROSS JOIN` with `ON` conditions. - -|**Mutations** -|**None.** VCL is strictly read-only. All writes go through the Octad API: `POST /api/octads` (create), `PUT /api/octads/:id` (update), `DELETE /api/octads/:id` (delete). This ensures cross-modal consistency. -|`INSERT INTO`, `UPDATE ... SET`, `DELETE FROM`, `MERGE`, `TRUNCATE`. - -|**Verification** -|`PROOF ()` clause requests a verifiable guarantee about the result. Six proof types: `EXISTENCE`, `INTEGRITY`, `CONSISTENCY`, `PROVENANCE`, `FRESHNESS`, `AUTHORIZATION`. No SQL equivalent exists. -|No equivalent. SQL has no mechanism for cryptographic or type-theoretic verification of query results. - -|**Pagination** -|`LIMIT n` and `OFFSET n` work identically to SQL. -|`LIMIT n` and `OFFSET n` (or vendor-specific: `TOP`, `FETCH FIRST`). -|=== - -== What VCL Lacks That SQL Has - -VCL is intentionally minimal. The following SQL features have no VCL equivalent: - -[cols="2,3",options="header"] -|=== -|SQL Feature |VCL Status - -|`GROUP BY` -|Not supported. Aggregation is performed application-side or through the Elixir orchestration layer. - -|`ORDER BY` -|Not supported. Result ordering is determined by the modality store (e.g., vector similarity ranking, graph traversal order, temporal chronological order). - -|Aggregate functions (`COUNT`, `SUM`, `AVG`, `MIN`, `MAX`) -|Not supported. VCL returns raw data; aggregation is a consumer responsibility. - -|Subqueries -|Not supported. Queries are flat. Compose results in application code or use the Elixir query router for multi-step workflows. - -|`DISTINCT` -|Not supported. Octad entities are inherently unique (UUID-addressed), so deduplication is rarely needed. - -|`HAVING` -|Not supported (requires `GROUP BY`). - -|`UNION` / `INTERSECT` / `EXCEPT` -|Not supported. Use multiple queries and combine results application-side. - -|`CREATE TABLE` / DDL -|Not supported. Schema is managed through the ReScript registry and Elixir `SchemaRegistry`. - -|Window functions (`ROW_NUMBER`, `RANK`, `OVER`) -|Not supported. - -|Stored procedures / triggers -|Not supported. Business logic lives in the Elixir orchestration layer. -|=== - -== What VCL Has That SQL Lacks - -[cols="2,3",options="header"] -|=== -|VCL Feature |Description - -|`PROOF` clause -|Requests a verifiable proof certificate alongside query results. Six proof types cover existence, integrity, consistency, provenance, freshness, and authorization guarantees. Enables dependent-type verification via the VCL-UT path. - -|Multimodal `SELECT` -|Select specific modalities (`GRAPH`, `VECTOR`, `TENSOR`, `SEMANTIC`, `DOCUMENT`, `TEMPORAL`) or all (`*`). No SQL equivalent for querying fundamentally different data representations of the same entity. - -|`FEDERATION` queries -|Target distributed VeriSimDB instances: `FROM FEDERATION `. Includes drift policies that govern how cross-instance consistency is enforced during query execution. - -|Drift policies -|`DRIFT POLICY ` attaches consistency enforcement rules to federated queries. Drift detection and repair are first-class query concerns. - -|Octad entity addressing -|`FROM HEXAD ` addresses a single cross-modal entity by its UUID. All six modalities are accessible through one query against one identifier. - -|Graph traversal syntax -|Cypher-inspired patterns directly in `WHERE`: `(h)-[:RELATES_TO]->(target)`. No need for separate graph query language or JOIN emulation. - -|Vector similarity -|`WHERE h.embedding SIMILAR TO [0.1, 0.2, ...]` performs nearest-neighbor search natively. SQL requires extensions (pgvector) or external systems. - -|Full-text predicates -|`CONTAINS(h.content, "term")` and `MATCHES(h.content, "pattern")` are built-in, backed by the Tantivy-powered document store. -|=== - -== Example Comparisons - -=== Retrieving an Entity - -**SQL:** -[source,sql] ----- -SELECT * -FROM entities -WHERE id = '550e8400-e29b-41d4-a716-446655440000'; ----- - -**VCL (Slipstream):** -[source,vcl] ----- -SELECT * -FROM HEXAD 550e8400-e29b-41d4-a716-446655440000 ----- - -The VCL version returns all six modalities for that octad. The SQL version returns only the columns in the `entities` table. - -=== Finding Related Entities - -**SQL (requires JOIN):** -[source,sql] ----- -SELECT e2.* -FROM entities e1 -JOIN relationships r ON e1.id = r.source_id -JOIN entities e2 ON r.target_id = e2.id -WHERE e1.id = '550e8400-e29b-41d4-a716-446655440000' - AND r.type = 'CITES'; ----- - -**VCL (graph traversal):** -[source,vcl] ----- -SELECT GRAPH -FROM HEXAD 550e8400-e29b-41d4-a716-446655440000 -WHERE (h)-[:CITES]->(target) ----- - -VCL eliminates the triple-table JOIN pattern entirely. Graph relationships are native. - -=== Full-Text Search - -**SQL (PostgreSQL example):** -[source,sql] ----- -SELECT * -FROM documents -WHERE to_tsvector('english', content) @@ to_tsquery('quantum & computing'); ----- - -**VCL:** -[source,vcl] ----- -SELECT DOCUMENT -FROM STORE research_papers -WHERE CONTAINS(h.content, "quantum computing") -LIMIT 20 ----- - -=== Similarity Search - -**SQL (requires pgvector extension):** -[source,sql] ----- -SELECT *, embedding <-> '[0.1, 0.2, 0.3, 0.4]' AS distance -FROM entities -ORDER BY embedding <-> '[0.1, 0.2, 0.3, 0.4]' -LIMIT 10; ----- - -**VCL:** -[source,vcl] ----- -SELECT VECTOR -FROM STORE research_papers -WHERE h.embedding SIMILAR TO [0.1, 0.2, 0.3, 0.4] -LIMIT 10 ----- - -=== Verified Query (No SQL Equivalent) - -**VCL-UT (dependent type path):** -[source,vcl] ----- -SELECT * -FROM HEXAD 550e8400-e29b-41d4-a716-446655440000 -PROOF INTEGRITY(DataIntegrityContract) ----- - -This returns the octad data _plus_ a cryptographic proof certificate asserting data integrity. SQL has no mechanism for this. - -== Summary - -[cols="2,1,1",options="header"] -|=== -|Capability |SQL |VCL - -|Relational queries |Yes |No -|Multimodal queries |No |Yes -|Mutations (INSERT/UPDATE/DELETE) |Yes |No (API only) -|JOINs |Yes |No (graph modality) -|Aggregation (GROUP BY, COUNT, SUM) |Yes |No -|Subqueries |Yes |No -|PROOF verification |No |Yes -|Federation with drift policies |No |Yes -|Graph traversal |No (or extensions) |Yes (native) -|Vector similarity |No (or extensions) |Yes (native) -|Full-text search |Extensions |Yes (native) -|Ordering (ORDER BY) |Yes |No -|Pagination (LIMIT/OFFSET) |Yes |Yes -|=== - -VCL is not a replacement for SQL. It is a purpose-built query language for a multimodal database that treats entities as six simultaneous representations rather than rows in flat tables. Use SQL when you need relational algebra. Use VCL when you need to query across modalities with optional formal verification. diff --git a/verisimdb/docs/vcl-vs-vcl-dt.adoc b/verisimdb/docs/vcl-vs-vcl-dt.adoc deleted file mode 100644 index dc526be8..00000000 --- a/verisimdb/docs/vcl-vs-vcl-dt.adoc +++ /dev/null @@ -1,271 +0,0 @@ -// SPDX-License-Identifier: CC-BY-SA-4.0 -// Copyright (c) 2026 Jonathan D.A. Jewell (hyperpolymath) - -= VCL Execution Modes & Safety Architecture -:toc: left -:toclevels: 3 -:sectnums: - -== Overview - -VCL has two distinct concerns, cleanly separated: - -1. **VCL** -- The query language itself. Parsing, basic types, execution. What users write. -2. **VCL-UT** -- The safety pipeline. Applies 10 progressive type-safety levels behind VCL. - Users write VCL; VCL-UT validates it. - -Within VCL execution, there are two modes, determined by the presence or absence -of a `PROOF` clause: - -* **Slipstream** -- Fast, progressive safety (VCL-UT levels apply automatically based - on query complexity — simple queries exit early at L1-L2, complex queries go deeper). -* **Proof-carrying** -- Formally verified queries. The `PROOF` clause triggers dependent - type checking, proof obligation generation, and a proof certificate attached to the - result. This corresponds to VCL-UT levels L9-L10. - -Both modes use the same parser and produce the same AST. They diverge at the query -router, which inspects the AST for a `PROOF` node and selects the appropriate pipeline. - -=== Historical Note: VCL-DT - -VCL-DT (Dependent Types) was the original name for the proof-carrying execution mode. -It is not a separate codebase or product — it is an internal execution path within VCL, -now folded into VCL-UT's L9 (Proof Attachment) and L10 (Cross-Cutting) levels. All -references to "VCL-DT" in older documentation should be read as "VCL proof-carrying mode" -or "VCL-UT L9-L10". - -== The Two-Tier Model - -[cols="1,2,3", options="header"] -|=== -| Tier | What | How - -| **VCL** -| The query language -| Parser (ReScript) → AST → basic validation. Everyone uses this. - -| **VCL-UT** -| The safety pipeline -| 10 progressive levels. Sits behind VCL. Applies automatically. - Simple queries short-circuit at L1-L2. Complex/federated/proof-carrying - queries go through L5-L10. No separate syntax — VCL-UT is invisible - to the user except when they add a `PROOF` clause. -|=== - -== VCL-UT Safety Levels (Progressive) - -The pipeline is progressive — each query enters at L1 and exits as soon as -remaining levels don't apply. A simple `SELECT` exits at L2. A `PROOF INTEGRITY` -query goes all the way to L10. - -[cols="^1,<3,<5,^1", options="header"] -|=== -| Level | Name | What It Prevents | Auto/Manual - -| L1 | Construction Safety | SQL/query injection | Auto -| L2 | Schema Pinning | Schema drift between query and database | Auto -| L3 | Resource Linearity | Connection/cursor leaks | Auto -| L4 | Session Protocols | State machine violations | Auto -| L5 | Effect Tracking | Unaudited side effects | Auto -| L6 | Scope Isolation | Namespace boundary violations | Auto -| L7 | Information Flow | Data lineage tracking failures | Auto -| L8 | Quantitative Bounds | Resource exhaustion (unbounded queries) | Auto -| L9 | Proof Attachment | Missing evidence for sensitive operations | `PROOF` clause -| L10 | Cross-Cutting | Composition bugs across query boundaries | `PROOF` clause -|=== - -Levels L1-L8 apply automatically (slipstream). Levels L9-L10 activate when the -query includes a `PROOF` clause — this is what was previously called "VCL-DT". - -== Slipstream Mode (Default) - -Use for: - -* **Analytics and exploration** -- Ad-hoc queries, dashboards, development -* **Non-critical reads** -- L1-L4 safety is automatic, no overhead -* **Performance-sensitive workloads** -- No proof generation cost - -Example: -[source,vcl] ----- -SELECT GRAPH, DOCUMENT -FROM STORE research_papers -WHERE CONTAINS(h.content, "neural network") -LIMIT 50 ----- - -This query enters VCL-UT at L1 (injection prevention), passes L2 (schema -check), and exits. L3-L10 don't apply — no connections to leak, no proofs -requested. - -== Proof-Carrying Mode (PROOF Clause) - -Use for: - -* **Security assertions** -- Verifying data integrity, access control enforcement -* **Compliance and audit** -- Proof certificates for regulators -* **Provenance chains** -- Cryptographic lineage evidence -* **Cross-instance federation** -- When trust is not assumed - -Example: -[source,vcl] ----- -SELECT * -FROM HEXAD 550e8400-e29b-41d4-a716-446655440000 -PROOF INTEGRITY(DataIntegrityContract) ----- - -This query passes through all 10 VCL-UT levels. At L9, proof obligations are -generated. At L10, the proof certificate is assembled and attached to the result. - -== The Six PROOF Types - -Every `PROOF` clause takes the form `PROOF ()`: - -=== EXISTENCE - -Verifies the requested octad entity exists and has data in the requested modalities. - -[source,vcl] ----- -PROOF EXISTENCE(ExistenceContract) ----- - -=== INTEGRITY - -Verifies returned data matches stored data exactly (Merkle root / hash chain). - -[source,vcl] ----- -PROOF INTEGRITY(DataIntegrityContract) ----- - -=== CONSISTENCY - -Verifies all modalities of the octad are mutually consistent (no drift). - -[source,vcl] ----- -PROOF CONSISTENCY(CrossModalConsistencyContract) ----- - -=== PROVENANCE - -Traces the chain of custody — creation, updates, drift repairs, normalizations. - -[source,vcl] ----- -PROOF PROVENANCE(DataLineageContract) ----- - -=== FRESHNESS - -Verifies data is no older than a specified time bound. - -[source,vcl] ----- -PROOF FRESHNESS(FreshnessContract) ----- - -=== AUTHORIZATION - -Verifies the requesting entity has appropriate permissions. - -[source,vcl] ----- -PROOF AUTHORIZATION(AccessControlContract) ----- - -== Performance - -[cols="2,2,2", options="header"] -|=== -| Phase | Slipstream (L1-L8) | Proof-carrying (L9-L10) - -| Parsing | Same | Same -| AST validation | Progressive check | Full dependent type check -| Proof obligations | Skipped | Generated from AST -| Query execution | Direct store | Store + witness collection -| Result assembly | Data only | Data + proof certificate -| Typical overhead | Baseline | 2x-10x (depends on proof type) -|=== - -The progressive pipeline means slipstream queries only pay for the levels -that apply — a simple SELECT is essentially free beyond parsing. - -== Architecture - -[source,text] ----- - ┌──────────────────┐ - │ VCL Query │ - │ (raw string) │ - └────────┬─────────┘ - │ - ┌────────▼─────────┐ - │ VCL Parser │ - │ (ReScript) │ - └────────┬─────────┘ - │ - ┌────────▼─────────┐ - │ AST │ - └────────┬─────────┘ - │ - ┌────────────▼────────────┐ - │ VCL-UT Pipeline │ - │ │ - │ L1: Construction ──┐ │ - │ L2: Schema ────────┤ │ - │ L3: Resource ──────┤ │ ← Exit early if - │ L4: Session ───────┤ │ remaining levels - │ L5: Effects ───────┤ │ don't apply - │ L6: Scope ─────────┤ │ - │ L7: Info Flow ─────┤ │ - │ L8: Bounds ────────┤ │ - │ L9: Proof* ────────┤ │ ← Only if PROOF clause - │ L10: Cross-Cut* ───┘ │ ← Only if PROOF clause - │ │ - └────────────┬────────────┘ - │ - ┌────────────▼────────────┐ - │ Slipstream: { data } │ - │ -- or -- │ - │ Proof: { data, proof: { │ - │ type, contract, │ - │ certificate } } │ - └─────────────────────────┘ ----- - -== Implementation Status - -[cols="1,3,1", options="header"] -|=== -| Phase | Work | Status - -| Phase 1 -| VCL parser handles PROOF clauses, AST includes proof nodes, query router branches correctly. -| Done - -| Phase 2 -| Type definitions for all six proof types with contract specifications. -| Done - -| Phase 3 -| Wire type checker into the proof execution path. Proof obligations generated from AST. -| Not started - -| Phase 4 -| Proof witness collection in each modality store (Rust core). -| Not started - -| Phase 5 -| Proof certificate assembly. Combine witnesses into verifiable certificate. -| Not started - -| Phase 6 -| Proof cache with TTL to avoid re-verification for unchanged data. -| Not started -|=== - -Phases 3-6 represent the remaining work to produce real, verifiable proof -certificates rather than placeholders. diff --git a/verisimdb/docs/zkp-and-sanctify-integration.adoc b/verisimdb/docs/zkp-and-sanctify-integration.adoc deleted file mode 100644 index 4488115f..00000000 --- a/verisimdb/docs/zkp-and-sanctify-integration.adoc +++ /dev/null @@ -1,65 +0,0 @@ -// SPDX-License-Identifier: CC-BY-SA-4.0 - -= ZK-Historiography: Proven & Sactify-PHP Integration - -This specification details the generation and verification of Non-Interactive Zero-Knowledge Proofs (NIZKs) within the VeriSimDB ecosystem. This allows for "Privacy-Preserving Veracity," where a third party can verify the logical integrity of a Octad without accessing the underlying data. - -== 1. The ZKP Generation Workflow - -The integration relies on `proven` to act as the arithmetic circuit generator and `sactify-php` to act as the secure transport and identity layer. - -*Circuit Definition*:: The `proven` contract (e.g., `CitationContract`) defines the constraints. - -*Witness Generation*:: The store-side logic provides the private data (the "witness") to the `proven` engine. - -*Proof Computation*:: `proven` generates a proof π that the private data satisfies the contract constraints. - -*Sactify Wrapping*:: The proof π is packaged into a `sactify-php` signature block. - -== 2. Technical Specification - -=== A. The Proven Witness - -The `proven` engine generates a succinct witness hash that represents the logical state of the Octad at the moment of verification. - -[source,javascript] ----- -// Witness generation logic -witness CitationWitness(h: Octad) { - apply CitationContract(h); - return extract_proof_witness(); -} ----- - -=== B. The Sactify Signature Structure - -`sactify-php` wraps the witness hash with the signer's identity and a timestamp to prevent replay attacks. - -[source,php] ----- -// sactify-php pseudo-implementation -$payload = [ - 'octad_id' => '0x12AB...', - 'contract_hash' => 'sha256(...)', - 'zk_witness' => $provenWitness, // The NIZK from proven - 'timestamp' => time() -]; - -$signature = Sactify::sign($payload, $privateKey); ----- - -== 3. Verification Without Data Access - -A consumer (e.g., a peer researcher or an automated auditor) receives the signed package. They can verify the claim as follows: - -. *Identity Check*: Use `sactify-php` to verify the signature against the custodian's public key in the ReScript Registry. -. *Logic Check*: Pass the `zk_witness` to the `proven` verification function. -. *Conclusion*: If both pass, the consumer is mathematically certain the data is valid according to the contract, despite never seeing the data itself. - -== 4. Security Considerations - -*Soundness*:: `proven` ensures that a false claim cannot generate a valid witness. - -*Zero-Knowledge*:: No information about the specific references or the text of the claim is leaked in the `zk_witness`. - -*Non-Repudiation*:: The `sactify-php` layer ensures the custodian cannot deny having verified the proof. diff --git a/verisimdb/elixir-orchestration/.formatter.exs b/verisimdb/elixir-orchestration/.formatter.exs deleted file mode 100644 index 42ccc563..00000000 --- a/verisimdb/elixir-orchestration/.formatter.exs +++ /dev/null @@ -1,5 +0,0 @@ -# SPDX-License-Identifier: MPL-2.0 - -[ - inputs: ["{mix,.formatter}.exs", "{config,lib,test}/**/*.{ex,exs}"] -] diff --git a/verisimdb/elixir-orchestration/config/config.exs b/verisimdb/elixir-orchestration/config/config.exs deleted file mode 100644 index fa600b88..00000000 --- a/verisimdb/elixir-orchestration/config/config.exs +++ /dev/null @@ -1,16 +0,0 @@ -# SPDX-License-Identifier: MPL-2.0 - -import Config - -# VeriSim configuration -config :verisim, - rust_core_url: "http://localhost:8080/api/v1", - rust_core_timeout: 30_000 - -# Logger configuration -config :logger, :console, - format: "$time $metadata[$level] $message\n", - metadata: [:request_id, :entity_id] - -# Import environment-specific config -import_config "#{config_env()}.exs" diff --git a/verisimdb/elixir-orchestration/config/dev.exs b/verisimdb/elixir-orchestration/config/dev.exs deleted file mode 100644 index 8fa0c311..00000000 --- a/verisimdb/elixir-orchestration/config/dev.exs +++ /dev/null @@ -1,10 +0,0 @@ -# SPDX-License-Identifier: MPL-2.0 - -import Config - -# Development-specific configuration -config :verisim, - rust_core_url: "http://localhost:8080/api/v1" - -config :logger, :console, - level: :debug diff --git a/verisimdb/elixir-orchestration/config/prod.exs b/verisimdb/elixir-orchestration/config/prod.exs deleted file mode 100644 index c221061e..00000000 --- a/verisimdb/elixir-orchestration/config/prod.exs +++ /dev/null @@ -1,11 +0,0 @@ -# SPDX-License-Identifier: MPL-2.0 - -import Config - -# Production configuration -config :verisim, - rust_core_url: System.get_env("VERISIM_RUST_CORE_URL") || "https://verisim-core:8080/api/v1", - rust_core_timeout: 60_000 - -config :logger, :console, - level: :info diff --git a/verisimdb/elixir-orchestration/config/runtime.exs b/verisimdb/elixir-orchestration/config/runtime.exs deleted file mode 100644 index f41de835..00000000 --- a/verisimdb/elixir-orchestration/config/runtime.exs +++ /dev/null @@ -1,9 +0,0 @@ -# SPDX-License-Identifier: MPL-2.0 - -import Config - -# Runtime configuration loaded at application start -# Read from env in all environments (dev, test, prod) -config :verisim, - rust_core_url: System.get_env("VERISIM_RUST_CORE_URL") || "http://[::1]:8080/api/v1", - rust_core_timeout: String.to_integer(System.get_env("VERISIM_RUST_CORE_TIMEOUT") || "30000") diff --git a/verisimdb/elixir-orchestration/config/test.exs b/verisimdb/elixir-orchestration/config/test.exs deleted file mode 100644 index e414a870..00000000 --- a/verisimdb/elixir-orchestration/config/test.exs +++ /dev/null @@ -1,12 +0,0 @@ -# SPDX-License-Identifier: MPL-2.0 - -import Config - -# Test configuration -config :verisim, - rust_core_url: "http://localhost:8081/api/v1", - rust_core_timeout: 5_000, - orch_api_port: 0 - -config :logger, :console, - level: :warning diff --git a/verisimdb/elixir-orchestration/lib/verisim.ex b/verisimdb/elixir-orchestration/lib/verisim.ex deleted file mode 100644 index cff277cb..00000000 --- a/verisimdb/elixir-orchestration/lib/verisim.ex +++ /dev/null @@ -1,67 +0,0 @@ -# SPDX-License-Identifier: MPL-2.0 - -defmodule VeriSim do - @moduledoc """ - VeriSim Orchestration - Elixir/OTP coordination layer for VeriSimDB. - - This module provides high-level orchestration for the VeriSimDB database, - coordinating between the Rust core and managing distributed operations. - - ## Architecture - - ``` - ┌─────────────────────────────────────────────────────────────┐ - │ Elixir Orchestration Layer │ - │ ├── VeriSim.EntityServer (GenServer per entity) │ - │ ├── VeriSim.DriftMonitor (drift detection coordinator) │ - │ ├── VeriSim.QueryRouter (distributes queries) │ - │ └── VeriSim.SchemaRegistry (type system coordinator) │ - │ ↓ HTTP/gRPC │ - ├─────────────────────────────────────────────────────────────┤ - │ Rust Core (verisim-api) │ - │ ├── 6 modality stores │ - │ ├── Octad management │ - │ └── Normalizer │ - └─────────────────────────────────────────────────────────────┘ - ``` - - ## Design Philosophy (Marr's Three Levels) - - - **Computational Level**: What problem are we solving? - → Maintain cross-modal consistency across 6 representations - - - **Algorithmic Level**: How do we solve it? - → OTP supervision trees for fault tolerance - → Process-per-entity for isolation - → Event-driven drift detection - - - **Implementational Level**: How is it built? - → Elixir/OTP for coordination - → Rust for performance-critical operations - → HTTP API for communication - """ - - @doc """ - Get the current VeriSim version. - """ - def version, do: "0.1.0" - - @doc """ - Health check - returns system status. - """ - def health do - %{ - status: :healthy, - version: version(), - rust_core: rust_core_status(), - uptime_seconds: System.monotonic_time(:second) - } - end - - defp rust_core_status do - case VeriSim.RustClient.health() do - {:ok, status} -> status - {:error, _} -> %{status: :unavailable} - end - end -end diff --git a/verisimdb/elixir-orchestration/lib/verisim/api/router.ex b/verisimdb/elixir-orchestration/lib/verisim/api/router.ex deleted file mode 100644 index 173284d2..00000000 --- a/verisimdb/elixir-orchestration/lib/verisim/api/router.ex +++ /dev/null @@ -1,167 +0,0 @@ -# SPDX-License-Identifier: MPL-2.0 - -defmodule VeriSim.Api.Router do - @moduledoc """ - Lightweight HTTP API router for VeriSimDB orchestration endpoints. - - Serves endpoints that are native to the Elixir orchestration layer - (telemetry, orchestration health, consensus status) rather than proxied - to the Rust core. Runs on a separate port (default 4080) from the Rust - core API (default 8080). - - ## Endpoints - - - `GET /health` — Orchestration layer health check - - `GET /telemetry` — Product telemetry report (opt-in, aggregate-only) - - `GET /telemetry/modality-heatmap` — Modality usage breakdown - - `GET /telemetry/query-patterns` — Query pattern distribution - - `GET /telemetry/drift` — Drift detection report - - `GET /telemetry/performance` — Query performance summary - - `GET /telemetry/federation` — Federation health metrics - - `GET /status` — Orchestration status (consensus, entity count, etc.) - - ## Configuration - - Port is configurable via: - - `config :verisim, orch_api_port: 4080` - - Environment variable: `VERISIM_ORCH_PORT` - - ## CORS - - All endpoints return `Access-Control-Allow-Origin: *` to allow PanLL - and other local tools to query without proxy configuration. - """ - - use Plug.Router - - plug :match - plug Plug.Parsers, parsers: [:json], pass: ["application/json"], json_decoder: Jason - plug :set_cors_headers - plug :dispatch - - # ── Health ──────────────────────────────────────────────────────────── - - get "/health" do - response = %{ - status: "ok", - layer: "orchestration", - node: Node.self() |> to_string(), - uptime_seconds: uptime_seconds(), - telemetry_enabled: VeriSim.Telemetry.Collector.enabled?() - } - - json_response(conn, 200, response) - end - - # ── Full Telemetry Report ───────────────────────────────────────────── - - get "/telemetry" do - if VeriSim.Telemetry.Collector.enabled?() do - json_string = VeriSim.Telemetry.Reporter.report_json() - conn - |> put_resp_content_type("application/json") - |> send_resp(200, json_string) - else - json_response(conn, 200, %{ - telemetry_enabled: false, - message: "Telemetry collection is disabled. Enable with VERISIM_TELEMETRY=true.", - privacy_notice: "When enabled, only aggregate counters are collected. No PII, no query content." - }) - end - end - - # ── Individual Telemetry Sections ───────────────────────────────────── - - get "/telemetry/modality-heatmap" do - json_response(conn, 200, VeriSim.Telemetry.Reporter.modality_heatmap()) - end - - get "/telemetry/query-patterns" do - json_response(conn, 200, VeriSim.Telemetry.Reporter.query_patterns()) - end - - get "/telemetry/drift" do - json_response(conn, 200, VeriSim.Telemetry.Reporter.drift_report()) - end - - get "/telemetry/performance" do - json_response(conn, 200, VeriSim.Telemetry.Reporter.performance_summary()) - end - - get "/telemetry/federation" do - json_response(conn, 200, VeriSim.Telemetry.Reporter.federation_health()) - end - - get "/telemetry/proof-types" do - json_response(conn, 200, VeriSim.Telemetry.Reporter.proof_type_usage()) - end - - get "/telemetry/entities" do - json_response(conn, 200, VeriSim.Telemetry.Reporter.entity_summary()) - end - - get "/telemetry/health" do - json_response(conn, 200, VeriSim.Telemetry.Reporter.health_report()) - end - - get "/telemetry/error-budget" do - json_response(conn, 200, VeriSim.Telemetry.Reporter.error_budget_report()) - end - - # ── Orchestration Status ────────────────────────────────────────────── - - get "/status" do - node_id = Application.get_env(:verisim, :kraft_node_id, "local") - - consensus_status = - try do - VeriSim.Consensus.KRaftNode.diagnostics(node_id) - catch - :exit, _ -> %{state: "unavailable"} - _kind, _reason -> %{state: "unavailable"} - end - - federation_peers = - try do - VeriSim.Federation.Resolver.list_peers() - catch - :exit, _ -> [] - _kind, _reason -> [] - end - - response = %{ - orchestration: "running", - consensus: consensus_status, - federation_adapters: length(federation_peers), - telemetry_enabled: VeriSim.Telemetry.Collector.enabled?() - } - - json_response(conn, 200, response) - end - - # ── Catch-all ───────────────────────────────────────────────────────── - - match _ do - json_response(conn, 404, %{error: "not_found", message: "Unknown endpoint"}) - end - - # ── Private helpers ─────────────────────────────────────────────────── - - defp json_response(conn, status, body) do - conn - |> put_resp_content_type("application/json") - |> send_resp(status, Jason.encode!(body)) - end - - defp set_cors_headers(conn, _opts) do - conn - |> put_resp_header("access-control-allow-origin", "*") - |> put_resp_header("access-control-allow-methods", "GET, OPTIONS") - |> put_resp_header("access-control-allow-headers", "content-type") - end - - defp uptime_seconds do - {uptime_ms, _} = :erlang.statistics(:wall_clock) - div(uptime_ms, 1000) - end -end diff --git a/verisimdb/elixir-orchestration/lib/verisim/application.ex b/verisimdb/elixir-orchestration/lib/verisim/application.ex deleted file mode 100644 index 41adde10..00000000 --- a/verisimdb/elixir-orchestration/lib/verisim/application.ex +++ /dev/null @@ -1,65 +0,0 @@ -# SPDX-License-Identifier: MPL-2.0 - -defmodule VeriSim.Application do - @moduledoc """ - VeriSim Application - OTP application entry point. - - Starts the supervision tree for VeriSim orchestration. - """ - - use Application - - @impl true - def start(_type, _args) do - children = [ - # Telemetry supervisor - VeriSim.Telemetry, - - # Registry for entity servers - {Registry, keys: :unique, name: VeriSim.EntityRegistry}, - - # Dynamic supervisor for entity servers - {DynamicSupervisor, - name: VeriSim.EntitySupervisor, - strategy: :one_for_one, - max_restarts: 100, - max_seconds: 60}, - - # Drift monitor - VeriSim.DriftMonitor, - - # Query router - VeriSim.QueryRouter, - - # Schema registry - VeriSim.SchemaRegistry, - - # Consensus registry (must start before KRaftNode) - {Registry, keys: :unique, name: VeriSim.Consensus.Registry}, - - # KRaft consensus node (single-node bootstrap by default) - {VeriSim.Consensus.KRaftNode, - node_id: Application.get_env(:verisim, :kraft_node_id, "local"), - peers: Application.get_env(:verisim, :kraft_peers, [])}, - - # Federation resolver - VeriSim.Federation.Resolver, - - # Health checker (periodic liveness probing) - VeriSim.HealthChecker, - - # Orchestration HTTP API (telemetry, status endpoints) - {Bandit, plug: VeriSim.Api.Router, port: orch_api_port()} - ] - - opts = [strategy: :rest_for_one, name: VeriSim.Supervisor, max_restarts: 10, max_seconds: 60] - Supervisor.start_link(children, opts) - end - - defp orch_api_port do - case System.get_env("VERISIM_ORCH_PORT") do - nil -> Application.get_env(:verisim, :orch_api_port, 4080) - port -> String.to_integer(port) - end - end -end diff --git a/verisimdb/elixir-orchestration/lib/verisim/consensus/kraft_node.ex b/verisimdb/elixir-orchestration/lib/verisim/consensus/kraft_node.ex deleted file mode 100644 index 29096508..00000000 --- a/verisimdb/elixir-orchestration/lib/verisim/consensus/kraft_node.ex +++ /dev/null @@ -1,874 +0,0 @@ -# SPDX-License-Identifier: MPL-2.0 - -defmodule VeriSim.Consensus.KRaftNode do - @moduledoc """ - KRaft Consensus Node — Elixir GenServer driving the ReScript Raft state machine. - - Each node in the cluster runs a KRaftNode GenServer that: - - Manages election timers and heartbeat scheduling - - Dispatches RPC messages (VoteRequest, AppendEntries) to peers - - Accepts client commands and replicates them via the leader - - Applies committed commands to the local Registry state - - Persists the Raft log via write-ahead log (WAL) - - ## Architecture - - ┌─────────────────────────────────────────────┐ - │ KRaftNode (GenServer) │ - │ ├── Election Timer (Process.send_after) │ - │ ├── Heartbeat Timer (leader only) │ - │ ├── Raft State (from MetadataLog.res) │ - │ └── Registry State (from Registry.res) │ - └─────────────────────────────────────────────┘ - ↕ RPC (GenServer.call) - ┌─────────────────────────────────────────────┐ - │ Peer KRaftNodes (same or remote) │ - └─────────────────────────────────────────────┘ - - ## Usage - - # Start a 3-node cluster - {:ok, _} = KRaftNode.start_link(node_id: "node-1", peers: ["node-2", "node-3"]) - {:ok, _} = KRaftNode.start_link(node_id: "node-2", peers: ["node-1", "node-3"]) - {:ok, _} = KRaftNode.start_link(node_id: "node-3", peers: ["node-1", "node-2"]) - - # Propose a command (routes to leader) - {:ok, index} = KRaftNode.propose("node-1", {:register_store, "store-1", "http://...", ["graph"]}) - """ - - use GenServer - require Logger - - alias VeriSim.Consensus.KRaftWAL - - @election_timeout_min 150 - @election_timeout_max 300 - @heartbeat_interval 50 - @tick_interval 10 - @snapshot_interval 1000 - - # --------------------------------------------------------------------------- - # Types - # --------------------------------------------------------------------------- - - defmodule State do - @moduledoc false - defstruct [ - :node_id, - :peers, - :role, - :current_term, - :voted_for, - :log, - :commit_index, - :last_applied, - :next_index, - :match_index, - :leader_id, - :votes_received, - :election_timer, - :heartbeat_timer, - :registry, - :pending_requests, - :election_count, - :wal_path - ] - end - - # --------------------------------------------------------------------------- - # Client API - # --------------------------------------------------------------------------- - - def start_link(opts) do - node_id = Keyword.fetch!(opts, :node_id) - GenServer.start_link(__MODULE__, opts, name: via(node_id)) - end - - @doc "Propose a command to the cluster (will forward to leader if needed)." - def propose(node_id, command) do - GenServer.call(via(node_id), {:propose, command}) - end - - @doc "Get the current diagnostics for a node." - def diagnostics(node_id) do - GenServer.call(via(node_id), :diagnostics) - end - - @doc "Get the current leader ID." - def leader(node_id) do - GenServer.call(via(node_id), :leader) - end - - @doc "Get the registry state." - def registry(node_id) do - GenServer.call(via(node_id), :registry) - end - - @doc """ - Add a server to the cluster dynamically. - - The new node starts as non-voting until caught up with the leader's log. - The command is replicated through the Raft log so all nodes learn about - the membership change atomically at the same commit index. - - ## Parameters - - - `node_id` — the node to send the request to (should be the leader) - - `new_node_id` — the ID of the node being added - - `new_peer_list` — the updated peer list for the new node (optional metadata) - """ - def add_server(node_id, new_node_id, new_peer_list \\ []) do - GenServer.call(via(node_id), {:propose, {:add_server, new_node_id, new_peer_list}}) - end - - @doc """ - Remove a server from the cluster dynamically. - - The removed node will no longer receive AppendEntries or vote requests - once the removal command commits. The removed node should be stopped - separately after the command commits. - - ## Parameters - - - `node_id` — the node to send the request to (should be the leader) - - `target_node_id` — the ID of the node being removed - """ - def remove_server(node_id, target_node_id) do - GenServer.call(via(node_id), {:propose, {:remove_server, target_node_id}}) - end - - @doc "Receive a vote request RPC from a candidate." - def request_vote(node_id, request) do - GenServer.call(via(node_id), {:request_vote, request}) - end - - @doc "Receive an AppendEntries RPC from a leader." - def append_entries(node_id, request) do - GenServer.call(via(node_id), {:append_entries, request}) - end - - # --------------------------------------------------------------------------- - # GenServer Callbacks - # --------------------------------------------------------------------------- - - @impl true - def init(opts) do - node_id = Keyword.fetch!(opts, :node_id) - peers = Keyword.get(opts, :peers, []) - wal_path = Keyword.get(opts, :wal_path, nil) - - # Initialize WAL directory and recover persisted state - KRaftWAL.init(wal_path) - {current_term, voted_for, log, registry, snapshot_index} = recover_from_wal(wal_path) - - state = %State{ - node_id: node_id, - peers: peers, - role: :follower, - current_term: current_term, - voted_for: voted_for, - log: log, - commit_index: snapshot_index, - last_applied: snapshot_index, - next_index: %{}, - match_index: %{}, - leader_id: nil, - votes_received: MapSet.new(), - election_timer: schedule_election_timeout(), - heartbeat_timer: nil, - registry: registry, - pending_requests: [], - election_count: 0, - wal_path: wal_path - } - - Logger.info( - "KRaft: node #{node_id} started as follower " <> - "(term=#{current_term}, log=#{length(log)} entries)" - ) - - {:ok, state} - end - - @impl true - def handle_call({:propose, command}, from, %{role: :leader} = state) do - entry = %{ - term: state.current_term, - index: length(state.log) + 1, - command: command, - timestamp: System.system_time(:millisecond) - } - - # Persist entry to WAL before replicating (durability guarantee) - KRaftWAL.append_entry(state.wal_path, entry) - - new_log = state.log ++ [entry] - pending = [{entry.index, from} | state.pending_requests] - - state = %{state | log: new_log, pending_requests: pending} - - # Replicate to followers - send_append_entries(state) - - # Try to advance commit immediately (handles single-node clusters where - # no AppendEntries responses will arrive but we already have quorum of 1) - state = maybe_advance_commit(state) - state = apply_committed(state) - state = reply_to_pending(state) - - {:noreply, state} - end - - def handle_call({:propose, _command}, _from, state) do - {:reply, {:error, {:not_leader, state.leader_id}}, state} - end - - def handle_call(:diagnostics, _from, state) do - diag = %{ - node_id: state.node_id, - role: state.role, - current_term: state.current_term, - commit_index: state.commit_index, - last_applied: state.last_applied, - log_length: length(state.log), - peer_count: length(state.peers), - leader_id: state.leader_id, - election_count: state.election_count, - pending_requests: length(state.pending_requests) - } - - {:reply, diag, state} - end - - def handle_call(:leader, _from, state) do - {:reply, state.leader_id, state} - end - - def handle_call(:registry, _from, state) do - {:reply, state.registry, state} - end - - def handle_call({:request_vote, request}, _from, state) do - {new_state, response} = handle_vote_request(state, request) - {:reply, response, new_state} - end - - def handle_call({:append_entries, request}, _from, state) do - {new_state, response} = handle_append_entries_rpc(state, request) - {:reply, response, new_state} - end - - @impl true - def handle_info(:election_timeout, %{role: role} = state) - when role in [:follower, :candidate] do - Logger.info("KRaft: node #{state.node_id} election timeout, starting election") - state = start_election(state) - {:noreply, state} - end - - def handle_info(:election_timeout, state) do - # Leaders ignore election timeouts - {:noreply, state} - end - - def handle_info(:heartbeat, %{role: :leader} = state) do - send_append_entries(state) - state = %{state | heartbeat_timer: schedule_heartbeat()} - {:noreply, state} - end - - def handle_info(:heartbeat, state) do - # Non-leaders ignore heartbeat timers - {:noreply, state} - end - - def handle_info({:vote_response, from_peer, response}, state) do - state = handle_vote_response(state, from_peer, response) - {:noreply, state} - end - - def handle_info({:append_entries_response, from_peer, response}, state) do - state = handle_ae_response(state, from_peer, response) - {:noreply, state} - end - - def handle_info(_msg, state), do: {:noreply, state} - - # --------------------------------------------------------------------------- - # Election Logic - # --------------------------------------------------------------------------- - - defp start_election(state) do - new_term = state.current_term + 1 - - Logger.info( - "KRaft: node #{state.node_id} starting election for term #{new_term}" - ) - - state = %{ - state - | role: :candidate, - current_term: new_term, - voted_for: state.node_id, - votes_received: MapSet.new([state.node_id]), - election_count: state.election_count + 1, - election_timer: schedule_election_timeout() - } - - # Persist term and vote BEFORE sending RPCs (Raft safety requirement) - KRaftWAL.persist_state(state.wal_path, new_term, state.node_id) - - # Send vote requests to all peers - last_log_index = length(state.log) - last_log_term = last_log_term(state) - - request = %{ - term: new_term, - candidate_id: state.node_id, - last_log_index: last_log_index, - last_log_term: last_log_term - } - - for peer <- state.peers do - Task.start(fn -> - try do - response = GenServer.call(via(peer), {:request_vote, request}, 1_000) - send(self_pid(state.node_id), {:vote_response, peer, response}) - rescue - _ -> :timeout - end - end) - end - - # Check if we already have quorum (e.g., single-node cluster with 0 peers) - quorum = div(length(state.peers) + 1, 2) + 1 - - if MapSet.size(state.votes_received) >= quorum do - become_leader(state) - else - state - end - end - - defp handle_vote_request(state, request) do - cond do - request.term < state.current_term -> - {state, %{term: state.current_term, vote_granted: false}} - - request.term > state.current_term -> - # Higher term — step down and grant vote if log is up-to-date - state = step_down(state, request.term) - - if log_up_to_date?(state, request) do - state = %{state | voted_for: request.candidate_id} - KRaftWAL.persist_state(state.wal_path, state.current_term, state.voted_for) - state = reset_election_timer(state) - {state, %{term: state.current_term, vote_granted: true}} - else - {state, %{term: state.current_term, vote_granted: false}} - end - - state.voted_for == nil or state.voted_for == request.candidate_id -> - if log_up_to_date?(state, request) do - state = %{state | voted_for: request.candidate_id} - KRaftWAL.persist_state(state.wal_path, state.current_term, state.voted_for) - state = reset_election_timer(state) - {state, %{term: state.current_term, vote_granted: true}} - else - {state, %{term: state.current_term, vote_granted: false}} - end - - true -> - {state, %{term: state.current_term, vote_granted: false}} - end - end - - defp handle_vote_response(state, from_peer, response) do - cond do - response.term > state.current_term -> - step_down(state, response.term) - - state.role == :candidate and response.vote_granted -> - votes = MapSet.put(state.votes_received, from_peer) - quorum = div(length(state.peers) + 1, 2) + 1 - - if MapSet.size(votes) >= quorum do - become_leader(%{state | votes_received: votes}) - else - %{state | votes_received: votes} - end - - true -> - state - end - end - - defp become_leader(state) do - Logger.info( - "KRaft: node #{state.node_id} became leader for term #{state.current_term}" - ) - - last_log_index = length(state.log) + 1 - - next_index = - state.peers - |> Enum.map(&{&1, last_log_index}) - |> Map.new() - - match_index = - state.peers - |> Enum.map(&{&1, 0}) - |> Map.new() - - # Cancel election timer, start heartbeat timer - if state.election_timer, do: Process.cancel_timer(state.election_timer) - - state = %{ - state - | role: :leader, - leader_id: state.node_id, - next_index: next_index, - match_index: match_index, - election_timer: nil, - heartbeat_timer: schedule_heartbeat() - } - - # Append NoOp to commit entries from previous terms - noop = %{ - term: state.current_term, - index: length(state.log) + 1, - command: :noop, - timestamp: System.system_time(:millisecond) - } - - KRaftWAL.append_entry(state.wal_path, noop) - state = %{state | log: state.log ++ [noop]} - - # Send initial heartbeats - send_append_entries(state) - - state - end - - # --------------------------------------------------------------------------- - # Log Replication - # --------------------------------------------------------------------------- - - defp send_append_entries(state) do - for peer <- state.peers do - next_idx = Map.get(state.next_index, peer, 1) - prev_log_index = next_idx - 1 - - prev_log_term = - if prev_log_index > 0 do - case Enum.at(state.log, prev_log_index - 1) do - nil -> 0 - entry -> entry.term - end - else - 0 - end - - entries = Enum.drop(state.log, next_idx - 1) - - request = %{ - term: state.current_term, - leader_id: state.node_id, - prev_log_index: prev_log_index, - prev_log_term: prev_log_term, - entries: entries, - leader_commit: state.commit_index - } - - node_id = state.node_id - - Task.start(fn -> - try do - response = GenServer.call(via(peer), {:append_entries, request}, 1_000) - send(self_pid(node_id), {:append_entries_response, peer, response}) - rescue - _ -> :timeout - end - end) - end - end - - defp handle_append_entries_rpc(state, request) do - cond do - request.term < state.current_term -> - {state, %{term: state.current_term, success: false, match_index: 0}} - - true -> - state = - if request.term > state.current_term do - step_down(state, request.term) - else - state - end - - state = %{state | leader_id: request.leader_id} - state = reset_election_timer(state) - - # Check prev log entry - prev_ok = - if request.prev_log_index == 0 do - true - else - case Enum.at(state.log, request.prev_log_index - 1) do - nil -> false - entry -> entry.term == request.prev_log_term - end - end - - if prev_ok do - # Append new entries (truncating conflicting suffix) - new_log = Enum.take(state.log, request.prev_log_index) ++ request.entries - new_commit = min(request.leader_commit, length(new_log)) - - # Persist received entries to WAL for crash recovery - if request.entries != [] do - if request.prev_log_index < length(state.log) do - # Log conflict — truncate WAL entries at conflicting indices first - KRaftWAL.truncate_after(state.wal_path, request.prev_log_index) - end - - KRaftWAL.append_entries(state.wal_path, request.entries) - end - - state = %{state | log: new_log, commit_index: new_commit} - state = apply_committed(state) - - {state, - %{ - term: state.current_term, - success: true, - match_index: length(state.log) - }} - else - {state, %{term: state.current_term, success: false, match_index: 0}} - end - end - end - - defp handle_ae_response(state, from_peer, response) do - cond do - response.term > state.current_term -> - step_down(state, response.term) - - state.role == :leader and response.success -> - next_index = Map.put(state.next_index, from_peer, response.match_index + 1) - match_index = Map.put(state.match_index, from_peer, response.match_index) - - state = %{state | next_index: next_index, match_index: match_index} - state = maybe_advance_commit(state) - state = apply_committed(state) - state = reply_to_pending(state) - - state - - state.role == :leader -> - # Decrement nextIndex and retry - current_next = Map.get(state.next_index, from_peer, 1) - next_index = Map.put(state.next_index, from_peer, max(current_next - 1, 1)) - %{state | next_index: next_index} - - true -> - state - end - end - - # --------------------------------------------------------------------------- - # Commit & Apply - # --------------------------------------------------------------------------- - - defp maybe_advance_commit(state) do - # Find highest N where majority of matchIndex[i] >= N - all_match = - (Map.values(state.match_index) ++ [length(state.log)]) - |> Enum.sort(:desc) - - quorum_idx = div(length(state.peers) + 1, 2) - new_commit = Enum.at(all_match, quorum_idx, state.commit_index) - - # Only commit entries from current term (Raft safety property) - can_commit = - case Enum.at(state.log, new_commit - 1) do - nil -> false - entry -> entry.term == state.current_term - end - - if can_commit and new_commit > state.commit_index do - %{state | commit_index: new_commit} - else - state - end - end - - defp apply_committed(state) do - if state.last_applied >= state.commit_index do - state - else - entries_to_apply = - state.log - |> Enum.slice(state.last_applied, state.commit_index - state.last_applied) - - {new_registry, new_peers} = - Enum.reduce(entries_to_apply, {state.registry, state.peers}, fn entry, {reg, peers} -> - new_reg = apply_command(reg, entry.command) - new_peers = apply_membership_change(peers, state.node_id, entry.command) - {new_reg, new_peers} - end) - - Logger.debug( - "KRaft: node #{state.node_id} applied #{length(entries_to_apply)} entries " <> - "(#{state.last_applied + 1}..#{state.commit_index})" - ) - - state = %{state | last_applied: state.commit_index, registry: new_registry, peers: new_peers} - - # Update leader tracking structures when peers change - state = - if state.role == :leader and state.peers != new_peers do - # Add next_index/match_index entries for any new peers - new_next_index = - Enum.reduce(new_peers, state.next_index, fn peer, ni -> - Map.put_new(ni, peer, length(state.log) + 1) - end) - - new_match_index = - Enum.reduce(new_peers, state.match_index, fn peer, mi -> - Map.put_new(mi, peer, 0) - end) - - # Remove entries for removed peers - removed = MapSet.difference(MapSet.new(Map.keys(state.next_index)), MapSet.new(new_peers)) - - new_next_index = Map.drop(new_next_index, MapSet.to_list(removed)) - new_match_index = Map.drop(new_match_index, MapSet.to_list(removed)) - - %{state | next_index: new_next_index, match_index: new_match_index} - else - state - end - - # Trigger snapshot periodically to bound WAL growth - maybe_trigger_snapshot(state) - end - end - - # Update the peer list when a membership change command is applied. - # The new node is added as non-voting initially (it appears in the peer - # list but needs to catch up with the log before it counts for quorum). - defp apply_membership_change(peers, self_id, {:add_server, new_node_id, _peer_list}) do - if new_node_id != self_id and new_node_id not in peers do - peers ++ [new_node_id] - else - peers - end - end - - defp apply_membership_change(peers, self_id, {:remove_server, target_node_id}) do - if target_node_id != self_id do - List.delete(peers, target_node_id) - else - # A node removing itself — leave the peer list intact; the node - # will be stopped separately after the command commits. - peers - end - end - - defp apply_membership_change(peers, _self_id, _command), do: peers - - defp maybe_trigger_snapshot(state) do - if state.wal_path != nil and state.last_applied > 0 and - rem(state.last_applied, @snapshot_interval) == 0 do - last_entry = Enum.at(state.log, state.last_applied - 1) - - if last_entry do - KRaftWAL.save_snapshot( - state.wal_path, - state.registry, - state.last_applied, - last_entry.term - ) - - Logger.info( - "KRaft: node #{state.node_id} saved snapshot at index #{state.last_applied}" - ) - end - end - - state - end - - defp apply_command(registry, {:register_store, store_id, endpoint, modalities}) do - store = %{ - store_id: store_id, - endpoint: endpoint, - modalities: modalities, - trust_level: 1.0, - last_seen: DateTime.utc_now(), - response_time_ms: nil - } - - put_in(registry, [:stores, store_id], store) - end - - defp apply_command(registry, {:unregister_store, store_id}) do - update_in(registry, [:stores], &Map.delete(&1, store_id)) - end - - defp apply_command(registry, {:map_octad, octad_id, locations}) do - mapping = %{ - octad_id: octad_id, - locations: locations, - primary_store: nil, - created: DateTime.utc_now(), - modified: DateTime.utc_now() - } - - put_in(registry, [:mappings, octad_id], mapping) - end - - defp apply_command(registry, {:unmap_octad, octad_id}) do - update_in(registry, [:mappings], &Map.delete(&1, octad_id)) - end - - defp apply_command(registry, {:update_trust, store_id, new_trust}) do - case get_in(registry, [:stores, store_id]) do - nil -> registry - store -> put_in(registry, [:stores, store_id], %{store | trust_level: new_trust}) - end - end - - defp apply_command(registry, {:add_server, node_id, _peer_list}) do - # Track membership in registry config for observability. - # The actual peer list update happens in the GenServer state (see below). - members = get_in(registry, [:config, :members]) || [] - - unless node_id in members do - put_in(registry, [:config, :members], members ++ [node_id]) - else - registry - end - end - - defp apply_command(registry, {:remove_server, node_id}) do - members = get_in(registry, [:config, :members]) || [] - put_in(registry, [:config, :members], List.delete(members, node_id)) - end - - defp apply_command(registry, :noop), do: registry - - defp reply_to_pending(state) do - {fulfilled, remaining} = - Enum.split_with(state.pending_requests, fn {index, _from} -> - index <= state.commit_index - end) - - Enum.each(fulfilled, fn {index, from} -> - GenServer.reply(from, {:ok, index}) - end) - - %{state | pending_requests: remaining} - end - - # --------------------------------------------------------------------------- - # Helpers - # --------------------------------------------------------------------------- - - defp step_down(state, new_term) do - if state.heartbeat_timer, do: Process.cancel_timer(state.heartbeat_timer) - - # Persist new term BEFORE responding (Raft safety requirement) - KRaftWAL.persist_state(state.wal_path, new_term, nil) - - %{ - state - | role: :follower, - current_term: new_term, - voted_for: nil, - votes_received: MapSet.new(), - heartbeat_timer: nil, - election_timer: schedule_election_timeout() - } - end - - defp reset_election_timer(state) do - if state.election_timer, do: Process.cancel_timer(state.election_timer) - %{state | election_timer: schedule_election_timeout()} - end - - defp log_up_to_date?(state, request) do - my_last_term = last_log_term(state) - my_last_index = length(state.log) - - cond do - request.last_log_term > my_last_term -> true - request.last_log_term == my_last_term and request.last_log_index >= my_last_index -> true - true -> false - end - end - - defp last_log_term(state) do - case List.last(state.log) do - nil -> 0 - entry -> entry.term - end - end - - defp schedule_election_timeout do - timeout = - @election_timeout_min + - :rand.uniform(@election_timeout_max - @election_timeout_min) - - Process.send_after(self(), :election_timeout, timeout) - end - - defp schedule_heartbeat do - Process.send_after(self(), :heartbeat, @heartbeat_interval) - end - - defp via(node_id) do - {:via, Registry, {VeriSim.Consensus.Registry, node_id}} - end - - defp self_pid(node_id) do - case Registry.lookup(VeriSim.Consensus.Registry, node_id) do - [{pid, _}] -> pid - _ -> self() - end - end - - defp initial_registry do - %{ - stores: %{}, - mappings: %{}, - config: %{ - min_trust_level: 0.5, - max_store_downtime_ms: 300_000, - replication_factor: 3, - consistency_mode: :quorum - } - } - end - - defp recover_from_wal(wal_path) do - case KRaftWAL.recover(wal_path) do - {:ok, nil} -> - # Fresh start — no WAL found - {0, nil, [], initial_registry(), 0} - - {:ok, recovered} -> - registry = recovered.registry || initial_registry() - - { - recovered.current_term, - recovered.voted_for, - recovered.log, - registry, - recovered.snapshot_index - } - end - end -end diff --git a/verisimdb/elixir-orchestration/lib/verisim/consensus/kraft_supervisor.ex b/verisimdb/elixir-orchestration/lib/verisim/consensus/kraft_supervisor.ex deleted file mode 100644 index 143aefb1..00000000 --- a/verisimdb/elixir-orchestration/lib/verisim/consensus/kraft_supervisor.ex +++ /dev/null @@ -1,48 +0,0 @@ -# SPDX-License-Identifier: MPL-2.0 - -defmodule VeriSim.Consensus.KRaftSupervisor do - @moduledoc """ - Supervisor for the KRaft consensus cluster. - - Starts a local Elixir Registry for node name resolution, then - starts the configured KRaft nodes under supervision. - - ## Usage - - # In application.ex - children = [ - {VeriSim.Consensus.KRaftSupervisor, nodes: [ - [node_id: "node-1", peers: ["node-2", "node-3"]], - [node_id: "node-2", peers: ["node-1", "node-3"]], - [node_id: "node-3", peers: ["node-1", "node-2"]], - ]} - ] - """ - - use Supervisor - - def start_link(opts) do - Supervisor.start_link(__MODULE__, opts, name: __MODULE__) - end - - @impl true - def init(opts) do - nodes = Keyword.get(opts, :nodes, []) - - children = - [ - # Process registry for node name resolution - {Registry, keys: :unique, name: VeriSim.Consensus.Registry} - ] ++ - Enum.map(nodes, fn node_opts -> - node_id = Keyword.fetch!(node_opts, :node_id) - - Supervisor.child_spec( - {VeriSim.Consensus.KRaftNode, node_opts}, - id: {VeriSim.Consensus.KRaftNode, node_id} - ) - end) - - Supervisor.init(children, strategy: :one_for_one) - end -end diff --git a/verisimdb/elixir-orchestration/lib/verisim/consensus/kraft_transport.ex b/verisimdb/elixir-orchestration/lib/verisim/consensus/kraft_transport.ex deleted file mode 100644 index c72f9ae6..00000000 --- a/verisimdb/elixir-orchestration/lib/verisim/consensus/kraft_transport.ex +++ /dev/null @@ -1,227 +0,0 @@ -# SPDX-License-Identifier: MPL-2.0 - -defmodule VeriSim.Consensus.KRaftTransport do - @moduledoc """ - Network transport abstraction for KRaft Raft consensus. - - Provides RPC delivery to peers via either local GenServer.call (same VM) - or HTTP (remote nodes). Automatically detects whether a peer is local - (registered in the Elixir Registry) or remote (requires HTTP). - - ## Peer Format - - Peers can be specified as: - - `"node-1"` — resolved via local Elixir Registry (same VM) - - `{"node-1", "http://host:4000"}` — resolved via HTTP (remote VM) - - ## HTTP Endpoints - - When using HTTP transport, the remote node must expose: - - `POST /raft/vote` — RequestVote RPC - - `POST /raft/append` — AppendEntries RPC - - `POST /raft/propose` — Client proposal (forwarded to leader) - - ## Architecture - - ┌──────────────────────────────────────────────────┐ - │ KRaftTransport │ - │ ├── Local: GenServer.call via Registry │ - │ └── Remote: HTTP POST via :httpc │ - └──────────────────────────────────────────────────┘ - ↕ ↕ - ┌──────────────────┐ ┌──────────────────────┐ - │ Same-VM Peers │ │ Remote Peers (HTTP) │ - └──────────────────┘ └──────────────────────┘ - """ - - require Logger - - @rpc_timeout 1_000 - @connect_timeout 500 - - # --------------------------------------------------------------------------- - # Public API - # --------------------------------------------------------------------------- - - @doc """ - Send a RequestVote RPC to a peer. - - Returns `{:ok, response}` or `{:error, reason}`. - The response is a map with `:term` and `:vote_granted` keys. - """ - def send_vote_request(peer, request) do - case resolve_peer(peer) do - {:local, node_id} -> - local_call(node_id, {:request_vote, request}) - - {:remote, _node_id, endpoint} -> - http_post(endpoint, "/raft/vote", request) - - {:error, reason} -> - {:error, reason} - end - end - - @doc """ - Send an AppendEntries RPC to a peer. - - Returns `{:ok, response}` or `{:error, reason}`. - The response is a map with `:term`, `:success`, and `:match_index` keys. - """ - def send_append_entries(peer, request) do - case resolve_peer(peer) do - {:local, node_id} -> - local_call(node_id, {:append_entries, request}) - - {:remote, _node_id, endpoint} -> - http_post(endpoint, "/raft/append", request) - - {:error, reason} -> - {:error, reason} - end - end - - @doc """ - Send a vote request to a peer asynchronously. - - Spawns a Task that sends the result back to the caller process - as `{:vote_response, peer_id, response}`. - """ - def async_vote_request(peer, request, reply_to) do - peer_id = peer_id(peer) - - Task.start(fn -> - case send_vote_request(peer, request) do - {:ok, response} -> - send(reply_to, {:vote_response, peer_id, response}) - - {:error, reason} -> - Logger.debug("KRaft transport: vote request to #{peer_id} failed: #{inspect(reason)}") - end - end) - end - - @doc """ - Send an AppendEntries RPC to a peer asynchronously. - - Spawns a Task that sends the result back to the caller process - as `{:append_entries_response, peer_id, response}`. - """ - def async_append_entries(peer, request, reply_to) do - peer_id = peer_id(peer) - - Task.start(fn -> - case send_append_entries(peer, request) do - {:ok, response} -> - send(reply_to, {:append_entries_response, peer_id, response}) - - {:error, reason} -> - Logger.debug( - "KRaft transport: append_entries to #{peer_id} failed: #{inspect(reason)}" - ) - end - end) - end - - @doc """ - Extract the node ID string from a peer specification. - """ - def peer_id({node_id, _endpoint}), do: node_id - def peer_id(node_id) when is_binary(node_id), do: node_id - - # --------------------------------------------------------------------------- - # Private: Peer Resolution - # --------------------------------------------------------------------------- - - defp resolve_peer({node_id, endpoint}) when is_binary(endpoint) do - # Explicit remote endpoint — always use HTTP - {:remote, node_id, endpoint} - end - - defp resolve_peer(node_id) when is_binary(node_id) do - # Check if peer is registered locally first - case Registry.lookup(VeriSim.Consensus.Registry, node_id) do - [{_pid, _}] -> {:local, node_id} - _ -> {:error, {:peer_not_found, node_id}} - end - end - - # --------------------------------------------------------------------------- - # Private: Local RPC (same VM) - # --------------------------------------------------------------------------- - - defp local_call(node_id, message) do - via = {:via, Registry, {VeriSim.Consensus.Registry, node_id}} - - try do - response = GenServer.call(via, message, @rpc_timeout) - {:ok, response} - catch - :exit, _ -> {:error, :timeout} - end - end - - # --------------------------------------------------------------------------- - # Private: HTTP RPC (remote VM) - # --------------------------------------------------------------------------- - - defp http_post(endpoint, path, request) do - url = String.to_charlist("#{endpoint}#{path}") - body = Jason.encode!(serialize_for_json(request)) - content_type = ~c"application/json" - http_opts = [timeout: @rpc_timeout, connect_timeout: @connect_timeout] - - case :httpc.request( - :post, - {url, [], content_type, body}, - http_opts, - [] - ) do - {:ok, {{_, 200, _}, _, response_body}} -> - parsed = Jason.decode!(to_string(response_body)) - {:ok, atomize_keys(parsed)} - - {:ok, {{_, status, _}, _, _}} -> - {:error, {:http_status, status}} - - {:error, reason} -> - {:error, {:http_error, reason}} - end - rescue - e -> {:error, {:http_exception, Exception.message(e)}} - end - - # --------------------------------------------------------------------------- - # Private: JSON Serialization Helpers - # --------------------------------------------------------------------------- - - @doc false - def serialize_for_json(map) when is_map(map) do - Map.new(map, fn - {k, v} when is_atom(k) -> {Atom.to_string(k), serialize_for_json(v)} - {k, v} -> {k, serialize_for_json(v)} - end) - end - - def serialize_for_json(list) when is_list(list) do - Enum.map(list, &serialize_for_json/1) - end - - def serialize_for_json(atom) when is_atom(atom) and not is_nil(atom) and not is_boolean(atom) do - Atom.to_string(atom) - end - - def serialize_for_json(other), do: other - - defp atomize_keys(map) when is_map(map) do - Map.new(map, fn {k, v} -> - key = if is_binary(k), do: String.to_existing_atom(k), else: k - {key, atomize_keys(v)} - end) - rescue - ArgumentError -> map - end - - defp atomize_keys(list) when is_list(list), do: Enum.map(list, &atomize_keys/1) - defp atomize_keys(other), do: other -end diff --git a/verisimdb/elixir-orchestration/lib/verisim/consensus/kraft_wal.ex b/verisimdb/elixir-orchestration/lib/verisim/consensus/kraft_wal.ex deleted file mode 100644 index 83afad6e..00000000 --- a/verisimdb/elixir-orchestration/lib/verisim/consensus/kraft_wal.ex +++ /dev/null @@ -1,410 +0,0 @@ -# SPDX-License-Identifier: MPL-2.0 - -defmodule VeriSim.Consensus.KRaftWAL do - @moduledoc """ - Write-Ahead Log for the KRaft Raft consensus implementation. - - Persists Raft log entries as newline-delimited JSON (JSONL) for crash recovery. - Also persists durable state (currentTerm, votedFor) which MUST survive restarts - to maintain Raft's safety guarantees. - - ## File Layout - - Given `wal_path = "/data/raft/node-1"`, the WAL creates: - - ``` - /data/raft/node-1/ - ├── wal.jsonl # Append-only log entries - ├── state.json # currentTerm + votedFor (overwritten atomically) - └── snapshot.json # Registry state at snapshot index (overwritten) - ``` - - ## WAL Format (wal.jsonl) - - Each line is a JSON object: - ```json - {"term":1,"index":1,"command":{"type":"register_store","store_id":"s1","endpoint":"http://...","modalities":["graph"]},"timestamp":1234567890} - ``` - - ## Recovery - - On startup: - 1. Read `state.json` to restore currentTerm and votedFor - 2. Read `snapshot.json` to restore registry state + last included index - 3. Read `wal.jsonl` to replay log entries after the snapshot index - """ - - require Logger - - @state_file "state.json" - @wal_file "wal.jsonl" - @snapshot_file "snapshot.json" - - # --------------------------------------------------------------------------- - # Public API - # --------------------------------------------------------------------------- - - @doc """ - Initialize the WAL directory. Creates the directory if needed. - Returns `:ok` or `{:error, reason}`. - """ - def init(wal_path) when is_binary(wal_path) do - case File.mkdir_p(wal_path) do - :ok -> :ok - {:error, reason} -> {:error, {:wal_init_failed, reason}} - end - end - def init(nil), do: :ok - - @doc """ - Persist the durable Raft state (currentTerm, votedFor). - - MUST be called before responding to any RPC that changes these values. - Uses atomic write (write to .tmp, then rename) to prevent corruption. - """ - def persist_state(wal_path, current_term, voted_for) when is_binary(wal_path) do - state = %{ - "current_term" => current_term, - "voted_for" => voted_for - } - - path = Path.join(wal_path, @state_file) - tmp_path = path <> ".tmp" - - case File.write(tmp_path, Jason.encode!(state)) do - :ok -> - File.rename(tmp_path, path) - {:error, reason} -> - Logger.error("KRaft WAL: failed to persist state: #{inspect(reason)}") - {:error, reason} - end - end - def persist_state(nil, _term, _voted_for), do: :ok - - @doc """ - Append a log entry to the WAL. The entry is fsync'd to ensure durability. - """ - def append_entry(wal_path, entry) when is_binary(wal_path) do - path = Path.join(wal_path, @wal_file) - json_line = Jason.encode!(serialize_entry(entry)) <> "\n" - - case File.write(path, json_line, [:append, :sync]) do - :ok -> :ok - {:error, reason} -> - Logger.error("KRaft WAL: failed to append entry: #{inspect(reason)}") - {:error, reason} - end - end - def append_entry(nil, _entry), do: :ok - - @doc """ - Append multiple log entries to the WAL in a single write. - """ - def append_entries(wal_path, entries) when is_binary(wal_path) and is_list(entries) do - if entries == [] do - :ok - else - path = Path.join(wal_path, @wal_file) - lines = Enum.map_join(entries, fn entry -> - Jason.encode!(serialize_entry(entry)) <> "\n" - end) - - case File.write(path, lines, [:append, :sync]) do - :ok -> :ok - {:error, reason} -> - Logger.error("KRaft WAL: failed to append entries: #{inspect(reason)}") - {:error, reason} - end - end - end - def append_entries(nil, _entries), do: :ok - - @doc """ - Save a snapshot of the registry state at a given index. - - After saving, truncates the WAL to remove entries at or before the snapshot index. - """ - def save_snapshot(wal_path, registry, last_included_index, last_included_term) when is_binary(wal_path) do - snapshot = %{ - "registry" => serialize_registry(registry), - "last_included_index" => last_included_index, - "last_included_term" => last_included_term, - "timestamp" => System.system_time(:millisecond) - } - - path = Path.join(wal_path, @snapshot_file) - tmp_path = path <> ".tmp" - - case File.write(tmp_path, Jason.encode!(snapshot, pretty: true)) do - :ok -> - File.rename(tmp_path, path) - # Truncate WAL to only contain entries after the snapshot - truncate_wal(wal_path, last_included_index) - {:error, reason} -> - Logger.error("KRaft WAL: failed to save snapshot: #{inspect(reason)}") - {:error, reason} - end - end - def save_snapshot(nil, _registry, _index, _term), do: :ok - - @doc """ - Recover Raft state from the WAL directory. - - Returns `{:ok, recovered_state}` where recovered_state contains: - - `:current_term` — persisted term - - `:voted_for` — persisted vote - - `:log` — recovered log entries (after snapshot) - - `:registry` — snapshot registry state (or empty) - - `:snapshot_index` — last included index from snapshot - - `:snapshot_term` — last included term from snapshot - - Returns `{:ok, nil}` if no WAL exists (fresh start). - """ - def recover(wal_path) when is_binary(wal_path) do - if File.dir?(wal_path) do - state = recover_state(wal_path) - {registry, snap_index, snap_term} = recover_snapshot(wal_path) - log = recover_log(wal_path, snap_index) - - {:ok, %{ - current_term: state[:current_term] || 0, - voted_for: state[:voted_for], - log: log, - registry: registry, - snapshot_index: snap_index, - snapshot_term: snap_term - }} - else - {:ok, nil} - end - end - def recover(nil), do: {:ok, nil} - - # --------------------------------------------------------------------------- - # Private: Recovery - # --------------------------------------------------------------------------- - - defp recover_state(wal_path) do - path = Path.join(wal_path, @state_file) - - case File.read(path) do - {:ok, data} -> - case Jason.decode(data) do - {:ok, %{"current_term" => term, "voted_for" => voted_for}} -> - %{current_term: term, voted_for: voted_for} - _ -> - %{} - end - {:error, _} -> - %{} - end - end - - defp recover_snapshot(wal_path) do - path = Path.join(wal_path, @snapshot_file) - - case File.read(path) do - {:ok, data} -> - case Jason.decode(data) do - {:ok, %{"registry" => reg, "last_included_index" => idx, "last_included_term" => term}} -> - {deserialize_registry(reg), idx, term} - _ -> - {nil, 0, 0} - end - {:error, _} -> - {nil, 0, 0} - end - end - - defp recover_log(wal_path, after_index) do - path = Path.join(wal_path, @wal_file) - - case File.read(path) do - {:ok, data} -> - data - |> String.split("\n", trim: true) - |> Enum.flat_map(fn line -> - case Jason.decode(line) do - {:ok, entry_map} -> - entry = deserialize_entry(entry_map) - if entry.index > after_index, do: [entry], else: [] - _ -> - Logger.warning("KRaft WAL: skipping corrupt line: #{String.slice(line, 0, 100)}") - [] - end - end) - {:error, _} -> - [] - end - end - - @doc """ - Truncate the WAL, keeping only entries with index <= `keep_up_to_index`. - - Used when a follower's log is truncated due to receiving conflicting entries - from a new leader. After calling this, `append_entries/2` can be used to - write the correct entries from the leader. - """ - def truncate_after(wal_path, keep_up_to_index) when is_binary(wal_path) do - path = Path.join(wal_path, @wal_file) - - case File.read(path) do - {:ok, data} -> - remaining = data - |> String.split("\n", trim: true) - |> Enum.filter(fn line -> - case Jason.decode(line) do - {:ok, %{"index" => idx}} -> idx <= keep_up_to_index - _ -> false - end - end) - |> Enum.join("\n") - - remaining = if remaining != "", do: remaining <> "\n", else: "" - File.write(path, remaining, [:sync]) - - {:error, _} -> :ok - end - end - def truncate_after(nil, _index), do: :ok - - # --------------------------------------------------------------------------- - # Private: Truncation (used by snapshotting) - # --------------------------------------------------------------------------- - - defp truncate_wal(wal_path, up_to_index) do - path = Path.join(wal_path, @wal_file) - - case File.read(path) do - {:ok, data} -> - remaining = data - |> String.split("\n", trim: true) - |> Enum.filter(fn line -> - case Jason.decode(line) do - {:ok, %{"index" => idx}} -> idx > up_to_index - _ -> false - end - end) - |> Enum.join("\n") - - remaining = if remaining != "", do: remaining <> "\n", else: "" - File.write(path, remaining) - - {:error, _} -> :ok - end - end - - # --------------------------------------------------------------------------- - # Private: Serialization - # --------------------------------------------------------------------------- - - defp serialize_entry(entry) do - %{ - "term" => entry.term, - "index" => entry.index, - "command" => serialize_command(entry.command), - "timestamp" => entry[:timestamp] || System.system_time(:millisecond) - } - end - - defp serialize_command(:noop), do: %{"type" => "noop"} - defp serialize_command({:register_store, store_id, endpoint, modalities}) do - %{"type" => "register_store", "store_id" => store_id, - "endpoint" => endpoint, "modalities" => modalities} - end - defp serialize_command({:unregister_store, store_id}) do - %{"type" => "unregister_store", "store_id" => store_id} - end - defp serialize_command({:map_octad, octad_id, locations}) do - %{"type" => "map_octad", "octad_id" => octad_id, - "locations" => locations} - end - defp serialize_command({:unmap_octad, octad_id}) do - %{"type" => "unmap_octad", "octad_id" => octad_id} - end - defp serialize_command({:update_trust, store_id, new_trust}) do - %{"type" => "update_trust", "store_id" => store_id, - "new_trust" => new_trust} - end - defp serialize_command(other) do - %{"type" => "unknown", "data" => inspect(other)} - end - - defp deserialize_entry(map) do - %{ - term: map["term"], - index: map["index"], - command: deserialize_command(map["command"]), - timestamp: map["timestamp"] - } - end - - defp deserialize_command(%{"type" => "noop"}), do: :noop - defp deserialize_command(%{"type" => "register_store"} = cmd) do - {:register_store, cmd["store_id"], cmd["endpoint"], cmd["modalities"]} - end - defp deserialize_command(%{"type" => "unregister_store"} = cmd) do - {:unregister_store, cmd["store_id"]} - end - defp deserialize_command(%{"type" => "map_octad"} = cmd) do - {:map_octad, cmd["octad_id"], cmd["locations"]} - end - defp deserialize_command(%{"type" => "unmap_octad"} = cmd) do - {:unmap_octad, cmd["octad_id"]} - end - defp deserialize_command(%{"type" => "update_trust"} = cmd) do - {:update_trust, cmd["store_id"], cmd["new_trust"]} - end - defp deserialize_command(_), do: :noop - - defp serialize_registry(registry) do - %{ - "stores" => Map.new(registry[:stores] || %{}, fn {k, v} -> - {k, %{ - "store_id" => v[:store_id] || k, - "endpoint" => v[:endpoint], - "modalities" => v[:modalities], - "trust_level" => v[:trust_level] - }} - end), - "mappings" => Map.new(registry[:mappings] || %{}, fn {k, v} -> - {k, %{ - "octad_id" => v[:octad_id] || k, - "locations" => v[:locations], - "primary_store" => v[:primary_store] - }} - end) - } - end - - defp deserialize_registry(nil), do: nil - defp deserialize_registry(map) do - %{ - stores: Map.new(map["stores"] || %{}, fn {k, v} -> - {k, %{ - store_id: v["store_id"] || k, - endpoint: v["endpoint"], - modalities: v["modalities"] || [], - trust_level: v["trust_level"] || 1.0, - last_seen: nil, - response_time_ms: nil - }} - end), - mappings: Map.new(map["mappings"] || %{}, fn {k, v} -> - {k, %{ - octad_id: v["octad_id"] || k, - locations: v["locations"], - primary_store: v["primary_store"], - created: nil, - modified: nil - }} - end), - config: %{ - min_trust_level: 0.5, - max_store_downtime_ms: 300_000, - replication_factor: 3, - consistency_mode: :quorum - } - } - end -end diff --git a/verisimdb/elixir-orchestration/lib/verisim/drift/drift_monitor.ex b/verisimdb/elixir-orchestration/lib/verisim/drift/drift_monitor.ex deleted file mode 100644 index c23f3b24..00000000 --- a/verisimdb/elixir-orchestration/lib/verisim/drift/drift_monitor.ex +++ /dev/null @@ -1,329 +0,0 @@ -# SPDX-License-Identifier: MPL-2.0 - -defmodule VeriSim.DriftMonitor do - @moduledoc """ - Drift Monitor - Coordinates drift detection across entities. - - This GenServer monitors the overall drift state of the system and - coordinates normalization when drift exceeds thresholds. - - ## Drift Types - - - `:semantic_vector` - Embedding diverged from semantic content - - `:graph_document` - Graph structure doesn't match document - - `:temporal_consistency` - Version history inconsistencies - - `:tensor` - Tensor representation diverged - - `:schema` - Schema violations detected - - `:quality` - Overall data quality issues - - ## Thresholds - - | Drift Type | Warning | Critical | - |------------|---------|----------| - | semantic_vector | 0.3 | 0.7 | - | graph_document | 0.4 | 0.8 | - | temporal | 0.2 | 0.6 | - | tensor | 0.35 | 0.75 | - | schema | 0.1 | 0.5 | - | quality | 0.25 | 0.65 | - """ - - use GenServer - require Logger - - alias VeriSim.{EntityServer, RustClient, Telemetry} - - # State structure - defstruct [ - :drift_scores, - :entity_drift, - :last_sweep, - :pending_normalizations, - :config - ] - - # Default configuration - @default_config %{ - sweep_interval_ms: 60_000, - max_concurrent_normalizations: 10, - thresholds: %{ - semantic_vector: %{warning: 0.3, critical: 0.7}, - graph_document: %{warning: 0.4, critical: 0.8}, - temporal_consistency: %{warning: 0.2, critical: 0.6}, - tensor: %{warning: 0.35, critical: 0.75}, - schema: %{warning: 0.1, critical: 0.5}, - quality: %{warning: 0.25, critical: 0.65} - } - } - - # Client API - - def start_link(opts \\ []) do - GenServer.start_link(__MODULE__, opts, name: __MODULE__) - end - - @doc """ - Report a drift event for an entity. - """ - def report_drift(entity_id, score, drift_type \\ :quality) do - GenServer.cast(__MODULE__, {:report_drift, entity_id, score, drift_type}) - end - - @doc """ - Notify that an entity has changed. - """ - def entity_changed(entity_id) do - GenServer.cast(__MODULE__, {:entity_changed, entity_id}) - end - - @doc """ - Get the current drift status. - """ - def status do - GenServer.call(__MODULE__, :status) - end - - @doc """ - Get drift history for an entity. - """ - def entity_history(entity_id) do - GenServer.call(__MODULE__, {:entity_history, entity_id}) - end - - @doc """ - Trigger a manual drift sweep. - """ - def sweep do - GenServer.cast(__MODULE__, :sweep) - end - - # Server Callbacks - - @impl true - def init(opts) do - config = Keyword.get(opts, :config, @default_config) - - state = %__MODULE__{ - drift_scores: %{}, - entity_drift: %{}, - last_sweep: DateTime.utc_now(), - pending_normalizations: MapSet.new(), - config: config - } - - schedule_sweep(config.sweep_interval_ms) - - Logger.info("DriftMonitor started with config: #{inspect(config)}") - {:ok, state} - end - - @impl true - def handle_cast({:report_drift, entity_id, score, drift_type}, state) do - Logger.debug("Drift reported for #{entity_id}: #{drift_type} = #{score}") - - # Emit telemetry event (aggregate counter only — no entity data captured). - Telemetry.emit_drift_detected(drift_type) - - new_entity_drift = - Map.update( - state.entity_drift, - entity_id, - %{drift_type => score}, - &Map.put(&1, drift_type, score) - ) - - new_drift_scores = - Map.update(state.drift_scores, drift_type, [score], &[score | Enum.take(&1, 99)]) - - new_state = %{state | - entity_drift: new_entity_drift, - drift_scores: new_drift_scores - } - - # Check if normalization is needed - new_state = maybe_trigger_normalization(new_state, entity_id, score, drift_type) - - {:noreply, new_state} - end - - @impl true - def handle_cast({:entity_changed, entity_id}, state) do - # Mark entity for drift check on next sweep - Logger.debug("Entity #{entity_id} changed, marking for drift check") - {:noreply, state} - end - - @impl true - def handle_cast(:sweep, state) do - new_state = perform_sweep(state) - {:noreply, new_state} - end - - @impl true - def handle_call(:status, _from, state) do - status = %{ - overall_health: calculate_overall_health(state), - drift_by_type: summarize_drift_by_type(state), - entities_with_drift: map_size(state.entity_drift), - pending_normalizations: MapSet.size(state.pending_normalizations), - last_sweep: state.last_sweep - } - {:reply, status, state} - end - - @impl true - def handle_call({:entity_history, entity_id}, _from, state) do - history = Map.get(state.entity_drift, entity_id, %{}) - {:reply, history, state} - end - - @impl true - def handle_info(:sweep, state) do - new_state = perform_sweep(state) - schedule_sweep(state.config.sweep_interval_ms) - {:noreply, new_state} - end - - @impl true - def handle_info({:normalization_complete, entity_id, result}, state) do - Logger.info("Normalization complete for #{entity_id}: #{result}") - - new_pending = MapSet.delete(state.pending_normalizations, entity_id) - - new_entity_drift = - if result == :success do - Map.delete(state.entity_drift, entity_id) - else - state.entity_drift - end - - {:noreply, %{state | - pending_normalizations: new_pending, - entity_drift: new_entity_drift - }} - end - - # Private Functions - - defp schedule_sweep(interval_ms) do - Process.send_after(self(), :sweep, interval_ms) - end - - defp perform_sweep(state) do - Logger.debug("Performing drift sweep across #{map_size(state.entity_drift)} entities") - - # Check each entity that has reported drift - new_state = state.entity_drift - |> Enum.reduce(state, fn {entity_id, drift_scores}, acc -> - # Check each drift type for this entity - Enum.reduce(drift_scores, acc, fn {drift_type, score}, inner_acc -> - # Re-evaluate thresholds (scores may have changed since last check) - maybe_trigger_normalization(inner_acc, entity_id, score, drift_type) - end) - end) - - # Pull aggregate drift metrics from Rust core - new_state = - case RustClient.drift_status() do - {:ok, rust_metrics} when is_list(rust_metrics) -> - updated_scores = - Enum.reduce(rust_metrics, new_state.drift_scores, fn metric, acc -> - drift_type = - case metric["drift_type"] do - "SemanticVectorDrift" -> :semantic_vector - "GraphDocumentDrift" -> :graph_document - "TemporalConsistencyDrift" -> :temporal_consistency - "TensorDrift" -> :tensor - "SchemaDrift" -> :schema - "QualityDrift" -> :quality - other -> - Logger.debug("Unknown drift type from Rust core: #{inspect(other)}") - nil - end - - if drift_type do - score = metric["current_score"] || 0.0 - Map.update(acc, drift_type, [score], &[score | Enum.take(&1, 99)]) - else - acc - end - end) - - %{new_state | drift_scores: updated_scores} - - {:ok, _unexpected} -> - # Non-list response (e.g. HTML from wrong server on port 8080) - Logger.debug("DriftMonitor: Rust core returned unexpected response, skipping") - new_state - - {:error, reason} -> - Logger.warning("DriftMonitor: failed to fetch Rust core drift status: #{inspect(reason)}") - new_state - end - - %{new_state | last_sweep: DateTime.utc_now()} - end - - defp maybe_trigger_normalization(state, entity_id, score, drift_type) do - thresholds = state.config.thresholds[drift_type] || %{warning: 0.5, critical: 0.8} - - cond do - score >= thresholds.critical and entity_id not in state.pending_normalizations -> - Logger.warning("Critical drift for #{entity_id}, triggering normalization") - trigger_normalization(state, entity_id) - - score >= thresholds.warning -> - Logger.info("Warning drift for #{entity_id}: #{score}") - state - - true -> - state - end - end - - defp trigger_normalization(state, entity_id) do - if MapSet.size(state.pending_normalizations) < state.config.max_concurrent_normalizations do - # Start async normalization - Task.start(fn -> - result = case EntityServer.normalize(entity_id) do - :ok -> :success - _ -> :failure - end - send(__MODULE__, {:normalization_complete, entity_id, result}) - end) - - %{state | pending_normalizations: MapSet.put(state.pending_normalizations, entity_id)} - else - Logger.warning("Max concurrent normalizations reached, deferring #{entity_id}") - state - end - end - - defp calculate_overall_health(state) do - if map_size(state.entity_drift) == 0 do - :healthy - else - max_score = - state.entity_drift - |> Map.values() - |> Enum.flat_map(&Map.values/1) - |> Enum.max(fn -> 0.0 end) - - cond do - max_score >= 0.8 -> :critical - max_score >= 0.5 -> :degraded - max_score >= 0.3 -> :warning - true -> :healthy - end - end - end - - defp summarize_drift_by_type(state) do - state.drift_scores - |> Enum.map(fn {type, scores} -> - avg = if length(scores) > 0, do: Enum.sum(scores) / length(scores), else: 0.0 - {type, %{average: avg, max: Enum.max(scores, fn -> 0.0 end), count: length(scores)}} - end) - |> Map.new() - end -end diff --git a/verisimdb/elixir-orchestration/lib/verisim/entity/entity_server.ex b/verisimdb/elixir-orchestration/lib/verisim/entity/entity_server.ex deleted file mode 100644 index 4324e3ef..00000000 --- a/verisimdb/elixir-orchestration/lib/verisim/entity/entity_server.ex +++ /dev/null @@ -1,252 +0,0 @@ -# SPDX-License-Identifier: MPL-2.0 - -defmodule VeriSim.EntityServer do - @moduledoc """ - GenServer representing a single Octad entity. - - Each entity has its own process for isolation and fault tolerance. - The EntityServer coordinates operations across all 6 modalities. - - ## Process-per-Entity Model - - Benefits: - - Isolation: Entity failures don't cascade - - Concurrency: Entities can be modified in parallel - - State: Entity state is encapsulated - - Supervision: OTP handles restarts - - ## State Structure - - ```elixir - %{ - id: "entity-uuid", - status: :active | :normalizing | :stale, - modalities: %{ - graph: true | false, - vector: true | false, - tensor: true | false, - semantic: true | false, - document: true | false, - temporal: true | false - }, - version: 1, - last_modified: ~U[2026-01-16 00:00:00Z], - drift_score: 0.0 - } - ``` - """ - - use GenServer - require Logger - - alias VeriSim.{DriftMonitor, RustClient} - - # Client API - - @doc """ - Start a new EntityServer for the given entity ID. - """ - def start_link(entity_id) do - GenServer.start_link(__MODULE__, entity_id, name: via_tuple(entity_id)) - end - - @doc """ - Get the current state of an entity. - """ - def get(entity_id) do - GenServer.call(via_tuple(entity_id), :get) - end - - @doc """ - Update the entity with new data. - """ - def update(entity_id, changes) do - GenServer.call(via_tuple(entity_id), {:update, changes}) - end - - @doc """ - Trigger normalization for the entity. - """ - def normalize(entity_id) do - GenServer.cast(via_tuple(entity_id), :normalize) - end - - @doc """ - Get the modality status for an entity. - """ - def modality_status(entity_id) do - GenServer.call(via_tuple(entity_id), :modality_status) - end - - # Server Callbacks - - @impl true - def init(entity_id) do - Logger.info("Starting EntityServer for #{entity_id}") - - state = %{ - id: entity_id, - status: :active, - modalities: %{ - graph: false, - vector: false, - tensor: false, - semantic: false, - document: false, - temporal: false - }, - version: 0, - last_modified: DateTime.utc_now(), - drift_score: 0.0 - } - - # Schedule periodic drift check - schedule_drift_check() - - {:ok, state} - end - - @impl true - def handle_call(:get, _from, state) do - {:reply, {:ok, state}, state} - end - - @impl true - def handle_call({:update, changes}, _from, state) do - new_state = - state - |> apply_changes(changes) - |> bump_version() - |> update_timestamp() - - # Notify drift monitor of change - DriftMonitor.entity_changed(state.id) - - {:reply, {:ok, new_state}, new_state} - end - - @impl true - def handle_call(:modality_status, _from, state) do - {:reply, {:ok, state.modalities}, state} - end - - @impl true - def handle_cast(:normalize, state) do - Logger.info("Starting normalization for #{state.id}") - - # Snapshot current state via temporal store before normalization - RustClient.post("/octads/#{state.id}/versions", %{ - version: state.version, - modalities: state.modalities, - drift_score: state.drift_score, - timestamp: DateTime.to_iso8601(DateTime.utc_now()) - }) - - new_state = %{state | status: :normalizing} - - # Determine which modalities have drifted and need normalization - drifted_modalities = - case RustClient.get_drift_score(state.id) do - {:ok, score} when is_map(score) -> - score - |> Enum.filter(fn {_k, v} -> is_number(v) and v > 0.3 end) - |> Enum.map(fn {k, _v} -> k end) - - _ -> - # If we can't determine drift per-modality, normalize all - [:graph, :vector, :tensor, :semantic, :document, :temporal] - end - - # Trigger async normalization via Rust core - server_pid = self() - - Task.start(fn -> - case RustClient.normalize(state.id) do - {:ok, _result} -> - send(server_pid, {:normalization_complete, :success, drifted_modalities}) - - {:error, reason} -> - Logger.error("Normalization failed for #{state.id}: #{inspect(reason)}") - send(server_pid, {:normalization_complete, :failure, drifted_modalities}) - end - end) - - {:noreply, new_state} - end - - @impl true - def handle_info(:check_drift, state) do - new_state = - case RustClient.get_drift_score(state.id) do - {:ok, score} -> - if score > 0.3 do - Logger.warning("High drift detected for #{state.id}: #{score}") - DriftMonitor.report_drift(state.id, score) - end - %{state | drift_score: score} - {:error, _} -> - state - end - - schedule_drift_check() - {:noreply, new_state} - end - - @impl true - def handle_info({:normalization_complete, result, modalities}, state) do - new_status = if result == :success, do: :active, else: :stale - - Logger.info( - "Normalization #{result} for #{state.id}, modalities: #{inspect(modalities)}" - ) - - new_state = - %{state | status: new_status} - |> bump_version() - |> update_timestamp() - - {:noreply, new_state} - end - - # Backwards-compatible catch for old-format messages - @impl true - def handle_info({:normalization_complete, result}, state) do - new_status = if result == :success, do: :active, else: :stale - - new_state = - %{state | status: new_status} - |> update_timestamp() - - {:noreply, new_state} - end - - # Private Functions - - defp via_tuple(entity_id) do - {:via, Registry, {VeriSim.EntityRegistry, entity_id}} - end - - defp apply_changes(state, changes) do - Enum.reduce(changes, state, fn - {:modality, modality, value}, acc -> - put_in(acc, [:modalities, modality], value) - {key, value}, acc when is_atom(key) -> - Map.put(acc, key, value) - _, acc -> - acc - end) - end - - defp bump_version(state) do - %{state | version: state.version + 1} - end - - defp update_timestamp(state) do - %{state | last_modified: DateTime.utc_now()} - end - - defp schedule_drift_check do - # Check drift every 30 seconds - Process.send_after(self(), :check_drift, 30_000) - end -end diff --git a/verisimdb/elixir-orchestration/lib/verisim/federation/adapter.ex b/verisimdb/elixir-orchestration/lib/verisim/federation/adapter.ex deleted file mode 100644 index 760892b1..00000000 --- a/verisimdb/elixir-orchestration/lib/verisim/federation/adapter.ex +++ /dev/null @@ -1,282 +0,0 @@ -# SPDX-License-Identifier: MPL-2.0 - -defmodule VeriSim.Federation.Adapter do - @moduledoc """ - Behaviour for heterogeneous federation adapters. - - VeriSimDB federates across heterogeneous databases — each peer in the - federation may be a different database system (ArangoDB, PostgreSQL, - Elasticsearch, or another VeriSimDB instance). Adapters translate between - VeriSimDB's octad modality model and each backend's native capabilities. - - ## Implementing an Adapter - - Each adapter must implement 5 callbacks: - - - `connect/1` — Establish or verify connectivity to the backend. - - `query/3` — Translate a VeriSimDB modality query into the backend's - native query language and return normalised results. - - `health_check/1` — Check backend availability and return latency. - - `supported_modalities/1` — Declare which VeriSimDB modalities the - backend can serve (e.g., PostgreSQL with pgvector supports `:vector`). - - `translate_results/2` — Normalise backend-specific results into - VeriSimDB's `FederationResult` format. - - ## Modality Mapping - - Adapters map VeriSimDB's 8 octad modalities to native capabilities: - - ┌─────────────┬──────────────┬─────────────────┬───────────────┐ - │ Modality │ ArangoDB │ PostgreSQL │ Elasticsearch │ - ├─────────────┼──────────────┼─────────────────┼───────────────┤ - │ :graph │ AQL traversal│ recursive CTE │ — │ - │ :vector │ — │ pgvector <=> │ dense_vector │ - │ :tensor │ — │ — │ — │ - │ :semantic │ document │ JSONB │ nested object │ - │ :document │ fulltext idx │ tsvector/GIN │ full-text │ - │ :temporal │ document │ tstzrange │ date_range │ - │ :provenance │ edge coll. │ audit table │ — │ - │ :spatial │ GeoJSON idx │ PostGIS │ geo_shape │ - └─────────────┴──────────────┴─────────────────┴───────────────┘ - - Extended adapters (10 additional backends): - - ┌─────────────┬──────────┬───────┬────────┬────────────┬───────────┐ - │ Modality │ MongoDB │ Redis │ DuckDB │ ClickHouse │ SurrealDB │ - ├─────────────┼──────────┼───────┼────────┼────────────┼───────────┤ - │ :graph │ $graphLkp│ Graph │ rCTE │ — │ native │ - │ :vector │ AtlasVS │ VSS │ HNSW │ annoy │ — │ - │ :tensor │ — │ — │ array │ — │ — │ - │ :semantic │ BSON │ JSON │ JSON │ JSON │ schema- │ - │ :document │ text idx │ FT │ FTS │ fulltext │ FTS │ - │ :temporal │ ISODate │ TS │ tstamp │ DateTime64 │ datetime │ - │ :provenance │ chg strm │ Strm │ — │ — │ — │ - │ :spatial │ 2dsphere │ — │ spat │ geo funcs │ — │ - └─────────────┴──────────┴───────┴────────┴────────────┴───────────┘ - - ┌─────────────┬──────────┬──────────┬──────────┬──────────────┐ - │ Modality │ SQLite │ Neo4j │ VectorDB │ InfluxDB │ - ├─────────────┼──────────┼──────────┼──────────┼──────────────┤ - │ :graph │ rCTE │ Cypher │ — │ — │ - │ :vector │ vss │ vec idx │ native │ — │ - │ :tensor │ — │ — │ — │ — │ - │ :semantic │ JSON1 │ props │ payload │ tags │ - │ :document │ FTS5 │ FT idx │ — │ — │ - │ :temporal │ datetime │ temporal │ ts filt │ native TS │ - │ :provenance │ — │ — │ — │ — │ - │ :spatial │ — │ spatial │ geo filt │ — │ - └─────────────┴──────────┴──────────┴──────────┴──────────────┘ - - ObjectStorage (MinIO/S3): :document, :temporal, :provenance, :semantic - - ## Peer Configuration - - When registering a peer, the `adapter_type` field selects the adapter: - - VeriSim.Federation.Resolver.register_peer("arango-prod", %{ - endpoint: "http://arango.internal:8529", - adapter_type: :arangodb, - adapter_config: %{ - database: "_system", - collection: "octads", - auth: {:basic, "root", "password"} - }, - modalities: [:graph, :document, :semantic, :spatial] - }) - - ## Result Format - - All adapters must return results conforming to `federation_result()`: - - %{ - source_store: "peer-id", - octad_id: "entity-id-or-equivalent", - score: 0.85, - drifted: false, - data: %{...}, # Raw data from the backend - response_time_ms: 42 - } - """ - - # --------------------------------------------------------------------------- - # Types - # --------------------------------------------------------------------------- - - @typedoc "Adapter configuration passed during peer registration." - @type adapter_config :: %{ - optional(:database) => String.t(), - optional(:collection) => String.t(), - optional(:index) => String.t(), - optional(:table) => String.t(), - optional(:schema) => String.t(), - optional(:auth) => auth(), - optional(atom()) => term() - } - - @typedoc "Authentication credentials for the backend." - @type auth :: - {:basic, String.t(), String.t()} - | {:bearer, String.t()} - | {:api_key, String.t()} - | :none - - @typedoc "Peer connection details passed to adapter callbacks." - @type peer_info :: %{ - store_id: String.t(), - endpoint: String.t(), - adapter_config: adapter_config() - } - - @typedoc "A VeriSimDB modality that can be queried." - @type modality :: - :graph - | :vector - | :tensor - | :semantic - | :document - | :temporal - | :provenance - | :spatial - - @typedoc "Query parameters passed to the adapter." - @type query_params :: %{ - required(:modalities) => [modality()], - required(:limit) => non_neg_integer(), - optional(:text_query) => String.t(), - optional(:vector_query) => [float()], - optional(:graph_pattern) => String.t(), - optional(:spatial_bounds) => map(), - optional(:temporal_range) => map(), - optional(:filters) => map() - } - - @typedoc "Normalised result from a federated query." - @type federation_result :: %{ - source_store: String.t(), - octad_id: String.t(), - score: float(), - drifted: boolean(), - data: map(), - response_time_ms: non_neg_integer() - } - - @typedoc "Health check result with latency measurement." - @type health_result :: - {:ok, non_neg_integer()} - | {:error, term()} - - # --------------------------------------------------------------------------- - # Callbacks - # --------------------------------------------------------------------------- - - @doc """ - Verify connectivity to the backend. - - Called during peer registration and periodically during health checks. - Should return `:ok` if the backend is reachable and properly configured, - or `{:error, reason}` otherwise. - """ - @callback connect(peer_info()) :: :ok | {:error, term()} - - @doc """ - Execute a query against the backend. - - Translates VeriSimDB's modality-based query into the backend's native - query language (AQL, SQL, Elasticsearch DSL, etc.), executes it, and - returns normalised results. - - The `query_params` map contains the requested modalities and any - modality-specific query parameters (text queries, vector embeddings, - graph traversal patterns, spatial bounds, etc.). - - Results must be normalised to `[federation_result()]` format. - """ - @callback query(peer_info(), query_params(), keyword()) :: - {:ok, [federation_result()]} | {:error, term()} - - @doc """ - Check backend health and return response latency in milliseconds. - - Used by the federation resolver to update peer trust levels. - A successful health check increases trust; failures decrease it. - """ - @callback health_check(peer_info()) :: health_result() - - @doc """ - Declare which VeriSimDB modalities this backend can serve. - - Called during peer registration to validate that the peer's declared - modalities are actually supported by the adapter. Returns the subset - of modalities the backend can serve. - - For example, a PostgreSQL instance with pgvector and PostGIS installed - might return `[:document, :vector, :spatial, :semantic, :temporal]`. - """ - @callback supported_modalities(adapter_config()) :: [modality()] - - @doc """ - Normalise backend-specific results into VeriSimDB's federation format. - - This is called internally by `query/3` but is exposed as a callback - so that custom result transforms can be tested independently. - """ - @callback translate_results([map()], peer_info()) :: [federation_result()] - - # --------------------------------------------------------------------------- - # Adapter Registry - # --------------------------------------------------------------------------- - - @doc """ - Look up the adapter module for a given adapter type atom. - - ## Examples - - iex> VeriSim.Federation.Adapter.module_for(:verisimdb) - {:ok, VeriSim.Federation.Adapters.VeriSimDB} - - iex> VeriSim.Federation.Adapter.module_for(:arangodb) - {:ok, VeriSim.Federation.Adapters.ArangoDB} - - iex> VeriSim.Federation.Adapter.module_for(:unknown) - {:error, :unknown_adapter} - """ - @spec module_for(atom()) :: {:ok, module()} | {:error, :unknown_adapter} - def module_for(:verisimdb), do: {:ok, VeriSim.Federation.Adapters.VeriSimDB} - def module_for(:arangodb), do: {:ok, VeriSim.Federation.Adapters.ArangoDB} - def module_for(:postgresql), do: {:ok, VeriSim.Federation.Adapters.PostgreSQL} - def module_for(:elasticsearch), do: {:ok, VeriSim.Federation.Adapters.Elasticsearch} - def module_for(:mongodb), do: {:ok, VeriSim.Federation.Adapters.MongoDB} - def module_for(:redis), do: {:ok, VeriSim.Federation.Adapters.Redis} - def module_for(:duckdb), do: {:ok, VeriSim.Federation.Adapters.DuckDB} - def module_for(:clickhouse), do: {:ok, VeriSim.Federation.Adapters.ClickHouse} - def module_for(:surrealdb), do: {:ok, VeriSim.Federation.Adapters.SurrealDB} - def module_for(:sqlite), do: {:ok, VeriSim.Federation.Adapters.SQLite} - def module_for(:neo4j), do: {:ok, VeriSim.Federation.Adapters.Neo4j} - def module_for(:vector_db), do: {:ok, VeriSim.Federation.Adapters.VectorDB} - def module_for(:influxdb), do: {:ok, VeriSim.Federation.Adapters.InfluxDB} - def module_for(:object_storage), do: {:ok, VeriSim.Federation.Adapters.ObjectStorage} - def module_for(_), do: {:error, :unknown_adapter} - - @doc """ - List all registered adapter types. - """ - @spec adapter_types() :: [atom()] - def adapter_types do - [ - :verisimdb, - :arangodb, - :postgresql, - :elasticsearch, - :mongodb, - :redis, - :duckdb, - :clickhouse, - :surrealdb, - :sqlite, - :neo4j, - :vector_db, - :influxdb, - :object_storage - ] - end -end diff --git a/verisimdb/elixir-orchestration/lib/verisim/federation/adapters/arangodb.ex b/verisimdb/elixir-orchestration/lib/verisim/federation/adapters/arangodb.ex deleted file mode 100644 index c86684d3..00000000 --- a/verisimdb/elixir-orchestration/lib/verisim/federation/adapters/arangodb.ex +++ /dev/null @@ -1,305 +0,0 @@ -# SPDX-License-Identifier: MPL-2.0 - -defmodule VeriSim.Federation.Adapters.ArangoDB do - @moduledoc """ - Federation adapter for ArangoDB. - - ArangoDB is a multi-model database supporting documents, graphs, and - key-value access. This adapter translates VeriSimDB modality queries - into AQL (ArangoDB Query Language) and normalises results into the - federation result format. - - ## Modality Mapping - - | VeriSimDB Modality | ArangoDB Capability | Implementation | - |--------------------|---------------------------|-----------------------------| - | `:graph` | Native graph traversal | `FOR v, e IN ... GRAPH` | - | `:document` | Fulltext index (Analyzer) | `ANALYZER(... "text_en")` | - | `:semantic` | Document attributes | JSON document fields | - | `:temporal` | Document with date fields | AQL date functions | - | `:provenance` | Edge collections | `FOR e IN edges FILTER ...` | - | `:spatial` | GeoJSON index | `GEO_DISTANCE(...)` filter | - - ArangoDB does not natively support vector similarity or tensor storage, - so `:vector` and `:tensor` modalities are not supported. - - ## Configuration - - %{ - database: "_system", # ArangoDB database name - collection: "octads", # Default document collection - graph_name: "octad_graph", # Named graph for traversals - auth: {:basic, "root", "password"} - } - - ## ArangoDB HTTP API - - Queries are sent to `POST /_db/{database}/_api/cursor` with AQL. - Health checks hit `GET /_api/version`. - """ - - @behaviour VeriSim.Federation.Adapter - - require Logger - - @default_timeout 10_000 - - # --------------------------------------------------------------------------- - # Callbacks - # --------------------------------------------------------------------------- - - @impl true - def connect(peer_info) do - case health_check(peer_info) do - {:ok, _latency} -> :ok - {:error, reason} -> {:error, reason} - end - end - - @impl true - def query(peer_info, query_params, opts \\ []) do - timeout = Keyword.get(opts, :timeout, @default_timeout) - modalities = Map.get(query_params, :modalities, []) - limit = Map.get(query_params, :limit, 100) - - start = System.monotonic_time(:millisecond) - - {aql, bind_vars} = build_aql(modalities, query_params, limit, peer_info) - - result = execute_aql(peer_info, aql, bind_vars, timeout) - - elapsed = System.monotonic_time(:millisecond) - start - - case result do - {:ok, raw_results} -> - normalised = - raw_results - |> translate_results(peer_info) - |> Enum.map(fn r -> Map.put(r, :response_time_ms, elapsed) end) - - {:ok, normalised} - - {:error, reason} -> - {:error, reason} - end - rescue - e -> - Logger.warning( - "ArangoDB adapter: exception querying #{peer_info.store_id}: #{inspect(e)}" - ) - - {:error, {:exception, e}} - end - - @impl true - def health_check(peer_info) do - config = peer_info.adapter_config - db = Map.get(config, :database, "_system") - url = "#{peer_info.endpoint}/_db/#{db}/_api/version" - - start = System.monotonic_time(:millisecond) - headers = auth_headers(config) - - case Req.get(url, headers: headers, receive_timeout: 5_000) do - {:ok, %Req.Response{status: 200}} -> - elapsed = System.monotonic_time(:millisecond) - start - {:ok, elapsed} - - {:ok, %Req.Response{status: status}} -> - {:error, {:unhealthy, status}} - - {:error, reason} -> - {:error, reason} - end - rescue - e -> {:error, {:exception, e}} - end - - @impl true - def supported_modalities(_adapter_config) do - [:graph, :document, :semantic, :temporal, :provenance, :spatial] - end - - @impl true - def translate_results(raw_results, peer_info) do - raw_results - |> List.wrap() - |> Enum.map(fn doc -> - %{ - source_store: peer_info.store_id, - octad_id: extract_id(doc), - score: doc["_score"] || doc["score"] || 0.0, - drifted: false, - data: doc, - response_time_ms: 0 - } - end) - end - - # --------------------------------------------------------------------------- - # Private — AQL Query Builder - # --------------------------------------------------------------------------- - - defp build_aql(modalities, query_params, limit, peer_info) do - config = peer_info.adapter_config - collection = Map.get(config, :collection, "octads") - - cond do - :graph in modalities && Map.has_key?(query_params, :graph_pattern) -> - graph_name = Map.get(config, :graph_name, "octad_graph") - start_vertex = query_params.graph_pattern - - aql = """ - FOR v, e, p IN 1..3 ANY @start_vertex GRAPH @graph_name - LIMIT @limit - RETURN MERGE(v, {_edge: e, _path_length: LENGTH(p.edges)}) - """ - - bind_vars = %{ - "start_vertex" => "#{collection}/#{start_vertex}", - "graph_name" => graph_name, - "limit" => limit - } - - {aql, bind_vars} - - :document in modalities && Map.has_key?(query_params, :text_query) -> - aql = """ - FOR doc IN #{collection} - FILTER ANALYZER(LIKE(doc.title, @query) OR LIKE(doc.body, @query), "text_en") - SORT BM25(doc) DESC - LIMIT @limit - RETURN MERGE(doc, {_score: BM25(doc)}) - """ - - bind_vars = %{ - "query" => "%#{query_params.text_query}%", - "limit" => limit - } - - {aql, bind_vars} - - :spatial in modalities && Map.has_key?(query_params, :spatial_bounds) -> - bounds = query_params.spatial_bounds - - aql = """ - FOR doc IN #{collection} - FILTER GEO_CONTAINS( - GEO_POLYGON([ - [@min_lon, @min_lat], - [@max_lon, @min_lat], - [@max_lon, @max_lat], - [@min_lon, @max_lat], - [@min_lon, @min_lat] - ]), - doc.location - ) - LIMIT @limit - RETURN doc - """ - - bind_vars = %{ - "min_lat" => bounds[:min_lat] || bounds["min_lat"] || 0.0, - "min_lon" => bounds[:min_lon] || bounds["min_lon"] || 0.0, - "max_lat" => bounds[:max_lat] || bounds["max_lat"] || 0.0, - "max_lon" => bounds[:max_lon] || bounds["max_lon"] || 0.0, - "limit" => limit - } - - {aql, bind_vars} - - :temporal in modalities && Map.has_key?(query_params, :temporal_range) -> - range = query_params.temporal_range - - aql = """ - FOR doc IN #{collection} - FILTER doc.created_at >= @start_time AND doc.created_at <= @end_time - SORT doc.created_at DESC - LIMIT @limit - RETURN doc - """ - - bind_vars = %{ - "start_time" => range[:start] || range["start"] || "", - "end_time" => range[:end] || range["end"] || "", - "limit" => limit - } - - {aql, bind_vars} - - :provenance in modalities -> - edge_collection = Map.get(config, :edge_collection, "provenance_edges") - - aql = """ - FOR e IN #{edge_collection} - SORT e.timestamp DESC - LIMIT @limit - RETURN MERGE(e, { - _from_doc: DOCUMENT(e._from), - _to_doc: DOCUMENT(e._to) - }) - """ - - bind_vars = %{"limit" => limit} - {aql, bind_vars} - - true -> - # Default: return documents from the collection - aql = """ - FOR doc IN #{collection} - SORT doc._key ASC - LIMIT @limit - RETURN doc - """ - - bind_vars = %{"limit" => limit} - {aql, bind_vars} - end - end - - defp execute_aql(peer_info, aql, bind_vars, timeout) do - config = peer_info.adapter_config - db = Map.get(config, :database, "_system") - url = "#{peer_info.endpoint}/_db/#{db}/_api/cursor" - headers = auth_headers(config) - - body = %{ - "query" => aql, - "bindVars" => bind_vars, - "batchSize" => 1000 - } - - case Req.post(url, json: body, headers: headers, receive_timeout: timeout) do - {:ok, %Req.Response{status: status, body: resp_body}} when status in 200..201 -> - results = resp_body["result"] || [] - {:ok, results} - - {:ok, %Req.Response{status: status, body: resp_body}} -> - error_msg = resp_body["errorMessage"] || "HTTP #{status}" - Logger.warning("ArangoDB adapter: query failed: #{error_msg}") - {:error, {:aql_error, status, error_msg}} - - {:error, reason} -> - {:error, reason} - end - end - - defp extract_id(doc) do - # ArangoDB uses _key or _id; normalise to a plain ID - doc["_key"] || doc["_id"] || doc["id"] || "unknown" - end - - defp auth_headers(config) do - case Map.get(config, :auth, :none) do - {:basic, user, pass} -> - encoded = Base.encode64("#{user}:#{pass}") - [{"Authorization", "Basic #{encoded}"}] - - {:bearer, token} -> - [{"Authorization", "Bearer #{token}"}] - - _ -> - [] - end - end -end diff --git a/verisimdb/elixir-orchestration/lib/verisim/federation/adapters/clickhouse.ex b/verisimdb/elixir-orchestration/lib/verisim/federation/adapters/clickhouse.ex deleted file mode 100644 index ba874219..00000000 --- a/verisimdb/elixir-orchestration/lib/verisim/federation/adapters/clickhouse.ex +++ /dev/null @@ -1,339 +0,0 @@ -# SPDX-License-Identifier: MPL-2.0 - -defmodule VeriSim.Federation.Adapters.ClickHouse do - @moduledoc """ - Federation adapter for ClickHouse. - - Translates VeriSimDB modality queries into ClickHouse SQL and executes - them via the ClickHouse HTTP interface (port 8123). ClickHouse is an - OLAP columnar database optimised for analytical queries over large - datasets, with built-in support for vector operations, full-text search, - and geospatial functions. - - ## Modality Mapping - - | VeriSimDB Modality | ClickHouse Capability | Extension/Feature Required | - |--------------------|-----------------------------|------------------------------| - | `:vector` | Array distance functions | Built-in (cosineDistance) | - | `:document` | Full-text index / ngramBF | Built-in (experimental) | - | `:temporal` | DateTime64 / toInterval | Built-in | - | `:spatial` | Geo functions (pointInPolygon) | Built-in | - | `:semantic` | JSON extraction | Built-in (JSONExtract*) | - - ClickHouse does not natively support graph traversal, tensor operations, - or provenance tracking, so `:graph`, `:tensor`, and `:provenance` - modalities are not supported. - - ## Configuration - - %{ - host: "clickhouse.internal", - port: 8123, - database: "verisimdb", - table: "octads", - auth: {:basic, "default", "password"} - } - - ## ClickHouse HTTP Interface - - ClickHouse exposes a native HTTP interface on port 8123. Queries are - sent as POST body to `/?database={db}&default_format=JSONEachRow`. - The root endpoint `GET /` returns `"Ok.\\n"` for health checks. - """ - - @behaviour VeriSim.Federation.Adapter - - require Logger - - @default_timeout 10_000 - - # --------------------------------------------------------------------------- - # Callbacks - # --------------------------------------------------------------------------- - - @impl true - def connect(peer_info) do - case health_check(peer_info) do - {:ok, _latency} -> :ok - {:error, reason} -> {:error, reason} - end - end - - @impl true - def query(peer_info, query_params, opts \\ []) do - timeout = Keyword.get(opts, :timeout, @default_timeout) - modalities = Map.get(query_params, :modalities, []) - limit = Map.get(query_params, :limit, 100) - - start = System.monotonic_time(:millisecond) - - sql = build_sql(modalities, query_params, limit, peer_info) - - result = execute_sql(peer_info, sql, timeout) - - elapsed = System.monotonic_time(:millisecond) - start - - case result do - {:ok, raw_results} -> - normalised = - raw_results - |> translate_results(peer_info) - |> Enum.map(fn r -> Map.put(r, :response_time_ms, elapsed) end) - - {:ok, normalised} - - {:error, reason} -> - {:error, reason} - end - rescue - e -> - Logger.warning( - "ClickHouse adapter: exception querying #{peer_info.store_id}: #{inspect(e)}" - ) - - {:error, {:exception, e}} - end - - @impl true - def health_check(peer_info) do - start = System.monotonic_time(:millisecond) - headers = auth_headers(peer_info.adapter_config) - - # ClickHouse HTTP interface: GET / returns "Ok.\n" - url = peer_info.endpoint - - case Req.get(url, headers: headers, receive_timeout: 5_000) do - {:ok, %Req.Response{status: 200, body: body}} -> - if String.starts_with?(String.trim(to_string(body)), "Ok") do - elapsed = System.monotonic_time(:millisecond) - start - {:ok, elapsed} - else - {:error, {:unexpected_health_response, body}} - end - - {:ok, %Req.Response{status: status}} -> - {:error, {:unhealthy, status}} - - {:error, reason} -> - {:error, reason} - end - rescue - e -> {:error, {:exception, e}} - end - - @impl true - def supported_modalities(_adapter_config) do - # ClickHouse supports these modalities natively — no extensions needed - [:vector, :document, :temporal, :spatial, :semantic] - end - - @impl true - def translate_results(raw_results, peer_info) do - raw_results - |> List.wrap() - |> Enum.map(fn row -> - %{ - source_store: peer_info.store_id, - octad_id: row["id"] || row["entity_id"] || row["_key"] || "unknown", - score: parse_score(row), - drifted: false, - data: row, - response_time_ms: 0 - } - end) - end - - # --------------------------------------------------------------------------- - # Private — ClickHouse SQL Builder - # --------------------------------------------------------------------------- - - defp build_sql(modalities, query_params, limit, peer_info) do - config = peer_info.adapter_config - db = Map.get(config, :database, "verisimdb") - table = Map.get(config, :table, "octads") - qualified_table = "#{db}.#{table}" - - cond do - :vector in modalities && Map.has_key?(query_params, :vector_query) -> - # ClickHouse cosineDistance on Array(Float32) columns - embedding = query_params.vector_query - embedding_str = "[" <> Enum.join(Enum.map(embedding, &to_string/1), ", ") <> "]" - distance_fn = Map.get(config, :distance_function, "cosineDistance") - - """ - SELECT *, - 1.0 - #{distance_fn}(embedding, #{embedding_str}) AS score - FROM #{qualified_table} - ORDER BY #{distance_fn}(embedding, #{embedding_str}) ASC - LIMIT #{limit} - FORMAT JSONEachRow - """ - - :document in modalities && Map.has_key?(query_params, :text_query) -> - # ClickHouse full-text search via hasToken / multiSearchAny / ngramSearch - text = escape_clickhouse_string(query_params.text_query) - tokens = String.split(text, ~r/\s+/, trim: true) - - token_conditions = - tokens - |> Enum.map(fn token -> "hasToken(lower(content), '#{String.downcase(token)}')" end) - |> Enum.join(" AND ") - - condition = if token_conditions == "", do: "1=1", else: token_conditions - - """ - SELECT *, - ngramSearch(lower(content), '#{String.downcase(text)}') AS score - FROM #{qualified_table} - WHERE #{condition} - ORDER BY score DESC - LIMIT #{limit} - FORMAT JSONEachRow - """ - - :temporal in modalities && Map.has_key?(query_params, :temporal_range) -> - range = query_params.temporal_range - start_time = escape_clickhouse_string(range[:start] || range["start"] || "") - end_time = escape_clickhouse_string(range[:end] || range["end"] || "") - - """ - SELECT *, 0.0 AS score - FROM #{qualified_table} - WHERE created_at >= toDateTime64('#{start_time}', 3) - AND created_at <= toDateTime64('#{end_time}', 3) - ORDER BY created_at DESC - LIMIT #{limit} - FORMAT JSONEachRow - """ - - :spatial in modalities && Map.has_key?(query_params, :spatial_bounds) -> - bounds = query_params.spatial_bounds - min_lon = bounds[:min_lon] || bounds["min_lon"] || 0.0 - min_lat = bounds[:min_lat] || bounds["min_lat"] || 0.0 - max_lon = bounds[:max_lon] || bounds["max_lon"] || 0.0 - max_lat = bounds[:max_lat] || bounds["max_lat"] || 0.0 - - """ - SELECT *, 0.0 AS score - FROM #{qualified_table} - WHERE pointInPolygon( - (longitude, latitude), - [(#{min_lon}, #{min_lat}), (#{max_lon}, #{min_lat}), - (#{max_lon}, #{max_lat}), (#{min_lon}, #{max_lat})] - ) - LIMIT #{limit} - FORMAT JSONEachRow - """ - - :semantic in modalities && Map.has_key?(query_params, :filters) -> - # JSONExtract functions for structured metadata queries - filters = query_params.filters - - where_clauses = - filters - |> Enum.map(fn {field, value} -> - "JSONExtractString(metadata, '#{escape_clickhouse_string(to_string(field))}') = '#{escape_clickhouse_string(to_string(value))}'" - end) - |> Enum.join(" AND ") - - where_clause = if where_clauses == "", do: "1=1", else: where_clauses - - """ - SELECT *, 0.0 AS score - FROM #{qualified_table} - WHERE #{where_clause} - LIMIT #{limit} - FORMAT JSONEachRow - """ - - true -> - # Default: paginated listing - """ - SELECT *, 0.0 AS score - FROM #{qualified_table} - ORDER BY id ASC - LIMIT #{limit} - FORMAT JSONEachRow - """ - end - end - - defp execute_sql(peer_info, sql, timeout) do - config = peer_info.adapter_config - db = Map.get(config, :database, "verisimdb") - headers = auth_headers(config) - - # ClickHouse HTTP interface: POST with SQL as body - url = "#{peer_info.endpoint}/?database=#{db}&default_format=JSONEachRow" - - case Req.post(url, body: sql, headers: headers, receive_timeout: timeout) do - {:ok, %Req.Response{status: 200, body: body}} when is_list(body) -> - {:ok, body} - - {:ok, %Req.Response{status: 200, body: body}} when is_binary(body) -> - # JSONEachRow format: one JSON object per line - rows = parse_json_each_row(body) - {:ok, rows} - - {:ok, %Req.Response{status: 200, body: body}} when is_map(body) -> - rows = body["data"] || [body] - {:ok, rows} - - {:ok, %Req.Response{status: status, body: body}} -> - error_msg = if is_binary(body), do: String.trim(body), else: inspect(body) - Logger.warning("ClickHouse adapter: query failed (#{status}): #{error_msg}") - {:error, {:sql_error, status, error_msg}} - - {:error, reason} -> - {:error, reason} - end - end - - # --------------------------------------------------------------------------- - # Private — Helpers - # --------------------------------------------------------------------------- - - defp parse_json_each_row(body) when is_binary(body) do - body - |> String.trim() - |> String.split("\n", trim: true) - |> Enum.map(fn line -> - case Jason.decode(line) do - {:ok, row} -> row - {:error, _} -> %{"raw" => line} - end - end) - end - - defp parse_score(row) do - case row["score"] do - score when is_number(score) -> score / 1 - _ -> 0.0 - end - end - - defp escape_clickhouse_string(str) when is_binary(str) do - str - |> String.replace("\\", "\\\\") - |> String.replace("'", "\\'") - end - - defp escape_clickhouse_string(str), do: to_string(str) - - defp auth_headers(config) do - case Map.get(config, :auth, :none) do - {:basic, user, pass} -> - encoded = Base.encode64("#{user}:#{pass}") - [{"Authorization", "Basic #{encoded}"}] - - {:bearer, token} -> - [{"Authorization", "Bearer #{token}"}] - - {:api_key, key} -> - [{"X-ClickHouse-Key", key}] - - _ -> - [] - end - end -end diff --git a/verisimdb/elixir-orchestration/lib/verisim/federation/adapters/duckdb.ex b/verisimdb/elixir-orchestration/lib/verisim/federation/adapters/duckdb.ex deleted file mode 100644 index 4799c099..00000000 --- a/verisimdb/elixir-orchestration/lib/verisim/federation/adapters/duckdb.ex +++ /dev/null @@ -1,361 +0,0 @@ -# SPDX-License-Identifier: MPL-2.0 - -defmodule VeriSim.Federation.Adapters.DuckDB do - @moduledoc """ - Federation adapter for DuckDB (with HNSW, FTS, and Spatial extensions). - - Translates VeriSimDB modality queries into DuckDB SQL and executes them - via an HTTP endpoint (DuckDB Web Shell, duckdb-wasm REST proxy, or a - custom HTTP wrapper around the DuckDB C/C++ library). DuckDB excels at - analytical workloads and can query Parquet, CSV, and JSON files directly. - - ## Modality Mapping - - | VeriSimDB Modality | DuckDB Capability | Extension Required | - |--------------------|------------------------------|--------------------| - | `:graph` | Recursive CTEs | Built-in | - | `:vector` | HNSW index, array distance | hnsw | - | `:document` | Full-text search (FTS) | fts | - | `:temporal` | TIMESTAMP / INTERVAL types | Built-in | - | `:spatial` | ST_* functions | spatial | - | `:semantic` | JSON extraction | Built-in (json) | - | `:tensor` | FLOAT[] array operations | Built-in | - - The `:provenance` modality has no direct DuckDB mapping and is not - supported by this adapter (DuckDB is an analytical engine without - change tracking). - - ## Configuration - - %{ - path: "/data/verisimdb.duckdb", # Database file or :memory - extensions: [:hnsw, :fts, :spatial], - table: "octads" - } - - ## HTTP Endpoint - - DuckDB does not natively expose an HTTP API. This adapter assumes a - lightweight HTTP wrapper (e.g., a Rust/Elixir sidecar) that accepts - SQL via `POST /query` and returns JSON results. - """ - - @behaviour VeriSim.Federation.Adapter - - require Logger - - @default_timeout 10_000 - - # --------------------------------------------------------------------------- - # Callbacks - # --------------------------------------------------------------------------- - - @impl true - def connect(peer_info) do - case health_check(peer_info) do - {:ok, _latency} -> :ok - {:error, reason} -> {:error, reason} - end - end - - @impl true - def query(peer_info, query_params, opts \\ []) do - timeout = Keyword.get(opts, :timeout, @default_timeout) - modalities = Map.get(query_params, :modalities, []) - limit = Map.get(query_params, :limit, 100) - - start = System.monotonic_time(:millisecond) - - {sql, params} = build_sql(modalities, query_params, limit, peer_info) - - result = execute_sql(peer_info, sql, params, timeout) - - elapsed = System.monotonic_time(:millisecond) - start - - case result do - {:ok, raw_results} -> - normalised = - raw_results - |> translate_results(peer_info) - |> Enum.map(fn r -> Map.put(r, :response_time_ms, elapsed) end) - - {:ok, normalised} - - {:error, reason} -> - {:error, reason} - end - rescue - e -> - Logger.warning( - "DuckDB adapter: exception querying #{peer_info.store_id}: #{inspect(e)}" - ) - - {:error, {:exception, e}} - end - - @impl true - def health_check(peer_info) do - start = System.monotonic_time(:millisecond) - - # DuckDB health check: execute a trivial query - result = execute_sql(peer_info, "SELECT 1 AS ok", %{}, 5_000) - - case result do - {:ok, _} -> - elapsed = System.monotonic_time(:millisecond) - start - {:ok, elapsed} - - {:error, reason} -> - Logger.warning( - "DuckDB adapter: health check failed for #{peer_info.store_id}: #{inspect(reason)}" - ) - - {:error, reason} - end - rescue - e -> {:error, {:exception, e}} - end - - @impl true - def supported_modalities(adapter_config) do - extensions = Map.get(adapter_config, :extensions, []) - - # DuckDB always supports: recursive CTEs (graph), timestamps (temporal), - # JSON (semantic), and FLOAT[] arrays (tensor) - base = [:graph, :temporal, :semantic, :tensor] - - base - |> maybe_add(:vector, :hnsw in extensions) - |> maybe_add(:document, :fts in extensions) - |> maybe_add(:spatial, :spatial in extensions) - end - - @impl true - def translate_results(raw_results, peer_info) do - raw_results - |> List.wrap() - |> Enum.map(fn row -> - %{ - source_store: peer_info.store_id, - octad_id: row["id"] || row["entity_id"] || row["_key"] || "unknown", - score: parse_score(row), - drifted: false, - data: row, - response_time_ms: 0 - } - end) - end - - # --------------------------------------------------------------------------- - # Private — DuckDB SQL Query Builder - # --------------------------------------------------------------------------- - - defp build_sql(modalities, query_params, limit, peer_info) do - config = peer_info.adapter_config - table = Map.get(config, :table, "octads") - - cond do - :vector in modalities && Map.has_key?(query_params, :vector_query) -> - # HNSW extension: array_distance for cosine similarity - embedding = query_params.vector_query - embedding_str = "[" <> Enum.join(Enum.map(embedding, &to_string/1), ", ") <> "]" - distance_metric = Map.get(config, :distance_metric, "cosine") - - sql = """ - SELECT *, - 1.0 - array_distance(embedding, $1::FLOAT[#{length(embedding)}], '#{distance_metric}') AS score - FROM #{table} - ORDER BY array_distance(embedding, $1::FLOAT[#{length(embedding)}], '#{distance_metric}') ASC - LIMIT $2 - """ - - {sql, %{"$1" => embedding_str, "$2" => limit}} - - :document in modalities && Map.has_key?(query_params, :text_query) -> - # FTS extension: fts_main_{table}.match_bm25 - text = query_params.text_query - fts_index = "fts_main_#{table}" - - sql = """ - SELECT *, - fts_main_#{table}.match_bm25(id, $1, fields := 'title, body, content') AS score - FROM #{table} - WHERE score IS NOT NULL - ORDER BY score DESC - LIMIT $2 - """ - - {sql, %{"$1" => text, "$2" => limit}} - - :spatial in modalities && Map.has_key?(query_params, :spatial_bounds) -> - # Spatial extension: ST_Within / ST_MakeEnvelope - bounds = query_params.spatial_bounds - - sql = """ - SELECT *, - 0.0 AS score - FROM #{table} - WHERE ST_Within( - geom, - ST_MakeEnvelope($1, $2, $3, $4) - ) - LIMIT $5 - """ - - {sql, %{ - "$1" => bounds[:min_lon] || bounds["min_lon"] || 0.0, - "$2" => bounds[:min_lat] || bounds["min_lat"] || 0.0, - "$3" => bounds[:max_lon] || bounds["max_lon"] || 0.0, - "$4" => bounds[:max_lat] || bounds["max_lat"] || 0.0, - "$5" => limit - }} - - :graph in modalities && Map.has_key?(query_params, :graph_pattern) -> - # Recursive CTE for graph traversal - edges_table = Map.get(config, :edges_table, "edges") - - sql = """ - WITH RECURSIVE traversal AS ( - SELECT id, 1 AS depth - FROM #{table} - WHERE id = $1 - - UNION ALL - - SELECT e.target_id, t.depth + 1 - FROM #{edges_table} e - INNER JOIN traversal t ON e.source_id = t.id - WHERE t.depth < 3 - ) - SELECT h.*, 0.0 AS score - FROM #{table} h - INNER JOIN traversal t ON h.id = t.id - LIMIT $2 - """ - - {sql, %{"$1" => query_params.graph_pattern, "$2" => limit}} - - :temporal in modalities && Map.has_key?(query_params, :temporal_range) -> - range = query_params.temporal_range - - sql = """ - SELECT *, 0.0 AS score - FROM #{table} - WHERE created_at >= CAST($1 AS TIMESTAMP) - AND created_at <= CAST($2 AS TIMESTAMP) - ORDER BY created_at DESC - LIMIT $3 - """ - - {sql, %{ - "$1" => range[:start] || range["start"] || "", - "$2" => range[:end] || range["end"] || "", - "$3" => limit - }} - - :tensor in modalities && Map.has_key?(query_params, :vector_query) -> - # DuckDB array operations for tensor similarity - embedding = query_params.vector_query - embedding_str = "[" <> Enum.join(Enum.map(embedding, &to_string/1), ", ") <> "]" - - sql = """ - SELECT *, - list_cosine_similarity(tensor_data, $1::FLOAT[]) AS score - FROM #{table} - WHERE tensor_data IS NOT NULL - ORDER BY score DESC - LIMIT $2 - """ - - {sql, %{"$1" => embedding_str, "$2" => limit}} - - :semantic in modalities && Map.has_key?(query_params, :filters) -> - # JSON extraction via DuckDB's built-in JSON support - filters = query_params.filters - - where_clauses = - filters - |> Enum.map(fn {field, value} -> - "json_extract_string(metadata, '$.#{field}') = '#{value}'" - end) - |> Enum.join(" AND ") - - where_clause = if where_clauses == "", do: "1=1", else: where_clauses - - sql = """ - SELECT *, 0.0 AS score - FROM #{table} - WHERE #{where_clause} - LIMIT $1 - """ - - {sql, %{"$1" => limit}} - - true -> - # Default: paginated listing - sql = """ - SELECT *, 0.0 AS score - FROM #{table} - ORDER BY id ASC - LIMIT $1 - """ - - {sql, %{"$1" => limit}} - end - end - - defp execute_sql(peer_info, sql, params, timeout) do - url = "#{peer_info.endpoint}/query" - headers = auth_headers(peer_info.adapter_config) - - body = %{ - "query" => sql, - "params" => params - } - - case Req.post(url, json: body, headers: headers, receive_timeout: timeout) do - {:ok, %Req.Response{status: status, body: resp_body}} when status in 200..299 -> - rows = resp_body["rows"] || resp_body["result"] || resp_body["data"] || resp_body - {:ok, List.wrap(rows)} - - {:ok, %Req.Response{status: status, body: resp_body}} -> - error_msg = resp_body["error"] || resp_body["message"] || "HTTP #{status}" - Logger.warning("DuckDB adapter: query failed: #{error_msg}") - {:error, {:sql_error, status, error_msg}} - - {:error, reason} -> - {:error, reason} - end - end - - # --------------------------------------------------------------------------- - # Private — Helpers - # --------------------------------------------------------------------------- - - defp parse_score(row) do - case row["score"] do - score when is_number(score) -> score / 1 - _ -> 0.0 - end - end - - defp maybe_add(list, item, true), do: list ++ [item] - defp maybe_add(list, _item, false), do: list - - defp auth_headers(config) do - case Map.get(config, :auth, :none) do - {:basic, user, pass} -> - encoded = Base.encode64("#{user}:#{pass}") - [{"Authorization", "Basic #{encoded}"}, {"Content-Type", "application/json"}] - - {:bearer, token} -> - [{"Authorization", "Bearer #{token}"}, {"Content-Type", "application/json"}] - - {:api_key, key} -> - [{"X-API-Key", key}, {"Content-Type", "application/json"}] - - _ -> - [{"Content-Type", "application/json"}] - end - end -end diff --git a/verisimdb/elixir-orchestration/lib/verisim/federation/adapters/elasticsearch.ex b/verisimdb/elixir-orchestration/lib/verisim/federation/adapters/elasticsearch.ex deleted file mode 100644 index de357145..00000000 --- a/verisimdb/elixir-orchestration/lib/verisim/federation/adapters/elasticsearch.ex +++ /dev/null @@ -1,322 +0,0 @@ -# SPDX-License-Identifier: MPL-2.0 - -defmodule VeriSim.Federation.Adapters.Elasticsearch do - @moduledoc """ - Federation adapter for Elasticsearch (and OpenSearch). - - Translates VeriSimDB modality queries into Elasticsearch Query DSL - and normalises results into the federation result format. Communicates - via the Elasticsearch REST API. - - ## Modality Mapping - - | VeriSimDB Modality | Elasticsearch Capability | Mapping Type | - |--------------------|---------------------------|-----------------------| - | `:document` | Full-text search | `text` + `match` | - | `:vector` | kNN / dense_vector | `dense_vector` + kNN | - | `:semantic` | Nested objects | `nested` / `object` | - | `:temporal` | Date range queries | `date` + `range` | - | `:spatial` | Geo queries | `geo_shape` / `geo_point` | - - Elasticsearch does not natively support graph traversal, tensors, or - provenance chains, so `:graph`, `:tensor`, and `:provenance` modalities - are not supported. - - ## Configuration - - %{ - index: "octads", # Default index name - auth: {:basic, "elastic", "password"}, - version: 8 # ES major version (7 or 8) - } - - ## OpenSearch Compatibility - - This adapter is compatible with OpenSearch (the AWS-managed fork). - Set `version: 7` for OpenSearch 1.x/2.x compatibility (uses - `_doc` type specifier where needed). - """ - - @behaviour VeriSim.Federation.Adapter - - require Logger - - @default_timeout 10_000 - - # --------------------------------------------------------------------------- - # Callbacks - # --------------------------------------------------------------------------- - - @impl true - def connect(peer_info) do - case health_check(peer_info) do - {:ok, _latency} -> :ok - {:error, reason} -> {:error, reason} - end - end - - @impl true - def query(peer_info, query_params, opts \\ []) do - timeout = Keyword.get(opts, :timeout, @default_timeout) - modalities = Map.get(query_params, :modalities, []) - limit = Map.get(query_params, :limit, 100) - - start = System.monotonic_time(:millisecond) - - dsl = build_query_dsl(modalities, query_params, limit, peer_info) - - result = execute_search(peer_info, dsl, timeout) - - elapsed = System.monotonic_time(:millisecond) - start - - case result do - {:ok, raw_results} -> - normalised = - raw_results - |> translate_results(peer_info) - |> Enum.map(fn r -> Map.put(r, :response_time_ms, elapsed) end) - - {:ok, normalised} - - {:error, reason} -> - {:error, reason} - end - rescue - e -> - Logger.warning( - "Elasticsearch adapter: exception querying #{peer_info.store_id}: #{inspect(e)}" - ) - - {:error, {:exception, e}} - end - - @impl true - def health_check(peer_info) do - url = "#{peer_info.endpoint}/_cluster/health" - start = System.monotonic_time(:millisecond) - headers = auth_headers(peer_info.adapter_config) - - case Req.get(url, headers: headers, receive_timeout: 5_000) do - {:ok, %Req.Response{status: 200, body: body}} -> - status = body["status"] - - if status in ["green", "yellow"] do - elapsed = System.monotonic_time(:millisecond) - start - {:ok, elapsed} - else - {:error, {:cluster_red, status}} - end - - {:ok, %Req.Response{status: status}} -> - {:error, {:unhealthy, status}} - - {:error, reason} -> - {:error, reason} - end - rescue - e -> {:error, {:exception, e}} - end - - @impl true - def supported_modalities(_adapter_config) do - [:document, :vector, :semantic, :temporal, :spatial] - end - - @impl true - def translate_results(raw_results, peer_info) do - raw_results - |> List.wrap() - |> Enum.map(fn hit -> - # Elasticsearch hits have _source, _id, _score structure - source = hit["_source"] || hit - id = hit["_id"] || source["id"] || "unknown" - score = hit["_score"] || 0.0 - - %{ - source_store: peer_info.store_id, - octad_id: id, - score: if(is_number(score), do: score, else: 0.0), - drifted: false, - data: source, - response_time_ms: 0 - } - end) - end - - # --------------------------------------------------------------------------- - # Private — Elasticsearch Query DSL Builder - # --------------------------------------------------------------------------- - - defp build_query_dsl(modalities, query_params, limit, peer_info) do - cond do - :vector in modalities && Map.has_key?(query_params, :vector_query) -> - # kNN search (Elasticsearch 8+) or script_score (ES 7) - version = get_in(peer_info, [:adapter_config, :version]) || 8 - - if version >= 8 do - build_knn_query(query_params.vector_query, limit) - else - build_script_score_vector_query(query_params.vector_query, limit) - end - - :document in modalities && Map.has_key?(query_params, :text_query) -> - build_text_query(query_params.text_query, limit) - - :spatial in modalities && Map.has_key?(query_params, :spatial_bounds) -> - build_spatial_query(query_params.spatial_bounds, limit) - - :temporal in modalities && Map.has_key?(query_params, :temporal_range) -> - build_temporal_query(query_params.temporal_range, limit) - - :semantic in modalities && Map.has_key?(query_params, :filters) -> - build_filter_query(query_params.filters, limit) - - true -> - # Default: match_all - %{ - "query" => %{"match_all" => %{}}, - "size" => limit - } - end - end - - # Elasticsearch 8+ native kNN search - defp build_knn_query(vector, limit) do - %{ - "knn" => %{ - "field" => "embedding", - "query_vector" => vector, - "k" => limit, - "num_candidates" => limit * 10 - }, - "size" => limit - } - end - - # Elasticsearch 7 / OpenSearch script_score fallback for vector search - defp build_script_score_vector_query(vector, limit) do - %{ - "query" => %{ - "script_score" => %{ - "query" => %{"match_all" => %{}}, - "script" => %{ - "source" => "cosineSimilarity(params.query_vector, 'embedding') + 1.0", - "params" => %{"query_vector" => vector} - } - } - }, - "size" => limit - } - end - - # Full-text search with multi_match across common text fields - defp build_text_query(text_query, limit) do - %{ - "query" => %{ - "multi_match" => %{ - "query" => text_query, - "fields" => ["title^3", "body^2", "content", "description"], - "type" => "best_fields", - "fuzziness" => "AUTO" - } - }, - "size" => limit - } - end - - # Geo bounding box query (PostGIS equivalent) - defp build_spatial_query(bounds, limit) do - %{ - "query" => %{ - "geo_bounding_box" => %{ - "location" => %{ - "top_left" => %{ - "lat" => bounds[:max_lat] || bounds["max_lat"] || 0.0, - "lon" => bounds[:min_lon] || bounds["min_lon"] || 0.0 - }, - "bottom_right" => %{ - "lat" => bounds[:min_lat] || bounds["min_lat"] || 0.0, - "lon" => bounds[:max_lon] || bounds["max_lon"] || 0.0 - } - } - } - }, - "size" => limit - } - end - - # Date range query - defp build_temporal_query(range, limit) do - %{ - "query" => %{ - "range" => %{ - "created_at" => %{ - "gte" => range[:start] || range["start"], - "lte" => range[:end] || range["end"], - "format" => "strict_date_optional_time" - } - } - }, - "sort" => [%{"created_at" => %{"order" => "desc"}}], - "size" => limit - } - end - - # Generic filter query for semantic/structured data - defp build_filter_query(filters, limit) do - must_clauses = - Enum.map(filters, fn {field, value} -> - %{"term" => %{field => value}} - end) - - %{ - "query" => %{ - "bool" => %{"must" => must_clauses} - }, - "size" => limit - } - end - - defp execute_search(peer_info, dsl, timeout) do - config = peer_info.adapter_config - index = Map.get(config, :index, "octads") - url = "#{peer_info.endpoint}/#{index}/_search" - headers = auth_headers(config) ++ [{"Content-Type", "application/json"}] - - case Req.post(url, json: dsl, headers: headers, receive_timeout: timeout) do - {:ok, %Req.Response{status: 200, body: body}} -> - hits = get_in(body, ["hits", "hits"]) || [] - {:ok, hits} - - {:ok, %Req.Response{status: status, body: body}} -> - error_type = get_in(body, ["error", "type"]) || "unknown" - error_reason = get_in(body, ["error", "reason"]) || "HTTP #{status}" - - Logger.warning( - "Elasticsearch adapter: search failed: #{error_type} — #{error_reason}" - ) - - {:error, {:es_error, status, error_type, error_reason}} - - {:error, reason} -> - {:error, reason} - end - end - - defp auth_headers(config) do - case Map.get(config, :auth, :none) do - {:basic, user, pass} -> - encoded = Base.encode64("#{user}:#{pass}") - [{"Authorization", "Basic #{encoded}"}] - - {:bearer, token} -> - [{"Authorization", "Bearer #{token}"}] - - {:api_key, key} -> - [{"Authorization", "ApiKey #{key}"}] - - _ -> - [] - end - end -end diff --git a/verisimdb/elixir-orchestration/lib/verisim/federation/adapters/influxdb.ex b/verisimdb/elixir-orchestration/lib/verisim/federation/adapters/influxdb.ex deleted file mode 100644 index ccde85c8..00000000 --- a/verisimdb/elixir-orchestration/lib/verisim/federation/adapters/influxdb.ex +++ /dev/null @@ -1,379 +0,0 @@ -# SPDX-License-Identifier: MPL-2.0 - -defmodule VeriSim.Federation.Adapters.InfluxDB do - @moduledoc """ - Federation adapter for InfluxDB 2.x. - - Translates VeriSimDB modality queries into Flux (InfluxDB's functional - query language) and executes them via the InfluxDB v2 HTTP API. InfluxDB - is purpose-built for time-series data and excels at temporal queries with - high write throughput and configurable retention policies. - - ## Modality Mapping - - | VeriSimDB Modality | InfluxDB Capability | Extension/Feature Required | - |--------------------|----------------------------|----------------------------| - | `:temporal` | Native time-series storage | Built-in | - | `:semantic` | Tag-based filtering | Built-in | - - InfluxDB is specialised for time-series workloads. It does not support - graph traversal, vector similarity, full-text search, tensor operations, - provenance tracking, or geospatial queries, so `:graph`, `:vector`, - `:document`, `:tensor`, `:provenance`, and `:spatial` modalities are not - supported. - - ## Configuration - - %{ - host: "influxdb.internal", - port: 8086, - org: "verisim-org", - bucket: "octads", - token: "your-influxdb-token", - measurement: "octad_events" - } - - ## InfluxDB v2 HTTP API - - Queries are sent to `POST /api/v2/query` with Flux as the request body. - Health checks hit `GET /health` which returns `{"status": "pass"}` when - the server is ready. Authentication uses Bearer tokens. - - ## Flux Query Language - - Flux is a functional data scripting language designed for InfluxDB: - - from(bucket: "octads") - |> range(start: -1h) - |> filter(fn: (r) => r._measurement == "octad_events") - |> sort(columns: ["_time"], desc: true) - |> limit(n: 100) - """ - - @behaviour VeriSim.Federation.Adapter - - require Logger - - @default_timeout 10_000 - - # --------------------------------------------------------------------------- - # Callbacks - # --------------------------------------------------------------------------- - - @impl true - def connect(peer_info) do - case health_check(peer_info) do - {:ok, _latency} -> :ok - {:error, reason} -> {:error, reason} - end - end - - @impl true - def query(peer_info, query_params, opts \\ []) do - timeout = Keyword.get(opts, :timeout, @default_timeout) - modalities = Map.get(query_params, :modalities, []) - limit = Map.get(query_params, :limit, 100) - - start = System.monotonic_time(:millisecond) - - flux = build_flux(modalities, query_params, limit, peer_info) - - result = execute_flux(peer_info, flux, timeout) - - elapsed = System.monotonic_time(:millisecond) - start - - case result do - {:ok, raw_results} -> - normalised = - raw_results - |> translate_results(peer_info) - |> Enum.map(fn r -> Map.put(r, :response_time_ms, elapsed) end) - - {:ok, normalised} - - {:error, reason} -> - {:error, reason} - end - rescue - e -> - Logger.warning( - "InfluxDB adapter: exception querying #{peer_info.store_id}: #{inspect(e)}" - ) - - {:error, {:exception, e}} - end - - @impl true - def health_check(peer_info) do - start = System.monotonic_time(:millisecond) - headers = auth_headers(peer_info.adapter_config) - - # InfluxDB v2 health endpoint: GET /health returns {"status": "pass"} - url = "#{peer_info.endpoint}/health" - - case Req.get(url, headers: headers, receive_timeout: 5_000) do - {:ok, %Req.Response{status: 200, body: body}} -> - status = if is_map(body), do: body["status"], else: body - - if status in ["pass", "ready"] do - elapsed = System.monotonic_time(:millisecond) - start - {:ok, elapsed} - else - {:error, {:unhealthy_status, status}} - end - - {:ok, %Req.Response{status: status}} -> - {:error, {:unhealthy, status}} - - {:error, reason} -> - {:error, reason} - end - rescue - e -> {:error, {:exception, e}} - end - - @impl true - def supported_modalities(_adapter_config) do - # InfluxDB is a time-series database — only temporal and semantic (tags) - [:temporal, :semantic] - end - - @impl true - def translate_results(raw_results, peer_info) do - raw_results - |> List.wrap() - |> Enum.map(fn row -> - %{ - source_store: peer_info.store_id, - octad_id: extract_id(row), - score: parse_score(row), - drifted: false, - data: row, - response_time_ms: 0 - } - end) - end - - # --------------------------------------------------------------------------- - # Private — Flux Query Builder - # --------------------------------------------------------------------------- - - defp build_flux(modalities, query_params, limit, peer_info) do - config = peer_info.adapter_config - bucket = Map.get(config, :bucket, "octads") - measurement = Map.get(config, :measurement, "octad_events") - - cond do - :temporal in modalities && Map.has_key?(query_params, :temporal_range) -> - # Time-range query — the core InfluxDB use case - range = query_params.temporal_range - start_time = range[:start] || range["start"] || "-24h" - end_time = range[:end] || range["end"] || "now()" - - # Determine if start/end are relative durations or absolute timestamps - start_expr = format_flux_time(start_time) - end_expr = format_flux_time(end_time) - - base_flux = """ - from(bucket: "#{bucket}") - |> range(start: #{start_expr}, stop: #{end_expr}) - |> filter(fn: (r) => r._measurement == "#{measurement}") - """ - - # Add tag filters if semantic modality is also requested - base_flux = - if :semantic in modalities && Map.has_key?(query_params, :filters) do - tag_filters = build_flux_tag_filters(query_params.filters) - base_flux <> tag_filters - else - base_flux - end - - base_flux <> - """ - |> sort(columns: ["_time"], desc: true) - |> limit(n: #{limit}) - """ - - :semantic in modalities && Map.has_key?(query_params, :filters) -> - # Tag-based filtering with default time range - tag_filters = build_flux_tag_filters(query_params.filters) - - """ - from(bucket: "#{bucket}") - |> range(start: -30d) - |> filter(fn: (r) => r._measurement == "#{measurement}") - #{tag_filters} - |> sort(columns: ["_time"], desc: true) - |> limit(n: #{limit}) - """ - - true -> - # Default: recent data from the bucket - """ - from(bucket: "#{bucket}") - |> range(start: -24h) - |> filter(fn: (r) => r._measurement == "#{measurement}") - |> sort(columns: ["_time"], desc: true) - |> limit(n: #{limit}) - """ - end - end - - defp build_flux_tag_filters(filters) do - filters - |> Enum.map(fn {tag, value} -> - tag_str = escape_flux(to_string(tag)) - val_str = escape_flux(to_string(value)) - " |> filter(fn: (r) => r.#{tag_str} == \"#{val_str}\")" - end) - |> Enum.join("\n") - end - - defp format_flux_time(time) when is_binary(time) do - cond do - # Relative duration: -1h, -24h, -7d, etc. - String.match?(time, ~r/^-\d+[smhdw]$/) -> time - # "now()" function call - time == "now()" -> "now()" - # Absolute ISO8601 timestamp - true -> ~s|#{time}| - end - end - - defp format_flux_time(time), do: to_string(time) - - defp execute_flux(peer_info, flux, timeout) do - config = peer_info.adapter_config - org = Map.get(config, :org, "verisim-org") - headers = auth_headers(config) - - # InfluxDB v2 query endpoint: POST /api/v2/query - url = "#{peer_info.endpoint}/api/v2/query?org=#{URI.encode(org)}" - - query_headers = - headers ++ - [ - {"Content-Type", "application/vnd.flux"}, - {"Accept", "application/json"} - ] - - case Req.post(url, body: flux, headers: query_headers, receive_timeout: timeout) do - {:ok, %Req.Response{status: 200, body: body}} when is_list(body) -> - {:ok, body} - - {:ok, %Req.Response{status: 200, body: body}} when is_binary(body) -> - # InfluxDB may return CSV or annotated CSV; parse to maps - rows = parse_flux_csv(body) - {:ok, rows} - - {:ok, %Req.Response{status: 200, body: body}} when is_map(body) -> - results = body["results"] || [body] - {:ok, results} - - {:ok, %Req.Response{status: status, body: body}} -> - error_msg = - cond do - is_map(body) -> body["message"] || body["error"] || "HTTP #{status}" - is_binary(body) -> String.trim(body) - true -> "HTTP #{status}" - end - - Logger.warning("InfluxDB adapter: query failed (#{status}): #{error_msg}") - {:error, {:flux_error, status, error_msg}} - - {:error, reason} -> - {:error, reason} - end - end - - # --------------------------------------------------------------------------- - # Private — Helpers - # --------------------------------------------------------------------------- - - defp parse_flux_csv(csv_body) when is_binary(csv_body) do - lines = String.split(csv_body, "\n", trim: true) - - case lines do - [] -> - [] - - [header_line | data_lines] -> - # Skip annotation lines (start with #) and empty lines - {headers, data} = extract_csv_headers_and_data(header_line, data_lines) - - Enum.map(data, fn line -> - values = String.split(line, ",") - - headers - |> Enum.zip(values) - |> Enum.reject(fn {h, _v} -> String.starts_with?(h, "#") end) - |> Map.new() - end) - end - end - - defp extract_csv_headers_and_data(first_line, rest) do - # InfluxDB annotated CSV has annotation rows starting with # - all_lines = [first_line | rest] - - non_annotation = Enum.reject(all_lines, &String.starts_with?(&1, "#")) - - case non_annotation do - [] -> {[], []} - [header | data] -> {String.split(header, ","), data} - end - end - - defp extract_id(row) do - # InfluxDB records are identified by measurement + tag set + timestamp - id_parts = - [ - row["_measurement"], - row["entity_id"] || row["id"], - row["_time"] - ] - |> Enum.reject(&is_nil/1) - - case id_parts do - [] -> "unknown" - parts -> Enum.join(parts, ":") - end - end - - defp parse_score(row) do - case row["_value"] || row["score"] do - score when is_number(score) -> score / 1 - _ -> 0.0 - end - end - - defp escape_flux(str) when is_binary(str) do - str - |> String.replace("\\", "\\\\") - |> String.replace("\"", "\\\"") - end - - defp auth_headers(config) do - token = Map.get(config, :token) - - case Map.get(config, :auth, :none) do - {:bearer, tok} -> - [{"Authorization", "Token #{tok}"}] - - {:api_key, key} -> - [{"Authorization", "Token #{key}"}] - - :none when is_binary(token) -> - # InfluxDB convention: token in config map directly - [{"Authorization", "Token #{token}"}] - - {:basic, user, pass} -> - encoded = Base.encode64("#{user}:#{pass}") - [{"Authorization", "Basic #{encoded}"}] - - _ -> - [] - end - end -end diff --git a/verisimdb/elixir-orchestration/lib/verisim/federation/adapters/mongodb.ex b/verisimdb/elixir-orchestration/lib/verisim/federation/adapters/mongodb.ex deleted file mode 100644 index c556d23d..00000000 --- a/verisimdb/elixir-orchestration/lib/verisim/federation/adapters/mongodb.ex +++ /dev/null @@ -1,389 +0,0 @@ -# SPDX-License-Identifier: MPL-2.0 - -defmodule VeriSim.Federation.Adapters.MongoDB do - @moduledoc """ - Federation adapter for MongoDB (with Atlas Vector Search and GeoJSON support). - - Translates VeriSimDB modality queries into MongoDB aggregation pipelines and - executes them via the MongoDB Data API (HTTP/JSON). Supports a wide range of - modalities through MongoDB's multi-model capabilities and Atlas extensions. - - ## Modality Mapping - - | VeriSimDB Modality | MongoDB Capability | Extension/Module Required | - |--------------------|-----------------------------|--------------------------------| - | `:graph` | `$graphLookup` / `DBRef` | Built-in | - | `:vector` | Atlas Vector Search | Atlas cluster with vector index | - | `:document` | `$text` search / Atlas FTS | Text index on collection | - | `:temporal` | ISODate range filters | Built-in | - | `:provenance` | Change streams / oplog | Replica set required | - | `:spatial` | `$geoNear` / `$geoWithin` | 2dsphere index | - | `:semantic` | Nested BSON documents | Built-in | - - The `:tensor` modality has no direct MongoDB mapping and is not supported - by this adapter. - - ## Configuration - - %{ - host: "cluster0.example.mongodb.net", - port: 27017, - database: "verisimdb", - collection: "octads", - auth: {:basic, "verisim_user", "password"}, - replica_set: "rs0", - data_api: true # Use MongoDB Data API (HTTP) instead of wire protocol - } - - ## MongoDB Data API - - When `data_api: true` (default), queries are sent via the MongoDB Atlas Data - API at `POST /action/aggregate`. When `data_api: false`, the adapter falls - back to a generic HTTP proxy endpoint compatible with mongosh-style queries. - - ## Health Check - - The adapter sends a `{ping: 1}` command via the Data API's `runCommand` - endpoint or checks the `/status` endpoint of a MongoDB HTTP proxy. - """ - - @behaviour VeriSim.Federation.Adapter - - require Logger - - @default_timeout 10_000 - - # --------------------------------------------------------------------------- - # Callbacks - # --------------------------------------------------------------------------- - - @impl true - def connect(peer_info) do - case health_check(peer_info) do - {:ok, _latency} -> :ok - {:error, reason} -> {:error, reason} - end - end - - @impl true - def query(peer_info, query_params, opts \\ []) do - timeout = Keyword.get(opts, :timeout, @default_timeout) - modalities = Map.get(query_params, :modalities, []) - limit = Map.get(query_params, :limit, 100) - - start = System.monotonic_time(:millisecond) - - pipeline = build_pipeline(modalities, query_params, limit, peer_info) - - result = execute_aggregate(peer_info, pipeline, timeout) - - elapsed = System.monotonic_time(:millisecond) - start - - case result do - {:ok, raw_results} -> - normalised = - raw_results - |> translate_results(peer_info) - |> Enum.map(fn r -> Map.put(r, :response_time_ms, elapsed) end) - - {:ok, normalised} - - {:error, reason} -> - {:error, reason} - end - rescue - e -> - Logger.warning( - "MongoDB adapter: exception querying #{peer_info.store_id}: #{inspect(e)}" - ) - - {:error, {:exception, e}} - end - - @impl true - def health_check(peer_info) do - config = peer_info.adapter_config - db = Map.get(config, :database, "verisimdb") - start = System.monotonic_time(:millisecond) - headers = auth_headers(config) - - # MongoDB Data API: run {ping: 1} command - url = "#{peer_info.endpoint}/action/runCommand" - - body = %{ - "database" => db, - "dataSource" => Map.get(config, :data_source, "Cluster0"), - "command" => %{"ping" => 1} - } - - case Req.post(url, json: body, headers: headers, receive_timeout: 5_000) do - {:ok, %Req.Response{status: status}} when status in 200..299 -> - elapsed = System.monotonic_time(:millisecond) - start - {:ok, elapsed} - - {:ok, %Req.Response{status: status}} -> - {:error, {:unhealthy, status}} - - {:error, reason} -> - {:error, reason} - end - rescue - e -> {:error, {:exception, e}} - end - - @impl true - def supported_modalities(adapter_config) do - # MongoDB supports most modalities natively; vector search requires Atlas - atlas_enabled = Map.get(adapter_config, :atlas, false) - has_replica_set = Map.has_key?(adapter_config, :replica_set) - has_geo_index = Map.get(adapter_config, :geo_index, true) - - base = [:graph, :document, :temporal, :semantic] - - base - |> maybe_add(:vector, atlas_enabled) - |> maybe_add(:provenance, has_replica_set) - |> maybe_add(:spatial, has_geo_index) - end - - @impl true - def translate_results(raw_results, peer_info) do - raw_results - |> List.wrap() - |> Enum.map(fn doc -> - %{ - source_store: peer_info.store_id, - octad_id: extract_id(doc), - score: parse_score(doc), - drifted: false, - data: doc, - response_time_ms: 0 - } - end) - end - - # --------------------------------------------------------------------------- - # Private — MongoDB Aggregation Pipeline Builder - # --------------------------------------------------------------------------- - - defp build_pipeline(modalities, query_params, limit, peer_info) do - config = peer_info.adapter_config - _collection = Map.get(config, :collection, "octads") - - cond do - :vector in modalities && Map.has_key?(query_params, :vector_query) -> - # Atlas Vector Search — $vectorSearch aggregation stage - embedding = query_params.vector_query - index_name = Map.get(config, :vector_index, "vector_index") - - [ - %{ - "$vectorSearch" => %{ - "index" => index_name, - "path" => "embedding", - "queryVector" => embedding, - "numCandidates" => limit * 10, - "limit" => limit - } - }, - %{ - "$addFields" => %{ - "score" => %{"$meta" => "vectorSearchScore"} - } - } - ] - - :document in modalities && Map.has_key?(query_params, :text_query) -> - # Full-text search via $text index or Atlas Search - text = query_params.text_query - - [ - %{ - "$match" => %{ - "$text" => %{"$search" => text} - } - }, - %{ - "$addFields" => %{ - "score" => %{"$meta" => "textScore"} - } - }, - %{"$sort" => %{"score" => -1}}, - %{"$limit" => limit} - ] - - :spatial in modalities && Map.has_key?(query_params, :spatial_bounds) -> - # GeoJSON $geoWithin or $geoNear - bounds = query_params.spatial_bounds - min_lon = bounds[:min_lon] || bounds["min_lon"] || 0.0 - min_lat = bounds[:min_lat] || bounds["min_lat"] || 0.0 - max_lon = bounds[:max_lon] || bounds["max_lon"] || 0.0 - max_lat = bounds[:max_lat] || bounds["max_lat"] || 0.0 - - [ - %{ - "$match" => %{ - "location" => %{ - "$geoWithin" => %{ - "$geometry" => %{ - "type" => "Polygon", - "coordinates" => [ - [ - [min_lon, min_lat], - [max_lon, min_lat], - [max_lon, max_lat], - [min_lon, max_lat], - [min_lon, min_lat] - ] - ] - } - } - } - } - }, - %{"$limit" => limit} - ] - - :graph in modalities && Map.has_key?(query_params, :graph_pattern) -> - # $graphLookup for recursive graph traversal - start_id = query_params.graph_pattern - edges_collection = Map.get(config, :edges_collection, "edges") - - [ - %{ - "$match" => %{"_id" => start_id} - }, - %{ - "$graphLookup" => %{ - "from" => edges_collection, - "startWith" => "$_id", - "connectFromField" => "target_id", - "connectToField" => "source_id", - "as" => "traversal", - "maxDepth" => 3, - "depthField" => "depth" - } - }, - %{"$limit" => limit} - ] - - :temporal in modalities && Map.has_key?(query_params, :temporal_range) -> - # Date range filter on ISODate fields - range = query_params.temporal_range - start_time = range[:start] || range["start"] || "" - end_time = range[:end] || range["end"] || "" - - [ - %{ - "$match" => %{ - "created_at" => %{ - "$gte" => %{"$date" => start_time}, - "$lte" => %{"$date" => end_time} - } - } - }, - %{"$sort" => %{"created_at" => -1}}, - %{"$limit" => limit} - ] - - :provenance in modalities -> - # Query change stream log or provenance collection - provenance_collection = Map.get(config, :provenance_collection, "provenance_log") - - [ - %{ - "$unionWith" => %{ - "coll" => provenance_collection, - "pipeline" => [ - %{"$sort" => %{"timestamp" => -1}}, - %{"$limit" => limit} - ] - } - }, - %{"$limit" => limit} - ] - - :semantic in modalities && Map.has_key?(query_params, :filters) -> - # Nested document query via dot notation - filters = query_params.filters - match_clause = Enum.into(filters, %{}, fn {k, v} -> {"metadata.#{k}", v} end) - - [ - %{"$match" => match_clause}, - %{"$limit" => limit} - ] - - true -> - # Default: return all documents sorted by _id - [ - %{"$sort" => %{"_id" => 1}}, - %{"$limit" => limit} - ] - end - end - - defp execute_aggregate(peer_info, pipeline, timeout) do - config = peer_info.adapter_config - db = Map.get(config, :database, "verisimdb") - collection = Map.get(config, :collection, "octads") - headers = auth_headers(config) - - url = "#{peer_info.endpoint}/action/aggregate" - - body = %{ - "database" => db, - "dataSource" => Map.get(config, :data_source, "Cluster0"), - "collection" => collection, - "pipeline" => pipeline - } - - case Req.post(url, json: body, headers: headers, receive_timeout: timeout) do - {:ok, %Req.Response{status: status, body: resp_body}} when status in 200..299 -> - documents = resp_body["documents"] || resp_body["cursor"]["firstBatch"] || [] - {:ok, documents} - - {:ok, %Req.Response{status: status, body: resp_body}} -> - error_msg = resp_body["error"] || resp_body["errorMessage"] || "HTTP #{status}" - Logger.warning("MongoDB adapter: aggregate failed: #{error_msg}") - {:error, {:mongo_error, status, error_msg}} - - {:error, reason} -> - {:error, reason} - end - end - - # --------------------------------------------------------------------------- - # Private — Helpers - # --------------------------------------------------------------------------- - - defp extract_id(doc) do - doc["_id"] || doc["id"] || doc["_key"] || "unknown" - end - - defp parse_score(doc) do - case doc["score"] do - score when is_number(score) -> score / 1 - _ -> 0.0 - end - end - - defp maybe_add(list, item, true), do: list ++ [item] - defp maybe_add(list, _item, false), do: list - - defp auth_headers(config) do - case Map.get(config, :auth, :none) do - {:basic, user, pass} -> - encoded = Base.encode64("#{user}:#{pass}") - [{"Authorization", "Basic #{encoded}"}, {"Content-Type", "application/json"}] - - {:bearer, token} -> - [{"Authorization", "Bearer #{token}"}, {"Content-Type", "application/json"}] - - {:api_key, key} -> - [{"api-key", key}, {"Content-Type", "application/json"}] - - _ -> - [{"Content-Type", "application/json"}] - end - end -end diff --git a/verisimdb/elixir-orchestration/lib/verisim/federation/adapters/neo4j.ex b/verisimdb/elixir-orchestration/lib/verisim/federation/adapters/neo4j.ex deleted file mode 100644 index ce7b473a..00000000 --- a/verisimdb/elixir-orchestration/lib/verisim/federation/adapters/neo4j.ex +++ /dev/null @@ -1,427 +0,0 @@ -# SPDX-License-Identifier: MPL-2.0 - -defmodule VeriSim.Federation.Adapters.Neo4j do - @moduledoc """ - Federation adapter for Neo4j. - - Translates VeriSimDB modality queries into Cypher and executes them via - the Neo4j HTTP Transactional API. Neo4j is the premier graph database, - offering native graph storage and processing with a rich query language - (Cypher), vector search indices, full-text indices, temporal types, and - spatial point types. - - ## Modality Mapping - - | VeriSimDB Modality | Neo4j Capability | Extension/Feature Required | - |--------------------|-------------------------------|----------------------------| - | `:graph` | Native Cypher traversal | Built-in | - | `:vector` | Vector index (5.11+) | Built-in (5.11+) | - | `:document` | Full-text index (Lucene) | Built-in | - | `:temporal` | Temporal types (date, datetime) | Built-in | - | `:spatial` | Point type, distance() | Built-in | - | `:semantic` | Node/relationship properties | Built-in | - - Neo4j does not support tensor operations or provenance chains natively, - so `:tensor` and `:provenance` modalities are not supported. Provenance - could be modelled as relationship chains if needed in future. - - ## Configuration - - %{ - host: "neo4j.internal", - port: 7474, - bolt_port: 7687, - database: "neo4j", - auth: {:basic, "neo4j", "password"}, - version: 5 # Major version (4 or 5) - } - - ## Neo4j HTTP API - - Queries are sent to `POST /db/{database}/tx/commit` as Cypher statements. - Health checks hit `GET /` which returns the server discovery document. - For vector search, Neo4j 5.11+ is required. - """ - - @behaviour VeriSim.Federation.Adapter - - require Logger - - @default_timeout 10_000 - - # --------------------------------------------------------------------------- - # Callbacks - # --------------------------------------------------------------------------- - - @impl true - def connect(peer_info) do - case health_check(peer_info) do - {:ok, _latency} -> :ok - {:error, reason} -> {:error, reason} - end - end - - @impl true - def query(peer_info, query_params, opts \\ []) do - timeout = Keyword.get(opts, :timeout, @default_timeout) - modalities = Map.get(query_params, :modalities, []) - limit = Map.get(query_params, :limit, 100) - - start = System.monotonic_time(:millisecond) - - {cypher, cypher_params} = build_cypher(modalities, query_params, limit, peer_info) - - result = execute_cypher(peer_info, cypher, cypher_params, timeout) - - elapsed = System.monotonic_time(:millisecond) - start - - case result do - {:ok, raw_results} -> - normalised = - raw_results - |> translate_results(peer_info) - |> Enum.map(fn r -> Map.put(r, :response_time_ms, elapsed) end) - - {:ok, normalised} - - {:error, reason} -> - {:error, reason} - end - rescue - e -> - Logger.warning( - "Neo4j adapter: exception querying #{peer_info.store_id}: #{inspect(e)}" - ) - - {:error, {:exception, e}} - end - - @impl true - def health_check(peer_info) do - start = System.monotonic_time(:millisecond) - headers = auth_headers(peer_info.adapter_config) - - # Neo4j discovery endpoint: GET / returns server info JSON - url = peer_info.endpoint - - case Req.get(url, headers: headers, receive_timeout: 5_000) do - {:ok, %Req.Response{status: 200, body: body}} when is_map(body) -> - # Neo4j returns {"bolt_routing": "...", "transaction": "...", ...} - if Map.has_key?(body, "bolt_routing") or Map.has_key?(body, "neo4j_version") do - elapsed = System.monotonic_time(:millisecond) - start - {:ok, elapsed} - else - # Could still be a valid Neo4j response; accept any 200 - elapsed = System.monotonic_time(:millisecond) - start - {:ok, elapsed} - end - - {:ok, %Req.Response{status: 200}} -> - elapsed = System.monotonic_time(:millisecond) - start - {:ok, elapsed} - - {:ok, %Req.Response{status: status}} -> - {:error, {:unhealthy, status}} - - {:error, reason} -> - {:error, reason} - end - rescue - e -> {:error, {:exception, e}} - end - - @impl true - def supported_modalities(adapter_config) do - version = Map.get(adapter_config, :version, 5) - - base = [:graph, :document, :temporal, :spatial, :semantic] - - # Vector search requires Neo4j 5.11+ - base - |> maybe_add(:vector, version >= 5) - end - - @impl true - def translate_results(raw_results, peer_info) do - raw_results - |> List.wrap() - |> Enum.map(fn row -> - # Neo4j transaction API returns {"row": [...], "meta": [...]} per result - node_data = extract_node_data(row) - - %{ - source_store: peer_info.store_id, - octad_id: node_data["id"] || node_data["elementId"] || node_data["_id"] || "unknown", - score: parse_score(node_data), - drifted: false, - data: node_data, - response_time_ms: 0 - } - end) - end - - # --------------------------------------------------------------------------- - # Private — Cypher Query Builder - # --------------------------------------------------------------------------- - - defp build_cypher(modalities, query_params, limit, peer_info) do - config = peer_info.adapter_config - label = Map.get(config, :label, "Octad") - - cond do - :graph in modalities && Map.has_key?(query_params, :graph_pattern) -> - # Cypher graph traversal: variable-length path patterns - start_id = query_params.graph_pattern - max_depth = Map.get(config, :max_depth, 3) - rel_type = Map.get(config, :relationship_type, "CONNECTED_TO") - - cypher = """ - MATCH (start:#{label} {id: $start_id}) - MATCH path = (start)-[:#{rel_type}*1..#{max_depth}]-(connected) - RETURN connected, length(path) AS depth, 0.0 AS score - ORDER BY depth ASC - LIMIT $limit - """ - - params = %{"start_id" => start_id, "limit" => limit} - {cypher, params} - - :vector in modalities && Map.has_key?(query_params, :vector_query) -> - # Neo4j 5.11+ vector index query - embedding = query_params.vector_query - index_name = Map.get(config, :vector_index, "octad_embedding_index") - - cypher = """ - CALL db.index.vector.queryNodes($index_name, $limit, $embedding) - YIELD node, score - RETURN node, score - ORDER BY score DESC - """ - - params = %{ - "index_name" => index_name, - "limit" => limit, - "embedding" => embedding - } - - {cypher, params} - - :document in modalities && Map.has_key?(query_params, :text_query) -> - # Neo4j full-text index (Lucene-backed) - text = query_params.text_query - index_name = Map.get(config, :fulltext_index, "octad_fulltext") - - cypher = """ - CALL db.index.fulltext.queryNodes($index_name, $query) - YIELD node, score - RETURN node, score - ORDER BY score DESC - LIMIT $limit - """ - - params = %{ - "index_name" => index_name, - "query" => text, - "limit" => limit - } - - {cypher, params} - - :spatial in modalities && Map.has_key?(query_params, :spatial_bounds) -> - # Neo4j point-based spatial queries - bounds = query_params.spatial_bounds - center_lat = ((bounds[:min_lat] || 0.0) + (bounds[:max_lat] || 0.0)) / 2 - center_lon = ((bounds[:min_lon] || 0.0) + (bounds[:max_lon] || 0.0)) / 2 - # Approximate radius from bounds (rough calculation) - radius_km = Map.get(query_params, :radius_km, 50.0) - - cypher = """ - MATCH (n:#{label}) - WHERE point.distance(n.location, point({latitude: $lat, longitude: $lon})) < $radius - RETURN n, point.distance(n.location, point({latitude: $lat, longitude: $lon})) AS distance, - 0.0 AS score - ORDER BY distance ASC - LIMIT $limit - """ - - params = %{ - "lat" => center_lat, - "lon" => center_lon, - "radius" => radius_km * 1000, - "limit" => limit - } - - {cypher, params} - - :temporal in modalities && Map.has_key?(query_params, :temporal_range) -> - range = query_params.temporal_range - - cypher = """ - MATCH (n:#{label}) - WHERE n.created_at >= datetime($start_time) - AND n.created_at <= datetime($end_time) - RETURN n, 0.0 AS score - ORDER BY n.created_at DESC - LIMIT $limit - """ - - params = %{ - "start_time" => range[:start] || range["start"] || "", - "end_time" => range[:end] || range["end"] || "", - "limit" => limit - } - - {cypher, params} - - :semantic in modalities && Map.has_key?(query_params, :filters) -> - # Node property filters - filters = query_params.filters - - where_clauses = - filters - |> Enum.with_index() - |> Enum.map(fn {{field, _value}, idx} -> - "n.#{field} = $filter_#{idx}" - end) - |> Enum.join(" AND ") - - where_clause = if where_clauses == "", do: "true", else: where_clauses - - filter_params = - filters - |> Enum.with_index() - |> Enum.into(%{}, fn {{_field, value}, idx} -> - {"filter_#{idx}", value} - end) - - cypher = """ - MATCH (n:#{label}) - WHERE #{where_clause} - RETURN n, 0.0 AS score - LIMIT $limit - """ - - params = Map.merge(filter_params, %{"limit" => limit}) - {cypher, params} - - true -> - # Default: return all nodes of the label - cypher = """ - MATCH (n:#{label}) - RETURN n, 0.0 AS score - ORDER BY n.id ASC - LIMIT $limit - """ - - params = %{"limit" => limit} - {cypher, params} - end - end - - defp execute_cypher(peer_info, cypher, cypher_params, timeout) do - config = peer_info.adapter_config - database = Map.get(config, :database, "neo4j") - headers = auth_headers(config) ++ [{"Content-Type", "application/json"}] - - # Neo4j Transaction API: POST /db/{database}/tx/commit - url = "#{peer_info.endpoint}/db/#{database}/tx/commit" - - body = %{ - "statements" => [ - %{ - "statement" => cypher, - "parameters" => cypher_params, - "resultDataContents" => ["row"] - } - ] - } - - case Req.post(url, json: body, headers: headers, receive_timeout: timeout) do - {:ok, %Req.Response{status: 200, body: resp_body}} -> - errors = resp_body["errors"] || [] - - if errors == [] do - results = extract_neo4j_results(resp_body) - {:ok, results} - else - error = List.first(errors) - error_msg = "#{error["code"]}: #{error["message"]}" - Logger.warning("Neo4j adapter: Cypher error: #{error_msg}") - {:error, {:cypher_error, error_msg}} - end - - {:ok, %Req.Response{status: status, body: resp_body}} -> - error_msg = resp_body["message"] || "HTTP #{status}" - Logger.warning("Neo4j adapter: request failed: #{error_msg}") - {:error, {:neo4j_error, status, error_msg}} - - {:error, reason} -> - {:error, reason} - end - end - - # --------------------------------------------------------------------------- - # Private — Helpers - # --------------------------------------------------------------------------- - - defp extract_neo4j_results(resp_body) do - results = resp_body["results"] || [] - - results - |> Enum.flat_map(fn result -> - columns = result["columns"] || [] - data = result["data"] || [] - - Enum.map(data, fn datum -> - row_values = datum["row"] || [] - Enum.zip(columns, row_values) |> Map.new() - end) - end) - end - - defp extract_node_data(row) when is_map(row) do - # Neo4j rows may contain node objects under column names like "n", "node", "connected" - # Extract the first map value that looks like node properties - node = - row - |> Map.values() - |> Enum.find(fn - v when is_map(v) -> true - _ -> false - end) - - case node do - nil -> row - node_map -> Map.merge(row, node_map) - end - end - - defp extract_node_data(row), do: %{"raw" => row} - - defp parse_score(row) do - case row["score"] do - score when is_number(score) -> score / 1 - _ -> 0.0 - end - end - - defp maybe_add(list, item, true), do: list ++ [item] - defp maybe_add(list, _item, false), do: list - - defp auth_headers(config) do - case Map.get(config, :auth, :none) do - {:basic, user, pass} -> - encoded = Base.encode64("#{user}:#{pass}") - [{"Authorization", "Basic #{encoded}"}] - - {:bearer, token} -> - [{"Authorization", "Bearer #{token}"}] - - {:api_key, key} -> - [{"X-API-Key", key}] - - _ -> - [] - end - end -end diff --git a/verisimdb/elixir-orchestration/lib/verisim/federation/adapters/object_storage.ex b/verisimdb/elixir-orchestration/lib/verisim/federation/adapters/object_storage.ex deleted file mode 100644 index efea1b2d..00000000 --- a/verisimdb/elixir-orchestration/lib/verisim/federation/adapters/object_storage.ex +++ /dev/null @@ -1,461 +0,0 @@ -# SPDX-License-Identifier: MPL-2.0 - -defmodule VeriSim.Federation.Adapters.ObjectStorage do - @moduledoc """ - Unified federation adapter for S3-compatible object storage: MinIO and - Amazon S3. - - Translates VeriSimDB modality queries into S3 API calls and normalises - results into the federation result format. Object storage is not a - traditional database, but it serves as a durable persistence layer for - VeriSimDB entities — particularly for document content, provenance audit - trails, and temporal versioning. - - ## Modality Mapping - - | VeriSimDB Modality | S3/MinIO Capability | Feature Required | - |--------------------|-------------------------------|----------------------------| - | `:document` | Object content + metadata FT | Custom metadata indexing | - | `:temporal` | Object versioning | Bucket versioning enabled | - | `:provenance` | Access logs / audit trails | Server access logging | - | `:semantic` | Object metadata / tags | Built-in (user metadata) | - - Object storage does not support graph traversal, vector similarity, - tensor operations, or geospatial queries, so `:graph`, `:vector`, - `:tensor`, and `:spatial` modalities are not supported. - - ## Configuration - - # MinIO - %{ - host: "minio.internal", - port: 9000, - bucket: "verisimdb-octads", - region: "us-east-1", - access_key: "minioadmin", - secret_key: "minioadmin", - backend: :minio - } - - # Amazon S3 - %{ - host: "s3.amazonaws.com", - port: 443, - bucket: "verisimdb-octads", - region: "eu-west-1", - access_key: "AKIA...", - secret_key: "...", - backend: :s3 - } - - ## S3 API Compatibility - - Both MinIO and Amazon S3 implement the S3 API. This adapter uses the - REST API via Req with AWS Signature V4 authentication. MinIO endpoints - use path-style URLs; S3 uses virtual-hosted-style by default. - - ## Query Model - - Since S3 is not a query engine, "queries" are implemented as: - - **ListObjectsV2**: List objects with prefix filtering - - **HeadObject**: Check existence and metadata - - **GetObject**: Retrieve content - - **GetObjectTagging**: Retrieve tags for semantic filtering - - **ListObjectVersions**: Temporal versioning queries - """ - - @behaviour VeriSim.Federation.Adapter - - require Logger - - @default_timeout 15_000 - - # --------------------------------------------------------------------------- - # Callbacks - # --------------------------------------------------------------------------- - - @impl true - def connect(peer_info) do - case health_check(peer_info) do - {:ok, _latency} -> :ok - {:error, reason} -> {:error, reason} - end - end - - @impl true - def query(peer_info, query_params, opts \\ []) do - timeout = Keyword.get(opts, :timeout, @default_timeout) - modalities = Map.get(query_params, :modalities, []) - limit = Map.get(query_params, :limit, 100) - - start = System.monotonic_time(:millisecond) - - result = execute_s3_query(peer_info, modalities, query_params, limit, timeout) - - elapsed = System.monotonic_time(:millisecond) - start - - case result do - {:ok, raw_results} -> - normalised = - raw_results - |> translate_results(peer_info) - |> Enum.map(fn r -> Map.put(r, :response_time_ms, elapsed) end) - - {:ok, normalised} - - {:error, reason} -> - {:error, reason} - end - rescue - e -> - Logger.warning( - "ObjectStorage adapter: exception querying #{peer_info.store_id}: #{inspect(e)}" - ) - - {:error, {:exception, e}} - end - - @impl true - def health_check(peer_info) do - config = peer_info.adapter_config - bucket = Map.get(config, :bucket, "verisimdb-octads") - start = System.monotonic_time(:millisecond) - - # Health check: HEAD bucket to verify it exists and is accessible - url = build_bucket_url(peer_info, bucket) - headers = auth_headers(config) - - case Req.head(url, headers: headers, receive_timeout: 5_000) do - {:ok, %Req.Response{status: status}} when status in [200, 301, 307] -> - elapsed = System.monotonic_time(:millisecond) - start - {:ok, elapsed} - - {:ok, %Req.Response{status: 404}} -> - {:error, {:bucket_not_found, bucket}} - - {:ok, %Req.Response{status: 403}} -> - {:error, {:access_denied, bucket}} - - {:ok, %Req.Response{status: status}} -> - {:error, {:unhealthy, status}} - - {:error, reason} -> - {:error, reason} - end - rescue - e -> {:error, {:exception, e}} - end - - @impl true - def supported_modalities(adapter_config) do - versioning_enabled = Map.get(adapter_config, :versioning, false) - logging_enabled = Map.get(adapter_config, :access_logging, false) - - base = [:document, :semantic] - - base - |> maybe_add(:temporal, versioning_enabled) - |> maybe_add(:provenance, logging_enabled) - end - - @impl true - def translate_results(raw_results, peer_info) do - raw_results - |> List.wrap() - |> Enum.map(fn obj -> - %{ - source_store: peer_info.store_id, - octad_id: extract_object_id(obj), - score: parse_score(obj), - drifted: false, - data: obj, - response_time_ms: 0 - } - end) - end - - # --------------------------------------------------------------------------- - # Private — S3 Query Execution - # --------------------------------------------------------------------------- - - defp execute_s3_query(peer_info, modalities, query_params, limit, timeout) do - config = peer_info.adapter_config - bucket = Map.get(config, :bucket, "verisimdb-octads") - - cond do - :temporal in modalities && Map.has_key?(query_params, :temporal_range) -> - # ListObjectVersions: retrieve versioned objects within a time range - list_object_versions(peer_info, bucket, query_params, limit, timeout) - - :document in modalities && Map.has_key?(query_params, :text_query) -> - # ListObjectsV2 with prefix matching, then filter by metadata - text = query_params.text_query - list_objects_with_prefix(peer_info, bucket, text, limit, timeout) - - :provenance in modalities -> - # List objects from the audit/provenance prefix - provenance_prefix = Map.get(config, :provenance_prefix, "provenance/") - list_objects_with_prefix(peer_info, bucket, provenance_prefix, limit, timeout) - - :semantic in modalities && Map.has_key?(query_params, :filters) -> - # ListObjectsV2 + HeadObject to filter by metadata/tags - list_and_filter_by_metadata(peer_info, bucket, query_params.filters, limit, timeout) - - true -> - # Default: list all objects in the bucket - list_objects_with_prefix(peer_info, bucket, "", limit, timeout) - end - end - - defp list_objects_with_prefix(peer_info, bucket, prefix, limit, timeout) do - url = build_bucket_url(peer_info, bucket) - headers = auth_headers(peer_info.adapter_config) - - query_params = %{ - "list-type" => "2", - "prefix" => prefix, - "max-keys" => to_string(limit) - } - - query_string = URI.encode_query(query_params) - full_url = "#{url}?#{query_string}" - - case Req.get(full_url, headers: headers, receive_timeout: timeout) do - {:ok, %Req.Response{status: 200, body: body}} -> - objects = parse_list_objects_response(body) - {:ok, objects} - - {:ok, %Req.Response{status: status, body: body}} -> - error_msg = extract_s3_error(body) || "HTTP #{status}" - Logger.warning("ObjectStorage adapter: ListObjects failed: #{error_msg}") - {:error, {:s3_error, status, error_msg}} - - {:error, reason} -> - {:error, reason} - end - end - - defp list_object_versions(peer_info, bucket, query_params, limit, timeout) do - config = peer_info.adapter_config - url = build_bucket_url(peer_info, bucket) - headers = auth_headers(config) - - range = query_params.temporal_range - prefix = Map.get(config, :prefix, "") - - params = %{ - "versions" => "", - "prefix" => prefix, - "max-keys" => to_string(limit) - } - - query_string = URI.encode_query(params) - full_url = "#{url}?#{query_string}" - - case Req.get(full_url, headers: headers, receive_timeout: timeout) do - {:ok, %Req.Response{status: 200, body: body}} -> - versions = parse_list_versions_response(body, range) - {:ok, versions} - - {:ok, %Req.Response{status: status, body: body}} -> - error_msg = extract_s3_error(body) || "HTTP #{status}" - {:error, {:s3_error, status, error_msg}} - - {:error, reason} -> - {:error, reason} - end - end - - defp list_and_filter_by_metadata(peer_info, bucket, filters, limit, timeout) do - # S3 does not support server-side metadata filtering. - # Strategy: list objects, then HEAD each to check metadata. - # This is expensive — limited to `limit * 2` candidates. - candidate_limit = min(limit * 2, 1000) - - case list_objects_with_prefix(peer_info, bucket, "", candidate_limit, timeout) do - {:ok, objects} -> - filtered = - objects - |> Enum.filter(fn obj -> - metadata = obj["metadata"] || %{} - tags = obj["tags"] || %{} - combined = Map.merge(metadata, tags) - - Enum.all?(filters, fn {key, value} -> - combined[to_string(key)] == to_string(value) - end) - end) - |> Enum.take(limit) - - {:ok, filtered} - - {:error, reason} -> - {:error, reason} - end - end - - # --------------------------------------------------------------------------- - # Private — Response Parsing - # --------------------------------------------------------------------------- - - defp parse_list_objects_response(body) when is_map(body) do - # JSON response (some S3 proxies return JSON) - contents = body["Contents"] || body["contents"] || [] - Enum.map(List.wrap(contents), &normalise_s3_object/1) - end - - defp parse_list_objects_response(body) when is_binary(body) do - # XML response (native S3/MinIO format) - # Extract elements from the XML - ~r/([^<]+)<\/Key>/ - |> Regex.scan(body) - |> Enum.map(fn [_full, key] -> - size = extract_xml_field(body, key, "Size") - last_modified = extract_xml_field(body, key, "LastModified") - - %{ - "key" => key, - "size" => size, - "last_modified" => last_modified, - "etag" => extract_xml_field(body, key, "ETag") - } - end) - end - - defp parse_list_objects_response(_body), do: [] - - defp parse_list_versions_response(body, range) when is_binary(body) do - # Parse XML ListObjectVersions response and filter by time range - start_time = range[:start] || range["start"] - end_time = range[:end] || range["end"] - - ~r/[\s\S]*?([^<]+)<\/Key>[\s\S]*?([^<]+)<\/LastModified>[\s\S]*?([^<]+)<\/VersionId>[\s\S]*?<\/Version>/ - |> Regex.scan(body) - |> Enum.map(fn [_full, key, modified, version_id] -> - %{ - "key" => key, - "last_modified" => modified, - "version_id" => version_id, - "is_version" => true - } - end) - |> Enum.filter(fn obj -> - modified = obj["last_modified"] || "" - - cond do - is_nil(start_time) and is_nil(end_time) -> true - is_nil(start_time) -> modified <= end_time - is_nil(end_time) -> modified >= start_time - true -> modified >= start_time and modified <= end_time - end - end) - end - - defp parse_list_versions_response(body, _range) when is_map(body) do - body["Versions"] || body["versions"] || [] - end - - defp parse_list_versions_response(_body, _range), do: [] - - defp extract_xml_field(xml, key, field) do - # Simple XML field extraction near a specific key - # This is a best-effort parser for S3 ListObjects XML - pattern = ~r/<#{field}>([^<]+)<\/#{field}>/ - - case Regex.scan(pattern, xml) do - [] -> nil - matches -> matches |> List.first() |> List.last() - end - end - - defp normalise_s3_object(obj) do - %{ - "key" => obj["Key"] || obj["key"] || "unknown", - "size" => obj["Size"] || obj["size"] || 0, - "last_modified" => obj["LastModified"] || obj["last_modified"] || "", - "etag" => obj["ETag"] || obj["etag"] || "", - "metadata" => obj["Metadata"] || obj["metadata"] || %{} - } - end - - defp extract_s3_error(body) when is_binary(body) do - case Regex.run(~r/([^<]+)<\/Message>/, body) do - [_, message] -> message - _ -> nil - end - end - - defp extract_s3_error(body) when is_map(body) do - body["Error"] || body["error"] || body["Message"] || body["message"] - end - - defp extract_s3_error(_), do: nil - - # --------------------------------------------------------------------------- - # Private — Helpers - # --------------------------------------------------------------------------- - - defp build_bucket_url(peer_info, bucket) do - config = peer_info.adapter_config - backend = Map.get(config, :backend, :minio) - - case backend do - :s3 -> - # Virtual-hosted-style URL for S3 - region = Map.get(config, :region, "us-east-1") - "https://#{bucket}.s3.#{region}.amazonaws.com" - - :minio -> - # Path-style URL for MinIO - "#{peer_info.endpoint}/#{bucket}" - - _ -> - "#{peer_info.endpoint}/#{bucket}" - end - end - - defp extract_object_id(obj) do - key = obj["key"] || obj["Key"] || "unknown" - - # Strip common prefixes and extensions to get a clean ID - key - |> String.replace(~r/^(octads|entities|objects)\//, "") - |> String.replace(~r/\.(json|cbor|bin)$/, "") - end - - defp parse_score(obj) do - case obj["score"] do - score when is_number(score) -> score / 1 - _ -> 0.0 - end - end - - defp maybe_add(list, item, true), do: list ++ [item] - defp maybe_add(list, _item, false), do: list - - defp auth_headers(config) do - case Map.get(config, :auth, :none) do - {:basic, user, pass} -> - encoded = Base.encode64("#{user}:#{pass}") - [{"Authorization", "Basic #{encoded}"}] - - {:bearer, token} -> - [{"Authorization", "Bearer #{token}"}] - - {:api_key, key} -> - [{"X-API-Key", key}] - - _ -> - # For S3/MinIO, proper auth requires AWS Signature V4 signing. - # In production, use an S3-aware HTTP client or middleware. - # For federation stubs, we pass access_key via header if available. - access_key = Map.get(config, :access_key) - - if access_key do - [{"X-Amz-Access-Key", access_key}] - else - [] - end - end - end -end diff --git a/verisimdb/elixir-orchestration/lib/verisim/federation/adapters/postgresql.ex b/verisimdb/elixir-orchestration/lib/verisim/federation/adapters/postgresql.ex deleted file mode 100644 index 01607967..00000000 --- a/verisimdb/elixir-orchestration/lib/verisim/federation/adapters/postgresql.ex +++ /dev/null @@ -1,431 +0,0 @@ -# SPDX-License-Identifier: MPL-2.0 - -defmodule VeriSim.Federation.Adapters.PostgreSQL do - @moduledoc """ - Federation adapter for PostgreSQL (with pgvector and PostGIS extensions). - - Translates VeriSimDB modality queries into SQL and executes them via - the PostgreSQL wire protocol (Postgrex). Supports rich modality mapping - when optional extensions are installed. - - ## Modality Mapping - - | VeriSimDB Modality | PostgreSQL Capability | Extension Required | - |--------------------|---------------------------|--------------------| - | `:document` | `tsvector` + GIN index | Built-in | - | `:vector` | `pgvector` cosine/L2 | pgvector | - | `:semantic` | JSONB columns | Built-in | - | `:temporal` | `tstzrange`, timestamps | Built-in | - | `:spatial` | PostGIS `geometry`/`geography` | PostGIS | - | `:graph` | Recursive CTEs | Built-in | - | `:provenance` | Audit table / triggers | Built-in | - - The `:tensor` modality has no direct PostgreSQL mapping and is not - supported by this adapter. - - ## Configuration - - %{ - host: "postgres.internal", - port: 5432, - database: "verisimdb", - schema: "public", - table: "octads", - auth: {:basic, "verisim", "password"}, - extensions: [:pgvector, :postgis] # Optional: declares installed extensions - } - - ## Connection Management - - This adapter uses Req for HTTP-based PostgreSQL proxies (e.g., PostgREST, - Supabase) or can be extended to use Postgrex for direct wire protocol. - The HTTP approach is used for consistency with other federation adapters - and to avoid adding Postgrex as a required dependency. - - For direct Postgrex connections (higher performance), configure: - - %{ - protocol: :wire, - host: "localhost", - port: 5432, - ... - } - """ - - @behaviour VeriSim.Federation.Adapter - - require Logger - - @default_timeout 10_000 - - # --------------------------------------------------------------------------- - # Callbacks - # --------------------------------------------------------------------------- - - @impl true - def connect(peer_info) do - case health_check(peer_info) do - {:ok, _latency} -> :ok - {:error, reason} -> {:error, reason} - end - end - - @impl true - def query(peer_info, query_params, opts \\ []) do - timeout = Keyword.get(opts, :timeout, @default_timeout) - modalities = Map.get(query_params, :modalities, []) - limit = Map.get(query_params, :limit, 100) - - start = System.monotonic_time(:millisecond) - - {sql, params} = build_sql(modalities, query_params, limit, peer_info) - - result = execute_sql(peer_info, sql, params, timeout) - - elapsed = System.monotonic_time(:millisecond) - start - - case result do - {:ok, raw_results} -> - normalised = - raw_results - |> translate_results(peer_info) - |> Enum.map(fn r -> Map.put(r, :response_time_ms, elapsed) end) - - {:ok, normalised} - - {:error, reason} -> - {:error, reason} - end - rescue - e -> - Logger.warning( - "PostgreSQL adapter: exception querying #{peer_info.store_id}: #{inspect(e)}" - ) - - {:error, {:exception, e}} - end - - @impl true - def health_check(peer_info) do - _config = peer_info.adapter_config - start = System.monotonic_time(:millisecond) - - # Health check: execute a trivial query to verify the connection - result = execute_sql(peer_info, "SELECT 1 AS ok", %{}, 5_000) - - case result do - {:ok, _} -> - elapsed = System.monotonic_time(:millisecond) - start - {:ok, elapsed} - - {:error, reason} -> - Logger.warning( - "PostgreSQL adapter: health check failed for #{peer_info.store_id}: #{inspect(reason)}" - ) - - {:error, reason} - end - rescue - e -> {:error, {:exception, e}} - end - - @impl true - def supported_modalities(adapter_config) do - extensions = Map.get(adapter_config, :extensions, []) - - base = [:document, :semantic, :temporal, :graph, :provenance] - - base - |> maybe_add(:vector, :pgvector in extensions) - |> maybe_add(:spatial, :postgis in extensions) - end - - @impl true - def translate_results(raw_results, peer_info) do - raw_results - |> List.wrap() - |> Enum.map(fn row -> - %{ - source_store: peer_info.store_id, - octad_id: row["id"] || row["entity_id"] || row["_key"] || "unknown", - score: parse_score(row), - drifted: false, - data: row, - response_time_ms: 0 - } - end) - end - - # --------------------------------------------------------------------------- - # Private — SQL Query Builder - # --------------------------------------------------------------------------- - - defp build_sql(modalities, query_params, limit, peer_info) do - config = peer_info.adapter_config - schema = Map.get(config, :schema, "public") - table = Map.get(config, :table, "octads") - qualified_table = "#{schema}.#{table}" - - cond do - :vector in modalities && Map.has_key?(query_params, :vector_query) -> - # pgvector cosine similarity search - embedding = query_params.vector_query - embedding_str = "[" <> Enum.join(embedding, ",") <> "]" - - sql = """ - SELECT *, - 1 - (embedding <=> $1::vector) AS score - FROM #{qualified_table} - ORDER BY embedding <=> $1::vector - LIMIT $2 - """ - - {sql, %{"$1" => embedding_str, "$2" => limit}} - - :document in modalities && Map.has_key?(query_params, :text_query) -> - # Full-text search via tsvector - sql = """ - SELECT *, - ts_rank_cd(search_vector, plainto_tsquery('english', $1)) AS score - FROM #{qualified_table} - WHERE search_vector @@ plainto_tsquery('english', $1) - ORDER BY score DESC - LIMIT $2 - """ - - {sql, %{"$1" => query_params.text_query, "$2" => limit}} - - :spatial in modalities && Map.has_key?(query_params, :spatial_bounds) -> - # PostGIS bounding box query - bounds = query_params.spatial_bounds - - sql = """ - SELECT *, - 0.0 AS score - FROM #{qualified_table} - WHERE ST_Within( - geom, - ST_MakeEnvelope($1, $2, $3, $4, 4326) - ) - LIMIT $5 - """ - - {sql, %{ - "$1" => bounds[:min_lon] || bounds["min_lon"] || 0.0, - "$2" => bounds[:min_lat] || bounds["min_lat"] || 0.0, - "$3" => bounds[:max_lon] || bounds["max_lon"] || 0.0, - "$4" => bounds[:max_lat] || bounds["max_lat"] || 0.0, - "$5" => limit - }} - - :graph in modalities && Map.has_key?(query_params, :graph_pattern) -> - # Recursive CTE for graph traversal - edges_table = Map.get(config, :edges_table, "#{schema}.edges") - - sql = """ - WITH RECURSIVE traversal AS ( - SELECT id, 1 AS depth - FROM #{qualified_table} - WHERE id = $1 - - UNION ALL - - SELECT e.target_id, t.depth + 1 - FROM #{edges_table} e - INNER JOIN traversal t ON e.source_id = t.id - WHERE t.depth < 3 - ) - SELECT h.*, 0.0 AS score - FROM #{qualified_table} h - INNER JOIN traversal t ON h.id = t.id - LIMIT $2 - """ - - {sql, %{"$1" => query_params.graph_pattern, "$2" => limit}} - - :temporal in modalities && Map.has_key?(query_params, :temporal_range) -> - range = query_params.temporal_range - - sql = """ - SELECT *, 0.0 AS score - FROM #{qualified_table} - WHERE created_at >= $1::timestamptz - AND created_at <= $2::timestamptz - ORDER BY created_at DESC - LIMIT $3 - """ - - {sql, %{ - "$1" => range[:start] || range["start"] || "", - "$2" => range[:end] || range["end"] || "", - "$3" => limit - }} - - :provenance in modalities -> - audit_table = Map.get(config, :audit_table, "#{schema}.audit_log") - - sql = """ - SELECT *, 0.0 AS score - FROM #{audit_table} - ORDER BY event_time DESC - LIMIT $1 - """ - - {sql, %{"$1" => limit}} - - true -> - # Default: paginated listing - sql = """ - SELECT *, 0.0 AS score - FROM #{qualified_table} - ORDER BY id ASC - LIMIT $1 - """ - - {sql, %{"$1" => limit}} - end - end - - defp execute_sql(peer_info, sql, params, timeout) do - config = peer_info.adapter_config - protocol = Map.get(config, :protocol, :http) - - case protocol do - :http -> - execute_via_http(peer_info, sql, params, timeout) - - :wire -> - # Direct Postgrex connection — requires Postgrex in mix.exs - execute_via_postgrex(peer_info, sql, params, timeout) - end - end - - # HTTP-based execution (PostgREST / pg-gateway / custom endpoint) - defp execute_via_http(peer_info, sql, params, timeout) do - url = "#{peer_info.endpoint}/query" - headers = auth_headers(peer_info.adapter_config) - - body = %{ - "query" => sql, - "params" => params - } - - case Req.post(url, json: body, headers: headers, receive_timeout: timeout) do - {:ok, %Req.Response{status: status, body: resp_body}} when status in 200..299 -> - rows = resp_body["rows"] || resp_body["result"] || resp_body - {:ok, List.wrap(rows)} - - {:ok, %Req.Response{status: status, body: resp_body}} -> - error_msg = resp_body["error"] || resp_body["message"] || "HTTP #{status}" - Logger.warning("PostgreSQL adapter: query failed: #{error_msg}") - {:error, {:sql_error, status, error_msg}} - - {:error, reason} -> - {:error, reason} - end - end - - # Direct Postgrex wire protocol — only used when :protocol is :wire. - # Requires Postgrex as an optional dependency in mix.exs. - # Falls back to HTTP if Postgrex is not available. - defp execute_via_postgrex(peer_info, sql, params, _timeout) do - config = peer_info.adapter_config - - postgrex_opts = [ - hostname: Map.get(config, :host, "localhost"), - port: Map.get(config, :port, 5432), - database: Map.get(config, :database, "verisimdb"), - username: extract_username(config), - password: extract_password(config) - ] - - # Use dynamic dispatch to avoid compile-time dependency on Postgrex. - # If Postgrex is not in mix.exs, the UndefinedFunctionError rescue - # below catches it and falls back to HTTP. - postgrex_mod = Module.concat([Postgrex]) - - case apply(postgrex_mod, :start_link, [postgrex_opts]) do - {:ok, conn} -> - param_values = params |> Map.values() - - case apply(postgrex_mod, :query, [conn, sql, param_values]) do - {:ok, result} -> - # Postgrex.Result has :columns and :rows fields - columns = Map.get(result, :columns, []) - rows = Map.get(result, :rows, []) - - results = - Enum.map(rows, fn row -> - columns |> Enum.zip(row) |> Map.new() - end) - - GenServer.stop(conn) - {:ok, results} - - {:error, reason} -> - GenServer.stop(conn) - {:error, {:postgrex_error, reason}} - end - - {:error, reason} -> - Logger.warning( - "PostgreSQL adapter: Postgrex connection failed for #{peer_info.store_id}, " <> - "falling back to HTTP: #{inspect(reason)}" - ) - - execute_via_http(peer_info, sql, params, @default_timeout) - end - rescue - UndefinedFunctionError -> - Logger.debug( - "PostgreSQL adapter: Postgrex not available, using HTTP for #{peer_info.store_id}" - ) - - execute_via_http(peer_info, sql, params, @default_timeout) - end - - # --------------------------------------------------------------------------- - # Private — Helpers - # --------------------------------------------------------------------------- - - defp parse_score(row) do - case row["score"] do - score when is_number(score) -> score / 1 - _ -> 0.0 - end - end - - defp maybe_add(list, item, true), do: list ++ [item] - defp maybe_add(list, _item, false), do: list - - defp extract_username(config) do - case Map.get(config, :auth, :none) do - {:basic, user, _pass} -> user - _ -> "postgres" - end - end - - defp extract_password(config) do - case Map.get(config, :auth, :none) do - {:basic, _user, pass} -> pass - _ -> "" - end - end - - defp auth_headers(config) do - case Map.get(config, :auth, :none) do - {:basic, user, pass} -> - encoded = Base.encode64("#{user}:#{pass}") - [{"Authorization", "Basic #{encoded}"}] - - {:bearer, token} -> - [{"Authorization", "Bearer #{token}"}] - - {:api_key, key} -> - [{"X-API-Key", key}] - - _ -> - [] - end - end -end diff --git a/verisimdb/elixir-orchestration/lib/verisim/federation/adapters/redis.ex b/verisimdb/elixir-orchestration/lib/verisim/federation/adapters/redis.ex deleted file mode 100644 index 053a5eb1..00000000 --- a/verisimdb/elixir-orchestration/lib/verisim/federation/adapters/redis.ex +++ /dev/null @@ -1,380 +0,0 @@ -# SPDX-License-Identifier: MPL-2.0 - -defmodule VeriSim.Federation.Adapters.Redis do - @moduledoc """ - Federation adapter for Redis (with RedisSearch, RedisGraph, RedisJSON, and - RedisTimeSeries modules). - - Translates VeriSimDB modality queries into Redis command sequences and - executes them via the Redis HTTP API (RedisInsight REST, Redis Cloud API, - or a custom HTTP-to-Redis bridge). Modality support depends on which Redis - modules are installed on the target instance. - - ## Modality Mapping - - | VeriSimDB Modality | Redis Capability | Module Required | - |--------------------|--------------------------|---------------------| - | `:graph` | RedisGraph (Cypher) | RedisGraph | - | `:vector` | Vector Similarity Search | RediSearch 2.4+ | - | `:document` | Full-text index (FT) | RediSearch | - | `:temporal` | Time-series data | RedisTimeSeries | - | `:provenance` | Redis Streams | Built-in (5.0+) | - | `:semantic` | JSON documents | RedisJSON | - - The `:tensor` and `:spatial` modalities have no direct Redis mapping and - are not supported by this adapter. Spatial queries could be partially - served via RediSearch GEO filters if needed in future. - - ## Configuration - - %{ - host: "redis.internal", - port: 6379, - database: 0, - auth: {:basic, "default", "password"}, - modules: [:redisearch, :redisgraph, :redisjson, :redistimeseries] - } - - ## Module Detection - - The `supported_modalities/1` callback checks the `modules` list in the - adapter configuration to determine which modalities are available. This - mirrors the PostgreSQL adapter's extension-based modality detection. - - ## Redis HTTP Bridge - - Redis does not natively expose an HTTP API. This adapter assumes one of: - - Redis Cloud REST API - - RedisInsight API - - A custom HTTP-to-Redis proxy (e.g., webdis, redis-rest) - - Commands are sent as JSON arrays to `POST /command` or equivalent. - """ - - @behaviour VeriSim.Federation.Adapter - - require Logger - - @default_timeout 10_000 - - # --------------------------------------------------------------------------- - # Callbacks - # --------------------------------------------------------------------------- - - @impl true - def connect(peer_info) do - case health_check(peer_info) do - {:ok, _latency} -> :ok - {:error, reason} -> {:error, reason} - end - end - - @impl true - def query(peer_info, query_params, opts \\ []) do - timeout = Keyword.get(opts, :timeout, @default_timeout) - modalities = Map.get(query_params, :modalities, []) - limit = Map.get(query_params, :limit, 100) - - start = System.monotonic_time(:millisecond) - - commands = build_commands(modalities, query_params, limit, peer_info) - - result = execute_commands(peer_info, commands, timeout) - - elapsed = System.monotonic_time(:millisecond) - start - - case result do - {:ok, raw_results} -> - normalised = - raw_results - |> translate_results(peer_info) - |> Enum.map(fn r -> Map.put(r, :response_time_ms, elapsed) end) - - {:ok, normalised} - - {:error, reason} -> - {:error, reason} - end - rescue - e -> - Logger.warning( - "Redis adapter: exception querying #{peer_info.store_id}: #{inspect(e)}" - ) - - {:error, {:exception, e}} - end - - @impl true - def health_check(peer_info) do - start = System.monotonic_time(:millisecond) - headers = auth_headers(peer_info.adapter_config) - - # Send PING command via HTTP bridge - url = "#{peer_info.endpoint}/command" - - body = %{"command" => ["PING"]} - - case Req.post(url, json: body, headers: headers, receive_timeout: 5_000) do - {:ok, %Req.Response{status: status, body: resp_body}} when status in 200..299 -> - # Redis PING returns "PONG" or {"result": "PONG"} - response = if is_map(resp_body), do: resp_body["result"], else: resp_body - - if response in ["PONG", "pong"] do - elapsed = System.monotonic_time(:millisecond) - start - {:ok, elapsed} - else - {:error, {:unexpected_ping_response, response}} - end - - {:ok, %Req.Response{status: status}} -> - {:error, {:unhealthy, status}} - - {:error, reason} -> - {:error, reason} - end - rescue - e -> {:error, {:exception, e}} - end - - @impl true - def supported_modalities(adapter_config) do - modules = Map.get(adapter_config, :modules, []) - - # Provenance via Redis Streams is built-in (Redis 5.0+), always available - base = [:provenance] - - base - |> maybe_add(:vector, :redisearch in modules) - |> maybe_add(:document, :redisearch in modules) - |> maybe_add(:graph, :redisgraph in modules) - |> maybe_add(:temporal, :redistimeseries in modules) - |> maybe_add(:semantic, :redisjson in modules) - end - - @impl true - def translate_results(raw_results, peer_info) do - raw_results - |> List.wrap() - |> Enum.map(fn row -> - # Redis results come in varied formats depending on the command. - # Normalise to a consistent map structure. - {id, data} = extract_id_and_data(row) - - %{ - source_store: peer_info.store_id, - octad_id: id, - score: parse_score(row), - drifted: false, - data: data, - response_time_ms: 0 - } - end) - end - - # --------------------------------------------------------------------------- - # Private — Redis Command Builder - # --------------------------------------------------------------------------- - - defp build_commands(modalities, query_params, limit, peer_info) do - config = peer_info.adapter_config - index_name = Map.get(config, :index_name, "octad_idx") - - cond do - :vector in modalities && Map.has_key?(query_params, :vector_query) -> - # RediSearch Vector Similarity Search (VSS) - # FT.SEARCH index_name "*=>[KNN limit @vec_field $BLOB]" PARAMS 2 BLOB - embedding = query_params.vector_query - vector_field = Map.get(config, :vector_field, "embedding") - blob = encode_vector_blob(embedding) - - %{ - "command" => [ - "FT.SEARCH", - index_name, - "*=>[KNN #{limit} @#{vector_field} $BLOB]", - "PARAMS", - "2", - "BLOB", - blob, - "SORTBY", - "__#{vector_field}_score", - "LIMIT", - "0", - to_string(limit), - "DIALECT", - "2" - ] - } - - :document in modalities && Map.has_key?(query_params, :text_query) -> - # RediSearch full-text query - text = query_params.text_query - - %{ - "command" => [ - "FT.SEARCH", - index_name, - text, - "LIMIT", - "0", - to_string(limit), - "WITHSCORES" - ] - } - - :graph in modalities && Map.has_key?(query_params, :graph_pattern) -> - # RedisGraph Cypher query - graph_key = Map.get(config, :graph_key, "octad_graph") - start_vertex = query_params.graph_pattern - - cypher = """ - MATCH (n)-[r*1..3]-(m) - WHERE n.id = '#{start_vertex}' - RETURN m - LIMIT #{limit} - """ - - %{ - "command" => ["GRAPH.QUERY", graph_key, String.trim(cypher)] - } - - :temporal in modalities && Map.has_key?(query_params, :temporal_range) -> - # RedisTimeSeries range query - range = query_params.temporal_range - ts_key = Map.get(config, :timeseries_key, "octad_ts") - start_ts = range[:start] || range["start"] || "-" - end_ts = range[:end] || range["end"] || "+" - - %{ - "command" => ["TS.RANGE", ts_key, to_string(start_ts), to_string(end_ts), "COUNT", to_string(limit)] - } - - :provenance in modalities -> - # Redis Streams — XREVRANGE for latest events - stream_key = Map.get(config, :stream_key, "octad_provenance") - - %{ - "command" => ["XREVRANGE", stream_key, "+", "-", "COUNT", to_string(limit)] - } - - :semantic in modalities && Map.has_key?(query_params, :filters) -> - # RedisJSON — JSON.GET with path filters - filters = query_params.filters - json_key_pattern = Map.get(config, :json_key_pattern, "octad:*") - - # Use FT.SEARCH with JSON filter if RediSearch is available - filter_clauses = - filters - |> Enum.map(fn {field, value} -> "@#{field}:{#{value}}" end) - |> Enum.join(" ") - - %{ - "command" => [ - "FT.SEARCH", - index_name, - filter_clauses, - "LIMIT", - "0", - to_string(limit) - ], - "_json_key_pattern" => json_key_pattern - } - - true -> - # Default: scan keys matching the collection pattern - key_pattern = Map.get(config, :key_pattern, "octad:*") - - %{ - "command" => ["SCAN", "0", "MATCH", key_pattern, "COUNT", to_string(limit)] - } - end - end - - defp execute_commands(peer_info, commands, timeout) do - url = "#{peer_info.endpoint}/command" - headers = auth_headers(peer_info.adapter_config) - - # Send the command(s) to the Redis HTTP bridge - body = Map.take(commands, ["command"]) - - case Req.post(url, json: body, headers: headers, receive_timeout: timeout) do - {:ok, %Req.Response{status: status, body: resp_body}} when status in 200..299 -> - results = parse_redis_response(resp_body) - {:ok, results} - - {:ok, %Req.Response{status: status, body: resp_body}} -> - error_msg = resp_body["error"] || "HTTP #{status}" - Logger.warning("Redis adapter: command failed: #{error_msg}") - {:error, {:redis_error, status, error_msg}} - - {:error, reason} -> - {:error, reason} - end - end - - # --------------------------------------------------------------------------- - # Private — Helpers - # --------------------------------------------------------------------------- - - defp parse_redis_response(resp_body) when is_map(resp_body) do - # Redis HTTP bridges typically return {"result": [...]} or {"data": [...]} - result = resp_body["result"] || resp_body["data"] || resp_body - List.wrap(result) - end - - defp parse_redis_response(resp_body) when is_list(resp_body), do: resp_body - defp parse_redis_response(resp_body), do: [resp_body] - - defp extract_id_and_data(row) when is_map(row) do - id = row["id"] || row["_id"] || row["key"] || "unknown" - {id, row} - end - - defp extract_id_and_data(row) when is_binary(row) do - {row, %{"raw" => row}} - end - - defp extract_id_and_data(row) do - {"unknown", %{"raw" => inspect(row)}} - end - - defp parse_score(row) when is_map(row) do - case row["score"] || row["__score"] do - score when is_number(score) -> score / 1 - score when is_binary(score) -> String.to_float(score) - _ -> 0.0 - end - rescue - _ -> 0.0 - end - - defp parse_score(_row), do: 0.0 - - defp encode_vector_blob(embedding) when is_list(embedding) do - # Encode float vector as Base64-encoded binary blob for RediSearch VSS - embedding - |> Enum.map(fn f -> <> end) - |> IO.iodata_to_binary() - |> Base.encode64() - end - - defp maybe_add(list, item, true), do: list ++ [item] - defp maybe_add(list, _item, false), do: list - - defp auth_headers(config) do - case Map.get(config, :auth, :none) do - {:basic, user, pass} -> - encoded = Base.encode64("#{user}:#{pass}") - [{"Authorization", "Basic #{encoded}"}, {"Content-Type", "application/json"}] - - {:bearer, token} -> - [{"Authorization", "Bearer #{token}"}, {"Content-Type", "application/json"}] - - {:api_key, key} -> - [{"X-API-Key", key}, {"Content-Type", "application/json"}] - - _ -> - [{"Content-Type", "application/json"}] - end - end -end diff --git a/verisimdb/elixir-orchestration/lib/verisim/federation/adapters/sqlite.ex b/verisimdb/elixir-orchestration/lib/verisim/federation/adapters/sqlite.ex deleted file mode 100644 index 308de062..00000000 --- a/verisimdb/elixir-orchestration/lib/verisim/federation/adapters/sqlite.ex +++ /dev/null @@ -1,331 +0,0 @@ -# SPDX-License-Identifier: MPL-2.0 - -defmodule VeriSim.Federation.Adapters.SQLite do - @moduledoc """ - Federation adapter for SQLite (with sqlite-vss and FTS5 extensions). - - Translates VeriSimDB modality queries into SQLite SQL and executes them - via an HTTP endpoint (a lightweight sidecar wrapping the SQLite C library). - SQLite is the world's most deployed database and excels at embedded, - edge, and single-node workloads with zero configuration. - - ## Modality Mapping - - | VeriSimDB Modality | SQLite Capability | Extension Required | - |--------------------|----------------------------|-------------------------| - | `:graph` | Recursive CTEs | Built-in | - | `:vector` | sqlite-vss (HNSW) | sqlite-vss | - | `:document` | FTS5 full-text search | fts5 (built-in compile) | - | `:temporal` | datetime() / julianday() | Built-in | - | `:semantic` | JSON1 functions | json1 (built-in) | - - SQLite does not support tensor operations, provenance chains, or - geospatial queries natively, so `:tensor`, `:provenance`, and `:spatial` - modalities are not supported. SpatiaLite could add spatial support in - future if needed. - - ## Configuration - - %{ - path: "/data/verisimdb.sqlite3", - extensions: [:vss, :fts5], - table: "octads" - } - - ## HTTP Endpoint - - SQLite has no network protocol. This adapter assumes a lightweight HTTP - wrapper (e.g., sqlite-web, sqld from Turso, or a custom Rust sidecar) - that accepts SQL via `POST /query` and returns JSON results. - - ## Edge Federation - - SQLite adapters are ideal for edge federation — running VeriSimDB - modality queries on devices, embedded systems, or serverless functions - where a full database server would be too heavy. - """ - - @behaviour VeriSim.Federation.Adapter - - require Logger - - @default_timeout 10_000 - - # --------------------------------------------------------------------------- - # Callbacks - # --------------------------------------------------------------------------- - - @impl true - def connect(peer_info) do - case health_check(peer_info) do - {:ok, _latency} -> :ok - {:error, reason} -> {:error, reason} - end - end - - @impl true - def query(peer_info, query_params, opts \\ []) do - timeout = Keyword.get(opts, :timeout, @default_timeout) - modalities = Map.get(query_params, :modalities, []) - limit = Map.get(query_params, :limit, 100) - - start = System.monotonic_time(:millisecond) - - {sql, params} = build_sql(modalities, query_params, limit, peer_info) - - result = execute_sql(peer_info, sql, params, timeout) - - elapsed = System.monotonic_time(:millisecond) - start - - case result do - {:ok, raw_results} -> - normalised = - raw_results - |> translate_results(peer_info) - |> Enum.map(fn r -> Map.put(r, :response_time_ms, elapsed) end) - - {:ok, normalised} - - {:error, reason} -> - {:error, reason} - end - rescue - e -> - Logger.warning( - "SQLite adapter: exception querying #{peer_info.store_id}: #{inspect(e)}" - ) - - {:error, {:exception, e}} - end - - @impl true - def health_check(peer_info) do - start = System.monotonic_time(:millisecond) - - # SQLite health check: execute a trivial query - result = execute_sql(peer_info, "SELECT 1 AS ok", %{}, 5_000) - - case result do - {:ok, _} -> - elapsed = System.monotonic_time(:millisecond) - start - {:ok, elapsed} - - {:error, reason} -> - Logger.warning( - "SQLite adapter: health check failed for #{peer_info.store_id}: #{inspect(reason)}" - ) - - {:error, reason} - end - rescue - e -> {:error, {:exception, e}} - end - - @impl true - def supported_modalities(adapter_config) do - extensions = Map.get(adapter_config, :extensions, []) - - # SQLite always supports: recursive CTEs (graph), datetime (temporal), - # and JSON1 (semantic — compiled in by default since SQLite 3.38) - base = [:graph, :temporal, :semantic] - - base - |> maybe_add(:vector, :vss in extensions) - |> maybe_add(:document, :fts5 in extensions) - end - - @impl true - def translate_results(raw_results, peer_info) do - raw_results - |> List.wrap() - |> Enum.map(fn row -> - %{ - source_store: peer_info.store_id, - octad_id: row["id"] || row["rowid"] || row["entity_id"] || "unknown", - score: parse_score(row), - drifted: false, - data: row, - response_time_ms: 0 - } - end) - end - - # --------------------------------------------------------------------------- - # Private — SQLite SQL Query Builder - # --------------------------------------------------------------------------- - - defp build_sql(modalities, query_params, limit, peer_info) do - config = peer_info.adapter_config - table = Map.get(config, :table, "octads") - - cond do - :vector in modalities && Map.has_key?(query_params, :vector_query) -> - # sqlite-vss: virtual table for vector similarity search - embedding = query_params.vector_query - embedding_json = Jason.encode!(embedding) - vss_table = Map.get(config, :vss_table, "vss_#{table}") - - sql = """ - SELECT h.*, v.distance AS score - FROM #{vss_table} v - INNER JOIN #{table} h ON h.rowid = v.rowid - WHERE vss_search(v.embedding, vss_search_params($1, $2)) - ORDER BY v.distance ASC - LIMIT $2 - """ - - {sql, %{"$1" => embedding_json, "$2" => limit}} - - :document in modalities && Map.has_key?(query_params, :text_query) -> - # FTS5: full-text search with MATCH syntax - text = query_params.text_query - fts_table = Map.get(config, :fts_table, "#{table}_fts") - - sql = """ - SELECT h.*, fts.rank AS score - FROM #{fts_table} fts - INNER JOIN #{table} h ON h.rowid = fts.rowid - WHERE #{fts_table} MATCH $1 - ORDER BY fts.rank - LIMIT $2 - """ - - {sql, %{"$1" => text, "$2" => limit}} - - :graph in modalities && Map.has_key?(query_params, :graph_pattern) -> - # Recursive CTE for graph traversal - edges_table = Map.get(config, :edges_table, "edges") - - sql = """ - WITH RECURSIVE traversal(id, depth) AS ( - SELECT id, 0 - FROM #{table} - WHERE id = $1 - - UNION ALL - - SELECT e.target_id, t.depth + 1 - FROM #{edges_table} e - INNER JOIN traversal t ON e.source_id = t.id - WHERE t.depth < 3 - ) - SELECT h.*, 0.0 AS score - FROM #{table} h - INNER JOIN traversal t ON h.id = t.id - LIMIT $2 - """ - - {sql, %{"$1" => query_params.graph_pattern, "$2" => limit}} - - :temporal in modalities && Map.has_key?(query_params, :temporal_range) -> - range = query_params.temporal_range - - sql = """ - SELECT *, 0.0 AS score - FROM #{table} - WHERE datetime(created_at) >= datetime($1) - AND datetime(created_at) <= datetime($2) - ORDER BY datetime(created_at) DESC - LIMIT $3 - """ - - {sql, %{ - "$1" => range[:start] || range["start"] || "", - "$2" => range[:end] || range["end"] || "", - "$3" => limit - }} - - :semantic in modalities && Map.has_key?(query_params, :filters) -> - # JSON1 extraction for structured metadata queries - filters = query_params.filters - - where_clauses = - filters - |> Enum.map(fn {field, value} -> - "json_extract(metadata, '$.#{field}') = '#{escape_sqlite(to_string(value))}'" - end) - |> Enum.join(" AND ") - - where_clause = if where_clauses == "", do: "1=1", else: where_clauses - - sql = """ - SELECT *, 0.0 AS score - FROM #{table} - WHERE #{where_clause} - LIMIT $1 - """ - - {sql, %{"$1" => limit}} - - true -> - # Default: paginated listing - sql = """ - SELECT *, 0.0 AS score - FROM #{table} - ORDER BY rowid ASC - LIMIT $1 - """ - - {sql, %{"$1" => limit}} - end - end - - defp execute_sql(peer_info, sql, params, timeout) do - url = "#{peer_info.endpoint}/query" - headers = auth_headers(peer_info.adapter_config) - - body = %{ - "query" => sql, - "params" => params - } - - case Req.post(url, json: body, headers: headers, receive_timeout: timeout) do - {:ok, %Req.Response{status: status, body: resp_body}} when status in 200..299 -> - rows = resp_body["rows"] || resp_body["result"] || resp_body["results"] || resp_body - {:ok, List.wrap(rows)} - - {:ok, %Req.Response{status: status, body: resp_body}} -> - error_msg = resp_body["error"] || resp_body["message"] || "HTTP #{status}" - Logger.warning("SQLite adapter: query failed: #{error_msg}") - {:error, {:sql_error, status, error_msg}} - - {:error, reason} -> - {:error, reason} - end - end - - # --------------------------------------------------------------------------- - # Private — Helpers - # --------------------------------------------------------------------------- - - defp parse_score(row) do - case row["score"] || row["rank"] || row["distance"] do - score when is_number(score) -> score / 1 - _ -> 0.0 - end - end - - defp maybe_add(list, item, true), do: list ++ [item] - defp maybe_add(list, _item, false), do: list - - defp escape_sqlite(str) when is_binary(str) do - String.replace(str, "'", "''") - end - - defp auth_headers(config) do - case Map.get(config, :auth, :none) do - {:basic, user, pass} -> - encoded = Base.encode64("#{user}:#{pass}") - [{"Authorization", "Basic #{encoded}"}, {"Content-Type", "application/json"}] - - {:bearer, token} -> - [{"Authorization", "Bearer #{token}"}, {"Content-Type", "application/json"}] - - {:api_key, key} -> - [{"X-API-Key", key}, {"Content-Type", "application/json"}] - - _ -> - [{"Content-Type", "application/json"}] - end - end -end diff --git a/verisimdb/elixir-orchestration/lib/verisim/federation/adapters/surrealdb.ex b/verisimdb/elixir-orchestration/lib/verisim/federation/adapters/surrealdb.ex deleted file mode 100644 index 2c56d490..00000000 --- a/verisimdb/elixir-orchestration/lib/verisim/federation/adapters/surrealdb.ex +++ /dev/null @@ -1,342 +0,0 @@ -# SPDX-License-Identifier: MPL-2.0 - -defmodule VeriSim.Federation.Adapters.SurrealDB do - @moduledoc """ - Federation adapter for SurrealDB. - - Translates VeriSimDB modality queries into SurrealQL and executes them - via the SurrealDB HTTP REST API. SurrealDB is a multi-model database - with native graph traversal, schemaless documents, and built-in - full-text search — making it a natural fit for several VeriSimDB modalities. - - ## Modality Mapping - - | VeriSimDB Modality | SurrealDB Capability | Extension/Feature Required | - |--------------------|------------------------------|----------------------------| - | `:graph` | RELATE / graph traversal | Built-in | - | `:document` | Full-text search (analyzers) | Built-in | - | `:temporal` | datetime type / duration | Built-in | - | `:semantic` | Schemaless nested records | Built-in | - - SurrealDB does not natively support vector similarity search, tensor - operations, provenance chains, or geospatial queries (though spatial - support is on the roadmap), so `:vector`, `:tensor`, `:provenance`, - and `:spatial` modalities are not supported. - - ## Configuration - - %{ - host: "surrealdb.internal", - port: 8000, - namespace: "verisim", - database: "production", - auth: {:basic, "root", "root"} - } - - ## SurrealDB HTTP API - - Queries are sent to `POST /sql` with SurrealQL as the request body. - The `NS` and `DB` headers select the namespace and database context. - Health checks hit `GET /health` which returns 200 when the server is - ready. - """ - - @behaviour VeriSim.Federation.Adapter - - require Logger - - @default_timeout 10_000 - - # --------------------------------------------------------------------------- - # Callbacks - # --------------------------------------------------------------------------- - - @impl true - def connect(peer_info) do - case health_check(peer_info) do - {:ok, _latency} -> :ok - {:error, reason} -> {:error, reason} - end - end - - @impl true - def query(peer_info, query_params, opts \\ []) do - timeout = Keyword.get(opts, :timeout, @default_timeout) - modalities = Map.get(query_params, :modalities, []) - limit = Map.get(query_params, :limit, 100) - - start = System.monotonic_time(:millisecond) - - surreal_ql = build_surreal_ql(modalities, query_params, limit, peer_info) - - result = execute_surreal_ql(peer_info, surreal_ql, timeout) - - elapsed = System.monotonic_time(:millisecond) - start - - case result do - {:ok, raw_results} -> - normalised = - raw_results - |> translate_results(peer_info) - |> Enum.map(fn r -> Map.put(r, :response_time_ms, elapsed) end) - - {:ok, normalised} - - {:error, reason} -> - {:error, reason} - end - rescue - e -> - Logger.warning( - "SurrealDB adapter: exception querying #{peer_info.store_id}: #{inspect(e)}" - ) - - {:error, {:exception, e}} - end - - @impl true - def health_check(peer_info) do - start = System.monotonic_time(:millisecond) - headers = auth_headers(peer_info.adapter_config) - - # SurrealDB health endpoint: GET /health returns 200 when ready - url = "#{peer_info.endpoint}/health" - - case Req.get(url, headers: headers, receive_timeout: 5_000) do - {:ok, %Req.Response{status: 200}} -> - elapsed = System.monotonic_time(:millisecond) - start - {:ok, elapsed} - - {:ok, %Req.Response{status: status}} -> - {:error, {:unhealthy, status}} - - {:error, reason} -> - {:error, reason} - end - rescue - e -> {:error, {:exception, e}} - end - - @impl true - def supported_modalities(_adapter_config) do - # SurrealDB supports graph, document, temporal, and semantic natively - [:graph, :document, :temporal, :semantic] - end - - @impl true - def translate_results(raw_results, peer_info) do - raw_results - |> List.wrap() - |> Enum.map(fn record -> - %{ - source_store: peer_info.store_id, - octad_id: extract_id(record), - score: parse_score(record), - drifted: false, - data: record, - response_time_ms: 0 - } - end) - end - - # --------------------------------------------------------------------------- - # Private — SurrealQL Query Builder - # --------------------------------------------------------------------------- - - defp build_surreal_ql(modalities, query_params, limit, peer_info) do - config = peer_info.adapter_config - table = Map.get(config, :table, "octads") - - cond do - :graph in modalities && Map.has_key?(query_params, :graph_pattern) -> - # SurrealDB graph traversal: record links and RELATE edges - start_vertex = query_params.graph_pattern - edge_table = Map.get(config, :edge_table, "connects") - max_depth = Map.get(config, :max_depth, 3) - - """ - SELECT * - FROM #{table}:#{escape_surreal(start_vertex)} - ->#{edge_table}->#{table} - LIMIT #{limit}; - """ - - :document in modalities && Map.has_key?(query_params, :text_query) -> - # SurrealDB full-text search via search analyzer - text = escape_surreal(query_params.text_query) - search_fields = Map.get(config, :search_fields, ["title", "body", "content"]) - - # Build OR conditions across searchable fields - conditions = - search_fields - |> Enum.map(fn field -> "string::contains(string::lowercase(#{field}), string::lowercase('#{text}'))" end) - |> Enum.join(" OR ") - - """ - SELECT *, search::score(0) AS score - FROM #{table} - WHERE #{conditions} - ORDER BY score DESC - LIMIT #{limit}; - """ - - :temporal in modalities && Map.has_key?(query_params, :temporal_range) -> - range = query_params.temporal_range - start_time = escape_surreal(range[:start] || range["start"] || "") - end_time = escape_surreal(range[:end] || range["end"] || "") - - """ - SELECT * - FROM #{table} - WHERE created_at >= d'#{start_time}' - AND created_at <= d'#{end_time}' - ORDER BY created_at DESC - LIMIT #{limit}; - """ - - :semantic in modalities && Map.has_key?(query_params, :filters) -> - # Schemaless nested record query - filters = query_params.filters - - where_clauses = - filters - |> Enum.map(fn {field, value} -> - "metadata.#{escape_surreal(to_string(field))} = '#{escape_surreal(to_string(value))}'" - end) - |> Enum.join(" AND ") - - where_clause = if where_clauses == "", do: "true", else: where_clauses - - """ - SELECT * - FROM #{table} - WHERE #{where_clause} - LIMIT #{limit}; - """ - - true -> - # Default: select all records from the table - """ - SELECT * - FROM #{table} - ORDER BY id ASC - LIMIT #{limit}; - """ - end - end - - defp execute_surreal_ql(peer_info, surreal_ql, timeout) do - config = peer_info.adapter_config - namespace = Map.get(config, :namespace, "verisim") - database = Map.get(config, :database, "production") - - url = "#{peer_info.endpoint}/sql" - - headers = - auth_headers(config) ++ - [ - {"NS", namespace}, - {"DB", database}, - {"Accept", "application/json"}, - {"Content-Type", "text/plain"} - ] - - case Req.post(url, body: surreal_ql, headers: headers, receive_timeout: timeout) do - {:ok, %Req.Response{status: 200, body: body}} -> - # SurrealDB returns an array of statement results: [{"result": [...], "status": "OK"}] - results = extract_surreal_results(body) - {:ok, results} - - {:ok, %Req.Response{status: status, body: body}} -> - error_msg = extract_surreal_error(body) || "HTTP #{status}" - Logger.warning("SurrealDB adapter: query failed: #{error_msg}") - {:error, {:surreal_error, status, error_msg}} - - {:error, reason} -> - {:error, reason} - end - end - - # --------------------------------------------------------------------------- - # Private — Helpers - # --------------------------------------------------------------------------- - - defp extract_surreal_results(body) when is_list(body) do - # SurrealDB returns [{%{"result" => [...], "status" => "OK"}}] - body - |> Enum.flat_map(fn statement -> - case statement do - %{"result" => results, "status" => "OK"} -> List.wrap(results) - %{"result" => results} -> List.wrap(results) - _ -> [] - end - end) - end - - defp extract_surreal_results(body) when is_map(body) do - body["result"] || [body] - end - - defp extract_surreal_results(_body), do: [] - - defp extract_surreal_error(body) when is_list(body) do - body - |> Enum.find_value(fn - %{"status" => status, "detail" => detail} when status != "OK" -> detail - %{"status" => status, "result" => result} when status != "OK" -> inspect(result) - _ -> nil - end) - end - - defp extract_surreal_error(body) when is_map(body) do - body["information"] || body["description"] || body["error"] - end - - defp extract_surreal_error(_), do: nil - - defp extract_id(record) do - case record["id"] do - # SurrealDB IDs are "table:id" format — extract the id part - id when is_binary(id) -> - case String.split(id, ":", parts: 2) do - [_table, record_id] -> record_id - _ -> id - end - - _ -> - record["_id"] || record["_key"] || "unknown" - end - end - - defp parse_score(record) do - case record["score"] do - score when is_number(score) -> score / 1 - _ -> 0.0 - end - end - - defp escape_surreal(str) when is_binary(str) do - str - |> String.replace("\\", "\\\\") - |> String.replace("'", "\\'") - end - - defp escape_surreal(str), do: to_string(str) - - defp auth_headers(config) do - case Map.get(config, :auth, :none) do - {:basic, user, pass} -> - encoded = Base.encode64("#{user}:#{pass}") - [{"Authorization", "Basic #{encoded}"}] - - {:bearer, token} -> - [{"Authorization", "Bearer #{token}"}] - - {:api_key, key} -> - [{"X-API-Key", key}] - - _ -> - [] - end - end -end diff --git a/verisimdb/elixir-orchestration/lib/verisim/federation/adapters/vector_db.ex b/verisimdb/elixir-orchestration/lib/verisim/federation/adapters/vector_db.ex deleted file mode 100644 index 161c809e..00000000 --- a/verisimdb/elixir-orchestration/lib/verisim/federation/adapters/vector_db.ex +++ /dev/null @@ -1,653 +0,0 @@ -# SPDX-License-Identifier: MPL-2.0 - -defmodule VeriSim.Federation.Adapters.VectorDB do - @moduledoc """ - Unified federation adapter for dedicated vector databases: Qdrant, Milvus, - and Weaviate. - - Rather than maintaining three separate adapters for purpose-built vector - databases, this unified adapter dispatches to backend-specific API formats - based on the `:backend` configuration. All three backends share the same - core capability — high-performance vector similarity search — but differ - in their API shapes, filtering syntax, and secondary features. - - ## Modality Mapping - - | VeriSimDB Modality | Capability | Backend Support | - |--------------------|-------------------------------|------------------------------| - | `:vector` | Native ANN similarity search | All (Qdrant, Milvus, Weaviate) | - | `:temporal` | Timestamp-based filtering | All (payload/property filter) | - | `:spatial` | Geo-distance filtering | Qdrant (geo payload), Weaviate | - | `:semantic` | Metadata/payload filtering | All (payload/property filter) | - - Dedicated vector databases do not support graph traversal, full-text - document search, tensor operations, or provenance chains, so `:graph`, - `:document`, `:tensor`, and `:provenance` modalities are not supported. - - ## Configuration - - # Qdrant - %{ - host: "qdrant.internal", - port: 6333, - collection: "octads", - backend: :qdrant, - auth: {:api_key, "your-api-key"} - } - - # Milvus - %{ - host: "milvus.internal", - port: 19530, - collection: "octads", - backend: :milvus, - auth: {:bearer, "token"} - } - - # Weaviate - %{ - host: "weaviate.internal", - port: 8080, - collection: "Octad", - backend: :weaviate, - auth: {:api_key, "your-api-key"} - } - - ## Backend Dispatch - - Each backend uses a different API format: - - **Qdrant**: REST API — `POST /collections/{name}/points/search` - - **Milvus**: REST API (v2) — `POST /v2/vectordb/entities/search` - - **Weaviate**: GraphQL API — `POST /v1/graphql` - """ - - @behaviour VeriSim.Federation.Adapter - - require Logger - - @default_timeout 10_000 - - # --------------------------------------------------------------------------- - # Callbacks - # --------------------------------------------------------------------------- - - @impl true - def connect(peer_info) do - case health_check(peer_info) do - {:ok, _latency} -> :ok - {:error, reason} -> {:error, reason} - end - end - - @impl true - def query(peer_info, query_params, opts \\ []) do - timeout = Keyword.get(opts, :timeout, @default_timeout) - modalities = Map.get(query_params, :modalities, []) - limit = Map.get(query_params, :limit, 100) - - start = System.monotonic_time(:millisecond) - - backend = get_backend(peer_info) - result = dispatch_query(backend, peer_info, modalities, query_params, limit, timeout) - - elapsed = System.monotonic_time(:millisecond) - start - - case result do - {:ok, raw_results} -> - normalised = - raw_results - |> translate_results(peer_info) - |> Enum.map(fn r -> Map.put(r, :response_time_ms, elapsed) end) - - {:ok, normalised} - - {:error, reason} -> - {:error, reason} - end - rescue - e -> - Logger.warning( - "VectorDB adapter: exception querying #{peer_info.store_id}: #{inspect(e)}" - ) - - {:error, {:exception, e}} - end - - @impl true - def health_check(peer_info) do - backend = get_backend(peer_info) - start = System.monotonic_time(:millisecond) - headers = auth_headers(peer_info.adapter_config) - - url = health_url(backend, peer_info) - - case Req.get(url, headers: headers, receive_timeout: 5_000) do - {:ok, %Req.Response{status: status}} when status in 200..299 -> - elapsed = System.monotonic_time(:millisecond) - start - {:ok, elapsed} - - {:ok, %Req.Response{status: status}} -> - {:error, {:unhealthy, status}} - - {:error, reason} -> - {:error, reason} - end - rescue - e -> {:error, {:exception, e}} - end - - @impl true - def supported_modalities(adapter_config) do - backend = Map.get(adapter_config, :backend, :qdrant) - - base = [:vector, :temporal, :semantic] - - # Qdrant and Weaviate support geo filtering - base - |> maybe_add(:spatial, backend in [:qdrant, :weaviate]) - end - - @impl true - def translate_results(raw_results, peer_info) do - backend = get_backend(peer_info) - - raw_results - |> List.wrap() - |> Enum.map(fn row -> - {id, score, data} = extract_result(backend, row) - - %{ - source_store: peer_info.store_id, - octad_id: id, - score: score, - drifted: false, - data: data, - response_time_ms: 0 - } - end) - end - - # --------------------------------------------------------------------------- - # Private — Backend Dispatch - # --------------------------------------------------------------------------- - - defp dispatch_query(:qdrant, peer_info, modalities, query_params, limit, timeout) do - query_qdrant(peer_info, modalities, query_params, limit, timeout) - end - - defp dispatch_query(:milvus, peer_info, modalities, query_params, limit, timeout) do - query_milvus(peer_info, modalities, query_params, limit, timeout) - end - - defp dispatch_query(:weaviate, peer_info, modalities, query_params, limit, timeout) do - query_weaviate(peer_info, modalities, query_params, limit, timeout) - end - - defp dispatch_query(backend, _peer_info, _modalities, _query_params, _limit, _timeout) do - {:error, {:unsupported_backend, backend}} - end - - # --------------------------------------------------------------------------- - # Private — Qdrant - # --------------------------------------------------------------------------- - - defp query_qdrant(peer_info, modalities, query_params, limit, timeout) do - config = peer_info.adapter_config - collection = Map.get(config, :collection, "octads") - headers = auth_headers(config) - - cond do - :vector in modalities && Map.has_key?(query_params, :vector_query) -> - # Qdrant: POST /collections/{name}/points/search - url = "#{peer_info.endpoint}/collections/#{collection}/points/search" - embedding = query_params.vector_query - - body = %{ - "vector" => embedding, - "limit" => limit, - "with_payload" => true - } - - # Add filters for temporal/spatial/semantic modalities - body = maybe_add_qdrant_filter(body, modalities, query_params) - - case Req.post(url, json: body, headers: headers, receive_timeout: timeout) do - {:ok, %Req.Response{status: 200, body: resp}} -> - results = resp["result"] || [] - {:ok, results} - - {:ok, %Req.Response{status: status, body: resp}} -> - {:error, {:qdrant_error, status, resp["status"] || "unknown"}} - - {:error, reason} -> - {:error, reason} - end - - :semantic in modalities && Map.has_key?(query_params, :filters) -> - # Qdrant: scroll with filter (no vector) - url = "#{peer_info.endpoint}/collections/#{collection}/points/scroll" - - body = %{ - "filter" => build_qdrant_filter(query_params), - "limit" => limit, - "with_payload" => true - } - - case Req.post(url, json: body, headers: headers, receive_timeout: timeout) do - {:ok, %Req.Response{status: 200, body: resp}} -> - results = resp["result"]["points"] || [] - {:ok, results} - - {:ok, %Req.Response{status: status, body: resp}} -> - {:error, {:qdrant_error, status, resp}} - - {:error, reason} -> - {:error, reason} - end - - true -> - # Default: scroll all points - url = "#{peer_info.endpoint}/collections/#{collection}/points/scroll" - - body = %{"limit" => limit, "with_payload" => true} - - case Req.post(url, json: body, headers: headers, receive_timeout: timeout) do - {:ok, %Req.Response{status: 200, body: resp}} -> - {:ok, resp["result"]["points"] || []} - - {:ok, %Req.Response{status: status, body: resp}} -> - {:error, {:qdrant_error, status, resp}} - - {:error, reason} -> - {:error, reason} - end - end - end - - defp maybe_add_qdrant_filter(body, modalities, query_params) do - filter = build_qdrant_filter_conditions(modalities, query_params) - - if filter == %{} do - body - else - Map.put(body, "filter", filter) - end - end - - defp build_qdrant_filter(query_params) do - filters = Map.get(query_params, :filters, %{}) - - must_conditions = - Enum.map(filters, fn {field, value} -> - %{"key" => to_string(field), "match" => %{"value" => value}} - end) - - %{"must" => must_conditions} - end - - defp build_qdrant_filter_conditions(modalities, query_params) do - conditions = [] - - conditions = - if :temporal in modalities && Map.has_key?(query_params, :temporal_range) do - range = query_params.temporal_range - - conditions ++ - [ - %{ - "key" => "created_at", - "range" => %{ - "gte" => range[:start] || range["start"], - "lte" => range[:end] || range["end"] - } - } - ] - else - conditions - end - - conditions = - if :spatial in modalities && Map.has_key?(query_params, :spatial_bounds) do - bounds = query_params.spatial_bounds - - conditions ++ - [ - %{ - "key" => "location", - "geo_bounding_box" => %{ - "top_left" => %{ - "lat" => bounds[:max_lat] || bounds["max_lat"] || 0.0, - "lon" => bounds[:min_lon] || bounds["min_lon"] || 0.0 - }, - "bottom_right" => %{ - "lat" => bounds[:min_lat] || bounds["min_lat"] || 0.0, - "lon" => bounds[:max_lon] || bounds["max_lon"] || 0.0 - } - } - } - ] - else - conditions - end - - conditions = - if :semantic in modalities && Map.has_key?(query_params, :filters) do - filter_conditions = - Enum.map(query_params.filters, fn {field, value} -> - %{"key" => to_string(field), "match" => %{"value" => value}} - end) - - conditions ++ filter_conditions - else - conditions - end - - if conditions == [], do: %{}, else: %{"must" => conditions} - end - - # --------------------------------------------------------------------------- - # Private — Milvus - # --------------------------------------------------------------------------- - - defp query_milvus(peer_info, modalities, query_params, limit, timeout) do - config = peer_info.adapter_config - collection = Map.get(config, :collection, "octads") - headers = auth_headers(config) ++ [{"Content-Type", "application/json"}] - - cond do - :vector in modalities && Map.has_key?(query_params, :vector_query) -> - # Milvus REST v2: POST /v2/vectordb/entities/search - url = "#{peer_info.endpoint}/v2/vectordb/entities/search" - embedding = query_params.vector_query - - body = %{ - "collectionName" => collection, - "data" => [embedding], - "limit" => limit, - "outputFields" => ["*"] - } - - # Add filter expression for temporal/semantic - body = maybe_add_milvus_filter(body, modalities, query_params) - - case Req.post(url, json: body, headers: headers, receive_timeout: timeout) do - {:ok, %Req.Response{status: 200, body: resp}} -> - results = resp["data"] || [] - {:ok, results} - - {:ok, %Req.Response{status: status, body: resp}} -> - {:error, {:milvus_error, status, resp["message"] || "unknown"}} - - {:error, reason} -> - {:error, reason} - end - - true -> - # Milvus: query without vector (filter only) - url = "#{peer_info.endpoint}/v2/vectordb/entities/query" - - filter_expr = build_milvus_filter(modalities, query_params) - - body = %{ - "collectionName" => collection, - "filter" => filter_expr, - "limit" => limit, - "outputFields" => ["*"] - } - - case Req.post(url, json: body, headers: headers, receive_timeout: timeout) do - {:ok, %Req.Response{status: 200, body: resp}} -> - {:ok, resp["data"] || []} - - {:ok, %Req.Response{status: status, body: resp}} -> - {:error, {:milvus_error, status, resp["message"] || "unknown"}} - - {:error, reason} -> - {:error, reason} - end - end - end - - defp maybe_add_milvus_filter(body, modalities, query_params) do - filter_expr = build_milvus_filter(modalities, query_params) - - if filter_expr == "" do - body - else - Map.put(body, "filter", filter_expr) - end - end - - defp build_milvus_filter(modalities, query_params) do - conditions = [] - - conditions = - if :temporal in modalities && Map.has_key?(query_params, :temporal_range) do - range = query_params.temporal_range - start_time = range[:start] || range["start"] || "" - end_time = range[:end] || range["end"] || "" - conditions ++ ["created_at >= '#{start_time}' AND created_at <= '#{end_time}'"] - else - conditions - end - - conditions = - if :semantic in modalities && Map.has_key?(query_params, :filters) do - filter_exprs = - Enum.map(query_params.filters, fn {field, value} -> - "#{field} == '#{value}'" - end) - - conditions ++ filter_exprs - else - conditions - end - - Enum.join(conditions, " AND ") - end - - # --------------------------------------------------------------------------- - # Private — Weaviate - # --------------------------------------------------------------------------- - - defp query_weaviate(peer_info, modalities, query_params, limit, timeout) do - config = peer_info.adapter_config - collection = Map.get(config, :collection, "Octad") - headers = auth_headers(config) ++ [{"Content-Type", "application/json"}] - - cond do - :vector in modalities && Map.has_key?(query_params, :vector_query) -> - # Weaviate GraphQL: nearVector search - url = "#{peer_info.endpoint}/v1/graphql" - embedding = query_params.vector_query - - where_filter = build_weaviate_where(modalities, query_params) - - graphql = build_weaviate_near_vector_query(collection, embedding, limit, where_filter) - - body = %{"query" => graphql} - - case Req.post(url, json: body, headers: headers, receive_timeout: timeout) do - {:ok, %Req.Response{status: 200, body: resp}} -> - results = - get_in(resp, ["data", "Get", collection]) || [] - - {:ok, results} - - {:ok, %Req.Response{status: status, body: resp}} -> - errors = resp["errors"] || [] - error_msg = Enum.map_join(errors, "; ", & &1["message"]) - {:error, {:weaviate_error, status, error_msg}} - - {:error, reason} -> - {:error, reason} - end - - true -> - # Weaviate: filtered query without vector - url = "#{peer_info.endpoint}/v1/graphql" - - where_filter = build_weaviate_where(modalities, query_params) - graphql = build_weaviate_get_query(collection, limit, where_filter) - - body = %{"query" => graphql} - - case Req.post(url, json: body, headers: headers, receive_timeout: timeout) do - {:ok, %Req.Response{status: 200, body: resp}} -> - {:ok, get_in(resp, ["data", "Get", collection]) || []} - - {:ok, %Req.Response{status: status, body: resp}} -> - errors = resp["errors"] || [] - error_msg = Enum.map_join(errors, "; ", & &1["message"]) - {:error, {:weaviate_error, status, error_msg}} - - {:error, reason} -> - {:error, reason} - end - end - end - - defp build_weaviate_near_vector_query(collection, embedding, limit, where_filter) do - embedding_str = Jason.encode!(embedding) - where_clause = if where_filter != "", do: ", where: #{where_filter}", else: "" - - """ - { Get { #{collection}( - nearVector: { vector: #{embedding_str} } - limit: #{limit} - #{where_clause} - ) { - _additional { id distance } - ... on #{collection} { _additional { id } } - } - } } - """ - end - - defp build_weaviate_get_query(collection, limit, where_filter) do - where_clause = if where_filter != "", do: ", where: #{where_filter}", else: "" - - """ - { Get { #{collection}( - limit: #{limit} - #{where_clause} - ) { - _additional { id } - } - } } - """ - end - - defp build_weaviate_where(modalities, query_params) do - conditions = [] - - conditions = - if :temporal in modalities && Map.has_key?(query_params, :temporal_range) do - range = query_params.temporal_range - start_time = range[:start] || range["start"] || "" - - conditions ++ - [ - ~s|{ path: ["created_at"], operator: GreaterThanEqual, valueDate: "#{start_time}" }| - ] - else - conditions - end - - conditions = - if :semantic in modalities && Map.has_key?(query_params, :filters) do - filter_conditions = - Enum.map(query_params.filters, fn {field, value} -> - ~s|{ path: ["#{field}"], operator: Equal, valueText: "#{value}" }| - end) - - conditions ++ filter_conditions - else - conditions - end - - case conditions do - [] -> "" - [single] -> single - multiple -> "{ operator: And, operands: [#{Enum.join(multiple, ", ")}] }" - end - end - - # --------------------------------------------------------------------------- - # Private — Helpers - # --------------------------------------------------------------------------- - - defp get_backend(peer_info) do - Map.get(peer_info.adapter_config, :backend, :qdrant) - end - - defp health_url(:qdrant, peer_info) do - "#{peer_info.endpoint}/readyz" - end - - defp health_url(:milvus, peer_info) do - "#{peer_info.endpoint}/v2/vectordb/collections/list" - end - - defp health_url(:weaviate, peer_info) do - "#{peer_info.endpoint}/v1/.well-known/ready" - end - - defp health_url(_backend, peer_info) do - "#{peer_info.endpoint}/health" - end - - defp extract_result(:qdrant, row) do - id = to_string(row["id"] || "unknown") - score = row["score"] || 0.0 - data = row["payload"] || row - {id, score, data} - end - - defp extract_result(:milvus, row) do - id = to_string(row["id"] || row["pk"] || "unknown") - score = row["distance"] || row["score"] || 0.0 - {id, score, row} - end - - defp extract_result(:weaviate, row) do - additional = row["_additional"] || %{} - id = additional["id"] || "unknown" - score = 1.0 - (additional["distance"] || 1.0) - data = Map.drop(row, ["_additional"]) - {id, score, data} - end - - defp extract_result(_backend, row) do - id = row["id"] || row["_id"] || "unknown" - score = parse_score(row) - {to_string(id), score, row} - end - - defp parse_score(row) do - case row["score"] || row["distance"] do - score when is_number(score) -> score / 1 - _ -> 0.0 - end - end - - defp maybe_add(list, item, true), do: list ++ [item] - defp maybe_add(list, _item, false), do: list - - defp auth_headers(config) do - case Map.get(config, :auth, :none) do - {:basic, user, pass} -> - encoded = Base.encode64("#{user}:#{pass}") - [{"Authorization", "Basic #{encoded}"}] - - {:bearer, token} -> - [{"Authorization", "Bearer #{token}"}] - - {:api_key, key} -> - [{"Authorization", "Bearer #{key}"}] - - _ -> - [] - end - end -end diff --git a/verisimdb/elixir-orchestration/lib/verisim/federation/adapters/verisimdb.ex b/verisimdb/elixir-orchestration/lib/verisim/federation/adapters/verisimdb.ex deleted file mode 100644 index 401bb7be..00000000 --- a/verisimdb/elixir-orchestration/lib/verisim/federation/adapters/verisimdb.ex +++ /dev/null @@ -1,231 +0,0 @@ -# SPDX-License-Identifier: MPL-2.0 - -defmodule VeriSim.Federation.Adapters.VeriSimDB do - @moduledoc """ - Federation adapter for VeriSimDB-to-VeriSimDB communication. - - This is the default adapter used when federating across VeriSimDB instances. - It communicates via the verisim-api HTTP endpoints, routing queries based - on requested modalities: - - - `:vector` → `POST /search/vector` (embedding similarity) - - `:graph` → `GET /search/related/:id` (graph traversal) - - `:document` / `:text` → `GET /search/text` (Tantivy full-text) - - `:spatial` → `POST /spatial/search/radius` (PostGIS-backed) - - default → `GET /octads` (paginated listing) - - ## Capabilities - - A VeriSimDB peer supports all 8 octad modalities natively: - Graph, Vector, Tensor, Semantic, Document, Temporal, Provenance, Spatial. - - ## Configuration - - No special `adapter_config` is required — the peer's `endpoint` URL - is sufficient. Optionally, a PSK can be provided for authenticated - federation: - - %{ - endpoint: "http://verisim-peer:8080/api/v1", - adapter_config: %{ - psk: "shared-secret-key" - } - } - """ - - @behaviour VeriSim.Federation.Adapter - - require Logger - - @default_timeout 10_000 - - # --------------------------------------------------------------------------- - # Callbacks - # --------------------------------------------------------------------------- - - @impl true - def connect(peer_info) do - case health_check(peer_info) do - {:ok, _latency} -> :ok - {:error, reason} -> {:error, reason} - end - end - - @impl true - def query(peer_info, query_params, opts \\ []) do - timeout = Keyword.get(opts, :timeout, @default_timeout) - limit = Map.get(query_params, :limit, 100) - modalities = Map.get(query_params, :modalities, []) - - start = System.monotonic_time(:millisecond) - - result = dispatch_query(peer_info.endpoint, modalities, query_params, limit, timeout) - - elapsed = System.monotonic_time(:millisecond) - start - - case result do - {:ok, raw_results} -> - normalised = - raw_results - |> translate_results(peer_info) - |> Enum.map(fn r -> Map.put(r, :response_time_ms, elapsed) end) - - {:ok, normalised} - - {:error, reason} -> - {:error, reason} - end - rescue - e -> - Logger.warning( - "VeriSimDB adapter: exception querying #{peer_info.store_id}: #{inspect(e)}" - ) - - {:error, {:exception, e}} - end - - @impl true - def health_check(peer_info) do - url = "#{peer_info.endpoint}/health" - start = System.monotonic_time(:millisecond) - - headers = auth_headers(peer_info) - - case Req.get(url, headers: headers, receive_timeout: 5_000) do - {:ok, %Req.Response{status: 200}} -> - elapsed = System.monotonic_time(:millisecond) - start - {:ok, elapsed} - - {:ok, %Req.Response{status: status}} -> - {:error, {:unhealthy, status}} - - {:error, reason} -> - {:error, reason} - end - rescue - e -> {:error, {:exception, e}} - end - - @impl true - def supported_modalities(_adapter_config) do - # VeriSimDB peers support all 8 octad modalities natively. - [:graph, :vector, :tensor, :semantic, :document, :temporal, :provenance, :spatial] - end - - @impl true - def translate_results(raw_results, peer_info) do - raw_results - |> List.wrap() - |> Enum.map(fn item -> - %{ - source_store: peer_info.store_id, - octad_id: item["id"] || item["entity_id"] || "unknown", - score: item["score"] || 0.0, - drifted: item["drifted"] || false, - data: item, - response_time_ms: 0 - } - end) - end - - # --------------------------------------------------------------------------- - # Private — Query Dispatch - # --------------------------------------------------------------------------- - - defp dispatch_query(endpoint, modalities, query_params, limit, timeout) do - headers = [] - - cond do - :vector in modalities && Map.has_key?(query_params, :vector_query) -> - url = "#{endpoint}/search/vector" - body = %{vector: query_params.vector_query, k: limit} - - case Req.post(url, json: body, headers: headers, receive_timeout: timeout) do - {:ok, %Req.Response{status: status, body: body}} when status in 200..299 -> - {:ok, extract_results(body)} - - {:ok, %Req.Response{status: status}} -> - {:error, {:http_error, status}} - - {:error, reason} -> - {:error, reason} - end - - :graph in modalities && Map.has_key?(query_params, :graph_pattern) -> - url = "#{endpoint}/search/related/#{query_params.graph_pattern}" - - case Req.get(url, params: [limit: limit], headers: headers, receive_timeout: timeout) do - {:ok, %Req.Response{status: status, body: body}} when status in 200..299 -> - {:ok, extract_results(body)} - - {:ok, %Req.Response{status: status}} -> - {:error, {:http_error, status}} - - {:error, reason} -> - {:error, reason} - end - - :document in modalities && Map.has_key?(query_params, :text_query) -> - url = "#{endpoint}/search/text" - - case Req.get(url, - params: [q: query_params.text_query, limit: limit], - headers: headers, - receive_timeout: timeout - ) do - {:ok, %Req.Response{status: status, body: body}} when status in 200..299 -> - {:ok, extract_results(body)} - - {:ok, %Req.Response{status: status}} -> - {:error, {:http_error, status}} - - {:error, reason} -> - {:error, reason} - end - - :spatial in modalities && Map.has_key?(query_params, :spatial_bounds) -> - url = "#{endpoint}/spatial/search/bounds" - body = Map.put(query_params.spatial_bounds, :limit, limit) - - case Req.post(url, json: body, headers: headers, receive_timeout: timeout) do - {:ok, %Req.Response{status: status, body: body}} when status in 200..299 -> - {:ok, extract_results(body)} - - {:ok, %Req.Response{status: status}} -> - {:error, {:http_error, status}} - - {:error, reason} -> - {:error, reason} - end - - true -> - # Default: paginated listing - url = "#{endpoint}/octads" - - case Req.get(url, params: [limit: limit], headers: headers, receive_timeout: timeout) do - {:ok, %Req.Response{status: status, body: body}} when status in 200..299 -> - {:ok, extract_results(body)} - - {:ok, %Req.Response{status: status}} -> - {:error, {:http_error, status}} - - {:error, reason} -> - {:error, reason} - end - end - end - - defp extract_results(body) when is_list(body), do: body - - defp extract_results(%{"results" => results}) when is_list(results), do: results - defp extract_results(%{"octads" => octads}) when is_list(octads), do: octads - defp extract_results(body) when is_map(body), do: [body] - defp extract_results(_), do: [] - - defp auth_headers(peer_info) do - case get_in(peer_info, [:adapter_config, :psk]) do - nil -> [] - psk -> [{"X-Federation-PSK", psk}] - end - end -end diff --git a/verisimdb/elixir-orchestration/lib/verisim/federation/resolver.ex b/verisimdb/elixir-orchestration/lib/verisim/federation/resolver.ex deleted file mode 100644 index ed2f7c0c..00000000 --- a/verisimdb/elixir-orchestration/lib/verisim/federation/resolver.ex +++ /dev/null @@ -1,466 +0,0 @@ -# SPDX-License-Identifier: MPL-2.0 - -defmodule VeriSim.Federation.Resolver do - @moduledoc """ - Federation Resolver — coordinates cross-instance queries across - heterogeneous database backends. - - Resolves federated query patterns to peer stores, dispatches parallel - queries via adapter modules, and aggregates results according to drift - policies. Supports VeriSimDB, ArangoDB, PostgreSQL, and Elasticsearch - peers in the same federation. - - ## Heterogeneous Federation - - Each peer declares an `adapter_type` at registration. The resolver - routes queries through the correct adapter module, which translates - VeriSimDB's modality-based queries into the backend's native language - (AQL, SQL, Elasticsearch DSL, or VeriSimDB HTTP API). - - ┌───────────────────────────────────────────────────┐ - │ VeriSim.Federation.Resolver │ - │ ┌─────────┬──────────┬───────────┬────────────┐ │ - │ │VeriSimDB│ ArangoDB │PostgreSQL │Elasticsearch│ │ - │ │ Adapter │ Adapter │ Adapter │ Adapter │ │ - │ └────┬────┴────┬─────┴─────┬─────┴─────┬──────┘ │ - └───────┼─────────┼───────────┼───────────┼─────────┘ - │ │ │ │ - verisim-api ArangoDB PostgreSQL Elasticsearch - │ HTTP API (wire/ REST API - │ PostgREST) - - ## Drift Policies - - - `:strict` — Only include stores with trust level above threshold. - - `:repair` — Include all, trigger normalization on drifted stores. - - `:tolerate` — Include all, annotate drifted results. - - `:latest` — Use only the most recent data from each store. - - ## Peer Registration - - ### VeriSimDB peer (backward-compatible 3-arity) - - register_peer("verisim-prod", "http://verisim:8080/api/v1", - [:graph, :vector, :document]) - - ### Heterogeneous peer (map-based config) - - register_peer("arango-prod", %{ - endpoint: "http://arango:8529", - adapter_type: :arangodb, - adapter_config: %{database: "_system", collection: "octads"}, - modalities: [:graph, :document, :semantic] - }) - """ - - use GenServer - require Logger - - alias VeriSim.Federation.Adapter - alias VeriSim.RustClient - - @default_timeout 10_000 - @health_check_interval 60_000 - @strict_trust_threshold 0.7 - - # --------------------------------------------------------------------------- - # Client API - # --------------------------------------------------------------------------- - - def start_link(opts \\ []) do - GenServer.start_link(__MODULE__, opts, name: __MODULE__) - end - - @doc """ - Register a peer store in the federation. - - ## 3-arity (backward-compatible — VeriSimDB peers) - - register_peer("store-id", "http://host:port/api/v1", [:graph, :vector]) - - ## 2-arity (heterogeneous — any adapter) - - register_peer("store-id", %{ - endpoint: "http://host:port", - adapter_type: :arangodb, - adapter_config: %{database: "_system"}, - modalities: [:graph, :document] - }) - """ - def register_peer(store_id, endpoint, modalities) - when is_binary(endpoint) and is_list(modalities) do - # Backward-compatible: wrap as a VeriSimDB adapter config - register_peer(store_id, %{ - endpoint: endpoint, - adapter_type: :verisimdb, - adapter_config: %{}, - modalities: modalities - }) - end - - def register_peer(store_id, %{} = config) do - GenServer.call(__MODULE__, {:register, store_id, config}) - end - - @doc "Remove a peer from the federation." - def deregister_peer(store_id) do - GenServer.call(__MODULE__, {:deregister, store_id}) - end - - @doc "List all known peers." - def list_peers do - GenServer.call(__MODULE__, :list_peers) - end - - @doc """ - Execute a federated query across matching stores. - - Dispatches to each peer's adapter module for query translation and - execution. Results are normalised into a common format regardless - of the backend database. - - ## Options - - `:drift_policy` — :strict | :repair | :tolerate | :latest (default: :tolerate) - - `:limit` — max results per store - - `:timeout` — query timeout in ms - - `:text_query` — full-text search query string - - `:vector_query` — embedding vector for similarity search - - `:graph_pattern` — graph traversal start vertex - - `:spatial_bounds` — bounding box for spatial queries - - `:temporal_range` — time range for temporal queries - """ - def query(pattern, modalities, opts \\ []) do - internal_timeout = Keyword.get(opts, :timeout, @default_timeout) - - # GenServer.call timeout must exceed the internal Task.yield_many timeout - # to allow the handler to finish processing (shutdown stale tasks, build response). - GenServer.call( - __MODULE__, - {:query, pattern, modalities, opts}, - internal_timeout + 5_000 - ) - end - - # --------------------------------------------------------------------------- - # GenServer Callbacks - # --------------------------------------------------------------------------- - - @impl true - def init(_opts) do - state = %{ - peers: %{}, - self_store_id: "local", - health_check_ref: schedule_health_check() - } - - {:ok, state} - end - - @impl true - def handle_call({:register, store_id, config}, _from, state) do - adapter_type = Map.get(config, :adapter_type, :verisimdb) - adapter_config = Map.get(config, :adapter_config, %{}) - endpoint = Map.fetch!(config, :endpoint) - declared_modalities = Map.get(config, :modalities, []) - - # Validate adapter type - case Adapter.module_for(adapter_type) do - {:ok, adapter_module} -> - # Validate modalities against adapter capabilities. - # Modalities may be strings or atoms — normalise for comparison. - supported = adapter_module.supported_modalities(adapter_config) - supported_strings = Enum.map(supported, &to_string/1) - - effective_modalities = - if declared_modalities == [] do - # No modalities declared — use all supported (as atoms) - Enum.map(supported, &to_string/1) - else - # Filter declared modalities to those the adapter supports - Enum.filter(declared_modalities, fn m -> - to_string(m) in supported_strings - end) - end - - unsupported = declared_modalities -- effective_modalities - - if unsupported != [] do - Logger.warning( - "Federation: peer #{store_id} (#{adapter_type}) declared unsupported " <> - "modalities #{inspect(unsupported)}, using #{inspect(effective_modalities)}" - ) - end - - peer = %{ - store_id: store_id, - endpoint: endpoint, - adapter_type: adapter_type, - adapter_module: adapter_module, - adapter_config: adapter_config, - modalities: effective_modalities, - trust_level: 1.0, - last_seen: DateTime.utc_now(), - response_time_ms: nil - } - - Logger.info( - "Federation: registered #{adapter_type} peer #{store_id} at #{endpoint} " <> - "with modalities #{inspect(effective_modalities)}" - ) - - new_peers = Map.put(state.peers, store_id, peer) - - # Also register VeriSimDB-type peers with the Rust API for Rust-side federation - if adapter_type == :verisimdb do - RustClient.post("/federation/register", %{ - store_id: store_id, - endpoint: endpoint, - modalities: effective_modalities - }) - end - - {:reply, :ok, %{state | peers: new_peers}} - - {:error, :unknown_adapter} -> - Logger.error("Federation: unknown adapter type #{inspect(adapter_type)} for #{store_id}") - {:reply, {:error, {:unknown_adapter, adapter_type}}, state} - end - end - - @impl true - def handle_call({:deregister, store_id}, _from, state) do - case Map.get(state.peers, store_id) do - nil -> - {:reply, {:error, :not_found}, state} - - peer -> - new_peers = Map.delete(state.peers, store_id) - Logger.info("Federation: deregistered peer #{store_id}") - - # Deregister VeriSimDB peers from Rust API - if peer.adapter_type == :verisimdb do - RustClient.post("/federation/deregister/#{store_id}", %{}) - end - - {:reply, :ok, %{state | peers: new_peers}} - end - end - - @impl true - def handle_call(:list_peers, _from, state) do - peers = - state.peers - |> Map.values() - |> Enum.map(fn peer -> - # Return a clean view without the adapter module reference - Map.drop(peer, [:adapter_module]) - end) - - {:reply, peers, state} - end - - @impl true - def handle_call({:query, pattern, modalities, opts}, _from, state) do - drift_policy = Keyword.get(opts, :drift_policy, :tolerate) - limit = Keyword.get(opts, :limit, 100) - timeout = Keyword.get(opts, :timeout, @default_timeout) - - # Build query_params from opts for adapter dispatch - query_params = %{ - modalities: modalities, - limit: limit - } - - # Add optional modality-specific query parameters - query_params = maybe_add_param(query_params, :text_query, Keyword.get(opts, :text_query)) - query_params = maybe_add_param(query_params, :vector_query, Keyword.get(opts, :vector_query)) - query_params = maybe_add_param(query_params, :graph_pattern, Keyword.get(opts, :graph_pattern)) - query_params = maybe_add_param(query_params, :spatial_bounds, Keyword.get(opts, :spatial_bounds)) - query_params = maybe_add_param(query_params, :temporal_range, Keyword.get(opts, :temporal_range)) - query_params = maybe_add_param(query_params, :filters, Keyword.get(opts, :filters)) - - # Resolve pattern to matching stores - matching = resolve_pattern(state.peers, pattern, modalities) - - # Apply drift policy filter - {included, excluded} = apply_drift_policy(matching, drift_policy) - - # Fan out queries in parallel — each peer dispatches through its adapter - tasks = - Enum.map(included, fn peer -> - Task.async(fn -> query_peer_via_adapter(peer, query_params, timeout: timeout) end) - end) - - # Collect results with timeout - results = - tasks - |> Task.yield_many(timeout) - |> Enum.flat_map(fn - {_task, {:ok, {:ok, results}}} -> - results - - {_task, {:ok, {:error, reason}}} -> - Logger.warning("Federation: peer query failed: #{inspect(reason)}") - [] - - {task, nil} -> - Task.shutdown(task, :brutal_kill) - Logger.warning("Federation: peer query timed out") - [] - end) - - # Handle repair policy: trigger normalization on drifted VeriSimDB stores - if drift_policy == :repair do - trigger_repairs(excluded) - end - - response = %{ - results: results, - stores_queried: Enum.map(included, & &1.store_id), - stores_excluded: Enum.map(excluded, & &1.store_id), - drift_policy: drift_policy - } - - {:reply, {:ok, response}, state} - end - - @impl true - def handle_info(:health_check, state) do - # Health-check all peers via their adapter modules - new_peers = - state.peers - |> Enum.map(fn {id, peer} -> - peer_info = build_peer_info(peer) - - case peer.adapter_module.health_check(peer_info) do - {:ok, response_time} -> - {id, %{peer | - last_seen: DateTime.utc_now(), - response_time_ms: response_time, - trust_level: min(peer.trust_level + 0.05, 1.0) - }} - - {:error, reason} -> - Logger.debug( - "Federation: health check failed for #{peer.adapter_type} peer #{id}: " <> - "#{inspect(reason)}" - ) - - {id, %{peer | - trust_level: max(peer.trust_level - 0.1, 0.0) - }} - end - end) - |> Map.new() - - {:noreply, %{state | peers: new_peers, health_check_ref: schedule_health_check()}} - end - - @impl true - def handle_info(_msg, state), do: {:noreply, state} - - # --------------------------------------------------------------------------- - # Private — Adapter Dispatch - # --------------------------------------------------------------------------- - - defp query_peer_via_adapter(peer, query_params, opts) do - peer_info = build_peer_info(peer) - peer.adapter_module.query(peer_info, query_params, opts) - end - - defp build_peer_info(peer) do - %{ - store_id: peer.store_id, - endpoint: peer.endpoint, - adapter_config: peer.adapter_config - } - end - - # --------------------------------------------------------------------------- - # Private — Pattern Matching & Drift Policies - # --------------------------------------------------------------------------- - - defp resolve_pattern(peers, pattern, required_modalities) do - peers - |> Map.values() - |> Enum.filter(fn peer -> - pattern_matches?(pattern, peer.store_id) && - Enum.all?(required_modalities, fn m -> - # Normalise to string for comparison (modalities may be atoms or strings) - m_str = to_string(m) - Enum.any?(peer.modalities, fn pm -> to_string(pm) == m_str end) - end) - end) - end - - defp pattern_matches?("*", _store_id), do: true - - defp pattern_matches?(pattern, store_id) do - if String.ends_with?(pattern, "/*") do - prefix = String.trim_trailing(pattern, "/*") - String.starts_with?(store_id, prefix) - else - pattern == store_id - end - end - - defp apply_drift_policy(peers, :strict) do - Enum.split_with(peers, fn peer -> - peer.trust_level >= @strict_trust_threshold - end) - end - - defp apply_drift_policy(peers, _policy) do - {peers, []} - end - - # --------------------------------------------------------------------------- - # Private — Repair Policy - # --------------------------------------------------------------------------- - - defp trigger_repairs(excluded_peers) do - Enum.each(excluded_peers, fn peer -> - Logger.info( - "Federation: triggering repair on drifted #{peer.adapter_type} store #{peer.store_id}" - ) - - # Only VeriSimDB peers support the normaliser endpoint - if peer.adapter_type == :verisimdb do - Task.start(fn -> - url = "#{peer.endpoint}/normalizer/trigger/all" - - case Req.post(url, json: %{}, receive_timeout: @default_timeout) do - {:ok, %Req.Response{status: status}} when status in 200..299 -> - Logger.info("Federation: repair triggered on #{peer.store_id}") - - {:ok, %Req.Response{status: status}} -> - Logger.warning( - "Federation: repair request to #{peer.store_id} returned #{status}" - ) - - {:error, reason} -> - Logger.warning( - "Federation: repair request to #{peer.store_id} failed: #{inspect(reason)}" - ) - end - end) - else - Logger.debug( - "Federation: skipping repair for non-VeriSimDB peer #{peer.store_id} " <> - "(#{peer.adapter_type} does not support normalisation)" - ) - end - end) - end - - # --------------------------------------------------------------------------- - # Private — Helpers - # --------------------------------------------------------------------------- - - defp maybe_add_param(params, _key, nil), do: params - defp maybe_add_param(params, key, value), do: Map.put(params, key, value) - - defp schedule_health_check do - Process.send_after(self(), :health_check, @health_check_interval) - end -end diff --git a/verisimdb/elixir-orchestration/lib/verisim/health_checker.ex b/verisimdb/elixir-orchestration/lib/verisim/health_checker.ex deleted file mode 100644 index a81342f5..00000000 --- a/verisimdb/elixir-orchestration/lib/verisim/health_checker.ex +++ /dev/null @@ -1,129 +0,0 @@ -# SPDX-License-Identifier: MPL-2.0 - -defmodule VeriSim.HealthChecker do - @moduledoc """ - Periodic health checker for VeriSimDB orchestration layer. - - Probes the Rust core reachability, GenServer liveness, and ETS cache health. - Reports status via telemetry events. - """ - - use GenServer - require Logger - - alias VeriSim.RustClient - - @check_interval 30_000 - - # Client API - - def start_link(opts \\ []) do - GenServer.start_link(__MODULE__, opts, name: __MODULE__) - end - - @doc "Get the latest health check result." - def status do - GenServer.call(__MODULE__, :status) - end - - # Server Callbacks - - @impl true - def init(_opts) do - state = %{ - rust_core: :unknown, - entity_registry: :unknown, - ets_cache: :unknown, - last_checked: nil, - check_count: 0 - } - - schedule_check() - {:ok, state} - end - - @impl true - def handle_call(:status, _from, state) do - {:reply, {:ok, state}, state} - end - - @impl true - def handle_info(:check, state) do - rust_status = check_rust_core() - registry_status = check_entity_registry() - ets_status = check_ets_cache() - - new_state = %{state | - rust_core: rust_status, - entity_registry: registry_status, - ets_cache: ets_status, - last_checked: DateTime.utc_now(), - check_count: state.check_count + 1 - } - - # Emit telemetry events - :telemetry.execute( - [:verisim, :health_check], - %{duration: 0}, - %{ - rust_core: rust_status, - entity_registry: registry_status, - ets_cache: ets_status - } - ) - - # Log warnings for unhealthy components - if rust_status != :healthy do - Logger.warning("Health check: Rust core is #{rust_status}") - end - - if registry_status != :healthy do - Logger.warning("Health check: Entity registry is #{registry_status}") - end - - if ets_status != :healthy do - Logger.warning("Health check: ETS cache is #{ets_status}") - end - - schedule_check() - {:noreply, new_state} - end - - @impl true - def handle_info(_msg, state), do: {:noreply, state} - - # Private - - defp check_rust_core do - case RustClient.get("/health", []) do - {:ok, %{"status" => status}} when status in ["healthy", "degraded"] -> :healthy - {:ok, _} -> :degraded - {:error, _} -> :unreachable - end - rescue - _ -> :unreachable - end - - defp check_entity_registry do - case Registry.count(VeriSim.EntityRegistry) do - count when is_integer(count) -> :healthy - _ -> :degraded - end - rescue - _ -> :unavailable - end - - defp check_ets_cache do - case :ets.info(:verisim_rust_client_cache) do - :undefined -> :unavailable - info when is_list(info) -> :healthy - _ -> :degraded - end - rescue - _ -> :unavailable - end - - defp schedule_check do - Process.send_after(self(), :check, @check_interval) - end -end diff --git a/verisimdb/elixir-orchestration/lib/verisim/hypatia/dispatch_bridge.ex b/verisimdb/elixir-orchestration/lib/verisim/hypatia/dispatch_bridge.ex deleted file mode 100644 index 1c466848..00000000 --- a/verisimdb/elixir-orchestration/lib/verisim/hypatia/dispatch_bridge.ex +++ /dev/null @@ -1,333 +0,0 @@ -# SPDX-License-Identifier: MPL-2.0 - -defmodule VeriSim.Hypatia.DispatchBridge do - @moduledoc """ - Bridge between VeriSimDB octad data and Hypatia's dispatch pipeline. - - Reads dispatch manifests (JSONL files) from verisimdb-data/dispatch/, - tracks execution status, and provides feedback to VeriSimDB for drift - tracking between scan cycles. - - ## Dispatch Lifecycle - - ┌─────────────────────────────────────────────────────────┐ - │ Hypatia Pipeline │ - │ PatternAnalyzer → TriangleRouter → FleetDispatcher │ - │ │ │ - │ dispatch/*.jsonl │ - └───────────────────────────────────────────┬─────────────┘ - │ - ┌───────────────────────────────────────────┴─────────────┐ - │ DispatchBridge (this module) │ - │ ├── read_pending/1 — read pending.jsonl │ - │ ├── read_dispatch_log/1 — read dispatch-*.jsonl │ - │ ├── summarize/1 — aggregate dispatch stats │ - │ ├── track_outcomes/1 — read outcomes/*.jsonl │ - │ └── feedback_to_drift/1 — feed outcomes back to VDB │ - └─────────────────────────────────────────────────────────┘ - - ## Usage - - # Read pending dispatch actions - {:ok, actions} = DispatchBridge.read_pending("/path/to/verisimdb-data") - - # Get dispatch summary - summary = DispatchBridge.summarize("/path/to/verisimdb-data") - - # Feed outcomes back for drift tracking - DispatchBridge.feedback_to_drift("/path/to/verisimdb-data") - """ - - require Logger - - alias VeriSim.Hypatia.ScanIngester - - @dispatch_dir "dispatch" - @outcomes_dir "outcomes" - @pending_file "pending.jsonl" - - # --------------------------------------------------------------------------- - # Read Dispatch Data - # --------------------------------------------------------------------------- - - @doc """ - Read all pending dispatch actions from `dispatch/pending.jsonl`. - - Returns `{:ok, actions}` where each action is a decoded JSON map with: - - `repo` — target repository - - `pattern` — canonical pattern ID - - `strategy` — dispatch strategy (auto_execute, review, report_only) - - `confidence` — recipe confidence score - - `mutation` — GraphQL mutation payload - """ - def read_pending(data_path) when is_binary(data_path) do - path = Path.join([data_path, @dispatch_dir, @pending_file]) - read_jsonl(path) - end - - @doc """ - Read dispatch log for a specific date (e.g., "2026-02-12"). - - Returns `{:ok, records}` or `{:error, reason}`. - """ - def read_dispatch_log(data_path, date) when is_binary(data_path) and is_binary(date) do - path = Path.join([data_path, @dispatch_dir, "dispatch-#{date}.jsonl"]) - read_jsonl(path) - end - - @doc """ - Read all dispatch logs from the dispatch directory. - - Returns `{:ok, records}` with all dispatch records across all log files. - """ - def read_all_dispatch_logs(data_path) when is_binary(data_path) do - dir = Path.join(data_path, @dispatch_dir) - - case File.ls(dir) do - {:ok, files} -> - records = - files - |> Enum.filter(&String.starts_with?(&1, "dispatch-")) - |> Enum.filter(&String.ends_with?(&1, ".jsonl")) - |> Enum.flat_map(fn file -> - path = Path.join(dir, file) - - case read_jsonl(path) do - {:ok, lines} -> lines - {:error, _} -> [] - end - end) - - {:ok, records} - - {:error, reason} -> - {:error, {:dir_read_error, reason}} - end - end - - # --------------------------------------------------------------------------- - # Read Outcomes - # --------------------------------------------------------------------------- - - @doc """ - Read all fix outcomes from `outcomes/*.jsonl`. - - Each outcome records whether a dispatched fix succeeded or failed, - and feeds back into the learning loop. - """ - def read_outcomes(data_path) when is_binary(data_path) do - dir = Path.join(data_path, @outcomes_dir) - - case File.ls(dir) do - {:ok, files} -> - outcomes = - files - |> Enum.filter(&String.ends_with?(&1, ".jsonl")) - |> Enum.flat_map(fn file -> - path = Path.join(dir, file) - - case read_jsonl(path) do - {:ok, lines} -> lines - {:error, _} -> [] - end - end) - - {:ok, outcomes} - - {:error, reason} -> - {:error, {:dir_read_error, reason}} - end - end - - # --------------------------------------------------------------------------- - # Summary - # --------------------------------------------------------------------------- - - @doc """ - Aggregate dispatch statistics from all logs and pending actions. - - Returns a summary map with counts per strategy, per-repo breakdown, - and outcome statistics. - """ - def summarize(data_path) when is_binary(data_path) do - pending = case read_pending(data_path) do - {:ok, p} -> p - _ -> [] - end - - dispatched = case read_all_dispatch_logs(data_path) do - {:ok, d} -> d - _ -> [] - end - - outcomes = case read_outcomes(data_path) do - {:ok, o} -> o - _ -> [] - end - - %{ - pending_count: length(pending), - dispatched_count: length(dispatched), - outcome_count: length(outcomes), - - by_strategy: group_by_field(dispatched, "strategy"), - by_repo: group_by_field(dispatched, "repo") |> top_n(20), - - outcome_success_rate: outcome_success_rate(outcomes), - - pending_by_strategy: group_by_field(pending, "strategy"), - - repos_with_pending: - pending - |> Enum.map(& &1["repo"]) - |> Enum.reject(&is_nil/1) - |> Enum.uniq() - |> length() - } - end - - # --------------------------------------------------------------------------- - # Feedback to VeriSimDB Drift - # --------------------------------------------------------------------------- - - @doc """ - Feed dispatch outcomes back into VeriSimDB for temporal drift tracking. - - For each repo with outcomes, compares the latest scan's weak point count - against earlier scans to detect improvement or regression (drift). - - Returns a list of `{repo, drift_direction, delta}` tuples. - """ - def feedback_to_drift(data_path) when is_binary(data_path) do - outcomes = case read_outcomes(data_path) do - {:ok, o} -> o - _ -> [] - end - - # Group outcomes by repo - repo_outcomes = - outcomes - |> Enum.group_by(& &1["repo"]) - - # For each repo with outcomes, check scan trends - repo_outcomes - |> Enum.map(fn {repo, repo_ocs} -> - successful = Enum.count(repo_ocs, &(&1["status"] == "success")) - total = length(repo_ocs) - - drift = cond do - successful == total -> :improving - successful >= total / 2 -> :stable - true -> :regressing - end - - {repo, drift, %{successful: successful, total: total}} - end) - end - - @doc """ - Create a VeriSimDB octad representing a dispatch batch summary. - - Stores dispatch metadata as a trackable entity for temporal analysis. - """ - def ingest_dispatch_summary(data_path) when is_binary(data_path) do - summary = summarize(data_path) - octad_id = "dispatch:summary:#{System.system_time(:second)}" - - octad_input = %{ - octad_id: octad_id, - metadata: %{ - type: "dispatch_summary", - timestamp: DateTime.utc_now() |> DateTime.to_iso8601() - }, - document: %{ - title: "Hypatia Dispatch Summary", - body: Jason.encode!(summary, pretty: true), - content_type: "application/json" - }, - temporal: %{ - timestamp: System.system_time(:millisecond), - version: "dispatch-v1", - event_type: "dispatch_summary" - }, - provenance: %{ - source: "hypatia-dispatch", - actor: "dispatch-bridge", - operation: "summarize" - } - } - - ScanIngester.ingest_scan(%{ - "assail_report" => %{ - "program_path" => "dispatch-summary", - "language" => "n/a", - "frameworks" => [], - "weak_points" => [] - } - }) - - # Also store the full summary in local ETS - ensure_ets_table() - :ets.insert(:hypatia_dispatch_summaries, {octad_id, octad_input}) - - {:ok, octad_id} - end - - # --------------------------------------------------------------------------- - # Private: JSONL Reading - # --------------------------------------------------------------------------- - - defp read_jsonl(path) do - case File.read(path) do - {:ok, data} -> - lines = - data - |> String.split("\n", trim: true) - |> Enum.flat_map(fn line -> - case Jason.decode(line) do - {:ok, parsed} -> [parsed] - {:error, _} -> [] - end - end) - - {:ok, lines} - - {:error, reason} -> - {:error, {:file_read_error, path, reason}} - end - end - - # --------------------------------------------------------------------------- - # Private: Aggregation Helpers - # --------------------------------------------------------------------------- - - defp group_by_field(records, field) do - records - |> Enum.group_by(& &1[field]) - |> Map.new(fn {key, items} -> {key || "unknown", length(items)} end) - end - - defp top_n(map, n) do - map - |> Enum.sort_by(fn {_k, v} -> v end, :desc) - |> Enum.take(n) - |> Map.new() - end - - defp outcome_success_rate([]), do: 0.0 - - defp outcome_success_rate(outcomes) do - successful = Enum.count(outcomes, &(&1["status"] == "success")) - Float.round(successful / length(outcomes) * 100, 1) - end - - defp ensure_ets_table do - case :ets.info(:hypatia_dispatch_summaries) do - :undefined -> - :ets.new(:hypatia_dispatch_summaries, [:named_table, :set, :public]) - - _ -> - :ok - end - end -end diff --git a/verisimdb/elixir-orchestration/lib/verisim/hypatia/pattern_query.ex b/verisimdb/elixir-orchestration/lib/verisim/hypatia/pattern_query.ex deleted file mode 100644 index 3f4f496e..00000000 --- a/verisimdb/elixir-orchestration/lib/verisim/hypatia/pattern_query.ex +++ /dev/null @@ -1,240 +0,0 @@ -# SPDX-License-Identifier: MPL-2.0 - -defmodule VeriSim.Hypatia.PatternQuery do - @moduledoc """ - Cross-repo pattern analytics for the Hypatia pipeline. - - Provides query functions over ingested scan data to identify patterns - that appear across multiple repositories, track severity distributions, - and detect temporal trends (scan-to-scan drift). - - ## Query Functions - - - `pipeline_health/0` — Overall pipeline status and metrics - - `cross_repo_patterns/1` — Patterns appearing in N+ repos - - `severity_distribution/0` — Severity breakdown across all scans - - `category_distribution/0` — Category breakdown across all scans - - `temporal_trends/1` — How a repo's scan results change over time - - `repos_by_severity/1` — Repos ranked by severity count - - `weakness_hotspots/0` — Files with highest weakness density - """ - - require Logger - - alias VeriSim.Hypatia.ScanIngester - - # --------------------------------------------------------------------------- - # Pipeline Health - # --------------------------------------------------------------------------- - - @doc """ - Returns overall pipeline health metrics. - - Example: - %{ - total_scans: 289, - total_weak_points: 3260, - repos_scanned: 289, - severity_distribution: %{"High" => 120, "Medium" => 2800, "Low" => 340}, - top_categories: [{"PanicPath", 1200}, {"UnsafeCode", 500}, ...] - } - """ - def pipeline_health do - scans = ScanIngester.list_scans() - - all_weak_points = extract_all_weak_points(scans) - - %{ - total_scans: length(scans), - total_weak_points: length(all_weak_points), - repos_scanned: scans |> Enum.map(&get_in(&1, [:metadata, :repo_name])) |> Enum.uniq() |> length(), - severity_distribution: severity_distribution(all_weak_points), - top_categories: top_categories(all_weak_points, 10), - average_weak_points_per_repo: safe_avg(length(all_weak_points), length(scans)), - last_scan_timestamp: latest_timestamp(scans) - } - end - - # --------------------------------------------------------------------------- - # Cross-Repo Patterns - # --------------------------------------------------------------------------- - - @doc """ - Finds weakness patterns that appear across `min_repos` or more repositories. - - Returns a list of `{pattern_key, repo_count, repos}` tuples, sorted by - repo count descending. - - The pattern key is `"category:severity"` (e.g., `"PanicPath:Medium"`). - """ - def cross_repo_patterns(min_repos \\ 3) do - scans = ScanIngester.list_scans() - - # Build pattern → [repo_names] map - pattern_repos = - scans - |> Enum.flat_map(fn scan -> - repo = get_in(scan, [:metadata, :repo_name]) || "unknown" - weak_points = extract_weak_points(scan) - - weak_points - |> Enum.map(fn wp -> - category = wp["category"] || "unknown" - severity = wp["severity"] || "unknown" - pattern_key = "#{category}:#{severity}" - {pattern_key, repo} - end) - end) - |> Enum.group_by(fn {key, _} -> key end, fn {_, repo} -> repo end) - |> Map.new(fn {key, repos} -> {key, Enum.uniq(repos)} end) - - pattern_repos - |> Enum.filter(fn {_key, repos} -> length(repos) >= min_repos end) - |> Enum.map(fn {key, repos} -> {key, length(repos), repos} end) - |> Enum.sort_by(fn {_key, count, _repos} -> count end, :desc) - end - - # --------------------------------------------------------------------------- - # Severity Distribution - # --------------------------------------------------------------------------- - - @doc """ - Returns a map of severity level → count across all scans. - """ - def severity_distribution do - ScanIngester.list_scans() - |> extract_all_weak_points() - |> severity_distribution() - end - - defp severity_distribution(weak_points) do - weak_points - |> Enum.group_by(& &1["severity"]) - |> Map.new(fn {severity, items} -> {severity || "unknown", length(items)} end) - end - - # --------------------------------------------------------------------------- - # Category Distribution - # --------------------------------------------------------------------------- - - @doc """ - Returns a map of category → count across all scans, sorted by count descending. - """ - def category_distribution do - ScanIngester.list_scans() - |> extract_all_weak_points() - |> Enum.group_by(& &1["category"]) - |> Map.new(fn {cat, items} -> {cat || "unknown", length(items)} end) - |> Enum.sort_by(fn {_cat, count} -> count end, :desc) - end - - # --------------------------------------------------------------------------- - # Temporal Trends - # --------------------------------------------------------------------------- - - @doc """ - Returns scan results for a specific repo over time, enabling drift detection. - - Each entry contains the scan timestamp and weak point count, sorted - chronologically. - """ - def temporal_trends(repo_name) when is_binary(repo_name) do - ScanIngester.list_scans() - |> Enum.filter(fn scan -> - get_in(scan, [:metadata, :repo_name]) == repo_name - end) - |> Enum.map(fn scan -> - %{ - octad_id: scan[:octad_id], - timestamp: get_in(scan, [:metadata, :scan_timestamp]), - weak_point_count: get_in(scan, [:metadata, :weak_point_count]) || 0, - severity_counts: get_in(scan, [:metadata, :severity_counts]) || %{} - } - end) - |> Enum.sort_by(& &1.timestamp) - end - - # --------------------------------------------------------------------------- - # Repos by Severity - # --------------------------------------------------------------------------- - - @doc """ - Returns repos ranked by total weakness count for a given severity level. - - Useful for identifying which repos need the most attention. - """ - def repos_by_severity(severity \\ "High") do - ScanIngester.list_scans() - |> Enum.map(fn scan -> - repo = get_in(scan, [:metadata, :repo_name]) || "unknown" - counts = get_in(scan, [:metadata, :severity_counts]) || %{} - count = counts[severity] || 0 - {repo, count} - end) - |> Enum.filter(fn {_repo, count} -> count > 0 end) - |> Enum.sort_by(fn {_repo, count} -> count end, :desc) - end - - # --------------------------------------------------------------------------- - # Weakness Hotspots - # --------------------------------------------------------------------------- - - @doc """ - Returns files ranked by weakness density across all repos. - - Identifies the most problematic files in the entire ecosystem. - """ - def weakness_hotspots do - ScanIngester.list_scans() - |> extract_all_weak_points() - |> Enum.group_by(& &1["location"]) - |> Map.new(fn {location, items} -> - {location || "unknown", %{count: length(items), categories: Enum.map(items, & &1["category"]) |> Enum.uniq()}} - end) - |> Enum.sort_by(fn {_loc, data} -> data.count end, :desc) - |> Enum.take(50) - end - - # --------------------------------------------------------------------------- - # Private Helpers - # --------------------------------------------------------------------------- - - defp extract_weak_points(scan) do - # Weak points might be in the document body or the original report - case scan do - %{document: %{body: body}} when is_binary(body) -> - # Parse weak points from the stored JSON report - case Jason.decode(body) do - {:ok, %{"assail_report" => %{"weak_points" => wps}}} -> wps - {:ok, %{"weak_points" => wps}} -> wps - _ -> [] - end - - _ -> - [] - end - end - - defp extract_all_weak_points(scans) do - Enum.flat_map(scans, &extract_weak_points/1) - end - - defp top_categories(weak_points, limit) do - weak_points - |> Enum.group_by(& &1["category"]) - |> Enum.map(fn {cat, items} -> {cat || "unknown", length(items)} end) - |> Enum.sort_by(fn {_cat, count} -> count end, :desc) - |> Enum.take(limit) - end - - defp safe_avg(_total, 0), do: 0.0 - defp safe_avg(total, count), do: Float.round(total / count, 1) - - defp latest_timestamp(scans) do - scans - |> Enum.map(&get_in(&1, [:metadata, :scan_timestamp])) - |> Enum.reject(&is_nil/1) - |> Enum.sort(:desc) - |> List.first() - end -end diff --git a/verisimdb/elixir-orchestration/lib/verisim/hypatia/scan_ingester.ex b/verisimdb/elixir-orchestration/lib/verisim/hypatia/scan_ingester.ex deleted file mode 100644 index b0fcc4c1..00000000 --- a/verisimdb/elixir-orchestration/lib/verisim/hypatia/scan_ingester.ex +++ /dev/null @@ -1,332 +0,0 @@ -# SPDX-License-Identifier: MPL-2.0 - -defmodule VeriSim.Hypatia.ScanIngester do - @moduledoc """ - Ingests panic-attack scan results and stores them as VeriSimDB octad entities. - - Each scan report becomes an octad entity with: - - **Document**: Full JSON report as searchable text - - **Graph**: file → weakness → recommendation triples - - **Temporal**: Scan timestamp (enables drift tracking across cycles) - - **Vector**: Weakness description embeddings (for similarity search) - - **Provenance**: Scanner identity, CI run origin, transformation chain - - **Semantic**: Severity annotations, category tags - - ## Data Flow - - panic-attack assail (JSON) - │ - ScanIngester.ingest_scan/1 - │ - ┌───────┴───────┐ - │ Octad Entity │ ← Document, Graph, Temporal, Vector, Provenance, Semantic - └───────┬───────┘ - │ - VeriSimDB storage - │ - Hypatia VCL queries - - ## Usage - - # Ingest a single scan result - {:ok, octad_id} = ScanIngester.ingest_scan(scan_json) - - # Ingest all scans from verisimdb-data/scans/ directory - {:ok, results} = ScanIngester.ingest_directory("/path/to/verisimdb-data/scans") - - # Ingest from a panic-attack JSON file - {:ok, octad_id} = ScanIngester.ingest_file("/path/to/scan.json") - """ - - require Logger - - alias VeriSim.RustClient - - @type scan_report :: %{ - optional(String.t()) => any() - } - - @type weak_point :: %{ - optional(String.t()) => any() - } - - # --------------------------------------------------------------------------- - # Public API - # --------------------------------------------------------------------------- - - @doc """ - Ingest a panic-attack scan result (decoded JSON map) into VeriSimDB. - - The scan report is expected to have an `assail_report` key containing: - - `program_path` — path to the scanned repository - - `language` — detected primary language - - `frameworks` — detected frameworks - - `weak_points` — list of weakness findings - - Returns `{:ok, octad_id}` or `{:error, reason}`. - """ - def ingest_scan(scan_report) when is_map(scan_report) do - report = scan_report["assail_report"] || scan_report - - repo_name = extract_repo_name(report["program_path"]) - octad_id = "scan:#{repo_name}:#{timestamp_id()}" - - octad_input = build_octad(octad_id, repo_name, report) - - case RustClient.create_octad(octad_input) do - {:ok, _result} -> - Logger.info("Hypatia: ingested scan for #{repo_name} as #{octad_id}") - {:ok, octad_id} - - {:error, reason} -> - Logger.warning( - "Hypatia: failed to ingest scan for #{repo_name} via Rust core, " <> - "storing locally: #{inspect(reason)}" - ) - - # Fallback: store in ETS for local querying - store_local(octad_id, octad_input) - {:ok, octad_id} - end - end - - def ingest_scan(_), do: {:error, :invalid_scan_format} - - @doc """ - Ingest a scan result from a JSON file on disk. - - Returns `{:ok, octad_id}` or `{:error, reason}`. - """ - def ingest_file(path) when is_binary(path) do - case File.read(path) do - {:ok, data} -> - case Jason.decode(data) do - {:ok, scan} -> ingest_scan(scan) - {:error, reason} -> {:error, {:json_parse_error, reason}} - end - - {:error, reason} -> - {:error, {:file_read_error, reason}} - end - end - - @doc """ - Ingest all scan JSON files from a directory. - - Returns `{:ok, results}` where results is a list of - `{filename, {:ok, octad_id}}` or `{filename, {:error, reason}}`. - """ - def ingest_directory(dir_path) when is_binary(dir_path) do - case File.ls(dir_path) do - {:ok, files} -> - results = - files - |> Enum.filter(&String.ends_with?(&1, ".json")) - |> Enum.map(fn file -> - path = Path.join(dir_path, file) - {file, ingest_file(path)} - end) - - successful = Enum.count(results, fn {_, result} -> match?({:ok, _}, result) end) - Logger.info("Hypatia: ingested #{successful}/#{length(results)} scan files from #{dir_path}") - - {:ok, results} - - {:error, reason} -> - {:error, {:dir_read_error, reason}} - end - end - - @doc """ - Query all ingested scans. Returns locally stored scans when Rust core is unavailable. - """ - def list_scans do - ensure_ets_table() - - :ets.tab2list(:hypatia_scans) - |> Enum.map(fn {id, data} -> Map.put(data, :octad_id, id) end) - end - - @doc """ - Get scan data for a specific repo name. - """ - def get_scan(repo_name) when is_binary(repo_name) do - ensure_ets_table() - - :ets.tab2list(:hypatia_scans) - |> Enum.find(fn {_id, data} -> - get_in(data, [:metadata, :repo_name]) == repo_name - end) - |> case do - {id, data} -> {:ok, Map.put(data, :octad_id, id)} - nil -> {:error, :not_found} - end - end - - # --------------------------------------------------------------------------- - # Private: Octad Construction - # --------------------------------------------------------------------------- - - defp build_octad(octad_id, repo_name, report) do - weak_points = report["weak_points"] || [] - language = report["language"] || "unknown" - frameworks = report["frameworks"] || [] - - %{ - octad_id: octad_id, - metadata: %{ - repo_name: repo_name, - language: language, - frameworks: frameworks, - scan_timestamp: DateTime.utc_now() |> DateTime.to_iso8601(), - weak_point_count: length(weak_points), - severity_counts: count_severities(weak_points) - }, - - # Document modality: full report as searchable text - document: %{ - title: "Panic-attack scan: #{repo_name}", - body: build_document_body(repo_name, language, weak_points), - content_type: "application/json" - }, - - # Graph modality: file → weakness → recommendation triples - graph: %{ - triples: build_graph_triples(octad_id, repo_name, weak_points) - }, - - # Temporal modality: scan timestamp for drift tracking - temporal: %{ - timestamp: System.system_time(:millisecond), - version: "scan-v1", - event_type: "panic_attack_scan" - }, - - # Provenance modality: scanner and origin tracking - provenance: %{ - source: "panic-attack", - actor: "hypatia-scan-workflow", - operation: "assail", - input_path: report["program_path"] - }, - - # Semantic modality: type annotations and severity tags - semantic: %{ - types: ["scan_result", "security_finding", "panic_attack_report"], - tags: extract_categories(weak_points), - severity_levels: Enum.map(weak_points, & &1["severity"]) |> Enum.uniq() - }, - - # Vector modality: embedding from weakness descriptions - vector: %{ - text_for_embedding: build_embedding_text(weak_points), - dimensions: nil - } - } - end - - defp build_document_body(repo_name, language, weak_points) do - weakness_text = - weak_points - |> Enum.map(fn wp -> - "#{wp["severity"]} #{wp["category"]} in #{wp["location"]}: #{wp["description"]}" - end) - |> Enum.join("\n") - - """ - Repository: #{repo_name} - Language: #{language} - Weak Points: #{length(weak_points)} - - #{weakness_text} - """ - end - - defp build_graph_triples(octad_id, repo_name, weak_points) do - repo_node = "repo:#{repo_name}" - - # repo → has_scan → scan (lists for JSON compatibility) - base = [[repo_node, "has_scan", octad_id]] - - # For each weak point: scan → has_weakness → weakness, weakness → in_file → file - weakness_triples = - weak_points - |> Enum.with_index() - |> Enum.flat_map(fn {wp, idx} -> - weakness_id = "#{octad_id}:wp:#{idx}" - file = wp["location"] || "unknown" - category = wp["category"] || "unknown" - - [ - [octad_id, "has_weakness", weakness_id], - [weakness_id, "in_file", "file:#{file}"], - [weakness_id, "has_category", "category:#{category}"], - [weakness_id, "has_severity", "severity:#{wp["severity"] || "unknown"}"] - ] - end) - - base ++ weakness_triples - end - - defp build_embedding_text(weak_points) do - weak_points - |> Enum.map(fn wp -> - "#{wp["category"]} #{wp["severity"]} #{wp["description"]}" - end) - |> Enum.join(" ") - |> String.slice(0, 2000) - end - - defp extract_categories(weak_points) do - weak_points - |> Enum.map(& &1["category"]) - |> Enum.reject(&is_nil/1) - |> Enum.uniq() - end - - defp count_severities(weak_points) do - weak_points - |> Enum.group_by(& &1["severity"]) - |> Map.new(fn {severity, items} -> {severity, length(items)} end) - end - - # --------------------------------------------------------------------------- - # Private: Helpers - # --------------------------------------------------------------------------- - - defp extract_repo_name(nil), do: "unknown" - - defp extract_repo_name(path) when is_binary(path) do - path - |> String.split("/") - |> List.last() - |> String.replace(~r/[^a-zA-Z0-9_-]/, "") - end - - defp timestamp_id do - DateTime.utc_now() - |> DateTime.to_iso8601(:basic) - |> String.replace(~r/[^0-9]/, "") - |> String.slice(0, 14) - end - - # --------------------------------------------------------------------------- - # Private: Local ETS Storage (fallback when Rust core unavailable) - # --------------------------------------------------------------------------- - - defp ensure_ets_table do - case :ets.info(:hypatia_scans) do - :undefined -> - :ets.new(:hypatia_scans, [:named_table, :set, :public]) - - _ -> - :ok - end - end - - defp store_local(octad_id, octad_input) do - ensure_ets_table() - :ets.insert(:hypatia_scans, {octad_id, octad_input}) - :ok - end -end diff --git a/verisimdb/elixir-orchestration/lib/verisim/nif_bridge.ex b/verisimdb/elixir-orchestration/lib/verisim/nif_bridge.ex deleted file mode 100644 index d574606b..00000000 --- a/verisimdb/elixir-orchestration/lib/verisim/nif_bridge.ex +++ /dev/null @@ -1,110 +0,0 @@ -# SPDX-License-Identifier: MPL-2.0 - -defmodule VeriSim.NifBridge do - @moduledoc """ - NIF bridge to the Rust core via Rustler. - - Provides direct in-process calls to VeriSimDB's Rust engine, bypassing HTTP - for same-node deployments. When the NIF is loaded, operations execute ~10-100x - faster than the HTTP transport. - - ## Loading - - The NIF shared library is loaded from `priv/native/libverisim_nif.so` (Linux) - or `priv/native/libverisim_nif.dylib` (macOS). If the library is not present - (e.g., in a pure-Elixir development setup), all functions return - `{:error, :nif_not_loaded}`. - - ## Transport Selection - - This module is not called directly. Instead, `VeriSim.Transport` selects the - transport based on `VERISIM_TRANSPORT`: - - VERISIM_TRANSPORT=http # Default: HTTP via VeriSim.RustClient - VERISIM_TRANSPORT=nif # Direct NIF calls (this module) - VERISIM_TRANSPORT=auto # NIF if loaded, HTTP fallback - - ## Functions - - All functions accept and return JSON strings for compatibility with the HTTP - transport (same serialisation format). - """ - - @on_load :load_nif - - @doc false - def load_nif do - nif_path = - :verisim - |> :code.priv_dir() - |> Path.join("native/libverisim_nif") - - case :erlang.load_nif(String.to_charlist(nif_path), 0) do - :ok -> :ok - {:error, {:load_failed, _}} -> :ok # NIF not available — stubs will be used - {:error, {:reload, _}} -> :ok # Already loaded - {:error, reason} -> - require Logger - Logger.debug("VeriSim.NifBridge: NIF not loaded (#{inspect(reason)})") - :ok - end - end - - @doc """ - Create a new octad entity from a JSON string. - - Returns `{:ok, json}` on success, `{:error, reason}` on failure. - """ - def create_octad(_json_input), do: {:error, :nif_not_loaded} - - @doc """ - Retrieve a octad by ID. - - Returns the full octad JSON with all 8 octad modalities. - """ - def get_octad(_octad_id), do: {:error, :nif_not_loaded} - - @doc """ - Delete a octad entity by ID. - """ - def delete_octad(_octad_id), do: {:error, :nif_not_loaded} - - @doc """ - Full-text search across the document modality. - """ - def search_text(_query, _limit), do: {:error, :nif_not_loaded} - - @doc """ - Vector similarity search. - - Accepts a JSON-encoded embedding vector and a k parameter. - """ - def search_vector(_embedding_json, _k), do: {:error, :nif_not_loaded} - - @doc """ - Paginated listing of octad entities. - """ - def list_octads(_limit, _offset), do: {:error, :nif_not_loaded} - - @doc """ - Get drift detection scores for a specific entity. - - Returns drift scores across all 8 octad modalities. - """ - def get_drift_score(_octad_id), do: {:error, :nif_not_loaded} - - @doc """ - Trigger normalisation (self-repair) for a drifted entity. - """ - def trigger_normalise(_octad_id), do: {:error, :nif_not_loaded} - - @doc """ - Check whether the NIF bridge is loaded and operational. - """ - def loaded? do - case get_octad("__health_check__") do - {:error, :nif_not_loaded} -> false - _ -> true - end - end -end diff --git a/verisimdb/elixir-orchestration/lib/verisim/query/query_router.ex b/verisimdb/elixir-orchestration/lib/verisim/query/query_router.ex deleted file mode 100644 index 178341b7..00000000 --- a/verisimdb/elixir-orchestration/lib/verisim/query/query_router.ex +++ /dev/null @@ -1,186 +0,0 @@ -# SPDX-License-Identifier: MPL-2.0 - -defmodule VeriSim.QueryRouter do - @moduledoc """ - Query Router - Distributes queries across modalities and nodes. - - Routes queries to the appropriate modality store based on query type, - and aggregates results from multiple sources when needed. - - ## Query Types - - - `:text` - Full-text search (Document modality) - - `:vector` - Similarity search (Vector modality) - - `:graph` - Relationship traversal (Graph modality) - - `:semantic` - Type-based queries (Semantic modality) - - `:temporal` - Time-based queries (Temporal modality) - - `:multi` - Cross-modal queries (multiple modalities) - """ - - use GenServer - require Logger - - alias VeriSim.RustClient - - # Client API - - def start_link(opts \\ []) do - GenServer.start_link(__MODULE__, opts, name: __MODULE__) - end - - @doc """ - Execute a query. - - ## Examples - - # Text search - QueryRouter.query(:text, "machine learning", limit: 10) - - # Vector similarity - QueryRouter.query(:vector, [0.1, 0.2, ...], k: 5) - - # Graph traversal - QueryRouter.query(:graph, %{start: "entity-1", predicate: "relates_to"}) - - # Multi-modal - QueryRouter.query(:multi, %{text: "AI", types: ["https://example.org/Paper"]}) - """ - def query(type, params, opts \\ []) do - GenServer.call(__MODULE__, {:query, type, params, opts}) - end - - @doc """ - Get query statistics. - """ - def stats do - GenServer.call(__MODULE__, :stats) - end - - # Server Callbacks - - @impl true - def init(_opts) do - state = %{ - query_count: 0, - query_by_type: %{}, - avg_latency_ms: 0.0, - total_latency_ms: 0 - } - {:ok, state} - end - - @impl true - def handle_call({:query, type, params, opts}, _from, state) do - start_time = System.monotonic_time(:millisecond) - - result = execute_query(type, params, opts) - - end_time = System.monotonic_time(:millisecond) - latency = end_time - start_time - - new_state = update_stats(state, type, latency) - - {:reply, result, new_state} - end - - @impl true - def handle_call(:stats, _from, state) do - stats = %{ - total_queries: state.query_count, - queries_by_type: state.query_by_type, - avg_latency_ms: state.avg_latency_ms - } - {:reply, stats, state} - end - - # Private Functions - - defp execute_query(:text, query, opts) when is_binary(query) do - limit = Keyword.get(opts, :limit, 10) - RustClient.search_text(query, limit) - end - - defp execute_query(:vector, vector, opts) when is_list(vector) do - k = Keyword.get(opts, :k, 10) - RustClient.search_vector(vector, k) - end - - defp execute_query(:graph, %{start: entity_id}, _opts) do - RustClient.get_related(entity_id) - end - - defp execute_query(:semantic, %{types: types}, opts) do - # Type-based query: search for entities matching a type IRI via text search - limit = Keyword.get(opts, :limit, 10) - - types - |> List.wrap() - |> Enum.flat_map(fn type_iri -> - case RustClient.search_text("type:#{type_iri}", limit) do - {:ok, results} -> results - {:error, _} -> [] - end - end) - |> Enum.uniq_by(& &1["id"]) - |> Enum.take(limit) - |> then(&{:ok, &1}) - end - - defp execute_query(:temporal, %{entity_id: entity_id} = params, _opts) do - # Get entity version at a specific timestamp - query_params = - case Map.get(params, :time) do - nil -> [] - time -> [at: to_string(time)] - end - - RustClient.get("/octads/#{entity_id}/versions", query_params) - end - - defp execute_query(:multi, params, opts) do - # Multi-modal query - combine results from multiple modalities - results = [] - - results = - if text = Map.get(params, :text) do - case execute_query(:text, text, opts) do - {:ok, text_results} -> results ++ text_results - _ -> results - end - else - results - end - - results = - if vector = Map.get(params, :vector) do - case execute_query(:vector, vector, opts) do - {:ok, vector_results} -> results ++ vector_results - _ -> results - end - else - results - end - - # Deduplicate and rank - {:ok, Enum.uniq_by(results, & &1["id"])} - end - - defp execute_query(type, _params, _opts) do - {:error, {:unknown_query_type, type}} - end - - defp update_stats(state, type, latency) do - new_count = state.query_count + 1 - new_total = state.total_latency_ms + latency - new_avg = new_total / new_count - - new_by_type = Map.update(state.query_by_type, type, 1, &(&1 + 1)) - - %{state | - query_count: new_count, - query_by_type: new_by_type, - avg_latency_ms: new_avg, - total_latency_ms: new_total - } - end -end diff --git a/verisimdb/elixir-orchestration/lib/verisim/query/vcl_bridge.ex b/verisimdb/elixir-orchestration/lib/verisim/query/vcl_bridge.ex deleted file mode 100644 index d288139f..00000000 --- a/verisimdb/elixir-orchestration/lib/verisim/query/vcl_bridge.ex +++ /dev/null @@ -1,744 +0,0 @@ -# SPDX-License-Identifier: MPL-2.0 - -defmodule VeriSim.Query.VCLBridge do - @moduledoc """ - Bridge between the ReScript VCL parser and Elixir VCL executor. - - Manages a long-running Deno/Node process that runs the compiled ReScript - VCL parser. Communication uses JSON over stdin/stdout with length-prefixed - framing for reliable message boundaries. - - ## Architecture - - VCL String ──► VCLBridge (GenServer) - │ - ▼ - Port (stdin/stdout JSON) - │ - ▼ - Deno process running compiled VCLParser.res - │ - ▼ - Parsed AST (JSON) ──► VCLExecutor - - ## Usage - - {:ok, ast} = VCLBridge.parse("SELECT GRAPH, VECTOR FROM HEXAD abc-123") - {:ok, results} = VCLExecutor.execute(ast) - """ - - use GenServer - require Logger - - @default_timeout 5_000 - @parser_script_path "vcl-bridge/vcl_parser_port.js" - - # --------------------------------------------------------------------------- - # Client API - # --------------------------------------------------------------------------- - - def start_link(opts \\ []) do - GenServer.start_link(__MODULE__, opts, name: __MODULE__) - end - - @doc """ - Parse a VCL query string into an AST map. - - Returns `{:ok, ast}` where ast is a map matching the VCLParser.AST types, - or `{:error, reason}` on parse failure. - """ - def parse(query_string, timeout \\ @default_timeout) do - GenServer.call(__MODULE__, {:parse, query_string}, timeout) - end - - @doc """ - Parse a slipstream query (no PROOF clause allowed). - """ - def parse_slipstream(query_string, timeout \\ @default_timeout) do - GenServer.call(__MODULE__, {:parse_slipstream, query_string}, timeout) - end - - @doc """ - Parse a dependent-type query (PROOF clause required). - """ - def parse_dependent(query_string, timeout \\ @default_timeout) do - GenServer.call(__MODULE__, {:parse_dependent, query_string}, timeout) - end - - @doc """ - Parse a VCL mutation (INSERT / UPDATE / DELETE). - """ - def parse_mutation(query_string, timeout \\ @default_timeout) do - GenServer.call(__MODULE__, {:parse_mutation, query_string}, timeout) - end - - @doc """ - Parse a VCL statement (query or mutation). - """ - def parse_statement(query_string, timeout \\ @default_timeout) do - GenServer.call(__MODULE__, {:parse_statement, query_string}, timeout) - end - - @doc """ - Parse and execute a VCL query string in one call. - Combines VCLBridge.parse/1 with VCLExecutor.execute/2. - """ - def parse_and_execute(query_string, opts \\ []) do - case parse(query_string) do - {:ok, ast} -> VeriSim.Query.VCLExecutor.execute(ast, opts) - {:error, _} = error -> error - end - end - - @doc """ - Parse and execute a VCL statement (query or mutation). - """ - def parse_and_execute_statement(query_string, opts \\ []) do - case parse_statement(query_string) do - {:ok, ast} -> VeriSim.Query.VCLExecutor.execute_statement(ast, opts) - {:error, _} = error -> error - end - end - - @doc """ - Type-check a parsed VCL-UT AST. - - Sends the AST to the ReScript type checker (VCLBidir.synthesizeQuery) - via the Deno/Node subprocess. Returns inferred types, proof obligations, - and composition strategy. - - Falls back to `{:error, :type_checker_unavailable}` if the subprocess is - not running — callers MUST NOT silently bypass this. - - ## Returns - - - `{:ok, type_info}` where type_info contains: - - `:proof_obligations` — list of proof specs from the type checker - - `:composition_strategy` — how proofs should be composed (:conjunction, :disjunction, etc.) - - `:inferred_types` — map of modality → field → type - - `{:error, :type_checker_unavailable}` — subprocess not running - - `{:error, reason}` — type checking failed (invalid query, type mismatch, etc.) - """ - def typecheck(ast, timeout \\ @default_timeout) do - GenServer.call(__MODULE__, {:typecheck, ast}, timeout) - end - - # --------------------------------------------------------------------------- - # GenServer callbacks - # --------------------------------------------------------------------------- - - @impl true - def init(opts) do - runtime = Keyword.get(opts, :runtime, detect_runtime()) - script_path = Keyword.get(opts, :script_path, resolve_script_path()) - - state = %{ - port: nil, - runtime: runtime, - script_path: script_path, - pending: %{}, - next_id: 1 - } - - case start_port(state) do - {:ok, port} -> - Logger.info("VCLBridge started with #{runtime}") - {:ok, %{state | port: port}} - - {:error, reason} -> - Logger.warning("VCLBridge: parser process unavailable (#{reason}), falling back to built-in parser") - {:ok, state} - end - end - - @impl true - def handle_call({:typecheck, _ast}, _from, %{port: nil} = state) do - # Type checking requires the ReScript subprocess — no fallback. - # Callers MUST handle this error; silently passing would defeat VCL-UT. - {:reply, {:error, :type_checker_unavailable}, state} - end - - @impl true - def handle_call({:typecheck, ast}, from, state) do - id = state.next_id - message = Jason.encode!(%{ - "id" => id, - "action" => "typecheck", - "ast" => ast - }) - - send_to_port(state.port, message) - pending = Map.put(state.pending, id, from) - {:noreply, %{state | pending: pending, next_id: id + 1}} - end - - @impl true - def handle_call({action, query_string}, _from, %{port: nil} = state) - when action in [:parse, :parse_slipstream, :parse_dependent, - :parse_mutation, :parse_statement] do - # Fallback: no external parser available, use built-in Elixir parser - result = builtin_parse(query_string, action) - {:reply, result, state} - end - - @impl true - def handle_call({action, query_string}, from, state) - when action in [:parse, :parse_slipstream, :parse_dependent, - :parse_mutation, :parse_statement] do - id = state.next_id - message = Jason.encode!(%{ - "id" => id, - "action" => Atom.to_string(action), - "query" => query_string - }) - - # Send length-prefixed message to port - send_to_port(state.port, message) - - pending = Map.put(state.pending, id, from) - {:noreply, %{state | pending: pending, next_id: id + 1}} - end - - @impl true - def handle_info({port, {:data, data}}, %{port: port} = state) do - case Jason.decode(IO.iodata_to_binary(data)) do - {:ok, %{"id" => id, "ok" => ast}} -> - case Map.pop(state.pending, id) do - {nil, _} -> {:noreply, state} - {from, pending} -> - GenServer.reply(from, {:ok, atomize_keys(ast)}) - {:noreply, %{state | pending: pending}} - end - - {:ok, %{"id" => id, "error" => reason}} -> - case Map.pop(state.pending, id) do - {nil, _} -> {:noreply, state} - {from, pending} -> - GenServer.reply(from, {:error, reason}) - {:noreply, %{state | pending: pending}} - end - - {:error, _} -> - Logger.warning("VCLBridge: received malformed data from parser port") - {:noreply, state} - end - end - - @impl true - def handle_info({port, {:exit_status, status}}, %{port: port} = state) do - Logger.warning("VCLBridge: parser process exited with status #{status}") - - # Reply to all pending requests with error - for {_id, from} <- state.pending do - GenServer.reply(from, {:error, :parser_crashed}) - end - - # Try to restart - case start_port(state) do - {:ok, new_port} -> - Logger.info("VCLBridge: parser process restarted") - {:noreply, %{state | port: new_port, pending: %{}}} - - {:error, _} -> - {:noreply, %{state | port: nil, pending: %{}}} - end - end - - @impl true - def handle_info(_msg, state), do: {:noreply, state} - - # --------------------------------------------------------------------------- - # Private - # --------------------------------------------------------------------------- - - defp start_port(state) do - runtime = state.runtime - script = state.script_path - - if runtime && File.exists?(script) do - try do - port = Port.open( - {:spawn_executable, runtime}, - [:binary, :exit_status, {:args, [script]}, {:line, 1_048_576}] - ) - {:ok, port} - rescue - e -> {:error, Exception.message(e)} - end - else - {:error, "runtime (#{runtime}) or script (#{script}) not found"} - end - end - - defp send_to_port(port, message) do - Port.command(port, message <> "\n") - end - - defp detect_runtime do - cond do - System.find_executable("deno") -> System.find_executable("deno") - System.find_executable("node") -> System.find_executable("node") - true -> nil - end - end - - defp resolve_script_path do - # Look relative to the project root - base = Application.get_env(:verisim, :project_root, File.cwd!()) - Path.join(base, @parser_script_path) - end - - defp atomize_keys(map) when is_map(map) do - Map.new(map, fn - {key, value} when is_binary(key) -> - {String.to_existing_atom(key), atomize_keys(value)} - {key, value} -> - {key, atomize_keys(value)} - end) - rescue - ArgumentError -> map - end - - defp atomize_keys(list) when is_list(list), do: Enum.map(list, &atomize_keys/1) - defp atomize_keys(other), do: other - - # --------------------------------------------------------------------------- - # Built-in Elixir Parser (fallback when Deno/Node unavailable) - # --------------------------------------------------------------------------- - - defp builtin_parse(query_string, action) do - query_string = String.trim(query_string) - - case action do - :parse_mutation -> - with {:ok, tokens} <- tokenize(query_string), - {:ok, mutation} <- parse_mutation_tokens(tokens) do - {:ok, mutation} - end - - :parse_statement -> - with {:ok, tokens} <- tokenize(query_string) do - first = tokens |> List.first() |> to_string() |> String.upcase() - case first do - cmd when cmd in ["INSERT", "UPDATE", "DELETE"] -> - with {:ok, mutation} <- parse_mutation_tokens(tokens) do - {:ok, %{TAG: "Mutation", _0: mutation}} - end - _ -> - with {:ok, ast} <- parse_tokens(tokens) do - {:ok, %{TAG: "Query", _0: ast}} - end - end - end - - _ -> - with {:ok, tokens} <- tokenize(query_string), - {:ok, ast} <- parse_tokens(tokens) do - case action do - :parse_slipstream -> - if ast[:proof], do: {:error, "Slipstream queries cannot have PROOF clause"}, else: {:ok, ast} - :parse_dependent -> - if ast[:proof], do: {:ok, ast}, else: {:error, "Dependent-type queries require PROOF clause"} - :parse -> - {:ok, ast} - end - end - end - end - - defp tokenize(input) do - # Reject null bytes before tokenising: they truncate C strings silently at the - # Rust FFI boundary, allowing entity IDs to be forged via truncation. - if String.contains?(input, "\0") do - {:error, "VCL query contains null bytes"} - else - # Simple whitespace-aware tokenizer - tokens = - input - |> String.replace(~r/\s+/, " ") - |> String.split(" ", trim: true) - - {:ok, tokens} - end - end - - defp parse_tokens(tokens) do - with {:ok, modalities, projections, aggregates, rest} <- parse_select_extended(tokens), - {:ok, source, rest} <- parse_from(rest), - {:ok, where_clause, rest} <- parse_where(rest), - {:ok, group_by, rest} <- parse_group_by(rest), - {:ok, having, rest} <- parse_having(rest), - {:ok, proof, rest} <- parse_proof(rest), - {:ok, order_by, rest} <- parse_order_by(rest), - {:ok, limit, rest} <- parse_limit(rest), - {:ok, offset, _rest} <- parse_offset(rest) do - {:ok, %{ - modalities: modalities, - projections: projections, - aggregates: aggregates, - source: source, - where: where_clause, - groupBy: group_by, - having: having, - proof: proof, - orderBy: order_by, - limit: limit, - offset: offset - }} - end - end - - defp parse_select(["SELECT" | rest]) do - {modalities, rest} = take_modalities(rest, []) - {:ok, modalities, rest} - end - - defp parse_select(_), do: {:error, "Expected SELECT"} - - defp take_modalities(["GRAPH" | rest], acc), do: take_modalities(strip_comma(rest), [:graph | acc]) - defp take_modalities(["VECTOR" | rest], acc), do: take_modalities(strip_comma(rest), [:vector | acc]) - defp take_modalities(["TENSOR" | rest], acc), do: take_modalities(strip_comma(rest), [:tensor | acc]) - defp take_modalities(["SEMANTIC" | rest], acc), do: take_modalities(strip_comma(rest), [:semantic | acc]) - defp take_modalities(["DOCUMENT" | rest], acc), do: take_modalities(strip_comma(rest), [:document | acc]) - defp take_modalities(["TEMPORAL" | rest], acc), do: take_modalities(strip_comma(rest), [:temporal | acc]) - defp take_modalities(["PROVENANCE" | rest], acc), do: take_modalities(strip_comma(rest), [:provenance | acc]) - defp take_modalities(["SPATIAL" | rest], acc), do: take_modalities(strip_comma(rest), [:spatial | acc]) - defp take_modalities(["*" | rest], acc), do: take_modalities(strip_comma(rest), [:all | acc]) - defp take_modalities(rest, acc), do: {Enum.reverse(acc), rest} - - defp strip_comma(["," <> token | rest]) when token != "" do - [token | rest] - end - defp strip_comma(["," | rest]), do: rest - defp strip_comma(rest), do: rest - - defp parse_from(["FROM", "HEXAD", uuid | rest]) do - {:ok, {:octad, uuid}, rest} - end - - defp parse_from(["FROM", "FEDERATION", pattern | rest]) do - {drift_policy, rest} = parse_drift_policy(rest) - {:ok, {:federation, pattern, drift_policy}, rest} - end - - defp parse_from(["FROM", "STORE", store_id | rest]) do - {:ok, {:store, store_id}, rest} - end - - defp parse_from(_), do: {:error, "Expected FROM clause"} - - defp parse_drift_policy(["WITH", "DRIFT", policy | rest]) do - drift = case String.upcase(policy) do - "STRICT" -> :strict - "REPAIR" -> :repair - "TOLERATE" -> :tolerate - "LATEST" -> :latest - _ -> nil - end - {drift, rest} - end - - defp parse_drift_policy(rest), do: {nil, rest} - - defp parse_where(["WHERE" | rest]) do - # Simplified: collect everything until PROOF, LIMIT, OFFSET, or end - {condition_tokens, rest} = Enum.split_while(rest, fn token -> - token not in ["PROOF", "LIMIT", "OFFSET"] - end) - - condition = if condition_tokens == [] do - nil - else - %{raw: Enum.join(condition_tokens, " ")} - end - - {:ok, condition, rest} - end - - defp parse_where(rest), do: {:ok, nil, rest} - - defp parse_proof(["PROOF" | rest]) do - {proof_tokens, rest} = Enum.split_while(rest, fn token -> - token not in ["LIMIT", "OFFSET"] - end) - - raw = Enum.join(proof_tokens, " ") - - # Split multi-proof specs on AND/OR connectors into a list. - # "EXISTENCE(a) AND PROVENANCE(b)" → [%{raw: "EXISTENCE(a)"}, %{raw: "PROVENANCE(b)"}] - specs = VeriSim.Query.VCLTypeChecker.parse_proof_specs(%{raw: raw}) - - proof = case specs do - [] -> %{raw: raw} - [single] -> single - multiple -> multiple - end - - {:ok, proof, rest} - end - - defp parse_proof(rest), do: {:ok, nil, rest} - - defp parse_limit(["LIMIT", n | rest]) do - case Integer.parse(n) do - {limit, _} -> {:ok, limit, rest} - :error -> {:error, "Invalid LIMIT value"} - end - end - - defp parse_limit(rest), do: {:ok, nil, rest} - - defp parse_offset(["OFFSET", n | rest]) do - case Integer.parse(n) do - {offset, _} -> {:ok, offset, rest} - :error -> {:error, "Invalid OFFSET value"} - end - end - - defp parse_offset(rest), do: {:ok, nil, rest} - - # Extended SELECT parser: handles MODALITY.field projections and aggregates - defp parse_select_extended(["SELECT" | rest]) do - {items, rest} = take_select_items(rest, [], [], []) - {:ok, elem(items, 0), elem(items, 1), elem(items, 2), rest} - end - - defp parse_select_extended(_), do: {:error, "Expected SELECT"} - - @modality_names ~w(GRAPH VECTOR TENSOR SEMANTIC DOCUMENT TEMPORAL PROVENANCE SPATIAL) - @aggregate_funcs ~w(COUNT SUM AVG MIN MAX) - - # Safe atom conversion using allowlist — prevents atom table exhaustion - @safe_atoms %{ - "graph" => :graph, "vector" => :vector, "tensor" => :tensor, - "semantic" => :semantic, "document" => :document, "temporal" => :temporal, - "provenance" => :provenance, "spatial" => :spatial, - "all" => :all, - "count" => :count, "sum" => :sum, "avg" => :avg, "min" => :min, "max" => :max - } - - defp safe_to_atom(str) when is_binary(str) do - downcased = String.downcase(str) - - case Map.fetch(@safe_atoms, downcased) do - {:ok, atom} -> - atom - - :error -> - raise ArgumentError, - "Unknown VCL token #{inspect(str)} — not in @safe_atoms allowlist" - end - end - - defp take_select_items(tokens, mods, projs, aggs) do - case tokens do - # COUNT(*) - [func, "(*)" | rest] when func in @aggregate_funcs -> - take_select_items(strip_comma(rest), mods, projs, [:count_all | aggs]) - - [func, "(", "*", ")" | rest] when func in @aggregate_funcs -> - take_select_items(strip_comma(rest), mods, projs, [:count_all | aggs]) - - # FUNC(MODALITY.field) - [func | rest] when func in @aggregate_funcs -> - case parse_aggregate_arg(rest) do - {:ok, mod, field, rest} -> - agg = {:aggregate_field, safe_to_atom(func), %{modality: mod, field: field}} - mod_atom = safe_to_atom(mod) - mods = if mod_atom in mods, do: mods, else: [mod_atom | mods] - take_select_items(strip_comma(rest), mods, projs, [agg | aggs]) - _ -> - {{Enum.reverse(mods), nilify(projs), nilify(aggs)}, tokens} - end - - # MODALITY.field (column projection) - [token | rest] -> - case String.split(token, ".", parts: 2) do - [mod_str, field] when mod_str in @modality_names -> - proj = %{modality: safe_to_atom(mod_str), field: field} - mod_atom = safe_to_atom(mod_str) - mods = if mod_atom in mods, do: mods, else: [mod_atom | mods] - take_select_items(strip_comma(rest), mods, [proj | projs], aggs) - - _ -> - # Try as bare modality - up = String.upcase(String.replace(token, ",", "")) - cond do - up in @modality_names -> - mod_atom = safe_to_atom(up) - take_select_items(strip_comma(rest), [mod_atom | mods], projs, aggs) - up == "*" -> - take_select_items(strip_comma(rest), [:all | mods], projs, aggs) - true -> - {{Enum.reverse(mods), nilify(projs), nilify(aggs)}, tokens} - end - end - - [] -> - {{Enum.reverse(mods), nilify(projs), nilify(aggs)}, []} - end - end - - defp parse_aggregate_arg(["(" <> rest_token | rest]) do - # Handle "(MODALITY.field)" — may be split across tokens - inner = String.trim_trailing(rest_token, ")") - case String.split(inner, ".", parts: 2) do - [mod, field] when mod in @modality_names -> - rest = case rest do - [")" | r] -> r - _ -> rest - end - {:ok, mod, field, rest} - _ -> :error - end - end - defp parse_aggregate_arg(_), do: :error - - defp nilify([]), do: nil - defp nilify(list), do: Enum.reverse(list) - - # GROUP BY parser - defp parse_group_by(["GROUP", "BY" | rest]) do - {fields, rest} = take_field_refs(rest, []) - {:ok, if(fields == [], do: nil, else: fields), rest} - end - - defp parse_group_by(rest), do: {:ok, nil, rest} - - defp take_field_refs([token | rest], acc) do - clean = String.replace(token, ",", "") - case String.split(clean, ".", parts: 2) do - [mod_str, field] when mod_str in @modality_names -> - ref = %{modality: safe_to_atom(mod_str), field: field} - take_field_refs(strip_comma(rest), [ref | acc]) - _ -> - {Enum.reverse(acc), [token | rest]} - end - end - - defp take_field_refs([], acc), do: {Enum.reverse(acc), []} - - # HAVING parser (collects tokens until ORDER/PROOF/LIMIT/OFFSET/end) - defp parse_having(["HAVING" | rest]) do - {condition_tokens, rest} = Enum.split_while(rest, fn token -> - String.upcase(token) not in ["ORDER", "PROOF", "LIMIT", "OFFSET"] - end) - - condition = if condition_tokens == [] do - nil - else - %{raw: Enum.join(condition_tokens, " ")} - end - - {:ok, condition, rest} - end - - defp parse_having(rest), do: {:ok, nil, rest} - - # ORDER BY parser - defp parse_order_by(["ORDER", "BY" | rest]) do - {items, rest} = take_order_items(rest, []) - {:ok, if(items == [], do: nil, else: items), rest} - end - - defp parse_order_by(rest), do: {:ok, nil, rest} - - defp take_order_items([token | rest], acc) do - clean = String.replace(token, ",", "") - case String.split(clean, ".", parts: 2) do - [mod_str, field] when mod_str in @modality_names -> - {direction, rest} = case rest do - ["ASC" | r] -> {:asc, strip_comma(r)} - ["DESC" | r] -> {:desc, strip_comma(r)} - ["ASC," <> _ | _] -> {:asc, strip_comma(rest)} - ["DESC," <> _ | _] -> {:desc, strip_comma(rest)} - _ -> {:asc, strip_comma(rest)} - end - - item = %{ - field: %{modality: safe_to_atom(mod_str), field: field}, - direction: direction - } - take_order_items(rest, [item | acc]) - - _ -> - {Enum.reverse(acc), [token | rest]} - end - end - - defp take_order_items([], acc), do: {Enum.reverse(acc), []} - - # --------------------------------------------------------------------------- - # Mutation Parser (INSERT / UPDATE / DELETE) - # --------------------------------------------------------------------------- - - defp parse_mutation_tokens(["INSERT", "HEXAD", "WITH" | rest]) do - {modality_data, rest} = take_modality_data(rest, []) - {:ok, proof, _rest} = parse_proof(rest) - {:ok, %{ - TAG: "Insert", - modalities: modality_data, - proof: proof - }} - end - - defp parse_mutation_tokens(["UPDATE", "HEXAD", uuid, "SET" | rest]) do - {sets, rest} = take_set_assignments(rest, []) - {:ok, proof, _rest} = parse_proof(rest) - {:ok, %{ - TAG: "Update", - octadId: uuid, - sets: sets, - proof: proof - }} - end - - defp parse_mutation_tokens(["DELETE", "HEXAD", uuid | rest]) do - {:ok, proof, _rest} = parse_proof(rest) - {:ok, %{ - TAG: "Delete", - octadId: uuid, - proof: proof - }} - end - - defp parse_mutation_tokens(_), do: {:error, "Expected INSERT, UPDATE, or DELETE"} - - # Parse modality data for INSERT: DOCUMENT(field=value, ...), VECTOR([...]), etc. - defp take_modality_data(tokens, acc) do - case tokens do - [mod | rest] when mod in @modality_names -> - case rest do - ["(" <> inner_start | rest2] -> - {inner_tokens, rest3} = collect_until_close_paren([inner_start | rest2], []) - data = %{modality: safe_to_atom(mod), raw: Enum.join(inner_tokens, " ")} - take_modality_data(strip_comma(rest3), [data | acc]) - _ -> - {Enum.reverse(acc), tokens} - end - _ -> - {Enum.reverse(acc), tokens} - end - end - - defp collect_until_close_paren([], acc), do: {Enum.reverse(acc), []} - defp collect_until_close_paren([token | rest], acc) do - if String.ends_with?(token, ")") do - cleaned = String.trim_trailing(token, ")") - if cleaned != "", do: {Enum.reverse([cleaned | acc]), rest}, else: {Enum.reverse(acc), rest} - else - collect_until_close_paren(rest, [token | acc]) - end - end - - # Parse SET assignments for UPDATE: field = value, field = value - defp take_set_assignments(tokens, acc) do - case tokens do - [field, "=", value | rest] -> - assignment = %{field: field, value: value} - take_set_assignments(strip_comma(rest), [assignment | acc]) - _ -> - {Enum.reverse(acc), tokens} - end - end -end diff --git a/verisimdb/elixir-orchestration/lib/verisim/query/vcl_executor.ex b/verisimdb/elixir-orchestration/lib/verisim/query/vcl_executor.ex deleted file mode 100644 index 53690674..00000000 --- a/verisimdb/elixir-orchestration/lib/verisim/query/vcl_executor.ex +++ /dev/null @@ -1,1812 +0,0 @@ -# SPDX-License-Identifier: MPL-2.0 - -defmodule VeriSim.Query.VCLExecutor do - @moduledoc """ - VCL Executor - Executes VCL queries and mutations parsed by the ReScript VCL parser. - - Supports three phases: - 1. Dependent-type queries with multi-proof composition - 2. Cross-modal correlation conditions - 3. Write path (INSERT / UPDATE / DELETE) - - ## Execution Pipeline - - 1. Parse VCL query (done by ReScript VCLParser) - 2. Type-check if PROOF clause present (VCLTypeChecker → VCLBidir) - 3. Generate execution plan (VCLExplain) - 4. Classify conditions (pushdown vs cross-modal) - 5. Route to appropriate modality stores - 6. Evaluate cross-modal conditions post-fetch - 7. Aggregate and return results - """ - - require Logger - - alias VeriSim.{QueryRouter, RustClient, Telemetry} - alias VeriSim.Query.VCLProofCertificate - - @doc """ - Execute a parsed VCL query. - - ## Parameters - - - `query_ast` - Parsed VCL query from ReScript VCLParser - - `opts` - Execution options - - ## Options - - - `:explain` - Return execution plan instead of results - - `:timeout` - Query timeout in milliseconds (default: 30000) - - `:verify_proof` - Verify ZKP proofs for dependent-type queries (default: true) - - ## Returns - - - `{:ok, results}` - Query results - - `{:error, reason}` - Query execution failed - """ - def execute(query_ast, opts \\ []) do - explain = Keyword.get(opts, :explain, false) - timeout = Keyword.get(opts, :timeout, 30_000) - - if explain do - {:ok, generate_explain_plan(query_ast)} - else - execute_query(query_ast, timeout) - end - end - - @doc """ - Execute a VCL mutation (INSERT / UPDATE / DELETE). - """ - def execute_mutation(mutation_ast, opts \\ []) do - timeout = Keyword.get(opts, :timeout, 30_000) - - case mutation_ast do - %{TAG: "Insert", modalities: modality_data, proof: proof} -> - execute_insert(modality_data, proof, timeout) - - %{TAG: "Update", octadId: octad_id, sets: sets, proof: proof} -> - execute_update(octad_id, sets, proof, timeout) - - %{TAG: "Delete", octadId: octad_id, proof: proof} -> - execute_delete(octad_id, proof, timeout) - - _ -> - {:error, {:invalid_mutation, "Unknown mutation type"}} - end - end - - @doc """ - Execute a VCL statement (query or mutation). - """ - def execute_statement(statement_ast, opts \\ []) do - case statement_ast do - %{TAG: "Query", _0: query} -> execute(query, opts) - %{TAG: "Mutation", _0: mutation} -> execute_mutation(mutation, opts) - # Map with :type key from built-in parser - %{type: :query} -> execute(statement_ast, opts) - %{type: :mutation, mutation: mutation} -> execute_mutation(mutation, opts) - _ -> execute(statement_ast, opts) - end - end - - @doc """ - Execute a VCL query string (includes parsing). - """ - def execute_string(query_string, opts \\ []) do - start_time = System.monotonic_time() - - result = - case VeriSim.Query.VCLBridge.parse(query_string) do - {:ok, ast} -> execute(ast, opts) - {:error, reason} -> {:error, {:parse_error, reason}} - end - - # Emit telemetry for product insights (aggregate-only, no query content). - duration = System.monotonic_time() - start_time - statement_type = classify_statement_type(query_string) - modalities = extract_modalities_from_string(query_string) - - case result do - {:ok, _} -> - Telemetry.emit_query_stop(duration, %{ - statement_type: statement_type, - modalities: modalities - }) - {:error, _} -> - Telemetry.emit_query_exception(%{statement_type: statement_type}) - end - - result - end - - # Classify a query string into a statement type for telemetry (no content captured). - defp classify_statement_type(query_string) do - upper = String.upcase(String.trim(query_string)) - cond do - String.starts_with?(upper, "SELECT") -> "SELECT" - String.starts_with?(upper, "INSERT") -> "INSERT" - String.starts_with?(upper, "UPDATE") -> "UPDATE" - String.starts_with?(upper, "DELETE") -> "DELETE" - String.starts_with?(upper, "EXPLAIN") -> "EXPLAIN" - String.starts_with?(upper, "SHOW") -> "SHOW" - String.starts_with?(upper, "SEARCH") -> "SEARCH" - String.starts_with?(upper, "COUNT") -> "COUNT" - true -> "OTHER" - end - end - - # Extract modality names from a query string for telemetry (names only, no content). - defp extract_modalities_from_string(query_string) do - upper = String.upcase(query_string) - ~w(GRAPH VECTOR TENSOR SEMANTIC DOCUMENT TEMPORAL PROVENANCE SPATIAL) - |> Enum.filter(&String.contains?(upper, &1)) - |> Enum.map(&String.downcase/1) - |> Enum.map(&String.to_existing_atom/1) - end - - # =========================================================================== - # Query Execution - # =========================================================================== - - defp execute_query(query_ast, timeout) do - modalities = extract_modalities(query_ast) - source = extract_source(query_ast) - where_clause = extract_where(query_ast) - proof_specs = extract_proof(query_ast) - limit = extract_limit(query_ast) - offset = extract_offset(query_ast) - order_by = extract_order_by(query_ast) - group_by = extract_group_by(query_ast) - aggregates = extract_aggregates(query_ast) - projections = extract_projections(query_ast) - - if proof_specs do - # VCL-UT path: type-check → execute → verify proofs → bundle certificate - execute_dt_query(query_ast, proof_specs, modalities, source, where_clause, - limit, offset, order_by, group_by, aggregates, projections, timeout) - else - # Slipstream path: no proofs, no type checking - execute_slipstream_query(modalities, source, where_clause, - limit, offset, order_by, group_by, aggregates, projections, timeout) - end - end - - # Slipstream execution — fast path, no proofs - defp execute_slipstream_query(modalities, source, where_clause, - limit, offset, order_by, group_by, aggregates, projections, timeout) do - {pushdown_conditions, cross_modal_conditions} = classify_conditions(where_clause) - - result = execute_by_source(source, modalities, pushdown_conditions, limit, offset, timeout) - - case result do - {:ok, rows} -> - rows - |> maybe_evaluate_cross_modal(cross_modal_conditions) - |> maybe_group_and_aggregate(group_by, aggregates) - |> maybe_order_by(order_by) - |> maybe_project_columns(projections) - |> then(&{:ok, &1}) - - error -> error - end - end - - # VCL-UT execution — type check, execute, verify proofs, bundle certificate - defp execute_dt_query(query_ast, proof_specs, modalities, source, where_clause, - limit, offset, order_by, group_by, aggregates, projections, timeout) do - alias VeriSim.Query.{VCLBridge, VCLTypeChecker} - - # Step 1: Type-check the query to get proof obligations and composition strategy. - # Tries three strategies in order: - # 1. ReScript bidirectional type checker (VCLBridge.typecheck — full formal system) - # 2. Elixir-native type checker (VCLTypeChecker — validates types, generates obligations) - # 3. Bare AST extraction (last resort — no validation, just structuring) - type_info = case VCLBridge.typecheck(query_ast) do - {:ok, info} -> - info - - {:error, :type_checker_unavailable} -> - # ReScript subprocess not running. Use the Elixir-native type checker - # which validates proof types, modality compatibility, and composition. - Logger.info("VCL-UT: Using Elixir-native type checker (ReScript subprocess unavailable)") - - case VCLTypeChecker.typecheck(query_ast) do - {:ok, info} -> - info - - {:error, reason} -> - # Native type checker rejected the query — this is a real type error. - Logger.error("VCL-UT: Type checking failed: #{inspect(reason)}") - nil - end - - {:error, reason} -> - Logger.error("VCL-UT: Type checking failed: #{inspect(reason)}") - nil - end - - if is_nil(type_info) do - {:error, {:type_check_failed, "VCL-UT query type checking failed"}} - else - # Step 2: Execute the query (get data) - {pushdown_conditions, cross_modal_conditions} = classify_conditions(where_clause) - data_result = execute_by_source(source, modalities, pushdown_conditions, limit, offset, timeout) - - case data_result do - {:ok, rows} -> - processed_rows = - rows - |> maybe_evaluate_cross_modal(cross_modal_conditions) - |> maybe_group_and_aggregate(group_by, aggregates) - |> maybe_order_by(order_by) - |> maybe_project_columns(projections) - - # Step 3: Verify all proof obligations - obligations = type_info[:proof_obligations] || type_info["proof_obligations"] || proof_specs - proof_result = verify_multi_proof(query_ast, obligations) - - case proof_result do - {:ok, artifacts} -> - # Step 4: Bundle data + proof certificate as ProvedResult. - # The certificate includes the proof artifacts returned by each - # verifier (e.g., hash commitments, Merkle proofs, entity existence - # confirmations) so downstream consumers can independently verify. - query_text = Map.get(query_ast, :raw, "") || "" - composition = type_info[:composition_strategy] || type_info["composition_strategy"] || :conjunction - - # Step 4b: Generate independently verifiable certificates for - # each proof obligation using VCLProofCertificate. Each artifact - # from the verifier becomes the witness for its corresponding - # obligation, yielding a hash-sealed certificate. - verifiable_certificates = - obligations - |> Enum.zip(List.wrap(artifacts)) - |> Enum.map(fn {obligation, artifact} -> - witness = if is_map(artifact), do: artifact, else: %{raw: artifact} - case VCLProofCertificate.generate_certificate(obligation, witness) do - {:ok, cert} -> cert - {:error, _} -> nil - end - end) - |> Enum.reject(&is_nil/1) - - proved_result = %{ - data: processed_rows, - proof_certificate: %{ - proofs: artifacts, - obligations: obligations, - composition: composition, - verified_at: DateTime.utc_now(), - query_hash: :crypto.hash(:sha256, to_string(query_text)) |> Base.encode16(case: :lower), - verifiable_certificates: verifiable_certificates - } - } - {:ok, proved_result} - - {:error, reason} -> - {:error, {:proof_verification_failed, reason}} - end - - error -> error - end - end - end - - # Shared: route query to the appropriate source - defp execute_by_source(source, modalities, pushdown_conditions, limit, offset, timeout) do - case source do - {:octad, entity_id} -> - execute_octad_query(entity_id, modalities, pushdown_conditions, limit, offset, timeout) - - {:federation, pattern, drift_policy} -> - execute_federation_query(pattern, drift_policy, modalities, pushdown_conditions, limit, offset, timeout) - - {:store, store_id} -> - execute_store_query(store_id, modalities, pushdown_conditions, limit, offset, timeout) - - :reflect -> - execute_reflect_query(modalities, pushdown_conditions, limit, offset, timeout) - end - end - - # =========================================================================== - # Phase 2: Condition Classification - # =========================================================================== - - defp classify_conditions(nil), do: {nil, []} - defp classify_conditions(%{raw: _} = condition), do: {condition, []} - defp classify_conditions(condition) when is_map(condition) do - case condition do - %{TAG: "CrossModalFieldCompare"} -> {nil, [condition]} - %{TAG: "ModalityDrift"} -> {nil, [condition]} - %{TAG: "ModalityExists"} -> {nil, [condition]} - %{TAG: "ModalityNotExists"} -> {nil, [condition]} - %{TAG: "ModalityConsistency"} -> {nil, [condition]} - %{TAG: "And", _0: left, _1: right} -> - {push_l, cross_l} = classify_conditions(left) - {push_r, cross_r} = classify_conditions(right) - pushdown = combine_pushdown(push_l, push_r, :and) - {pushdown, cross_l ++ cross_r} - %{TAG: "Or", _0: left, _1: right} -> - {push_l, cross_l} = classify_conditions(left) - {push_r, cross_r} = classify_conditions(right) - pushdown = combine_pushdown(push_l, push_r, :or) - {pushdown, cross_l ++ cross_r} - %{TAG: "Not", _0: inner} -> - {push, cross} = classify_conditions(inner) - {push, cross} - _ -> - # Simple condition: pushdown - {condition, []} - end - end - defp classify_conditions(condition), do: {condition, []} - - defp combine_pushdown(nil, nil, _op), do: nil - defp combine_pushdown(a, nil, _op), do: a - defp combine_pushdown(nil, b, _op), do: b - defp combine_pushdown(a, b, :and), do: %{TAG: "And", _0: a, _1: b} - defp combine_pushdown(a, b, :or), do: %{TAG: "Or", _0: a, _1: b} - - # =========================================================================== - # Phase 2: Cross-Modal Evaluation - # =========================================================================== - - defp maybe_evaluate_cross_modal(rows, []), do: rows - defp maybe_evaluate_cross_modal(rows, cross_modal_conditions) do - Enum.filter(rows, fn octad -> - Enum.all?(cross_modal_conditions, fn condition -> - evaluate_cross_modal(octad, condition) - end) - end) - end - - defp evaluate_cross_modal(octad, condition) do - case condition do - %{TAG: "CrossModalFieldCompare", - _0: mod1, _1: field1, _2: op, _3: mod2, _4: field2} -> - val1 = get_modality_field(octad, mod1, field1) - val2 = get_modality_field(octad, mod2, field2) - compare_values_with_op(val1, op, val2) - - %{TAG: "ModalityDrift", _0: mod1, _1: mod2, _2: threshold} -> - drift = compute_modality_drift(octad, mod1, mod2) - drift > threshold - - %{TAG: "ModalityExists", _0: modality} -> - has_modality_data?(octad, modality) - - %{TAG: "ModalityNotExists", _0: modality} -> - not has_modality_data?(octad, modality) - - %{TAG: "ModalityConsistency", _0: mod1, _1: mod2, _2: metric} -> - compute_consistency(octad, mod1, mod2, metric) > 0.0 - - _ -> true - end - end - - defp get_modality_field(octad, modality, field) do - mod_str = modality_to_string(modality) - mod_data = Map.get(octad, mod_str, %{}) - Map.get(mod_data, field) || Map.get(octad, "#{mod_str}.#{field}") - end - - defp has_modality_data?(octad, modality) do - mod_str = modality_to_string(modality) - case Map.get(octad, mod_str) do - nil -> false - data when data == %{} -> false - _ -> true - end - end - - defp compute_modality_drift(octad, mod1, mod2) do - # Compute drift between two modality representations. - # Uses the Rust drift API when both modalities have data, - # falling back to embedding-based comparison. - mod1_str = modality_to_string(mod1) - mod2_str = modality_to_string(mod2) - - case {Map.get(octad, mod1_str), Map.get(octad, mod2_str)} do - {nil, _} -> 1.0 # Missing modality = maximum drift - {_, nil} -> 1.0 - - {data1, data2} -> - # Try to get drift from the Rust drift detector via octad ID - octad_id = Map.get(octad, "id") || Map.get(octad, :id) - - case octad_id && RustClient.get_drift_score(octad_id) do - {:ok, score} when is_number(score) -> - score - - _ -> - # Fallback: compute local drift from modality data - vec1 = extract_embedding_from_modality(data1) - vec2 = extract_embedding_from_modality(data2) - compute_cosine_distance(vec1, vec2) - end - end - end - - defp extract_embedding_from_modality(data) when is_map(data) do - # Extract a numeric vector from modality data for comparison. - # Vector modality stores embeddings directly; others use hash fingerprints. - cond do - is_list(data["embedding"]) -> data["embedding"] - is_list(data["vector"]) -> data["vector"] - is_binary(data["content"]) -> content_fingerprint(data["content"]) - true -> [] - end - end - defp extract_embedding_from_modality(data) when is_list(data), do: data - defp extract_embedding_from_modality(_), do: [] - - defp content_fingerprint(text) when is_binary(text) do - # Simple hash-based fingerprint for non-vector modalities. - # Produces a 4-element vector from character distribution. - bytes = :binary.bin_to_list(text) - len = max(length(bytes), 1) - quartiles = Enum.chunk_every(bytes, max(div(len, 4), 1)) - - Enum.map(quartiles |> Enum.take(4), fn chunk -> - Enum.sum(chunk) / max(length(chunk), 1) / 255.0 - end) - end - - defp compute_cosine_distance([], _), do: 0.5 # Unknown = moderate drift - defp compute_cosine_distance(_, []), do: 0.5 - defp compute_cosine_distance(vec1, vec2) do - # Cosine distance: 1 - cosine_similarity. Range: [0.0, 2.0], normalized to [0.0, 1.0]. - {dot, mag1, mag2} = Enum.zip(vec1, vec2) - |> Enum.reduce({0.0, 0.0, 0.0}, fn {a, b}, {d, m1, m2} -> - {d + a * b, m1 + a * a, m2 + b * b} - end) - - denom = :math.sqrt(mag1) * :math.sqrt(mag2) - - if denom > 0.0 do - similarity = dot / denom - # Clamp and normalize to [0.0, 1.0] - min(max(1.0 - similarity, 0.0), 1.0) - else - 1.0 # Zero vectors = max drift - end - end - - defp compute_consistency(octad, mod1, mod2, metric) do - # Compute consistency score between two modalities using the specified metric. - # Returns a score in [0.0, 1.0] where 1.0 = perfectly consistent. - mod1_str = modality_to_string(mod1) - mod2_str = modality_to_string(mod2) - - data1 = Map.get(octad, mod1_str) - data2 = Map.get(octad, mod2_str) - - case {data1, data2} do - {nil, _} -> 0.0 - {_, nil} -> 0.0 - - {d1, d2} -> - vec1 = extract_embedding_from_modality(d1) - vec2 = extract_embedding_from_modality(d2) - - case metric do - m when m in ["COSINE", :cosine, %{TAG: "Cosine"}] -> - cosine_similarity(vec1, vec2) - - m when m in ["EUCLIDEAN", :euclidean, %{TAG: "Euclidean"}] -> - euclidean_similarity(vec1, vec2) - - m when m in ["DOT_PRODUCT", :dot_product, %{TAG: "DotProduct"}] -> - dot_product_similarity(vec1, vec2) - - m when m in ["JACCARD", :jaccard, %{TAG: "Jaccard"}] -> - jaccard_similarity(d1, d2) - - _ -> - cosine_similarity(vec1, vec2) # Default to cosine - end - end - end - - defp cosine_similarity([], _), do: 0.0 - defp cosine_similarity(_, []), do: 0.0 - defp cosine_similarity(vec1, vec2) do - {dot, mag1, mag2} = Enum.zip(vec1, vec2) - |> Enum.reduce({0.0, 0.0, 0.0}, fn {a, b}, {d, m1, m2} -> - {d + a * b, m1 + a * a, m2 + b * b} - end) - - denom = :math.sqrt(mag1) * :math.sqrt(mag2) - if denom > 0.0, do: max(dot / denom, 0.0), else: 0.0 - end - - defp euclidean_similarity([], _), do: 0.0 - defp euclidean_similarity(_, []), do: 0.0 - defp euclidean_similarity(vec1, vec2) do - dist = Enum.zip(vec1, vec2) - |> Enum.reduce(0.0, fn {a, b}, acc -> acc + (a - b) * (a - b) end) - |> :math.sqrt() - - # Convert distance to similarity: 1 / (1 + distance) - 1.0 / (1.0 + dist) - end - - defp dot_product_similarity([], _), do: 0.0 - defp dot_product_similarity(_, []), do: 0.0 - defp dot_product_similarity(vec1, vec2) do - dot = Enum.zip(vec1, vec2) |> Enum.reduce(0.0, fn {a, b}, acc -> acc + a * b end) - # Normalize to [0.0, 1.0] using sigmoid - 1.0 / (1.0 + :math.exp(-dot)) - end - - defp jaccard_similarity(d1, d2) when is_map(d1) and is_map(d2) do - # Jaccard: |intersection| / |union| of map keys - keys1 = MapSet.new(Map.keys(d1)) - keys2 = MapSet.new(Map.keys(d2)) - intersection = MapSet.intersection(keys1, keys2) |> MapSet.size() - union = MapSet.union(keys1, keys2) |> MapSet.size() - if union > 0, do: intersection / union, else: 0.0 - end - defp jaccard_similarity(_, _), do: 0.0 - - defp compare_values_with_op(val1, op, val2) when is_number(val1) and is_number(val2) do - case op do - "==" -> val1 == val2 - "!=" -> val1 != val2 - ">" -> val1 > val2 - "<" -> val1 < val2 - ">=" -> val1 >= val2 - "<=" -> val1 <= val2 - %{TAG: "Eq"} -> val1 == val2 - %{TAG: "Neq"} -> val1 != val2 - %{TAG: "Gt"} -> val1 > val2 - %{TAG: "Lt"} -> val1 < val2 - %{TAG: "Gte"} -> val1 >= val2 - %{TAG: "Lte"} -> val1 <= val2 - _ -> false - end - end - defp compare_values_with_op(val1, op, val2) when is_binary(val1) and is_binary(val2) do - case op do - "==" -> val1 == val2 - "!=" -> val1 != val2 - %{TAG: "Eq"} -> val1 == val2 - %{TAG: "Neq"} -> val1 != val2 - _ -> false - end - end - defp compare_values_with_op(_val1, _op, _val2), do: false - - defp modality_to_string(mod) when is_binary(mod), do: String.downcase(mod) - defp modality_to_string(mod) when is_atom(mod), do: Atom.to_string(mod) - defp modality_to_string(%{TAG: tag}), do: String.downcase(tag) - defp modality_to_string(_), do: "unknown" - - # =========================================================================== - # Phase 3: Mutation Execution - # =========================================================================== - - defp execute_insert(modality_data, proof, _timeout) do - proof_result = if proof, do: verify_multi_proof(nil, proof), else: {:ok, []} - - case proof_result do - {:error, reason} -> - {:error, {:write_proof_failed, reason}} - - {:ok, _artifacts} -> - case RustClient.create_octad(modality_data) do - {:ok, octad_id} -> {:ok, %{octad_id: octad_id, operation: :insert}} - {:error, reason} -> {:error, {:insert_failed, reason}} - end - end - rescue - e -> {:error, {:insert_failed, Exception.message(e)}} - end - - defp execute_update(octad_id, sets, proof, _timeout) do - proof_result = if proof, do: verify_multi_proof(nil, proof), else: {:ok, []} - - case proof_result do - {:error, reason} -> - {:error, {:write_proof_failed, reason}} - - {:ok, _artifacts} -> - field_updates = Enum.map(sets, fn {field_ref, value} -> - {field_ref, value} - end) - - case RustClient.update_octad(octad_id, field_updates) do - {:ok, _} -> {:ok, %{octad_id: octad_id, operation: :update, fields_updated: length(sets)}} - {:error, reason} -> {:error, {:update_failed, reason}} - end - end - rescue - e -> {:error, {:update_failed, Exception.message(e)}} - end - - defp execute_delete(octad_id, proof, _timeout) do - proof_result = if proof, do: verify_multi_proof(nil, proof), else: {:ok, []} - - case proof_result do - {:error, reason} -> - {:error, {:write_proof_failed, reason}} - - {:ok, _artifacts} -> - case RustClient.delete_octad(octad_id) do - {:ok, _} -> {:ok, %{octad_id: octad_id, operation: :delete}} - {:error, reason} -> {:error, {:delete_failed, reason}} - end - end - rescue - e -> {:error, {:delete_failed, Exception.message(e)}} - end - - # =========================================================================== - # Multi-Proof Verification - # =========================================================================== - - defp verify_multi_proof(_query_ast, proof_specs) when is_list(proof_specs) do - # Verify each proof in the composition, collecting proof artifacts. - # Each verify_single_proof/1 returns {:ok, artifact} or {:error, reason}. - results = Enum.map(proof_specs, fn spec -> - verify_single_proof(spec) - end) - - case Enum.find(results, &match?({:error, _}, &1)) do - nil -> - # All proofs passed — collect the artifacts for the ProvedResult certificate - artifacts = Enum.map(results, fn {:ok, artifact} -> artifact end) - {:ok, artifacts} - - error -> - error - end - end - defp verify_multi_proof(_query_ast, nil), do: {:ok, []} - defp verify_multi_proof(_query_ast, proof_specs) do - # Non-list proof specs are a safety violation — they must not silently pass. - # This catch-all previously returned :ok, which meant malformed proof specs - # (e.g., a bare map or atom) would bypass verification entirely. - {:error, {:invalid_proof_specs, "Expected a list of proof specifications, got: #{inspect(proof_specs)}"}} - end - - defp verify_single_proof(proof_spec) do - # Verify a single proof obligation against the VeriSimDB contract registry. - # Validates proof type, contract existence, and parameter compatibility. - # Full ZKP witness generation requires the verisim-semantic crate. - proof_type = extract_proof_type(proof_spec) - contract_name = extract_contract_name(proof_spec) - - case proof_type do - :existence -> - # Existence proofs verify the octad exists and is accessible. - entity_id = contract_name || extract_entity_from_proof(proof_spec) - - if entity_id do - case RustClient.get_octad(entity_id) do - {:ok, octad} -> - {:ok, %{type: :existence, entity_id: entity_id, verified: true, - status: Map.get(octad, "status", %{})}} - {:error, :not_found} -> - {:error, {:existence_failed, "Entity '#{entity_id}' does not exist"}} - {:error, reason} -> - {:error, {:existence_check_failed, reason}} - end - else - {:error, {:missing_entity, "Existence proof requires an entity ID or contract name"}} - end - - :citation -> - # Citation proofs verify the citation chain is valid - if contract_name do - case validate_contract_exists(contract_name) do - {:ok, artifact} -> {:ok, artifact} - {:error, _} = err -> err - end - else - {:error, {:missing_contract, "Citation proof requires a contract name"}} - end - - :access -> - # Access proofs verify the user has rights via the Rust RBAC module. - entity_id = contract_name || extract_entity_from_proof(proof_spec) - - if entity_id do - case RustClient.get("/auth/check/#{entity_id}") do - {:ok, %{status: 200, body: %{"authorized" => true}}} -> - {:ok, %{type: :access, entity_id: entity_id, authorized: true}} - {:ok, %{status: 200, body: %{"authorized" => false}}} -> - {:error, {:access_denied, "Not authorized to access '#{entity_id}'"}} - {:ok, %{status: 403}} -> - {:error, {:access_denied, "Not authorized to access '#{entity_id}'"}} - {:error, reason} -> - {:error, {:access_check_failed, reason}} - end - else - # No specific entity — log a warning but allow for global queries. - Logger.warning("VCL-UT: Access proof without entity ID — global query, skipping entity-level check") - {:ok, %{type: :access, entity_id: nil, authorized: true, scope: :global}} - end - - :integrity -> - # Integrity proofs verify data has not been tampered with. - # Route to the Rust proofs/generate endpoint for Merkle proof generation. - if contract_name do - case RustClient.post("/proofs/generate", %{ - type: "integrity", - contract: contract_name, - privacy_level: "public" - }) do - {:ok, %{status: 200, body: %{"success" => true} = body}} -> - {:ok, %{type: :integrity, contract: contract_name, verified: true, - proof_data: Map.get(body, "proof")}} - {:ok, %{status: 200, body: %{"success" => false, "error" => reason}}} -> - {:error, {:integrity_failed, reason}} - {:ok, %{status: 200, body: %{"error" => reason}}} -> - {:error, {:integrity_failed, reason}} - {:error, reason} -> - {:error, {:integrity_check_failed, reason}} - end - else - {:error, {:missing_contract, "Integrity proof requires a contract name"}} - end - - :provenance -> - # Provenance proofs verify lineage is verifiable via the provenance store. - # An entity/contract reference is REQUIRED. - entity_id = contract_name || extract_entity_from_proof(proof_spec) - - if entity_id do - case RustClient.verify_provenance(entity_id) do - {:ok, %{"has_provenance" => true, "chain_valid" => true} = body} -> - {:ok, %{type: :provenance, entity_id: entity_id, chain_valid: true, - chain_length: Map.get(body, "chain_length", 0)}} - - {:ok, %{"has_provenance" => true, "chain_valid" => false}} -> - {:error, {:provenance_chain_broken, entity_id}} - - {:ok, %{"has_provenance" => false}} -> - {:error, {:no_provenance, "Entity '#{entity_id}' has no provenance chain"}} - - {:error, :not_found} -> - {:error, {:entity_not_found, entity_id}} - - {:error, reason} -> - {:error, {:provenance_verify_failed, reason}} - end - else - {:error, {:missing_entity, "Provenance proof requires an entity ID or contract name"}} - end - - :consistency -> - # Consistency proofs verify that two or more modalities are in agreement. - # Uses the Rust drift API to get the cross-modal drift score and compares - # against the threshold (default: 0.3). Requires an entity ID. - entity_id = contract_name || extract_entity_from_proof(proof_spec) - - if entity_id do - case RustClient.get_drift_score(entity_id) do - {:ok, score} when is_number(score) -> - threshold = Map.get(proof_spec, :threshold, 0.3) - if score <= threshold do - {:ok, %{type: :consistency, entity_id: entity_id, drift_score: score, - threshold: threshold, consistent: true}} - else - {:error, {:consistency_failed, - "Entity '#{entity_id}' drift score #{score} exceeds threshold #{threshold}"}} - end - - {:error, reason} -> - {:error, {:consistency_check_failed, reason}} - end - else - {:error, {:missing_entity, "Consistency proof requires an entity ID"}} - end - - :freshness -> - # Freshness proofs verify that entity data is recent enough. - # Checks the temporal modality's last-modified timestamp against - # a maximum age (default: 1 hour = 3_600_000ms). - entity_id = contract_name || extract_entity_from_proof(proof_spec) - - if entity_id do - case RustClient.get_octad(entity_id) do - {:ok, octad} -> - max_age_ms = Map.get(proof_spec, :max_age_ms, 3_600_000) - temporal = Map.get(octad, "temporal", %{}) - last_modified = Map.get(temporal, "last_modified") || - Map.get(temporal, "updated_at") || - Map.get(octad, "updated_at") - - if last_modified do - age_ms = case DateTime.from_iso8601(to_string(last_modified)) do - {:ok, dt, _} -> - DateTime.diff(DateTime.utc_now(), dt, :millisecond) - _ -> - # If timestamp is a unix epoch, convert - if is_number(last_modified) do - now_ms = System.system_time(:millisecond) - now_ms - trunc(last_modified) - else - max_age_ms + 1 # Unknown format — treat as stale - end - end - - if age_ms <= max_age_ms do - {:ok, %{type: :freshness, entity_id: entity_id, age_ms: age_ms, - max_age_ms: max_age_ms, fresh: true}} - else - {:error, {:freshness_expired, - "Entity '#{entity_id}' is #{age_ms}ms old, exceeds max age #{max_age_ms}ms"}} - end - else - {:error, {:no_temporal_data, - "Entity '#{entity_id}' has no temporal/timestamp data for freshness check"}} - end - - {:error, :not_found} -> - {:error, {:entity_not_found, entity_id}} - - {:error, reason} -> - {:error, {:freshness_check_failed, reason}} - end - else - {:error, {:missing_entity, "Freshness proof requires an entity ID"}} - end - - :custom -> - # Custom ZKP proofs route to the circuit-specific proof generation endpoint - if contract_name do - claim = Map.get(proof_spec, :claim, "") - - case RustClient.post("/proofs/generate-with-circuit", %{ - circuit: contract_name, - claim: claim, - privacy_level: Map.get(proof_spec, :privacy_level, "public") - }) do - {:ok, %{status: 200, body: %{"success" => true} = body}} -> - {:ok, %{type: :custom, circuit: contract_name, verified: true, - proof_data: Map.get(body, "proof")}} - {:ok, %{status: 200, body: %{"error" => reason}}} -> - {:error, {:custom_proof_failed, reason}} - {:error, reason} -> - {:error, {:custom_proof_failed, reason}} - end - else - {:error, {:missing_contract, "Custom proof requires a circuit/contract name"}} - end - - :zkp -> - # ZKP proofs route to the zkp_bridge via the Rust API. - # When the type checker provides witness_fields and circuit, include them - # for richer proof generation (privacy-aware with blinding). - privacy_level = Map.get(proof_spec, :privacy_level, "public") - claim = Map.get(proof_spec, :claim, "") - witness = Map.get(proof_spec, :witness_fields, []) - circuit = Map.get(proof_spec, :circuit) - - request = %{claim: claim, privacy_level: privacy_level} - request = if witness != [], do: Map.put(request, :witness, witness), else: request - request = if circuit, do: Map.put(request, :circuit_name, circuit), else: request - - case RustClient.post("/proofs/generate", request) do - {:ok, %{status: 200, body: %{"success" => true} = body}} -> - {:ok, %{type: :zkp, verified: true, privacy_level: privacy_level, - proof_data: Map.get(body, "proof")}} - {:ok, %{status: 200, body: %{"error" => reason}}} -> - {:error, {:zkp_failed, reason}} - {:error, reason} -> - {:error, {:zkp_unavailable, reason}} - end - - :proven -> - # Proven proofs verify against certificates from the proven library - claim = Map.get(proof_spec, :claim, "") - case RustClient.post("/proofs/generate", %{claim: claim, privacy_level: "public"}) do - {:ok, %{status: 200, body: %{"success" => true} = body}} -> - {:ok, %{type: :proven, verified: true, - proof_data: Map.get(body, "proof")}} - _ -> - {:error, {:proven_unavailable, "proven certificate verification failed"}} - end - - :sanctify -> - # Sanctify proofs verify security contracts from sanctify-php - if contract_name do - case validate_contract_exists(contract_name) do - {:ok, artifact} -> {:ok, artifact} - {:error, _} = err -> err - end - else - {:error, {:missing_contract, "Sanctify proof requires a contract name"}} - end - - _ -> - {:error, {:unknown_proof_type, proof_type}} - end - end - - defp extract_proof_type(%{proofType: type}), do: normalize_proof_type(type) - defp extract_proof_type(%{TAG: tag}), do: normalize_proof_type(tag) - defp extract_proof_type(%{raw: raw}) when is_binary(raw) do - # Raw proof strings look like "EXISTENCE(entity-001)" or "EXISTENCE entity-001". - # Extract just the proof type name (before any parens or whitespace). - raw - |> String.split(~r/[\s(]/, parts: 2) - |> List.first() - |> normalize_proof_type() - end - defp extract_proof_type(_), do: :unknown - - defp normalize_proof_type("EXISTENCE"), do: :existence - defp normalize_proof_type("CITATION"), do: :citation - defp normalize_proof_type("ACCESS"), do: :access - defp normalize_proof_type("INTEGRITY"), do: :integrity - defp normalize_proof_type("CONSISTENCY"), do: :consistency - defp normalize_proof_type("PROVENANCE"), do: :provenance - defp normalize_proof_type("FRESHNESS"), do: :freshness - defp normalize_proof_type("CUSTOM"), do: :custom - defp normalize_proof_type("ZKP"), do: :zkp - defp normalize_proof_type("PROVEN"), do: :proven - defp normalize_proof_type("SANCTIFY"), do: :sanctify - defp normalize_proof_type(%{TAG: tag}), do: normalize_proof_type(tag) - defp normalize_proof_type(atom) when is_atom(atom), do: atom - defp normalize_proof_type(str) when is_binary(str) do - try do - String.downcase(str) |> String.to_existing_atom() - rescue - ArgumentError -> :unknown - end - end - defp normalize_proof_type(_), do: :unknown - - defp extract_contract_name(%{contractName: name}), do: name - defp extract_contract_name(%{contract: name}), do: name - defp extract_contract_name(%{raw: raw}) when is_binary(raw) do - # Extract contract name from raw proof spec: "INTEGRITY(my_contract)" - case Regex.run(~r/\(([^)]+)\)/, raw) do - [_, name] -> String.trim(name) - _ -> nil - end - end - defp extract_contract_name(_), do: nil - - defp validate_contract_exists(contract_name) do - # Check if the contract exists in the semantic store via search. - # If the store is unreachable, the proof MUST fail — silently passing - # would defeat the purpose of verified queries. - case RustClient.search_text("contract:#{contract_name}", 1) do - {:ok, results} when is_list(results) and length(results) > 0 -> - {:ok, %{type: :contract_verified, contract: contract_name, verified: true}} - {:ok, %{"results" => [_ | _]}} -> - {:ok, %{type: :contract_verified, contract: contract_name, verified: true}} - {:ok, _} -> - {:error, {:contract_not_found, contract_name}} - {:error, reason} -> - {:error, {:contract_verification_failed, reason}} - end - end - - defp extract_entity_from_proof(%{entity_id: id}), do: id - defp extract_entity_from_proof(%{entityId: id}), do: id - defp extract_entity_from_proof(%{octad_id: id}), do: id - defp extract_entity_from_proof(%{octadId: id}), do: id - defp extract_entity_from_proof(_), do: nil - - # --------------------------------------------------------------------------- - # Proof obligation structuring (fallback when type checker unavailable) - # --------------------------------------------------------------------------- - - # Proof obligation structuring and composition determination are now handled - # by VeriSim.Query.VCLTypeChecker, which provides richer validation and - # enrichment (witness fields, circuit names, time estimates). - - # =========================================================================== - # Query execution by source type - # =========================================================================== - - defp execute_octad_query(entity_id, modalities, where_clause, limit, offset, _timeout) do - case RustClient.get_octad(entity_id) do - {:ok, octad} -> - filtered = filter_octad(octad, modalities, where_clause) - paginated = paginate_results([filtered], limit, offset) - {:ok, paginated} - - {:error, reason} -> - {:error, reason} - end - end - - defp execute_federation_query(pattern, drift_policy, modalities, where_clause, limit, offset, timeout) do - Logger.info("Federation query: pattern=#{inspect(pattern)}, drift=#{inspect(drift_policy)}") - - # Delegate to Rust federation API which handles peer discovery and fan-out - federation_params = %{ - pattern: pattern, - drift_policy: drift_policy_to_string(drift_policy), - modalities: Enum.map(modalities, &to_string/1), - limit: limit || 100, - offset: offset || 0, - timeout_ms: timeout - } - - # Add query parameters if WHERE clause exists - federation_params = if where_clause do - text = case where_clause do - %{raw: raw} -> raw - _ -> nil - end - if text, do: Map.put(federation_params, :text_query, text), else: federation_params - else - federation_params - end - - case RustClient.post("/federation/query", federation_params) do - {:ok, %{status: 200, body: body}} when is_list(body) -> - {:ok, body} - - {:ok, %{status: 200, body: %{"results" => results}}} when is_list(results) -> - {:ok, results} - - {:ok, %{status: 200, body: body}} when is_map(body) -> - {:ok, Map.get(body, "results", [])} - - {:ok, %{status: status, body: body}} -> - Logger.warning("Federation query returned status #{status}: #{inspect(body)}") - {:error, {:federation_error, status, body}} - - {:error, reason} -> - Logger.warning("Federation query failed: #{inspect(reason)}") - # Graceful degradation: return empty results instead of crashing - {:ok, []} - end - end - - defp drift_policy_to_string(:strict), do: "strict" - defp drift_policy_to_string(:repair), do: "repair" - defp drift_policy_to_string(:tolerate), do: "tolerate" - defp drift_policy_to_string(:latest), do: "latest" - defp drift_policy_to_string(nil), do: "tolerate" - defp drift_policy_to_string(s) when is_binary(s), do: s - - defp execute_store_query(store_id, modalities, where_clause, limit, _offset, _timeout) do - Logger.info("Store query: store_id=#{store_id}") - - query_type = determine_query_type(modalities, where_clause) - - case query_type do - :text -> - text_query = extract_text_query(where_clause) - QueryRouter.query(:text, text_query, limit: limit || 10) - - :vector -> - {embedding, _threshold} = extract_vector_query(where_clause) - QueryRouter.query(:vector, embedding, k: limit || 10) - - :graph -> - graph_params = extract_graph_query(where_clause) - QueryRouter.query(:graph, graph_params) - - :provenance -> - provenance_params = extract_provenance_query(where_clause) - execute_provenance_query(provenance_params, limit) - - :spatial -> - spatial_params = extract_spatial_query(where_clause) - execute_spatial_query(spatial_params, limit) - - :multi -> - params = extract_multi_modal_params(modalities, where_clause) - QueryRouter.query(:multi, params, limit: limit || 10) - - _ -> - {:error, :unsupported_query_type} - end - end - - defp execute_reflect_query(modalities, where_clause, limit, _offset, _timeout) do - # REFLECT queries the query store itself — meta-circular homoiconicity. - # Stored queries are octads, so we search them via the /queries/similar - # and /search/text endpoints filtered to type=vcl_query. - Logger.info("REFLECT query: querying the query store") - - text_query = extract_text_query(where_clause) - search_limit = limit || 20 - - # Search stored queries by text similarity - result = if text_query != "" do - case RustClient.post("/search/text", %{q: "type:vcl_query #{text_query}", limit: search_limit}) do - {:ok, %{status: 200, body: %{"results" => results}}} -> {:ok, results} - {:ok, %{status: 200, body: body}} when is_list(body) -> {:ok, body} - {:ok, %{status: 200, body: body}} when is_map(body) -> {:ok, Map.get(body, "results", [])} - {:error, reason} -> {:error, {:reflect_query_failed, reason}} - _ -> {:ok, []} - end - else - # No text filter — list all stored queries - case RustClient.post("/search/text", %{q: "type:vcl_query", limit: search_limit}) do - {:ok, %{status: 200, body: %{"results" => results}}} -> {:ok, results} - {:ok, %{status: 200, body: body}} when is_list(body) -> {:ok, body} - {:ok, %{status: 200, body: body}} when is_map(body) -> {:ok, Map.get(body, "results", [])} - {:error, reason} -> {:error, {:reflect_query_failed, reason}} - _ -> {:ok, []} - end - end - - # Enrich results with query-specific metadata - case result do - {:ok, rows} -> - enriched = Enum.map(rows, fn row -> - row - |> Map.put("_source", "reflect") - |> Map.put("_type", "stored_query") - |> maybe_filter_modalities(modalities) - end) - {:ok, enriched} - - error -> error - end - end - - defp maybe_filter_modalities(row, [:all]), do: row - defp maybe_filter_modalities(row, modalities) do - mod_strings = Enum.map(modalities, &to_string/1) - base_keys = ["id", "_source", "_type", "status"] - keep_keys = base_keys ++ mod_strings - Map.take(row, keep_keys ++ Map.keys(row) |> Enum.filter(fn k -> - Enum.any?(mod_strings, &String.starts_with?(k, &1)) - end)) - end - - defp filter_octad(octad, modalities, _where_clause) do - if :all in modalities do - octad - else - Map.take(octad, modalities |> Enum.map(&to_string/1)) - end - end - - defp paginate_results(results, nil, nil), do: results - defp paginate_results(results, limit, nil), do: Enum.take(results, limit) - defp paginate_results(results, nil, offset), do: Enum.drop(results, offset) - defp paginate_results(results, limit, offset) do - results - |> Enum.drop(offset) - |> Enum.take(limit) - end - - defp determine_query_type(modalities, where_clause) do - cond do - has_fulltext_condition?(where_clause) -> :text - has_vector_condition?(where_clause) -> :vector - has_graph_pattern?(where_clause) -> :graph - has_provenance_condition?(where_clause) -> :provenance - has_spatial_condition?(where_clause) -> :spatial - length(modalities) > 1 -> :multi - true -> :multi - end - end - - # AST-walking condition detectors: check TAG values for modality-specific conditions - - defp has_fulltext_condition?(nil), do: false - defp has_fulltext_condition?(%{raw: raw}) when is_binary(raw) do - upper = String.upcase(raw) - String.contains?(upper, "FULLTEXT") or String.contains?(upper, "CONTAINS") or - String.contains?(upper, "MATCHES") - end - defp has_fulltext_condition?(%{TAG: tag}) when tag in ["FulltextContains", "FulltextMatches", "DocumentCondition"], do: true - defp has_fulltext_condition?(%{TAG: "And", _0: left, _1: right}), do: has_fulltext_condition?(left) or has_fulltext_condition?(right) - defp has_fulltext_condition?(%{TAG: "Or", _0: left, _1: right}), do: has_fulltext_condition?(left) or has_fulltext_condition?(right) - defp has_fulltext_condition?(%{TAG: "Not", _0: inner}), do: has_fulltext_condition?(inner) - defp has_fulltext_condition?(_), do: false - - defp has_vector_condition?(nil), do: false - defp has_vector_condition?(%{raw: raw}) when is_binary(raw) do - upper = String.upcase(raw) - String.contains?(upper, "SIMILAR") or String.contains?(upper, "NEAREST") - end - defp has_vector_condition?(%{TAG: tag}) when tag in ["VectorSimilar", "VectorNearest", "VectorCondition"], do: true - defp has_vector_condition?(%{TAG: "And", _0: left, _1: right}), do: has_vector_condition?(left) or has_vector_condition?(right) - defp has_vector_condition?(%{TAG: "Or", _0: left, _1: right}), do: has_vector_condition?(left) or has_vector_condition?(right) - defp has_vector_condition?(%{TAG: "Not", _0: inner}), do: has_vector_condition?(inner) - defp has_vector_condition?(_), do: false - - defp has_graph_pattern?(nil), do: false - defp has_graph_pattern?(%{raw: raw}) when is_binary(raw) do - # Graph patterns use SPARQL-like syntax with arrow edges - String.contains?(raw, "->") or String.contains?(raw, "-[") - end - defp has_graph_pattern?(%{TAG: tag}) when tag in ["SparqlPattern", "PathPattern", "GraphCondition"], do: true - defp has_graph_pattern?(%{TAG: "And", _0: left, _1: right}), do: has_graph_pattern?(left) or has_graph_pattern?(right) - defp has_graph_pattern?(%{TAG: "Or", _0: left, _1: right}), do: has_graph_pattern?(left) or has_graph_pattern?(right) - defp has_graph_pattern?(%{TAG: "Not", _0: inner}), do: has_graph_pattern?(inner) - defp has_graph_pattern?(_), do: false - - # Provenance condition detection: actor, origin, chain_valid, event_type queries - defp has_provenance_condition?(nil), do: false - defp has_provenance_condition?(%{raw: raw}) when is_binary(raw) do - upper = String.upcase(raw) - String.contains?(upper, "PROVENANCE.") or String.contains?(upper, "CHAIN_VALID") or - String.contains?(upper, "CHAIN_LENGTH") - end - defp has_provenance_condition?(%{TAG: tag}) when tag in [ - "ProvenanceActor", "ProvenanceOrigin", "ProvenanceChainValid", - "ProvenanceEventType", "ProvenanceCondition" - ], do: true - defp has_provenance_condition?(%{TAG: "And", _0: left, _1: right}), do: has_provenance_condition?(left) or has_provenance_condition?(right) - defp has_provenance_condition?(%{TAG: "Or", _0: left, _1: right}), do: has_provenance_condition?(left) or has_provenance_condition?(right) - defp has_provenance_condition?(%{TAG: "Not", _0: inner}), do: has_provenance_condition?(inner) - defp has_provenance_condition?(%{modality: mod}) when mod in [:provenance, "PROVENANCE", "provenance"], do: true - defp has_provenance_condition?(_), do: false - - # Spatial condition detection: radius, bounding box, nearest queries - defp has_spatial_condition?(nil), do: false - defp has_spatial_condition?(%{raw: raw}) when is_binary(raw) do - upper = String.upcase(raw) - String.contains?(upper, "WITHIN RADIUS") or String.contains?(upper, "WITHIN BOUNDS") or - String.contains?(upper, "SPATIAL.") or String.contains?(upper, "NEAREST") - end - defp has_spatial_condition?(%{TAG: tag}) when tag in [ - "SpatialRadius", "SpatialBounds", "SpatialNearest", - "SpatialCondition", "WithinRadius", "WithinBounds" - ], do: true - defp has_spatial_condition?(%{TAG: "And", _0: left, _1: right}), do: has_spatial_condition?(left) or has_spatial_condition?(right) - defp has_spatial_condition?(%{TAG: "Or", _0: left, _1: right}), do: has_spatial_condition?(left) or has_spatial_condition?(right) - defp has_spatial_condition?(%{TAG: "Not", _0: inner}), do: has_spatial_condition?(inner) - defp has_spatial_condition?(%{modality: mod}) when mod in [:spatial, "SPATIAL", "spatial"], do: true - defp has_spatial_condition?(_), do: false - - # AST-walking query extractors: pull actual values from parsed conditions - - defp extract_text_query(nil), do: "" - defp extract_text_query(%{raw: raw}) when is_binary(raw) do - # Extract text between quotes from raw WHERE clause: FULLTEXT CONTAINS 'search terms' - case Regex.run(~r/'([^']*)'/, raw) do - [_, text] -> text - _ -> - # Try without quotes: FULLTEXT CONTAINS keyword - case Regex.run(~r/(?:CONTAINS|MATCHES)\s+(.+?)(?:\s+AND|\s+OR|\s*$)/i, raw) do - [_, text] -> String.trim(text) - _ -> raw # Use the raw clause as search text - end - end - end - defp extract_text_query(%{TAG: "FulltextContains", _0: text}), do: text - defp extract_text_query(%{TAG: "FulltextMatches", _0: pattern}), do: pattern - defp extract_text_query(%{TAG: "And", _0: left, _1: right}) do - case {has_fulltext_condition?(left), has_fulltext_condition?(right)} do - {true, _} -> extract_text_query(left) - {_, true} -> extract_text_query(right) - _ -> "" - end - end - defp extract_text_query(_), do: "" - - defp extract_vector_query(nil), do: {[], 0.9} - defp extract_vector_query(%{raw: raw}) when is_binary(raw) do - # Extract vector literal [0.1, 0.2, ...] and optional WITHIN threshold - vector = case Regex.run(~r/\[([0-9.,\s-]+)\]/, raw) do - [_, nums] -> - nums - |> String.split(",") - |> Enum.map(&String.trim/1) - |> Enum.flat_map(fn s -> - case Float.parse(s) do - {f, _} -> [f] - :error -> [] - end - end) - _ -> [] - end - - threshold = case Regex.run(~r/WITHIN\s+([0-9.]+)/i, raw) do - [_, t] -> - case Float.parse(t) do - {f, _} -> f - :error -> 0.9 - end - _ -> 0.9 - end - - {vector, threshold} - end - defp extract_vector_query(%{TAG: "VectorSimilar", _0: _field, _1: vector, _2: threshold}) do - {vector, threshold || 0.9} - end - defp extract_vector_query(%{TAG: "VectorNearest", _0: _field, _1: k}) do - {[], k} # k-nearest doesn't have a vector, uses the field's own embedding - end - defp extract_vector_query(%{TAG: "And", _0: left, _1: right}) do - case {has_vector_condition?(left), has_vector_condition?(right)} do - {true, _} -> extract_vector_query(left) - {_, true} -> extract_vector_query(right) - _ -> {[], 0.9} - end - end - defp extract_vector_query(_), do: {[], 0.9} - - defp extract_graph_query(nil), do: %{} - defp extract_graph_query(%{raw: raw}) when is_binary(raw) do - # Parse simple SPARQL-like patterns from raw WHERE clause - %{raw_pattern: raw} - end - defp extract_graph_query(%{TAG: "SparqlPattern", _0: node1, _1: edge, _2: node2}) do - %{subject: node1, predicate: edge, object: node2} - end - defp extract_graph_query(%{TAG: "PathPattern", _0: start, _1: edge, _2: finish}) do - %{start: start, edge: edge, finish: finish, traversal: true} - end - defp extract_graph_query(%{TAG: "And", _0: left, _1: right}) do - case {has_graph_pattern?(left), has_graph_pattern?(right)} do - {true, _} -> extract_graph_query(left) - {_, true} -> extract_graph_query(right) - _ -> %{} - end - end - defp extract_graph_query(_), do: %{} - - # --------------------------------------------------------------------------- - # Provenance query extraction and execution - # --------------------------------------------------------------------------- - - defp extract_provenance_query(nil), do: %{} - defp extract_provenance_query(%{raw: raw}) when is_binary(raw) do - params = %{} - - # Extract actor filter: PROVENANCE.actor = 'someone' - params = case Regex.run(~r/(?:PROVENANCE\.)?actor\s*=\s*'([^']+)'/i, raw) do - [_, actor] -> Map.put(params, :actor, actor) - _ -> params - end - - # Extract origin filter: PROVENANCE.origin = 'source' - params = case Regex.run(~r/(?:PROVENANCE\.)?origin\s*=\s*'([^']+)'/i, raw) do - [_, origin] -> Map.put(params, :origin, origin) - _ -> params - end - - # Extract chain_valid filter: PROVENANCE.chain_valid = true - params = case Regex.run(~r/(?:PROVENANCE\.)?chain_valid\s*=\s*(true|false)/i, raw) do - [_, val] -> Map.put(params, :chain_valid, String.downcase(val) == "true") - _ -> params - end - - # Extract event_type filter: PROVENANCE.event_type = 'Modified' - params = case Regex.run(~r/(?:PROVENANCE\.)?event_type\s*=\s*'([^']+)'/i, raw) do - [_, et] -> Map.put(params, :event_type, et) - _ -> params - end - - params - end - defp extract_provenance_query(%{TAG: "ProvenanceActor", _0: actor}), do: %{actor: actor} - defp extract_provenance_query(%{TAG: "ProvenanceOrigin", _0: origin}), do: %{origin: origin} - defp extract_provenance_query(%{TAG: "ProvenanceChainValid", _0: valid}), do: %{chain_valid: valid} - defp extract_provenance_query(%{TAG: "ProvenanceEventType", _0: event_type}), do: %{event_type: event_type} - defp extract_provenance_query(%{TAG: "And", _0: left, _1: right}) do - Map.merge(extract_provenance_query(left), extract_provenance_query(right)) - end - defp extract_provenance_query(%{modality: mod, field: field, value: value}) - when mod in [:provenance, "PROVENANCE", "provenance"] do - %{String.to_existing_atom(field) => value} - end - defp extract_provenance_query(_), do: %{} - - defp execute_provenance_query(params, limit) do - # Route to dedicated RustClient provenance methods for caching and telemetry - cond do - Map.has_key?(params, :entity_id) -> - # Get the full provenance chain for a specific entity - case RustClient.get_provenance_chain(params.entity_id) do - {:ok, body} -> {:ok, wrap_provenance_results(body)} - {:error, :not_found} -> {:ok, []} - {:error, reason} -> {:error, {:provenance_query_failed, reason}} - end - - Map.has_key?(params, :verify_entity_id) -> - # Verify chain integrity for a specific entity - case RustClient.verify_provenance(params.verify_entity_id) do - {:ok, body} -> {:ok, wrap_provenance_results([body])} - {:error, :not_found} -> {:ok, []} - {:error, reason} -> {:error, {:provenance_query_failed, reason}} - end - - Map.has_key?(params, :actor) -> - # Search by actor — uses the provenance search-by-actor endpoint - case RustClient.post("/provenance/search", %{actor: params.actor, limit: limit || 50}) do - {:ok, %{status: 200, body: body}} -> {:ok, wrap_provenance_results(body)} - {:error, reason} -> {:error, {:provenance_query_failed, reason}} - _ -> {:ok, []} - end - - Map.has_key?(params, :chain_valid) -> - # Verify chain integrity across all provenance chains - case RustClient.get("/provenance/verify-all") do - {:ok, %{status: 200, body: body}} when is_list(body) -> - filtered = if params.chain_valid do - Enum.filter(body, &(&1["chain_valid"] == true)) - else - Enum.filter(body, &(&1["chain_valid"] == false)) - end - {:ok, wrap_provenance_results(filtered)} - {:ok, %{status: 200, body: body}} when is_map(body) -> - {:ok, wrap_provenance_results(Map.get(body, "results", []))} - {:error, reason} -> {:error, {:provenance_query_failed, reason}} - _ -> {:ok, []} - end - - true -> - # Generic provenance search with all params - case RustClient.post("/provenance/search", Map.put(params, :limit, limit || 50)) do - {:ok, %{status: 200, body: body}} -> {:ok, wrap_provenance_results(body)} - {:error, reason} -> {:error, {:provenance_query_failed, reason}} - _ -> {:ok, []} - end - end - end - - defp wrap_provenance_results(body) when is_list(body) do - Enum.map(body, fn item -> - %{"provenance" => item, "_source" => "provenance"} - end) - end - defp wrap_provenance_results(body) when is_map(body) do - results = Map.get(body, "results", [body]) - wrap_provenance_results(results) - end - defp wrap_provenance_results(_), do: [] - - # --------------------------------------------------------------------------- - # Spatial query extraction and execution - # --------------------------------------------------------------------------- - - defp extract_spatial_query(nil), do: %{} - defp extract_spatial_query(%{raw: raw}) when is_binary(raw) do - upper = String.upcase(raw) - - cond do - # WITHIN RADIUS(lat, lon, radius_km) - String.contains?(upper, "WITHIN RADIUS") -> - case Regex.run(~r/WITHIN\s+RADIUS\s*\(\s*([0-9.\-]+)\s*,\s*([0-9.\-]+)\s*,\s*([0-9.]+)\s*\)/i, raw) do - [_, lat, lon, radius] -> - %{type: :radius, - latitude: parse_float_safe(lat), - longitude: parse_float_safe(lon), - radius_km: parse_float_safe(radius)} - _ -> %{} - end - - # WITHIN BOUNDS(min_lat, min_lon, max_lat, max_lon) - String.contains?(upper, "WITHIN BOUNDS") -> - case Regex.run(~r/WITHIN\s+BOUNDS\s*\(\s*([0-9.\-]+)\s*,\s*([0-9.\-]+)\s*,\s*([0-9.\-]+)\s*,\s*([0-9.\-]+)\s*\)/i, raw) do - [_, min_lat, min_lon, max_lat, max_lon] -> - %{type: :bounds, - min_lat: parse_float_safe(min_lat), - min_lon: parse_float_safe(min_lon), - max_lat: parse_float_safe(max_lat), - max_lon: parse_float_safe(max_lon)} - _ -> %{} - end - - # NEAREST(lat, lon, k) - String.contains?(upper, "NEAREST") -> - case Regex.run(~r/NEAREST\s*\(\s*([0-9.\-]+)\s*,\s*([0-9.\-]+)\s*,\s*([0-9]+)\s*\)/i, raw) do - [_, lat, lon, k] -> - %{type: :nearest, - latitude: parse_float_safe(lat), - longitude: parse_float_safe(lon), - k: String.to_integer(k)} - _ -> %{} - end - - true -> %{} - end - end - defp extract_spatial_query(%{TAG: "SpatialRadius", _0: lat, _1: lon, _2: radius}) do - %{type: :radius, latitude: lat, longitude: lon, radius_km: radius} - end - defp extract_spatial_query(%{TAG: "SpatialBounds", _0: min_lat, _1: min_lon, _2: max_lat, _3: max_lon}) do - %{type: :bounds, min_lat: min_lat, min_lon: min_lon, max_lat: max_lat, max_lon: max_lon} - end - defp extract_spatial_query(%{TAG: "SpatialNearest", _0: lat, _1: lon, _2: k}) do - %{type: :nearest, latitude: lat, longitude: lon, k: k} - end - defp extract_spatial_query(%{TAG: "And", _0: left, _1: right}) do - case {has_spatial_condition?(left), has_spatial_condition?(right)} do - {true, _} -> extract_spatial_query(left) - {_, true} -> extract_spatial_query(right) - _ -> %{} - end - end - defp extract_spatial_query(_), do: %{} - - defp execute_spatial_query(params, limit) do - # Route to dedicated RustClient spatial methods for caching and telemetry - case Map.get(params, :type) do - :radius -> - case RustClient.search_spatial_radius( - params.latitude, - params.longitude, - params.radius_km, - limit: limit || 50 - ) do - {:ok, results} -> {:ok, wrap_spatial_results(results)} - {:error, reason} -> {:error, {:spatial_query_failed, reason}} - end - - :bounds -> - case RustClient.search_spatial_bounds( - params.min_lat, - params.min_lon, - params.max_lat, - params.max_lon, - limit: limit || 50 - ) do - {:ok, results} -> {:ok, wrap_spatial_results(results)} - {:error, reason} -> {:error, {:spatial_query_failed, reason}} - end - - :nearest -> - k = params[:k] || limit || 10 - - case RustClient.search_spatial_nearest( - params.latitude, - params.longitude, - k - ) do - {:ok, results} -> {:ok, wrap_spatial_results(results)} - {:error, reason} -> {:error, {:spatial_query_failed, reason}} - end - - _ -> - {:error, {:invalid_spatial_query, "Missing spatial query type (radius, bounds, or nearest)"}} - end - end - - defp wrap_spatial_results(body) when is_list(body) do - Enum.map(body, fn item -> - %{"spatial" => item, "_source" => "spatial", - "distance_km" => Map.get(item, "distance_km")} - end) - end - defp wrap_spatial_results(body) when is_map(body) do - results = Map.get(body, "results", [body]) - wrap_spatial_results(results) - end - defp wrap_spatial_results(_), do: [] - - defp parse_float_safe(str) when is_binary(str) do - case Float.parse(str) do - {f, _} -> f - :error -> 0.0 - end - end - defp parse_float_safe(num) when is_number(num), do: num / 1 - defp parse_float_safe(_), do: 0.0 - - defp extract_multi_modal_params(modalities, where_clause) do - %{modalities: modalities, conditions: where_clause} - end - - defp generate_explain_plan(query_ast) do - # Generate a cost-aware execution plan based on actual query structure. - # Tries the Rust verisim-planner first; falls back to local estimation. - modalities = extract_modalities(query_ast) - source = extract_source(query_ast) - where_clause = extract_where(query_ast) - proof_specs = extract_proof(query_ast) - group_by = extract_group_by(query_ast) - {_pushdown, cross_modal} = classify_conditions(where_clause) - - # Try Rust planner API for optimized plan - case RustClient.post("/query/explain", %{query_ast: query_ast}) do - {:ok, %{status: 200, body: plan}} -> - plan - - _ -> - # Fallback: local cost estimation - steps = [] - - # Step 1: Parse (already done) - steps = [%{operation: "Parse VCL", cost_ms: 1, notes: "Already completed"} | steps] - - # Step 2: Type check (if proofs present) - steps = if proof_specs do - proof_count = length(proof_specs) - [%{operation: "Type check + proof verification", cost_ms: 5 * proof_count, - notes: "#{proof_count} proof obligation(s)"} | steps] - else - steps - end - - # Step 3: Route to stores - {source_type, source_cost, source_notes} = case source do - {:octad, id} -> {"Octad lookup", 2, "Direct ID: #{id}"} - {:federation, pattern, _} -> {"Federation fan-out", 100, "Pattern: #{inspect(pattern)}"} - {:store, id} -> {"Store query", 15, "Store: #{id}"} - :reflect -> {"REFLECT (meta-query)", 20, "Query the query store itself"} - end - steps = [%{operation: source_type, cost_ms: source_cost, notes: source_notes} | steps] - - # Step 4: Modality queries - modality_cost = length(modalities) * 10 - steps = [%{operation: "Query #{length(modalities)} modality store(s)", - cost_ms: modality_cost, - modalities: Enum.map(modalities, &to_string/1)} | steps] - - # Step 5: Where clause evaluation - where_cost = if where_clause, do: 5, else: 0 - steps = if where_clause do - query_type = determine_query_type(modalities, where_clause) - [%{operation: "Evaluate WHERE (#{query_type})", cost_ms: where_cost} | steps] - else - steps - end - - # Step 6: Cross-modal evaluation (if any) - steps = if cross_modal != [] do - [%{operation: "Cross-modal evaluation", cost_ms: 20 * length(cross_modal), - conditions: length(cross_modal), - notes: "Post-fetch filter across modalities"} | steps] - else - steps - end - - # Step 7: Aggregation (if GROUP BY present) - steps = if group_by do - [%{operation: "Group + Aggregate", cost_ms: 8, - group_fields: length(group_by)} | steps] - else - steps - end - - steps = Enum.reverse(steps) - total = Enum.reduce(steps, 0, fn step, acc -> acc + Map.get(step, :cost_ms, 0) end) - - %{ - strategy: if(cross_modal != [], do: :two_phase, else: :sequential), - steps: steps, - total_cost_ms: total, - modalities_queried: modalities, - has_cross_modal: cross_modal != [], - has_proof: proof_specs != nil - } - end - end - - # --------------------------------------------------------------------------- - # Post-processing: GROUP BY, Aggregation, ORDER BY, Projection - # --------------------------------------------------------------------------- - - defp maybe_group_and_aggregate(rows, nil, _aggregates), do: rows - defp maybe_group_and_aggregate(rows, _group_by, nil), do: rows - defp maybe_group_and_aggregate(rows, group_by, aggregates) do - grouped = Enum.group_by(rows, fn row -> - Enum.map(group_by, fn %{modality: mod, field: field} -> - get_in(row, [to_string(mod), field]) - end) - end) - - Enum.map(grouped, fn {group_key, group_rows} -> - base = group_by - |> Enum.zip(group_key) - |> Enum.into(%{}, fn {%{modality: mod, field: field}, val} -> - {"#{mod}.#{field}", val} - end) - - Enum.reduce(aggregates, base, fn agg, acc -> - case agg do - :count_all -> - Map.put(acc, "COUNT(*)", length(group_rows)) - - {:aggregate_field, func, %{modality: mod, field: field}} -> - values = Enum.map(group_rows, fn row -> - get_in(row, [to_string(mod), field]) || 0 - end) - - result = case func do - :count -> length(values) - :sum -> Enum.sum(values) - :avg -> - if length(values) > 0, do: Enum.sum(values) / length(values), else: 0 - :min -> Enum.min(values, fn -> 0 end) - :max -> Enum.max(values, fn -> 0 end) - end - - label = "#{String.upcase(to_string(func))}(#{mod}.#{field})" - Map.put(acc, label, result) - end - end) - end) - end - - defp maybe_order_by(rows, nil), do: rows - defp maybe_order_by(rows, order_items) do - Enum.sort_by(rows, fn row -> - Enum.map(order_items, fn %{field: %{modality: mod, field: field}} -> - get_in(row, [to_string(mod), field]) || get_in(row, ["#{mod}.#{field}"]) - end) - end, fn a, b -> - order_items - |> Enum.zip(Enum.zip(a, b)) - |> Enum.reduce_while(:eq, fn {item, {va, vb}}, _acc -> - cmp = compare_values(va, vb) - direction = Map.get(item, :direction, :asc) - effective = if direction == :desc, do: invert_cmp(cmp), else: cmp - case effective do - :eq -> {:cont, :eq} - :lt -> {:halt, true} - :gt -> {:halt, false} - end - end) - |> case do - :eq -> true - bool -> bool - end - end) - end - - defp compare_values(a, b) when is_number(a) and is_number(b) do - cond do - a < b -> :lt - a > b -> :gt - true -> :eq - end - end - defp compare_values(a, b) when is_binary(a) and is_binary(b) do - cond do - a < b -> :lt - a > b -> :gt - true -> :eq - end - end - defp compare_values(_a, _b), do: :eq - - defp invert_cmp(:lt), do: :gt - defp invert_cmp(:gt), do: :lt - defp invert_cmp(:eq), do: :eq - - defp maybe_project_columns(rows, nil), do: rows - defp maybe_project_columns(rows, projections) do - Enum.map(rows, fn row -> - Enum.into(projections, %{}, fn %{modality: mod, field: field} -> - value = get_in(row, [to_string(mod), field]) || get_in(row, ["#{mod}.#{field}"]) - {"#{mod}.#{field}", value} - end) - end) - end - - # --------------------------------------------------------------------------- - # AST extraction helpers - # --------------------------------------------------------------------------- - - defp extract_modalities(query_ast), do: Map.get(query_ast, :modalities, [:all]) - defp extract_source(query_ast) do - case Map.get(query_ast, :source, {:octad, "default"}) do - %{TAG: "Reflect"} -> :reflect - :reflect -> :reflect - "REFLECT" -> :reflect - other -> other - end - end - defp extract_where(query_ast), do: Map.get(query_ast, :where, nil) - - defp extract_proof(query_ast) do - case Map.get(query_ast, :proof, nil) do - nil -> nil - proof when is_list(proof) -> proof - proof when is_map(proof) -> [proof] # backward compat: single proof → list - end - end - - defp extract_limit(query_ast), do: Map.get(query_ast, :limit, nil) - defp extract_offset(query_ast), do: Map.get(query_ast, :offset, nil) - defp extract_order_by(query_ast), do: Map.get(query_ast, :orderBy, nil) - defp extract_group_by(query_ast), do: Map.get(query_ast, :groupBy, nil) - defp extract_aggregates(query_ast), do: Map.get(query_ast, :aggregates, nil) - defp extract_projections(query_ast), do: Map.get(query_ast, :projections, nil) -end diff --git a/verisimdb/elixir-orchestration/lib/verisim/query/vcl_proof_certificate.ex b/verisimdb/elixir-orchestration/lib/verisim/query/vcl_proof_certificate.ex deleted file mode 100644 index a9bc62fd..00000000 --- a/verisimdb/elixir-orchestration/lib/verisim/query/vcl_proof_certificate.ex +++ /dev/null @@ -1,190 +0,0 @@ -# SPDX-License-Identifier: MPL-2.0 - -defmodule VeriSim.Query.VCLProofCertificate do - @moduledoc """ - VCL-UT proof certificate generation and verification. - - After the VCL type checker validates a query's proof obligations and the - executor verifies the proofs at runtime, this module produces independently - verifiable certificates. Each certificate bundles the proof obligation, - witness data, and a SHA-256 integrity hash so that any party can later - confirm that a proof was satisfied without re-executing the query. - - ## Certificate structure - - ``` - %{ - type: :existence, # proof type atom - obligation: %{...}, # structured obligation from type checker - witness: %{...}, # witness data gathered during execution - timestamp: ~U[2026-02-28 12:00:00Z], # when the proof was verified - hash: <> # integrity hash of (obligation ++ witness ++ timestamp) - } - ``` - - ## Privacy guarantees - - Certificates contain only proof metadata (types, hashes, timestamps) — never - query content, entity data, or PII. The witness fields are structural (e.g., - `octad_id`, `drift_score`, `chain_length`) not content-bearing. - - ## Usage - - # After type checking succeeds and proof is verified: - obligation = %{type: :existence, witness_fields: ["octad_id", "timestamp", "modality_count"], ...} - witness = %{"octad_id" => "entity-001", "timestamp" => "2026-02-28T12:00:00Z", "modality_count" => 8} - - {:ok, cert} = VCLProofCertificate.generate_certificate(obligation, witness) - :ok = VCLProofCertificate.verify_certificate(cert) - - ## Batch certificates - - obligations_and_witnesses = [{obl1, wit1}, {obl2, wit2}] - {:ok, certs} = VCLProofCertificate.generate_batch(obligations_and_witnesses) - :ok = VCLProofCertificate.verify_batch(certs) - """ - - @doc """ - Generate a proof certificate from a type-checked obligation and runtime witness. - - The certificate includes a SHA-256 hash computed over the canonical - representation of the obligation, witness, and timestamp. This hash - serves as an integrity seal — any tampering with the certificate will - cause `verify_certificate/1` to fail. - - ## Parameters - - - `obligation` — proof obligation map from `VCLTypeChecker.typecheck/1` - - `witness` — witness data map gathered during proof execution - - ## Returns - - - `{:ok, certificate}` — valid certificate with integrity hash - - `{:error, reason}` — obligation or witness is malformed - """ - def generate_certificate(obligation, witness) when is_map(obligation) and is_map(witness) do - proof_type = obligation[:type] || obligation["type"] - - if proof_type == nil do - {:error, {:invalid_obligation, "Obligation must have a :type field"}} - else - timestamp = DateTime.utc_now() - hash = compute_hash(obligation, witness, timestamp) - - certificate = %{ - type: proof_type, - obligation: obligation, - witness: witness, - timestamp: timestamp, - hash: hash - } - - {:ok, certificate} - end - end - - def generate_certificate(_obligation, _witness) do - {:error, {:invalid_arguments, "Obligation and witness must both be maps"}} - end - - @doc """ - Verify a proof certificate's integrity by recomputing its hash. - - Checks that the hash stored in the certificate matches a fresh SHA-256 - computation over the same obligation, witness, and timestamp. This - confirms the certificate has not been tampered with since generation. - - ## Parameters - - - `certificate` — a certificate map previously returned by `generate_certificate/2` - - ## Returns - - - `:ok` — certificate is valid and untampered - - `{:error, :invalid_hash}` — hash mismatch (certificate was modified) - - `{:error, :malformed_certificate}` — missing required fields - """ - def verify_certificate(%{ - obligation: obligation, - witness: witness, - timestamp: timestamp, - hash: stored_hash - }) - when is_map(obligation) and is_map(witness) and is_binary(stored_hash) do - recomputed = compute_hash(obligation, witness, timestamp) - - if recomputed == stored_hash do - :ok - else - {:error, :invalid_hash} - end - end - - def verify_certificate(_), do: {:error, :malformed_certificate} - - @doc """ - Generate certificates for a batch of obligation/witness pairs. - - Useful when a VCL-UT query has multiple PROOF clauses (e.g., - `PROOF EXISTENCE(x) AND PROVENANCE(x)`). - - ## Parameters - - - `pairs` — list of `{obligation, witness}` tuples - - ## Returns - - - `{:ok, certificates}` — list of valid certificates - - `{:error, reason}` — if any pair fails - """ - def generate_batch(pairs) when is_list(pairs) do - results = - Enum.reduce_while(pairs, {:ok, []}, fn {obligation, witness}, {:ok, acc} -> - case generate_certificate(obligation, witness) do - {:ok, cert} -> {:cont, {:ok, acc ++ [cert]}} - {:error, _} = err -> {:halt, err} - end - end) - - results - end - - @doc """ - Verify a batch of certificates. Returns `:ok` only if all pass. - - ## Parameters - - - `certificates` — list of certificate maps - - ## Returns - - - `:ok` — all certificates verified - - `{:error, {:batch_failure, index, reason}}` — certificate at `index` failed - """ - def verify_batch(certificates) when is_list(certificates) do - certificates - |> Enum.with_index() - |> Enum.reduce_while(:ok, fn {cert, index}, :ok -> - case verify_certificate(cert) do - :ok -> {:cont, :ok} - {:error, reason} -> {:halt, {:error, {:batch_failure, index, reason}}} - end - end) - end - - # --------------------------------------------------------------------------- - # Private: Hash computation - # --------------------------------------------------------------------------- - - # Compute SHA-256 over the canonical serialisation of obligation, witness, - # and timestamp. We use `:erlang.term_to_binary/1` for deterministic - # serialisation of Elixir terms, then hash the concatenated binaries. - defp compute_hash(obligation, witness, timestamp) do - canonical = - :erlang.term_to_binary(obligation) <> - :erlang.term_to_binary(witness) <> - :erlang.term_to_binary(timestamp) - - :crypto.hash(:sha256, canonical) - end -end diff --git a/verisimdb/elixir-orchestration/lib/verisim/query/vcl_type_checker.ex b/verisimdb/elixir-orchestration/lib/verisim/query/vcl_type_checker.ex deleted file mode 100644 index 76485c66..00000000 --- a/verisimdb/elixir-orchestration/lib/verisim/query/vcl_type_checker.ex +++ /dev/null @@ -1,353 +0,0 @@ -# SPDX-License-Identifier: MPL-2.0 - -defmodule VeriSim.Query.VCLTypeChecker do - @moduledoc """ - Elixir-native VCL-UT type checker. - - Provides lightweight type checking for VCL queries with PROOF clauses when the - ReScript subprocess (VCLBidir) is unavailable. This ensures that VCL-UT queries - are always type-checked before execution — never silently downgraded to slipstream. - - ## Design - - The ReScript type checker (VCLBidir.res) is the canonical implementation with - bidirectional type inference and full subtyping. This Elixir module implements a - pragmatic subset of the same rules: - - 1. **Modality validation** — queried modalities match proof requirements - 2. **Proof type validation** — proof types are known and well-formed - 3. **Multi-proof composition** — composition strategies are valid - 4. **Contract reference validation** — proof specs that need contracts have them - 5. **Proof obligation generation** — structured obligations with witness fields - - Falls back to the ReScript type checker when it's available (via VCLBridge.typecheck/2). - - ## Proof Types and Required Modalities - - | Proof Type | Required Modalities | Circuit Name | - |-------------|---------------------------|-----------------------| - | EXISTENCE | any | existence-proof-v1 | - | INTEGRITY | semantic | integrity-proof-v1 | - | CONSISTENCY | 2+ modalities | consistency-proof-v1 | - | PROVENANCE | provenance | provenance-proof-v1 | - | FRESHNESS | temporal | freshness-proof-v1 | - | ACCESS | any | access-control-v1 | - | CITATION | document or semantic | citation-proof-v1 | - | CUSTOM | varies (circuit-defined) | (user-specified) | - | ZKP | any | (privacy-aware) | - | PROVEN | semantic | proven-cert-v1 | - | SANCTIFY | semantic | sanctify-v1 | - """ - - require Logger - - @known_proof_types ~w( - existence integrity consistency provenance freshness access citation - custom zkp proven sanctify - )a - - @proof_required_modalities %{ - existence: [], - integrity: [:semantic], - consistency: [], - provenance: [:provenance], - freshness: [:temporal], - access: [], - citation: [:document, :semantic], - custom: [], - zkp: [], - proven: [:semantic], - sanctify: [:semantic] - } - - @proof_witness_fields %{ - existence: ["octad_id", "timestamp", "modality_count"], - integrity: ["content_hash", "merkle_root", "schema_version"], - consistency: ["modality_a", "modality_b", "drift_score", "threshold"], - provenance: ["chain_hash", "chain_length", "origin", "actor_trail"], - freshness: ["last_modified", "max_age_ms", "version_count"], - access: ["principal_id", "resource_id", "permission_set"], - citation: ["source_ids", "citation_chain", "reference_count"], - custom: ["circuit_inputs"], - zkp: ["claim", "blinding_nonce"], - proven: ["certificate_hash", "proof_data"], - sanctify: ["contract_hash", "security_level"] - } - - @proof_circuits %{ - existence: "existence-proof-v1", - integrity: "integrity-proof-v1", - consistency: "consistency-proof-v1", - provenance: "provenance-proof-v1", - freshness: "freshness-proof-v1", - access: "access-control-v1", - citation: "citation-proof-v1", - custom: nil, - zkp: nil, - proven: "proven-cert-v1", - sanctify: "sanctify-v1" - } - - @proof_time_estimates_ms %{ - existence: 50, - integrity: 200, - consistency: 250, - provenance: 300, - freshness: 100, - access: 150, - citation: 100, - custom: 500, - zkp: 400, - proven: 200, - sanctify: 200 - } - - # --------------------------------------------------------------------------- - # Public API - # --------------------------------------------------------------------------- - - @doc """ - Type-check a VCL-UT query AST. - - Validates proof types, modality compatibility, and composition rules. - Returns structured proof obligations ready for the executor. - - ## Returns - - - `{:ok, type_info}` with: - - `:proof_obligations` — list of structured obligation maps - - `:composition_strategy` — how proofs compose (:conjunction, :independent, :sequential) - - `:inferred_types` — modality type map (empty in native checker) - - `:total_estimated_ms` — sum of per-proof time estimates - - `:is_parallelizable` — whether proofs can run concurrently - - `{:error, reason}` — type checking failed - """ - def typecheck(query_ast) do - modalities = extract_modalities(query_ast) - proof_specs = extract_and_split_proofs(query_ast) - - with :ok <- validate_modalities(modalities), - :ok <- validate_proof_specs(proof_specs), - :ok <- validate_modality_compatibility(proof_specs, modalities), - {:ok, obligations} <- generate_obligations(proof_specs), - {:ok, composition} <- determine_composition(obligations) do - total_ms = Enum.reduce(obligations, 0, & &1.estimated_time_ms + &2) - - {:ok, %{ - proof_obligations: obligations, - composition_strategy: composition, - inferred_types: %{}, - total_estimated_ms: total_ms, - is_parallelizable: composition == :independent - }} - end - end - - @doc """ - Parse a raw PROOF clause string into a list of individual proof spec maps. - - Splits on AND/OR connectors and extracts proof type + contract name from - each spec. Handles the built-in parser's `%{raw: "..."}` format. - - ## Examples - - iex> VCLTypeChecker.parse_proof_specs(%{raw: "EXISTENCE(entity-001) AND PROVENANCE(entity-001)"}) - [ - %{proofType: "EXISTENCE", contractName: "entity-001", raw: "EXISTENCE(entity-001)"}, - %{proofType: "PROVENANCE", contractName: "entity-001", raw: "PROVENANCE(entity-001)"} - ] - - iex> VCLTypeChecker.parse_proof_specs(%{raw: "INTEGRITY(my_contract)"}) - [%{proofType: "INTEGRITY", contractName: "my_contract", raw: "INTEGRITY(my_contract)"}] - """ - def parse_proof_specs(nil), do: [] - def parse_proof_specs(specs) when is_list(specs) do - Enum.flat_map(specs, &parse_proof_specs/1) - end - def parse_proof_specs(%{raw: raw}) when is_binary(raw) do - # Split on AND/OR connectors (case-insensitive), preserving each proof spec - raw - |> String.split(~r/\s+AND\s+|\s+OR\s+/i) - |> Enum.map(&String.trim/1) - |> Enum.reject(&(&1 == "")) - |> Enum.map(fn spec_str -> - {proof_type, contract_name} = parse_single_proof_spec(spec_str) - %{ - proofType: proof_type, - contractName: contract_name, - raw: spec_str - } - end) - end - def parse_proof_specs(%{proofType: _} = spec), do: [spec] - def parse_proof_specs(%{TAG: _} = spec), do: [spec] - def parse_proof_specs(_), do: [] - - # --------------------------------------------------------------------------- - # Private: Validation - # --------------------------------------------------------------------------- - - defp validate_modalities([]), do: {:error, {:invalid_query, "No modalities specified"}} - defp validate_modalities(_modalities), do: :ok - - defp validate_proof_specs([]) do - {:error, {:missing_proof, "VCL-UT query requires at least one PROOF specification"}} - end - defp validate_proof_specs(specs) do - Enum.reduce_while(specs, :ok, fn spec, _acc -> - proof_type = normalize_proof_type(spec) - cond do - proof_type == :unknown -> - raw = Map.get(spec, :raw, Map.get(spec, :proofType, inspect(spec))) - {:halt, {:error, {:unknown_proof_type, "Unknown proof type: #{raw}"}}} - - needs_contract?(proof_type) and not has_contract?(spec) -> - {:halt, {:error, {:missing_contract, - "#{proof_type_name(proof_type)} proof requires a contract/entity reference"}}} - - true -> - {:cont, :ok} - end - end) - end - - defp validate_modality_compatibility(proof_specs, modalities) do - # For each proof, check that required modalities are being queried - # (or that :all is in the modality list) - has_all = :all in modalities - - Enum.reduce_while(proof_specs, :ok, fn spec, _acc -> - proof_type = normalize_proof_type(spec) - required = Map.get(@proof_required_modalities, proof_type, []) - - if has_all or required == [] or Enum.any?(required, &(&1 in modalities)) do - {:cont, :ok} - else - required_str = required |> Enum.map(&to_string/1) |> Enum.join(", ") - {:halt, {:error, {:modality_mismatch, - "#{proof_type_name(proof_type)} proof requires #{required_str} modality " <> - "but query only selects #{inspect(modalities)}"}}} - end - end) - end - - # --------------------------------------------------------------------------- - # Private: Obligation Generation - # --------------------------------------------------------------------------- - - defp generate_obligations(proof_specs) do - obligations = Enum.map(proof_specs, fn spec -> - proof_type = normalize_proof_type(spec) - contract = extract_contract(spec) - - %{ - type: proof_type, - proofType: proof_type |> Atom.to_string() |> String.upcase(), - contract: contract, - contractName: contract, - witness_fields: Map.get(@proof_witness_fields, proof_type, []), - circuit: circuit_for(proof_type, spec), - estimated_time_ms: Map.get(@proof_time_estimates_ms, proof_type, 200), - required_modalities: Map.get(@proof_required_modalities, proof_type, []) - } - end) - - {:ok, obligations} - end - - defp determine_composition([]), do: {:ok, :independent} - defp determine_composition([_single]), do: {:ok, :independent} - defp determine_composition(obligations) do - types = Enum.map(obligations, & &1.type) |> MapSet.new() - - # Provenance + Citation must be sequential (citation validated first) - if :provenance in types and :citation in types do - {:ok, :sequential} - else - {:ok, :conjunction} - end - end - - # --------------------------------------------------------------------------- - # Private: Proof spec parsing helpers - # --------------------------------------------------------------------------- - - defp extract_and_split_proofs(query_ast) do - proof = query_ast[:proof] || query_ast["proof"] - parse_proof_specs(proof) - end - - defp extract_modalities(query_ast) do - query_ast[:modalities] || query_ast["modalities"] || [:all] - end - - defp parse_single_proof_spec(spec_str) do - # Handle "PROOF_TYPE(contract)" and "PROOF_TYPE contract" formats - case Regex.run(~r/^([A-Z_]+)\(([^)]*)\)$/, String.trim(spec_str)) do - [_, proof_type, contract] -> - {proof_type, String.trim(contract)} - _ -> - # Try space-separated: "EXISTENCE entity-001" - case String.split(String.trim(spec_str), ~r/\s+/, parts: 2) do - [proof_type, contract] -> {proof_type, String.trim(contract)} - [proof_type] -> {proof_type, nil} - _ -> {spec_str, nil} - end - end - end - - defp normalize_proof_type(%{proofType: type}), do: do_normalize(type) - defp normalize_proof_type(%{TAG: tag}), do: do_normalize(tag) - defp normalize_proof_type(%{type: type}) when is_atom(type), do: type - defp normalize_proof_type(%{raw: raw}) when is_binary(raw) do - raw - |> String.split(~r/[\s(]/, parts: 2) - |> List.first() - |> do_normalize() - end - defp normalize_proof_type(_), do: :unknown - - defp do_normalize(str) when is_binary(str) do - # String.to_existing_atom/1 raises ArgumentError for atoms not yet interned. - # Rescue it so unknown proof type names degrade to :unknown instead of crashing. - atom = - str - |> String.downcase() - |> String.to_existing_atom() - - if atom in @known_proof_types, do: atom, else: :unknown - rescue - ArgumentError -> :unknown - end - defp do_normalize(atom) when is_atom(atom) do - if atom in @known_proof_types, do: atom, else: :unknown - end - defp do_normalize(_), do: :unknown - - defp extract_contract(%{contractName: name}) when is_binary(name) and name != "", do: name - defp extract_contract(%{contract: name}) when is_binary(name) and name != "", do: name - defp extract_contract(%{raw: raw}) when is_binary(raw) do - case Regex.run(~r/\(([^)]+)\)/, raw) do - [_, name] -> String.trim(name) - _ -> nil - end - end - defp extract_contract(_), do: nil - - defp needs_contract?(type) when type in [:integrity, :citation, :custom, :sanctify], do: true - defp needs_contract?(_), do: false - - defp has_contract?(spec) do - contract = extract_contract(spec) - is_binary(contract) and contract != "" - end - - defp circuit_for(:custom, spec) do - extract_contract(spec) || "custom-circuit" - end - defp circuit_for(type, _spec), do: Map.get(@proof_circuits, type) - - defp proof_type_name(type) do - type |> Atom.to_string() |> String.upcase() - end -end diff --git a/verisimdb/elixir-orchestration/lib/verisim/rust_client.ex b/verisimdb/elixir-orchestration/lib/verisim/rust_client.ex deleted file mode 100644 index b3c8650b..00000000 --- a/verisimdb/elixir-orchestration/lib/verisim/rust_client.ex +++ /dev/null @@ -1,460 +0,0 @@ -# SPDX-License-Identifier: MPL-2.0 - -defmodule VeriSim.RustClient do - @moduledoc """ - HTTP client for communicating with the Rust core (verisim-api). - - This module provides a typed interface to the Rust HTTP API, - handling serialization, error handling, and retries. - """ - - require Logger - - @default_base_url "http://localhost:8080/api/v1" - @default_timeout 30_000 - @cache_table :verisim_rust_client_cache - @cache_ttl_ms 30_000 # 30 seconds default TTL - - # --------------------------------------------------------------------------- - # ETS Cache — transparent read-through caching for octad lookups and searches - # --------------------------------------------------------------------------- - - @doc """ - Initialize the ETS cache table. Call from application supervisor or init. - """ - def init_cache do - if :ets.info(@cache_table) == :undefined do - :ets.new(@cache_table, [:set, :public, :named_table, read_concurrency: true]) - end - :ok - end - - @doc """ - Clear all cached entries. - """ - def clear_cache do - if :ets.info(@cache_table) != :undefined do - :ets.delete_all_objects(@cache_table) - end - :ok - end - - @doc """ - Invalidate a specific cache key (e.g., after a write operation). - """ - def invalidate_cache(key) do - if :ets.info(@cache_table) != :undefined do - :ets.delete(@cache_table, key) - end - :ok - end - - defp cache_get(key) do - if :ets.info(@cache_table) != :undefined do - case :ets.lookup(@cache_table, key) do - [{^key, value, expiry}] -> - if System.monotonic_time(:millisecond) < expiry do - {:hit, value} - else - :ets.delete(@cache_table, key) - :miss - end - _ -> :miss - end - else - :miss - end - end - - defp cache_put(key, value, ttl_ms \\ @cache_ttl_ms) do - if :ets.info(@cache_table) != :undefined do - expiry = System.monotonic_time(:millisecond) + ttl_ms - :ets.insert(@cache_table, {key, value, expiry}) - end - value - end - - # Configuration - - def base_url do - Application.get_env(:verisim, :rust_core_url, @default_base_url) - end - - def timeout do - Application.get_env(:verisim, :rust_core_timeout, @default_timeout) - end - - # Health Check - - @doc """ - Check the health of the Rust core. - """ - def health do - case get("/health") do - {:ok, %{status: 200, body: body}} -> - {:ok, body} - {:ok, %{status: status}} -> - {:error, {:unhealthy, status}} - {:error, reason} -> - {:error, reason} - end - end - - # Octad Operations - - @doc """ - Create a new octad entity. - """ - def create_octad(input) do - case post("/octads", input) do - {:ok, %{status: 200, body: body}} -> {:ok, body} - {:ok, %{status: 201, body: body}} -> {:ok, body} - {:ok, %{status: status, body: body}} -> {:error, {status, body}} - {:error, reason} -> {:error, reason} - end - end - - @doc """ - Get a octad by ID. - """ - def get_octad(entity_id) do - cache_key = {:octad, entity_id} - - case cache_get(cache_key) do - {:hit, cached} -> {:ok, cached} - :miss -> - case get("/octads/#{entity_id}") do - {:ok, %{status: 200, body: body}} -> - cache_put(cache_key, body) - {:ok, body} - {:ok, %{status: 404}} -> {:error, :not_found} - {:ok, %{status: status, body: body}} -> {:error, {status, body}} - {:error, reason} -> {:error, reason} - end - end - end - - @doc """ - Update a octad. - """ - def update_octad(entity_id, changes) do - case put("/octads/#{entity_id}", changes) do - {:ok, %{status: 200, body: body}} -> - invalidate_cache({:octad, entity_id}) - invalidate_cache({:drift, entity_id}) - {:ok, body} - {:ok, %{status: 404}} -> {:error, :not_found} - {:ok, %{status: status, body: body}} -> {:error, {status, body}} - {:error, reason} -> {:error, reason} - end - end - - @doc """ - Delete a octad. - """ - def delete_octad(entity_id) do - case delete("/octads/#{entity_id}") do - {:ok, %{status: 204}} -> - invalidate_cache({:octad, entity_id}) - invalidate_cache({:drift, entity_id}) - :ok - {:ok, %{status: 200}} -> - invalidate_cache({:octad, entity_id}) - invalidate_cache({:drift, entity_id}) - :ok - {:ok, %{status: 404}} -> {:error, :not_found} - {:ok, %{status: status, body: body}} -> {:error, {status, body}} - {:error, reason} -> {:error, reason} - end - end - - # Search Operations - - @doc """ - Search for octads by text. - """ - def search_text(query, limit \\ 10) do - case get("/search/text", q: query, limit: limit) do - {:ok, %{status: 200, body: body}} -> {:ok, body} - {:ok, %{status: status, body: body}} -> {:error, {status, body}} - {:error, reason} -> {:error, reason} - end - end - - @doc """ - Search for octads by vector similarity. - """ - def search_vector(vector, k \\ 10) do - case post("/search/vector", %{vector: vector, k: k}) do - {:ok, %{status: 200, body: body}} -> {:ok, body} - {:ok, %{status: status, body: body}} -> {:error, {status, body}} - {:error, reason} -> {:error, reason} - end - end - - @doc """ - Get related octads (graph query). - """ - def get_related(entity_id) do - case get("/search/related/#{entity_id}") do - {:ok, %{status: 200, body: body}} -> {:ok, body} - {:ok, %{status: 404}} -> {:error, :not_found} - {:ok, %{status: status, body: body}} -> {:error, {status, body}} - {:error, reason} -> {:error, reason} - end - end - - # Drift and Normalization - - @doc """ - Get the drift score for an entity. - """ - def get_drift_score(entity_id) do - cache_key = {:drift, entity_id} - - case cache_get(cache_key) do - {:hit, cached} -> {:ok, cached} - :miss -> - case get("/drift/entity/#{entity_id}") do - {:ok, %{status: 200, body: %{"score" => score}}} -> - cache_put(cache_key, score, 10_000) # Short TTL — drift changes frequently - {:ok, score} - {:ok, %{status: 404}} -> {:ok, 0.0} - {:ok, %{status: status, body: body}} -> {:error, {status, body}} - {:error, reason} -> {:error, reason} - end - end - end - - @doc """ - Get overall drift status. - """ - def drift_status do - case get("/drift/status") do - {:ok, %{status: 200, body: body}} -> {:ok, body} - {:ok, %{status: status, body: body}} -> {:error, {status, body}} - {:error, reason} -> {:error, reason} - end - end - - @doc """ - Trigger normalization for an entity. - """ - def normalize(entity_id) do - case post("/normalizer/trigger/#{entity_id}", %{}) do - {:ok, %{status: 202}} -> :ok - {:ok, %{status: 200}} -> :ok - {:ok, %{status: 404}} -> {:error, :not_found} - {:ok, %{status: status, body: body}} -> {:error, {status, body}} - {:error, reason} -> {:error, reason} - end - end - - @doc """ - Get normalizer status. - """ - def normalizer_status do - case get("/normalizer/status") do - {:ok, %{status: 200, body: body}} -> {:ok, body} - {:ok, %{status: status, body: body}} -> {:error, {status, body}} - {:error, reason} -> {:error, reason} - end - end - - # Provenance Operations - - @doc """ - Retrieve the full provenance chain for an entity. - """ - def get_provenance_chain(entity_id) do - cache_key = {:provenance_chain, entity_id} - - case cache_get(cache_key) do - {:hit, cached} -> - {:ok, cached} - - :miss -> - case get("/provenance/#{entity_id}") do - {:ok, %{status: 200, body: body}} -> - cache_put(cache_key, body, 15_000) - {:ok, body} - - {:ok, %{status: 404}} -> - {:error, :not_found} - - {:ok, %{status: status, body: body}} -> - {:error, {status, body}} - - {:error, reason} -> - {:error, reason} - end - end - end - - @doc """ - Record a new provenance event for an entity. - """ - def record_provenance(entity_id, event) do - case post("/provenance/#{entity_id}/record", event) do - {:ok, %{status: 200, body: body}} -> - invalidate_cache({:provenance_chain, entity_id}) - {:ok, body} - - {:ok, %{status: status, body: body}} -> - {:error, {status, body}} - - {:error, reason} -> - {:error, reason} - end - end - - @doc """ - Verify the hash-chain integrity for an entity's provenance. - """ - def verify_provenance(entity_id) do - case get("/provenance/#{entity_id}/verify") do - {:ok, %{status: 200, body: body}} -> {:ok, body} - {:ok, %{status: 404}} -> {:error, :not_found} - {:ok, %{status: status, body: body}} -> {:error, {status, body}} - {:error, reason} -> {:error, reason} - end - end - - # Spatial Operations - - @doc """ - Search for entities within a given radius of a point. - """ - def search_spatial_radius(latitude, longitude, radius_km, opts \\ []) do - body = %{ - latitude: latitude, - longitude: longitude, - radius_km: radius_km, - limit: Keyword.get(opts, :limit, 50) - } - - case post("/spatial/search/radius", body) do - {:ok, %{status: 200, body: results}} -> {:ok, results} - {:ok, %{status: status, body: body_resp}} -> {:error, {status, body_resp}} - {:error, reason} -> {:error, reason} - end - end - - @doc """ - Search for entities within a bounding box. - """ - def search_spatial_bounds(min_lat, min_lon, max_lat, max_lon, opts \\ []) do - body = %{ - min_lat: min_lat, - min_lon: min_lon, - max_lat: max_lat, - max_lon: max_lon, - limit: Keyword.get(opts, :limit, 50) - } - - case post("/spatial/search/bounds", body) do - {:ok, %{status: 200, body: results}} -> {:ok, results} - {:ok, %{status: status, body: body_resp}} -> {:error, {status, body_resp}} - {:error, reason} -> {:error, reason} - end - end - - @doc """ - Find the k nearest entities to a given point. - """ - def search_spatial_nearest(latitude, longitude, k \\ 10) do - body = %{ - latitude: latitude, - longitude: longitude, - k: k - } - - case post("/spatial/search/nearest", body) do - {:ok, %{status: 200, body: results}} -> {:ok, results} - {:ok, %{status: status, body: body_resp}} -> {:error, {status, body_resp}} - {:error, reason} -> {:error, reason} - end - end - - # HTTP Helpers (public: get/2 and post/2 used by Federation.Resolver) - - def get(path, params \\ []) do - url = base_url() <> path - - case Req.get(url, - params: params, - receive_timeout: timeout(), - decode_body: true - ) do - {:ok, resp} -> validate_json_response(resp) - {:error, reason} -> {:error, reason} - end - rescue - e -> {:error, {:request_failed, e}} - end - - def post(path, body) do - url = base_url() <> path - - case Req.post(url, - json: body, - receive_timeout: timeout(), - decode_body: true - ) do - {:ok, resp} -> validate_json_response(resp) - {:error, reason} -> {:error, reason} - end - rescue - e -> {:error, {:request_failed, e}} - end - - defp put(path, body) do - url = base_url() <> path - - case Req.put(url, - json: body, - receive_timeout: timeout(), - decode_body: true - ) do - {:ok, resp} -> validate_json_response(resp) - {:error, reason} -> {:error, reason} - end - rescue - e -> {:error, {:request_failed, e}} - end - - defp delete(path) do - url = base_url() <> path - - case Req.delete(url, - receive_timeout: timeout(), - decode_body: true - ) do - {:ok, resp} -> validate_json_response(resp) - {:error, reason} -> {:error, reason} - end - rescue - e -> {:error, {:request_failed, e}} - end - - # Validates that the response body is decoded JSON (map or list), not a raw - # string. When a non-VeriSimDB server (e.g. nginx) sits on the configured - # port, Req returns status 200 with an HTML string body. Without this check, - # upstream code crashes trying to enumerate/map over a binary. - defp validate_json_response(%{status: status, body: body} = resp) - when is_map(body) or is_list(body) do - {:ok, resp} - end - - defp validate_json_response(%{status: status, body: body}) - when is_binary(body) and status >= 200 and status < 300 do - Logger.warning("Rust core returned non-JSON response (status #{status}): #{String.slice(body, 0, 120)}...") - {:error, {:non_json_response, status, String.slice(body, 0, 500)}} - end - - defp validate_json_response(%{status: status, body: body}) when is_binary(body) do - {:error, {status, body}} - end - - defp validate_json_response(resp), do: {:ok, resp} -end diff --git a/verisimdb/elixir-orchestration/lib/verisim/schema/schema_registry.ex b/verisimdb/elixir-orchestration/lib/verisim/schema/schema_registry.ex deleted file mode 100644 index 1dfaec08..00000000 --- a/verisimdb/elixir-orchestration/lib/verisim/schema/schema_registry.ex +++ /dev/null @@ -1,254 +0,0 @@ -# SPDX-License-Identifier: MPL-2.0 - -defmodule VeriSim.SchemaRegistry do - @moduledoc """ - Schema Registry - Manages type definitions and constraints. - - Coordinates the semantic type system across the cluster, - ensuring consistent type validation and constraint checking. - - ## Type Hierarchy - - Types form a hierarchy: - - `verisim:Entity` - Base type for all entities - - `verisim:Document` - Entities with document content - - `verisim:Node` - Graph nodes - - Custom types defined by users - - ## Constraints - - Constraints are validated during entity creation/update: - - Required properties - - Property patterns (regex) - - Range constraints (min/max) - - Custom validators - """ - - use GenServer - require Logger - - # Client API - - def start_link(opts \\ []) do - GenServer.start_link(__MODULE__, opts, name: __MODULE__) - end - - @doc """ - Register a new type. - - ## Example - - SchemaRegistry.register_type(%{ - iri: "https://example.org/Person", - label: "Person", - supertypes: ["verisim:Entity"], - constraints: [ - %{name: "name_required", kind: {:required, "name"}, message: "Name is required"} - ] - }) - """ - def register_type(type_def) do - GenServer.call(__MODULE__, {:register_type, type_def}) - end - - @doc """ - Get a type definition by IRI. - """ - def get_type(iri) do - GenServer.call(__MODULE__, {:get_type, iri}) - end - - @doc """ - List all registered types. - """ - def list_types do - GenServer.call(__MODULE__, :list_types) - end - - @doc """ - Validate an entity against its declared types. - """ - def validate(entity) do - GenServer.call(__MODULE__, {:validate, entity}) - end - - @doc """ - Get the type hierarchy (supertypes) for a type. - """ - def type_hierarchy(iri) do - GenServer.call(__MODULE__, {:type_hierarchy, iri}) - end - - # Server Callbacks - - @impl true - def init(_opts) do - state = %{ - types: %{}, - type_cache: %{} - } - - # Register built-in types - state = register_builtin_types(state) - - {:ok, state} - end - - @impl true - def handle_call({:register_type, type_def}, _from, state) do - iri = type_def.iri - - if Map.has_key?(state.types, iri) do - {:reply, {:error, :already_exists}, state} - else - new_types = Map.put(state.types, iri, type_def) - new_state = %{state | types: new_types, type_cache: %{}} - - Logger.info("Registered type: #{iri}") - {:reply, :ok, new_state} - end - end - - @impl true - def handle_call({:get_type, iri}, _from, state) do - result = Map.get(state.types, iri) - {:reply, result, state} - end - - @impl true - def handle_call(:list_types, _from, state) do - types = Map.keys(state.types) - {:reply, types, state} - end - - @impl true - def handle_call({:validate, entity}, _from, state) do - types = Map.get(entity, :types, []) - - violations = - types - |> Enum.flat_map(fn type_iri -> - case Map.get(state.types, type_iri) do - nil -> [] - type_def -> validate_constraints(entity, type_def) - end - end) - - result = if Enum.empty?(violations), do: :ok, else: {:error, violations} - {:reply, result, state} - end - - @impl true - def handle_call({:type_hierarchy, iri}, _from, state) do - hierarchy = compute_hierarchy(iri, state.types, []) - {:reply, hierarchy, state} - end - - # Private Functions - - defp register_builtin_types(state) do - builtin_types = [ - %{ - iri: "verisim:Entity", - label: "Entity", - supertypes: [], - constraints: [] - }, - %{ - iri: "verisim:Document", - label: "Document", - supertypes: ["verisim:Entity"], - constraints: [ - %{name: "title_required", kind: {:required, "title"}, message: "Documents must have a title"} - ] - }, - %{ - iri: "verisim:Node", - label: "Graph Node", - supertypes: ["verisim:Entity"], - constraints: [] - }, - %{ - iri: "verisim:TimeSeries", - label: "Time Series", - supertypes: ["verisim:Entity"], - constraints: [] - } - ] - - new_types = - Enum.reduce(builtin_types, state.types, fn type, acc -> - Map.put(acc, type.iri, type) - end) - - %{state | types: new_types} - end - - defp validate_constraints(entity, type_def) do - type_def.constraints - |> Enum.map(fn constraint -> - case validate_constraint(entity, constraint) do - :ok -> nil - {:error, msg} -> msg - end - end) - |> Enum.reject(&is_nil/1) - end - - defp validate_constraint(entity, %{kind: {:required, property}, message: msg}) do - properties = Map.get(entity, :properties, %{}) - if Map.has_key?(properties, property) do - :ok - else - {:error, msg} - end - end - - defp validate_constraint(entity, %{kind: {:pattern, property, pattern}, message: msg}) do - properties = Map.get(entity, :properties, %{}) - case Map.get(properties, property) do - nil -> :ok - value -> - if Regex.match?(~r/#{pattern}/, value) do - :ok - else - {:error, msg} - end - end - end - - defp validate_constraint(entity, %{kind: {:range, property, min, max}, message: msg}) do - properties = Map.get(entity, :properties, %{}) - case Map.get(properties, property) do - nil -> :ok - value when is_number(value) -> - if (is_nil(min) or value >= min) and (is_nil(max) or value <= max) do - :ok - else - {:error, msg} - end - _ -> :ok - end - end - - defp validate_constraint(_entity, %{kind: kind, message: _msg}) do - {:error, "Unknown constraint type: #{inspect(kind)}"} - end - - defp validate_constraint(_entity, _constraint) do - {:error, "Malformed constraint: missing kind or message"} - end - - defp compute_hierarchy(iri, types, visited) do - if iri in visited do - [] # Prevent cycles - else - case Map.get(types, iri) do - nil -> [] - type_def -> - supertypes = type_def.supertypes || [] - [iri | Enum.flat_map(supertypes, &compute_hierarchy(&1, types, [iri | visited]))] - end - end - end -end diff --git a/verisimdb/elixir-orchestration/lib/verisim/telemetry.ex b/verisimdb/elixir-orchestration/lib/verisim/telemetry.ex deleted file mode 100644 index 209b91a5..00000000 --- a/verisimdb/elixir-orchestration/lib/verisim/telemetry.ex +++ /dev/null @@ -1,230 +0,0 @@ -# SPDX-License-Identifier: MPL-2.0 - -defmodule VeriSim.Telemetry do - @moduledoc """ - Telemetry — metrics, observability, and product development insights for VeriSim. - - Defines telemetry events, metric collectors, and supervises the opt-in - product telemetry collector. The collector aggregates counters and - distributions for query patterns, modality usage, drift frequency, - federation health, and VCL-UT proof adoption. - - ## Opt-in telemetry - - Set `VERISIM_TELEMETRY=true` or `config :verisim, telemetry_enabled: true` - to enable product telemetry collection. Disabled by default. - - All collected data is aggregate-only (counters, rates, distributions). - No query content, entity data, or PII is ever captured. - """ - - use Supervisor - require Logger - - def start_link(arg) do - Supervisor.start_link(__MODULE__, arg, name: __MODULE__) - end - - @impl true - def init(_arg) do - children = [ - {:telemetry_poller, measurements: periodic_measurements(), period: 10_000}, - # Product telemetry collector (opt-in, ETS-backed) - VeriSim.Telemetry.Collector - ] - - Supervisor.init(children, strategy: :one_for_one) - end - - @doc """ - Returns the list of telemetry events emitted by VeriSimDB. - """ - def events do - [ - # Entity events - [:verisim, :entity, :create], - [:verisim, :entity, :update], - [:verisim, :entity, :delete], - - # Query events - [:verisim, :query, :start], - [:verisim, :query, :stop], - [:verisim, :query, :exception], - - # Drift events - [:verisim, :drift, :detected], - [:verisim, :drift, :normalized], - - # Federation events - [:verisim, :federation, :query], - - # Proof events - [:verisim, :proof, :verified], - - # Rust client events - [:verisim, :rust_client, :request], - [:verisim, :rust_client, :response] - ] - end - - # ── Convenience helpers for emitting telemetry events ─────────────────── - - @doc "Emit a query completion event with timing and metadata." - def emit_query_stop(duration_native, metadata \\ %{}) do - :telemetry.execute( - [:verisim, :query, :stop], - %{duration: duration_native}, - metadata - ) - end - - @doc "Emit a query exception event." - def emit_query_exception(metadata \\ %{}) do - :telemetry.execute([:verisim, :query, :exception], %{}, metadata) - end - - @doc "Emit an entity creation event." - def emit_entity_create(metadata \\ %{}) do - :telemetry.execute([:verisim, :entity, :create], %{}, metadata) - end - - @doc "Emit an entity deletion event." - def emit_entity_delete(metadata \\ %{}) do - :telemetry.execute([:verisim, :entity, :delete], %{}, metadata) - end - - @doc "Emit a drift detection event for a specific modality." - def emit_drift_detected(modality, metadata \\ %{}) do - :telemetry.execute( - [:verisim, :drift, :detected], - %{}, - Map.put(metadata, :modality, modality) - ) - end - - @doc "Emit a normalisation event with success/failure status." - def emit_drift_normalized(success?, metadata \\ %{}) do - :telemetry.execute( - [:verisim, :drift, :normalized], - %{success: success?}, - metadata - ) - end - - @doc "Emit a federation query event." - def emit_federation_query(peer, metadata \\ %{}) do - :telemetry.execute( - [:verisim, :federation, :query], - %{}, - Map.put(metadata, :peer, peer) - ) - end - - @doc "Emit a proof verification event." - def emit_proof_verified(proof_type, metadata \\ %{}) do - :telemetry.execute( - [:verisim, :proof, :verified], - %{}, - Map.put(metadata, :proof_type, proof_type) - ) - end - - @doc """ - Returns metric definitions for Telemetry.Metrics. - """ - def metrics do - [ - # Entity metrics - Telemetry.Metrics.counter("verisim.entity.create.count"), - Telemetry.Metrics.counter("verisim.entity.update.count"), - Telemetry.Metrics.counter("verisim.entity.delete.count"), - - # Query metrics - Telemetry.Metrics.counter("verisim.query.count"), - Telemetry.Metrics.distribution("verisim.query.duration", - unit: {:native, :millisecond} - ), - - # Drift metrics - Telemetry.Metrics.last_value("verisim.drift.score"), - Telemetry.Metrics.counter("verisim.drift.detected.count"), - Telemetry.Metrics.counter("verisim.drift.normalized.count"), - - # Rust client metrics - Telemetry.Metrics.counter("verisim.rust_client.request.count"), - Telemetry.Metrics.distribution("verisim.rust_client.response.duration", - unit: {:native, :millisecond} - ), - Telemetry.Metrics.counter("verisim.rust_client.error.count"), - - # System metrics - Telemetry.Metrics.last_value("vm.memory.total", unit: :byte), - Telemetry.Metrics.last_value("vm.total_run_queue_lengths.total"), - Telemetry.Metrics.last_value("vm.system_counts.process_count") - ] - end - - defp periodic_measurements do - [ - # VM measurements - {__MODULE__, :measure_vm_memory, []}, - {__MODULE__, :measure_vm_queues, []}, - {__MODULE__, :measure_vm_processes, []}, - - # VeriSim measurements - {__MODULE__, :measure_entity_count, []}, - {__MODULE__, :measure_drift_status, []} - ] - end - - @doc false - def measure_vm_memory do - memory = :erlang.memory() - :telemetry.execute([:vm, :memory], %{total: memory[:total]}, %{}) - end - - @doc false - def measure_vm_queues do - total = :erlang.statistics(:total_run_queue_lengths) - :telemetry.execute([:vm, :total_run_queue_lengths], %{total: total}, %{}) - end - - @doc false - def measure_vm_processes do - count = :erlang.system_info(:process_count) - :telemetry.execute([:vm, :system_counts], %{process_count: count}, %{}) - end - - @doc false - def measure_entity_count do - # Count entity servers - count = - case Registry.count(VeriSim.EntityRegistry) do - n when is_integer(n) -> n - _ -> 0 - end - - :telemetry.execute([:verisim, :entities], %{count: count}, %{}) - end - - @doc false - def measure_drift_status do - # Get drift status from monitor - case GenServer.whereis(VeriSim.DriftMonitor) do - nil -> :ok - _pid -> - case VeriSim.DriftMonitor.status() do - %{overall_health: health} -> - score = health_to_score(health) - :telemetry.execute([:verisim, :drift], %{score: score}, %{}) - _ -> :ok - end - end - end - - defp health_to_score(:healthy), do: 0.0 - defp health_to_score(:warning), do: 0.3 - defp health_to_score(:degraded), do: 0.6 - defp health_to_score(:critical), do: 0.9 - defp health_to_score(_), do: 0.0 -end diff --git a/verisimdb/elixir-orchestration/lib/verisim/telemetry/collector.ex b/verisimdb/elixir-orchestration/lib/verisim/telemetry/collector.ex deleted file mode 100644 index d55729f4..00000000 --- a/verisimdb/elixir-orchestration/lib/verisim/telemetry/collector.ex +++ /dev/null @@ -1,343 +0,0 @@ -# SPDX-License-Identifier: MPL-2.0 - -defmodule VeriSim.Telemetry.Collector do - @moduledoc """ - ETS-backed telemetry event collector. - - Listens to `:telemetry` events emitted throughout VeriSimDB and aggregates - them into counters, distributions, and rate metrics stored in ETS. This - data powers the product insights reporter and the PanLL telemetry dashboard. - - ## Privacy guarantees - - - **Opt-in only**: telemetry collection must be explicitly enabled via - `VERISIM_TELEMETRY=true` or application config `telemetry_enabled: true` - - **No PII**: never captures query content, entity data, or user identifiers - - **Aggregate only**: stores counts, sums, min/max — never individual records - - **Local first**: all data stays on the machine until explicitly exported - - ## Collected metrics - - | Metric | Type | Source | - |--------|------|--------| - | query_count | counter | VCL executor | - | query_duration_sum | sum | VCL executor | - | query_duration_min/max | gauge | VCL executor | - | modality_usage | counter map | VCL executor / query router | - | query_pattern | counter map | VCL executor (SELECT/INSERT/DELETE/SEARCH/SHOW) | - | drift_detected_count | counter | drift monitor | - | drift_modality_breakdown | counter map | drift monitor | - | normalise_count | counter | normaliser | - | normalise_success_count | counter | normaliser | - | federation_query_count | counter | federation resolver | - | federation_peer_errors | counter map | federation resolver | - | proof_type_usage | counter map | VCL-UT executor | - | entity_created_count | counter | entity server | - | entity_deleted_count | counter | entity server | - """ - - use GenServer - require Logger - - @table :verisim_telemetry - @reset_interval :timer.hours(24) - @health_snapshot_interval :timer.seconds(10) - @error_budget_reset_interval :timer.hours(1) - - # ── Public API ────────────────────────────────────────────────────────── - - @doc "Start the collector, optionally linked to a supervisor." - def start_link(opts \\ []) do - GenServer.start_link(__MODULE__, opts, name: __MODULE__) - end - - @doc "Check whether telemetry collection is enabled." - def enabled? do - Application.get_env(:verisim, :telemetry_enabled, false) || - System.get_env("VERISIM_TELEMETRY") == "true" - end - - @doc "Return a snapshot of all collected metrics as a map." - def snapshot do - if :ets.whereis(@table) != :undefined do - @table - |> :ets.tab2list() - |> Enum.into(%{}) - else - %{} - end - end - - @doc "Increment a counter metric by the given amount (default 1)." - def increment(key, amount \\ 1) do - if enabled?() and :ets.whereis(@table) != :undefined do - :ets.update_counter(@table, key, {2, amount}, {key, 0}) - end - end - - @doc "Record a distribution value (tracks count, sum, min, max)." - def record_distribution(key, value) when is_number(value) do - if enabled?() and :ets.whereis(@table) != :undefined do - GenServer.cast(__MODULE__, {:record_distribution, key, value}) - end - end - - @doc "Increment a counter within a counter map (e.g., modality usage)." - def increment_map(map_key, sub_key, amount \\ 1) do - if enabled?() and :ets.whereis(@table) != :undefined do - composite_key = {map_key, sub_key} - :ets.update_counter(@table, composite_key, {2, amount}, {composite_key, 0}) - end - end - - @doc "Reset all collected metrics. Used for testing or periodic rollover." - def reset do - if :ets.whereis(@table) != :undefined do - :ets.delete_all_objects(@table) - :ok - end - end - - @doc "Return the collection start time (when metrics were last reset)." - def collection_start do - GenServer.call(__MODULE__, :collection_start) - end - - # ── GenServer callbacks ───────────────────────────────────────────────── - - @impl true - def init(_opts) do - table = :ets.new(@table, [:named_table, :public, :set, read_concurrency: true]) - - if enabled?() do - attach_handlers() - Logger.info("[VeriSim.Telemetry.Collector] Telemetry collection ENABLED") - else - Logger.debug("[VeriSim.Telemetry.Collector] Telemetry collection disabled (opt-in)") - end - - # Schedule periodic reset to prevent unbounded growth. - Process.send_after(self(), :periodic_reset, @reset_interval) - - # Schedule periodic health snapshots (every 10s) - if enabled?() do - Process.send_after(self(), :health_snapshot, @health_snapshot_interval) - Process.send_after(self(), :error_budget_reset, @error_budget_reset_interval) - end - - {:ok, %{table: table, started_at: DateTime.utc_now()}} - end - - @impl true - def handle_cast({:record_distribution, key, value}, state) do - count_key = {key, :count} - sum_key = {key, :sum} - min_key = {key, :min} - max_key = {key, :max} - - :ets.update_counter(@table, count_key, {2, 1}, {count_key, 0}) - :ets.update_counter(@table, sum_key, {2, trunc(value * 1000)}, {sum_key, 0}) - - # Min/max require compare-and-swap - case :ets.lookup(@table, min_key) do - [{^min_key, current}] when value < current / 1000 -> - :ets.insert(@table, {min_key, trunc(value * 1000)}) - [] -> - :ets.insert(@table, {min_key, trunc(value * 1000)}) - _ -> :ok - end - - case :ets.lookup(@table, max_key) do - [{^max_key, current}] when value > current / 1000 -> - :ets.insert(@table, {max_key, trunc(value * 1000)}) - [] -> - :ets.insert(@table, {max_key, trunc(value * 1000)}) - _ -> :ok - end - - {:noreply, state} - end - - @impl true - def handle_call(:collection_start, _from, state) do - {:reply, state.started_at, state} - end - - @impl true - def handle_info(:periodic_reset, state) do - if enabled?() do - Logger.info("[VeriSim.Telemetry.Collector] Periodic metric reset (24h rollover)") - :ets.delete_all_objects(@table) - end - - Process.send_after(self(), :periodic_reset, @reset_interval) - {:noreply, %{state | started_at: DateTime.utc_now()}} - end - - def handle_info(:health_snapshot, state) do - capture_health_snapshot() - Process.send_after(self(), :health_snapshot, @health_snapshot_interval) - {:noreply, state} - end - - def handle_info(:error_budget_reset, state) do - if enabled?() and :ets.whereis(@table) != :undefined do - # Reset error budget counters (hourly rolling window) - :ets.insert(@table, {:error_budget_total, 0}) - - @table - |> :ets.tab2list() - |> Enum.filter(fn - {{:error_budget_by_type, _}, _} -> true - _ -> false - end) - |> Enum.each(fn {key, _} -> :ets.insert(@table, {key, 0}) end) - end - - Process.send_after(self(), :error_budget_reset, @error_budget_reset_interval) - {:noreply, state} - end - - # ── Health metrics snapshot ──────────────────────────────────────────── - - @doc """ - Capture a point-in-time health snapshot: memory, process count, uptime. - - Called periodically (every 10s) by the collector GenServer to track - system health without any PII or content data. - """ - def capture_health_snapshot do - if enabled?() and :ets.whereis(@table) != :undefined do - # Memory metrics from :erlang.memory/0 (in bytes) - memory = :erlang.memory() - :ets.insert(@table, {:health_memory_total_bytes, memory[:total]}) - :ets.insert(@table, {:health_memory_processes_bytes, memory[:processes]}) - :ets.insert(@table, {:health_memory_ets_bytes, memory[:ets]}) - :ets.insert(@table, {:health_memory_binary_bytes, memory[:binary]}) - - # Process count - :ets.insert(@table, {:health_process_count, length(Process.list())}) - - # Uptime in seconds - {uptime_ms, _} = :erlang.statistics(:wall_clock) - :ets.insert(@table, {:health_uptime_seconds, div(uptime_ms, 1000)}) - - # Scheduler utilisation (1-second sample) - :ets.insert(@table, {:health_scheduler_count, :erlang.system_info(:schedulers_online)}) - - # Timestamp of last health check - :ets.insert(@table, {:health_last_checked, DateTime.utc_now() |> DateTime.to_iso8601()}) - end - end - - @doc """ - Record an error by type for error budget tracking. - - Error budget resets hourly. Tracks: timeout, connection_error, - parse_error, proof_failure, federation_error, internal_error. - """ - def record_error(error_type) when is_atom(error_type) do - increment(:error_budget_total) - increment_map(:error_budget_by_type, error_type) - end - - # ── Telemetry event handlers ──────────────────────────────────────────── - - defp attach_handlers do - events = [ - # Query events - {[:verisim, :query, :stop], &__MODULE__.handle_query_stop/4}, - {[:verisim, :query, :exception], &__MODULE__.handle_query_exception/4}, - - # Entity events - {[:verisim, :entity, :create], &__MODULE__.handle_entity_create/4}, - {[:verisim, :entity, :delete], &__MODULE__.handle_entity_delete/4}, - - # Drift events - {[:verisim, :drift, :detected], &__MODULE__.handle_drift_detected/4}, - {[:verisim, :drift, :normalized], &__MODULE__.handle_drift_normalized/4}, - - # Federation events - {[:verisim, :federation, :query], &__MODULE__.handle_federation_query/4}, - - # Proof events - {[:verisim, :proof, :verified], &__MODULE__.handle_proof_verified/4} - ] - - Enum.each(events, fn {event, handler} -> - handler_id = "verisim_collector_#{Enum.join(event, "_")}" - :telemetry.attach(handler_id, event, handler, nil) - end) - end - - @doc false - def handle_query_stop(_event, measurements, metadata, _config) do - increment(:query_count) - - if duration = Map.get(measurements, :duration) do - duration_ms = System.convert_time_unit(duration, :native, :millisecond) - record_distribution(:query_duration, duration_ms) - end - - if pattern = Map.get(metadata, :statement_type) do - increment_map(:query_pattern, pattern) - end - - if modalities = Map.get(metadata, :modalities) do - Enum.each(List.wrap(modalities), fn modality -> - increment_map(:modality_usage, modality) - end) - end - end - - @doc false - def handle_query_exception(_event, _measurements, _metadata, _config) do - increment(:query_error_count) - end - - @doc false - def handle_entity_create(_event, _measurements, _metadata, _config) do - increment(:entity_created_count) - end - - @doc false - def handle_entity_delete(_event, _measurements, _metadata, _config) do - increment(:entity_deleted_count) - end - - @doc false - def handle_drift_detected(_event, _measurements, metadata, _config) do - increment(:drift_detected_count) - - if modality = Map.get(metadata, :modality) do - increment_map(:drift_modality_breakdown, modality) - end - end - - @doc false - def handle_drift_normalized(_event, measurements, _metadata, _config) do - increment(:normalise_count) - - if Map.get(measurements, :success, false) do - increment(:normalise_success_count) - end - end - - @doc false - def handle_federation_query(_event, _measurements, metadata, _config) do - increment(:federation_query_count) - - if peer = Map.get(metadata, :peer) do - if Map.get(metadata, :error) do - increment_map(:federation_peer_errors, peer) - end - end - end - - @doc false - def handle_proof_verified(_event, _measurements, metadata, _config) do - if proof_type = Map.get(metadata, :proof_type) do - increment_map(:proof_type_usage, proof_type) - end - end -end diff --git a/verisimdb/elixir-orchestration/lib/verisim/telemetry/reporter.ex b/verisimdb/elixir-orchestration/lib/verisim/telemetry/reporter.ex deleted file mode 100644 index f6d87167..00000000 --- a/verisimdb/elixir-orchestration/lib/verisim/telemetry/reporter.ex +++ /dev/null @@ -1,350 +0,0 @@ -# SPDX-License-Identifier: MPL-2.0 - -defmodule VeriSim.Telemetry.Reporter do - @moduledoc """ - Telemetry reporter — aggregates raw collector metrics into structured - product development insights and exports them as JSON. - - ## Insight categories - - 1. **Modality Heatmap** — which of the 8 octad modalities see real usage - 2. **Query Pattern Distribution** — read vs. write vs. drift vs. proof queries - 3. **Performance Summary** — query latency percentiles (from distribution) - 4. **Drift Report** — frequency, modality breakdown, normalisation success rate - 5. **Federation Health** — peer error rates - 6. **VCL-UT Adoption** — proof type usage distribution - - ## Privacy guarantees - - All data returned by this module is aggregate-only. No query content, entity - data, or personally identifiable information is ever included. The reporter - reads from the collector's ETS table, which itself only stores counters and - distribution summaries. - - ## Usage - - # Full report as map - VeriSim.Telemetry.Reporter.report() - - # JSON string for HTTP endpoint or PanLL - VeriSim.Telemetry.Reporter.report_json() - - # Individual insight - VeriSim.Telemetry.Reporter.modality_heatmap() - """ - - alias VeriSim.Telemetry.Collector - - @modalities ~w(graph vector tensor semantic document temporal provenance spatial)a - - @doc """ - Generate a full product insights report as a structured map. - - Returns a map with keys: `:meta`, `:modality_heatmap`, `:query_patterns`, - `:performance`, `:drift`, `:federation`, `:proof_types`. - """ - def report do - snapshot = Collector.snapshot() - started_at = safe_collection_start() - - %{ - meta: %{ - generated_at: DateTime.utc_now() |> DateTime.to_iso8601(), - collection_started: started_at |> DateTime.to_iso8601(), - telemetry_enabled: Collector.enabled?(), - privacy_notice: - "This report contains aggregate metrics only. " <> - "No query content, entity data, or PII is included." - }, - modality_heatmap: modality_heatmap(snapshot), - query_patterns: query_patterns(snapshot), - performance: performance_summary(snapshot), - drift: drift_report(snapshot), - federation: federation_health(snapshot), - proof_types: proof_type_usage(snapshot), - entities: entity_summary(snapshot), - health: health_report(snapshot), - error_budget: error_budget_report(snapshot) - } - end - - @doc "Generate the full report as a JSON string." - def report_json do - report() |> Jason.encode!(pretty: true) - end - - @doc """ - Modality heatmap — shows how much each of the 8 octad modalities is used - in queries. Helps identify which modalities are core to users vs. underused. - """ - def modality_heatmap(snapshot \\ nil) do - snapshot = snapshot || Collector.snapshot() - - usage_map = - @modalities - |> Enum.map(fn modality -> - count = Map.get(snapshot, {:modality_usage, modality}, 0) - {modality, count} - end) - |> Enum.into(%{}) - - total = Enum.sum(Map.values(usage_map)) - - percentages = - if total > 0 do - usage_map - |> Enum.map(fn {mod, count} -> - {mod, Float.round(count / total * 100, 1)} - end) - |> Enum.into(%{}) - else - usage_map |> Enum.map(fn {mod, _} -> {mod, 0.0} end) |> Enum.into(%{}) - end - - %{ - counts: usage_map, - percentages: percentages, - total_modality_queries: total, - most_used: most_used_modality(usage_map), - least_used: least_used_modality(usage_map) - } - end - - @doc """ - Query pattern distribution — what kinds of queries are being run. - Helps understand read-heavy vs. write-heavy vs. analytics workloads. - """ - def query_patterns(snapshot \\ nil) do - snapshot = snapshot || Collector.snapshot() - - patterns = - snapshot - |> Enum.filter(fn - {{:query_pattern, _}, _} -> true - _ -> false - end) - |> Enum.map(fn {{:query_pattern, pattern}, count} -> {pattern, count} end) - |> Enum.into(%{}) - - total = Map.get(snapshot, :query_count, 0) - errors = Map.get(snapshot, :query_error_count, 0) - - %{ - total_queries: total, - error_count: errors, - error_rate: if(total > 0, do: Float.round(errors / total * 100, 2), else: 0.0), - by_type: patterns - } - end - - @doc """ - Performance summary — query latency statistics derived from the - distribution tracker in the collector. Shows count, average, min, max. - """ - def performance_summary(snapshot \\ nil) do - snapshot = snapshot || Collector.snapshot() - - count = Map.get(snapshot, {:query_duration, :count}, 0) - sum_millis = Map.get(snapshot, {:query_duration, :sum}, 0) / 1000 - min_millis = Map.get(snapshot, {:query_duration, :min}, 0) / 1000 - max_millis = Map.get(snapshot, {:query_duration, :max}, 0) / 1000 - - avg = if count > 0, do: Float.round(sum_millis / count, 2), else: 0.0 - - %{ - query_count: count, - avg_duration_ms: avg, - min_duration_ms: Float.round(min_millis, 2), - max_duration_ms: Float.round(max_millis, 2), - total_duration_ms: Float.round(sum_millis, 2) - } - end - - @doc """ - Drift report — how often drift is detected, which modalities drift most, - and how successful normalisation is at repairing it. - """ - def drift_report(snapshot \\ nil) do - snapshot = snapshot || Collector.snapshot() - - detected = Map.get(snapshot, :drift_detected_count, 0) - normalised = Map.get(snapshot, :normalise_count, 0) - normalise_success = Map.get(snapshot, :normalise_success_count, 0) - - modality_breakdown = - snapshot - |> Enum.filter(fn - {{:drift_modality_breakdown, _}, _} -> true - _ -> false - end) - |> Enum.map(fn {{:drift_modality_breakdown, mod}, count} -> {mod, count} end) - |> Enum.into(%{}) - - success_rate = - if normalised > 0, do: Float.round(normalise_success / normalised * 100, 1), else: 0.0 - - %{ - drift_detected_count: detected, - normalise_attempts: normalised, - normalise_success_count: normalise_success, - normalise_success_rate: success_rate, - modality_breakdown: modality_breakdown, - most_drifted: most_drifted_modality(modality_breakdown) - } - end - - @doc """ - Federation health — tracks errors per federated peer to identify - unreliable backends. - """ - def federation_health(snapshot \\ nil) do - snapshot = snapshot || Collector.snapshot() - - total = Map.get(snapshot, :federation_query_count, 0) - - peer_errors = - snapshot - |> Enum.filter(fn - {{:federation_peer_errors, _}, _} -> true - _ -> false - end) - |> Enum.map(fn {{:federation_peer_errors, peer}, count} -> {peer, count} end) - |> Enum.into(%{}) - - %{ - total_federation_queries: total, - peer_errors: peer_errors - } - end - - @doc """ - Proof type usage — which VCL-UT proof types are used. Tracks adoption - of dependent type features. - """ - def proof_type_usage(snapshot \\ nil) do - snapshot = snapshot || Collector.snapshot() - - types = - snapshot - |> Enum.filter(fn - {{:proof_type_usage, _}, _} -> true - _ -> false - end) - |> Enum.map(fn {{:proof_type_usage, proof_type}, count} -> {proof_type, count} end) - |> Enum.into(%{}) - - total = Enum.sum(Map.values(types)) - - %{ - total_proofs: total, - by_type: types, - vcl_dt_active: total > 0 - } - end - - @doc "Entity creation/deletion summary." - def entity_summary(snapshot \\ nil) do - snapshot = snapshot || Collector.snapshot() - - %{ - created: Map.get(snapshot, :entity_created_count, 0), - deleted: Map.get(snapshot, :entity_deleted_count, 0) - } - end - - @doc """ - Health report — system resource usage and uptime. Uses data from the - periodic health snapshot (captured every 10 seconds by the collector). - - All values are aggregate system metrics — no PII, no query content. - """ - def health_report(snapshot \\ nil) do - snapshot = snapshot || Collector.snapshot() - - %{ - memory: %{ - total_bytes: Map.get(snapshot, :health_memory_total_bytes, 0), - processes_bytes: Map.get(snapshot, :health_memory_processes_bytes, 0), - ets_bytes: Map.get(snapshot, :health_memory_ets_bytes, 0), - binary_bytes: Map.get(snapshot, :health_memory_binary_bytes, 0), - total_mb: Float.round(Map.get(snapshot, :health_memory_total_bytes, 0) / 1_048_576, 1) - }, - process_count: Map.get(snapshot, :health_process_count, 0), - uptime_seconds: Map.get(snapshot, :health_uptime_seconds, 0), - scheduler_count: Map.get(snapshot, :health_scheduler_count, 0), - last_checked: Map.get(snapshot, :health_last_checked, nil) - } - end - - @doc """ - Error budget report — tracks error counts by type within a rolling - 1-hour window. Resets hourly. Used for reliability monitoring and SLO - tracking. - - Error types: timeout, connection_error, parse_error, proof_failure, - federation_error, internal_error. - """ - def error_budget_report(snapshot \\ nil) do - snapshot = snapshot || Collector.snapshot() - - total_errors = Map.get(snapshot, :error_budget_total, 0) - total_queries = Map.get(snapshot, :query_count, 0) - - by_type = - snapshot - |> Enum.filter(fn - {{:error_budget_by_type, _}, _} -> true - _ -> false - end) - |> Enum.map(fn {{:error_budget_by_type, error_type}, count} -> {error_type, count} end) - |> Enum.into(%{}) - - error_rate = - if total_queries > 0 do - Float.round(total_errors / total_queries * 100, 3) - else - 0.0 - end - - %{ - total_errors: total_errors, - error_rate_percent: error_rate, - by_type: by_type, - budget_window: "1 hour (rolling)" - } - end - - # ── Private helpers ───────────────────────────────────────────────────── - - defp most_used_modality(usage_map) do - case Enum.max_by(usage_map, fn {_, count} -> count end, fn -> {nil, 0} end) do - {nil, _} -> nil - {mod, 0} -> if Enum.all?(usage_map, fn {_, c} -> c == 0 end), do: nil, else: mod - {mod, _} -> mod - end - end - - defp least_used_modality(usage_map) do - nonzero = Enum.filter(usage_map, fn {_, count} -> count > 0 end) - - case nonzero do - [] -> nil - list -> list |> Enum.min_by(fn {_, count} -> count end) |> elem(0) - end - end - - defp most_drifted_modality(breakdown) do - case Enum.max_by(breakdown, fn {_, count} -> count end, fn -> {nil, 0} end) do - {nil, _} -> nil - {mod, _} -> mod - end - end - - defp safe_collection_start do - try do - Collector.collection_start() - catch - :exit, _ -> DateTime.utc_now() - end - end -end diff --git a/verisimdb/elixir-orchestration/lib/verisim/transport.ex b/verisimdb/elixir-orchestration/lib/verisim/transport.ex deleted file mode 100644 index 0b1eb9f3..00000000 --- a/verisimdb/elixir-orchestration/lib/verisim/transport.ex +++ /dev/null @@ -1,195 +0,0 @@ -# SPDX-License-Identifier: MPL-2.0 - -defmodule VeriSim.Transport do - @moduledoc """ - Transport selection for VeriSimDB Rust core communication. - - Selects between HTTP (verisim-api) and NIF (in-process) transports based on - the `VERISIM_TRANSPORT` environment variable: - - VERISIM_TRANSPORT=http # Default: HTTP via VeriSim.RustClient - VERISIM_TRANSPORT=nif # Direct NIF calls via VeriSim.NifBridge - VERISIM_TRANSPORT=auto # NIF if loaded, HTTP fallback - - ## Usage - - Higher-level modules (EntityServer, DriftMonitor, QueryRouter) should call - `VeriSim.Transport` instead of `VeriSim.RustClient` directly. The transport - module delegates to the appropriate backend transparently. - - ## Architecture - - ┌──────────────────────┐ - │ VeriSim.Transport │ ← Unified interface - ├──────────┬───────────┤ - │ NifBridge│ RustClient│ ← Backend selection - │ (in-proc)│ (HTTP) │ - └──────────┴───────────┘ - """ - - require Logger - - alias VeriSim.{NifBridge, RustClient} - - @doc """ - Returns the active transport mode: `:http`, `:nif`, or `:auto`. - """ - def mode do - case System.get_env("VERISIM_TRANSPORT", "http") do - "nif" -> :nif - "auto" -> :auto - _ -> :http - end - end - - @doc """ - Returns true if the NIF bridge is loaded and operational. - """ - def nif_available? do - try do - NifBridge.loaded?() - rescue - _ -> false - end - end - - @doc """ - Determine whether to use NIF for this call based on the transport mode. - """ - def use_nif? do - case mode() do - :nif -> true - :auto -> nif_available?() - :http -> false - end - end - - # --------------------------------------------------------------------------- - # Delegated Operations - # --------------------------------------------------------------------------- - - @doc """ - Check health of the Rust core. - """ - def health do - if use_nif?() do - {:ok, %{"status" => "healthy", "transport" => "nif"}} - else - RustClient.health() - end - end - - @doc """ - Create a new octad entity. - """ - def create_octad(input) do - if use_nif?() do - json = Jason.encode!(input) - case NifBridge.create_octad(json) do - result when is_binary(result) -> {:ok, Jason.decode!(result)} - {:error, reason} -> {:error, reason} - end - else - RustClient.create_octad(input) - end - end - - @doc """ - Get a octad by ID. - """ - def get_octad(entity_id) do - if use_nif?() do - case NifBridge.get_octad(entity_id) do - result when is_binary(result) -> {:ok, Jason.decode!(result)} - {:error, reason} -> {:error, reason} - end - else - RustClient.get_octad(entity_id) - end - end - - @doc """ - Delete a octad entity. - """ - def delete_octad(entity_id) do - if use_nif?() do - case NifBridge.delete_octad(entity_id) do - result when is_binary(result) -> :ok - {:error, reason} -> {:error, reason} - end - else - RustClient.delete_octad(entity_id) - end - end - - @doc """ - Full-text search across the document modality. - """ - def search_text(query, limit \\ 10) do - if use_nif?() do - case NifBridge.search_text(query, limit) do - result when is_binary(result) -> {:ok, Jason.decode!(result)} - {:error, reason} -> {:error, reason} - end - else - RustClient.search_text(query, limit) - end - end - - @doc """ - Vector similarity search. - """ - def search_vector(vector, k \\ 10) do - if use_nif?() do - embedding_json = Jason.encode!(vector) - case NifBridge.search_vector(embedding_json, k) do - result when is_binary(result) -> {:ok, Jason.decode!(result)} - {:error, reason} -> {:error, reason} - end - else - RustClient.search_vector(vector, k) - end - end - - @doc """ - Paginated listing of octad entities. - """ - def list_octads(limit \\ 50, offset \\ 0) do - if use_nif?() do - case NifBridge.list_octads(limit, offset) do - result when is_binary(result) -> {:ok, Jason.decode!(result)} - {:error, reason} -> {:error, reason} - end - else - RustClient.list_octads(limit, offset) - end - end - - @doc """ - Get drift scores for a specific entity. - """ - def get_drift_score(entity_id) do - if use_nif?() do - case NifBridge.get_drift_score(entity_id) do - result when is_binary(result) -> {:ok, Jason.decode!(result)} - {:error, reason} -> {:error, reason} - end - else - RustClient.get_drift_score(entity_id) - end - end - - @doc """ - Trigger normalisation for a drifted entity. - """ - def trigger_normalise(entity_id) do - if use_nif?() do - case NifBridge.trigger_normalise(entity_id) do - result when is_binary(result) -> {:ok, Jason.decode!(result)} - {:error, reason} -> {:error, reason} - end - else - RustClient.trigger_normalization(entity_id) - end - end -end diff --git a/verisimdb/elixir-orchestration/mix.exs b/verisimdb/elixir-orchestration/mix.exs deleted file mode 100644 index af2c1630..00000000 --- a/verisimdb/elixir-orchestration/mix.exs +++ /dev/null @@ -1,87 +0,0 @@ -# SPDX-License-Identifier: MPL-2.0 - -defmodule VeriSim.MixProject do - use Mix.Project - - def project do - [ - app: :verisim, - version: "0.1.0", - elixir: "~> 1.17", - start_permanent: Mix.env() == :prod, - elixirc_paths: elixirc_paths(Mix.env()), - deps: deps(), - aliases: aliases(), - releases: [ - verisim: [ - include_executables_for: [:unix], - applications: [runtime_tools: :permanent] - ] - ], - - # Docs - name: "VeriSim Orchestration", - source_url: "https://gitlab.com/hyperpolymath/verisimdb", - docs: [ - main: "VeriSim", - extras: ["README.md"] - ] - ] - end - - def application do - [ - extra_applications: [:logger], - mod: {VeriSim.Application, []} - ] - end - - defp elixirc_paths(:test), do: ["lib", "test/support"] - defp elixirc_paths(_), do: ["lib"] - - defp deps do - [ - # HTTP client for Rust core communication - {:req, "~> 0.5"}, - - # HTTP server for orchestration API (telemetry, status) - {:bandit, "~> 1.6"}, - - # JSON encoding/decoding - {:jason, "~> 1.4"}, - - # Telemetry and metrics - {:telemetry, "~> 1.2"}, - {:telemetry_metrics, "~> 1.0"}, - {:telemetry_poller, "~> 1.0"}, - - # Process registry - {:horde, "~> 0.9"}, - - # Testing - {:ex_machina, "~> 2.7", only: :test}, - {:mox, "~> 1.0", only: :test}, - {:stream_data, "~> 1.1", only: [:test, :dev]}, - - # Optional: native protocol adapters for federation - # These are only needed when using :wire protocol instead of HTTP - {:postgrex, "~> 0.19", optional: true}, - {:redix, "~> 1.5", optional: true}, - {:exqlite, "~> 0.27", optional: true}, - {:bolt_sips, "~> 2.0", optional: true}, - - # Development - {:ex_doc, "~> 0.34", only: :dev, runtime: false}, - {:credo, "~> 1.7", only: [:dev, :test], runtime: false}, - {:dialyxir, "~> 1.4", only: [:dev, :test], runtime: false} - ] - end - - defp aliases do - [ - setup: ["deps.get"], - test: ["test"], - "test.watch": ["test.watch"] - ] - end -end diff --git a/verisimdb/elixir-orchestration/test/fixtures/vcl/basic-select.expected.json b/verisimdb/elixir-orchestration/test/fixtures/vcl/basic-select.expected.json deleted file mode 100644 index 444ebd40..00000000 --- a/verisimdb/elixir-orchestration/test/fixtures/vcl/basic-select.expected.json +++ /dev/null @@ -1,9 +0,0 @@ -{ - "type": "query", - "modalities": ["all"], - "source": {"type": "octad_store"}, - "limit": 10, - "offset": null, - "where": null, - "proof": null -} diff --git a/verisimdb/elixir-orchestration/test/fixtures/vcl/basic-select.vcl b/verisimdb/elixir-orchestration/test/fixtures/vcl/basic-select.vcl deleted file mode 100644 index 3142a9a4..00000000 --- a/verisimdb/elixir-orchestration/test/fixtures/vcl/basic-select.vcl +++ /dev/null @@ -1 +0,0 @@ -SELECT * FROM hexads LIMIT 10 diff --git a/verisimdb/elixir-orchestration/test/fixtures/vcl/drift-query.expected.json b/verisimdb/elixir-orchestration/test/fixtures/vcl/drift-query.expected.json deleted file mode 100644 index 70bec76b..00000000 --- a/verisimdb/elixir-orchestration/test/fixtures/vcl/drift-query.expected.json +++ /dev/null @@ -1,9 +0,0 @@ -{ - "type": "query", - "modalities": ["all"], - "source": {"type": "octad_store"}, - "where": {"field": "drift_score", "op": ">", "value": 0.3}, - "order_by": {"field": "drift_score", "direction": "DESC"}, - "limit": 20, - "proof": null -} diff --git a/verisimdb/elixir-orchestration/test/fixtures/vcl/drift-query.vcl b/verisimdb/elixir-orchestration/test/fixtures/vcl/drift-query.vcl deleted file mode 100644 index 95c52b37..00000000 --- a/verisimdb/elixir-orchestration/test/fixtures/vcl/drift-query.vcl +++ /dev/null @@ -1 +0,0 @@ -SELECT id, drift_score FROM hexads WHERE drift_score > 0.3 ORDER BY drift_score DESC LIMIT 20 diff --git a/verisimdb/elixir-orchestration/test/fixtures/vcl/multi-proof.expected.json b/verisimdb/elixir-orchestration/test/fixtures/vcl/multi-proof.expected.json deleted file mode 100644 index 017a7f3f..00000000 --- a/verisimdb/elixir-orchestration/test/fixtures/vcl/multi-proof.expected.json +++ /dev/null @@ -1,8 +0,0 @@ -{ - "type": "query", - "modalities": ["all"], - "source": {"type": "octad_store"}, - "where": {"field": "id", "op": "=", "value": "entity-001"}, - "proof": {"raw": "EXISTENCE(entity-001) AND PROVENANCE(entity-001)"}, - "limit": null -} diff --git a/verisimdb/elixir-orchestration/test/fixtures/vcl/multi-proof.vcl b/verisimdb/elixir-orchestration/test/fixtures/vcl/multi-proof.vcl deleted file mode 100644 index 3d05748b..00000000 --- a/verisimdb/elixir-orchestration/test/fixtures/vcl/multi-proof.vcl +++ /dev/null @@ -1 +0,0 @@ -SELECT * FROM hexads WHERE id = 'entity-001' PROOF EXISTENCE(entity-001) AND PROVENANCE(entity-001) diff --git a/verisimdb/elixir-orchestration/test/fixtures/vcl/proof-existence.expected.json b/verisimdb/elixir-orchestration/test/fixtures/vcl/proof-existence.expected.json deleted file mode 100644 index 7f29adba..00000000 --- a/verisimdb/elixir-orchestration/test/fixtures/vcl/proof-existence.expected.json +++ /dev/null @@ -1,8 +0,0 @@ -{ - "type": "query", - "modalities": ["graph", "semantic"], - "source": {"type": "octad_store"}, - "where": {"field": "id", "op": "=", "value": "entity-001"}, - "proof": {"raw": "EXISTENCE(entity-001)"}, - "limit": null -} diff --git a/verisimdb/elixir-orchestration/test/fixtures/vcl/proof-existence.vcl b/verisimdb/elixir-orchestration/test/fixtures/vcl/proof-existence.vcl deleted file mode 100644 index c7ed839d..00000000 --- a/verisimdb/elixir-orchestration/test/fixtures/vcl/proof-existence.vcl +++ /dev/null @@ -1 +0,0 @@ -SELECT graph, semantic FROM hexads WHERE id = 'entity-001' PROOF EXISTENCE(entity-001) diff --git a/verisimdb/elixir-orchestration/test/fixtures/vcl/vector-search.expected.json b/verisimdb/elixir-orchestration/test/fixtures/vcl/vector-search.expected.json deleted file mode 100644 index ce68dd93..00000000 --- a/verisimdb/elixir-orchestration/test/fixtures/vcl/vector-search.expected.json +++ /dev/null @@ -1,8 +0,0 @@ -{ - "type": "query", - "modalities": ["vector"], - "source": {"type": "vector_search", "vector": [0.1, 0.2, 0.3, 0.4]}, - "limit": 5, - "where": null, - "proof": null -} diff --git a/verisimdb/elixir-orchestration/test/fixtures/vcl/vector-search.vcl b/verisimdb/elixir-orchestration/test/fixtures/vcl/vector-search.vcl deleted file mode 100644 index 949e4540..00000000 --- a/verisimdb/elixir-orchestration/test/fixtures/vcl/vector-search.vcl +++ /dev/null @@ -1 +0,0 @@ -SEARCH VECTOR SIMILAR TO [0.1, 0.2, 0.3, 0.4] LIMIT 5 diff --git a/verisimdb/elixir-orchestration/test/integration_test.exs b/verisimdb/elixir-orchestration/test/integration_test.exs deleted file mode 100644 index dd0ec31b..00000000 --- a/verisimdb/elixir-orchestration/test/integration_test.exs +++ /dev/null @@ -1,352 +0,0 @@ -# SPDX-License-Identifier: MPL-2.0 - -defmodule VeriSim.IntegrationTest do - @moduledoc """ - Integration tests for the full VeriSimDB stack. - - Tests: - - Elixir orchestration layer - - Communication with Rust core - - VCL query execution - - Drift detection and normalization - - Cross-modal queries - """ - - use ExUnit.Case, async: false - - alias VeriSim.{ - QueryRouter, - RustClient, - DriftMonitor, - SchemaRegistry, - EntityServer, - Query.VCLExecutor - } - - @moduletag :integration - - setup do - # Start the application - {:ok, _} = Application.ensure_all_started(:verisim) - - # Wait for Rust core to be ready - wait_for_rust_core() - - :ok - end - - describe "Rust Client" do - test "health check" do - case RustClient.health() do - {:ok, _health} -> - assert true - - {:error, _reason} -> - # Rust core may not be running, skip test - :ok - end - end - - test "create and get octad" do - input = %{ - title: "Test Octad", - body: "Integration test octad", - embedding: List.duplicate(0.5, 384) - } - - case RustClient.create_octad(input) do - {:ok, %{"id" => entity_id}} -> - assert is_binary(entity_id) - - case RustClient.get_octad(entity_id) do - {:ok, octad} -> - assert octad["id"] == entity_id - assert octad["document"]["title"] == "Test Octad" - - {:error, :not_found} -> - flunk("Octad should exist after creation") - - {:error, reason} -> - flunk("Failed to get octad: #{inspect(reason)}") - end - - {:error, _reason} -> - # Rust core not available - :ok - end - end - end - - describe "Query Router" do - test "routes text queries" do - stats = QueryRouter.stats() - assert is_map(stats) - assert Map.has_key?(stats, :total_queries) - end - - test "handles multi-modal queries" do - params = %{ - text: "machine learning", - types: ["https://example.org/Document"] - } - - result = QueryRouter.query(:multi, params, limit: 10) - - case result do - {:ok, _results} -> assert true - {:error, _} -> :ok - end - end - end - - describe "Drift Monitor" do - test "tracks drift events" do - entity_id = "test-entity-#{:rand.uniform(1000)}" - - DriftMonitor.report_drift(entity_id, 0.5, :semantic_vector) - DriftMonitor.entity_changed(entity_id) - - # Give it time to process - Process.sleep(100) - - status = DriftMonitor.status() - assert status[:overall_health] in [:healthy, :warning, :degraded, :critical] - assert is_integer(status[:entities_with_drift]) - end - - test "triggers normalization on critical drift" do - entity_id = "critical-drift-#{:rand.uniform(1000)}" - - # Report critical drift - DriftMonitor.report_drift(entity_id, 0.9, :quality) - - Process.sleep(200) - - status = DriftMonitor.status() - # Normalization might have been triggered - assert is_integer(status[:pending_normalizations]) - end - end - - describe "Schema Registry" do - test "registers and retrieves types" do - type_def = %{ - iri: "https://test.example.org/TestType", - label: "Test Type", - supertypes: ["verisim:Entity"], - constraints: [ - %{ - name: "title_required", - kind: {:required, "title"}, - message: "Title is required" - } - ] - } - - assert :ok == SchemaRegistry.register_type(type_def) - - retrieved = SchemaRegistry.get_type("https://test.example.org/TestType") - assert retrieved != nil - assert retrieved.label == "Test Type" - end - - test "validates entities against type constraints" do - # Valid entity - valid_entity = %{ - types: ["verisim:Document"], - properties: %{ - "title" => "Valid Document" - } - } - - assert :ok == SchemaRegistry.validate(valid_entity) - - # Invalid entity (missing required field) - invalid_entity = %{ - types: ["verisim:Document"], - properties: %{} - } - - case SchemaRegistry.validate(invalid_entity) do - {:error, violations} -> - assert length(violations) > 0 - - :ok -> - # Schema validation might be lenient - assert true - end - end - - test "computes type hierarchy" do - hierarchy = SchemaRegistry.type_hierarchy("verisim:Document") - assert is_list(hierarchy) - assert "verisim:Document" in hierarchy - assert "verisim:Entity" in hierarchy - end - end - - describe "VCL Executor" do - test "generates explain plans" do - query_ast = %{ - modalities: [:graph, :vector], - source: {:octad, "test-id"}, - where: nil, - proof: nil, - limit: 10, - offset: 0 - } - - {:ok, plan} = VCLExecutor.execute(query_ast, explain: true) - - assert is_map(plan) - assert Map.has_key?(plan, :strategy) - assert Map.has_key?(plan, :steps) - end - - test "executes octad queries" do - query_ast = %{ - modalities: [:document], - source: {:octad, "nonexistent-id"}, - where: nil, - proof: nil, - limit: 10, - offset: 0 - } - - result = VCLExecutor.execute(query_ast) - - # Expect error for nonexistent octad - assert match?({:error, _}, result) - end - end - - describe "Entity Server" do - test "manages entity lifecycle" do - entity_id = "entity-server-test-#{:rand.uniform(10000)}" - - # Start entity server - {:ok, _pid} = EntityServer.start_link(entity_id) - - # Get initial state - {:ok, state} = EntityServer.get(entity_id) - - assert state.id == entity_id - assert state.status == :active - assert state.version == 0 - - # Update entity - {:ok, new_state} = EntityServer.update(entity_id, [ - {:modality, :document, true} - ]) - - assert new_state.modalities.document == true - end - - test "handles normalization requests" do - entity_id = "normalization-test-#{:rand.uniform(10000)}" - - {:ok, _pid} = EntityServer.start_link(entity_id) - - # Trigger normalization - :ok = EntityServer.normalize(entity_id) - - # Normalization is async, just verify it doesn't crash - Process.sleep(100) - - {:ok, state} = EntityServer.get(entity_id) - assert is_map(state) - end - end - - describe "Full Stack Integration" do - test "create octad via Rust, query via Elixir" do - # Create octad via Rust client - input = %{ - title: "Full Stack Test", - body: "Testing end-to-end integration", - embedding: List.duplicate(0.7, 384), - types: ["verisim:Document"] - } - - case RustClient.create_octad(input) do - {:ok, %{"id" => entity_id}} -> - # Query via Elixir QueryRouter - case RustClient.get_octad(entity_id) do - {:ok, octad} -> - assert octad["document"]["title"] == "Full Stack Test" - - {:error, reason} -> - flunk("Failed to retrieve octad: #{inspect(reason)}") - end - - {:error, _reason} -> - # Rust core not available - :ok - end - end - - test "cross-modal search" do - # Test combining text search with vector similarity - params = %{ - text: "integration test", - vector: List.duplicate(0.6, 384) - } - - result = QueryRouter.query(:multi, params, limit: 5) - - case result do - {:ok, results} -> - assert is_list(results) - - {:error, _} -> - # Expected if no matching data - :ok - end - end - - test "drift detection across stack" do - entity_id = "drift-integration-#{:rand.uniform(10000)}" - - # Create octad with potential drift - input = %{ - title: "Drift Test", - body: "Testing drift detection across stack", - embedding: List.duplicate(0.8, 384) - } - - case RustClient.create_octad(input) do - {:ok, %{"id" => ^entity_id}} -> - # Check drift via Elixir - case RustClient.get_drift_score(entity_id) do - {:ok, score} -> - assert is_float(score) - assert score >= 0.0 and score <= 1.0 - - {:error, _} -> - :ok - end - - {:error, _} -> - :ok - end - end - end - - # Helper functions - - defp wait_for_rust_core(retries \\ 5) do - case RustClient.health() do - {:ok, _} -> - :ok - - {:error, _} -> - if retries > 0 do - Process.sleep(1000) - wait_for_rust_core(retries - 1) - else - # Rust core not available, tests will be skipped - :ok - end - end - end -end diff --git a/verisimdb/elixir-orchestration/test/support/vcl_test_helpers.ex b/verisimdb/elixir-orchestration/test/support/vcl_test_helpers.ex deleted file mode 100644 index a8bce1d7..00000000 --- a/verisimdb/elixir-orchestration/test/support/vcl_test_helpers.ex +++ /dev/null @@ -1,317 +0,0 @@ -# SPDX-License-Identifier: MPL-2.0 - -defmodule VeriSim.Test.VCLTestHelpers do - @moduledoc """ - Shared test helpers for VCL integration tests. - - Provides: - - VCLBridge startup and teardown - - Octad fixtures for each modality combination - - AST assertion helpers - - Query execution wrappers with Rust-core-unavailable handling - - Cross-modal condition builders - - ## Usage - - use VeriSim.Test.VCLTestHelpers - - This imports all helper functions and sets up the VCLBridge GenServer in - `setup_all`. Tests can then call `parse!/1`, `execute_safely/2`, and - fixture builders without boilerplate. - """ - - alias VeriSim.Query.{VCLBridge, VCLExecutor} - - @doc """ - Ensure VCLBridge GenServer is running. Returns the PID. - - Safe to call multiple times — if already started, returns the existing PID. - """ - def ensure_bridge_started do - case VCLBridge.start_link([]) do - {:ok, pid} -> pid - {:error, {:already_started, pid}} -> pid - end - end - - @doc """ - Parse a VCL query string, raising on failure. - - Returns the parsed AST map. Useful in tests where a parse failure - means the test itself is broken, not the feature under test. - """ - def parse!(query_string) do - case VCLBridge.parse(query_string) do - {:ok, ast} -> ast - {:error, reason} -> raise "VCL parse failed: #{inspect(reason)}\n Query: #{query_string}" - end - end - - @doc """ - Parse a VCL statement (query or mutation), raising on failure. - """ - def parse_statement!(query_string) do - case VCLBridge.parse_statement(query_string) do - {:ok, ast} -> ast - {:error, reason} -> raise "VCL statement parse failed: #{inspect(reason)}\n Query: #{query_string}" - end - end - - @doc """ - Execute a VCL query string, handling Rust-core-unavailable gracefully. - - Returns `{:ok, result}`, `{:error, reason}`, or `{:unavailable, reason}` - when the Rust core is not running (rescues connection errors). - """ - def execute_safely(query_string, opts \\ []) do - opts = Keyword.put_new(opts, :timeout, 2_000) - - try do - VCLExecutor.execute_string(query_string, opts) - rescue - e -> - {:unavailable, Exception.message(e)} - end - end - - @doc """ - Execute a parsed VCL AST, handling Rust-core-unavailable gracefully. - """ - def execute_ast_safely(ast, opts \\ []) do - opts = Keyword.put_new(opts, :timeout, 2_000) - - try do - VCLExecutor.execute(ast, opts) - rescue - e -> - {:unavailable, Exception.message(e)} - end - end - - @doc """ - Execute a VCL statement AST (query or mutation), handling errors gracefully. - """ - def execute_statement_safely(ast, opts \\ []) do - opts = Keyword.put_new(opts, :timeout, 2_000) - - try do - VCLExecutor.execute_statement(ast, opts) - rescue - e -> - {:unavailable, Exception.message(e)} - end - end - - @doc """ - Assert that a result is either a successful response or an error due to - Rust core being unavailable. Fails only on unexpected crashes. - - Returns `:ok_result` if successful, `:rust_unavailable` if the Rust core - is not running (expected in CI / unit-test environments). - """ - def assert_ok_or_rust_unavailable(result) do - case result do - {:ok, _} -> - :ok_result - - {:error, reason} when is_atom(reason) -> - :rust_unavailable - - {:error, {tag, _}} when tag in [ - :connection_refused, :econnrefused, :connect_timeout, - :timeout, :req_error, :mint_error - ] -> - :rust_unavailable - - {:unavailable, _} -> - :rust_unavailable - - {:error, reason} -> - # Rust core returned an error (e.g., :not_found) — this is still a - # valid response, meaning the Rust core IS running. The error might - # be expected (nonexistent entity) or unexpected. - {:error_from_rust, reason} - - other -> - raise "Unexpected result shape: #{inspect(other)}" - end - end - - # =========================================================================== - # Fixtures: pre-built VCL query strings for common test scenarios - # =========================================================================== - - @doc "All 8 modality names as atoms." - def all_modalities do - [:graph, :vector, :tensor, :semantic, :document, :temporal, :provenance, :spatial] - end - - @doc "All 8 modality names as uppercase strings (for VCL SELECT)." - def all_modality_names do - ~w(GRAPH VECTOR TENSOR SEMANTIC DOCUMENT TEMPORAL PROVENANCE SPATIAL) - end - - @doc "Build a SELECT query for a single modality from a octad." - def single_modality_query(modality, entity_id \\ "test-entity-001") do - mod_upper = modality |> to_string() |> String.upcase() - "SELECT #{mod_upper}.* FROM HEXAD '#{entity_id}'" - end - - @doc "Build a SELECT query for multiple modalities from a octad." - def multi_modality_query(modalities, entity_id \\ "test-entity-001") do - projection = modalities - |> Enum.map(fn m -> "#{m |> to_string() |> String.upcase()}.*" end) - |> Enum.join(", ") - "SELECT #{projection} FROM HEXAD '#{entity_id}'" - end - - @doc "Build a SELECT * (all modalities) query from a octad." - def all_modality_query(entity_id \\ "test-entity-001") do - "SELECT * FROM HEXAD '#{entity_id}'" - end - - @doc "Build a federation query." - def federation_query(pattern \\ "/*", drift_policy \\ nil) do - base = "SELECT * FROM FEDERATION #{pattern}" - if drift_policy do - "#{base} WITH DRIFT #{drift_policy |> to_string() |> String.upcase()}" - else - base - end - end - - @doc "Build an INSERT mutation." - def insert_mutation(title, body) do - "INSERT HEXAD WITH DOCUMENT(title = '#{title}', body = '#{body}')" - end - - @doc "Build an UPDATE mutation." - def update_mutation(entity_id, field, value) do - "UPDATE HEXAD '#{entity_id}' SET #{field} = '#{value}'" - end - - @doc "Build a DELETE mutation." - def delete_mutation(entity_id) do - "DELETE HEXAD '#{entity_id}'" - end - - # =========================================================================== - # Cross-modal condition AST builders - # =========================================================================== - - @doc "Build a CrossModalFieldCompare condition AST node." - def cross_modal_compare(mod1, field1, op, mod2, field2) do - %{ - TAG: "CrossModalFieldCompare", - _0: to_string(mod1), - _1: to_string(field1), - _2: op, - _3: to_string(mod2), - _4: to_string(field2) - } - end - - @doc "Build a ModalityDrift condition AST node." - def modality_drift(mod1, mod2, threshold) do - %{TAG: "ModalityDrift", _0: to_string(mod1), _1: to_string(mod2), _2: threshold} - end - - @doc "Build a ModalityExists condition AST node." - def modality_exists(modality) do - %{TAG: "ModalityExists", _0: to_string(modality)} - end - - @doc "Build a ModalityNotExists condition AST node." - def modality_not_exists(modality) do - %{TAG: "ModalityNotExists", _0: to_string(modality)} - end - - @doc "Build a ModalityConsistency condition AST node." - def modality_consistency(mod1, mod2, metric) do - %{TAG: "ModalityConsistency", _0: to_string(mod1), _1: to_string(mod2), _2: metric} - end - - @doc "Build an And condition combining two sub-conditions." - def and_condition(left, right) do - %{TAG: "And", _0: left, _1: right} - end - - @doc "Build an Or condition combining two sub-conditions." - def or_condition(left, right) do - %{TAG: "Or", _0: left, _1: right} - end - - # =========================================================================== - # AST assertion helpers - # =========================================================================== - - @doc "Assert that an AST contains the expected modalities." - def assert_modalities(ast, expected) do - modalities = ast[:modalities] || ast["modalities"] || [] - expected_set = MapSet.new(expected) - actual_set = MapSet.new(modalities) - - unless MapSet.equal?(expected_set, actual_set) do - raise ExUnit.AssertionError, - message: "Modalities mismatch", - left: Enum.sort(modalities), - right: Enum.sort(expected) - end - end - - @doc "Assert that an AST has a specific source type." - def assert_source(ast, expected_type) do - source = ast[:source] || ast["source"] - - case {expected_type, source} do - {:octad, {:octad, _id}} -> :ok - {:federation, {:federation, _, _}} -> :ok - {:store, {:store, _id}} -> :ok - _ -> - raise ExUnit.AssertionError, - message: "Source type mismatch", - left: source, - right: expected_type - end - end - - @doc "Assert that an AST has a WHERE clause present." - def assert_has_where(ast) do - where = ast[:where] || ast["where"] - unless where do - raise ExUnit.AssertionError, - message: "Expected WHERE clause to be present, got nil" - end - end - - @doc "Assert that an AST has a PROOF clause present." - def assert_has_proof(ast) do - proof = ast[:proof] || ast["proof"] - unless proof do - raise ExUnit.AssertionError, - message: "Expected PROOF clause to be present, got nil" - end - end - - @doc "Assert that an AST has LIMIT set." - def assert_limit(ast, expected_limit) do - limit = ast[:limit] || ast["limit"] - unless limit == expected_limit do - raise ExUnit.AssertionError, - message: "LIMIT mismatch", - left: limit, - right: expected_limit - end - end - - @doc "Assert that an AST has OFFSET set." - def assert_offset(ast, expected_offset) do - offset = ast[:offset] || ast["offset"] - unless offset == expected_offset do - raise ExUnit.AssertionError, - message: "OFFSET mismatch", - left: offset, - right: expected_offset - end - end -end diff --git a/verisimdb/elixir-orchestration/test/test_helper.exs b/verisimdb/elixir-orchestration/test/test_helper.exs deleted file mode 100644 index bfa29171..00000000 --- a/verisimdb/elixir-orchestration/test/test_helper.exs +++ /dev/null @@ -1,3 +0,0 @@ -# SPDX-License-Identifier: MPL-2.0 - -ExUnit.start() diff --git a/verisimdb/elixir-orchestration/test/verisim/api/router_test.exs b/verisimdb/elixir-orchestration/test/verisim/api/router_test.exs deleted file mode 100644 index a8520de7..00000000 --- a/verisimdb/elixir-orchestration/test/verisim/api/router_test.exs +++ /dev/null @@ -1,287 +0,0 @@ -# SPDX-License-Identifier: MPL-2.0 - -defmodule VeriSim.Api.RouterTest do - @moduledoc """ - Tests for the orchestration HTTP API router. - - Uses Plug.Test to invoke endpoints directly without needing a running - HTTP server, so these tests work regardless of port configuration. - """ - - use ExUnit.Case, async: true - - @opts VeriSim.Api.Router.init([]) - - defp call(conn) do - VeriSim.Api.Router.call(conn, @opts) - end - - # ── Health ──────────────────────────────────────────────────────────── - - describe "GET /health" do - test "returns 200 with orchestration status" do - conn = - Plug.Test.conn(:get, "/health") - |> call() - - assert conn.status == 200 - assert {"content-type", "application/json; charset=utf-8"} in conn.resp_headers - - body = Jason.decode!(conn.resp_body) - assert body["status"] == "ok" - assert body["layer"] == "orchestration" - assert is_integer(body["uptime_seconds"]) - assert is_boolean(body["telemetry_enabled"]) - end - - test "returns CORS headers" do - conn = - Plug.Test.conn(:get, "/health") - |> call() - - assert {"access-control-allow-origin", "*"} in conn.resp_headers - end - end - - # ── Telemetry ───────────────────────────────────────────────────────── - - describe "GET /telemetry" do - test "returns telemetry report or disabled message" do - conn = - Plug.Test.conn(:get, "/telemetry") - |> call() - - assert conn.status == 200 - body = Jason.decode!(conn.resp_body) - - # Either full report (if enabled) or disabled message - assert Map.has_key?(body, "telemetry_enabled") or Map.has_key?(body, "meta") - end - end - - describe "GET /telemetry/modality-heatmap" do - test "returns modality heatmap data" do - conn = - Plug.Test.conn(:get, "/telemetry/modality-heatmap") - |> call() - - assert conn.status == 200 - body = Jason.decode!(conn.resp_body) - assert Map.has_key?(body, "counts") - assert Map.has_key?(body, "percentages") - assert Map.has_key?(body, "total_modality_queries") - end - end - - describe "GET /telemetry/query-patterns" do - test "returns query pattern distribution" do - conn = - Plug.Test.conn(:get, "/telemetry/query-patterns") - |> call() - - assert conn.status == 200 - body = Jason.decode!(conn.resp_body) - assert Map.has_key?(body, "total_queries") - assert Map.has_key?(body, "error_rate") - assert Map.has_key?(body, "by_type") - end - end - - describe "GET /telemetry/drift" do - test "returns drift report" do - conn = - Plug.Test.conn(:get, "/telemetry/drift") - |> call() - - assert conn.status == 200 - body = Jason.decode!(conn.resp_body) - assert Map.has_key?(body, "drift_detected_count") - assert Map.has_key?(body, "normalise_success_rate") - end - end - - describe "GET /telemetry/performance" do - test "returns performance summary" do - conn = - Plug.Test.conn(:get, "/telemetry/performance") - |> call() - - assert conn.status == 200 - body = Jason.decode!(conn.resp_body) - assert Map.has_key?(body, "query_count") - assert Map.has_key?(body, "avg_duration_ms") - end - end - - describe "GET /telemetry/federation" do - test "returns federation health" do - conn = - Plug.Test.conn(:get, "/telemetry/federation") - |> call() - - assert conn.status == 200 - body = Jason.decode!(conn.resp_body) - assert Map.has_key?(body, "total_federation_queries") - assert Map.has_key?(body, "peer_errors") - end - end - - describe "GET /telemetry/proof-types" do - test "returns proof type usage" do - conn = - Plug.Test.conn(:get, "/telemetry/proof-types") - |> call() - - assert conn.status == 200 - body = Jason.decode!(conn.resp_body) - assert Map.has_key?(body, "total_proofs") - assert Map.has_key?(body, "by_type") - assert Map.has_key?(body, "vcl_dt_active") - end - end - - describe "GET /telemetry/entities" do - test "returns entity summary" do - conn = - Plug.Test.conn(:get, "/telemetry/entities") - |> call() - - assert conn.status == 200 - body = Jason.decode!(conn.resp_body) - assert Map.has_key?(body, "created") - assert Map.has_key?(body, "deleted") - end - end - - # ── Status ──────────────────────────────────────────────────────────── - - describe "GET /status" do - test "returns orchestration status" do - conn = - Plug.Test.conn(:get, "/status") - |> call() - - assert conn.status == 200 - body = Jason.decode!(conn.resp_body) - assert body["orchestration"] == "running" - assert Map.has_key?(body, "consensus") - assert Map.has_key?(body, "telemetry_enabled") - end - end - - # ── Telemetry End-to-End Validation ────────────────────────────────── - - describe "telemetry endpoint structure validation" do - test "all telemetry sub-sections return valid JSON with expected keys" do - endpoints = [ - {"/telemetry/modality-heatmap", ["counts", "percentages", "total_modality_queries"]}, - {"/telemetry/query-patterns", ["total_queries", "error_count", "error_rate", "by_type"]}, - {"/telemetry/drift", ["drift_detected_count", "normalise_attempts", "normalise_success_count", "normalise_success_rate", "modality_breakdown"]}, - {"/telemetry/performance", ["query_count", "avg_duration_ms", "min_duration_ms", "max_duration_ms", "total_duration_ms"]}, - {"/telemetry/federation", ["total_federation_queries", "peer_errors"]}, - {"/telemetry/proof-types", ["total_proofs", "by_type", "vcl_dt_active"]}, - {"/telemetry/entities", ["created", "deleted"]} - ] - - for {path, expected_keys} <- endpoints do - conn = - Plug.Test.conn(:get, path) - |> call() - - assert conn.status == 200, "#{path} returned status #{conn.status}" - body = Jason.decode!(conn.resp_body) - - for key <- expected_keys do - assert Map.has_key?(body, key), - "#{path} missing expected key '#{key}'. Got: #{inspect(Map.keys(body))}" - end - end - end - - test "full telemetry report has all top-level sections for PanLL consumption" do - conn = - Plug.Test.conn(:get, "/telemetry") - |> call() - - assert conn.status == 200 - body = Jason.decode!(conn.resp_body) - - # Either full report or disabled message — both are valid - if Map.has_key?(body, "meta") do - # Full report structure matches what PanLL Model.res expects - expected_sections = ["meta", "modality_heatmap", "query_patterns", - "performance", "drift", "federation", "proof_types", "entities"] - - for section <- expected_sections do - assert Map.has_key?(body, section), - "Full telemetry report missing '#{section}' section" - end - - # Meta section has privacy notice - assert is_binary(body["meta"]["privacy_notice"]) - else - # Disabled message - assert body["telemetry_enabled"] == false - end - end - - test "all telemetry endpoints return numeric values (not nil or string)" do - conn = - Plug.Test.conn(:get, "/telemetry/performance") - |> call() - - body = Jason.decode!(conn.resp_body) - - assert is_number(body["query_count"]) - assert is_number(body["avg_duration_ms"]) - assert is_number(body["min_duration_ms"]) - assert is_number(body["max_duration_ms"]) - assert is_number(body["total_duration_ms"]) - end - - test "modality heatmap covers all 8 octad modalities" do - conn = - Plug.Test.conn(:get, "/telemetry/modality-heatmap") - |> call() - - body = Jason.decode!(conn.resp_body) - octad_modalities = ~w(graph vector tensor semantic document temporal provenance spatial) - - for modality <- octad_modalities do - assert Map.has_key?(body["counts"], modality), - "Modality heatmap missing '#{modality}' in counts" - assert Map.has_key?(body["percentages"], modality), - "Modality heatmap missing '#{modality}' in percentages" - end - end - - test "all telemetry endpoints return CORS headers" do - endpoints = ["/telemetry", "/telemetry/modality-heatmap", "/telemetry/query-patterns", - "/telemetry/drift", "/telemetry/performance", "/telemetry/federation", - "/telemetry/proof-types", "/telemetry/entities"] - - for path <- endpoints do - conn = - Plug.Test.conn(:get, path) - |> call() - - assert {"access-control-allow-origin", "*"} in conn.resp_headers, - "#{path} missing CORS header" - end - end - end - - # ── 404 ─────────────────────────────────────────────────────────────── - - describe "unknown routes" do - test "returns 404 with error message" do - conn = - Plug.Test.conn(:get, "/nonexistent") - |> call() - - assert conn.status == 404 - body = Jason.decode!(conn.resp_body) - assert body["error"] == "not_found" - end - end -end diff --git a/verisimdb/elixir-orchestration/test/verisim/aspect/concurrency_test.exs b/verisimdb/elixir-orchestration/test/verisim/aspect/concurrency_test.exs deleted file mode 100644 index 09066b4a..00000000 --- a/verisimdb/elixir-orchestration/test/verisim/aspect/concurrency_test.exs +++ /dev/null @@ -1,444 +0,0 @@ -# SPDX-License-Identifier: MPL-2.0 -# Author: Jonathan D.A. Jewell -# -# Concurrency aspect tests for VeriSimDB. -# -# Validates that the in-process components of VeriSimDB handle concurrent -# access correctly, with no data corruption, deadlocks, or silent data loss. -# -# Test categories: -# -# 1. Concurrent entity writes — last-write-wins or CRDT merge, never corrupt. -# 2. Parallel VCL queries — all queries complete without contention. -# 3. Concurrent Kraft proposals — at most one accepted per term slot. -# 4. DriftMonitor under load — concurrent drift reports do not corrupt state. -# 5. SchemaRegistry concurrency — concurrent type registrations are serialised. -# -# All tests are in-process. No external databases, containers, or TCP. - -defmodule VeriSim.Aspect.ConcurrencyTest do - @moduledoc """ - Concurrency tests for the VeriSimDB in-process stack. - - All concurrent operations are performed from Elixir Tasks, which map to - BEAM processes. The BEAM scheduler ensures preemptive multi-tasking, making - these tests meaningful even without OS-level thread races. - """ - - use ExUnit.Case, async: false - - alias VeriSim.{DriftMonitor, SchemaRegistry} - alias VeriSim.Consensus.KRaftNode - alias VeriSim.Query.{VCLBridge, VCLExecutor} - alias VeriSim.Test.VCLTestHelpers, as: H - - # Number of concurrent writers / readers for load tests. - @concurrency 20 - - # Timeout for collecting Task results (ms). - @task_timeout 5_000 - - setup_all do - _pid = H.ensure_bridge_started() - :ok - end - - # =========================================================================== - # 1. Concurrent Entity Writes - # - # @concurrency Tasks simultaneously write to the same entity ID via the - # EntityServer. The final state must be consistent — no corrupt maps, no - # nil fields, version must be a non-negative integer. - # =========================================================================== - - describe "concurrent writes to EntityServer" do - test "concurrent updates to same entity produce consistent final state" do - alias VeriSim.EntityServer - - entity_id = "conc-write-#{System.unique_integer([:positive])}" - {:ok, _pid} = EntityServer.start_link(entity_id) - - # Spawn @concurrency concurrent update tasks. - tasks = - for _i <- 1..@concurrency do - Task.async(fn -> - EntityServer.update(entity_id, [{:modality, :document, true}]) - end) - end - - results = Task.await_many(tasks, @task_timeout) - - # Every task must receive a result (no timeout, no crash). - assert length(results) == @concurrency - - Enum.each(results, fn result -> - assert match?({:ok, _state}, result) or match?({:error, _}, result), - "Unexpected result from concurrent update: #{inspect(result)}" - end) - - # Final state must be structurally valid. - {:ok, final_state} = EntityServer.get(entity_id) - assert is_map(final_state) - assert final_state.id == entity_id - assert is_integer(final_state.version) and final_state.version >= 0 - assert final_state.status == :active - end - - test "concurrent writes to different entities do not interfere" do - alias VeriSim.EntityServer - - # Start @concurrency independent entities. - entity_ids = - for i <- 1..@concurrency do - id = "conc-isolated-#{i}-#{System.unique_integer([:positive])}" - {:ok, _pid} = EntityServer.start_link(id) - id - end - - # Each task writes only to its own entity. - tasks = - for id <- entity_ids do - Task.async(fn -> - EntityServer.update(id, [{:modality, :vector, true}]) - end) - end - - results = Task.await_many(tasks, @task_timeout) - - # All tasks complete. - assert length(results) == @concurrency - - # Each entity's final state reflects its own writes, not another's. - for id <- entity_ids do - {:ok, state} = EntityServer.get(id) - assert state.id == id - assert state.modalities.vector == true - end - end - end - - # =========================================================================== - # 2. Parallel VCL Queries - # - # @concurrency Tasks simultaneously execute VCL queries through the executor. - # All must complete; none must crash or block indefinitely. - # =========================================================================== - - describe "parallel VCL queries" do - test "concurrent VCL parse calls produce consistent results" do - query = "SELECT DOCUMENT.* FROM HEXAD 'entity-001' LIMIT 10" - - tasks = - for _i <- 1..@concurrency do - Task.async(fn -> - VCLBridge.parse(query) - end) - end - - results = Task.await_many(tasks, @task_timeout) - - # Every parse must return the same AST (parsers must be pure functions). - assert length(results) == @concurrency - - unique_results = - results - |> Enum.filter(&match?({:ok, _}, &1)) - |> Enum.map(fn {:ok, ast} -> ast end) - |> Enum.uniq() - - # All successes should be structurally identical. - if length(unique_results) > 1 do - flunk("Concurrent parse of identical query produced #{length(unique_results)} different ASTs") - end - end - - test "concurrent VCL executions do not contend on internal state" do - # Build a simple no-proof query AST (avoids Rust-core dependency). - ast = %{ - modalities: [:document], - source: {:octad, "concurrent-query-test"}, - where: nil, - proof: nil, - limit: 1, - offset: 0 - } - - tasks = - for _i <- 1..@concurrency do - Task.async(fn -> - VCLExecutor.execute(ast) - end) - end - - results = Task.await_many(tasks, @task_timeout) - - # All queries must produce a structured result. - assert length(results) == @concurrency - - Enum.each(results, fn result -> - assert match?({:ok, _}, result) or match?({:error, _}, result), - "Concurrent executor returned invalid result: #{inspect(result)}" - end) - end - - test "concurrent explain plan generation is deterministic" do - ast = %{ - modalities: [:graph, :vector, :document], - source: {:octad, "explain-concurrent"}, - where: nil, - proof: nil, - limit: 5, - offset: 0 - } - - tasks = - for _i <- 1..@concurrency do - Task.async(fn -> - VCLExecutor.execute(ast, explain: true) - end) - end - - results = Task.await_many(tasks, @task_timeout) - - plans = - results - |> Enum.filter(&match?({:ok, _}, &1)) - |> Enum.map(fn {:ok, plan} -> plan end) - - # All explain plans must be structurally identical. - unique_plans = Enum.uniq_by(plans, fn p -> {p[:strategy], length(p[:steps] || [])} end) - - assert length(unique_plans) <= 1, - "Concurrent explain produced divergent plans: #{inspect(unique_plans)}" - end - end - - # =========================================================================== - # 3. Concurrent Kraft Proposals - # - # Multiple tasks simultaneously propose to the leader. All must complete - # (success or redirect), and the committed log must be internally consistent. - # =========================================================================== - - describe "concurrent Kraft proposals" do - test "concurrent proposals to leader do not corrupt registry" do - suffix = System.unique_integer([:positive]) - node_id = "conc-kraft-#{suffix}" - {:ok, pid} = KRaftNode.start_link(node_id: node_id, peers: []) - - # Wait for election. - Process.sleep(500) - assert KRaftNode.diagnostics(node_id).role == :leader - - # Propose @concurrency unique stores concurrently. - tasks = - for i <- 1..@concurrency do - Task.async(fn -> - KRaftNode.propose( - node_id, - {:register_store, "store-#{i}-#{suffix}", "http://localhost:#{9000 + i}", ["graph"]} - ) - end) - end - - results = Task.await_many(tasks, @task_timeout) - - # All must complete. - assert length(results) == @concurrency - - # Every result must be {:ok, index} — all proposals to a single leader - # should be accepted (no network partition). - Enum.each(results, fn result -> - assert match?({:ok, _index}, result), - "Concurrent proposal failed unexpectedly: #{inspect(result)}" - end) - - # Give time for all entries to be applied. - Process.sleep(200) - - # The registry must contain all @concurrency stores. - registry = KRaftNode.registry(node_id) - registered_count = map_size(registry.stores) - - assert registered_count == @concurrency, - "Expected #{@concurrency} stores in registry, found #{registered_count}" - - GenServer.stop(pid, :normal, 1_000) - catch - :exit, _ -> :ok - end - - test "log index monotonically increases under concurrent proposals" do - suffix = System.unique_integer([:positive]) - node_id = "conc-mono-#{suffix}" - {:ok, pid} = KRaftNode.start_link(node_id: node_id, peers: []) - - Process.sleep(500) - - # Collect log indices from concurrent proposals. - tasks = - for i <- 1..10 do - Task.async(fn -> - KRaftNode.propose( - node_id, - {:register_store, "mono-store-#{i}-#{suffix}", "http://localhost:#{9100 + i}", ["vector"]} - ) - end) - end - - results = Task.await_many(tasks, @task_timeout) - - indices = - results - |> Enum.filter(&match?({:ok, _}, &1)) - |> Enum.map(fn {:ok, index} -> index end) - |> Enum.sort() - - # Indices must be unique and strictly ascending (no duplicates). - assert indices == Enum.uniq(indices), - "Duplicate log indices detected: #{inspect(indices)}" - - GenServer.stop(pid, :normal, 1_000) - catch - :exit, _ -> :ok - end - end - - # =========================================================================== - # 4. DriftMonitor Under Concurrent Load - # - # @concurrency Tasks simultaneously report drift events. The monitor's - # health snapshot must remain structurally valid throughout. - # =========================================================================== - - describe "DriftMonitor under concurrent load" do - test "concurrent drift reports do not corrupt monitor state" do - tasks = - for i <- 1..@concurrency do - Task.async(fn -> - entity_id = "drift-conc-#{i}-#{System.unique_integer([:positive])}" - drift_score = :rand.uniform() * 0.5 - DriftMonitor.report_drift(entity_id, drift_score, :semantic_vector) - end) - end - - Task.await_many(tasks, @task_timeout) - - # Allow the monitor to process all events. - Process.sleep(100) - - # The monitor's status must remain structurally valid. - status = DriftMonitor.status() - assert is_map(status) - assert status[:overall_health] in [:healthy, :warning, :degraded, :critical], - "Invalid overall_health after concurrent drift reports: #{inspect(status[:overall_health])}" - assert is_integer(status[:entities_with_drift]) and status[:entities_with_drift] >= 0, - "entities_with_drift is non-integer: #{inspect(status[:entities_with_drift])}" - end - - test "concurrent entity_changed notifications do not deadlock" do - tasks = - for i <- 1..@concurrency do - Task.async(fn -> - entity_id = "changed-conc-#{i}" - DriftMonitor.entity_changed(entity_id) - end) - end - - # If this times out it indicates a deadlock in the monitor. - Task.await_many(tasks, @task_timeout) - - # Verification: monitor is still responsive after the storm. - status = DriftMonitor.status() - assert is_map(status) - end - end - - # =========================================================================== - # 5. SchemaRegistry Concurrent Type Registration - # - # Concurrent registrations of different types must all succeed (or be - # de-duplicated cleanly), and the registry must reflect every registration. - # =========================================================================== - - describe "SchemaRegistry concurrent type registration" do - test "concurrent type registrations are serialised without data loss" do - suffix = System.unique_integer([:positive]) - - type_defs = - for i <- 1..@concurrency do - %{ - iri: "https://concurrent.test.org/Type#{i}-#{suffix}", - label: "Concurrent Type #{i}", - supertypes: ["verisim:Entity"], - constraints: [] - } - end - - tasks = - for type_def <- type_defs do - Task.async(fn -> - SchemaRegistry.register_type(type_def) - end) - end - - results = Task.await_many(tasks, @task_timeout) - - # All registrations must succeed. - Enum.each(results, fn result -> - assert result == :ok, - "Type registration failed: #{inspect(result)}" - end) - - # All types must be retrievable from the registry. - for i <- 1..@concurrency do - iri = "https://concurrent.test.org/Type#{i}-#{suffix}" - type = SchemaRegistry.get_type(iri) - assert not is_nil(type), - "Type #{iri} missing from registry after concurrent registration" - assert type.label == "Concurrent Type #{i}" - end - end - - test "concurrent registration of the same type is handled safely (no crash)" do - # NOTE: SchemaRegistry.register_type/1 returns {:error, :already_exists} - # for duplicate IRIs rather than :ok (idempotent upsert). This is the - # current implementation behaviour. The test verifies the minimum safety - # bar: concurrent duplicate registrations must not crash, corrupt state, - # or return unexpected values — only :ok or {:error, :already_exists}. - suffix = System.unique_integer([:positive]) - iri = "https://idempotent.test.org/SharedType-#{suffix}" - - type_def = %{ - iri: iri, - label: "Shared Type", - supertypes: ["verisim:Entity"], - constraints: [] - } - - # Register once first to ensure the IRI exists. - assert :ok = SchemaRegistry.register_type(type_def) - - tasks = - for _i <- 1..@concurrency do - Task.async(fn -> - SchemaRegistry.register_type(type_def) - end) - end - - results = Task.await_many(tasks, @task_timeout) - - # All results must be :ok (first registration) or {:error, :already_exists}. - # None may be a crash, timeout, or unexpected value. - Enum.each(results, fn result -> - assert result == :ok or result == {:error, :already_exists}, - "Concurrent duplicate registration returned unexpected value: #{inspect(result)}" - end) - - # The type must still be present exactly once after the concurrent storm. - type = SchemaRegistry.get_type(iri) - assert not is_nil(type) - assert type.label == "Shared Type" - end - end -end diff --git a/verisimdb/elixir-orchestration/test/verisim/aspect/security_test.exs b/verisimdb/elixir-orchestration/test/verisim/aspect/security_test.exs deleted file mode 100644 index 0068e01c..00000000 --- a/verisimdb/elixir-orchestration/test/verisim/aspect/security_test.exs +++ /dev/null @@ -1,436 +0,0 @@ -# SPDX-License-Identifier: MPL-2.0 -# Author: Jonathan D.A. Jewell -# -# VCL security aspect tests. -# -# Validates that the VCL execution pipeline is hardened against the attack -# surface of a multi-tenant distributed database: -# -# 1. VCL Injection — crafted query strings do not escape their parse -# context and execute unintended operations. -# 2. Unauthorised Access — requests without valid authentication are -# rejected with a clear error, not silently -# truncated or granted partial data. -# 3. Cross-Tenant Isolation — an authenticated tenant cannot read or -# mutate octads belonging to a different tenant, -# even with a well-formed VCL query. -# 4. Error Disclosure — error responses do not leak internal paths, -# stack traces, or schema metadata to the caller. -# -# No real network connections are made. All adapter/auth calls are -# exercised through the in-process Elixir layer. - -defmodule VeriSim.Aspect.SecurityTest do - @moduledoc """ - Security aspect tests for VCL and federation. - - Covers: - - VCL injection through crafted query strings - - Unauthorised access rejection - - Cross-tenant isolation enforcement - - Error message hygiene (no internal detail leakage) - """ - - use ExUnit.Case, async: false - - alias VeriSim.Query.{VCLBridge, VCLExecutor, VCLTypeChecker} - alias VeriSim.Test.VCLTestHelpers, as: H - - setup_all do - pid = H.ensure_bridge_started() - %{bridge_pid: pid} - end - - # =========================================================================== - # 1. VCL Injection - # - # The VCL parser must treat all string literals as opaque data, not - # executable code. Injection attempts should either: - # a) Fail at the parse stage with {:error, _}, or - # b) Parse successfully as a literal string (correctly escaping the payload). - # - # In either case, the injected payload must not execute as SQL/VCL commands. - # =========================================================================== - - describe "VCL injection: crafted query strings are neutralised" do - test "single-quote injection in entity ID cannot escape string context" do - # Classic SQL injection attempt embedded in a VCL entity ID. - payload = "entity'; DROP TABLE octads; --" - query = "SELECT * FROM HEXAD '#{payload}'" - - result = VCLBridge.parse(query) - - case result do - {:error, _reason} -> - # Parser correctly rejected the malformed entity ID — ideal behaviour. - assert true - - {:ok, ast} -> - # If it parsed, the entity ID must be an opaque string, not executed. - source = ast[:source] - # The source must be a tagged tuple — not a raw command sequence. - assert match?({:octad, _id}, source), - "Injected payload should be treated as a literal entity ID, got: #{inspect(source)}" - - # The entity ID string must not contain unescaped statement terminators - # that could be forwarded to a downstream query engine. - case source do - {:octad, id} -> - # We accept any wrapping of the ID; the key invariant is that the - # parser did not execute the injected commands. - assert is_binary(id) or is_nil(id) - - _ -> - :ok - end - end - end - - test "semicolon injection in WHERE clause is rejected or escaped" do - # Attempt to inject a second VCL statement via a WHERE literal. - query = ~S(SELECT DOCUMENT.* FROM HEXAD 'entity-001' WHERE DOCUMENT.title = 'a'; DELETE HEXAD 'entity-001') - - result = VCLBridge.parse(query) - - # The parser must not produce an AST that represents two statements - # when given a single-statement query string. - case result do - {:error, _reason} -> - # Correctly rejected - assert true - - {:ok, ast} -> - # If it parsed, it must be a single SELECT statement. - # The presence of a mutation field on a query AST would be a bug. - refute Map.has_key?(ast, :mutation), - "Parser produced a mutation in a SELECT query — injection risk" - refute Map.has_key?(ast, :delete), - "Parser produced a delete op in a SELECT query — injection risk" - end - end - - test "null-byte injection in entity ID is handled without process crash" do - # NOTE: The VCL built-in parser does NOT strip null bytes from entity IDs - # (as of 2026-04-04). The null byte survives into the AST. This is a - # known gap: downstream stores must sanitise or reject IDs containing - # null bytes to prevent C-string truncation vulnerabilities at the FFI - # layer. Filed as a TODO in TEST-NEEDS.md. - # - # This test verifies the MINIMUM safety bar: the parser must not crash - # or raise an exception. Null-byte sanitisation at the parser layer is - # a P1 hardening task. - query = "SELECT * FROM HEXAD 'entity\x00malicious'" - - result = - try do - VCLBridge.parse(query) - rescue - e -> {:crashed, e} - end - - # Must not raise — structured result is mandatory. - refute match?({:crashed, _}, result), - "VCL parser raised an exception on null-byte input: #{inspect(result)}" - - assert match?({:ok, _}, result) or match?({:error, _}, result), - "VCL parser must return {:ok, _} or {:error, _} for null-byte input" - end - - test "excessively long query string does not crash the parser" do - # Fuzz the parser with a query that is 1 MiB long. - long_payload = String.duplicate("a", 1_048_576) - query = "SELECT DOCUMENT.* FROM HEXAD '#{long_payload}'" - - # The parser must return either {:ok, _} or {:error, _} — never raise. - result = VCLBridge.parse(query) - assert match?({:ok, _}, result) or match?({:error, _}, result), - "Parser must not raise on oversized input" - end - - test "deeply nested WHERE condition does not trigger stack overflow" do - # Craft a deeply nested AND condition (100 levels). - inner = "DOCUMENT.x > 0" - nested = Enum.reduce(1..100, inner, fn _, acc -> "(#{acc}) AND DOCUMENT.x > 0" end) - query = "SELECT DOCUMENT.* FROM HEXAD 'entity-001' WHERE #{nested}" - - result = VCLBridge.parse(query) - assert match?({:ok, _}, result) or match?({:error, _}, result), - "Parser must not overflow on deeply nested conditions" - end - end - - # =========================================================================== - # 2. Unauthorised Access - # - # The VCL executor must reject requests that lack a valid authentication - # context. The shape of authentication is a `:tenant_id` or `:auth_token` - # in the execution options; absences must be handled gracefully. - # =========================================================================== - - describe "unauthorised access: missing authentication is rejected" do - test "execute_string without auth context produces error, not data" do - # The executor may or may not enforce auth at the Elixir layer (Rust core - # handles auth enforcement when running). We verify that the execution path - # does not panic and that, when the Rust core returns an auth error, it is - # propagated correctly. - query = "SELECT * FROM HEXAD 'entity-001'" - - result = VCLExecutor.execute_string(query, auth_token: nil) - - # Must be either {:ok, []} (no data — no auth, no results), or - # {:error, reason} (explicit rejection). Must NEVER return data for - # a nil-auth request if the system is in auth-enforcing mode. - case result do - {:ok, _results} -> - # Auth not enforced at Elixir layer (enforcement is in Rust core / - # svalinn gateway) — this is an acceptable outcome. Tests for the - # Rust-level auth enforcement belong in the Rust integration tests. - assert true - - {:error, reason} -> - # Rejection with any reason is acceptable. - assert not is_nil(reason) - end - end - - test "VCL type checker does not expose schema internals in error messages" do - # The type checker should reject unknown proof types without revealing - # internal proof-obligation structure or schema implementation details. - # - # NOTE: As of 2026-04-04, VCLTypeChecker.do_normalize/1 calls - # :erlang.binary_to_existing_atom/1 on the proof type string, which raises - # ArgumentError for atoms not previously interned (e.g., "xattack"). - # This is a hardening gap: the type checker should guard with - # :erlang.binary_to_atom/2 or a safe whitelist lookup, not crash. - # We wrap in a try to verify the minimum bar: no unhandled crash propagates. - query = "SELECT GRAPH.* FROM HEXAD 'entity-001' PROOF XATTACK(entity-001)" - - typecheck_result = - case VCLBridge.parse(query) do - {:ok, ast} -> - try do - VCLTypeChecker.typecheck(ast) - rescue - ArgumentError -> {:error, {:unknown_proof_type, "XATTACK"}} - e -> {:error, {:unexpected_exception, inspect(e)}} - end - - {:error, _reason} -> - # Parser rejected the unknown proof type — ideal. - {:error, :parse_rejected} - end - - case typecheck_result do - {:error, reason} -> - reason_str = inspect(reason) - - # Error must not contain filesystem paths. - refute reason_str =~ ~r{/home/|/var/|/mnt/|\.ex:|\.exs:}, - "Error leaks filesystem path: #{reason_str}" - - # Error must not contain internal module names that reveal architecture. - refute reason_str =~ "VeriSim.Query.VCLTypeChecker.Impl", - "Error leaks internal module: #{reason_str}" - - {:ok, _} -> - # If typecheck succeeded, there is nothing to leak. - :ok - end - end - - test "parse error messages do not disclose internal grammar details" do - # Deliberately invalid query. - result = VCLBridge.parse("INVALID QUERY SYNTAX $$$ @@@ ###") - - case result do - {:error, reason} -> - reason_str = inspect(reason) - - # Must not contain raw Erlang stacktraces or internal module paths. - refute reason_str =~ ":erlang.apply", - "Error leaks Erlang stacktrace" - - # Must be a structured error, not a raw exception dump. - assert is_atom(reason) or is_tuple(reason) or is_binary(reason), - "Error must be a structured value, got: #{inspect(reason)}" - - {:ok, _} -> - # If the parser accepted it (unlikely), it did not crash — fine. - :ok - end - end - end - - # =========================================================================== - # 3. Cross-Tenant Isolation - # - # A VCL query for tenant A must not return data belonging to tenant B. - # The isolation mechanism is namespace-prefixed octad IDs: tenant IDs are - # encoded in the entity ID prefix (e.g., "tenant-A::entity-001"). - # - # We verify that the query router and executor respect the source entity ID - # and do not widen the query to return data from other namespaces. - # =========================================================================== - - describe "cross-tenant isolation: tenant namespaces are respected" do - test "query for tenant-A entity does not return tenant-B data" do - # Construct AST queries for two different tenant namespaces. - ast_a = %{ - modalities: [:document], - source: {:octad, "tenant-a::entity-001"}, - where: nil, - proof: nil, - limit: 100, - offset: 0 - } - - ast_b = %{ - modalities: [:document], - source: {:octad, "tenant-b::entity-001"}, - where: nil, - proof: nil, - limit: 100, - offset: 0 - } - - result_a = VCLExecutor.execute(ast_a) - result_b = VCLExecutor.execute(ast_b) - - # Both may return {:error, :not_found} (Rust core not running), or - # {:ok, results}. The key invariant is that result_a must not contain - # any item whose ID is prefixed with "tenant-b::". - case result_a do - {:ok, results} -> - Enum.each(results, fn item -> - id = item["id"] || item[:id] || "" - refute String.starts_with?(id, "tenant-b::"), - "Cross-tenant leak: tenant-A result contains tenant-B item: #{inspect(item)}" - end) - - {:error, _} -> - # Rust core unavailable or entity not found — no leak possible. - assert true - end - - # Symmetric check. - case result_b do - {:ok, results} -> - Enum.each(results, fn item -> - id = item["id"] || item[:id] || "" - refute String.starts_with?(id, "tenant-a::"), - "Cross-tenant leak: tenant-B result contains tenant-A item: #{inspect(item)}" - end) - - {:error, _} -> - assert true - end - end - - test "federation query wildcard does not return all tenants' stores" do - # A FEDERATION query with pattern /* must not be interpreted as - # "return data from every registered tenant's store". - query = "SELECT * FROM FEDERATION /*" - result = H.execute_safely(query) - - # We only verify this does not crash and does not return a non-list result. - case result do - {:ok, items} -> - assert is_list(items) - - {:error, _reason} -> - # Acceptable — no stores registered or Rust core down. - assert true - end - end - - test "INSERT with tenant-A prefix cannot write to tenant-B namespace" do - # Build an INSERT AST with an explicit tenant-A document. - mutation_ast = %{ - TAG: "Mutation", - _0: %{ - TAG: "Insert", - modalities: %{ - document: %{title: "tenant-a doc", body: "test"}, - id_prefix: "tenant-a::" - }, - proof: nil - } - } - - result = VCLExecutor.execute_mutation(mutation_ast[:_0]) - - case result do - {:ok, created} -> - id = created["id"] || created[:id] || "" - - # If the system set an ID, it must not be in the tenant-b namespace. - refute String.starts_with?(id, "tenant-b::"), - "INSERT created entity in wrong tenant namespace: #{inspect(id)}" - - {:error, _} -> - # Rust core unavailable or mutation rejected — fine. - assert true - end - end - end - - # =========================================================================== - # 4. Error Disclosure Hygiene - # - # Error responses propagated to callers must be sanitised: they should - # contain enough information for the caller to understand what went wrong - # without exposing internal paths, secrets, or architecture details. - # =========================================================================== - - describe "error disclosure: error messages do not leak internals" do - test "executing a query against a non-existent store returns clean error" do - ast = %{ - modalities: [:document], - source: {:store, "nonexistent-store-0xdeadbeef"}, - where: nil, - proof: nil, - limit: 10, - offset: 0 - } - - result = VCLExecutor.execute(ast) - - case result do - {:error, reason} -> - reason_str = inspect(reason) - - # Must not contain filesystem paths or host-level details. - refute reason_str =~ ~r{/var/mnt|/home/hyper|nvme[01]}, - "Error leaks filesystem or device path: #{reason_str}" - - # Must be a structured error term. - assert not is_nil(reason) - - {:ok, _} -> - # Returns empty results — acceptable. - assert true - end - end - - test "malformed proof spec produces structured error without stacktrace" do - # Pass a proof spec that is missing required fields. - query_ast = %{ - modalities: [:graph], - proof: [%{raw: "EXISTENCE()"}] # Missing contract name - } - - result = VCLTypeChecker.typecheck(query_ast) - - case result do - {:error, reason} -> - # Error must be a structured term, not a raw Exception struct. - refute match?(%{__exception__: true, __struct__: _}, reason), - "Error should not be a raw exception struct (leaks internals)" - - {:ok, _} -> - # If it succeeded despite the malformed spec, that's acceptable. - assert true - end - end - end -end diff --git a/verisimdb/elixir-orchestration/test/verisim/consensus/kraft_node_test.exs b/verisimdb/elixir-orchestration/test/verisim/consensus/kraft_node_test.exs deleted file mode 100644 index e10b2f56..00000000 --- a/verisimdb/elixir-orchestration/test/verisim/consensus/kraft_node_test.exs +++ /dev/null @@ -1,302 +0,0 @@ -# SPDX-License-Identifier: MPL-2.0 - -defmodule VeriSim.Consensus.KRaftNodeTest do - use ExUnit.Case, async: false - - alias VeriSim.Consensus.KRaftNode - - # The Consensus.Registry is started by the Application supervisor. - # KRaft nodes use unique via-tuple names, so each test can start - # its own nodes without conflict. - - setup do - # Track started nodes for cleanup - nodes = [] - on_exit(fn -> - # Nodes are cleaned up by ExUnit since we use start_supervised for each - :ok - end) - {:ok, nodes: nodes} - end - - defp start_kraft(node_id, peers \\ []) do - # Use start_link directly since Registry is already running via the app. - # We can't use start_supervised! because it expects a child spec compatible - # with the test supervisor, and the via-tuple registration goes through the - # app's Consensus.Registry. - {:ok, pid} = KRaftNode.start_link(node_id: node_id, peers: peers) - pid - end - - defp stop_kraft(pid) do - if Process.alive?(pid), do: GenServer.stop(pid, :normal, 1_000) - catch - :exit, _ -> :ok - end - - # Poll `fun` every `interval_ms` ms until it returns a truthy value, or - # fail after `timeout_ms` ms have elapsed. Used for membership-change - # assertions where the commit + apply cycle can be slower under CI load. - defp poll_until(fun, timeout_ms \\ 2_000, interval_ms \\ 50) do - deadline = System.monotonic_time(:millisecond) + timeout_ms - - Enum.reduce_while(Stream.repeatedly(fn -> :poll end), nil, fn _, _acc -> - if fun.() do - {:halt, :ok} - else - remaining = deadline - System.monotonic_time(:millisecond) - - if remaining <= 0 do - {:halt, :timeout} - else - Process.sleep(min(interval_ms, remaining)) - {:cont, nil} - end - end - end) - end - - describe "single-node leader election" do - test "node with 0 peers becomes leader within 500ms" do - node_id = "solo-#{System.unique_integer([:positive])}" - pid = start_kraft(node_id) - - # Wait for election timeout (max 300ms) + margin - Process.sleep(500) - - diag = KRaftNode.diagnostics(node_id) - assert diag.role == :leader - assert diag.leader_id == node_id - - stop_kraft(pid) - end - end - - describe "3-node cluster election" do - test "exactly 1 leader emerges" do - suffix = System.unique_integer([:positive]) - ids = ["n1-#{suffix}", "n2-#{suffix}", "n3-#{suffix}"] - - pids = - for id <- ids do - peers = Enum.reject(ids, &(&1 == id)) - start_kraft(id, peers) - end - - # Allow enough time for election (timeouts are 150-300ms, but under CI - # load or when the system is busy, elections can take several rounds). - # Retry up to 5 times with increasing delays to handle slow convergence. - {leaders, roles} = - Enum.reduce_while(1..5, {0, []}, fn attempt, _acc -> - wait = 500 * attempt - Process.sleep(wait) - - roles = - Enum.map(ids, fn id -> - KRaftNode.diagnostics(id).role - end) - - leaders = Enum.count(roles, &(&1 == :leader)) - - if leaders == 1 do - {:halt, {leaders, roles}} - else - {:cont, {leaders, roles}} - end - end) - - assert leaders == 1, "Expected 1 leader, got #{leaders}. Roles: #{inspect(roles)}" - - Enum.each(pids, &stop_kraft/1) - end - end - - describe "command proposal" do - test "leader accepts register_store command" do - node_id = "cmd-#{System.unique_integer([:positive])}" - pid = start_kraft(node_id) - - Process.sleep(500) - - command = {:register_store, "store-1", "http://localhost:9000", ["graph", "vector"]} - assert {:ok, _index} = KRaftNode.propose(node_id, command) - - stop_kraft(pid) - end - end - - describe "registry state after commit" do - test "register_store appears in registry" do - node_id = "reg-#{System.unique_integer([:positive])}" - pid = start_kraft(node_id) - - Process.sleep(500) - - command = {:register_store, "my-store", "http://localhost:9001", ["document"]} - {:ok, _} = KRaftNode.propose(node_id, command) - - # Give time for commit + apply - Process.sleep(100) - - registry = KRaftNode.registry(node_id) - assert Map.has_key?(registry.stores, "my-store") - assert registry.stores["my-store"].endpoint == "http://localhost:9001" - assert registry.stores["my-store"].modalities == ["document"] - - stop_kraft(pid) - end - end - - describe "non-leader redirect" do - test "proposal to follower returns {:error, {:not_leader, leader_id}}" do - suffix = System.unique_integer([:positive]) - ids = ["f1-#{suffix}", "f2-#{suffix}", "f3-#{suffix}"] - - pids = - for id <- ids do - peers = Enum.reject(ids, &(&1 == id)) - start_kraft(id, peers) - end - - Process.sleep(1_000) - - # Find a follower - {follower_id, _} = - ids - |> Enum.map(fn id -> {id, KRaftNode.diagnostics(id)} end) - |> Enum.find(fn {_id, diag} -> diag.role == :follower end) - - result = KRaftNode.propose(follower_id, {:register_store, "s", "http://x", []}) - assert {:error, {:not_leader, _leader}} = result - - Enum.each(pids, &stop_kraft/1) - end - end - - describe "dynamic membership — add_server" do - test "leader accepts add_server and new peer appears in registry members" do - node_id = "add-#{System.unique_integer([:positive])}" - pid = start_kraft(node_id) - - Process.sleep(500) - - assert {:ok, _index} = KRaftNode.add_server(node_id, "new-peer-1", []) - - Process.sleep(100) - - registry = KRaftNode.registry(node_id) - members = get_in(registry, [:config, :members]) || [] - assert "new-peer-1" in members - - stop_kraft(pid) - end - end - - describe "dynamic membership — remove_server" do - test "leader accepts remove_server and peer is removed from registry members" do - node_id = "rm-#{System.unique_integer([:positive])}" - pid = start_kraft(node_id) - - Process.sleep(500) - - # Add then remove - {:ok, _} = KRaftNode.add_server(node_id, "ephemeral-peer", []) - - # Wait for add to be committed before issuing remove - :ok = - poll_until(fn -> - members = get_in(KRaftNode.registry(node_id), [:config, :members]) || [] - "ephemeral-peer" in members - end) - - {:ok, _} = KRaftNode.remove_server(node_id, "ephemeral-peer") - - # Poll until the commit + state-machine apply removes the member - result = - poll_until(fn -> - members = get_in(KRaftNode.registry(node_id), [:config, :members]) || [] - "ephemeral-peer" not in members - end) - - assert result == :ok, "ephemeral-peer was not removed from registry within 2 s" - - stop_kraft(pid) - end - end - - describe "dynamic membership — removed node stops participating" do - test "removed peer no longer in diagnostics peer_count after commit" do - suffix = System.unique_integer([:positive]) - leader_id = "dyn-l-#{suffix}" - follower_id = "dyn-f-#{suffix}" - - leader_pid = start_kraft(leader_id, [follower_id]) - follower_pid = start_kraft(follower_id, [leader_id]) - - # Poll until exactly one leader has been elected (up to 3 s under CI load) - result = - poll_until( - fn -> - roles = Enum.map([leader_id, follower_id], &KRaftNode.diagnostics(&1).role) - Enum.count(roles, &(&1 == :leader)) == 1 - end, - 3_000, - 100 - ) - - assert result == :ok, "Cluster did not elect a leader within 3 s" - - # Find which node is actually the leader - leader_diag = KRaftNode.diagnostics(leader_id) - {actual_leader, actual_follower, actual_follower_pid} = - if leader_diag.role == :leader do - {leader_id, follower_id, follower_pid} - else - {follower_id, leader_id, leader_pid} - end - - # Remove the follower via the leader - {:ok, _} = KRaftNode.remove_server(actual_leader, actual_follower) - - # Poll until the leader's peer list has been updated (commit + apply) - remove_result = - poll_until( - fn -> - KRaftNode.diagnostics(actual_leader).peer_count == 0 - end, - 2_000, - 50 - ) - - assert remove_result == :ok, - "Leader peer_count did not drop to 0 within 2 s after remove_server" - - stop_kraft(leader_pid) - stop_kraft(follower_pid) - end - end - - describe "diagnostics" do - test "returns expected fields" do - node_id = "diag-#{System.unique_integer([:positive])}" - pid = start_kraft(node_id) - - Process.sleep(500) - - diag = KRaftNode.diagnostics(node_id) - - assert is_binary(diag.node_id) - assert diag.node_id == node_id - assert diag.role in [:leader, :follower, :candidate] - assert is_integer(diag.current_term) - assert is_integer(diag.commit_index) - assert is_integer(diag.last_applied) - assert is_integer(diag.log_length) - assert is_integer(diag.peer_count) - assert is_integer(diag.election_count) - assert is_integer(diag.pending_requests) - - stop_kraft(pid) - end - end -end diff --git a/verisimdb/elixir-orchestration/test/verisim/consensus/kraft_property_test.exs b/verisimdb/elixir-orchestration/test/verisim/consensus/kraft_property_test.exs deleted file mode 100644 index 34244c28..00000000 --- a/verisimdb/elixir-orchestration/test/verisim/consensus/kraft_property_test.exs +++ /dev/null @@ -1,328 +0,0 @@ -# SPDX-License-Identifier: MPL-2.0 -# Author: Jonathan D.A. Jewell -# -# Kraft consensus P2P property-based tests. -# -# Validates the safety and liveness properties of the in-process KRaft -# consensus implementation against the canonical Raft guarantees: -# -# 1. Election Safety — at most one leader per term. -# 2. Log Matching — committed entries replicate to all live nodes. -# 3. State Machine — all nodes apply the same sequence of entries. -# 4. Partition Tolerance — writes succeed with quorum, reject below quorum. -# -# All tests are in-process — no Docker, no TCP, no external processes. -# The KRaftNode GenServer communicates via GenServer.call/cast over the -# local Consensus.Registry, so the full Raft state machine runs in-memory. - -defmodule VeriSim.Consensus.KRaftPropertyTest do - @moduledoc """ - Property-based tests for the KRaft consensus layer. - - Uses ExUnitProperties / StreamData to drive arbitrary cluster sizes - and command sequences, verifying that Raft's core invariants hold - across a wide variety of schedules. - - ## P2P Properties Covered - - - **Leader uniqueness**: in any cluster of 1..7 nodes, after election - converges, exactly one node holds the `:leader` role. - - **Log replication**: commands proposed to the leader appear in the - committed registry of all follower nodes within a bounded window. - - **Partition tolerance**: a cluster of N nodes accepts writes when a - quorum (⌈N/2⌉ + 1) of nodes is reachable; it rejects writes below - quorum (simulated by isolating the leader from its peers). - - **Idempotent read-your-writes**: a command proposed by the leader is - immediately visible in that node's own registry. - """ - - use ExUnit.Case, async: false - use ExUnitProperties - - alias VeriSim.Consensus.KRaftNode - - # Maximum wall-clock time (ms) we wait for an election to converge. - # KRaft election timeouts are 150–300 ms; allow 5 rounds under CI load. - @election_convergence_ms 2_000 - - # Maximum wall-clock time (ms) we wait for a command to replicate. - @replication_window_ms 500 - - # --------------------------------------------------------------------------- - # Generators - # --------------------------------------------------------------------------- - - # Generates a cluster of between 1 and 5 nodes (odd sizes only, to avoid - # split-brain ambiguity in the quorum calculation tests). - defp cluster_size_gen do - member_of([1, 3, 5]) - end - - # Generates a valid store-registration command payload. - defp store_command_gen do - gen all name <- string(:alphanumeric, min_length: 4, max_length: 12), - port <- integer(9_000..9_999), - modality <- member_of(~w(graph vector document tensor semantic temporal spatial provenance)) do - {:register_store, "store-#{name}", "http://localhost:#{port}", [modality]} - end - end - - # --------------------------------------------------------------------------- - # Helpers - # --------------------------------------------------------------------------- - - # Start a symmetric cluster: every node knows about all other nodes. - defp start_cluster(size) do - suffix = System.unique_integer([:positive]) - ids = for i <- 1..size, do: "prop-n#{i}-#{suffix}" - - pids = - for id <- ids do - peers = Enum.reject(ids, &(&1 == id)) - {:ok, pid} = KRaftNode.start_link(node_id: id, peers: peers) - pid - end - - {ids, pids} - end - - # Stop all nodes in a cluster, ignoring already-dead processes. - defp stop_cluster(pids) do - Enum.each(pids, fn pid -> - if Process.alive?(pid) do - GenServer.stop(pid, :normal, 1_000) - end - end) - catch - :exit, _ -> :ok - end - - # Wait until the cluster has elected exactly one leader, or raise on timeout. - defp await_leader(ids, timeout_ms \\ @election_convergence_ms) do - deadline = System.monotonic_time(:millisecond) + timeout_ms - - Stream.repeatedly(fn -> - roles = Enum.map(ids, fn id -> KRaftNode.diagnostics(id).role end) - leader_count = Enum.count(roles, &(&1 == :leader)) - {leader_count, roles} - end) - |> Enum.find_value(fn {count, roles} -> - cond do - count == 1 -> - {count, roles} - - System.monotonic_time(:millisecond) > deadline -> - raise "Election did not converge within #{timeout_ms} ms. Roles: #{inspect(roles)}" - - true -> - Process.sleep(50) - nil - end - end) - end - - # Find the current leader ID from a list of node IDs. - defp find_leader(ids) do - Enum.find(ids, fn id -> KRaftNode.diagnostics(id).role == :leader end) - end - - # --------------------------------------------------------------------------- - # Property 1: Leader Uniqueness - # - # For any cluster size in {1, 3, 5}, after election converges exactly one - # node is the leader, and that leader's diagnostics are self-consistent. - # --------------------------------------------------------------------------- - - property "leader uniqueness: exactly one leader per cluster after convergence" do - check all size <- cluster_size_gen(), - max_runs: 5 do - {ids, pids} = start_cluster(size) - - {leader_count, _roles} = await_leader(ids) - assert leader_count == 1, - "Expected exactly 1 leader in #{size}-node cluster" - - leader_id = find_leader(ids) - diag = KRaftNode.diagnostics(leader_id) - - # The leader must know it is the leader and must agree on its own ID. - assert diag.role == :leader - assert diag.leader_id == leader_id - assert is_integer(diag.current_term) and diag.current_term >= 1 - - stop_cluster(pids) - end - end - - # --------------------------------------------------------------------------- - # Property 2: Log Replication - # - # A command proposed to the leader eventually appears in the registry of - # all nodes in the cluster, within @replication_window_ms. - # --------------------------------------------------------------------------- - - property "log replication: commands written to leader appear on all nodes" do - check all size <- cluster_size_gen(), - command <- store_command_gen(), - max_runs: 5 do - {ids, pids} = start_cluster(size) - await_leader(ids) - - leader_id = find_leader(ids) - {:register_store, store_name, endpoint, modalities} = command - - {:ok, _index} = KRaftNode.propose(leader_id, command) - - # Yield time for replication to complete on all nodes. - Process.sleep(@replication_window_ms) - - # Every node's registry must contain the store after replication. - for id <- ids do - registry = KRaftNode.registry(id) - assert Map.has_key?(registry.stores, store_name), - "Node #{id} missing store #{store_name} after replication. " <> - "Registry: #{inspect(Map.keys(registry.stores))}" - - store_entry = registry.stores[store_name] - assert store_entry.endpoint == endpoint - assert store_entry.modalities == modalities - end - - stop_cluster(pids) - end - end - - # --------------------------------------------------------------------------- - # Property 3: State Machine Safety - # - # Multiple sequential commands proposed to the leader produce a consistent - # final registry state that is identical on all nodes. - # --------------------------------------------------------------------------- - - property "state machine: sequential commands produce consistent final state on all nodes" do - check all size <- cluster_size_gen(), - commands <- list_of(store_command_gen(), min_length: 2, max_length: 5), - max_runs: 5 do - {ids, pids} = start_cluster(size) - await_leader(ids) - - leader_id = find_leader(ids) - - # Propose all commands sequentially via the leader. - Enum.each(commands, fn command -> - {:ok, _index} = KRaftNode.propose(leader_id, command) - end) - - # Allow all entries to replicate. - Process.sleep(@replication_window_ms) - - # Collect final registry states. - registries = Enum.map(ids, fn id -> KRaftNode.registry(id) end) - - # All nodes must agree on the same set of store names. - store_name_sets = Enum.map(registries, fn reg -> MapSet.new(Map.keys(reg.stores)) end) - [first | rest] = store_name_sets - - Enum.each(rest, fn other -> - assert MapSet.equal?(first, other), - "Nodes disagree on store set: #{inspect(first)} vs #{inspect(other)}" - end) - - stop_cluster(pids) - end - end - - # --------------------------------------------------------------------------- - # Property 4: Partition Tolerance — Quorum Write - # - # In a 3-node cluster, a command submitted when all 3 nodes are alive - # (quorum = 2, all 3 available) must succeed. - # - # We verify liveness (no error) when a majority is available. - # We simulate "below quorum" by checking that a proposal to a non-leader - # redirects rather than blocking indefinitely. - # --------------------------------------------------------------------------- - - property "partition tolerance: write to leader in full cluster always succeeds" do - check all command <- store_command_gen(), - max_runs: 10 do - # Use a fixed 3-node cluster for partition tests. - {ids, pids} = start_cluster(3) - await_leader(ids) - - leader_id = find_leader(ids) - result = KRaftNode.propose(leader_id, command) - - # The write must succeed (quorum of 2 is trivially met with all 3 running). - assert match?({:ok, _index}, result), - "Expected {:ok, index} but got #{inspect(result)}" - - stop_cluster(pids) - end - end - - property "partition tolerance: proposal to follower redirects to leader (not silent drop)" do - check all command <- store_command_gen(), - max_runs: 5 do - {ids, pids} = start_cluster(3) - await_leader(ids) - - # Find a follower. - follower_id = Enum.find(ids, fn id -> KRaftNode.diagnostics(id).role == :follower end) - - if is_nil(follower_id) do - # 1-node clusters have no followers — skip this check. - stop_cluster(pids) - else - result = KRaftNode.propose(follower_id, command) - - # A follower must redirect, not accept the write on behalf of the leader. - # Valid responses: {:ok, _} (auto-forwarded) or {:error, {:not_leader, _}}. - assert match?({:ok, _}, result) or match?({:error, {:not_leader, _}}, result), - "Unexpected follower response: #{inspect(result)}" - - stop_cluster(pids) - end - end - end - - # --------------------------------------------------------------------------- - # Idempotent Read-Your-Writes - # - # After a leader commits a command, it is immediately visible in the - # leader's own registry — no external wait required. - # --------------------------------------------------------------------------- - - test "idempotent read-your-writes: committed command visible in leader registry immediately" do - suffix = System.unique_integer([:positive]) - node_id = "ryw-#{suffix}" - {:ok, pid} = KRaftNode.start_link(node_id: node_id, peers: []) - - # Wait for single-node election. - Process.sleep(500) - assert KRaftNode.diagnostics(node_id).role == :leader - - store_name = "ryw-store-#{suffix}" - command = {:register_store, store_name, "http://localhost:9876", ["document"]} - {:ok, _index} = KRaftNode.propose(node_id, command) - - # Give the single node time to apply the commit. - Process.sleep(100) - - registry = KRaftNode.registry(node_id) - assert Map.has_key?(registry.stores, store_name), - "Leader should see its own committed write immediately" - - stop_kraft(pid) - end - - # --------------------------------------------------------------------------- - # Helpers - # --------------------------------------------------------------------------- - - defp stop_kraft(pid) do - if Process.alive?(pid), do: GenServer.stop(pid, :normal, 1_000) - catch - :exit, _ -> :ok - end -end diff --git a/verisimdb/elixir-orchestration/test/verisim/consensus/kraft_recovery_test.exs b/verisimdb/elixir-orchestration/test/verisim/consensus/kraft_recovery_test.exs deleted file mode 100644 index f28e5812..00000000 --- a/verisimdb/elixir-orchestration/test/verisim/consensus/kraft_recovery_test.exs +++ /dev/null @@ -1,267 +0,0 @@ -# SPDX-License-Identifier: MPL-2.0 - -defmodule VeriSim.Consensus.KRaftRecoveryTest do - @moduledoc """ - Tests for KRaft node WAL integration and crash recovery. - - Verifies that KRaft nodes: - 1. Persist durable state (currentTerm, votedFor) via WAL - 2. Persist log entries via WAL - 3. Recover correctly after restart (simulated crash) - 4. Preserve registry state through recovery - """ - - use ExUnit.Case, async: false - - alias VeriSim.Consensus.KRaftNode - alias VeriSim.Consensus.KRaftWAL - - setup do - dir = Path.join(System.tmp_dir!(), "kraft_recovery_#{System.unique_integer([:positive])}") - on_exit(fn -> File.rm_rf!(dir) end) - {:ok, wal_path: dir} - end - - defp unique_id(prefix) do - "#{prefix}-#{System.unique_integer([:positive])}" - end - - defp start_node(node_id, opts \\ []) do - {:ok, pid} = KRaftNode.start_link([node_id: node_id] ++ opts) - pid - end - - defp stop_node(pid) do - if Process.alive?(pid), do: GenServer.stop(pid, :normal, 1_000) - catch - :exit, _ -> :ok - end - - # =========================================================================== - # WAL persistence during normal operation - # =========================================================================== - - describe "WAL persistence during normal operation" do - test "persists term and votedFor when node starts election", %{wal_path: wal_path} do - node_id = unique_id("wal-election") - pid = start_node(node_id, wal_path: wal_path) - - # Wait for election (single node becomes leader) - Process.sleep(500) - - # Check WAL has persisted state - {:ok, recovered} = KRaftWAL.recover(wal_path) - assert recovered.current_term > 0 - assert recovered.voted_for == node_id - - stop_node(pid) - end - - test "persists log entries when commands are proposed", %{wal_path: wal_path} do - node_id = unique_id("wal-propose") - pid = start_node(node_id, wal_path: wal_path) - - # Wait for leader election - Process.sleep(500) - - # Propose a command - command = {:register_store, "s1", "http://localhost:9000", ["graph"]} - {:ok, _index} = KRaftNode.propose(node_id, command) - - # Check WAL has the entry (plus the noop from leader election) - {:ok, recovered} = KRaftWAL.recover(wal_path) - assert length(recovered.log) >= 2 - - # The last entry should be our register_store command - last_entry = List.last(recovered.log) - assert last_entry.command == command - - stop_node(pid) - end - - test "persists noop entry when node becomes leader", %{wal_path: wal_path} do - node_id = unique_id("wal-noop") - pid = start_node(node_id, wal_path: wal_path) - - Process.sleep(500) - - {:ok, recovered} = KRaftWAL.recover(wal_path) - # Leader appends a noop entry on election - noop_entries = Enum.filter(recovered.log, &(&1.command == :noop)) - assert length(noop_entries) >= 1 - - stop_node(pid) - end - end - - # =========================================================================== - # Crash recovery — single node - # =========================================================================== - - describe "crash recovery — single node" do - test "recovers term after restart", %{wal_path: wal_path} do - node_id = unique_id("recover-term") - - # Phase 1: Start node, let it elect itself - pid = start_node(node_id, wal_path: wal_path) - Process.sleep(500) - - diag1 = KRaftNode.diagnostics(node_id) - term_before = diag1.current_term - assert term_before > 0 - - # "Crash" the node - stop_node(pid) - Process.sleep(100) - - # Phase 2: Restart with same WAL path - pid2 = start_node(node_id, wal_path: wal_path) - Process.sleep(100) - - # Should recover with at least the previous term - diag2 = KRaftNode.diagnostics(node_id) - assert diag2.current_term >= term_before - - stop_node(pid2) - end - - test "recovers log entries after restart", %{wal_path: wal_path} do - node_id = unique_id("recover-log") - - # Phase 1: Start, elect, propose commands - pid = start_node(node_id, wal_path: wal_path) - Process.sleep(500) - - {:ok, _} = KRaftNode.propose(node_id, {:register_store, "s1", "http://a:8080", ["graph"]}) - {:ok, _} = KRaftNode.propose(node_id, {:register_store, "s2", "http://b:8080", ["vector"]}) - - diag1 = KRaftNode.diagnostics(node_id) - log_length_before = diag1.log_length - - stop_node(pid) - Process.sleep(100) - - # Phase 2: Restart and check recovered log - pid2 = start_node(node_id, wal_path: wal_path) - Process.sleep(100) - - diag2 = KRaftNode.diagnostics(node_id) - # Log length should match (noop + 2 commands) - assert diag2.log_length == log_length_before - - stop_node(pid2) - end - - test "recovers registry state after restart", %{wal_path: wal_path} do - node_id = unique_id("recover-registry") - - # Phase 1: Start, elect, register stores - pid = start_node(node_id, wal_path: wal_path) - Process.sleep(500) - - {:ok, _} = KRaftNode.propose(node_id, {:register_store, "s1", "http://a:8080", ["graph"]}) - {:ok, _} = KRaftNode.propose(node_id, {:register_store, "s2", "http://b:8080", ["vector"]}) - Process.sleep(100) - - registry_before = KRaftNode.registry(node_id) - assert Map.has_key?(registry_before.stores, "s1") - assert Map.has_key?(registry_before.stores, "s2") - - stop_node(pid) - Process.sleep(100) - - # Phase 2: Restart — the node needs to re-apply committed entries - # Since there's no snapshot, it replays from the WAL - pid2 = start_node(node_id, wal_path: wal_path) - Process.sleep(500) - - # After re-election and re-applying log, registry should have the stores - # Note: The recovered node starts as follower, needs to re-elect and - # re-commit entries. In a single-node cluster it will re-elect itself. - # After election, it appends a new noop and re-commits. - diag = KRaftNode.diagnostics(node_id) - - # The node should have re-elected and the log entries should be present - assert diag.log_length >= 2 - - stop_node(pid2) - end - end - - # =========================================================================== - # WAL with no persistence (nil path) - # =========================================================================== - - describe "WAL with no persistence (nil path)" do - test "node operates normally without WAL", %{} do - node_id = unique_id("no-wal") - pid = start_node(node_id, wal_path: nil) - - Process.sleep(500) - - diag = KRaftNode.diagnostics(node_id) - assert diag.role == :leader - - {:ok, _} = KRaftNode.propose(node_id, {:register_store, "s1", "http://x", []}) - - registry = KRaftNode.registry(node_id) - assert Map.has_key?(registry.stores, "s1") - - stop_node(pid) - end - end - - # =========================================================================== - # Multi-node WAL integration - # =========================================================================== - - describe "multi-node WAL integration" do - test "3-node cluster persists state across all nodes" do - suffix = System.unique_integer([:positive]) - ids = ["w1-#{suffix}", "w2-#{suffix}", "w3-#{suffix}"] - - wal_paths = - Enum.map(ids, fn id -> - path = Path.join(System.tmp_dir!(), "kraft_multi_#{id}") - File.rm_rf!(path) - path - end) - - on_exit(fn -> Enum.each(wal_paths, &File.rm_rf!/1) end) - - pids = - Enum.zip(ids, wal_paths) - |> Enum.map(fn {id, wal_path} -> - peers = Enum.reject(ids, &(&1 == id)) - start_node(id, peers: peers, wal_path: wal_path) - end) - - # Wait for election - Process.sleep(1_500) - - # Find the leader - {leader_id, _} = - ids - |> Enum.map(fn id -> {id, KRaftNode.diagnostics(id)} end) - |> Enum.find(fn {_id, diag} -> diag.role == :leader end) - - # Propose a command through the leader - {:ok, _} = - KRaftNode.propose( - leader_id, - {:register_store, "cluster-store", "http://cluster:8080", ["graph"]} - ) - - Process.sleep(500) - - # All nodes should have WAL data - Enum.each(wal_paths, fn wal_path -> - {:ok, recovered} = KRaftWAL.recover(wal_path) - assert recovered != nil - assert recovered.current_term > 0 - end) - - Enum.each(pids, &stop_node/1) - end - end -end diff --git a/verisimdb/elixir-orchestration/test/verisim/consensus/kraft_transport_test.exs b/verisimdb/elixir-orchestration/test/verisim/consensus/kraft_transport_test.exs deleted file mode 100644 index f5749db1..00000000 --- a/verisimdb/elixir-orchestration/test/verisim/consensus/kraft_transport_test.exs +++ /dev/null @@ -1,192 +0,0 @@ -# SPDX-License-Identifier: MPL-2.0 - -defmodule VeriSim.Consensus.KRaftTransportTest do - @moduledoc """ - Tests for the KRaft network transport abstraction. - - Verifies that the transport: - 1. Resolves local peers via Elixir Registry - 2. Handles remote peer tuple format - 3. Serializes/deserializes requests correctly for JSON transport - 4. Sends async RPC messages - """ - - use ExUnit.Case, async: false - - alias VeriSim.Consensus.KRaftTransport - alias VeriSim.Consensus.KRaftNode - - # =========================================================================== - # peer_id/1 - # =========================================================================== - - describe "peer_id/1" do - test "extracts ID from string" do - assert KRaftTransport.peer_id("node-1") == "node-1" - end - - test "extracts ID from tuple" do - assert KRaftTransport.peer_id({"node-1", "http://host:4000"}) == "node-1" - end - end - - # =========================================================================== - # JSON serialization - # =========================================================================== - - describe "serialize_for_json/1" do - test "converts atom keys to strings" do - result = KRaftTransport.serialize_for_json(%{term: 5, leader_id: "node-1"}) - assert result == %{"term" => 5, "leader_id" => "node-1"} - end - - test "converts atom values to strings" do - result = KRaftTransport.serialize_for_json(%{role: :leader}) - assert result == %{"role" => "leader"} - end - - test "preserves nil and boolean values" do - result = KRaftTransport.serialize_for_json(%{active: true, voted_for: nil}) - assert result == %{"active" => true, "voted_for" => nil} - end - - test "serializes nested maps" do - result = - KRaftTransport.serialize_for_json(%{ - term: 1, - entries: [%{term: 1, index: 1, command: :noop}] - }) - - assert result["entries"] == [%{"term" => 1, "index" => 1, "command" => "noop"}] - end - end - - # =========================================================================== - # Local transport — send_vote_request/2 - # =========================================================================== - - describe "send_vote_request/2 — local" do - test "sends vote request to a local KRaft node" do - node_id = "transport-vote-#{System.unique_integer([:positive])}" - {:ok, pid} = KRaftNode.start_link(node_id: node_id, peers: []) - - Process.sleep(500) - - request = %{ - term: 100, - candidate_id: "other-node", - last_log_index: 0, - last_log_term: 0 - } - - result = KRaftTransport.send_vote_request(node_id, request) - assert {:ok, response} = result - assert is_map(response) - assert Map.has_key?(response, :term) or Map.has_key?(response, "term") - - GenServer.stop(pid, :normal, 1_000) - end - - test "returns error for non-existent local peer" do - result = KRaftTransport.send_vote_request("nonexistent-node-xyz", %{}) - assert {:error, _reason} = result - end - end - - # =========================================================================== - # Local transport — send_append_entries/2 - # =========================================================================== - - describe "send_append_entries/2 — local" do - test "sends append_entries to a local KRaft node" do - node_id = "transport-ae-#{System.unique_integer([:positive])}" - {:ok, pid} = KRaftNode.start_link(node_id: node_id, peers: []) - - Process.sleep(500) - - request = %{ - term: 100, - leader_id: "other-leader", - prev_log_index: 0, - prev_log_term: 0, - entries: [], - leader_commit: 0 - } - - result = KRaftTransport.send_append_entries(node_id, request) - assert {:ok, response} = result - assert is_map(response) - - GenServer.stop(pid, :normal, 1_000) - end - end - - # =========================================================================== - # Remote peer detection - # =========================================================================== - - describe "remote peer detection" do - test "tuple peer is treated as remote" do - # This will fail with connection refused, but it should attempt HTTP - result = - KRaftTransport.send_vote_request( - {"remote-node", "http://127.0.0.1:59999"}, - %{term: 1, candidate_id: "me", last_log_index: 0, last_log_term: 0} - ) - - assert {:error, _reason} = result - end - end - - # =========================================================================== - # Async RPC - # =========================================================================== - - describe "async_vote_request/3" do - test "sends response back to caller process" do - node_id = "transport-async-#{System.unique_integer([:positive])}" - {:ok, pid} = KRaftNode.start_link(node_id: node_id, peers: []) - - Process.sleep(500) - - request = %{ - term: 100, - candidate_id: "other-node", - last_log_index: 0, - last_log_term: 0 - } - - KRaftTransport.async_vote_request(node_id, request, self()) - - assert_receive {:vote_response, ^node_id, response}, 2_000 - assert is_map(response) - - GenServer.stop(pid, :normal, 1_000) - end - end - - describe "async_append_entries/3" do - test "sends response back to caller process" do - node_id = "transport-async-ae-#{System.unique_integer([:positive])}" - {:ok, pid} = KRaftNode.start_link(node_id: node_id, peers: []) - - Process.sleep(500) - - request = %{ - term: 100, - leader_id: "other-leader", - prev_log_index: 0, - prev_log_term: 0, - entries: [], - leader_commit: 0 - } - - KRaftTransport.async_append_entries(node_id, request, self()) - - assert_receive {:append_entries_response, ^node_id, response}, 2_000 - assert is_map(response) - - GenServer.stop(pid, :normal, 1_000) - end - end -end diff --git a/verisimdb/elixir-orchestration/test/verisim/consensus/kraft_wal_test.exs b/verisimdb/elixir-orchestration/test/verisim/consensus/kraft_wal_test.exs deleted file mode 100644 index 9b159dbe..00000000 --- a/verisimdb/elixir-orchestration/test/verisim/consensus/kraft_wal_test.exs +++ /dev/null @@ -1,495 +0,0 @@ -# SPDX-License-Identifier: MPL-2.0 - -defmodule VeriSim.Consensus.KRaftWALTest do - @moduledoc """ - Tests for the KRaft Write-Ahead Log. - - Verifies that the WAL correctly: - 1. Persists and recovers durable Raft state (currentTerm, votedFor) - 2. Appends and recovers log entries - 3. Handles snapshots with WAL truncation - 4. Recovers correctly after simulated crashes - 5. Handles edge cases (empty WAL, corrupt data, missing files) - """ - - use ExUnit.Case, async: true - - alias VeriSim.Consensus.KRaftWAL - - setup do - # Create a unique temp directory for each test - dir = Path.join(System.tmp_dir!(), "kraft_wal_test_#{System.unique_integer([:positive])}") - on_exit(fn -> File.rm_rf!(dir) end) - {:ok, wal_path: dir} - end - - # =========================================================================== - # init/1 - # =========================================================================== - - describe "init/1" do - test "creates WAL directory", %{wal_path: wal_path} do - assert :ok = KRaftWAL.init(wal_path) - assert File.dir?(wal_path) - end - - test "succeeds if directory already exists", %{wal_path: wal_path} do - File.mkdir_p!(wal_path) - assert :ok = KRaftWAL.init(wal_path) - end - - test "nil path returns :ok" do - assert :ok = KRaftWAL.init(nil) - end - end - - # =========================================================================== - # persist_state/3 and recover/1 — durable state - # =========================================================================== - - describe "persist_state/3 + recover/1 — durable state" do - test "persists and recovers currentTerm and votedFor", %{wal_path: wal_path} do - KRaftWAL.init(wal_path) - KRaftWAL.persist_state(wal_path, 5, "node-2") - - {:ok, recovered} = KRaftWAL.recover(wal_path) - assert recovered.current_term == 5 - assert recovered.voted_for == "node-2" - end - - test "persists nil votedFor", %{wal_path: wal_path} do - KRaftWAL.init(wal_path) - KRaftWAL.persist_state(wal_path, 3, nil) - - {:ok, recovered} = KRaftWAL.recover(wal_path) - assert recovered.current_term == 3 - assert recovered.voted_for == nil - end - - test "overwrites previous state", %{wal_path: wal_path} do - KRaftWAL.init(wal_path) - KRaftWAL.persist_state(wal_path, 1, "node-1") - KRaftWAL.persist_state(wal_path, 2, "node-3") - - {:ok, recovered} = KRaftWAL.recover(wal_path) - assert recovered.current_term == 2 - assert recovered.voted_for == "node-3" - end - - test "nil path is a no-op" do - assert :ok = KRaftWAL.persist_state(nil, 5, "node-2") - end - end - - # =========================================================================== - # append_entry/2 and recover/1 — log entries - # =========================================================================== - - describe "append_entry/2 + recover/1 — log entries" do - test "appends and recovers a single entry", %{wal_path: wal_path} do - KRaftWAL.init(wal_path) - - entry = %{ - term: 1, - index: 1, - command: {:register_store, "s1", "http://localhost:9000", ["graph"]}, - timestamp: 1_000_000 - } - - assert :ok = KRaftWAL.append_entry(wal_path, entry) - - {:ok, recovered} = KRaftWAL.recover(wal_path) - assert length(recovered.log) == 1 - - [recovered_entry] = recovered.log - assert recovered_entry.term == 1 - assert recovered_entry.index == 1 - assert recovered_entry.command == {:register_store, "s1", "http://localhost:9000", ["graph"]} - end - - test "appends multiple entries sequentially", %{wal_path: wal_path} do - KRaftWAL.init(wal_path) - - for i <- 1..5 do - entry = %{ - term: 1, - index: i, - command: :noop, - timestamp: 1_000_000 + i - } - - KRaftWAL.append_entry(wal_path, entry) - end - - {:ok, recovered} = KRaftWAL.recover(wal_path) - assert length(recovered.log) == 5 - assert Enum.map(recovered.log, & &1.index) == [1, 2, 3, 4, 5] - end - - test "nil path is a no-op" do - entry = %{term: 1, index: 1, command: :noop} - assert :ok = KRaftWAL.append_entry(nil, entry) - end - end - - # =========================================================================== - # append_entries/2 — batch append - # =========================================================================== - - describe "append_entries/2 — batch append" do - test "appends multiple entries in one write", %{wal_path: wal_path} do - KRaftWAL.init(wal_path) - - entries = - for i <- 1..3 do - %{term: 1, index: i, command: :noop, timestamp: 1_000_000 + i} - end - - assert :ok = KRaftWAL.append_entries(wal_path, entries) - - {:ok, recovered} = KRaftWAL.recover(wal_path) - assert length(recovered.log) == 3 - end - - test "empty list is a no-op", %{wal_path: wal_path} do - KRaftWAL.init(wal_path) - assert :ok = KRaftWAL.append_entries(wal_path, []) - end - end - - # =========================================================================== - # Command serialization round-trip - # =========================================================================== - - describe "command serialization round-trip" do - test "noop", %{wal_path: wal_path} do - KRaftWAL.init(wal_path) - entry = %{term: 1, index: 1, command: :noop} - KRaftWAL.append_entry(wal_path, entry) - - {:ok, recovered} = KRaftWAL.recover(wal_path) - assert hd(recovered.log).command == :noop - end - - test "register_store", %{wal_path: wal_path} do - KRaftWAL.init(wal_path) - - entry = %{ - term: 1, - index: 1, - command: {:register_store, "store-1", "http://host:8080", ["graph", "vector"]} - } - - KRaftWAL.append_entry(wal_path, entry) - - {:ok, recovered} = KRaftWAL.recover(wal_path) - assert hd(recovered.log).command == - {:register_store, "store-1", "http://host:8080", ["graph", "vector"]} - end - - test "unregister_store", %{wal_path: wal_path} do - KRaftWAL.init(wal_path) - entry = %{term: 1, index: 1, command: {:unregister_store, "store-1"}} - KRaftWAL.append_entry(wal_path, entry) - - {:ok, recovered} = KRaftWAL.recover(wal_path) - assert hd(recovered.log).command == {:unregister_store, "store-1"} - end - - test "map_octad", %{wal_path: wal_path} do - KRaftWAL.init(wal_path) - locations = ["store-1", "store-2", "store-3"] - entry = %{term: 1, index: 1, command: {:map_octad, "hex-1", locations}} - KRaftWAL.append_entry(wal_path, entry) - - {:ok, recovered} = KRaftWAL.recover(wal_path) - assert hd(recovered.log).command == {:map_octad, "hex-1", locations} - end - - test "unmap_octad", %{wal_path: wal_path} do - KRaftWAL.init(wal_path) - entry = %{term: 1, index: 1, command: {:unmap_octad, "hex-1"}} - KRaftWAL.append_entry(wal_path, entry) - - {:ok, recovered} = KRaftWAL.recover(wal_path) - assert hd(recovered.log).command == {:unmap_octad, "hex-1"} - end - - test "update_trust", %{wal_path: wal_path} do - KRaftWAL.init(wal_path) - entry = %{term: 1, index: 1, command: {:update_trust, "store-1", 0.85}} - KRaftWAL.append_entry(wal_path, entry) - - {:ok, recovered} = KRaftWAL.recover(wal_path) - assert hd(recovered.log).command == {:update_trust, "store-1", 0.85} - end - end - - # =========================================================================== - # save_snapshot/4 + recover/1 - # =========================================================================== - - describe "save_snapshot/4 + recover/1" do - test "saves and recovers snapshot with registry state", %{wal_path: wal_path} do - KRaftWAL.init(wal_path) - - # Write 5 entries - for i <- 1..5 do - entry = %{term: 1, index: i, command: :noop} - KRaftWAL.append_entry(wal_path, entry) - end - - # Create a snapshot at index 3 - registry = %{ - stores: %{ - "s1" => %{ - store_id: "s1", - endpoint: "http://localhost:9000", - modalities: ["graph"], - trust_level: 0.95 - } - }, - mappings: %{} - } - - KRaftWAL.save_snapshot(wal_path, registry, 3, 1) - - {:ok, recovered} = KRaftWAL.recover(wal_path) - - # Snapshot state - assert recovered.snapshot_index == 3 - assert recovered.snapshot_term == 1 - assert recovered.registry.stores["s1"].endpoint == "http://localhost:9000" - assert recovered.registry.stores["s1"].trust_level == 0.95 - - # Only entries after snapshot should remain in log - assert length(recovered.log) == 2 - assert Enum.map(recovered.log, & &1.index) == [4, 5] - end - - test "snapshot truncates WAL entries at or before snapshot index", %{wal_path: wal_path} do - KRaftWAL.init(wal_path) - - for i <- 1..10 do - entry = %{term: 1, index: i, command: :noop} - KRaftWAL.append_entry(wal_path, entry) - end - - KRaftWAL.save_snapshot(wal_path, %{stores: %{}, mappings: %{}}, 7, 1) - - {:ok, recovered} = KRaftWAL.recover(wal_path) - assert length(recovered.log) == 3 - assert Enum.map(recovered.log, & &1.index) == [8, 9, 10] - end - end - - # =========================================================================== - # truncate_after/2 - # =========================================================================== - - describe "truncate_after/2" do - test "keeps entries up to given index", %{wal_path: wal_path} do - KRaftWAL.init(wal_path) - - for i <- 1..5 do - entry = %{term: 1, index: i, command: :noop} - KRaftWAL.append_entry(wal_path, entry) - end - - KRaftWAL.truncate_after(wal_path, 3) - - {:ok, recovered} = KRaftWAL.recover(wal_path) - assert length(recovered.log) == 3 - assert Enum.map(recovered.log, & &1.index) == [1, 2, 3] - end - - test "truncate to 0 removes all entries", %{wal_path: wal_path} do - KRaftWAL.init(wal_path) - - for i <- 1..3 do - entry = %{term: 1, index: i, command: :noop} - KRaftWAL.append_entry(wal_path, entry) - end - - KRaftWAL.truncate_after(wal_path, 0) - - {:ok, recovered} = KRaftWAL.recover(wal_path) - assert recovered.log == [] - end - - test "truncate + append simulates follower log conflict resolution", %{wal_path: wal_path} do - KRaftWAL.init(wal_path) - - # Follower has entries from term 1 - for i <- 1..5 do - entry = %{term: 1, index: i, command: :noop} - KRaftWAL.append_entry(wal_path, entry) - end - - # New leader sends entries starting at index 3 with term 2 - KRaftWAL.truncate_after(wal_path, 2) - - new_entries = - for i <- 3..6 do - %{term: 2, index: i, command: :noop} - end - - KRaftWAL.append_entries(wal_path, new_entries) - - {:ok, recovered} = KRaftWAL.recover(wal_path) - assert length(recovered.log) == 6 - assert Enum.map(recovered.log, & &1.index) == [1, 2, 3, 4, 5, 6] - # First 2 entries from term 1, remaining from term 2 - assert Enum.map(recovered.log, & &1.term) == [1, 1, 2, 2, 2, 2] - end - - test "nil path is a no-op" do - assert :ok = KRaftWAL.truncate_after(nil, 5) - end - end - - # =========================================================================== - # Recovery edge cases - # =========================================================================== - - describe "recovery edge cases" do - test "recover from non-existent directory returns nil", %{wal_path: wal_path} do - {:ok, result} = KRaftWAL.recover(wal_path) - assert result == nil - end - - test "recover from empty directory returns defaults", %{wal_path: wal_path} do - KRaftWAL.init(wal_path) - - {:ok, recovered} = KRaftWAL.recover(wal_path) - assert recovered.current_term == 0 - assert recovered.voted_for == nil - assert recovered.log == [] - assert recovered.registry == nil - assert recovered.snapshot_index == 0 - assert recovered.snapshot_term == 0 - end - - test "recover skips corrupt WAL lines", %{wal_path: wal_path} do - KRaftWAL.init(wal_path) - - # Write a valid entry - entry = %{term: 1, index: 1, command: :noop} - KRaftWAL.append_entry(wal_path, entry) - - # Append a corrupt line directly - wal_file = Path.join(wal_path, "wal.jsonl") - File.write!(wal_file, "this is not valid json\n", [:append]) - - # Write another valid entry - entry2 = %{term: 1, index: 2, command: :noop} - KRaftWAL.append_entry(wal_path, entry2) - - {:ok, recovered} = KRaftWAL.recover(wal_path) - # Should recover 2 valid entries, skipping the corrupt line - assert length(recovered.log) == 2 - end - - test "recover with nil path returns nil" do - {:ok, result} = KRaftWAL.recover(nil) - assert result == nil - end - end - - # =========================================================================== - # Full crash recovery simulation - # =========================================================================== - - describe "full crash recovery simulation" do - test "recovers complete state after simulated crash", %{wal_path: wal_path} do - # Phase 1: Normal operation - KRaftWAL.init(wal_path) - KRaftWAL.persist_state(wal_path, 3, "node-2") - - entries = [ - %{term: 1, index: 1, command: {:register_store, "s1", "http://a:8080", ["graph"]}}, - %{term: 1, index: 2, command: {:register_store, "s2", "http://b:8080", ["vector"]}}, - %{term: 2, index: 3, command: {:map_octad, "h1", ["s1", "s2"]}}, - %{term: 3, index: 4, command: :noop}, - %{term: 3, index: 5, command: {:update_trust, "s1", 0.9}} - ] - - for entry <- entries, do: KRaftWAL.append_entry(wal_path, entry) - - # Phase 2: "Crash" — just forget everything in memory - - # Phase 3: Recovery - {:ok, recovered} = KRaftWAL.recover(wal_path) - - assert recovered.current_term == 3 - assert recovered.voted_for == "node-2" - assert length(recovered.log) == 5 - - # Verify all commands survived - commands = Enum.map(recovered.log, & &1.command) - - assert Enum.at(commands, 0) == - {:register_store, "s1", "http://a:8080", ["graph"]} - - assert Enum.at(commands, 1) == - {:register_store, "s2", "http://b:8080", ["vector"]} - - assert Enum.at(commands, 2) == {:map_octad, "h1", ["s1", "s2"]} - assert Enum.at(commands, 3) == :noop - assert Enum.at(commands, 4) == {:update_trust, "s1", 0.9} - end - - test "recovers after snapshot + additional entries", %{wal_path: wal_path} do - KRaftWAL.init(wal_path) - KRaftWAL.persist_state(wal_path, 5, nil) - - # Write entries 1-10 - for i <- 1..10 do - entry = %{term: div(i - 1, 3) + 1, index: i, command: :noop} - KRaftWAL.append_entry(wal_path, entry) - end - - # Snapshot at index 7 - registry = %{ - stores: %{ - "s1" => %{ - store_id: "s1", - endpoint: "http://host:8080", - modalities: ["graph", "vector"], - trust_level: 0.85 - } - }, - mappings: %{ - "h1" => %{ - octad_id: "h1", - locations: ["s1"], - primary_store: "s1" - } - } - } - - KRaftWAL.save_snapshot(wal_path, registry, 7, 3) - - # Add entries 11-13 after snapshot - for i <- 11..13 do - entry = %{term: 5, index: i, command: :noop} - KRaftWAL.append_entry(wal_path, entry) - end - - # "Crash" and recover - {:ok, recovered} = KRaftWAL.recover(wal_path) - - assert recovered.current_term == 5 - assert recovered.snapshot_index == 7 - assert recovered.snapshot_term == 3 - - # Log should contain entries 8-13 (after snapshot) - assert length(recovered.log) == 6 - assert Enum.map(recovered.log, & &1.index) == [8, 9, 10, 11, 12, 13] - - # Registry from snapshot - assert recovered.registry.stores["s1"].trust_level == 0.85 - assert recovered.registry.mappings["h1"].primary_store == "s1" - end - end -end diff --git a/verisimdb/elixir-orchestration/test/verisim/e2e_verisimdb_test.exs b/verisimdb/elixir-orchestration/test/verisim/e2e_verisimdb_test.exs deleted file mode 100644 index 918e72f2..00000000 --- a/verisimdb/elixir-orchestration/test/verisim/e2e_verisimdb_test.exs +++ /dev/null @@ -1,410 +0,0 @@ -# SPDX-License-Identifier: MPL-2.0 -# Author: Jonathan D.A. Jewell -# -# VeriSimDB E2E Tests — Full lifecycle via the Elixir orchestration layer. -# -# These tests exercise the complete observable behaviour of VeriSimDB from -# the perspective of an external caller: they write data through the public -# API, read it back, and assert on the observed results. They are designed -# to run with and without the Rust core: -# -# - Rust available: full round-trip with real data. -# - Rust unavailable: orchestration layer correctness is still verified -# (query routing, schema validation, drift monitoring lifecycle). -# -# Test categories: -# -# 1. Lifecycle — write data → read back → verify consistency. -# 2. VCL — store data → execute VCL query → validate results. -# 3. Schema — create schema → insert conforming data → -# reject non-conforming data. -# 4. Error handling — invalid VCL syntax → clear error, connection -# failure → graceful degradation. - -defmodule VeriSimDB.E2ETest do - @moduledoc """ - End-to-end tests for VeriSimDB. - - Exercises the complete stack: Elixir orchestration → Rust core - (when available) → VCL layer → schema registry → drift monitor. - """ - - use ExUnit.Case, async: false - - alias VeriSim.{ - EntityServer, - DriftMonitor, - SchemaRegistry, - RustClient, - QueryRouter - } - alias VeriSim.Query.{VCLBridge, VCLExecutor} - alias VeriSim.Test.VCLTestHelpers, as: H - - # Tag for tests that require the Rust core. - @moduletag :e2e - - setup_all do - {:ok, _} = Application.ensure_all_started(:verisim) - - # Ensure VCLBridge GenServer is running (it is not started by the - # application supervisor in test mode since it depends on Deno, which - # may not be available; H.ensure_bridge_started/0 starts it safely). - _bridge_pid = H.ensure_bridge_started() - - # Determine whether Rust core is reachable. - rust_available = - case RustClient.health() do - {:ok, _} -> true - {:error, _} -> false - end - - %{rust_available: rust_available} - end - - # =========================================================================== - # 1. Lifecycle: write data → read back → verify consistency - # =========================================================================== - - describe "lifecycle: write, read back, verify consistency" do - test "create entity via EntityServer and read it back", %{rust_available: _rust} do - entity_id = "e2e-lifecycle-#{System.unique_integer([:positive])}" - - # Start an entity server for this entity. - {:ok, _pid} = EntityServer.start_link(entity_id) - - # Verify initial state. - {:ok, initial_state} = EntityServer.get(entity_id) - assert initial_state.id == entity_id - assert initial_state.status == :active - assert initial_state.version == 0 - - # Update the entity to enable document and vector modalities. - {:ok, updated_state} = EntityServer.update(entity_id, [ - {:modality, :document, true}, - {:modality, :vector, true} - ]) - - assert updated_state.modalities.document == true - assert updated_state.modalities.vector == true - assert updated_state.version >= 1 - - # Read it back again to confirm persistence within the session. - {:ok, read_back} = EntityServer.get(entity_id) - assert read_back.id == entity_id - assert read_back.modalities.document == true - end - - test "create entity via RustClient, read back from same session", %{rust_available: rust_available} do - if not rust_available do - # Rust core unavailable — test the graceful degradation path instead. - result = RustClient.create_octad(%{ - title: "E2E Lifecycle Test", - body: "Testing full lifecycle without Rust core", - embedding: List.duplicate(0.5, 384) - }) - - # Must return a structured error, not raise. - assert match?({:error, _}, result), - "Expected {:error, _} when Rust core unavailable, got: #{inspect(result)}" - else - input = %{ - title: "E2E Lifecycle Test", - body: "Testing full lifecycle: write → read → verify", - embedding: List.duplicate(0.3, 384), - types: ["verisim:Document"] - } - - {:ok, %{"id" => entity_id}} = RustClient.create_octad(input) - assert is_binary(entity_id) - - # Read back. - {:ok, octad} = RustClient.get_octad(entity_id) - assert octad["id"] == entity_id - assert octad["document"]["title"] == "E2E Lifecycle Test" - assert octad["document"]["body"] == "Testing full lifecycle: write → read → verify" - end - end - - test "drift monitoring is notified when entity changes", _ctx do - entity_id = "e2e-drift-lifecycle-#{System.unique_integer([:positive])}" - - # Report a moderate drift event (should not trigger immediate normalization). - DriftMonitor.report_drift(entity_id, 0.3, :semantic_vector) - DriftMonitor.entity_changed(entity_id) - - Process.sleep(100) - - status = DriftMonitor.status() - assert status[:overall_health] in [:healthy, :warning, :degraded, :critical], - "DriftMonitor returned invalid health status: #{inspect(status[:overall_health])}" - end - end - - # =========================================================================== - # 2. VCL: store data → execute VCL query → validate results - # =========================================================================== - - describe "VCL: end-to-end query execution" do - test "VCL parse → execute produces structured result for valid query", _ctx do - query = "SELECT DOCUMENT.* FROM HEXAD 'e2e-vcl-001' LIMIT 5" - - # Parse the query. - {:ok, ast} = VCLBridge.parse(query) - assert is_map(ast) - assert :document in (ast[:modalities] || []) - - # Execute through the full pipeline. - result = VCLExecutor.execute(ast) - - # Must be structured — either results or a clean error. - assert match?({:ok, _}, result) or match?({:error, _}, result), - "VCL execute returned non-structured result: #{inspect(result)}" - end - - test "VCL multi-modal query selects correct modalities", _ctx do - query = "SELECT GRAPH.*, VECTOR.* FROM HEXAD 'e2e-multimodal-001' LIMIT 3" - {:ok, ast} = VCLBridge.parse(query) - - assert :graph in (ast[:modalities] || []) - assert :vector in (ast[:modalities] || []) - refute :document in (ast[:modalities] || []), - "SELECT GRAPH, VECTOR should not include :document" - end - - test "VCL INSERT → read-back round-trip (when Rust available)", %{rust_available: rust_available} do - if not rust_available do - # Execute the mutation path, verify graceful error. - mutation_query = "INSERT HEXAD WITH DOCUMENT(title = 'E2E Insert Test', body = 'VCL insert round-trip')" - - result = VCLBridge.parse_statement(mutation_query) - case result do - {:ok, ast} -> - exec_result = VCLExecutor.execute_statement(ast) - assert match?({:ok, _}, exec_result) or match?({:error, _}, exec_result) - - {:error, _} -> - # Parse failure accepted — Deno VCL parser may not be running. - assert true - end - else - # Full round-trip with Rust. - mutation_query = "INSERT HEXAD WITH DOCUMENT(title = 'E2E Insert Test', body = 'VCL insert round-trip')" - - case VCLBridge.parse_statement(mutation_query) do - {:ok, ast} -> - case VCLExecutor.execute_statement(ast) do - {:ok, %{"id" => entity_id}} -> - # Read the entity back via RustClient. - {:ok, octad} = RustClient.get_octad(entity_id) - assert octad["document"]["title"] == "E2E Insert Test" - - {:ok, _other} -> - # Different return shape — acceptable. - assert true - - {:error, _reason} -> - # Insert failed — acceptable if schema constraints not met. - assert true - end - - {:error, _} -> - assert true - end - end - end - - test "VCL EXPLAIN returns plan with required fields", _ctx do - ast = %{ - modalities: [:document, :vector], - source: {:octad, "e2e-explain-001"}, - where: nil, - proof: nil, - limit: 10, - offset: 0 - } - - {:ok, plan} = VCLExecutor.execute(ast, explain: true) - - assert is_map(plan) - assert Map.has_key?(plan, :strategy), - "Explain plan missing :strategy field: #{inspect(plan)}" - assert Map.has_key?(plan, :steps), - "Explain plan missing :steps field: #{inspect(plan)}" - assert is_list(plan[:steps]) - end - - test "VCL error path: invalid syntax returns structured error", _ctx do - query = "COMPLETELY INVALID VCL $$$ @@@" - - result = VCLBridge.parse(query) - - # Must be {:error, reason}, not a raise. - assert match?({:error, _}, result), - "Expected parse error for invalid VCL, got: #{inspect(result)}" - end - - test "VCL error path: nonexistent store returns clean error", _ctx do - ast = %{ - modalities: [:document], - source: {:store, "definitely-not-a-store-#{System.unique_integer([:positive])}"}, - where: nil, - proof: nil, - limit: 10, - offset: 0 - } - - result = VCLExecutor.execute(ast) - - # Must be {:ok, []} (empty) or {:error, _} — never crash. - assert match?({:ok, _}, result) or match?({:error, _}, result) - end - end - - # =========================================================================== - # 3. Schema: create schema → insert conforming → reject non-conforming - # =========================================================================== - - describe "schema: type registration, validation, hierarchy" do - test "register a new type and retrieve it" do - suffix = System.unique_integer([:positive]) - - type_def = %{ - iri: "https://e2e.test.org/E2EDocument-#{suffix}", - label: "E2E Document Type", - supertypes: ["verisim:Entity"], - constraints: [ - %{ - name: "title_required", - kind: {:required, "title"}, - message: "E2E documents must have a title" - } - ] - } - - assert :ok == SchemaRegistry.register_type(type_def) - - # Retrieve and verify. - retrieved = SchemaRegistry.get_type("https://e2e.test.org/E2EDocument-#{suffix}") - assert not is_nil(retrieved), - "Registered type should be retrievable immediately" - assert retrieved.label == "E2E Document Type" - end - - test "valid entity passes schema validation" do - valid_entity = %{ - types: ["verisim:Document"], - properties: %{"title" => "A valid document title"} - } - - assert :ok == SchemaRegistry.validate(valid_entity), - "Valid entity should pass schema validation" - end - - test "entity missing required field fails validation" do - # verisim:Document has an implicit :required constraint for :title - # (via the built-in schema). If the implementation doesn't enforce this, - # it returns :ok — we accept that and note it as a gap. - invalid_entity = %{ - types: ["verisim:Document"], - properties: %{} - } - - case SchemaRegistry.validate(invalid_entity) do - {:error, violations} -> - assert is_list(violations) and length(violations) > 0, - "Expected at least one violation for entity missing required field" - - :ok -> - # Schema may not enforce :required at this layer — acceptable. - # Document in TEST-NEEDS.md as a known gap. - assert true - end - end - - test "type hierarchy includes entity root" do - hierarchy = SchemaRegistry.type_hierarchy("verisim:Document") - assert is_list(hierarchy) - assert "verisim:Document" in hierarchy - assert "verisim:Entity" in hierarchy, - "Type hierarchy must include verisim:Entity as root" - end - - test "unknown type returns empty or nil hierarchy without crash" do - result = SchemaRegistry.type_hierarchy("https://unknown.example.org/NeverRegistered") - - # Must be a list (possibly empty) or nil — never crash. - assert is_list(result) or is_nil(result), - "Unknown type hierarchy must return list or nil, got: #{inspect(result)}" - end - end - - # =========================================================================== - # 4. Error Handling: graceful degradation on failures - # =========================================================================== - - describe "error handling: graceful degradation" do - test "RustClient returns structured error when core is unavailable" do - # Force a connection attempt to a non-existent endpoint. - # We test this indirectly: if the Rust core is down, create_octad - # should return {:error, _}, not raise. - result = - try do - RustClient.create_octad(%{ - title: "Error test", - body: "Testing graceful degradation", - embedding: List.duplicate(0.0, 384) - }) - rescue - e -> {:crashed, e} - end - - refute match?({:crashed, _}, result), - "RustClient raised an exception instead of returning {:error, _}: #{inspect(result)}" - - assert match?({:ok, _}, result) or match?({:error, _}, result) - end - - test "QueryRouter returns structured result even with no stores registered" do - params = %{text: "test query", vector: List.duplicate(0.1, 384)} - - result = QueryRouter.query(:multi, params, limit: 5) - - assert match?({:ok, _}, result) or match?({:error, _}, result), - "QueryRouter must not crash when no stores registered" - end - - test "DriftMonitor handles extreme drift score without crashing" do - entity_id = "e2e-extreme-drift-#{System.unique_integer([:positive])}" - - # These should be clamped or handled, not crash. - assert :ok == DriftMonitor.report_drift(entity_id, 1.0, :quality) - assert :ok == DriftMonitor.report_drift(entity_id, 0.0, :quality) - - Process.sleep(50) - status = DriftMonitor.status() - assert is_map(status) - end - - test "VCL executor handles nil AST fields without crash" do - # AST with minimal fields — should fail gracefully, not crash. - # We wrap in try/rescue because the executor may raise a FunctionClauseError - # when source is nil (depending on the implementation). The critical - # invariant is that the calling process does not crash silently. - minimal_ast = %{modalities: [], source: nil} - - result = - try do - VCLExecutor.execute(minimal_ast) - rescue - e -> {:error, {:executor_raised, inspect(e)}} - catch - :exit, reason -> {:error, {:executor_exit, inspect(reason)}} - end - - # Any structured result is acceptable; the process must not die. - assert match?({:ok, _}, result) or match?({:error, _}, result), - "Executor must not produce unexpected value for nil AST: #{inspect(result)}" - end - end -end diff --git a/verisimdb/elixir-orchestration/test/verisim/federation/adapter_test.exs b/verisimdb/elixir-orchestration/test/verisim/federation/adapter_test.exs deleted file mode 100644 index d11ae1f5..00000000 --- a/verisimdb/elixir-orchestration/test/verisim/federation/adapter_test.exs +++ /dev/null @@ -1,428 +0,0 @@ -# SPDX-License-Identifier: MPL-2.0 - -defmodule VeriSim.Federation.AdapterTest do - @moduledoc """ - Tests for the heterogeneous federation adapter system. - - Tests adapter behaviour compliance, modality declarations, result - normalisation, and resolver integration with mixed adapter types. - """ - - use ExUnit.Case, async: false - - alias VeriSim.Federation.Adapter - alias VeriSim.Federation.Adapters.{ArangoDB, Elasticsearch, PostgreSQL, VeriSimDB} - alias VeriSim.Federation.Adapters.{MongoDB, Redis, DuckDB, ClickHouse, SurrealDB} - alias VeriSim.Federation.Adapters.{SQLite, Neo4j, VectorDB, InfluxDB, ObjectStorage} - alias VeriSim.Federation.Resolver - - # --------------------------------------------------------------------------- - # Adapter Registry - # --------------------------------------------------------------------------- - - describe "Adapter.module_for/1" do - test "resolves all known adapter types" do - assert {:ok, VeriSimDB} = Adapter.module_for(:verisimdb) - assert {:ok, ArangoDB} = Adapter.module_for(:arangodb) - assert {:ok, PostgreSQL} = Adapter.module_for(:postgresql) - assert {:ok, Elasticsearch} = Adapter.module_for(:elasticsearch) - assert {:ok, MongoDB} = Adapter.module_for(:mongodb) - assert {:ok, Redis} = Adapter.module_for(:redis) - assert {:ok, DuckDB} = Adapter.module_for(:duckdb) - assert {:ok, ClickHouse} = Adapter.module_for(:clickhouse) - assert {:ok, SurrealDB} = Adapter.module_for(:surrealdb) - assert {:ok, SQLite} = Adapter.module_for(:sqlite) - assert {:ok, Neo4j} = Adapter.module_for(:neo4j) - assert {:ok, VectorDB} = Adapter.module_for(:vector_db) - assert {:ok, InfluxDB} = Adapter.module_for(:influxdb) - assert {:ok, ObjectStorage} = Adapter.module_for(:object_storage) - end - - test "returns error for unknown adapter type" do - assert {:error, :unknown_adapter} = Adapter.module_for(:unknown) - assert {:error, :unknown_adapter} = Adapter.module_for(:cassandra) - end - - test "adapter_types/0 lists all 14 registered types" do - types = Adapter.adapter_types() - assert :verisimdb in types - assert :arangodb in types - assert :postgresql in types - assert :elasticsearch in types - assert :mongodb in types - assert :redis in types - assert :duckdb in types - assert :clickhouse in types - assert :surrealdb in types - assert :sqlite in types - assert :neo4j in types - assert :vector_db in types - assert :influxdb in types - assert :object_storage in types - assert length(types) == 14 - end - end - - # --------------------------------------------------------------------------- - # Supported Modalities - # --------------------------------------------------------------------------- - - describe "supported_modalities/1" do - test "VeriSimDB supports all 8 octad modalities" do - modalities = VeriSimDB.supported_modalities(%{}) - - assert :graph in modalities - assert :vector in modalities - assert :tensor in modalities - assert :semantic in modalities - assert :document in modalities - assert :temporal in modalities - assert :provenance in modalities - assert :spatial in modalities - assert length(modalities) == 8 - end - - test "ArangoDB supports graph, document, semantic, temporal, provenance, spatial" do - modalities = ArangoDB.supported_modalities(%{}) - - assert :graph in modalities - assert :document in modalities - assert :semantic in modalities - assert :temporal in modalities - assert :provenance in modalities - assert :spatial in modalities - - # ArangoDB does NOT support vector or tensor - refute :vector in modalities - refute :tensor in modalities - assert length(modalities) == 6 - end - - test "PostgreSQL base modalities without extensions" do - modalities = PostgreSQL.supported_modalities(%{}) - - assert :document in modalities - assert :semantic in modalities - assert :temporal in modalities - assert :graph in modalities - assert :provenance in modalities - - # Without extensions, no vector or spatial - refute :vector in modalities - refute :spatial in modalities - assert length(modalities) == 5 - end - - test "PostgreSQL with pgvector adds vector modality" do - modalities = PostgreSQL.supported_modalities(%{extensions: [:pgvector]}) - - assert :vector in modalities - assert :document in modalities - refute :spatial in modalities - assert length(modalities) == 6 - end - - test "PostgreSQL with PostGIS adds spatial modality" do - modalities = PostgreSQL.supported_modalities(%{extensions: [:postgis]}) - - assert :spatial in modalities - assert :document in modalities - refute :vector in modalities - assert length(modalities) == 6 - end - - test "PostgreSQL with both pgvector and PostGIS" do - modalities = PostgreSQL.supported_modalities(%{extensions: [:pgvector, :postgis]}) - - assert :vector in modalities - assert :spatial in modalities - assert length(modalities) == 7 - end - - test "Elasticsearch supports document, vector, semantic, temporal, spatial" do - modalities = Elasticsearch.supported_modalities(%{}) - - assert :document in modalities - assert :vector in modalities - assert :semantic in modalities - assert :temporal in modalities - assert :spatial in modalities - - # Elasticsearch does NOT support graph, tensor, or provenance - refute :graph in modalities - refute :tensor in modalities - refute :provenance in modalities - assert length(modalities) == 5 - end - end - - # --------------------------------------------------------------------------- - # Result Normalisation - # --------------------------------------------------------------------------- - - describe "translate_results/2" do - @peer_info %{store_id: "test-peer", endpoint: "http://test:8080", adapter_config: %{}} - - test "VeriSimDB normalises results with id and score" do - raw = [%{"id" => "abc-123", "score" => 0.95, "title" => "Test"}] - - [result] = VeriSimDB.translate_results(raw, @peer_info) - - assert result.source_store == "test-peer" - assert result.octad_id == "abc-123" - assert result.score == 0.95 - assert result.drifted == false - assert result.data == hd(raw) - end - - test "VeriSimDB handles missing id gracefully" do - raw = [%{"title" => "No ID"}] - - [result] = VeriSimDB.translate_results(raw, @peer_info) - - assert result.octad_id == "unknown" - assert result.score == 0.0 - end - - test "ArangoDB extracts _key from ArangoDB documents" do - raw = [%{"_key" => "doc-456", "_id" => "octads/doc-456", "_score" => 1.5}] - - [result] = ArangoDB.translate_results(raw, @peer_info) - - assert result.octad_id == "doc-456" - assert result.score == 1.5 - assert result.source_store == "test-peer" - end - - test "ArangoDB falls back to _id when _key missing" do - raw = [%{"_id" => "octads/fallback-789"}] - - [result] = ArangoDB.translate_results(raw, @peer_info) - - assert result.octad_id == "octads/fallback-789" - end - - test "PostgreSQL normalises row results" do - raw = [%{"id" => "pg-001", "score" => 0.88, "title" => "PostgreSQL doc"}] - - [result] = PostgreSQL.translate_results(raw, @peer_info) - - assert result.octad_id == "pg-001" - assert result.score == 0.88 - end - - test "PostgreSQL handles entity_id field" do - raw = [%{"entity_id" => "pg-alt", "score" => 0.5}] - - [result] = PostgreSQL.translate_results(raw, @peer_info) - - assert result.octad_id == "pg-alt" - end - - test "Elasticsearch extracts from _source and _id" do - raw = [ - %{ - "_id" => "es-doc-1", - "_score" => 2.3, - "_source" => %{"title" => "ES document", "body" => "Content"} - } - ] - - [result] = Elasticsearch.translate_results(raw, @peer_info) - - assert result.octad_id == "es-doc-1" - assert result.score == 2.3 - assert result.data == %{"title" => "ES document", "body" => "Content"} - end - - test "Elasticsearch handles null _score" do - raw = [%{"_id" => "es-null-score", "_score" => nil, "_source" => %{}}] - - [result] = Elasticsearch.translate_results(raw, @peer_info) - - assert result.score == 0.0 - end - - test "all adapters handle empty result lists" do - assert VeriSimDB.translate_results([], @peer_info) == [] - assert ArangoDB.translate_results([], @peer_info) == [] - assert PostgreSQL.translate_results([], @peer_info) == [] - assert Elasticsearch.translate_results([], @peer_info) == [] - end - end - - # --------------------------------------------------------------------------- - # Resolver Integration — Mixed Adapter Types - # --------------------------------------------------------------------------- - - describe "resolver with heterogeneous peers" do - setup do - # Clear any peers from previous tests - for peer <- Resolver.list_peers() do - Resolver.deregister_peer(peer.store_id) - end - - :ok - end - - test "register VeriSimDB peer via 3-arity (backward-compatible)" do - :ok = Resolver.register_peer("verisim-1", "http://v1:8080", ["graph", "vector"]) - - [peer] = Resolver.list_peers() - assert peer.store_id == "verisim-1" - assert peer.adapter_type == :verisimdb - assert "graph" in peer.modalities - assert "vector" in peer.modalities - end - - test "register ArangoDB peer via 2-arity map" do - :ok = - Resolver.register_peer("arango-1", %{ - endpoint: "http://arango:8529", - adapter_type: :arangodb, - adapter_config: %{database: "_system", collection: "entities"}, - modalities: ["graph", "document", "semantic"] - }) - - [peer] = Resolver.list_peers() - assert peer.store_id == "arango-1" - assert peer.adapter_type == :arangodb - assert "graph" in peer.modalities - assert "document" in peer.modalities - assert "semantic" in peer.modalities - end - - test "register PostgreSQL peer with extension-based modality validation" do - :ok = - Resolver.register_peer("pg-1", %{ - endpoint: "http://pg-proxy:3000", - adapter_type: :postgresql, - adapter_config: %{ - database: "verisimdb", - table: "octads", - extensions: [:pgvector, :postgis] - }, - modalities: ["document", "vector", "spatial", "tensor"] - }) - - [peer] = Resolver.list_peers() - assert peer.store_id == "pg-1" - assert peer.adapter_type == :postgresql - - # tensor is NOT supported by PostgreSQL — should be filtered out - refute "tensor" in peer.modalities - assert "document" in peer.modalities - assert "vector" in peer.modalities - assert "spatial" in peer.modalities - end - - test "register Elasticsearch peer" do - :ok = - Resolver.register_peer("es-1", %{ - endpoint: "http://elastic:9200", - adapter_type: :elasticsearch, - adapter_config: %{index: "octads", version: 8}, - modalities: ["document", "vector"] - }) - - [peer] = Resolver.list_peers() - assert peer.store_id == "es-1" - assert peer.adapter_type == :elasticsearch - assert "document" in peer.modalities - assert "vector" in peer.modalities - end - - test "reject unknown adapter type" do - result = - Resolver.register_peer("bad-adapter", %{ - endpoint: "http://mystery:1234", - adapter_type: :cassandra, - modalities: ["document"] - }) - - assert {:error, {:unknown_adapter, :cassandra}} = result - assert Resolver.list_peers() == [] - end - - test "mixed adapter types in federation query" do - # Register a VeriSimDB peer and an ArangoDB peer - :ok = Resolver.register_peer("verisim-peer", "http://v:8080", ["document"]) - - :ok = - Resolver.register_peer("arango-peer", %{ - endpoint: "http://a:8529", - adapter_type: :arangodb, - adapter_config: %{database: "_system"}, - modalities: ["document", "graph"] - }) - - # Query for document modality — both peers should match - {:ok, response} = Resolver.query("*", ["document"], timeout: 2_000) - - assert length(response.stores_queried) == 2 - assert "verisim-peer" in response.stores_queried - assert "arango-peer" in response.stores_queried - end - - test "modality filtering across heterogeneous peers" do - :ok = - Resolver.register_peer("es-peer", %{ - endpoint: "http://es:9200", - adapter_type: :elasticsearch, - adapter_config: %{index: "data"}, - modalities: ["document", "vector"] - }) - - :ok = - Resolver.register_peer("arango-peer", %{ - endpoint: "http://arango:8529", - adapter_type: :arangodb, - adapter_config: %{database: "_system"}, - modalities: ["graph", "document"] - }) - - # Query for vector — only ES supports it - {:ok, response} = Resolver.query("*", ["vector"], timeout: 2_000) - - assert length(response.stores_queried) == 1 - assert "es-peer" in response.stores_queried - end - - test "deregister heterogeneous peer" do - :ok = - Resolver.register_peer("to-remove", %{ - endpoint: "http://arango:8529", - adapter_type: :arangodb, - adapter_config: %{}, - modalities: ["document"] - }) - - assert length(Resolver.list_peers()) == 1 - - :ok = Resolver.deregister_peer("to-remove") - assert Resolver.list_peers() == [] - end - - test "deregister non-existent peer returns error" do - result = Resolver.deregister_peer("ghost-peer") - assert {:error, :not_found} = result - end - - test "peer with no declared modalities gets adapter defaults" do - :ok = - Resolver.register_peer("default-mods", %{ - endpoint: "http://arango:8529", - adapter_type: :arangodb, - adapter_config: %{}, - modalities: [] - }) - - [peer] = Resolver.list_peers() - - # ArangoDB defaults: graph, document, semantic, temporal, provenance, spatial - assert length(peer.modalities) == 6 - assert "graph" in peer.modalities - assert "document" in peer.modalities - end - end -end diff --git a/verisimdb/elixir-orchestration/test/verisim/federation/adapters/clickhouse_test.exs b/verisimdb/elixir-orchestration/test/verisim/federation/adapters/clickhouse_test.exs deleted file mode 100644 index 771148e7..00000000 --- a/verisimdb/elixir-orchestration/test/verisim/federation/adapters/clickhouse_test.exs +++ /dev/null @@ -1,52 +0,0 @@ -# SPDX-License-Identifier: MPL-2.0 - -defmodule VeriSim.Federation.Adapters.ClickHouseTest do - @moduledoc """ - Tests for the ClickHouse federation adapter. - - Validates modality declarations, ClickHouse HTTP protocol integration, - and result normalisation from ClickHouse JSON output format. - """ - - use ExUnit.Case, async: true - - alias VeriSim.Federation.Adapters.ClickHouse - - @peer_info %{ - store_id: "ch-test", - endpoint: "http://clickhouse:8123", - adapter_config: %{database: "verisimdb", table: "octads"} - } - - describe "supported_modalities/1" do - test "returns 5 supported modalities" do - modalities = ClickHouse.supported_modalities(%{}) - - assert :vector in modalities - assert :document in modalities - assert :temporal in modalities - assert :spatial in modalities - assert :semantic in modalities - - refute :graph in modalities - refute :tensor in modalities - refute :provenance in modalities - end - end - - describe "translate_results/2" do - test "normalises ClickHouse JSON row format" do - raw = [%{"id" => "ch-001", "score" => 0.65, "created_at" => "2026-02-28T12:00:00Z"}] - - [result] = ClickHouse.translate_results(raw, @peer_info) - - assert result.octad_id == "ch-001" - assert result.score == 0.65 - assert result.source_store == "ch-test" - end - - test "handles empty results" do - assert ClickHouse.translate_results([], @peer_info) == [] - end - end -end diff --git a/verisimdb/elixir-orchestration/test/verisim/federation/adapters/duckdb_test.exs b/verisimdb/elixir-orchestration/test/verisim/federation/adapters/duckdb_test.exs deleted file mode 100644 index cb6350cc..00000000 --- a/verisimdb/elixir-orchestration/test/verisim/federation/adapters/duckdb_test.exs +++ /dev/null @@ -1,58 +0,0 @@ -# SPDX-License-Identifier: MPL-2.0 - -defmodule VeriSim.Federation.Adapters.DuckDBTest do - @moduledoc """ - Tests for the DuckDB federation adapter. - - Validates extension-dependent modality declarations, SQL query - construction for DuckDB-specific syntax, and result normalisation. - """ - - use ExUnit.Case, async: true - - alias VeriSim.Federation.Adapters.DuckDB - - @peer_info %{ - store_id: "duckdb-test", - endpoint: "http://duckdb:8080", - adapter_config: %{path: ":memory:", table: "octads"} - } - - describe "supported_modalities/1" do - test "base modalities without extensions" do - modalities = DuckDB.supported_modalities(%{extensions: []}) - - assert :graph in modalities - assert :temporal in modalities - assert :semantic in modalities - assert :tensor in modalities - refute :document in modalities - refute :vector in modalities - refute :spatial in modalities - end - - test "with all extensions" do - modalities = DuckDB.supported_modalities(%{extensions: [:hnsw, :fts, :spatial]}) - - assert :vector in modalities - assert :spatial in modalities - assert length(modalities) == 7 - end - end - - describe "translate_results/2" do - test "normalises SQL row results" do - raw = [%{"id" => "duck-001", "score" => 0.77, "title" => "Analytics"}] - - [result] = DuckDB.translate_results(raw, @peer_info) - - assert result.octad_id == "duck-001" - assert result.score == 0.77 - assert result.source_store == "duckdb-test" - end - - test "handles empty results" do - assert DuckDB.translate_results([], @peer_info) == [] - end - end -end diff --git a/verisimdb/elixir-orchestration/test/verisim/federation/adapters/influxdb_test.exs b/verisimdb/elixir-orchestration/test/verisim/federation/adapters/influxdb_test.exs deleted file mode 100644 index b3e7772c..00000000 --- a/verisimdb/elixir-orchestration/test/verisim/federation/adapters/influxdb_test.exs +++ /dev/null @@ -1,59 +0,0 @@ -# SPDX-License-Identifier: MPL-2.0 - -defmodule VeriSim.Federation.Adapters.InfluxDBTest do - @moduledoc """ - Tests for the InfluxDB federation adapter. - - Validates modality declarations (time-series specialist), - Flux query construction, and result normalisation from - InfluxDB's annotated CSV or JSON response format. - """ - - use ExUnit.Case, async: true - - alias VeriSim.Federation.Adapters.InfluxDB - - @peer_info %{ - store_id: "influx-test", - endpoint: "http://influxdb:8086", - adapter_config: %{org: "verisim", bucket: "octads", token: "test-token"} - } - - describe "supported_modalities/1" do - test "returns 2 supported modalities" do - modalities = InfluxDB.supported_modalities(%{}) - - assert :temporal in modalities - assert :semantic in modalities - - refute :graph in modalities - refute :vector in modalities - refute :tensor in modalities - refute :document in modalities - refute :provenance in modalities - refute :spatial in modalities - end - end - - describe "translate_results/2" do - test "normalises InfluxDB time-series records" do - raw = [ - %{ - "_measurement" => "octad_events", - "_time" => "2026-02-28T12:00:00Z", - "_value" => 42.5, - "entity_id" => "influx-001" - } - ] - - [result] = InfluxDB.translate_results(raw, @peer_info) - - assert result.octad_id == "octad_events:influx-001:2026-02-28T12:00:00Z" - assert result.source_store == "influx-test" - end - - test "handles empty results" do - assert InfluxDB.translate_results([], @peer_info) == [] - end - end -end diff --git a/verisimdb/elixir-orchestration/test/verisim/federation/adapters/integration/clickhouse_integration_test.exs b/verisimdb/elixir-orchestration/test/verisim/federation/adapters/integration/clickhouse_integration_test.exs deleted file mode 100644 index efbda3ab..00000000 --- a/verisimdb/elixir-orchestration/test/verisim/federation/adapters/integration/clickhouse_integration_test.exs +++ /dev/null @@ -1,352 +0,0 @@ -# SPDX-License-Identifier: MPL-2.0 - -defmodule VeriSim.Federation.Adapters.ClickHouseIntegrationTest do - @moduledoc """ - Integration tests for the ClickHouse federation adapter. - - Runs against a real ClickHouse server from the test-infra container - stack. The seed script `clickhouse-init.sql` pre-loads: - - - `verisimdb.octads` table: 3 rows (octad-test-001, -002, -003) with - title, content, entity_type, drift_status, spatial coordinates, tags - - `verisimdb.modalities` table: 11 rows (modality records per octad) - - `verisimdb.drift_scores` table: 5 rows (time-series drift measurements) - - `verisimdb.provenance_events` table: 4 rows - - 3 materialized views: mv_drift_status_counts, mv_modality_distribution, - mv_avg_drift_by_type - - ## Test Infrastructure - - Requires the test-infra stack running: - - cd connectors/test-infra && selur-compose up -d - - ClickHouse is exposed on localhost:8123 (HTTP) and localhost:9000 (native). - - ## Running - - mix test --include integration test/verisim/federation/adapters/integration/clickhouse_integration_test.exs - - Author: Jonathan D.A. Jewell - """ - - use ExUnit.Case, async: false - - alias VeriSim.Federation.Adapters.ClickHouse - - @moduletag :integration - - # --------------------------------------------------------------------------- - # Configuration - # --------------------------------------------------------------------------- - - @clickhouse_url System.get_env("VERISIM_CLICKHOUSE_URL", "http://localhost:8123") - - @peer_info %{ - store_id: "ch-integration", - endpoint: @clickhouse_url, - adapter_config: %{ - database: "verisimdb", - table: "octads" - } - } - - @integration_prefix "octad-integration" - - # --------------------------------------------------------------------------- - # Setup / Teardown - # --------------------------------------------------------------------------- - - setup_all do - case ClickHouse.health_check(@peer_info) do - {:ok, latency_ms} -> - {:ok, %{latency_ms: latency_ms}} - - {:error, reason} -> - {:ok, %{skip_reason: reason}} - end - end - - setup %{} = context do - if Map.has_key?(context, :skip_reason) do - {:ok, Map.put(context, :skip, true)} - else - {:ok, context} - end - end - - # --------------------------------------------------------------------------- - # 1. Connection Tests - # --------------------------------------------------------------------------- - - describe "connection to real ClickHouse" do - test "connect/1 succeeds against running instance", context do - skip_if_unavailable(context) - - result = ClickHouse.connect(@peer_info) - assert result == :ok - end - - test "health_check/1 returns 'Ok.' with latency", context do - skip_if_unavailable(context) - - assert {:ok, latency_ms} = ClickHouse.health_check(@peer_info) - assert is_integer(latency_ms) - assert latency_ms >= 0 - end - end - - # --------------------------------------------------------------------------- - # 2. Read / Query — Verify Seed Data - # --------------------------------------------------------------------------- - - describe "querying seeded octads table" do - test "SELECT * returns all 3 seeded rows", context do - skip_if_unavailable(context) - - query_params = %{modalities: [], limit: 100} - assert {:ok, results} = ClickHouse.query(@peer_info, query_params) - - # The seed script inserts 3 octad rows - assert length(results) >= 3 - - Enum.each(results, fn result -> - assert result.source_store == "ch-integration" - assert is_binary(result.octad_id) - assert is_number(result.score) - assert result.drifted == false - assert is_map(result.data) - end) - end - - test "results contain expected seed data fields", context do - skip_if_unavailable(context) - - query_params = %{modalities: [], limit: 10} - assert {:ok, results} = ClickHouse.query(@peer_info, query_params) - - # Find octad-test-001 in results - test_001 = Enum.find(results, fn r -> r.octad_id == "octad-test-001" end) - - if test_001 do - assert test_001.data["title"] == "Introduction to Cross-Modal Consistency" - assert test_001.data["entity_type"] == "Article" - assert test_001.data["drift_status"] in ["healthy", 0] - end - end - end - - # --------------------------------------------------------------------------- - # 3. Full-Text Search (Document Modality) - # --------------------------------------------------------------------------- - - describe "full-text search via ClickHouse" do - test "text search for 'drift' returns matching rows", context do - skip_if_unavailable(context) - - query_params = %{ - modalities: [:document], - text_query: "drift", - limit: 10 - } - - assert {:ok, results} = ClickHouse.query(@peer_info, query_params) - # octad-test-002 title contains "Drift", octad-test-003 content mentions drift - assert length(results) >= 1 - end - - test "text search for 'VeriSimDB' returns matching rows", context do - skip_if_unavailable(context) - - query_params = %{ - modalities: [:document], - text_query: "verisimdb", - limit: 10 - } - - assert {:ok, results} = ClickHouse.query(@peer_info, query_params) - assert length(results) >= 1 - end - end - - # --------------------------------------------------------------------------- - # 4. Materialized View Queries - # --------------------------------------------------------------------------- - - describe "materialized view queries" do - test "drift status distribution view has data", context do - skip_if_unavailable(context) - - # Query the materialized view directly via a custom peer config - mv_peer = %{ - @peer_info - | adapter_config: %{database: "verisimdb", table: "mv_drift_status_counts"} - } - - query_params = %{modalities: [], limit: 100} - - case ClickHouse.query(mv_peer, query_params) do - {:ok, results} -> - # The materialized view should have rows from the 5 drift_scores inserts - assert is_list(results) - - {:error, _reason} -> - # Materialized views may have different column structure - assert true - end - end - end - - # --------------------------------------------------------------------------- - # 5. Aggregation Queries (Drift Scores) - # --------------------------------------------------------------------------- - - describe "aggregation on drift_scores table" do - test "querying drift_scores table returns seeded measurements", context do - skip_if_unavailable(context) - - drift_peer = %{ - @peer_info - | adapter_config: %{database: "verisimdb", table: "drift_scores"} - } - - query_params = %{modalities: [], limit: 100} - assert {:ok, results} = ClickHouse.query(drift_peer, query_params) - - # Seed script inserts 5 drift score rows - assert length(results) >= 5 - end - - test "temporal range query on drift_scores filters by measured_at", context do - skip_if_unavailable(context) - - drift_peer = %{ - @peer_info - | adapter_config: %{database: "verisimdb", table: "drift_scores"} - } - - now = DateTime.utc_now() - two_hours_ago = DateTime.add(now, -7200, :second) - - query_params = %{ - modalities: [:temporal], - temporal_range: %{ - start: DateTime.to_iso8601(two_hours_ago), - end: DateTime.to_iso8601(now) - }, - limit: 100 - } - - assert {:ok, results} = ClickHouse.query(drift_peer, query_params) - assert length(results) >= 1 - end - end - - # --------------------------------------------------------------------------- - # 6. Spatial Queries - # --------------------------------------------------------------------------- - - describe "spatial queries via ClickHouse geo functions" do - test "pointInPolygon query around London returns octad-test-001", context do - skip_if_unavailable(context) - - # octad-test-001 has lat=51.5074, lon=-0.1278 (London) - query_params = %{ - modalities: [:spatial], - spatial_bounds: %{ - min_lon: -1.0, - min_lat: 51.0, - max_lon: 1.0, - max_lat: 52.0 - }, - limit: 10 - } - - assert {:ok, results} = ClickHouse.query(@peer_info, query_params) - assert length(results) >= 1 - end - end - - # --------------------------------------------------------------------------- - # 7. Write + Read-Back - # --------------------------------------------------------------------------- - - describe "INSERT and SELECT round-trip" do - test "translate_results normalises ClickHouse JSONEachRow format", context do - skip_if_unavailable(context) - - test_id = "#{@integration_prefix}-ch-#{System.unique_integer([:positive])}" - - raw_row = %{ - "id" => test_id, - "title" => "Integration Test Row", - "entity_type" => "TestArticle", - "score" => 0.73, - "created_at" => DateTime.to_iso8601(DateTime.utc_now()) - } - - [normalised] = ClickHouse.translate_results([raw_row], @peer_info) - - assert normalised.source_store == "ch-integration" - assert normalised.octad_id == test_id - assert normalised.score == 0.73 - assert normalised.drifted == false - end - end - - # --------------------------------------------------------------------------- - # 8. Error Handling - # --------------------------------------------------------------------------- - - describe "error handling against real ClickHouse" do - test "querying a nonexistent table returns an error", context do - skip_if_unavailable(context) - - bad_peer = %{ - @peer_info - | adapter_config: %{database: "verisimdb", table: "nonexistent_table_xyz"} - } - - query_params = %{modalities: [], limit: 10} - - assert {:error, _reason} = ClickHouse.query(bad_peer, query_params) - end - - test "connecting to an unreachable endpoint returns an error" do - unreachable_peer = %{ - store_id: "ch-unreachable", - endpoint: "http://localhost:59996", - adapter_config: %{database: "verisimdb", table: "octads"} - } - - assert {:error, _reason} = ClickHouse.connect(unreachable_peer) - end - end - - # --------------------------------------------------------------------------- - # 9. Modality Support - # --------------------------------------------------------------------------- - - describe "modality support declarations" do - test "ClickHouse supports 5 modalities without extensions" do - modalities = ClickHouse.supported_modalities(%{}) - - assert :vector in modalities - assert :document in modalities - assert :temporal in modalities - assert :spatial in modalities - assert :semantic in modalities - - refute :graph in modalities - refute :tensor in modalities - refute :provenance in modalities - end - end - - # --------------------------------------------------------------------------- - # Helpers - # --------------------------------------------------------------------------- - - defp skip_if_unavailable(%{skip: true}), do: flunk("ClickHouse not available — start test-infra stack") - defp skip_if_unavailable(_context), do: :ok -end diff --git a/verisimdb/elixir-orchestration/test/verisim/federation/adapters/integration/influxdb_integration_test.exs b/verisimdb/elixir-orchestration/test/verisim/federation/adapters/integration/influxdb_integration_test.exs deleted file mode 100644 index 9c0a7734..00000000 --- a/verisimdb/elixir-orchestration/test/verisim/federation/adapters/integration/influxdb_integration_test.exs +++ /dev/null @@ -1,374 +0,0 @@ -# SPDX-License-Identifier: MPL-2.0 - -defmodule VeriSim.Federation.Adapters.InfluxDBIntegrationTest do - @moduledoc """ - Integration tests for the InfluxDB federation adapter. - - Runs against a real InfluxDB 2.x instance from the test-infra container - stack. The seed script `influxdb-init.sh` pre-loads: - - - Organisation: `verisimdb` - - Auth token: `verisim-test-token-do-not-use-in-production` - - Buckets: `metrics` (auto-setup), `drift_scores` (30d retention), - `federation_health` (7d retention) - - `drift_scores` bucket: 7 data points across 3 octads - - Measurement: `drift` with tags `octad_id`, `status` - - Fields: `semantic_vector`, `graph_document`, `temporal`, `overall` - - `metrics` bucket: ~20 query latency data points - - Measurement: `query_latency` with tags `service`, `query_type` - - `federation_health` bucket: 7 adapter health checks + 3 sync events - - Measurements: `adapter_health`, `federation_sync` - - ## Test Infrastructure - - Requires the test-infra stack running: - - cd connectors/test-infra && selur-compose up -d - - InfluxDB is exposed on localhost:8086. - - ## Running - - mix test --include integration test/verisim/federation/adapters/integration/influxdb_integration_test.exs - - Author: Jonathan D.A. Jewell - """ - - use ExUnit.Case, async: false - - alias VeriSim.Federation.Adapters.InfluxDB - - @moduletag :integration - - # --------------------------------------------------------------------------- - # Configuration - # --------------------------------------------------------------------------- - - @influxdb_url System.get_env("VERISIM_INFLUXDB_URL", "http://localhost:8086") - @influxdb_token System.get_env( - "VERISIM_INFLUXDB_TOKEN", - "verisim-test-token-do-not-use-in-production" - ) - @influxdb_org System.get_env("VERISIM_INFLUXDB_ORG", "verisimdb") - - @peer_info %{ - store_id: "influx-integration", - endpoint: @influxdb_url, - adapter_config: %{ - org: @influxdb_org, - bucket: "drift_scores", - token: @influxdb_token, - measurement: "drift" - } - } - - @integration_prefix "octad-integration" - - # --------------------------------------------------------------------------- - # Setup / Teardown - # --------------------------------------------------------------------------- - - setup_all do - case InfluxDB.health_check(@peer_info) do - {:ok, latency_ms} -> - {:ok, %{latency_ms: latency_ms}} - - {:error, reason} -> - {:ok, %{skip_reason: reason}} - end - end - - setup %{} = context do - if Map.has_key?(context, :skip_reason) do - {:ok, Map.put(context, :skip, true)} - else - {:ok, context} - end - end - - # --------------------------------------------------------------------------- - # 1. Connection Tests - # --------------------------------------------------------------------------- - - describe "connection to real InfluxDB" do - test "connect/1 succeeds against running instance", context do - skip_if_unavailable(context) - - result = InfluxDB.connect(@peer_info) - assert result == :ok - end - - test "health_check/1 returns 'pass' status with latency", context do - skip_if_unavailable(context) - - assert {:ok, latency_ms} = InfluxDB.health_check(@peer_info) - assert is_integer(latency_ms) - assert latency_ms >= 0 - end - end - - # --------------------------------------------------------------------------- - # 2. Read / Query — Verify Seed Data (drift_scores bucket) - # --------------------------------------------------------------------------- - - describe "querying seeded drift_scores bucket" do - test "default query returns recent drift data points", context do - skip_if_unavailable(context) - - query_params = %{modalities: [], limit: 100} - assert {:ok, results} = InfluxDB.query(@peer_info, query_params) - - # The seed script writes 7 drift data points - # Default Flux queries range(start: -24h), so all points should be within range - assert is_list(results) - - Enum.each(results, fn result -> - assert result.source_store == "influx-integration" - assert is_binary(result.octad_id) - assert result.drifted == false - end) - end - end - - # --------------------------------------------------------------------------- - # 3. Time Range Queries - # --------------------------------------------------------------------------- - - describe "temporal range queries" do - test "querying with time range filter returns drift data within window", context do - skip_if_unavailable(context) - - query_params = %{ - modalities: [:temporal], - temporal_range: %{ - start: "-2h", - end: "now()" - }, - limit: 100 - } - - assert {:ok, results} = InfluxDB.query(@peer_info, query_params) - # All 7 seed data points were written within the last hour - assert is_list(results) - end - - test "querying with absolute timestamps returns matching points", context do - skip_if_unavailable(context) - - now = DateTime.utc_now() - three_hours_ago = DateTime.add(now, -3 * 3600, :second) - - query_params = %{ - modalities: [:temporal], - temporal_range: %{ - start: DateTime.to_iso8601(three_hours_ago), - end: DateTime.to_iso8601(now) - }, - limit: 100 - } - - assert {:ok, results} = InfluxDB.query(@peer_info, query_params) - assert is_list(results) - end - end - - # --------------------------------------------------------------------------- - # 4. Aggregation Queries (Mean, Max via Semantic Filters) - # --------------------------------------------------------------------------- - - describe "aggregation and tag-based filtering" do - test "filtering by octad_id tag returns specific octad drift scores", context do - skip_if_unavailable(context) - - query_params = %{ - modalities: [:temporal, :semantic], - temporal_range: %{ - start: "-24h", - end: "now()" - }, - filters: %{"octad_id" => "octad-test-001"}, - limit: 100 - } - - assert {:ok, results} = InfluxDB.query(@peer_info, query_params) - assert is_list(results) - end - - test "filtering by status=drifted returns only drifted octads", context do - skip_if_unavailable(context) - - query_params = %{ - modalities: [:semantic], - filters: %{"status" => "drifted"}, - limit: 100 - } - - assert {:ok, results} = InfluxDB.query(@peer_info, query_params) - assert is_list(results) - end - end - - # --------------------------------------------------------------------------- - # 5. Query Latency Metrics (metrics bucket) - # --------------------------------------------------------------------------- - - describe "querying metrics bucket" do - test "query_latency measurements are accessible", context do - skip_if_unavailable(context) - - metrics_peer = %{ - @peer_info - | adapter_config: %{ - org: @influxdb_org, - bucket: "metrics", - token: @influxdb_token, - measurement: "query_latency" - } - } - - query_params = %{ - modalities: [:temporal], - temporal_range: %{start: "-2h", end: "now()"}, - limit: 50 - } - - assert {:ok, results} = InfluxDB.query(metrics_peer, query_params) - assert is_list(results) - end - end - - # --------------------------------------------------------------------------- - # 6. Federation Health Metrics (federation_health bucket) - # --------------------------------------------------------------------------- - - describe "querying federation_health bucket" do - test "adapter_health measurements are accessible", context do - skip_if_unavailable(context) - - health_peer = %{ - @peer_info - | adapter_config: %{ - org: @influxdb_org, - bucket: "federation_health", - token: @influxdb_token, - measurement: "adapter_health" - } - } - - query_params = %{ - modalities: [:temporal], - temporal_range: %{start: "-1h", end: "now()"}, - limit: 50 - } - - assert {:ok, results} = InfluxDB.query(health_peer, query_params) - assert is_list(results) - end - end - - # --------------------------------------------------------------------------- - # 7. Write + Read-Back - # --------------------------------------------------------------------------- - - describe "write and read-back cycle" do - test "translate_results normalises InfluxDB time-series record format", context do - skip_if_unavailable(context) - - _test_id = "#{@integration_prefix}-influx-#{System.unique_integer([:positive])}" - - # Simulate InfluxDB Flux CSV / JSON response format - raw_record = %{ - "_measurement" => "drift", - "_time" => "2026-02-28T12:00:00Z", - "_value" => 0.045, - "entity_id" => "octad-test-001", - "octad_id" => "octad-test-001", - "status" => "healthy" - } - - [normalised] = InfluxDB.translate_results([raw_record], @peer_info) - - assert normalised.source_store == "influx-integration" - # InfluxDB adapter builds composite ID: measurement:entity_id:_time - assert normalised.octad_id == "drift:octad-test-001:2026-02-28T12:00:00Z" - assert normalised.drifted == false - end - end - - # --------------------------------------------------------------------------- - # 8. Error Handling - # --------------------------------------------------------------------------- - - describe "error handling against real InfluxDB" do - test "querying a nonexistent bucket returns an error", context do - skip_if_unavailable(context) - - bad_peer = %{ - @peer_info - | adapter_config: %{ - org: @influxdb_org, - bucket: "nonexistent_bucket_xyz", - token: @influxdb_token, - measurement: "drift" - } - } - - query_params = %{modalities: [], limit: 10} - - result = InfluxDB.query(bad_peer, query_params) - # InfluxDB returns an error for nonexistent buckets - assert match?({:error, _}, result) or match?({:ok, []}, result) - end - - test "querying with an invalid token returns an error", context do - skip_if_unavailable(context) - - bad_token_peer = %{ - @peer_info - | adapter_config: Map.put(@peer_info.adapter_config, :token, "invalid-token-xyz") - } - - query_params = %{modalities: [], limit: 10} - - result = InfluxDB.query(bad_token_peer, query_params) - assert match?({:error, _}, result) - end - - test "connecting to an unreachable endpoint returns an error" do - unreachable_peer = %{ - store_id: "influx-unreachable", - endpoint: "http://localhost:59994", - adapter_config: %{org: "verisimdb", bucket: "drift_scores", token: "test"} - } - - assert {:error, _reason} = InfluxDB.connect(unreachable_peer) - end - end - - # --------------------------------------------------------------------------- - # 9. Modality Support - # --------------------------------------------------------------------------- - - describe "modality support declarations" do - test "InfluxDB supports only temporal and semantic modalities" do - modalities = InfluxDB.supported_modalities(%{}) - - assert :temporal in modalities - assert :semantic in modalities - - refute :graph in modalities - refute :vector in modalities - refute :document in modalities - refute :tensor in modalities - refute :provenance in modalities - refute :spatial in modalities - end - end - - # --------------------------------------------------------------------------- - # Helpers - # --------------------------------------------------------------------------- - - defp skip_if_unavailable(%{skip: true}), do: flunk("InfluxDB not available — start test-infra stack") - defp skip_if_unavailable(_context), do: :ok -end diff --git a/verisimdb/elixir-orchestration/test/verisim/federation/adapters/integration/mongodb_integration_test.exs b/verisimdb/elixir-orchestration/test/verisim/federation/adapters/integration/mongodb_integration_test.exs deleted file mode 100644 index e8cae8cc..00000000 --- a/verisimdb/elixir-orchestration/test/verisim/federation/adapters/integration/mongodb_integration_test.exs +++ /dev/null @@ -1,401 +0,0 @@ -# SPDX-License-Identifier: MPL-2.0 - -defmodule VeriSim.Federation.Adapters.MongoDBIntegrationTest do - @moduledoc """ - Integration tests for the MongoDB federation adapter. - - Runs against a real MongoDB 7+ replica set (rs0) from the test-infra - container stack. The seed script `mongodb-init.js` pre-loads 3 octad - documents, 2 drift score documents, and 2 provenance events into the - `verisimdb` database. - - ## Test Infrastructure - - Requires the test-infra stack running: - - cd connectors/test-infra && selur-compose up -d - - MongoDB is exposed on localhost:27017 with replica set `rs0`. - - ## Seed Data Summary - - - `octads` collection: 3 documents (octad-test-001, -002, -003) - - Each has modalities array with document, vector, spatial, temporal, - graph, provenance, and semantic sub-documents - - Indexes: unique on `id`, text on `modalities.data.content`/`title`, - 2dsphere on `modalities.data.location`, compound on `created_at`/`updated_at` - - `drift_scores` collection: 2 documents - - `provenance_events` collection: 2 documents - - ## Running - - mix test --include integration test/verisim/federation/adapters/integration/mongodb_integration_test.exs - - Author: Jonathan D.A. Jewell - """ - - use ExUnit.Case, async: false - - alias VeriSim.Federation.Adapters.MongoDB - - @moduletag :integration - - # --------------------------------------------------------------------------- - # Configuration - # --------------------------------------------------------------------------- - - @mongodb_url System.get_env("VERISIM_MONGODB_URL", "mongodb://localhost:27017") - - @peer_info %{ - store_id: "mongo-integration", - endpoint: @mongodb_url, - adapter_config: %{ - database: "verisimdb", - collection: "octads", - replica_set: "rs0", - geo_index: true, - data_source: "Cluster0" - } - } - - # Prefix for integration test data — avoids collision with seed data (octad-test-*) - @integration_prefix "octad-integration" - - # --------------------------------------------------------------------------- - # Setup / Teardown - # --------------------------------------------------------------------------- - - setup_all do - # Verify MongoDB is reachable before running the suite. - # If the connection fails, all tests in this module are skipped - # rather than producing confusing error messages. - case MongoDB.health_check(@peer_info) do - {:ok, latency_ms} -> - {:ok, %{latency_ms: latency_ms}} - - {:error, reason} -> - {:ok, %{skip_reason: reason}} - end - end - - setup %{} = context do - if Map.has_key?(context, :skip_reason) do - {:ok, Map.put(context, :skip, true)} - else - {:ok, context} - end - end - - # --------------------------------------------------------------------------- - # 1. Connection Tests - # --------------------------------------------------------------------------- - - describe "connection to real MongoDB" do - test "connect/1 succeeds against running instance", context do - skip_if_unavailable(context) - - result = MongoDB.connect(@peer_info) - assert result == :ok - end - - test "health_check/1 returns latency in milliseconds", context do - skip_if_unavailable(context) - - assert {:ok, latency_ms} = MongoDB.health_check(@peer_info) - assert is_integer(latency_ms) - assert latency_ms >= 0 - end - end - - # --------------------------------------------------------------------------- - # 2. Read / Query — Verify Seed Data - # --------------------------------------------------------------------------- - - describe "querying seeded octads collection" do - test "default query returns all 3 seeded documents", context do - skip_if_unavailable(context) - - query_params = %{modalities: [], limit: 100} - assert {:ok, results} = MongoDB.query(@peer_info, query_params) - - # The seed script inserts 3 octad documents - assert length(results) >= 3 - - # All results should be normalised with the correct source_store - Enum.each(results, fn result -> - assert result.source_store == "mongo-integration" - assert is_binary(result.octad_id) - assert is_float(result.score) or is_integer(result.score) - assert result.drifted == false - assert is_map(result.data) - assert is_integer(result.response_time_ms) - end) - end - - test "text search across document modality returns matching octads", context do - skip_if_unavailable(context) - - # Seed document octad-test-001 contains "cross-modal consistency" - query_params = %{ - modalities: [:document], - text_query: "consistency modality", - limit: 10 - } - - assert {:ok, results} = MongoDB.query(@peer_info, query_params) - assert length(results) >= 1 - - # Text search results should have a non-zero score from textScore - first = List.first(results) - assert first.score >= 0 - end - - test "temporal range query filters by created_at", context do - skip_if_unavailable(context) - - # Seed data uses relative timestamps — query a wide window to capture all - now = DateTime.utc_now() - two_days_ago = DateTime.add(now, -2 * 86400, :second) - - query_params = %{ - modalities: [:temporal], - temporal_range: %{ - start: DateTime.to_iso8601(two_days_ago), - end: DateTime.to_iso8601(now) - }, - limit: 100 - } - - assert {:ok, results} = MongoDB.query(@peer_info, query_params) - # Should find at least some of the 3 seeded octads created within 1 day - assert length(results) >= 1 - end - - test "semantic (metadata) filter query returns matching documents", context do - skip_if_unavailable(context) - - # Octad-test-001 has modality type "document" with specific content - query_params = %{ - modalities: [:semantic], - filters: %{"type" => "document"}, - limit: 10 - } - - assert {:ok, results} = MongoDB.query(@peer_info, query_params) - assert is_list(results) - end - end - - # --------------------------------------------------------------------------- - # 3. Spatial Modality Query - # --------------------------------------------------------------------------- - - describe "geospatial queries (2dsphere)" do - test "geoWithin query on London coordinates returns octad-test-001", context do - skip_if_unavailable(context) - - # octad-test-001 is located at [-0.1278, 51.5074] (London) - # Query a bounding box around London - query_params = %{ - modalities: [:spatial], - spatial_bounds: %{ - min_lon: -1.0, - min_lat: 51.0, - max_lon: 1.0, - max_lat: 52.0 - }, - limit: 10 - } - - assert {:ok, results} = MongoDB.query(@peer_info, query_params) - # Should find at least octad-test-001 which has London coordinates - assert length(results) >= 1 - end - - test "geoWithin query on empty region returns no results", context do - skip_if_unavailable(context) - - # Query an area in the middle of the ocean where no octads exist - query_params = %{ - modalities: [:spatial], - spatial_bounds: %{ - min_lon: 170.0, - min_lat: -60.0, - max_lon: 175.0, - max_lat: -55.0 - }, - limit: 10 - } - - assert {:ok, results} = MongoDB.query(@peer_info, query_params) - assert results == [] - end - end - - # --------------------------------------------------------------------------- - # 4. Write + Read-Back - # --------------------------------------------------------------------------- - - describe "write and read-back cycle" do - test "inserting a new octad document and querying it back", context do - skip_if_unavailable(context) - - test_id = "#{@integration_prefix}-write-#{System.unique_integer([:positive])}" - - # Write a new document via the adapter's query mechanism - # (The adapter uses the MongoDB Data API — we test round-trip through - # the adapter's translate_results/2 on the read side.) - # - # Since the adapter is read-oriented (query/3), we verify that - # translate_results/2 correctly normalises a raw MongoDB document. - raw_doc = %{ - "_id" => test_id, - "id" => test_id, - "title" => "Integration Test Document", - "content" => "Created by MongoDBIntegrationTest", - "score" => 0.77, - "created_at" => DateTime.to_iso8601(DateTime.utc_now()) - } - - [normalised] = MongoDB.translate_results([raw_doc], @peer_info) - - assert normalised.source_store == "mongo-integration" - assert normalised.octad_id == test_id - assert normalised.score == 0.77 - assert normalised.drifted == false - assert normalised.data["title"] == "Integration Test Document" - end - end - - # --------------------------------------------------------------------------- - # 5. Compound Index Queries - # --------------------------------------------------------------------------- - - describe "compound index and graph queries" do - test "graph traversal query with $graphLookup returns results", context do - skip_if_unavailable(context) - - # octad-test-001 has a relates_to relationship to octad-test-002 - query_params = %{ - modalities: [:graph], - graph_pattern: "octad-test-001", - limit: 10 - } - - assert {:ok, results} = MongoDB.query(@peer_info, query_params) - assert is_list(results) - end - end - - # --------------------------------------------------------------------------- - # 6. Change Stream / Provenance - # --------------------------------------------------------------------------- - - describe "provenance modality (change stream)" do - test "provenance query returns provenance log entries", context do - skip_if_unavailable(context) - - # The seed script creates a provenance_events collection with 2 events - provenance_peer = %{ - @peer_info - | adapter_config: - Map.merge(@peer_info.adapter_config, %{ - provenance_collection: "provenance_events", - replica_set: "rs0" - }) - } - - query_params = %{ - modalities: [:provenance], - limit: 10 - } - - assert {:ok, results} = MongoDB.query(provenance_peer, query_params) - assert is_list(results) - end - - test "supported_modalities includes :provenance when replica_set configured" do - config_with_rs = %{replica_set: "rs0", atlas: false, geo_index: true} - modalities = MongoDB.supported_modalities(config_with_rs) - - assert :provenance in modalities - end - end - - # --------------------------------------------------------------------------- - # 7. Error Handling - # --------------------------------------------------------------------------- - - describe "error handling against real MongoDB" do - test "querying a nonexistent collection returns an error or empty results", context do - skip_if_unavailable(context) - - bad_peer = %{ - @peer_info - | adapter_config: Map.put(@peer_info.adapter_config, :collection, "nonexistent_collection_xyz") - } - - query_params = %{modalities: [], limit: 10} - - # The adapter should either return an error tuple or an empty result - # (MongoDB allows queries on nonexistent collections — they return empty) - case MongoDB.query(bad_peer, query_params) do - {:ok, results} -> - assert results == [] - - {:error, _reason} -> - # Also acceptable — some configurations may reject this - assert true - end - end - - test "connecting to an unreachable endpoint returns an error" do - unreachable_peer = %{ - store_id: "mongo-unreachable", - endpoint: "http://localhost:59999", - adapter_config: %{database: "verisimdb"} - } - - assert {:error, _reason} = MongoDB.connect(unreachable_peer) - end - end - - # --------------------------------------------------------------------------- - # 8. Result Normalisation (Against Real Data) - # --------------------------------------------------------------------------- - - describe "translate_results/2 with real MongoDB document shapes" do - test "normalises a document with ObjectId _id" do - raw = [%{"_id" => "507f1f77bcf86cd799439011", "title" => "Real doc", "score" => 0.93}] - - [result] = MongoDB.translate_results(raw, @peer_info) - - assert result.octad_id == "507f1f77bcf86cd799439011" - assert result.score == 0.93 - end - - test "normalises a document with nested modalities array" do - raw = [ - %{ - "_id" => "octad-test-001", - "id" => "octad-test-001", - "modalities" => [ - %{"type" => "document", "data" => %{"title" => "Test"}}, - %{"type" => "vector", "data" => %{"embedding" => [0.1, 0.2]}} - ] - } - ] - - [result] = MongoDB.translate_results(raw, @peer_info) - assert result.octad_id == "octad-test-001" - assert is_list(result.data["modalities"]) - end - end - - # --------------------------------------------------------------------------- - # Helpers - # --------------------------------------------------------------------------- - - defp skip_if_unavailable(%{skip: true}), do: flunk("MongoDB not available — start test-infra stack") - defp skip_if_unavailable(_context), do: :ok -end diff --git a/verisimdb/elixir-orchestration/test/verisim/federation/adapters/integration/neo4j_integration_test.exs b/verisimdb/elixir-orchestration/test/verisim/federation/adapters/integration/neo4j_integration_test.exs deleted file mode 100644 index 30d9c4df..00000000 --- a/verisimdb/elixir-orchestration/test/verisim/federation/adapters/integration/neo4j_integration_test.exs +++ /dev/null @@ -1,359 +0,0 @@ -# SPDX-License-Identifier: MPL-2.0 - -defmodule VeriSim.Federation.Adapters.Neo4jIntegrationTest do - @moduledoc """ - Integration tests for the Neo4j federation adapter. - - Runs against a real Neo4j 5.x instance from the test-infra container - stack. The seed script `neo4j-init.cypher` pre-loads: - - - 4 Octad nodes (octad-test-001 through -004) with labels, properties, - and spatial coordinates - - 6 relationships: RELATES_TO, CITES, PART_OF - - 2 ProvenanceEvent nodes chained with FOLLOWED_BY and HAS_PROVENANCE - - 3 OntologyType nodes with IS_TYPE and SUBCLASS_OF relationships - - Fulltext index `octad_fulltext` on [title, content] - - Constraints: unique on Octad.id, ProvenanceEvent.event_id - - ## Test Infrastructure - - Requires the test-infra stack running: - - cd connectors/test-infra && selur-compose up -d - - Neo4j is exposed on localhost:7474 (HTTP) and localhost:7687 (Bolt). - Default credentials: neo4j/neo4j (or as configured in compose.toml). - - ## Running - - mix test --include integration test/verisim/federation/adapters/integration/neo4j_integration_test.exs - - Author: Jonathan D.A. Jewell - """ - - use ExUnit.Case, async: false - - alias VeriSim.Federation.Adapters.Neo4j - - @moduletag :integration - - # --------------------------------------------------------------------------- - # Configuration - # --------------------------------------------------------------------------- - - @neo4j_http System.get_env("VERISIM_NEO4J_HTTP", "http://localhost:7474") - @neo4j_user System.get_env("VERISIM_NEO4J_USER", "neo4j") - @neo4j_pass System.get_env("VERISIM_NEO4J_PASS", "neo4j") - - @peer_info %{ - store_id: "neo4j-integration", - endpoint: @neo4j_http, - adapter_config: %{ - database: "neo4j", - auth: {:basic, @neo4j_user, @neo4j_pass}, - label: "Octad", - version: 5, - relationship_type: "RELATES_TO", - fulltext_index: "octad_fulltext", - max_depth: 3 - } - } - - @integration_prefix "octad-integration" - - # --------------------------------------------------------------------------- - # Setup / Teardown - # --------------------------------------------------------------------------- - - setup_all do - case Neo4j.health_check(@peer_info) do - {:ok, latency_ms} -> - {:ok, %{latency_ms: latency_ms}} - - {:error, reason} -> - {:ok, %{skip_reason: reason}} - end - end - - setup %{} = context do - if Map.has_key?(context, :skip_reason) do - {:ok, Map.put(context, :skip, true)} - else - {:ok, context} - end - end - - # --------------------------------------------------------------------------- - # 1. Connection Tests - # --------------------------------------------------------------------------- - - describe "connection to real Neo4j" do - test "connect/1 succeeds against running instance", context do - skip_if_unavailable(context) - - result = Neo4j.connect(@peer_info) - assert result == :ok - end - - test "health_check/1 returns latency in milliseconds", context do - skip_if_unavailable(context) - - assert {:ok, latency_ms} = Neo4j.health_check(@peer_info) - assert is_integer(latency_ms) - assert latency_ms >= 0 - end - end - - # --------------------------------------------------------------------------- - # 2. Read / Query — Verify Seed Data - # --------------------------------------------------------------------------- - - describe "querying seeded Octad nodes" do - test "default MATCH query returns all 4 seeded Octad nodes", context do - skip_if_unavailable(context) - - query_params = %{modalities: [], limit: 100} - assert {:ok, results} = Neo4j.query(@peer_info, query_params) - - # The seed script creates 4 Octad nodes - assert length(results) >= 4 - - Enum.each(results, fn result -> - assert result.source_store == "neo4j-integration" - assert is_binary(result.octad_id) or result.octad_id == "unknown" - assert is_number(result.score) - assert result.drifted == false - assert is_map(result.data) - end) - end - end - - # --------------------------------------------------------------------------- - # 3. Fulltext Index Search - # --------------------------------------------------------------------------- - - describe "fulltext index search" do - test "fulltext search for 'drift' returns matching nodes", context do - skip_if_unavailable(context) - - # octad-test-002 has title "Drift Detection Algorithms" - # octad-test-003 has content mentioning "drift" - query_params = %{ - modalities: [:document], - text_query: "drift", - limit: 10 - } - - assert {:ok, results} = Neo4j.query(@peer_info, query_params) - assert length(results) >= 1 - - # Fulltext search should return Lucene scores - Enum.each(results, fn r -> - assert is_number(r.score) - end) - end - - test "fulltext search for 'normalisation' returns octad-test-003", context do - skip_if_unavailable(context) - - query_params = %{ - modalities: [:document], - text_query: "normalisation", - limit: 10 - } - - assert {:ok, results} = Neo4j.query(@peer_info, query_params) - assert length(results) >= 1 - end - end - - # --------------------------------------------------------------------------- - # 4. Relationship Traversal - # --------------------------------------------------------------------------- - - describe "graph relationship traversal" do - test "RELATES_TO traversal from octad-test-001 finds connected nodes", context do - skip_if_unavailable(context) - - query_params = %{ - modalities: [:graph], - graph_pattern: "octad-test-001", - limit: 10 - } - - assert {:ok, results} = Neo4j.query(@peer_info, query_params) - # octad-test-001 has RELATES_TO -> octad-test-002 (bidirectional), - # CITES -> octad-test-003, PART_OF -> octad-test-004 - assert length(results) >= 1 - end - - test "traversal from octad-test-004 finds all PART_OF components", context do - skip_if_unavailable(context) - - # Query with reversed relationship direction or multi-hop - part_of_peer = %{ - @peer_info - | adapter_config: Map.put(@peer_info.adapter_config, :relationship_type, "PART_OF") - } - - query_params = %{ - modalities: [:graph], - graph_pattern: "octad-test-001", - limit: 10 - } - - assert {:ok, results} = Neo4j.query(part_of_peer, query_params) - assert is_list(results) - end - end - - # --------------------------------------------------------------------------- - # 5. Provenance Chain Query - # --------------------------------------------------------------------------- - - describe "provenance chain traversal" do - test "semantic query for provenance events retrieves matching nodes", context do - skip_if_unavailable(context) - - # Query ProvenanceEvent nodes via semantic filter on entity_type - # (Since Neo4j adapter uses Octad label by default, we can filter - # by properties on Octad nodes that link to provenance events) - query_params = %{ - modalities: [:semantic], - filters: %{"drift_status" => "healthy"}, - limit: 10 - } - - assert {:ok, results} = Neo4j.query(@peer_info, query_params) - # octad-test-001, octad-test-003, octad-test-004 are 'healthy' - assert length(results) >= 1 - end - end - - # --------------------------------------------------------------------------- - # 6. Write + Read-Back - # --------------------------------------------------------------------------- - - describe "write and read-back cycle" do - test "translate_results correctly normalises Neo4j row format", context do - skip_if_unavailable(context) - - test_id = "#{@integration_prefix}-neo4j-#{System.unique_integer([:positive])}" - - # Simulate Neo4j transaction API row format - raw_row = %{ - "n" => %{ - "id" => test_id, - "title" => "Integration Test Node", - "entity_type" => "TestArticle", - "drift_status" => "healthy" - }, - "score" => 0.91 - } - - [normalised] = Neo4j.translate_results([raw_row], @peer_info) - - assert normalised.source_store == "neo4j-integration" - assert normalised.score == 0.91 - assert normalised.drifted == false - end - end - - # --------------------------------------------------------------------------- - # 7. Temporal Queries - # --------------------------------------------------------------------------- - - describe "temporal queries with datetime" do - test "temporal range query filters Octad nodes by created_at", context do - skip_if_unavailable(context) - - # All seed nodes were created relative to now() — query a wide range - now = DateTime.utc_now() - two_days_ago = DateTime.add(now, -2 * 86400, :second) - - query_params = %{ - modalities: [:temporal], - temporal_range: %{ - start: DateTime.to_iso8601(two_days_ago), - end: DateTime.to_iso8601(now) - }, - limit: 100 - } - - assert {:ok, results} = Neo4j.query(@peer_info, query_params) - assert length(results) >= 1 - end - end - - # --------------------------------------------------------------------------- - # 8. Error Handling - # --------------------------------------------------------------------------- - - describe "error handling against real Neo4j" do - test "invalid Cypher syntax returns an error", context do - skip_if_unavailable(context) - - # We cannot inject arbitrary Cypher through the adapter's query/3 API, - # but we can verify that a query with a nonexistent label returns empty. - nonexistent_peer = %{ - @peer_info - | adapter_config: Map.put(@peer_info.adapter_config, :label, "NonexistentLabel999") - } - - query_params = %{modalities: [], limit: 10} - - case Neo4j.query(nonexistent_peer, query_params) do - {:ok, results} -> - # Querying a nonexistent label returns empty results in Neo4j - assert results == [] - - {:error, _reason} -> - assert true - end - end - - test "connecting to an unreachable endpoint returns an error" do - unreachable_peer = %{ - store_id: "neo4j-unreachable", - endpoint: "http://localhost:59997", - adapter_config: %{database: "neo4j"} - } - - assert {:error, _reason} = Neo4j.connect(unreachable_peer) - end - end - - # --------------------------------------------------------------------------- - # 9. Modality Support - # --------------------------------------------------------------------------- - - describe "modality support declarations" do - test "Neo4j 5+ supports 6 modalities including vector" do - modalities = Neo4j.supported_modalities(%{version: 5}) - - assert :graph in modalities - assert :vector in modalities - assert :document in modalities - assert :temporal in modalities - assert :spatial in modalities - assert :semantic in modalities - - refute :tensor in modalities - refute :provenance in modalities - end - - test "Neo4j 4.x does not support vector search" do - modalities = Neo4j.supported_modalities(%{version: 4}) - - assert :graph in modalities - refute :vector in modalities - end - end - - # --------------------------------------------------------------------------- - # Helpers - # --------------------------------------------------------------------------- - - defp skip_if_unavailable(%{skip: true}), do: flunk("Neo4j not available — start test-infra stack") - defp skip_if_unavailable(_context), do: :ok -end diff --git a/verisimdb/elixir-orchestration/test/verisim/federation/adapters/integration/object_storage_integration_test.exs b/verisimdb/elixir-orchestration/test/verisim/federation/adapters/integration/object_storage_integration_test.exs deleted file mode 100644 index 6c77a455..00000000 --- a/verisimdb/elixir-orchestration/test/verisim/federation/adapters/integration/object_storage_integration_test.exs +++ /dev/null @@ -1,381 +0,0 @@ -# SPDX-License-Identifier: MPL-2.0 - -defmodule VeriSim.Federation.Adapters.ObjectStorageIntegrationTest do - @moduledoc """ - Integration tests for the ObjectStorage (MinIO/S3) federation adapter. - - Runs against a real MinIO instance from the test-infra container stack. - The seed script `minio-init.sh` pre-loads: - - - Buckets: `verisimdb-objects`, `verisimdb-backups`, `verisimdb-embeddings` - - `verisimdb-objects` bucket: - - `octads/octad-test-001/metadata.json` — octad metadata - - `octads/octad-test-001/document.txt` — document modality content - - `octads/octad-test-001/provenance.cbor` — provenance placeholder - - `octads/octad-test-002/metadata.json` — octad metadata - - `octads/octad-test-002/document.txt` — document modality content - - `verisimdb-embeddings` bucket: - - `octad-test-001.bin` — embedding binary blob (512 bytes) - - `octad-test-002.bin` — embedding binary blob (512 bytes) - - `verisimdb-backups` bucket: - - `snapshots/snap-test-001/metadata.json` — backup snapshot metadata - - Anonymous download policy on `verisimdb-objects` - - ## Test Infrastructure - - Requires the test-infra stack running: - - cd connectors/test-infra && selur-compose up -d - - MinIO is exposed on localhost:9002 (API) and localhost:9001 (console). - Credentials: verisim / verisim-test-password (as set in compose.toml). - - ## Running - - mix test --include integration test/verisim/federation/adapters/integration/object_storage_integration_test.exs - - Author: Jonathan D.A. Jewell - """ - - use ExUnit.Case, async: false - - alias VeriSim.Federation.Adapters.ObjectStorage - - @moduletag :integration - - # --------------------------------------------------------------------------- - # Configuration - # --------------------------------------------------------------------------- - - @minio_url System.get_env("VERISIM_MINIO_URL", "http://localhost:9002") - @minio_access_key System.get_env("VERISIM_MINIO_ACCESS_KEY", "verisim") - @minio_secret_key System.get_env("VERISIM_MINIO_SECRET_KEY", "verisim-test-password") - - @peer_info %{ - store_id: "minio-integration", - endpoint: @minio_url, - adapter_config: %{ - bucket: "verisimdb-objects", - backend: :minio, - region: "us-east-1", - access_key: @minio_access_key, - secret_key: @minio_secret_key - } - } - - @integration_prefix "octad-integration" - - # --------------------------------------------------------------------------- - # Setup / Teardown - # --------------------------------------------------------------------------- - - setup_all do - case ObjectStorage.health_check(@peer_info) do - {:ok, latency_ms} -> - {:ok, %{latency_ms: latency_ms}} - - {:error, reason} -> - {:ok, %{skip_reason: reason}} - end - end - - setup %{} = context do - if Map.has_key?(context, :skip_reason) do - {:ok, Map.put(context, :skip, true)} - else - {:ok, context} - end - end - - # --------------------------------------------------------------------------- - # 1. Connection Tests - # --------------------------------------------------------------------------- - - describe "connection to real MinIO" do - test "connect/1 succeeds against running instance", context do - skip_if_unavailable(context) - - result = ObjectStorage.connect(@peer_info) - assert result == :ok - end - - test "health_check/1 returns HEAD bucket success with latency", context do - skip_if_unavailable(context) - - assert {:ok, latency_ms} = ObjectStorage.health_check(@peer_info) - assert is_integer(latency_ms) - assert latency_ms >= 0 - end - end - - # --------------------------------------------------------------------------- - # 2. List Objects — Verify Seed Data - # --------------------------------------------------------------------------- - - describe "listing objects in verisimdb-objects bucket" do - test "default query lists objects in the bucket", context do - skip_if_unavailable(context) - - query_params = %{modalities: [], limit: 100} - assert {:ok, results} = ObjectStorage.query(@peer_info, query_params) - - # The seed script uploads 5 objects into verisimdb-objects - assert is_list(results) - assert length(results) >= 5 - - Enum.each(results, fn result -> - assert result.source_store == "minio-integration" - assert is_binary(result.octad_id) - assert result.drifted == false - end) - end - - test "listing with prefix 'octads/' returns octad-related objects", context do - skip_if_unavailable(context) - - query_params = %{ - modalities: [:document], - text_query: "octads/", - limit: 100 - } - - assert {:ok, results} = ObjectStorage.query(@peer_info, query_params) - assert is_list(results) - # Should find at least the 5 objects under octads/ prefix - assert length(results) >= 2 - end - end - - # --------------------------------------------------------------------------- - # 3. Get Specific Object - # --------------------------------------------------------------------------- - - describe "accessing specific objects" do - test "translate_results normalises S3 ListObjects format", context do - skip_if_unavailable(context) - - raw_obj = %{ - "Key" => "octads/octad-test-001/metadata.json", - "LastModified" => "2026-02-28T00:00:00Z", - "Size" => 512, - "ETag" => "\"abc123def456\"" - } - - [normalised] = ObjectStorage.translate_results([raw_obj], @peer_info) - - assert normalised.source_store == "minio-integration" - # The adapter extracts ID by stripping prefix and extension - assert normalised.octad_id == "octad-test-001/metadata" - assert normalised.drifted == false - end - - test "translate_results handles normalised (lowercase) S3 keys", context do - skip_if_unavailable(context) - - raw_obj = %{ - "key" => "octads/octad-test-002/document.txt", - "last_modified" => "2026-02-27T23:00:00Z", - "size" => 256, - "etag" => "\"xyz789\"" - } - - [normalised] = ObjectStorage.translate_results([raw_obj], @peer_info) - - assert normalised.source_store == "minio-integration" - assert is_binary(normalised.octad_id) - end - end - - # --------------------------------------------------------------------------- - # 4. Write + Read-Back - # --------------------------------------------------------------------------- - - describe "put and get object round-trip" do - test "translate_results correctly normalises a new object entry", context do - skip_if_unavailable(context) - - test_id = "#{@integration_prefix}-minio-#{System.unique_integer([:positive])}" - - raw_obj = %{ - "Key" => "octads/#{test_id}/metadata.json", - "LastModified" => DateTime.to_iso8601(DateTime.utc_now()), - "Size" => 128, - "ETag" => "\"integration-test-etag\"" - } - - [normalised] = ObjectStorage.translate_results([raw_obj], @peer_info) - - assert normalised.source_store == "minio-integration" - assert String.contains?(normalised.octad_id, test_id) - assert normalised.drifted == false - end - end - - # --------------------------------------------------------------------------- - # 5. Embeddings Bucket - # --------------------------------------------------------------------------- - - describe "accessing verisimdb-embeddings bucket" do - test "listing embeddings bucket returns binary blobs", context do - skip_if_unavailable(context) - - embeddings_peer = %{ - @peer_info - | adapter_config: Map.put(@peer_info.adapter_config, :bucket, "verisimdb-embeddings") - } - - query_params = %{modalities: [], limit: 100} - - case ObjectStorage.query(embeddings_peer, query_params) do - {:ok, results} -> - # Should find 2 embedding binary files - assert length(results) >= 2 - - {:error, _reason} -> - # Bucket may not be accessible with anonymous policy - assert true - end - end - end - - # --------------------------------------------------------------------------- - # 6. Backups Bucket - # --------------------------------------------------------------------------- - - describe "accessing verisimdb-backups bucket" do - test "listing backups bucket returns snapshot metadata", context do - skip_if_unavailable(context) - - backups_peer = %{ - @peer_info - | adapter_config: Map.put(@peer_info.adapter_config, :bucket, "verisimdb-backups") - } - - query_params = %{modalities: [], limit: 100} - - case ObjectStorage.query(backups_peer, query_params) do - {:ok, results} -> - # Should find at least the 1 snapshot metadata file - assert length(results) >= 1 - - {:error, _reason} -> - assert true - end - end - end - - # --------------------------------------------------------------------------- - # 7. Presigned URL Generation (Adapter-Level Concept) - # --------------------------------------------------------------------------- - - describe "presigned URL generation concept" do - test "adapter endpoint URL is correctly constructed for MinIO path-style", context do - skip_if_unavailable(context) - - # Verify the adapter uses path-style URLs for MinIO (not virtual-hosted) - # This is implicitly tested by the fact that queries work, but we can - # verify the URL construction logic is correct for the :minio backend. - config = @peer_info.adapter_config - assert config.backend == :minio - - # MinIO path-style URL: http://localhost:9002/verisimdb-objects - expected_base = "#{@minio_url}/verisimdb-objects" - assert String.starts_with?(expected_base, @minio_url) - end - end - - # --------------------------------------------------------------------------- - # 8. Error Handling - # --------------------------------------------------------------------------- - - describe "error handling against real MinIO" do - test "accessing a nonexistent object via translate_results handles gracefully" do - raw_obj = %{ - "Key" => "octads/nonexistent-object/metadata.json", - "LastModified" => "2026-02-28T00:00:00Z", - "Size" => 0, - "ETag" => "\"\"" - } - - [normalised] = ObjectStorage.translate_results([raw_obj], @peer_info) - - # translate_results should still produce a valid normalised result - assert normalised.source_store == "minio-integration" - assert is_binary(normalised.octad_id) - end - - test "health_check on nonexistent bucket returns bucket_not_found error" do - bad_peer = %{ - @peer_info - | adapter_config: Map.put(@peer_info.adapter_config, :bucket, "nonexistent-bucket-xyz") - } - - result = ObjectStorage.health_check(bad_peer) - - case result do - {:error, {:bucket_not_found, _}} -> assert true - {:error, {:access_denied, _}} -> assert true - {:error, {:unhealthy, _}} -> assert true - {:error, _other} -> assert true - {:ok, _} -> assert true - end - end - - test "connecting to an unreachable endpoint returns an error" do - unreachable_peer = %{ - store_id: "minio-unreachable", - endpoint: "http://localhost:59993", - adapter_config: %{bucket: "verisimdb-objects", backend: :minio} - } - - assert {:error, _reason} = ObjectStorage.connect(unreachable_peer) - end - end - - # --------------------------------------------------------------------------- - # 9. Modality Support - # --------------------------------------------------------------------------- - - describe "modality support declarations" do - test "base ObjectStorage supports document and semantic" do - modalities = ObjectStorage.supported_modalities(%{}) - - assert :document in modalities - assert :semantic in modalities - - refute :temporal in modalities - refute :provenance in modalities - end - - test "versioning enabled adds temporal modality" do - config = %{versioning: true} - modalities = ObjectStorage.supported_modalities(config) - - assert :temporal in modalities - end - - test "access logging enabled adds provenance modality" do - config = %{access_logging: true} - modalities = ObjectStorage.supported_modalities(config) - - assert :provenance in modalities - end - - test "full config enables all 4 modalities" do - config = %{versioning: true, access_logging: true} - modalities = ObjectStorage.supported_modalities(config) - - assert :document in modalities - assert :semantic in modalities - assert :temporal in modalities - assert :provenance in modalities - end - end - - # --------------------------------------------------------------------------- - # Helpers - # --------------------------------------------------------------------------- - - defp skip_if_unavailable(%{skip: true}), do: flunk("MinIO not available — start test-infra stack") - defp skip_if_unavailable(_context), do: :ok -end diff --git a/verisimdb/elixir-orchestration/test/verisim/federation/adapters/integration/redis_integration_test.exs b/verisimdb/elixir-orchestration/test/verisim/federation/adapters/integration/redis_integration_test.exs deleted file mode 100644 index 7f097b33..00000000 --- a/verisimdb/elixir-orchestration/test/verisim/federation/adapters/integration/redis_integration_test.exs +++ /dev/null @@ -1,358 +0,0 @@ -# SPDX-License-Identifier: MPL-2.0 - -defmodule VeriSim.Federation.Adapters.RedisIntegrationTest do - @moduledoc """ - Integration tests for the Redis federation adapter. - - Runs against a real Redis Stack instance from the test-infra container - stack. Redis Stack bundles RediSearch, RedisJSON, RedisTimeSeries, and - RedisGraph modules. The seed script `redis-init.sh` pre-loads: - - - 3 octad JSON documents (octad:test-001, octad:test-002, octad:test-003) - - 2 drift score documents (drift:test-001, drift:test-002) - - 2 RediSearch indexes (idx:octads, idx:drift) - - 3 RedisTimeSeries keys with sample data points - - ## Test Infrastructure - - Requires the test-infra stack running: - - cd connectors/test-infra && selur-compose up -d - - Redis Stack is exposed on localhost:6379. - - ## Running - - mix test --include integration test/verisim/federation/adapters/integration/redis_integration_test.exs - - Author: Jonathan D.A. Jewell - """ - - use ExUnit.Case, async: false - - alias VeriSim.Federation.Adapters.Redis - - @moduletag :integration - - # --------------------------------------------------------------------------- - # Configuration - # --------------------------------------------------------------------------- - - @redis_url System.get_env("VERISIM_REDIS_URL", "http://localhost:6379") - - @peer_info %{ - store_id: "redis-integration", - endpoint: @redis_url, - adapter_config: %{ - database: 0, - modules: [:redisearch, :redisgraph, :redisjson, :redistimeseries], - index_name: "idx:octads", - json_key_pattern: "octad:*" - } - } - - # Prefix for integration test data - @integration_prefix "octad-integration" - - # --------------------------------------------------------------------------- - # Setup / Teardown - # --------------------------------------------------------------------------- - - setup_all do - case Redis.health_check(@peer_info) do - {:ok, latency_ms} -> - {:ok, %{latency_ms: latency_ms}} - - {:error, reason} -> - {:ok, %{skip_reason: reason}} - end - end - - setup %{} = context do - if Map.has_key?(context, :skip_reason) do - {:ok, Map.put(context, :skip, true)} - else - {:ok, context} - end - end - - # --------------------------------------------------------------------------- - # 1. Connection Tests - # --------------------------------------------------------------------------- - - describe "connection to real Redis Stack" do - test "connect/1 succeeds against running instance", context do - skip_if_unavailable(context) - - result = Redis.connect(@peer_info) - assert result == :ok - end - - test "health_check/1 returns PONG with latency", context do - skip_if_unavailable(context) - - assert {:ok, latency_ms} = Redis.health_check(@peer_info) - assert is_integer(latency_ms) - assert latency_ms >= 0 - end - end - - # --------------------------------------------------------------------------- - # 2. RediSearch Full-Text Search (FT.SEARCH) - # --------------------------------------------------------------------------- - - describe "FT.SEARCH on idx:octads index" do - test "searching for 'consistency' returns matching octads", context do - skip_if_unavailable(context) - - query_params = %{ - modalities: [:document], - text_query: "consistency", - limit: 10 - } - - assert {:ok, results} = Redis.query(@peer_info, query_params) - assert is_list(results) - - # All results should be properly normalised - Enum.each(results, fn result -> - assert result.source_store == "redis-integration" - assert is_binary(result.octad_id) - assert is_number(result.score) - end) - end - - test "searching with wildcard '*' returns all indexed documents", context do - skip_if_unavailable(context) - - query_params = %{ - modalities: [:document], - text_query: "*", - limit: 100 - } - - assert {:ok, results} = Redis.query(@peer_info, query_params) - # Seed loads 3 octad documents into the index - assert length(results) >= 3 - end - end - - # --------------------------------------------------------------------------- - # 3. RedisJSON (JSON.GET) - # --------------------------------------------------------------------------- - - describe "RedisJSON document access" do - test "translate_results normalises a JSON document shape", context do - skip_if_unavailable(context) - - # Simulate a JSON document as returned by the Redis HTTP bridge - raw = [ - %{ - "id" => "octad:test-001", - "score" => 0.85, - "payload" => %{ - "title" => "Introduction to Cross-Modal Consistency", - "version" => 3 - } - } - ] - - [result] = Redis.translate_results(raw, @peer_info) - - assert result.source_store == "redis-integration" - assert result.octad_id == "octad:test-001" - assert result.score == 0.85 - assert result.drifted == false - end - end - - # --------------------------------------------------------------------------- - # 4. RedisTimeSeries (TS.RANGE) - # --------------------------------------------------------------------------- - - describe "RedisTimeSeries temporal queries" do - test "temporal range query returns time-series data", context do - skip_if_unavailable(context) - - # The seed script creates ts:drift:test-001:overall with sample data - ts_peer = %{ - @peer_info - | adapter_config: - Map.merge(@peer_info.adapter_config, %{ - timeseries_key: "ts:drift:test-001:overall" - }) - } - - query_params = %{ - modalities: [:temporal], - temporal_range: %{ - start: "-", - end: "+" - }, - limit: 100 - } - - assert {:ok, results} = Redis.query(ts_peer, query_params) - assert is_list(results) - end - end - - # --------------------------------------------------------------------------- - # 5. Write + Read-Back via translate_results - # --------------------------------------------------------------------------- - - describe "write and read-back cycle" do - test "translate_results correctly normalises integration test data", context do - skip_if_unavailable(context) - - test_id = "#{@integration_prefix}-redis-#{System.unique_integer([:positive])}" - - raw_doc = %{ - "id" => test_id, - "score" => 0.62, - "payload" => %{ - "title" => "Integration test octad", - "content" => "Written by RedisIntegrationTest" - } - } - - [normalised] = Redis.translate_results([raw_doc], @peer_info) - - assert normalised.source_store == "redis-integration" - assert normalised.octad_id == test_id - assert normalised.score == 0.62 - assert normalised.drifted == false - end - end - - # --------------------------------------------------------------------------- - # 6. Provenance (Redis Streams) - # --------------------------------------------------------------------------- - - describe "Redis Streams provenance queries" do - test "provenance query via XREVRANGE returns stream entries", context do - skip_if_unavailable(context) - - stream_peer = %{ - @peer_info - | adapter_config: - Map.merge(@peer_info.adapter_config, %{ - stream_key: "octad_provenance" - }) - } - - query_params = %{ - modalities: [:provenance], - limit: 10 - } - - assert {:ok, results} = Redis.query(stream_peer, query_params) - assert is_list(results) - end - - test "supported_modalities includes :provenance (Streams are built-in)" do - # Provenance via Redis Streams is always available - modalities = Redis.supported_modalities(%{modules: []}) - assert :provenance in modalities - end - end - - # --------------------------------------------------------------------------- - # 7. Vector Similarity Search - # --------------------------------------------------------------------------- - - describe "vector similarity search (RediSearch VSS)" do - test "supported_modalities includes :vector when redisearch module present" do - config = %{modules: [:redisearch]} - modalities = Redis.supported_modalities(config) - assert :vector in modalities - end - - test "vector query builds correct FT.SEARCH command", context do - skip_if_unavailable(context) - - # This tests that the adapter can construct and attempt a vector query. - # The actual search may return an error if no vector index is configured - # on the test instance, which is acceptable. - query_params = %{ - modalities: [:vector], - vector_query: [0.1, 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8], - limit: 5 - } - - result = Redis.query(@peer_info, query_params) - # Either succeeds or returns an error (no vector index in seed data) - assert match?({:ok, _}, result) or match?({:error, _}, result) - end - end - - # --------------------------------------------------------------------------- - # 8. Error Handling - # --------------------------------------------------------------------------- - - describe "error handling against real Redis" do - test "querying a nonexistent index returns an error", context do - skip_if_unavailable(context) - - bad_peer = %{ - @peer_info - | adapter_config: Map.put(@peer_info.adapter_config, :index_name, "idx:nonexistent_xyz") - } - - query_params = %{ - modalities: [:document], - text_query: "test", - limit: 10 - } - - result = Redis.query(bad_peer, query_params) - # Should return an error or empty results - assert match?({:ok, _}, result) or match?({:error, _}, result) - end - - test "connecting to an unreachable endpoint returns an error" do - unreachable_peer = %{ - store_id: "redis-unreachable", - endpoint: "http://localhost:59998", - adapter_config: %{database: 0, modules: []} - } - - assert {:error, _reason} = Redis.connect(unreachable_peer) - end - end - - # --------------------------------------------------------------------------- - # 9. Module-Dependent Modality Detection - # --------------------------------------------------------------------------- - - describe "module-dependent modality detection" do - test "all modules yields full modality set" do - config = %{modules: [:redisgraph, :redisearch, :redisjson, :redistimeseries]} - modalities = Redis.supported_modalities(config) - - assert :graph in modalities - assert :document in modalities - assert :semantic in modalities - assert :temporal in modalities - assert :vector in modalities - assert :provenance in modalities - end - - test "missing modules reduces available modalities" do - config = %{modules: []} - modalities = Redis.supported_modalities(config) - - # Only provenance (Streams) is built-in - assert modalities == [:provenance] - refute :graph in modalities - refute :vector in modalities - end - end - - # --------------------------------------------------------------------------- - # Helpers - # --------------------------------------------------------------------------- - - defp skip_if_unavailable(%{skip: true}), do: flunk("Redis Stack not available — start test-infra stack") - defp skip_if_unavailable(_context), do: :ok -end diff --git a/verisimdb/elixir-orchestration/test/verisim/federation/adapters/integration/surrealdb_integration_test.exs b/verisimdb/elixir-orchestration/test/verisim/federation/adapters/integration/surrealdb_integration_test.exs deleted file mode 100644 index 8b6e1b9b..00000000 --- a/verisimdb/elixir-orchestration/test/verisim/federation/adapters/integration/surrealdb_integration_test.exs +++ /dev/null @@ -1,373 +0,0 @@ -# SPDX-License-Identifier: MPL-2.0 - -defmodule VeriSim.Federation.Adapters.SurrealDBIntegrationTest do - @moduledoc """ - Integration tests for the SurrealDB federation adapter. - - Runs against a real SurrealDB instance from the test-infra container - stack. The seed script `surrealdb-init.surql` pre-loads: - - - Namespace `verisimdb`, database `test` - - `octads` table: 4 schemafull records (test001-test004) with title, - content, entity_type, drift_status, drift_score, tags, timestamps - - Edge tables: `relates_to` (2), `cites` (1), `part_of` (3), `derived_from` (1) - - `modalities` table: 5 records - - `drift_scores` table: 2 records - - `provenance_events` table: 2 records - - Fulltext index `idx_octads_fulltext` with BM25 scoring - - Field-level indexes on entity_type, drift_status, created_at - - ## Test Infrastructure - - Requires the test-infra stack running: - - cd connectors/test-infra && selur-compose up -d - - SurrealDB is exposed on localhost:8000. - Default credentials: root/root (test only). - - ## Running - - mix test --include integration test/verisim/federation/adapters/integration/surrealdb_integration_test.exs - - Author: Jonathan D.A. Jewell - """ - - use ExUnit.Case, async: false - - alias VeriSim.Federation.Adapters.SurrealDB - - @moduletag :integration - - # --------------------------------------------------------------------------- - # Configuration - # --------------------------------------------------------------------------- - - @surrealdb_url System.get_env("VERISIM_SURREALDB_URL", "http://localhost:8000") - - @peer_info %{ - store_id: "surreal-integration", - endpoint: @surrealdb_url, - adapter_config: %{ - namespace: "verisimdb", - database: "test", - table: "octads", - auth: {:basic, "root", "root"}, - edge_table: "relates_to", - search_fields: ["title", "content"], - max_depth: 3 - } - } - - @integration_prefix "octad-integration" - - # --------------------------------------------------------------------------- - # Setup / Teardown - # --------------------------------------------------------------------------- - - setup_all do - case SurrealDB.health_check(@peer_info) do - {:ok, latency_ms} -> - {:ok, %{latency_ms: latency_ms}} - - {:error, reason} -> - {:ok, %{skip_reason: reason}} - end - end - - setup %{} = context do - if Map.has_key?(context, :skip_reason) do - {:ok, Map.put(context, :skip, true)} - else - {:ok, context} - end - end - - # --------------------------------------------------------------------------- - # 1. Connection Tests - # --------------------------------------------------------------------------- - - describe "connection to real SurrealDB" do - test "connect/1 succeeds against running instance", context do - skip_if_unavailable(context) - - result = SurrealDB.connect(@peer_info) - assert result == :ok - end - - test "health_check/1 returns 200 with latency", context do - skip_if_unavailable(context) - - assert {:ok, latency_ms} = SurrealDB.health_check(@peer_info) - assert is_integer(latency_ms) - assert latency_ms >= 0 - end - end - - # --------------------------------------------------------------------------- - # 2. Read / Query — Verify Seed Data - # --------------------------------------------------------------------------- - - describe "querying seeded octads table" do - test "SELECT * returns all 4 seeded records", context do - skip_if_unavailable(context) - - query_params = %{modalities: [], limit: 100} - assert {:ok, results} = SurrealDB.query(@peer_info, query_params) - - # The seed script creates 4 octad records - assert length(results) >= 4 - - Enum.each(results, fn result -> - assert result.source_store == "surreal-integration" - assert is_binary(result.octad_id) - assert is_number(result.score) - assert result.drifted == false - assert is_map(result.data) - end) - end - - test "SurrealDB record IDs are correctly stripped of table prefix", context do - skip_if_unavailable(context) - - query_params = %{modalities: [], limit: 10} - assert {:ok, results} = SurrealDB.query(@peer_info, query_params) - - # SurrealDB IDs are "octads:test001" — adapter should extract "test001" - ids = Enum.map(results, & &1.octad_id) - - # At least one of the seeded IDs should be present (without table prefix) - assert Enum.any?(ids, fn id -> - id in ["test001", "test002", "test003", "test004"] - end) - end - end - - # --------------------------------------------------------------------------- - # 3. Edge Traversal (Graph Modality) - # --------------------------------------------------------------------------- - - describe "graph edge traversal" do - test "relates_to traversal from test001 finds connected octads", context do - skip_if_unavailable(context) - - query_params = %{ - modalities: [:graph], - graph_pattern: "test001", - limit: 10 - } - - assert {:ok, results} = SurrealDB.query(@peer_info, query_params) - # test001 -> relates_to -> test002 - assert is_list(results) - end - - test "derived_from traversal finds derivation chain", context do - skip_if_unavailable(context) - - derived_peer = %{ - @peer_info - | adapter_config: Map.put(@peer_info.adapter_config, :edge_table, "derived_from") - } - - query_params = %{ - modalities: [:graph], - graph_pattern: "test003", - limit: 10 - } - - assert {:ok, results} = SurrealDB.query(derived_peer, query_params) - # test003 -> derived_from -> test001 - assert is_list(results) - end - - test "part_of traversal from test001 finds parent concept", context do - skip_if_unavailable(context) - - part_of_peer = %{ - @peer_info - | adapter_config: Map.put(@peer_info.adapter_config, :edge_table, "part_of") - } - - query_params = %{ - modalities: [:graph], - graph_pattern: "test001", - limit: 10 - } - - assert {:ok, results} = SurrealDB.query(part_of_peer, query_params) - assert is_list(results) - end - end - - # --------------------------------------------------------------------------- - # 4. Fulltext Search - # --------------------------------------------------------------------------- - - describe "fulltext search via SurrealDB analyzer" do - test "text search for 'consistency' returns matching records", context do - skip_if_unavailable(context) - - query_params = %{ - modalities: [:document], - text_query: "consistency", - limit: 10 - } - - assert {:ok, results} = SurrealDB.query(@peer_info, query_params) - # octad test001 title contains "Consistency" - assert length(results) >= 1 - end - - test "text search for 'federation' returns octad test004", context do - skip_if_unavailable(context) - - query_params = %{ - modalities: [:document], - text_query: "federation", - limit: 10 - } - - assert {:ok, results} = SurrealDB.query(@peer_info, query_params) - assert length(results) >= 1 - end - end - - # --------------------------------------------------------------------------- - # 5. Temporal Queries - # --------------------------------------------------------------------------- - - describe "temporal queries with datetime" do - test "temporal range query filters records by created_at", context do - skip_if_unavailable(context) - - now = DateTime.utc_now() - two_days_ago = DateTime.add(now, -2 * 86400, :second) - - query_params = %{ - modalities: [:temporal], - temporal_range: %{ - start: DateTime.to_iso8601(two_days_ago), - end: DateTime.to_iso8601(now) - }, - limit: 100 - } - - assert {:ok, results} = SurrealDB.query(@peer_info, query_params) - assert length(results) >= 1 - end - end - - # --------------------------------------------------------------------------- - # 6. Write + Read-Back - # --------------------------------------------------------------------------- - - describe "CREATE and SELECT round-trip" do - test "translate_results correctly normalises SurrealDB record format", context do - skip_if_unavailable(context) - - test_id = "#{@integration_prefix}-surreal-#{System.unique_integer([:positive])}" - - # Simulate SurrealDB response format - raw_record = %{ - "id" => "octads:#{test_id}", - "title" => "Integration Test Record", - "entity_type" => "TestArticle", - "drift_status" => "healthy", - "score" => 0.82 - } - - [normalised] = SurrealDB.translate_results([raw_record], @peer_info) - - assert normalised.source_store == "surreal-integration" - # The adapter strips the "octads:" table prefix from the ID - assert normalised.octad_id == test_id - assert normalised.score == 0.82 - assert normalised.drifted == false - end - end - - # --------------------------------------------------------------------------- - # 7. Semantic (Metadata) Queries - # --------------------------------------------------------------------------- - - describe "semantic metadata queries" do - test "filtering by metadata.entity_type returns matching records", context do - skip_if_unavailable(context) - - query_params = %{ - modalities: [:semantic], - filters: %{"entity_type" => "Article"}, - limit: 10 - } - - assert {:ok, results} = SurrealDB.query(@peer_info, query_params) - # test001 and test003 are entity_type "Article" - assert is_list(results) - end - end - - # --------------------------------------------------------------------------- - # 8. Error Handling - # --------------------------------------------------------------------------- - - describe "error handling against real SurrealDB" do - test "invalid SurrealQL returns an error", context do - skip_if_unavailable(context) - - # Query a nonexistent table — SurrealDB may return empty or error - bad_peer = %{ - @peer_info - | adapter_config: Map.put(@peer_info.adapter_config, :table, "nonexistent_table_xyz") - } - - query_params = %{modalities: [], limit: 10} - - case SurrealDB.query(bad_peer, query_params) do - {:ok, results} -> - # SurrealDB may return empty for nonexistent tables - assert results == [] - - {:error, _reason} -> - assert true - end - end - - test "connecting to an unreachable endpoint returns an error" do - unreachable_peer = %{ - store_id: "surreal-unreachable", - endpoint: "http://localhost:59995", - adapter_config: %{namespace: "verisimdb", database: "test"} - } - - assert {:error, _reason} = SurrealDB.connect(unreachable_peer) - end - end - - # --------------------------------------------------------------------------- - # 9. Modality Support - # --------------------------------------------------------------------------- - - describe "modality support declarations" do - test "SurrealDB supports 4 modalities" do - modalities = SurrealDB.supported_modalities(%{}) - - assert :graph in modalities - assert :document in modalities - assert :temporal in modalities - assert :semantic in modalities - - refute :vector in modalities - refute :tensor in modalities - refute :provenance in modalities - refute :spatial in modalities - end - end - - # --------------------------------------------------------------------------- - # Helpers - # --------------------------------------------------------------------------- - - defp skip_if_unavailable(%{skip: true}), do: flunk("SurrealDB not available — start test-infra stack") - defp skip_if_unavailable(_context), do: :ok -end diff --git a/verisimdb/elixir-orchestration/test/verisim/federation/adapters/mongodb_test.exs b/verisimdb/elixir-orchestration/test/verisim/federation/adapters/mongodb_test.exs deleted file mode 100644 index 8a222171..00000000 --- a/verisimdb/elixir-orchestration/test/verisim/federation/adapters/mongodb_test.exs +++ /dev/null @@ -1,93 +0,0 @@ -# SPDX-License-Identifier: MPL-2.0 - -defmodule VeriSim.Federation.Adapters.MongoDBTest do - @moduledoc """ - Tests for the MongoDB federation adapter. - - Validates modality declarations, result normalisation from MongoDB - document format, and aggregation pipeline construction patterns. - """ - - use ExUnit.Case, async: true - - alias VeriSim.Federation.Adapters.MongoDB - - @peer_info %{ - store_id: "mongo-test", - endpoint: "http://mongo:27017", - adapter_config: %{database: "verisimdb", collection: "octads"} - } - - # --------------------------------------------------------------------------- - # Supported Modalities - # --------------------------------------------------------------------------- - - describe "supported_modalities/1" do - test "returns 5 base supported modalities (without Atlas or replica set)" do - modalities = MongoDB.supported_modalities(%{}) - - assert :graph in modalities - assert :document in modalities - assert :temporal in modalities - assert :spatial in modalities - assert :semantic in modalities - - # :vector requires atlas: true, :provenance requires replica_set key - refute :vector in modalities - refute :provenance in modalities - refute :tensor in modalities - end - - test "vector requires atlas: true" do - modalities = MongoDB.supported_modalities(%{atlas: false}) - refute :vector in modalities - - modalities_with_atlas = MongoDB.supported_modalities(%{atlas: true}) - assert :vector in modalities_with_atlas - end - end - - # --------------------------------------------------------------------------- - # Result Normalisation - # --------------------------------------------------------------------------- - - describe "translate_results/2" do - test "extracts _id from MongoDB documents" do - raw = [%{"_id" => "64a1b2c3d4e5f67890abcdef", "title" => "Test", "score" => 0.9}] - - [result] = MongoDB.translate_results(raw, @peer_info) - - assert result.source_store == "mongo-test" - assert result.octad_id == "64a1b2c3d4e5f67890abcdef" - assert result.score == 0.9 - assert result.drifted == false - end - - test "handles missing _id gracefully" do - raw = [%{"title" => "No ID"}] - - [result] = MongoDB.translate_results(raw, @peer_info) - assert result.octad_id == "unknown" - end - - test "handles empty result list" do - assert MongoDB.translate_results([], @peer_info) == [] - end - - test "normalises Atlas Vector Search results with score" do - # The adapter's parse_score/1 checks doc["score"], which is set by the - # $addFields stage in the vector search pipeline (vectorSearchScore meta). - # The translated results contain the "score" key, not "searchScore". - raw = [ - %{ - "_id" => "doc-1", - "score" => 0.95, - "title" => "Vector match" - } - ] - - [result] = MongoDB.translate_results(raw, @peer_info) - assert result.score == 0.95 - end - end -end diff --git a/verisimdb/elixir-orchestration/test/verisim/federation/adapters/neo4j_test.exs b/verisimdb/elixir-orchestration/test/verisim/federation/adapters/neo4j_test.exs deleted file mode 100644 index a2a49e89..00000000 --- a/verisimdb/elixir-orchestration/test/verisim/federation/adapters/neo4j_test.exs +++ /dev/null @@ -1,59 +0,0 @@ -# SPDX-License-Identifier: MPL-2.0 - -defmodule VeriSim.Federation.Adapters.Neo4jTest do - @moduledoc """ - Tests for the Neo4j federation adapter. - - Validates modality declarations, Cypher query construction, - and result normalisation from Neo4j's transactional HTTP response format. - """ - - use ExUnit.Case, async: true - - alias VeriSim.Federation.Adapters.Neo4j - - @peer_info %{ - store_id: "neo4j-test", - endpoint: "http://neo4j:7474", - adapter_config: %{database: "neo4j"} - } - - describe "supported_modalities/1" do - test "returns 6 supported modalities" do - modalities = Neo4j.supported_modalities(%{}) - - assert :graph in modalities - assert :vector in modalities - assert :document in modalities - assert :temporal in modalities - assert :spatial in modalities - assert :semantic in modalities - - refute :tensor in modalities - refute :provenance in modalities - end - end - - describe "translate_results/2" do - test "extracts Neo4j node properties" do - raw = [ - %{ - "id" => "neo4j-001", - "labels" => ["Octad"], - "properties" => %{"title" => "Graph entity"}, - "score" => 0.92 - } - ] - - [result] = Neo4j.translate_results(raw, @peer_info) - - assert result.octad_id == "neo4j-001" - assert result.score == 0.92 - assert result.source_store == "neo4j-test" - end - - test "handles empty results" do - assert Neo4j.translate_results([], @peer_info) == [] - end - end -end diff --git a/verisimdb/elixir-orchestration/test/verisim/federation/adapters/object_storage_test.exs b/verisimdb/elixir-orchestration/test/verisim/federation/adapters/object_storage_test.exs deleted file mode 100644 index 8431df32..00000000 --- a/verisimdb/elixir-orchestration/test/verisim/federation/adapters/object_storage_test.exs +++ /dev/null @@ -1,59 +0,0 @@ -# SPDX-License-Identifier: MPL-2.0 - -defmodule VeriSim.Federation.Adapters.ObjectStorageTest do - @moduledoc """ - Tests for the unified ObjectStorage federation adapter. - - Validates modality declarations for MinIO/S3 backends, - result normalisation from S3 API response format, and - metadata-based query patterns. - """ - - use ExUnit.Case, async: true - - alias VeriSim.Federation.Adapters.ObjectStorage - - @peer_info %{ - store_id: "minio-test", - endpoint: "http://minio:9000", - adapter_config: %{bucket: "verisim-octads", backend: :minio} - } - - describe "supported_modalities/1" do - test "returns 2 base supported modalities" do - modalities = ObjectStorage.supported_modalities(%{}) - - assert :document in modalities - assert :semantic in modalities - - refute :temporal in modalities - refute :provenance in modalities - refute :graph in modalities - refute :vector in modalities - refute :tensor in modalities - refute :spatial in modalities - end - end - - describe "translate_results/2" do - test "normalises S3 ListObjects results" do - raw = [ - %{ - "Key" => "octads/obj-001.json", - "LastModified" => "2026-02-28T12:00:00Z", - "Size" => 1024, - "ETag" => "\"abc123\"" - } - ] - - [result] = ObjectStorage.translate_results(raw, @peer_info) - - assert result.octad_id == "obj-001" - assert result.source_store == "minio-test" - end - - test "handles empty results" do - assert ObjectStorage.translate_results([], @peer_info) == [] - end - end -end diff --git a/verisimdb/elixir-orchestration/test/verisim/federation/adapters/redis_test.exs b/verisimdb/elixir-orchestration/test/verisim/federation/adapters/redis_test.exs deleted file mode 100644 index 5e5b7f0d..00000000 --- a/verisimdb/elixir-orchestration/test/verisim/federation/adapters/redis_test.exs +++ /dev/null @@ -1,76 +0,0 @@ -# SPDX-License-Identifier: MPL-2.0 - -defmodule VeriSim.Federation.Adapters.RedisTest do - @moduledoc """ - Tests for the Redis federation adapter. - - Validates module-dependent modality declarations, result normalisation - from Redis command responses, and RediSearch/RedisJSON integration. - """ - - use ExUnit.Case, async: true - - alias VeriSim.Federation.Adapters.Redis - - @peer_info %{ - store_id: "redis-test", - endpoint: "http://redis:6379", - adapter_config: %{database: 0} - } - - # --------------------------------------------------------------------------- - # Supported Modalities - # --------------------------------------------------------------------------- - - describe "supported_modalities/1" do - test "with no modules declared returns provenance (Redis Streams built-in)" do - modalities = Redis.supported_modalities(%{modules: []}) - assert modalities == [:provenance] - end - - test "with all modules returns full set" do - config = %{ - modules: [:redisgraph, :redisearch, :redisjson, :redistimeseries] - } - - modalities = Redis.supported_modalities(config) - - assert :graph in modalities - assert :document in modalities - assert :semantic in modalities - assert :temporal in modalities - assert :vector in modalities - assert :provenance in modalities - end - - test "vector requires redisearch module" do - modalities = Redis.supported_modalities(%{modules: [:redisjson]}) - refute :vector in modalities - end - - test "graph requires redisgraph module" do - modalities = Redis.supported_modalities(%{modules: [:redisearch]}) - refute :graph in modalities - end - end - - # --------------------------------------------------------------------------- - # Result Normalisation - # --------------------------------------------------------------------------- - - describe "translate_results/2" do - test "normalises RediSearch results" do - raw = [%{"id" => "redis:key:1", "score" => 0.8, "payload" => %{"title" => "Test"}}] - - [result] = Redis.translate_results(raw, @peer_info) - - assert result.source_store == "redis-test" - assert result.octad_id == "redis:key:1" - assert result.score == 0.8 - end - - test "handles empty results" do - assert Redis.translate_results([], @peer_info) == [] - end - end -end diff --git a/verisimdb/elixir-orchestration/test/verisim/federation/adapters/sqlite_test.exs b/verisimdb/elixir-orchestration/test/verisim/federation/adapters/sqlite_test.exs deleted file mode 100644 index 4d02ad84..00000000 --- a/verisimdb/elixir-orchestration/test/verisim/federation/adapters/sqlite_test.exs +++ /dev/null @@ -1,58 +0,0 @@ -# SPDX-License-Identifier: MPL-2.0 - -defmodule VeriSim.Federation.Adapters.SQLiteTest do - @moduledoc """ - Tests for the SQLite federation adapter. - - Validates extension-dependent modality declarations, FTS5 query - construction, and result normalisation from SQLite row format. - """ - - use ExUnit.Case, async: true - - alias VeriSim.Federation.Adapters.SQLite - - @peer_info %{ - store_id: "sqlite-test", - endpoint: "http://sqlite-proxy:8080", - adapter_config: %{path: "/data/verisim.db", table: "octads"} - } - - describe "supported_modalities/1" do - test "base modalities without extensions" do - modalities = SQLite.supported_modalities(%{extensions: []}) - - assert :graph in modalities - assert :temporal in modalities - assert :semantic in modalities - refute :document in modalities - refute :vector in modalities - end - - test "with sqlite-vss adds vector modality" do - modalities = SQLite.supported_modalities(%{extensions: [:vss]}) - assert :vector in modalities - end - - test "FTS5 adds document modality when enabled" do - modalities = SQLite.supported_modalities(%{extensions: [:fts5]}) - assert :document in modalities - end - end - - describe "translate_results/2" do - test "normalises SQLite row results" do - raw = [%{"id" => "sqlite-001", "score" => 0.72, "title" => "Embedded"}] - - [result] = SQLite.translate_results(raw, @peer_info) - - assert result.octad_id == "sqlite-001" - assert result.score == 0.72 - assert result.source_store == "sqlite-test" - end - - test "handles empty results" do - assert SQLite.translate_results([], @peer_info) == [] - end - end -end diff --git a/verisimdb/elixir-orchestration/test/verisim/federation/adapters/surrealdb_test.exs b/verisimdb/elixir-orchestration/test/verisim/federation/adapters/surrealdb_test.exs deleted file mode 100644 index 9c81c246..00000000 --- a/verisimdb/elixir-orchestration/test/verisim/federation/adapters/surrealdb_test.exs +++ /dev/null @@ -1,52 +0,0 @@ -# SPDX-License-Identifier: MPL-2.0 - -defmodule VeriSim.Federation.Adapters.SurrealDBTest do - @moduledoc """ - Tests for the SurrealDB federation adapter. - - Validates modality declarations, SurrealQL query construction, - and result normalisation from SurrealDB's multi-model response format. - """ - - use ExUnit.Case, async: true - - alias VeriSim.Federation.Adapters.SurrealDB - - @peer_info %{ - store_id: "surreal-test", - endpoint: "http://surrealdb:8000", - adapter_config: %{namespace: "verisim", database: "main"} - } - - describe "supported_modalities/1" do - test "returns 4 supported modalities" do - modalities = SurrealDB.supported_modalities(%{}) - - assert :graph in modalities - assert :document in modalities - assert :temporal in modalities - assert :semantic in modalities - - refute :vector in modalities - refute :tensor in modalities - refute :provenance in modalities - refute :spatial in modalities - end - end - - describe "translate_results/2" do - test "extracts SurrealDB record ID" do - raw = [%{"id" => "octads:abc123", "title" => "Test", "score" => 0.88}] - - [result] = SurrealDB.translate_results(raw, @peer_info) - - assert result.octad_id == "abc123" - assert result.score == 0.88 - assert result.source_store == "surreal-test" - end - - test "handles empty results" do - assert SurrealDB.translate_results([], @peer_info) == [] - end - end -end diff --git a/verisimdb/elixir-orchestration/test/verisim/federation/adapters/vector_db_test.exs b/verisimdb/elixir-orchestration/test/verisim/federation/adapters/vector_db_test.exs deleted file mode 100644 index 6c477dae..00000000 --- a/verisimdb/elixir-orchestration/test/verisim/federation/adapters/vector_db_test.exs +++ /dev/null @@ -1,59 +0,0 @@ -# SPDX-License-Identifier: MPL-2.0 - -defmodule VeriSim.Federation.Adapters.VectorDBTest do - @moduledoc """ - Tests for the unified VectorDB federation adapter. - - Validates modality declarations across Qdrant/Milvus/Weaviate backends, - result normalisation from each backend's response format, and backend - dispatch routing. - """ - - use ExUnit.Case, async: true - - alias VeriSim.Federation.Adapters.VectorDB - - @peer_info %{ - store_id: "vector-test", - endpoint: "http://qdrant:6333", - adapter_config: %{collection: "octads", backend: :qdrant} - } - - describe "supported_modalities/1" do - test "returns 4 supported modalities" do - modalities = VectorDB.supported_modalities(%{}) - - assert :vector in modalities - assert :temporal in modalities - assert :spatial in modalities - assert :semantic in modalities - - refute :graph in modalities - refute :document in modalities - refute :tensor in modalities - refute :provenance in modalities - end - end - - describe "translate_results/2" do - test "normalises Qdrant point results" do - raw = [ - %{ - "id" => "vec-001", - "score" => 0.98, - "payload" => %{"title" => "Vector match"} - } - ] - - [result] = VectorDB.translate_results(raw, @peer_info) - - assert result.octad_id == "vec-001" - assert result.score == 0.98 - assert result.source_store == "vector-test" - end - - test "handles empty results" do - assert VectorDB.translate_results([], @peer_info) == [] - end - end -end diff --git a/verisimdb/elixir-orchestration/test/verisim/federation/resolver_test.exs b/verisimdb/elixir-orchestration/test/verisim/federation/resolver_test.exs deleted file mode 100644 index 0ae1372b..00000000 --- a/verisimdb/elixir-orchestration/test/verisim/federation/resolver_test.exs +++ /dev/null @@ -1,105 +0,0 @@ -# SPDX-License-Identifier: MPL-2.0 - -defmodule VeriSim.Federation.ResolverTest do - use ExUnit.Case, async: false - - alias VeriSim.Federation.Resolver - - # The Resolver is started by the Application supervisor as a named - # GenServer (__MODULE__). Tests use the app-managed instance and - # clean up peers between tests. - - setup do - # Clear any peers from previous tests - for peer <- Resolver.list_peers() do - Resolver.deregister_peer(peer.store_id) - end - - :ok - end - - describe "register and list peers" do - test "registered peer appears in list_peers" do - :ok = Resolver.register_peer("store-a", "http://a.local:8080", ["graph", "vector"]) - - peers = Resolver.list_peers() - assert length(peers) == 1 - - [peer] = peers - assert peer.store_id == "store-a" - assert peer.endpoint == "http://a.local:8080" - assert peer.modalities == ["graph", "vector"] - assert peer.trust_level == 1.0 - end - end - - describe "deregister peer" do - test "deregistered peer is removed from list" do - :ok = Resolver.register_peer("store-b", "http://b.local:8080", ["document"]) - assert length(Resolver.list_peers()) == 1 - - :ok = Resolver.deregister_peer("store-b") - assert Resolver.list_peers() == [] - end - end - - describe "query with no peers" do - test "returns empty results without crashing" do - {:ok, response} = Resolver.query("*", ["document"]) - - assert response.results == [] - assert response.stores_queried == [] - assert response.stores_excluded == [] - assert response.drift_policy == :tolerate - end - end - - describe "pattern matching" do - test "wildcard pattern matches all peers" do - :ok = Resolver.register_peer("prod/us-1", "http://us1:8080", ["graph"]) - :ok = Resolver.register_peer("prod/eu-1", "http://eu1:8080", ["graph"]) - :ok = Resolver.register_peer("dev/local", "http://local:8080", ["graph"]) - - {:ok, response} = Resolver.query("*", ["graph"], timeout: 2_000) - - # All 3 stores should be queried (will fail HTTP, but listed as queried) - assert length(response.stores_queried) == 3 - end - - test "prefix pattern filters correctly" do - :ok = Resolver.register_peer("prod/us-1", "http://us1:8080", ["graph"]) - :ok = Resolver.register_peer("prod/eu-1", "http://eu1:8080", ["graph"]) - :ok = Resolver.register_peer("dev/local", "http://local:8080", ["graph"]) - - {:ok, response} = Resolver.query("prod/*", ["graph"], timeout: 2_000) - - assert length(response.stores_queried) == 2 - assert "dev/local" not in response.stores_queried - end - end - - describe "drift policy strict" do - test "excludes peers below trust threshold" do - :ok = Resolver.register_peer("trusted", "http://trusted:8080", ["graph"]) - :ok = Resolver.register_peer("untrusted", "http://untrusted:8080", ["graph"]) - - # Both start at trust 1.0 (above 0.7 threshold), so both should be queried - {:ok, response} = Resolver.query("*", ["graph"], drift_policy: :strict, timeout: 2_000) - - assert length(response.stores_queried) == 2 - assert response.stores_excluded == [] - end - end - - describe "modality filtering" do - test "only queries peers that support required modalities" do - :ok = Resolver.register_peer("graph-only", "http://g:8080", ["graph"]) - :ok = Resolver.register_peer("full-stack", "http://f:8080", ["graph", "vector", "document"]) - - {:ok, response} = Resolver.query("*", ["vector"], timeout: 2_000) - - assert length(response.stores_queried) == 1 - assert "full-stack" in response.stores_queried - end - end -end diff --git a/verisimdb/elixir-orchestration/test/verisim/hypatia/dispatch_bridge_test.exs b/verisimdb/elixir-orchestration/test/verisim/hypatia/dispatch_bridge_test.exs deleted file mode 100644 index 0d267e70..00000000 --- a/verisimdb/elixir-orchestration/test/verisim/hypatia/dispatch_bridge_test.exs +++ /dev/null @@ -1,217 +0,0 @@ -# SPDX-License-Identifier: MPL-2.0 - -defmodule VeriSim.Hypatia.DispatchBridgeTest do - @moduledoc """ - Tests for the Hypatia dispatch bridge module. - - Verifies that the bridge correctly: - 1. Reads pending dispatch actions from JSONL files - 2. Reads dispatch logs and outcomes - 3. Computes dispatch summaries - 4. Feeds outcomes back for drift tracking - """ - - use ExUnit.Case, async: true - - alias VeriSim.Hypatia.DispatchBridge - - setup do - dir = Path.join(System.tmp_dir!(), "hypatia_dispatch_#{System.unique_integer([:positive])}") - dispatch_dir = Path.join(dir, "dispatch") - outcomes_dir = Path.join(dir, "outcomes") - File.mkdir_p!(dispatch_dir) - File.mkdir_p!(outcomes_dir) - - on_exit(fn -> File.rm_rf!(dir) end) - {:ok, data_path: dir, dispatch_dir: dispatch_dir, outcomes_dir: outcomes_dir} - end - - # =========================================================================== - # Sample Data Helpers - # =========================================================================== - - defp write_jsonl(path, records) do - lines = Enum.map_join(records, "\n", &Jason.encode!/1) - File.write!(path, lines <> "\n") - end - - defp sample_pending do - [ - %{"repo" => "echidna", "pattern" => "PA001", "strategy" => "auto_execute", "confidence" => 0.98, - "mutation" => "mutation { createPR(repo: \"echidna\") }"}, - %{"repo" => "ambientops", "pattern" => "PA003", "strategy" => "review", "confidence" => 0.87, - "mutation" => "mutation { createPR(repo: \"ambientops\") }"}, - %{"repo" => "verisimdb", "pattern" => "PA001", "strategy" => "auto_execute", "confidence" => 0.96, - "mutation" => "mutation { createPR(repo: \"verisimdb\") }"} - ] - end - - defp sample_dispatch_log do - [ - %{"repo" => "echidna", "pattern" => "PA001", "strategy" => "auto_execute", - "status" => "dispatched", "timestamp" => "2026-02-12T10:00:00Z"}, - %{"repo" => "proven", "pattern" => "PA005", "strategy" => "report_only", - "status" => "dispatched", "timestamp" => "2026-02-12T10:01:00Z"}, - %{"repo" => "ambientops", "pattern" => "PA003", "strategy" => "review", - "status" => "dispatched", "timestamp" => "2026-02-12T10:02:00Z"} - ] - end - - defp sample_outcomes do - [ - %{"repo" => "echidna", "pattern" => "PA001", "status" => "success", - "timestamp" => "2026-02-12T11:00:00Z"}, - %{"repo" => "proven", "pattern" => "PA005", "status" => "success", - "timestamp" => "2026-02-12T11:01:00Z"}, - %{"repo" => "ambientops", "pattern" => "PA003", "status" => "failure", - "timestamp" => "2026-02-12T11:02:00Z"} - ] - end - - # =========================================================================== - # read_pending/1 - # =========================================================================== - - describe "read_pending/1" do - test "reads pending actions from JSONL", %{data_path: dp, dispatch_dir: dd} do - write_jsonl(Path.join(dd, "pending.jsonl"), sample_pending()) - - assert {:ok, actions} = DispatchBridge.read_pending(dp) - assert length(actions) == 3 - assert hd(actions)["repo"] == "echidna" - end - - test "returns error when file missing", %{data_path: dp} do - assert {:error, _} = DispatchBridge.read_pending(dp) - end - end - - # =========================================================================== - # read_dispatch_log/2 - # =========================================================================== - - describe "read_dispatch_log/2" do - test "reads dispatch log for a specific date", %{data_path: dp, dispatch_dir: dd} do - write_jsonl(Path.join(dd, "dispatch-2026-02-12.jsonl"), sample_dispatch_log()) - - assert {:ok, records} = DispatchBridge.read_dispatch_log(dp, "2026-02-12") - assert length(records) == 3 - end - - test "returns error for non-existent date", %{data_path: dp} do - assert {:error, _} = DispatchBridge.read_dispatch_log(dp, "2099-01-01") - end - end - - # =========================================================================== - # read_all_dispatch_logs/1 - # =========================================================================== - - describe "read_all_dispatch_logs/1" do - test "reads all dispatch logs", %{data_path: dp, dispatch_dir: dd} do - write_jsonl(Path.join(dd, "dispatch-2026-02-12.jsonl"), sample_dispatch_log()) - write_jsonl(Path.join(dd, "dispatch-2026-02-13.jsonl"), [ - %{"repo" => "verisimdb", "pattern" => "PA001", "strategy" => "auto_execute"} - ]) - - assert {:ok, records} = DispatchBridge.read_all_dispatch_logs(dp) - assert length(records) == 4 - end - - test "ignores non-dispatch files", %{data_path: dp, dispatch_dir: dd} do - write_jsonl(Path.join(dd, "dispatch-2026-02-12.jsonl"), sample_dispatch_log()) - File.write!(Path.join(dd, "pending.jsonl"), "") - - assert {:ok, records} = DispatchBridge.read_all_dispatch_logs(dp) - assert length(records) == 3 - end - end - - # =========================================================================== - # read_outcomes/1 - # =========================================================================== - - describe "read_outcomes/1" do - test "reads outcome files", %{data_path: dp, outcomes_dir: od} do - write_jsonl(Path.join(od, "2026-02.jsonl"), sample_outcomes()) - - assert {:ok, outcomes} = DispatchBridge.read_outcomes(dp) - assert length(outcomes) == 3 - end - end - - # =========================================================================== - # summarize/1 - # =========================================================================== - - describe "summarize/1" do - test "aggregates dispatch statistics", %{ - data_path: dp, - dispatch_dir: dd, - outcomes_dir: od - } do - write_jsonl(Path.join(dd, "pending.jsonl"), sample_pending()) - write_jsonl(Path.join(dd, "dispatch-2026-02-12.jsonl"), sample_dispatch_log()) - write_jsonl(Path.join(od, "2026-02.jsonl"), sample_outcomes()) - - summary = DispatchBridge.summarize(dp) - - assert summary.pending_count == 3 - assert summary.dispatched_count == 3 - assert summary.outcome_count == 3 - assert summary.outcome_success_rate == 66.7 - - assert summary.by_strategy["auto_execute"] == 1 - assert summary.by_strategy["review"] == 1 - assert summary.by_strategy["report_only"] == 1 - - assert summary.repos_with_pending == 3 - end - - test "handles empty data gracefully", %{data_path: dp} do - summary = DispatchBridge.summarize(dp) - - assert summary.pending_count == 0 - assert summary.dispatched_count == 0 - assert summary.outcome_count == 0 - end - end - - # =========================================================================== - # feedback_to_drift/1 - # =========================================================================== - - describe "feedback_to_drift/1" do - test "computes drift direction from outcomes", %{data_path: dp, outcomes_dir: od} do - outcomes = [ - %{"repo" => "echidna", "status" => "success"}, - %{"repo" => "echidna", "status" => "success"}, - %{"repo" => "proven", "status" => "success"}, - %{"repo" => "proven", "status" => "failure"}, - %{"repo" => "ambientops", "status" => "failure"}, - %{"repo" => "ambientops", "status" => "failure"} - ] - - write_jsonl(Path.join(od, "outcomes.jsonl"), outcomes) - - drift = DispatchBridge.feedback_to_drift(dp) - - # echidna: 2/2 success → improving - echidna = Enum.find(drift, fn {repo, _, _} -> repo == "echidna" end) - assert {_, :improving, %{successful: 2, total: 2}} = echidna - - # proven: 1/2 success → stable - proven = Enum.find(drift, fn {repo, _, _} -> repo == "proven" end) - assert {_, :stable, %{successful: 1, total: 2}} = proven - - # ambientops: 0/2 success → regressing - ambientops = Enum.find(drift, fn {repo, _, _} -> repo == "ambientops" end) - assert {_, :regressing, %{successful: 0, total: 2}} = ambientops - end - - test "returns empty when no outcomes", %{data_path: dp} do - drift = DispatchBridge.feedback_to_drift(dp) - assert drift == [] - end - end -end diff --git a/verisimdb/elixir-orchestration/test/verisim/hypatia/pattern_query_test.exs b/verisimdb/elixir-orchestration/test/verisim/hypatia/pattern_query_test.exs deleted file mode 100644 index 47323b53..00000000 --- a/verisimdb/elixir-orchestration/test/verisim/hypatia/pattern_query_test.exs +++ /dev/null @@ -1,159 +0,0 @@ -# SPDX-License-Identifier: MPL-2.0 - -defmodule VeriSim.Hypatia.PatternQueryTest do - @moduledoc """ - Tests for the Hypatia cross-repo pattern analytics module. - - Verifies pipeline health, cross-repo patterns, severity distributions, - and temporal trend queries against ingested scan data. - """ - - use ExUnit.Case, async: false - - alias VeriSim.Hypatia.ScanIngester - alias VeriSim.Hypatia.PatternQuery - - setup do - # Clean ETS between tests - case :ets.info(:hypatia_scans) do - :undefined -> :ok - _ -> :ets.delete_all_objects(:hypatia_scans) - end - - # Ingest some test scans - for {repo, lang, wps} <- test_data() do - scan = %{ - "assail_report" => %{ - "program_path" => "/repos/#{repo}", - "language" => lang, - "frameworks" => [], - "weak_points" => wps - } - } - - ScanIngester.ingest_scan(scan) - end - - :ok - end - - defp test_data do - [ - {"repo-alpha", "rust", [ - %{"category" => "PanicPath", "location" => "src/main.rs", "severity" => "Medium", "description" => "unwrap"}, - %{"category" => "UnsafeCode", "location" => "src/ffi.rs", "severity" => "High", "description" => "unsafe block"}, - %{"category" => "PanicPath", "location" => "src/lib.rs", "severity" => "Medium", "description" => "expect"} - ]}, - {"repo-beta", "elixir", [ - %{"category" => "PanicPath", "location" => "lib/worker.ex", "severity" => "Medium", "description" => "raise"}, - %{"category" => "InputValidation", "location" => "lib/api.ex", "severity" => "High", "description" => "unsanitized"} - ]}, - {"repo-gamma", "rust", [ - %{"category" => "PanicPath", "location" => "src/core.rs", "severity" => "Medium", "description" => "unwrap"}, - %{"category" => "PanicPath", "location" => "src/net.rs", "severity" => "High", "description" => "index panic"}, - %{"category" => "UnsafeCode", "location" => "src/sys.rs", "severity" => "High", "description" => "transmute"} - ]}, - {"repo-delta", "javascript", [ - %{"category" => "InputValidation", "location" => "src/handler.js", "severity" => "High", "description" => "XSS"}, - %{"category" => "PanicPath", "location" => "src/index.js", "severity" => "Low", "description" => "throw"} - ]} - ] - end - - # =========================================================================== - # pipeline_health/0 - # =========================================================================== - - describe "pipeline_health/0" do - test "returns correct total counts" do - health = PatternQuery.pipeline_health() - - assert health.total_scans == 4 - assert health.repos_scanned == 4 - assert health.total_weak_points == 0 # Weak points are in document body, not directly accessible - end - end - - # =========================================================================== - # cross_repo_patterns/1 - # =========================================================================== - - describe "cross_repo_patterns/1" do - test "returns empty when no shared patterns meet threshold" do - # Our test data has PanicPath:Medium in alpha, beta, gamma, delta - # but pattern extraction depends on document body parsing - patterns = PatternQuery.cross_repo_patterns(100) - assert patterns == [] - end - end - - # =========================================================================== - # severity_distribution/0 - # =========================================================================== - - describe "severity_distribution/0" do - test "returns a map" do - dist = PatternQuery.severity_distribution() - assert is_map(dist) - end - end - - # =========================================================================== - # category_distribution/0 - # =========================================================================== - - describe "category_distribution/0" do - test "returns a sorted list" do - dist = PatternQuery.category_distribution() - assert is_list(dist) - end - end - - # =========================================================================== - # temporal_trends/1 - # =========================================================================== - - describe "temporal_trends/1" do - test "returns scan history for a known repo" do - trends = PatternQuery.temporal_trends("repo-alpha") - - assert length(trends) == 1 - assert hd(trends).weak_point_count == 3 - end - - test "returns empty for unknown repo" do - trends = PatternQuery.temporal_trends("nonexistent") - assert trends == [] - end - end - - # =========================================================================== - # repos_by_severity/1 - # =========================================================================== - - describe "repos_by_severity/1" do - test "ranks repos by High severity" do - repos = PatternQuery.repos_by_severity("High") - - # repo-alpha has 1 High, repo-beta has 1, repo-gamma has 2, repo-delta has 1 - assert is_list(repos) - - if length(repos) > 0 do - {top_repo, top_count} = hd(repos) - assert is_binary(top_repo) - assert is_integer(top_count) - end - end - end - - # =========================================================================== - # weakness_hotspots/0 - # =========================================================================== - - describe "weakness_hotspots/0" do - test "returns a list of hotspots" do - hotspots = PatternQuery.weakness_hotspots() - assert is_list(hotspots) - end - end -end diff --git a/verisimdb/elixir-orchestration/test/verisim/hypatia/scan_ingester_test.exs b/verisimdb/elixir-orchestration/test/verisim/hypatia/scan_ingester_test.exs deleted file mode 100644 index 6e697f4a..00000000 --- a/verisimdb/elixir-orchestration/test/verisim/hypatia/scan_ingester_test.exs +++ /dev/null @@ -1,258 +0,0 @@ -# SPDX-License-Identifier: MPL-2.0 - -defmodule VeriSim.Hypatia.ScanIngesterTest do - @moduledoc """ - Tests for the Hypatia scan ingestion module. - - Verifies that the ingester correctly: - 1. Parses panic-attack JSON scan results - 2. Builds octad octad entities with all modalities - 3. Handles various input formats and edge cases - 4. Ingests from files and directories - 5. Stores data locally when Rust core is unavailable - """ - - use ExUnit.Case, async: false - - alias VeriSim.Hypatia.ScanIngester - - setup do - # Clean up ETS table between tests - case :ets.info(:hypatia_scans) do - :undefined -> :ok - _ -> :ets.delete_all_objects(:hypatia_scans) - end - - :ok - end - - # =========================================================================== - # Sample Data - # =========================================================================== - - defp sample_scan do - %{ - "assail_report" => %{ - "program_path" => "/var$REPOS_DIR/protocol-squisher", - "language" => "rust", - "frameworks" => ["WebServer"], - "weak_points" => [ - %{ - "category" => "PanicPath", - "location" => "src/main.rs", - "severity" => "Medium", - "description" => "11 unwrap/expect calls in src/main.rs", - "recommended_attack" => ["memory", "disk"] - }, - %{ - "category" => "UnsafeCode", - "location" => "src/ffi.rs", - "severity" => "High", - "description" => "unsafe block without SAFETY comment", - "recommended_attack" => ["memory"] - } - ] - } - } - end - - defp sample_scan_flat do - %{ - "program_path" => "/repos/my-app", - "language" => "elixir", - "frameworks" => ["Phoenix"], - "weak_points" => [ - %{ - "category" => "HotCodeReload", - "location" => "lib/my_app/worker.ex", - "severity" => "Low", - "description" => "Hot code reload in production" - } - ] - } - end - - # =========================================================================== - # ingest_scan/1 - # =========================================================================== - - describe "ingest_scan/1" do - test "ingests standard panic-attack scan result" do - assert {:ok, octad_id} = ScanIngester.ingest_scan(sample_scan()) - assert String.starts_with?(octad_id, "scan:protocol-squisher:") - end - - test "ingests flat scan format (without assail_report wrapper)" do - assert {:ok, octad_id} = ScanIngester.ingest_scan(sample_scan_flat()) - assert String.starts_with?(octad_id, "scan:my-app:") - end - - test "builds octad with correct metadata" do - {:ok, octad_id} = ScanIngester.ingest_scan(sample_scan()) - - scans = ScanIngester.list_scans() - scan = Enum.find(scans, &(&1.octad_id == octad_id)) - - assert scan.metadata.repo_name == "protocol-squisher" - assert scan.metadata.language == "rust" - assert scan.metadata.frameworks == ["WebServer"] - assert scan.metadata.weak_point_count == 2 - assert scan.metadata.severity_counts == %{"Medium" => 1, "High" => 1} - end - - test "builds document modality with searchable text" do - {:ok, octad_id} = ScanIngester.ingest_scan(sample_scan()) - - scans = ScanIngester.list_scans() - scan = Enum.find(scans, &(&1.octad_id == octad_id)) - - assert scan.document.title =~ "protocol-squisher" - assert scan.document.body =~ "PanicPath" - assert scan.document.body =~ "UnsafeCode" - assert scan.document.body =~ "src/main.rs" - end - - test "builds graph triples" do - {:ok, octad_id} = ScanIngester.ingest_scan(sample_scan()) - - scans = ScanIngester.list_scans() - scan = Enum.find(scans, &(&1.octad_id == octad_id)) - - triples = scan.graph.triples - assert length(triples) > 0 - - # Should have repo → has_scan triple - assert Enum.any?(triples, fn [s, p, _o] -> - s == "repo:protocol-squisher" and p == "has_scan" - end) - - # Should have weakness → in_file triples - assert Enum.any?(triples, fn [_s, p, _o] -> p == "in_file" end) - end - - test "builds semantic modality with categories" do - {:ok, octad_id} = ScanIngester.ingest_scan(sample_scan()) - - scans = ScanIngester.list_scans() - scan = Enum.find(scans, &(&1.octad_id == octad_id)) - - assert "PanicPath" in scan.semantic.tags - assert "UnsafeCode" in scan.semantic.tags - assert "scan_result" in scan.semantic.types - end - - test "builds provenance modality" do - {:ok, octad_id} = ScanIngester.ingest_scan(sample_scan()) - - scans = ScanIngester.list_scans() - scan = Enum.find(scans, &(&1.octad_id == octad_id)) - - assert scan.provenance.source == "panic-attack" - assert scan.provenance.operation == "assail" - end - - test "rejects non-map input" do - assert {:error, :invalid_scan_format} = ScanIngester.ingest_scan("not a map") - assert {:error, :invalid_scan_format} = ScanIngester.ingest_scan(42) - end - end - - # =========================================================================== - # ingest_file/1 - # =========================================================================== - - describe "ingest_file/1" do - setup do - dir = Path.join(System.tmp_dir!(), "hypatia_test_#{System.unique_integer([:positive])}") - File.mkdir_p!(dir) - on_exit(fn -> File.rm_rf!(dir) end) - {:ok, dir: dir} - end - - test "ingests from a JSON file", %{dir: dir} do - path = Path.join(dir, "test-scan.json") - File.write!(path, Jason.encode!(sample_scan())) - - assert {:ok, octad_id} = ScanIngester.ingest_file(path) - assert String.starts_with?(octad_id, "scan:") - end - - test "returns error for non-existent file" do - assert {:error, {:file_read_error, _}} = ScanIngester.ingest_file("/nonexistent/path.json") - end - - test "returns error for invalid JSON", %{dir: dir} do - path = Path.join(dir, "bad.json") - File.write!(path, "not valid json {{{") - - assert {:error, {:json_parse_error, _}} = ScanIngester.ingest_file(path) - end - end - - # =========================================================================== - # ingest_directory/1 - # =========================================================================== - - describe "ingest_directory/1" do - setup do - dir = Path.join(System.tmp_dir!(), "hypatia_dir_#{System.unique_integer([:positive])}") - File.mkdir_p!(dir) - on_exit(fn -> File.rm_rf!(dir) end) - {:ok, dir: dir} - end - - test "ingests all JSON files from directory", %{dir: dir} do - for i <- 1..3 do - scan = put_in(sample_scan(), ["assail_report", "program_path"], "/repos/repo-#{i}") - File.write!(Path.join(dir, "repo-#{i}.json"), Jason.encode!(scan)) - end - - assert {:ok, results} = ScanIngester.ingest_directory(dir) - assert length(results) == 3 - assert Enum.all?(results, fn {_f, result} -> match?({:ok, _}, result) end) - end - - test "skips non-JSON files", %{dir: dir} do - File.write!(Path.join(dir, "readme.txt"), "ignore me") - File.write!(Path.join(dir, "scan.json"), Jason.encode!(sample_scan())) - - assert {:ok, results} = ScanIngester.ingest_directory(dir) - assert length(results) == 1 - end - - test "returns error for non-existent directory" do - assert {:error, {:dir_read_error, _}} = ScanIngester.ingest_directory("/nonexistent/dir") - end - end - - # =========================================================================== - # list_scans/0 and get_scan/1 - # =========================================================================== - - describe "list_scans/0" do - test "returns empty list when no scans ingested" do - assert ScanIngester.list_scans() == [] - end - - test "returns all ingested scans" do - ScanIngester.ingest_scan(sample_scan()) - ScanIngester.ingest_scan(sample_scan_flat()) - - scans = ScanIngester.list_scans() - assert length(scans) == 2 - end - end - - describe "get_scan/1" do - test "finds scan by repo name" do - ScanIngester.ingest_scan(sample_scan()) - - assert {:ok, scan} = ScanIngester.get_scan("protocol-squisher") - assert scan.metadata.language == "rust" - end - - test "returns error for unknown repo" do - assert {:error, :not_found} = ScanIngester.get_scan("nonexistent-repo") - end - end -end diff --git a/verisimdb/elixir-orchestration/test/verisim/query/vcl_crossmodal_test.exs b/verisimdb/elixir-orchestration/test/verisim/query/vcl_crossmodal_test.exs deleted file mode 100644 index 4a23dac7..00000000 --- a/verisimdb/elixir-orchestration/test/verisim/query/vcl_crossmodal_test.exs +++ /dev/null @@ -1,280 +0,0 @@ -# SPDX-License-Identifier: MPL-2.0 - -defmodule VeriSim.Query.VCLCrossModalTest do - @moduledoc """ - Cross-modal condition tests for VCL queries. - - Exercises the condition classification and cross-modal evaluation logic in - `VCLExecutor`. Cross-modal conditions are NOT pushed down to individual - modality stores — they are evaluated post-fetch by comparing data across - two or more modalities on the same octad. - - ## Cross-modal condition types tested - - 1. **CrossModalFieldCompare** — Compare fields across modalities - `WHERE DOCUMENT.severity > GRAPH.importance` - - 2. **ModalityDrift** — Detect drift between modality representations - `WHERE DRIFT(VECTOR, DOCUMENT) > 0.3` - - 3. **ModalityExists / ModalityNotExists** — Check modality population - `WHERE SPATIAL EXISTS AND TENSOR NOT EXISTS` - - 4. **ModalityConsistency (cosine)** — Consistency via cosine similarity - `WHERE CONSISTENT(VECTOR, SEMANTIC) USING COSINE > 0.8` - - 5. **ModalityConsistency (jaccard)** — Consistency via Jaccard index - `WHERE CONSISTENT(GRAPH, DOCUMENT) USING JACCARD > 0.5` - - These tests verify condition classification (pushdown vs cross-modal), - AST structure, and evaluation logic — they do NOT require the Rust core - to be running. - """ - - use ExUnit.Case, async: false - - alias VeriSim.Query.VCLBridge - alias VeriSim.Test.VCLTestHelpers, as: H - - setup_all do - pid = H.ensure_bridge_started() - %{bridge_pid: pid} - end - - # =========================================================================== - # 1. CrossModalFieldCompare - # =========================================================================== - - describe "CrossModalFieldCompare" do - test "WHERE clause with cross-modal field comparison parses correctly" do - # The built-in parser stores raw WHERE text; cross-modal classification - # happens during execution. Verify the raw clause is preserved. - query = "SELECT * FROM HEXAD 'entity-001' WHERE DOCUMENT.severity > GRAPH.importance" - ast = H.parse!(query) - - H.assert_has_where(ast) - raw = ast[:where][:raw] - assert raw =~ "DOCUMENT.severity" - assert raw =~ "GRAPH.importance" - end - - test "CrossModalFieldCompare AST node evaluates correctly on mock octad" do - # Build a cross-modal condition AST node and verify its structure - condition = H.cross_modal_compare("document", "severity", ">", "graph", "importance") - - # The evaluate_cross_modal function is private, so we test via the - # classifier + filter pipeline by constructing a full query execution. - # Here, we verify the AST structure is well-formed. - assert condition[:TAG] == "CrossModalFieldCompare" - assert condition[:_0] == "document" - assert condition[:_2] == ">" - assert condition[:_3] == "graph" - end - - test "combined cross-modal And condition builds correct AST" do - cond1 = H.cross_modal_compare("document", "severity", ">", "graph", "importance") - cond2 = H.cross_modal_compare("vector", "dimension", "==", "tensor", "rank") - combined = H.and_condition(cond1, cond2) - - assert combined[:TAG] == "And" - assert combined[:_0][:TAG] == "CrossModalFieldCompare" - assert combined[:_1][:TAG] == "CrossModalFieldCompare" - end - end - - # =========================================================================== - # 2. ModalityDrift - # =========================================================================== - - describe "ModalityDrift" do - test "DRIFT condition in WHERE clause parses correctly" do - query = "SELECT * FROM HEXAD 'entity-001' WHERE DRIFT(VECTOR, DOCUMENT) > 0.3" - ast = H.parse!(query) - - H.assert_has_where(ast) - assert ast[:where][:raw] =~ "DRIFT" - assert ast[:where][:raw] =~ "VECTOR" - assert ast[:where][:raw] =~ "DOCUMENT" - end - - test "ModalityDrift AST node has correct structure" do - drift = H.modality_drift("vector", "document", 0.3) - - assert drift[:TAG] == "ModalityDrift" - assert drift[:_0] == "vector" - assert drift[:_1] == "document" - assert drift[:_2] == 0.3 - end - - test "drift query executes without crashing" do - query = "SELECT * FROM HEXAD 'entity-001' WHERE DRIFT(VECTOR, DOCUMENT) > 0.3" - result = H.execute_safely(query) - assert elem(result, 0) in [:ok, :error, :unavailable] - end - end - - # =========================================================================== - # 3. ModalityExists / ModalityNotExists - # =========================================================================== - - describe "ModalityExists / ModalityNotExists" do - test "EXISTS condition parses correctly" do - query = "SELECT * FROM HEXAD 'entity-001' WHERE PROVENANCE EXISTS" - ast = H.parse!(query) - - H.assert_has_where(ast) - assert ast[:where][:raw] =~ "PROVENANCE" - assert ast[:where][:raw] =~ "EXISTS" - end - - test "combined EXISTS and NOT EXISTS parses correctly" do - query = "SELECT * FROM HEXAD 'entity-001' WHERE SPATIAL EXISTS AND TENSOR NOT EXISTS" - ast = H.parse!(query) - - H.assert_has_where(ast) - raw = ast[:where][:raw] - assert raw =~ "SPATIAL" - assert raw =~ "TENSOR" - assert raw =~ "NOT EXISTS" - end - - test "ModalityExists AST node has correct structure" do - exists = H.modality_exists("provenance") - not_exists = H.modality_not_exists("tensor") - - assert exists[:TAG] == "ModalityExists" - assert exists[:_0] == "provenance" - assert not_exists[:TAG] == "ModalityNotExists" - assert not_exists[:_0] == "tensor" - end - - test "exists query executes without crashing" do - query = "SELECT * FROM HEXAD 'entity-001' WHERE SPATIAL EXISTS AND TENSOR NOT EXISTS" - result = H.execute_safely(query) - assert elem(result, 0) in [:ok, :error, :unavailable] - end - end - - # =========================================================================== - # 4. ModalityConsistency (cosine) - # =========================================================================== - - describe "ModalityConsistency with cosine metric" do - test "CONSISTENT condition with COSINE metric parses correctly" do - query = "SELECT * FROM HEXAD 'entity-001' WHERE CONSISTENT(VECTOR, SEMANTIC) USING COSINE > 0.8" - ast = H.parse!(query) - - H.assert_has_where(ast) - raw = ast[:where][:raw] - assert raw =~ "CONSISTENT" - assert raw =~ "COSINE" - end - - test "ModalityConsistency AST node (cosine) has correct structure" do - consistency = H.modality_consistency("vector", "semantic", "COSINE") - - assert consistency[:TAG] == "ModalityConsistency" - assert consistency[:_0] == "vector" - assert consistency[:_1] == "semantic" - assert consistency[:_2] == "COSINE" - end - - test "cosine consistency query executes without crashing" do - query = "SELECT * FROM HEXAD 'entity-001' WHERE CONSISTENT(VECTOR, SEMANTIC) USING COSINE > 0.8" - result = H.execute_safely(query) - assert elem(result, 0) in [:ok, :error, :unavailable] - end - end - - # =========================================================================== - # 5. ModalityConsistency (jaccard) - # =========================================================================== - - describe "ModalityConsistency with jaccard metric" do - test "CONSISTENT condition with JACCARD metric parses correctly" do - query = "SELECT * FROM HEXAD 'entity-001' WHERE CONSISTENT(GRAPH, DOCUMENT) USING JACCARD > 0.5" - ast = H.parse!(query) - - H.assert_has_where(ast) - raw = ast[:where][:raw] - assert raw =~ "CONSISTENT" - assert raw =~ "JACCARD" - end - - test "ModalityConsistency AST node (jaccard) has correct structure" do - consistency = H.modality_consistency("graph", "document", "JACCARD") - - assert consistency[:TAG] == "ModalityConsistency" - assert consistency[:_0] == "graph" - assert consistency[:_1] == "document" - assert consistency[:_2] == "JACCARD" - end - - test "jaccard consistency query executes without crashing" do - query = "SELECT * FROM HEXAD 'entity-001' WHERE CONSISTENT(GRAPH, DOCUMENT) USING JACCARD > 0.5" - result = H.execute_safely(query) - assert elem(result, 0) in [:ok, :error, :unavailable] - end - end - - # =========================================================================== - # 6. Condition classification (unit-level) - # =========================================================================== - - describe "condition classification" do - test "raw WHERE clause with simple field condition is classified as pushdown" do - # Simple conditions (single-modality) should be pushed down to the store - query = "SELECT DOCUMENT.* FROM HEXAD 'entity-001' WHERE DOCUMENT.title = 'Test'" - ast = H.parse!(query) - - # The raw WHERE clause should be preserved for pushdown - assert ast[:where][:raw] =~ "DOCUMENT.title" - end - - test "multiple cross-modal conditions combine via AND" do - cond1 = H.modality_exists("vector") - cond2 = H.modality_drift("vector", "document", 0.5) - combined = H.and_condition(cond1, cond2) - - assert combined[:TAG] == "And" - assert combined[:_0][:TAG] == "ModalityExists" - assert combined[:_1][:TAG] == "ModalityDrift" - end - - test "OR-combined cross-modal conditions" do - cond1 = H.modality_exists("spatial") - cond2 = H.modality_exists("provenance") - combined = H.or_condition(cond1, cond2) - - assert combined[:TAG] == "Or" - assert combined[:_0][:_0] == "spatial" - assert combined[:_1][:_0] == "provenance" - end - end - - # =========================================================================== - # 7. Complex compound conditions - # =========================================================================== - - describe "compound cross-modal conditions" do - test "drift AND consistency in same query parses correctly" do - query = "SELECT * FROM HEXAD 'entity-001' WHERE DRIFT(VECTOR, DOCUMENT) > 0.3 AND CONSISTENT(GRAPH, SEMANTIC) USING COSINE > 0.8" - ast = H.parse!(query) - - H.assert_has_where(ast) - raw = ast[:where][:raw] - assert raw =~ "DRIFT" - assert raw =~ "CONSISTENT" - end - - test "existence check with field comparison parses correctly" do - query = "SELECT * FROM HEXAD 'entity-001' WHERE PROVENANCE EXISTS AND DOCUMENT.severity > 5" - ast = H.parse!(query) - - H.assert_has_where(ast) - raw = ast[:where][:raw] - assert raw =~ "PROVENANCE EXISTS" - assert raw =~ "DOCUMENT.severity" - end - end -end diff --git a/verisimdb/elixir-orchestration/test/verisim/query/vcl_dt_integration_test.exs b/verisimdb/elixir-orchestration/test/verisim/query/vcl_dt_integration_test.exs deleted file mode 100644 index 6302f165..00000000 --- a/verisimdb/elixir-orchestration/test/verisim/query/vcl_dt_integration_test.exs +++ /dev/null @@ -1,422 +0,0 @@ -# SPDX-License-Identifier: MPL-2.0 - -defmodule VeriSim.Query.VCLDTIntegrationTest do - @moduledoc """ - VCL-UT (dependent type) integration tests. - - Exercises the full VCL-UT pipeline: parse → type-check → execute → verify - proofs → bundle ProvedResult. Tests verify both the happy path (proofs pass - when data exists) and the error paths (proofs fail correctly). - - ## Test categories - - 1. **Individual proof type execution** — each of 6 proof types routes to - the correct Rust endpoint and returns a structured artifact or error - 2. **Multi-proof composition** — multiple proofs in a single query - 3. **Invalid proof rejection** — tampered or malformed data fails verification - 4. **Proof downgrade rejection** — VCL-UT cannot silently become slipstream - 5. **ProvedResult structure** — data + proof_certificate with all required fields - 6. **Fallback type extraction** — when type checker unavailable, obligations - are extracted from the AST directly with correct structure - - Tests gracefully handle Rust core unavailability: when the Rust core is - not running, proof verification errors are EXPECTED (never silently passed). - """ - - use ExUnit.Case, async: false - - alias VeriSim.Query.{VCLBridge, VCLExecutor} - alias VeriSim.Test.VCLTestHelpers, as: H - - setup_all do - pid = H.ensure_bridge_started() - %{bridge_pid: pid} - end - - # =========================================================================== - # 1. Individual proof type execution - # =========================================================================== - - describe "existence proof" do - test "parses PROOF EXISTENCE clause and routes to verification" do - query = "SELECT GRAPH.* FROM HEXAD 'entity-001' PROOF EXISTENCE(entity-001)" - ast = H.parse!(query) - - H.assert_has_proof(ast) - - # Execute: should attempt proof verification and return error or proved result - result = H.execute_safely(query) - assert_proof_result_or_error(result) - end - end - - describe "provenance proof" do - test "parses PROOF PROVENANCE clause and routes to provenance verifier" do - query = "SELECT PROVENANCE.* FROM HEXAD 'entity-001' PROOF PROVENANCE(entity-001)" - ast = H.parse!(query) - - H.assert_has_proof(ast) - - result = H.execute_safely(query) - assert_proof_result_or_error(result) - end - end - - describe "integrity proof" do - test "parses PROOF INTEGRITY clause and routes to proof generator" do - query = "SELECT GRAPH.* FROM HEXAD 'entity-001' PROOF INTEGRITY(my_contract)" - ast = H.parse!(query) - - H.assert_has_proof(ast) - - result = H.execute_safely(query) - assert_proof_result_or_error(result) - end - end - - describe "access proof" do - test "parses PROOF ACCESS clause and routes to auth endpoint" do - query = "SELECT GRAPH.* FROM HEXAD 'entity-001' PROOF ACCESS(entity-001)" - ast = H.parse!(query) - - H.assert_has_proof(ast) - - result = H.execute_safely(query) - assert_proof_result_or_error(result) - end - end - - describe "citation proof" do - test "parses PROOF CITATION clause and validates contract existence" do - query = "SELECT GRAPH.* FROM HEXAD 'entity-001' PROOF CITATION(my_citation)" - ast = H.parse!(query) - - H.assert_has_proof(ast) - - result = H.execute_safely(query) - assert_proof_result_or_error(result) - end - end - - describe "zkp proof" do - test "parses PROOF ZKP clause and routes to zkp_bridge" do - query = "SELECT GRAPH.* FROM HEXAD 'entity-001' PROOF ZKP(claim_123)" - ast = H.parse!(query) - - H.assert_has_proof(ast) - - result = H.execute_safely(query) - assert_proof_result_or_error(result) - end - end - - # =========================================================================== - # 2. Multi-proof composition - # =========================================================================== - - describe "multi-proof composition" do - test "two proofs combined with AND parse correctly" do - query = "SELECT * FROM HEXAD 'entity-001' PROOF EXISTENCE(entity-001) AND PROVENANCE(entity-001)" - ast = H.parse!(query) - - H.assert_has_proof(ast) - - # The proof field should contain multiple proof specs - proof = ast[:proof] - assert proof != nil - end - - test "multi-proof execution routes each proof independently" do - query = "SELECT * FROM HEXAD 'entity-001' PROOF EXISTENCE(entity-001) AND PROVENANCE(entity-001)" - - result = H.execute_safely(query) - assert_proof_result_or_error(result) - end - end - - # =========================================================================== - # 3. Invalid proof rejection - # =========================================================================== - - describe "invalid proof rejection" do - test "unknown proof type is rejected" do - query = "SELECT * FROM HEXAD 'entity-001' PROOF BOGUS_TYPE(entity)" - - result = H.execute_safely(query) - - case result do - {:error, _} -> assert true # Expected: unknown proof type fails - {:unavailable, _} -> assert true # Connection error - {:ok, _} -> flunk("Unknown proof type should not produce a valid result") - end - end - - test "proof without required contract name produces specific error" do - # INTEGRITY requires a contract name — passing just whitespace should fail - ast = %{ - modalities: [:graph], - source: {:octad, "entity-001"}, - where: nil, - proof: [%{proofType: "INTEGRITY"}], # No contractName - limit: nil, - offset: nil, - orderBy: nil, - groupBy: nil, - having: nil, - aggregates: nil, - projections: nil - } - - result = H.execute_ast_safely(ast) - - case result do - {:error, {:proof_verification_failed, {:missing_contract, _}}} -> assert true - {:error, _} -> assert true # Any error is acceptable - {:unavailable, _} -> assert true - {:ok, _} -> flunk("Integrity proof without contract should fail") - end - end - end - - # =========================================================================== - # 4. Proof downgrade rejection - # =========================================================================== - - describe "proof downgrade rejection" do - test "VCL-UT query with PROOF clause cannot silently become slipstream" do - # A query with PROOF must go through the VCL-UT path and either - # return a ProvedResult or fail with a proof-related error. - query = "SELECT GRAPH.* FROM HEXAD 'entity-001' PROOF EXISTENCE(entity-001)" - - result = H.execute_safely(query) - - case result do - {:ok, proved_result} when is_map(proved_result) -> - # Must have a proof_certificate — NOT bare data - assert Map.has_key?(proved_result, :proof_certificate) or - Map.has_key?(proved_result, "proof_certificate"), - "VCL-UT result must include proof_certificate, got: #{inspect(Map.keys(proved_result))}" - - {:error, _} -> - # Proof verification error is acceptable (Rust core unavailable) - assert true - - {:unavailable, _} -> - assert true - end - end - - test "parse_dependent rejects queries without PROOF clause" do - result = VCLBridge.parse_dependent("SELECT GRAPH.* FROM HEXAD 'abc-123'") - - case result do - {:error, msg} -> - assert msg =~ "PROOF" or msg =~ "proof" or msg =~ "dependent" - - {:ok, _} -> - # If the parser somehow returns ok without PROOF, that's a bug - # (but the built-in parser might not enforce this in all paths) - :ok - end - end - - test "parse_slipstream rejects queries with PROOF clause" do - result = VCLBridge.parse_slipstream( - "SELECT GRAPH.* FROM HEXAD 'abc-123' PROOF EXISTENCE(abc-123)" - ) - - case result do - {:error, msg} -> - assert msg =~ "PROOF" or msg =~ "proof" or msg =~ "Slipstream" - - {:ok, _} -> - :ok - end - end - end - - # =========================================================================== - # 5. ProvedResult structure - # =========================================================================== - - describe "ProvedResult structure" do - test "VCL-UT result has expected shape with proof_certificate" do - # Build a well-formed AST that will go through the VCL-UT path - ast = %{ - modalities: [:graph], - source: {:octad, "entity-001"}, - where: nil, - proof: [%{proofType: "EXISTENCE", contractName: "entity-001"}], - limit: nil, - offset: nil, - orderBy: nil, - groupBy: nil, - having: nil, - aggregates: nil, - projections: nil - } - - result = H.execute_ast_safely(ast) - - case result do - {:ok, proved_result} when is_map(proved_result) -> - # Verify the ProvedResult structure - assert Map.has_key?(proved_result, :data) or Map.has_key?(proved_result, "data"), - "ProvedResult must have :data field" - assert Map.has_key?(proved_result, :proof_certificate) or - Map.has_key?(proved_result, "proof_certificate"), - "ProvedResult must have :proof_certificate field" - - cert = proved_result[:proof_certificate] || proved_result["proof_certificate"] - if cert do - assert Map.has_key?(cert, :proofs) or Map.has_key?(cert, "proofs"), - "Certificate must have :proofs field" - assert Map.has_key?(cert, :composition) or Map.has_key?(cert, "composition"), - "Certificate must have :composition field" - assert Map.has_key?(cert, :verified_at) or Map.has_key?(cert, "verified_at"), - "Certificate must have :verified_at field" - assert Map.has_key?(cert, :query_hash) or Map.has_key?(cert, "query_hash"), - "Certificate must have :query_hash field" - end - - {:error, _} -> - # Proof verification error (Rust unavailable) — expected in test env - assert true - - {:unavailable, _} -> - assert true - end - end - - test "proof_certificate includes obligations list" do - ast = %{ - modalities: [:graph], - source: {:octad, "entity-001"}, - where: nil, - proof: [%{proofType: "EXISTENCE", contractName: "entity-001"}], - limit: nil, - offset: nil, - orderBy: nil, - groupBy: nil, - having: nil, - aggregates: nil, - projections: nil - } - - result = H.execute_ast_safely(ast) - - case result do - {:ok, %{proof_certificate: cert}} -> - assert Map.has_key?(cert, :obligations), - "Certificate should include :obligations list" - - _ -> - # Rust unavailable — skip structure check - :ok - end - end - end - - # =========================================================================== - # 6. Fallback type extraction - # =========================================================================== - - describe "fallback type extraction" do - test "type checker unavailability falls back to AST-based obligation extraction" do - # When the type checker is unavailable, the executor should extract - # proof obligations from the AST directly and still attempt verification. - query = "SELECT GRAPH.* FROM HEXAD 'entity-001' PROOF EXISTENCE(entity-001)" - - result = H.execute_safely(query) - - # The result should be either a proof error (Rust unavailable) or - # a valid ProvedResult. It must NOT be a type_check_failed error, - # because the fallback should handle type_checker_unavailable. - case result do - {:error, {:type_check_failed, _}} -> - flunk("Fallback type extraction should handle type_checker_unavailable") - - {:error, _} -> - # Any other error is acceptable (proof verification, connection, etc.) - assert true - - {:ok, _} -> - assert true - - {:unavailable, _} -> - assert true - end - end - - test "fallback extracts proof type and contract from raw proof spec" do - # The built-in parser produces %{raw: "EXISTENCE(entity-001)"}. - # The fallback should extract type=:existence, contract="entity-001". - query = "SELECT * FROM HEXAD 'entity-001' PROOF EXISTENCE(entity-001)" - ast = H.parse!(query) - - # Verify the proof field is set - assert ast[:proof] != nil - - # Execute — should not fail with "unknown proof type" - result = H.execute_safely(query) - - case result do - {:error, {:proof_verification_failed, {:unknown_proof_type, _}}} -> - flunk("Fallback should correctly extract proof type from raw spec") - - _ -> - # Any other result (success, connection error, proof error) is fine - assert true - end - end - end - - # =========================================================================== - # 7. Explain plan for VCL-UT queries - # =========================================================================== - - describe "VCL-UT explain plan" do - test "explain: true on a VCL-UT query returns plan without executing proofs" do - query = "SELECT GRAPH.* FROM HEXAD 'entity-001' PROOF EXISTENCE(entity-001)" - ast = H.parse!(query) - - # explain: true should return the execution plan, NOT execute proofs - {:ok, plan} = VCLExecutor.execute(ast, explain: true) - - assert is_map(plan) - assert Map.has_key?(plan, :strategy) - end - end - - # =========================================================================== - # Helpers - # =========================================================================== - - # Assert that a result is either a valid ProvedResult, a proof error, - # or a connection/unavailability error. Never a crash. - defp assert_proof_result_or_error(result) do - case result do - {:ok, proved_result} when is_map(proved_result) -> - # Valid ProvedResult — should have proof_certificate - assert Map.has_key?(proved_result, :proof_certificate) or - Map.has_key?(proved_result, "proof_certificate") or - # Or it might be a plain data result if the proof silently passed (bug!) - Map.has_key?(proved_result, :data) or - Map.has_key?(proved_result, "data") - - {:ok, _other} -> - # Some other ok result — acceptable - assert true - - {:error, _reason} -> - # Expected when Rust core is not running - assert true - - {:unavailable, _} -> - assert true - - other -> - flunk("Unexpected result: #{inspect(other)}") - end - end -end diff --git a/verisimdb/elixir-orchestration/test/verisim/query/vcl_dt_test.exs b/verisimdb/elixir-orchestration/test/verisim/query/vcl_dt_test.exs deleted file mode 100644 index acd6daed..00000000 --- a/verisimdb/elixir-orchestration/test/verisim/query/vcl_dt_test.exs +++ /dev/null @@ -1,263 +0,0 @@ -# SPDX-License-Identifier: MPL-2.0 - -defmodule VeriSim.Query.VCLDTTest do - @moduledoc """ - VCL-UT (dependent type) integration tests. - - Tests the proof verification pipeline: type checking → execution → proof - verification → ProvedResult bundling. - - Without the Rust core running, these tests verify that the Elixir-side wiring - is correct and that proof failures propagate as errors (NOT silent passes). - - ## Security invariants tested - - 1. Proofs MUST fail when the Rust core is unreachable — never silently pass. - 2. Each proof type returns a specific error on failure, not a generic :ok. - 3. Invalid or malformed proof data is rejected, not silently accepted. - 4. Non-list proof_specs are rejected (previously silently passed). - 5. Provenance proofs without an entity ID are rejected (previously silently passed). - """ - - use ExUnit.Case, async: false - - alias VeriSim.Query.{VCLBridge, VCLExecutor} - - setup_all do - case VCLBridge.start_link([]) do - {:ok, pid} -> %{bridge_pid: pid} - {:error, {:already_started, pid}} -> %{bridge_pid: pid} - end - end - - # =========================================================================== - # Type checker integration - # =========================================================================== - - describe "VCLBridge.typecheck/1" do - test "returns :type_checker_unavailable when no Deno subprocess" do - # Without the Deno subprocess, typecheck must return an explicit error - {:ok, ast} = VCLBridge.parse("SELECT GRAPH.* FROM HEXAD 'abc-123'") - result = VCLBridge.typecheck(ast) - assert result == {:error, :type_checker_unavailable} - end - end - - # =========================================================================== - # VCL-UT query execution (PROOF clause) - # =========================================================================== - - describe "VCL-UT execution path" do - test "parse_statement handles PROOF clause in AST" do - # The built-in parser should handle PROOF clauses - result = - VCLBridge.parse_statement( - "SELECT GRAPH.* FROM HEXAD 'abc-123' PROOF EXISTENCE(abc-123)" - ) - - case result do - {:ok, ast} -> - assert is_map(ast) - - {:error, _} -> - # Built-in parser may not support PROOF in statement mode — acceptable - :ok - end - end - - test "execute_string with PROOF returns error when Rust unavailable" do - # VCL-UT queries MUST fail if proofs cannot be verified — - # they should NOT silently return unproven data. - result = - try do - VCLExecutor.execute_string( - "SELECT GRAPH.* FROM HEXAD 'abc-123' PROOF EXISTENCE(abc-123)", - timeout: 1_000 - ) - rescue - _ -> {:error, :rust_core_unavailable} - end - - case result do - {:error, _} -> - # Expected: proof verification fails because Rust core is not running - assert true - - {:ok, proved_result} when is_map(proved_result) -> - # If somehow we get a result, it MUST have a proof certificate - assert Map.has_key?(proved_result, :proof_certificate) or - Map.has_key?(proved_result, "proof_certificate") - end - end - end - - # =========================================================================== - # Proof verification — individual proof types - # =========================================================================== - - describe "proof verification types" do - test "existence proof fails when Rust core unavailable" do - # Without Rust core, existence check should fail (not silently pass) - result = - try do - VCLExecutor.execute_string( - "SELECT GRAPH.* FROM HEXAD 'abc-123' PROOF EXISTENCE(abc-123)", - timeout: 1_000 - ) - rescue - _ -> {:error, :rust_core_unavailable} - end - - assert elem(result, 0) == :error - end - - test "provenance proof fails when Rust core unavailable" do - result = - try do - VCLExecutor.execute_string( - "SELECT PROVENANCE.* FROM HEXAD 'abc-123' PROOF PROVENANCE(abc-123)", - timeout: 1_000 - ) - rescue - _ -> {:error, :rust_core_unavailable} - end - - # Must be an error (Rust core not running) — NOT a silent pass - assert elem(result, 0) == :error - end - - test "integrity proof fails when Rust core unavailable" do - result = - try do - VCLExecutor.execute_string( - "SELECT GRAPH.* FROM HEXAD 'abc-123' PROOF INTEGRITY(my_contract)", - timeout: 1_000 - ) - rescue - _ -> {:error, :rust_core_unavailable} - end - - # Must be an error (Rust core not running) — NOT a silent pass - assert elem(result, 0) == :error - end - - test "access proof fails when Rust core unavailable and entity specified" do - result = - try do - VCLExecutor.execute_string( - "SELECT GRAPH.* FROM HEXAD 'abc-123' PROOF ACCESS(abc-123)", - timeout: 1_000 - ) - rescue - _ -> {:error, :rust_core_unavailable} - end - - assert elem(result, 0) == :error - end - - test "citation proof fails when Rust core unavailable" do - result = - try do - VCLExecutor.execute_string( - "SELECT GRAPH.* FROM HEXAD 'abc-123' PROOF CITATION(my_citation)", - timeout: 1_000 - ) - rescue - _ -> {:error, :rust_core_unavailable} - end - - assert elem(result, 0) == :error - end - - test "zkp proof fails when Rust core unavailable" do - result = - try do - VCLExecutor.execute_string( - "SELECT GRAPH.* FROM HEXAD 'abc-123' PROOF ZKP(claim_123)", - timeout: 1_000 - ) - rescue - _ -> {:error, :rust_core_unavailable} - end - - assert elem(result, 0) == :error - end - end - - # =========================================================================== - # Silent-pass regression tests (these bugs were fixed) - # =========================================================================== - - describe "silent-pass bug regressions" do - test "provenance proof without entity ID must fail, not silently pass" do - # REGRESSION: Previously, provenance proofs without a contract_name - # returned :ok instead of failing. This was a security violation — - # a provenance proof that cannot identify its target entity is meaningless. - result = - try do - VCLExecutor.execute_string( - "SELECT PROVENANCE.* FROM HEXAD 'abc-123' PROOF PROVENANCE(abc-123)", - timeout: 1_000 - ) - rescue - _ -> {:error, :rust_core_unavailable} - end - - # Whether the Rust core is up or down, this must NEVER silently pass. - # It should either fail (Rust unavailable) or succeed with a real proof. - assert elem(result, 0) == :error or - (elem(result, 0) == :ok and - is_map(elem(result, 1)) and - Map.has_key?(elem(result, 1), :proof_certificate)) - end - - test "unknown proof type must fail, not silently pass" do - # A completely unknown proof type should produce an error. - result = - try do - VCLExecutor.execute_string( - "SELECT GRAPH.* FROM HEXAD 'abc-123' PROOF BOGUS_TYPE(entity)", - timeout: 1_000 - ) - rescue - _ -> {:error, :unknown_proof_type} - end - - assert elem(result, 0) == :error - end - end - - # =========================================================================== - # ProvedResult structure - # =========================================================================== - - describe "ProvedResult structure" do - test "VCL-UT queries should produce proved results with certificate" do - # This test documents the expected shape of VCL-UT results. - # In production (with Rust running), the result should be: - # %{ - # data: [...], - # proof_certificate: %{ - # proofs: [...], - # composition: :conjunction, - # verified_at: ~U[...], - # query_hash: "sha256hex..." - # } - # } - # - # For now, we verify the executor code path doesn't crash. - result = - try do - VCLExecutor.execute_string( - "SELECT GRAPH.* FROM HEXAD 'abc-123' PROOF EXISTENCE(abc-123)", - timeout: 1_000 - ) - rescue - _ -> {:error, :rust_core_unavailable} - end - - assert is_tuple(result) - assert elem(result, 0) in [:ok, :error] - end - end -end diff --git a/verisimdb/elixir-orchestration/test/verisim/query/vcl_e2e_test.exs b/verisimdb/elixir-orchestration/test/verisim/query/vcl_e2e_test.exs deleted file mode 100644 index 3f8e5ae5..00000000 --- a/verisimdb/elixir-orchestration/test/verisim/query/vcl_e2e_test.exs +++ /dev/null @@ -1,613 +0,0 @@ -# SPDX-License-Identifier: MPL-2.0 - -defmodule VeriSim.Query.VCLE2ETest do - @moduledoc """ - VCL end-to-end tests exercising the full pipeline: - - VCL string → parse → typecheck → plan → execute → result - - These tests validate the VCL-SPEC contract by exercising every major - language feature through the complete pipeline. They are designed to - pass both with and without the Rust core running: - - - **Rust available**: full data round-trip, real results - - **Rust unavailable**: parse + typecheck + plan verified, execution - errors accepted gracefully - - ## Test categories - - 1. **Full pipeline round-trip** — parse → typecheck → execute for each query shape - 2. **Proof certificate round-trip** — typecheck → generate cert → verify cert - 3. **VCL-SPEC grammar coverage** — every production in vcl-grammar.ebnf tested - 4. **Error paths** — malformed queries, invalid proof types, bad modality combos - 5. **Cross-modal condition evaluation** — drift, consistency, exists checks - 6. **Mutation pipeline** — INSERT/UPDATE/DELETE parse → execute routing - 7. **Pagination invariants** — LIMIT/OFFSET/ORDER BY preserved through pipeline - 8. **Federation with drift policies** — all 4 policies parsed and routed - """ - - use ExUnit.Case, async: false - - alias VeriSim.Query.{VCLBridge, VCLExecutor, VCLTypeChecker, VCLProofCertificate} - alias VeriSim.Test.VCLTestHelpers, as: H - - setup_all do - pid = H.ensure_bridge_started() - %{bridge_pid: pid} - end - - # =========================================================================== - # 1. Full pipeline round-trip: parse → typecheck → execute - # =========================================================================== - - describe "full pipeline: SELECT with all 8 modalities" do - test "SELECT * produces AST with all 8 modalities and executes without crash" do - query = "SELECT * FROM HEXAD 'entity-001'" - ast = H.parse!(query) - - # AST shape: must have modalities and source - assert is_map(ast) - # Built-in parser includes quotes in entity ID - assert {:octad, _entity_id} = ast[:source] - - # Execute through full pipeline - result = H.execute_safely(query) - H.assert_ok_or_rust_unavailable(result) - end - - test "explicit 8-modality SELECT matches star expansion" do - query = """ - SELECT GRAPH.*, VECTOR.*, TENSOR.*, SEMANTIC.*, - DOCUMENT.*, TEMPORAL.*, PROVENANCE.*, SPATIAL.* - FROM HEXAD 'entity-001' - """ - - ast = H.parse!(query) - H.assert_modalities(ast, [:graph, :vector, :tensor, :semantic, - :document, :temporal, :provenance, :spatial]) - end - end - - describe "full pipeline: WHERE conditions through execution" do - test "field comparison condition is preserved through parse → execute" do - query = "SELECT DOCUMENT.* FROM HEXAD 'entity-001' WHERE DOCUMENT.severity > 5" - ast = H.parse!(query) - - H.assert_has_where(ast) - assert ast[:modalities] == [:document] - - result = H.execute_safely(query) - H.assert_ok_or_rust_unavailable(result) - end - - test "CONTAINS full-text search condition is parsed and routed" do - query = "SELECT DOCUMENT.* FROM HEXAD 'entity-001' WHERE DOCUMENT CONTAINS 'security'" - ast = H.parse!(query) - - H.assert_has_where(ast) - - result = H.execute_safely(query) - H.assert_ok_or_rust_unavailable(result) - end - - test "VECTOR SIMILAR TO condition is parsed with embedding" do - query = "SELECT VECTOR.* FROM HEXAD 'entity-001' WHERE VECTOR SIMILAR TO [0.1, 0.2, 0.3]" - ast = H.parse!(query) - - H.assert_has_where(ast) - assert :vector in (ast[:modalities] || []) - - result = H.execute_safely(query) - H.assert_ok_or_rust_unavailable(result) - end - - test "WITHIN RADIUS spatial condition is parsed" do - query = "SELECT SPATIAL.* FROM HEXAD 'entity-001' WHERE WITHIN RADIUS(51.5074, -0.1278, 5000)" - ast = H.parse!(query) - - H.assert_has_where(ast) - assert :spatial in (ast[:modalities] || []) - end - - test "DRIFT cross-modal condition is parsed with threshold" do - query = "SELECT * FROM HEXAD 'entity-001' WHERE DRIFT(VECTOR, DOCUMENT) > 0.3" - ast = H.parse!(query) - - H.assert_has_where(ast) - end - - test "CONSISTENT cross-modal condition with COSINE metric" do - query = "SELECT * FROM HEXAD 'entity-001' WHERE CONSISTENT(VECTOR, SEMANTIC) USING COSINE > 0.8" - ast = H.parse!(query) - - H.assert_has_where(ast) - end - - test "modality existence check conditions" do - query = "SELECT * FROM HEXAD 'entity-001' WHERE PROVENANCE EXISTS AND TENSOR NOT EXISTS" - ast = H.parse!(query) - - H.assert_has_where(ast) - end - end - - # =========================================================================== - # 2. Proof certificate round-trip - # =========================================================================== - - describe "proof certificate round-trip: typecheck → generate → verify" do - @proof_types ~w(existence integrity consistency provenance freshness - access citation zkp proven sanctify)a - - for proof_type <- @proof_types do - @pt proof_type - @pt_upper @pt |> Atom.to_string() |> String.upcase() - - test "#{@pt_upper} proof: typecheck → certificate → verify" do - # Build a VCL-UT query with this proof type - contract = if @pt in [:integrity, :citation, :custom, :sanctify] do - "(my_contract)" - else - "(entity-001)" - end - - query = "SELECT GRAPH.* FROM HEXAD 'entity-001' PROOF #{@pt_upper}#{contract}" - ast = H.parse!(query) - H.assert_has_proof(ast) - - # Type check returns {:ok, %{proof_obligations: [...], ...}} - case VCLTypeChecker.typecheck(ast) do - {:ok, tc_result} -> - obligations = tc_result[:proof_obligations] || tc_result.proof_obligations - assert is_list(obligations) - assert length(obligations) >= 1 - - obligation = hd(obligations) - assert obligation[:type] == @pt - - # Generate certificate with mock witness data - witness = build_mock_witness(@pt) - {:ok, cert} = VCLProofCertificate.generate_certificate(obligation, witness) - - # Certificate structure - assert cert.type == @pt - assert is_binary(cert.hash) - assert byte_size(cert.hash) == 32 # SHA-256 - - # Verify round-trip - assert :ok == VCLProofCertificate.verify_certificate(cert) - - {:error, _reason} -> - # Some proof types may fail if modality requirements not met - # (e.g., INTEGRITY needs semantic in SELECT) - :ok - end - end - end - - test "multi-proof AND composition: typecheck produces multiple obligations" do - query = "SELECT * FROM HEXAD 'entity-001' PROOF EXISTENCE(entity-001) AND PROVENANCE(entity-001)" - ast = H.parse!(query) - H.assert_has_proof(ast) - - {:ok, tc_result} = VCLTypeChecker.typecheck(ast) - obligations = tc_result[:proof_obligations] || tc_result.proof_obligations - assert is_list(obligations) - assert length(obligations) >= 2 - - # Generate batch certificates - witnesses = Enum.map(obligations, fn obl -> - {obl, build_mock_witness(obl[:type])} - end) - - {:ok, certs} = VCLProofCertificate.generate_batch(witnesses) - assert length(certs) >= 2 - assert :ok == VCLProofCertificate.verify_batch(certs) - end - - test "tampered certificate fails verification" do - obligation = %{ - type: :existence, - proofType: "EXISTENCE", - contract: "entity-001", - contractName: "entity-001", - witness_fields: ["octad_id", "timestamp", "modality_count"], - circuit: "existence-proof-v1" - } - - witness = %{ - "octad_id" => "entity-001", - "timestamp" => DateTime.to_iso8601(DateTime.utc_now()), - "modality_count" => 8 - } - - {:ok, cert} = VCLProofCertificate.generate_certificate(obligation, witness) - assert :ok == VCLProofCertificate.verify_certificate(cert) - - # Tamper with the witness - tampered = %{cert | witness: Map.put(cert.witness, "modality_count", 999)} - assert {:error, :invalid_hash} == VCLProofCertificate.verify_certificate(tampered) - end - end - - # =========================================================================== - # 3. VCL-SPEC grammar coverage - # =========================================================================== - - describe "VCL-SPEC grammar coverage" do - test "SELECT with column projection (not star)" do - query = "SELECT DOCUMENT.title, DOCUMENT.body FROM HEXAD 'entity-001'" - ast = H.parse!(query) - assert :document in (ast[:modalities] || []) - end - - test "SELECT with LIMIT and OFFSET" do - query = "SELECT GRAPH.* FROM HEXAD 'entity-001' LIMIT 10 OFFSET 5" - ast = H.parse!(query) - - H.assert_limit(ast, 10) - H.assert_offset(ast, 5) - end - - test "SELECT with ORDER BY preserves LIMIT" do - # Built-in parser may not parse ORDER BY into AST, but should still - # parse the query without error and preserve LIMIT - query = "SELECT DOCUMENT.* FROM HEXAD 'entity-001' LIMIT 20" - ast = H.parse!(query) - H.assert_limit(ast, 20) - end - - test "nested AND/OR in WHERE" do - query = """ - SELECT * FROM HEXAD 'entity-001' - WHERE DOCUMENT.severity > 5 AND PROVENANCE EXISTS - """ - ast = H.parse!(query) - H.assert_has_where(ast) - end - end - - # =========================================================================== - # 4. Error paths - # =========================================================================== - - describe "error paths" do - test "empty query string returns parse error" do - assert {:error, _reason} = VCLBridge.parse("") - end - - test "gibberish returns parse error" do - assert {:error, _reason} = VCLBridge.parse("FROBNICATE THE WIDGET") - end - - test "unknown proof type is rejected by type checker" do - query = "SELECT GRAPH.* FROM HEXAD 'entity-001' PROOF TELEPORT(entity-001)" - case VCLBridge.parse(query) do - {:ok, ast} -> - result = VCLTypeChecker.typecheck(ast) - assert {:error, _reason} = result - - {:error, _} -> - # Parser may also reject unknown proof types - :ok - end - end - - test "INTEGRITY proof without semantic modality is rejected" do - query = "SELECT GRAPH.* FROM HEXAD 'entity-001' PROOF INTEGRITY(my_contract)" - case VCLBridge.parse(query) do - {:ok, ast} -> - case VCLTypeChecker.typecheck(ast) do - {:error, reason} -> - # reason may be a tuple like {:modality_mismatch, "..."} - reason_str = inspect(reason) - assert reason_str =~ "semantic" or reason_str =~ "modality" - - {:ok, _} -> - # Some implementations allow it and check at execution time - :ok - end - - {:error, _} -> :ok - end - end - - test "PROVENANCE proof without provenance modality is rejected" do - query = "SELECT GRAPH.* FROM HEXAD 'entity-001' PROOF PROVENANCE(entity-001)" - case VCLBridge.parse(query) do - {:ok, ast} -> - case VCLTypeChecker.typecheck(ast) do - {:error, reason} -> - reason_str = inspect(reason) - assert reason_str =~ "provenance" or reason_str =~ "modality" - - {:ok, _} -> :ok - end - - {:error, _} -> :ok - end - end - - test "SELECT from nonexistent store returns error" do - query = "SELECT * FROM STORE 'nonexistent-store-99'" - result = H.execute_safely(query) - - case H.assert_ok_or_rust_unavailable(result) do - :rust_unavailable -> :ok - {:error_from_rust, _} -> :ok - :ok_result -> :ok # May return empty results - end - end - end - - # =========================================================================== - # 5. Cross-modal condition evaluation - # =========================================================================== - - describe "cross-modal conditions through execution" do - test "CrossModalFieldCompare is classified as cross-modal (not pushdown)" do - ast = %{ - modalities: [:document, :graph], - source: {:octad, "entity-001"}, - where: H.cross_modal_compare(:document, :severity, ">", :graph, :importance), - limit: nil, - offset: nil, - order_by: nil, - proof: nil - } - - result = H.execute_ast_safely(ast) - H.assert_ok_or_rust_unavailable(result) - end - - test "ModalityDrift condition is classified as cross-modal" do - ast = %{ - modalities: [:vector, :document], - source: {:octad, "entity-001"}, - where: H.modality_drift(:vector, :document, 0.3), - limit: nil, - offset: nil, - order_by: nil, - proof: nil - } - - result = H.execute_ast_safely(ast) - H.assert_ok_or_rust_unavailable(result) - end - - test "ModalityConsistency with JACCARD metric" do - ast = %{ - modalities: [:graph, :document], - source: {:octad, "entity-001"}, - where: H.modality_consistency(:graph, :document, "JACCARD"), - limit: nil, - offset: nil, - order_by: nil, - proof: nil - } - - result = H.execute_ast_safely(ast) - H.assert_ok_or_rust_unavailable(result) - end - - test "compound AND condition with existence + drift" do - ast = %{ - modalities: [:graph, :vector, :document, :provenance], - source: {:octad, "entity-001"}, - where: H.and_condition( - H.modality_exists(:provenance), - H.modality_drift(:vector, :document, 0.5) - ), - limit: nil, - offset: nil, - order_by: nil, - proof: nil - } - - result = H.execute_ast_safely(ast) - H.assert_ok_or_rust_unavailable(result) - end - end - - # =========================================================================== - # 6. Mutation pipeline - # =========================================================================== - - describe "mutation pipeline" do - test "INSERT HEXAD parses via parse_statement" do - query = "INSERT HEXAD WITH DOCUMENT(title = 'Test Entity', body = 'Integration test body')" - {:ok, ast} = VCLBridge.parse_statement(query) - - # Mutation AST shape: %{TAG: "Mutation", _0: %{TAG: "Insert", ...}} - assert ast[:TAG] == "Mutation" or ast[:type] == :insert or ast[:mutation_type] == :insert - inner = ast[:_0] || ast - assert inner[:TAG] == "Insert" or inner[:type] == :insert - - result = H.execute_statement_safely(ast) - H.assert_ok_or_rust_unavailable(result) - end - - test "UPDATE HEXAD parses with SET clause" do - query = "UPDATE HEXAD 'entity-001' SET DOCUMENT.title = 'Updated Title'" - {:ok, ast} = VCLBridge.parse_statement(query) - - assert ast[:TAG] == "Mutation" or ast[:type] == :update or ast[:mutation_type] == :update - inner = ast[:_0] || ast - assert inner[:TAG] == "Update" or inner[:type] == :update - - result = H.execute_statement_safely(ast) - H.assert_ok_or_rust_unavailable(result) - end - - test "DELETE HEXAD parses with entity ID" do - query = "DELETE HEXAD 'entity-to-delete'" - {:ok, ast} = VCLBridge.parse_statement(query) - - assert ast[:TAG] == "Mutation" or ast[:type] == :delete or ast[:mutation_type] == :delete - inner = ast[:_0] || ast - assert inner[:TAG] == "Delete" or inner[:type] == :delete - - result = H.execute_statement_safely(ast) - H.assert_ok_or_rust_unavailable(result) - end - end - - # =========================================================================== - # 7. Pagination invariants - # =========================================================================== - - describe "pagination invariants through pipeline" do - test "LIMIT is preserved from parse through execution" do - query = "SELECT GRAPH.* FROM HEXAD 'entity-001' LIMIT 42" - ast = H.parse!(query) - H.assert_limit(ast, 42) - - # Execute and verify limit was passed to Rust - result = H.execute_safely(query) - H.assert_ok_or_rust_unavailable(result) - end - - test "OFFSET is preserved from parse through execution" do - query = "SELECT GRAPH.* FROM HEXAD 'entity-001' LIMIT 10 OFFSET 25" - ast = H.parse!(query) - H.assert_limit(ast, 10) - H.assert_offset(ast, 25) - end - - test "large LIMIT value is preserved" do - query = "SELECT DOCUMENT.* FROM HEXAD 'entity-001' LIMIT 100" - ast = H.parse!(query) - H.assert_limit(ast, 100) - end - - test "zero LIMIT returns empty result" do - query = "SELECT GRAPH.* FROM HEXAD 'entity-001' LIMIT 0" - ast = H.parse!(query) - H.assert_limit(ast, 0) - end - end - - # =========================================================================== - # 8. Federation with drift policies - # =========================================================================== - - describe "federation queries with drift policies" do - for policy <- ~w(STRICT REPAIR TOLERATE LATEST) do - @policy policy - - test "federation with WITH DRIFT #{@policy} parses and routes" do - query = "SELECT * FROM FEDERATION /* WITH DRIFT #{@policy}" - ast = H.parse!(query) - - {:federation, _pattern, _drift} = ast[:source] - - result = H.execute_safely(query) - H.assert_ok_or_rust_unavailable(result) - end - end - - test "federation without drift policy defaults to TOLERATE" do - query = "SELECT * FROM FEDERATION /*" - ast = H.parse!(query) - - case ast[:source] do - {:federation, _pattern, drift_policy} -> - # Default policy should be tolerate or nil - assert drift_policy in [nil, :tolerate, "TOLERATE", "tolerate"] - - {:federation, _pattern} -> - # No drift policy field at all — acceptable - :ok - end - end - end - - # =========================================================================== - # Private: mock witness builders - # =========================================================================== - - defp build_mock_witness(:existence) do - %{ - "octad_id" => "entity-001", - "timestamp" => DateTime.to_iso8601(DateTime.utc_now()), - "modality_count" => 8 - } - end - - defp build_mock_witness(:integrity) do - %{ - "content_hash" => Base.encode16(:crypto.hash(:sha256, "test"), case: :lower), - "merkle_root" => Base.encode16(:crypto.hash(:sha256, "root"), case: :lower), - "schema_version" => "1.0.0" - } - end - - defp build_mock_witness(:consistency) do - %{ - "modality_a" => "vector", - "modality_b" => "semantic", - "drift_score" => 0.05, - "threshold" => 0.3 - } - end - - defp build_mock_witness(:provenance) do - %{ - "chain_hash" => Base.encode16(:crypto.hash(:sha256, "chain"), case: :lower), - "chain_length" => 5, - "origin" => "scanner-v1", - "actor_trail" => ["user-1", "system", "validator"] - } - end - - defp build_mock_witness(:freshness) do - %{ - "last_modified" => DateTime.to_iso8601(DateTime.utc_now()), - "max_age_ms" => 60_000, - "version_count" => 3 - } - end - - defp build_mock_witness(:access) do - %{ - "principal_id" => "user-001", - "resource_id" => "entity-001", - "permission_set" => ["read", "write"] - } - end - - defp build_mock_witness(:citation) do - %{ - "source_ids" => ["src-001", "src-002"], - "citation_chain" => ["doc-a", "doc-b"], - "reference_count" => 2 - } - end - - defp build_mock_witness(:custom) do - %{"circuit_inputs" => %{"custom_field" => "custom_value"}} - end - - defp build_mock_witness(:zkp) do - %{ - "claim" => "entity-001 has property X", - "blinding_nonce" => Base.encode16(:crypto.strong_rand_bytes(16), case: :lower) - } - end - - defp build_mock_witness(:proven) do - %{ - "certificate_hash" => Base.encode16(:crypto.hash(:sha256, "cert"), case: :lower), - "proof_data" => Base.encode64("mock-proof-blob") - } - end - - defp build_mock_witness(:sanctify) do - %{ - "contract_hash" => Base.encode16(:crypto.hash(:sha256, "contract"), case: :lower), - "security_level" => "high" - } - end - - defp build_mock_witness(_other), do: %{} -end diff --git a/verisimdb/elixir-orchestration/test/verisim/query/vcl_integration_test.exs b/verisimdb/elixir-orchestration/test/verisim/query/vcl_integration_test.exs deleted file mode 100644 index f2d3b337..00000000 --- a/verisimdb/elixir-orchestration/test/verisim/query/vcl_integration_test.exs +++ /dev/null @@ -1,343 +0,0 @@ -# SPDX-License-Identifier: MPL-2.0 - -defmodule VeriSim.Query.VCLIntegrationTest do - @moduledoc """ - VCL Slipstream end-to-end integration tests. - - Exercises the full parse → classify → execute → aggregate pipeline for the - Slipstream (no-proof) path across all 8 octad modalities. - - ## Test categories - - 1. **Single-modality queries** (8 tests) — one SELECT per modality - 2. **Multi-modality queries** (3 tests) — combined projections - 3. **WHERE condition queries** (5 tests) — pushdown and cross-modal filters - 4. **Mutations** (3 tests) — INSERT / UPDATE / DELETE - 5. **Aggregation** (2 tests) — COUNT, AVG with GROUP BY - 6. **Pagination & sorting** (2 tests) — ORDER BY, LIMIT, OFFSET - 7. **Federation queries** (2 tests) — federation source with drift policies - 8. **REFLECT queries** (1 test) — meta-circular query store - - Tests gracefully handle Rust core unavailability: they verify the parse → - route path is correct and that no crashes occur, accepting connection errors - as valid in CI environments where the Rust core is not running. - """ - - use ExUnit.Case, async: false - - alias VeriSim.Query.{VCLBridge, VCLExecutor} - alias VeriSim.Test.VCLTestHelpers, as: H - - setup_all do - pid = H.ensure_bridge_started() - %{bridge_pid: pid} - end - - # =========================================================================== - # 1. Single-modality SELECT queries (8 tests) - # =========================================================================== - - describe "single-modality SELECT queries" do - for modality <- ~w(GRAPH VECTOR TENSOR SEMANTIC DOCUMENT TEMPORAL PROVENANCE SPATIAL) do - @mod modality - - test "parses and routes SELECT #{@mod}.* FROM HEXAD" do - query = "SELECT #{@mod}.* FROM HEXAD 'entity-001'" - ast = H.parse!(query) - - assert is_map(ast) - expected_atom = @mod |> String.downcase() |> String.to_existing_atom() - assert expected_atom in (ast[:modalities] || []) - H.assert_source(ast, :octad) - - # Execute: should not crash regardless of Rust core availability - result = H.execute_safely(query) - assert elem(result, 0) in [:ok, :error, :unavailable] - end - end - end - - # =========================================================================== - # 2. Multi-modality SELECT queries (3 tests) - # =========================================================================== - - describe "multi-modality SELECT queries" do - test "three modalities: GRAPH, VECTOR, DOCUMENT" do - query = "SELECT GRAPH.*, VECTOR.*, DOCUMENT.* FROM HEXAD 'entity-001'" - ast = H.parse!(query) - - assert :graph in ast[:modalities] - assert :vector in ast[:modalities] - assert :document in ast[:modalities] - assert length(ast[:modalities]) == 3 - end - - test "all 8 modalities via wildcard SELECT *" do - query = "SELECT * FROM HEXAD 'entity-001'" - ast = H.parse!(query) - - # Wildcard produces :all in the modalities list - assert :all in ast[:modalities] - end - - test "new octad modalities: PROVENANCE and SPATIAL" do - query = "SELECT PROVENANCE.*, SPATIAL.* FROM HEXAD 'entity-001'" - ast = H.parse!(query) - - assert :provenance in ast[:modalities] - assert :spatial in ast[:modalities] - end - end - - # =========================================================================== - # 3. WHERE condition queries (5 tests) - # =========================================================================== - - describe "WHERE condition queries" do - test "field comparison: WHERE DOCUMENT.severity > 5" do - query = "SELECT DOCUMENT.* FROM HEXAD 'entity-001' WHERE DOCUMENT.severity > 5" - ast = H.parse!(query) - - H.assert_has_where(ast) - assert ast[:where][:raw] =~ "DOCUMENT.severity" - end - - test "fulltext: WHERE DOCUMENT CONTAINS 'security'" do - query = "SELECT DOCUMENT.* FROM HEXAD 'entity-001' WHERE DOCUMENT CONTAINS 'security'" - ast = H.parse!(query) - - H.assert_has_where(ast) - assert ast[:where][:raw] =~ "CONTAINS" - end - - test "vector similarity: WHERE VECTOR SIMILAR TO [0.1, 0.2, 0.3]" do - query = "SELECT VECTOR.* FROM HEXAD 'entity-001' WHERE VECTOR SIMILAR TO [0.1, 0.2, 0.3]" - ast = H.parse!(query) - - H.assert_has_where(ast) - assert ast[:where][:raw] =~ "SIMILAR" - end - - test "cross-modal drift: WHERE DRIFT(VECTOR, DOCUMENT) > 0.3" do - query = "SELECT * FROM HEXAD 'entity-001' WHERE DRIFT(VECTOR, DOCUMENT) > 0.3" - ast = H.parse!(query) - - H.assert_has_where(ast) - assert ast[:where][:raw] =~ "DRIFT" - end - - test "modality existence: WHERE PROVENANCE EXISTS AND TENSOR NOT EXISTS" do - # The built-in parser stores the raw WHERE text — cross-modal classification - # happens at execution time, not parse time. - query = "SELECT * FROM HEXAD 'entity-001' WHERE PROVENANCE EXISTS AND TENSOR NOT EXISTS" - ast = H.parse!(query) - - H.assert_has_where(ast) - raw = ast[:where][:raw] - assert raw =~ "PROVENANCE" - assert raw =~ "TENSOR" - end - end - - # =========================================================================== - # 4. Mutations (3 tests) - # =========================================================================== - - describe "mutation parsing and routing" do - test "INSERT HEXAD WITH DOCUMENT(...)" do - query = "INSERT HEXAD WITH DOCUMENT(title = 'New Entity', body = 'Test body')" - ast = H.parse_statement!(query) - - assert ast[:TAG] == "Mutation" - mutation = ast[:_0] - assert mutation[:TAG] == "Insert" - end - - test "UPDATE HEXAD SET DOCUMENT.title" do - query = "UPDATE HEXAD 'entity-001' SET DOCUMENT.title = 'Updated Title'" - ast = H.parse_statement!(query) - - assert ast[:TAG] == "Mutation" - mutation = ast[:_0] - assert mutation[:TAG] == "Update" - assert mutation[:octadId] == "'entity-001'" - end - - test "DELETE HEXAD" do - query = "DELETE HEXAD 'entity-001'" - ast = H.parse_statement!(query) - - assert ast[:TAG] == "Mutation" - mutation = ast[:_0] - assert mutation[:TAG] == "Delete" - end - end - - describe "mutation execution" do - test "INSERT routes to RustClient.create_octad without crashing" do - query = "INSERT HEXAD WITH DOCUMENT(title = 'Test Insert', body = 'Body')" - ast = H.parse_statement!(query) - - result = H.execute_statement_safely(ast) - # Should return :ok (Rust running) or :error (Rust unavailable) — never crash - assert elem(result, 0) in [:ok, :error, :unavailable] - end - - test "UPDATE routes to RustClient.update_octad without crashing" do - query = "UPDATE HEXAD 'entity-001' SET DOCUMENT.title = 'New Title'" - ast = H.parse_statement!(query) - - result = H.execute_statement_safely(ast) - assert elem(result, 0) in [:ok, :error, :unavailable] - end - - test "DELETE routes to RustClient.delete_octad without crashing" do - query = "DELETE HEXAD 'entity-001'" - ast = H.parse_statement!(query) - - result = H.execute_statement_safely(ast) - assert elem(result, 0) in [:ok, :error, :unavailable] - end - end - - # =========================================================================== - # 5. Aggregation (2 tests) - # =========================================================================== - - describe "aggregation queries" do - test "SELECT COUNT(*) FROM FEDERATION parses or fails gracefully" do - # COUNT(*) is an aggregate function — the built-in parser may not recognize - # it as a modality during SELECT parsing, which is acceptable. The full - # ReScript parser handles aggregates natively. - result = VCLBridge.parse("SELECT COUNT(*) FROM FEDERATION /*") - - case result do - {:ok, ast} -> - assert is_map(ast) - - {:error, _reason} -> - # Built-in parser doesn't support aggregate-only projections — acceptable. - # When the full ReScript parser is available, this test should pass. - :ok - end - end - - test "GROUP BY clause is preserved in AST" do - query = "SELECT GRAPH.* FROM HEXAD 'entity-001' LIMIT 5" - ast = H.parse!(query) - - # The basic parser stores groupBy as nil when not present - assert Map.has_key?(ast, :groupBy) - end - end - - # =========================================================================== - # 6. Pagination & sorting (2 tests) - # =========================================================================== - - describe "pagination and sorting" do - test "LIMIT and OFFSET are parsed correctly" do - query = "SELECT * FROM HEXAD 'entity-001' LIMIT 10 OFFSET 5" - ast = H.parse!(query) - - H.assert_limit(ast, 10) - H.assert_offset(ast, 5) - end - - test "ORDER BY clause is preserved in AST" do - query = "SELECT DOCUMENT.* FROM HEXAD 'entity-001' ORDER BY DOCUMENT.title ASC LIMIT 20" - ast = H.parse!(query) - - # The built-in parser stores orderBy - assert Map.has_key?(ast, :orderBy) - H.assert_limit(ast, 20) - end - end - - # =========================================================================== - # 7. Federation queries (2 tests) - # =========================================================================== - - describe "federation queries" do - test "basic federation query with wildcard pattern" do - query = "SELECT * FROM FEDERATION /*" - ast = H.parse!(query) - - H.assert_source(ast, :federation) - end - - test "federation with drift policy" do - query = "SELECT * FROM FEDERATION /* WITH DRIFT STRICT" - ast = H.parse!(query) - - H.assert_source(ast, :federation) - - # The source tuple should include the drift policy - {:federation, _pattern, drift_policy} = ast[:source] - assert drift_policy == :strict - end - - test "federation query executes without crashing" do - query = "SELECT * FROM FEDERATION /*" - - result = H.execute_safely(query) - assert elem(result, 0) in [:ok, :error, :unavailable] - end - end - - # =========================================================================== - # 8. REFLECT query (1 test) - # =========================================================================== - - describe "REFLECT queries (meta-circular)" do - test "parses REFLECT query — queries the query store itself" do - # REFLECT is a source type: SELECT ... FROM REFLECT - # The built-in parser may not support this directly, but it should not crash. - result = VCLBridge.parse("SELECT * FROM REFLECT") - - case result do - {:ok, ast} -> - assert is_map(ast) - - {:error, _reason} -> - # Built-in parser may not support REFLECT source — acceptable - :ok - end - end - end - - # =========================================================================== - # 9. Explain plan (1 test) - # =========================================================================== - - describe "explain plans" do - test "explain: true returns execution plan structure" do - query = "SELECT GRAPH.*, DOCUMENT.* FROM HEXAD 'entity-001'" - ast = H.parse!(query) - - {:ok, plan} = VCLExecutor.execute(ast, explain: true) - - assert is_map(plan) - assert Map.has_key?(plan, :strategy) - assert Map.has_key?(plan, :steps) - end - end - - # =========================================================================== - # 10. Error handling (3 tests) - # =========================================================================== - - describe "error handling" do - test "empty string returns parse error" do - assert {:error, {:parse_error, _}} = VCLExecutor.execute_string("") - end - - test "nonsense returns parse error" do - assert {:error, {:parse_error, _}} = VCLExecutor.execute_string("JABBERWOCKY SNARK") - end - - test "missing FROM clause returns parse error" do - assert {:error, {:parse_error, _}} = VCLExecutor.execute_string("SELECT GRAPH.*") - end - end -end diff --git a/verisimdb/elixir-orchestration/test/verisim/query/vcl_proof_certificate_test.exs b/verisimdb/elixir-orchestration/test/verisim/query/vcl_proof_certificate_test.exs deleted file mode 100644 index dc59aac5..00000000 --- a/verisimdb/elixir-orchestration/test/verisim/query/vcl_proof_certificate_test.exs +++ /dev/null @@ -1,160 +0,0 @@ -# SPDX-License-Identifier: MPL-2.0 - -defmodule VeriSim.Query.VCLProofCertificateTest do - @moduledoc """ - Tests for VCL-UT proof certificate generation and verification. - - Verifies that: - 1. Certificates are generated with correct structure and SHA-256 hash - 2. Valid certificates pass verification - 3. Tampered certificates fail verification - 4. Batch generation and verification work correctly - 5. Edge cases (missing type, non-map args) are handled gracefully - """ - - use ExUnit.Case, async: true - - alias VeriSim.Query.VCLProofCertificate - - # --------------------------------------------------------------------------- - # Fixtures - # --------------------------------------------------------------------------- - - defp existence_obligation do - %{ - type: :existence, - proofType: "EXISTENCE", - contract: "entity-001", - contractName: "entity-001", - witness_fields: ["octad_id", "timestamp", "modality_count"], - circuit: "existence-proof-v1", - estimated_time_ms: 50, - required_modalities: [] - } - end - - defp existence_witness do - %{ - "octad_id" => "entity-001", - "timestamp" => "2026-02-28T12:00:00Z", - "modality_count" => 8 - } - end - - defp provenance_obligation do - %{ - type: :provenance, - proofType: "PROVENANCE", - contract: "entity-002", - contractName: "entity-002", - witness_fields: ["chain_hash", "chain_length", "origin", "actor_trail"], - circuit: "provenance-proof-v1", - estimated_time_ms: 300, - required_modalities: [:provenance] - } - end - - defp provenance_witness do - %{ - "chain_hash" => "abc123def456", - "chain_length" => 5, - "origin" => "import-pipeline", - "actor_trail" => ["ingester", "normalizer", "validator"] - } - end - - # --------------------------------------------------------------------------- - # Test: Certificate generation - # --------------------------------------------------------------------------- - - describe "generate_certificate/2" do - test "produces a certificate with all required fields" do - {:ok, cert} = VCLProofCertificate.generate_certificate(existence_obligation(), existence_witness()) - - assert cert.type == :existence - assert cert.obligation == existence_obligation() - assert cert.witness == existence_witness() - assert %DateTime{} = cert.timestamp - assert is_binary(cert.hash) - assert byte_size(cert.hash) == 32 # SHA-256 = 32 bytes - end - - test "rejects obligation without :type field" do - bad_obligation = %{proofType: "EXISTENCE", contract: "x"} - assert {:error, {:invalid_obligation, _}} = VCLProofCertificate.generate_certificate(bad_obligation, %{}) - end - - test "rejects non-map arguments" do - assert {:error, {:invalid_arguments, _}} = VCLProofCertificate.generate_certificate("not a map", %{}) - assert {:error, {:invalid_arguments, _}} = VCLProofCertificate.generate_certificate(%{type: :existence}, "not a map") - end - end - - # --------------------------------------------------------------------------- - # Test: Certificate verification - # --------------------------------------------------------------------------- - - describe "verify_certificate/1" do - test "valid certificate passes verification" do - {:ok, cert} = VCLProofCertificate.generate_certificate(existence_obligation(), existence_witness()) - assert :ok = VCLProofCertificate.verify_certificate(cert) - end - - test "tampered obligation causes hash mismatch" do - {:ok, cert} = VCLProofCertificate.generate_certificate(existence_obligation(), existence_witness()) - - tampered = %{cert | obligation: %{cert.obligation | type: :integrity}} - assert {:error, :invalid_hash} = VCLProofCertificate.verify_certificate(tampered) - end - - test "tampered witness causes hash mismatch" do - {:ok, cert} = VCLProofCertificate.generate_certificate(existence_obligation(), existence_witness()) - - tampered = %{cert | witness: %{"octad_id" => "evil-entity"}} - assert {:error, :invalid_hash} = VCLProofCertificate.verify_certificate(tampered) - end - - test "malformed certificate returns error" do - assert {:error, :malformed_certificate} = VCLProofCertificate.verify_certificate(%{}) - assert {:error, :malformed_certificate} = VCLProofCertificate.verify_certificate(nil) - assert {:error, :malformed_certificate} = VCLProofCertificate.verify_certificate("not a cert") - end - end - - # --------------------------------------------------------------------------- - # Test: Batch operations - # --------------------------------------------------------------------------- - - describe "generate_batch/1 and verify_batch/1" do - test "batch generation produces certificates for all pairs" do - pairs = [ - {existence_obligation(), existence_witness()}, - {provenance_obligation(), provenance_witness()} - ] - - {:ok, certs} = VCLProofCertificate.generate_batch(pairs) - assert length(certs) == 2 - assert Enum.at(certs, 0).type == :existence - assert Enum.at(certs, 1).type == :provenance - - # All certificates should verify - assert :ok = VCLProofCertificate.verify_batch(certs) - end - - test "batch verification fails with index on tampered certificate" do - pairs = [ - {existence_obligation(), existence_witness()}, - {provenance_obligation(), provenance_witness()} - ] - - {:ok, certs} = VCLProofCertificate.generate_batch(pairs) - - # Tamper with the second certificate - tampered_second = %{Enum.at(certs, 1) | witness: %{"chain_hash" => "tampered"}} - tampered_batch = List.replace_at(certs, 1, tampered_second) - - assert {:error, {:batch_failure, 1, :invalid_hash}} = - VCLProofCertificate.verify_batch(tampered_batch) - end - end -end diff --git a/verisimdb/elixir-orchestration/test/verisim/query/vcl_property_test.exs b/verisimdb/elixir-orchestration/test/verisim/query/vcl_property_test.exs deleted file mode 100644 index 1ce6f536..00000000 --- a/verisimdb/elixir-orchestration/test/verisim/query/vcl_property_test.exs +++ /dev/null @@ -1,117 +0,0 @@ -# SPDX-License-Identifier: MPL-2.0 - -defmodule VeriSim.Query.VCLPropertyTest do - @moduledoc """ - Property-based tests for the VCL type checker using StreamData. - - Generates random valid proof specifications and verifies that the type - checker handles them correctly — no crashes, proper error messages for - invalid inputs, and structural invariants on output. - """ - - use ExUnit.Case, async: true - use ExUnitProperties - - alias VeriSim.Query.VCLTypeChecker - - # --------------------------------------------------------------------------- - # Generators - # --------------------------------------------------------------------------- - - defp valid_proof_type do - member_of(~w(existence integrity consistency provenance freshness access citation custom zkp proven sanctify)) - end - - defp valid_modality do - member_of(~w(graph vector tensor semantic document temporal provenance spatial)a) - end - - defp valid_contract_name do - string(:alphanumeric, min_length: 3, max_length: 20) - |> map(fn s -> "entity-#{s}" end) - end - - defp proof_spec_raw do - gen all proof_type <- valid_proof_type(), - contract <- valid_contract_name() do - %{raw: "#{String.upcase(proof_type)}(#{contract})"} - end - end - - defp valid_query_ast do - gen all modalities <- list_of(valid_modality(), min_length: 1, max_length: 4), - proof_specs <- list_of(proof_spec_raw(), min_length: 1, max_length: 3) do - %{ - modalities: Enum.uniq(modalities), - proof: proof_specs - } - end - end - - # --------------------------------------------------------------------------- - # Properties - # --------------------------------------------------------------------------- - - property "type checker never crashes on valid proof specs" do - check all query_ast <- valid_query_ast() do - result = VCLTypeChecker.typecheck(query_ast) - assert match?({:ok, _}, result) or match?({:error, _}, result) - end - end - - property "type checker output has required fields when successful" do - check all query_ast <- valid_query_ast() do - case VCLTypeChecker.typecheck(query_ast) do - {:ok, info} -> - assert is_list(info.proof_obligations) - assert info.composition_strategy in [:independent, :conjunction, :sequential] - assert is_map(info.inferred_types) - assert is_number(info.total_estimated_ms) - assert is_boolean(info.is_parallelizable) - - {:error, _} -> - # Modality mismatch or other valid rejection — that's fine - :ok - end - end - end - - property "parse_proof_specs round-trips without data loss" do - check all proof_type <- valid_proof_type(), - contract <- valid_contract_name() do - raw = "#{String.upcase(proof_type)}(#{contract})" - specs = VCLTypeChecker.parse_proof_specs(%{raw: raw}) - - assert length(specs) == 1 - [spec] = specs - assert spec.proofType == String.upcase(proof_type) - assert spec.contractName == contract - end - end - - property "multi-proof specs split correctly on AND/OR" do - check all types <- list_of(valid_proof_type(), min_length: 2, max_length: 4), - contracts <- list_of(valid_contract_name(), length: length(types)) do - parts = Enum.zip(types, contracts) |> Enum.map(fn {t, c} -> "#{String.upcase(t)}(#{c})" end) - raw = Enum.join(parts, " AND ") - - specs = VCLTypeChecker.parse_proof_specs(%{raw: raw}) - assert length(specs) == length(types) - end - end - - property "unknown proof types always rejected" do - # Prefix with "XBOGUS" to guarantee the generated string never collides - # with a known proof type (existence, integrity, consistency, etc.) - check all suffix <- string(:alphanumeric, min_length: 3, max_length: 10) do - bad_type = "XBOGUS#{suffix}" - - query_ast = %{ - modalities: [:all], - proof: [%{raw: "#{bad_type}(test)"}] - } - - assert {:error, {:unknown_proof_type, _}} = VCLTypeChecker.typecheck(query_ast) - end - end -end diff --git a/verisimdb/elixir-orchestration/test/verisim/query/vcl_test.exs b/verisimdb/elixir-orchestration/test/verisim/query/vcl_test.exs deleted file mode 100644 index a219dd9a..00000000 --- a/verisimdb/elixir-orchestration/test/verisim/query/vcl_test.exs +++ /dev/null @@ -1,273 +0,0 @@ -# SPDX-License-Identifier: MPL-2.0 - -defmodule VeriSim.Query.VCLTest do - @moduledoc """ - VCL end-to-end tests for parsing, routing, and execution. - - Tests the VCL Slipstream path (no PROOF clause) across all 8 octad modalities. - Uses the built-in Elixir fallback parser via VCLBridge.parse/1. - """ - - use ExUnit.Case, async: false - - alias VeriSim.Query.{VCLBridge, VCLExecutor} - - setup_all do - # Start VCLBridge GenServer (will fall back to built-in parser without Deno) - case VCLBridge.start_link([]) do - {:ok, pid} -> %{bridge_pid: pid} - {:error, {:already_started, pid}} -> %{bridge_pid: pid} - end - end - - # =========================================================================== - # Parse tests — verify the bridge produces a well-formed AST - # =========================================================================== - - describe "VCLBridge.parse/1 — SELECT queries" do - test "parses single-modality GRAPH query" do - {:ok, ast} = VCLBridge.parse("SELECT GRAPH.* FROM HEXAD 'abc-123'") - assert is_map(ast) - assert ast[:modalities] || ast["modalities"] - end - - test "parses single-modality VECTOR query" do - {:ok, ast} = VCLBridge.parse("SELECT VECTOR.* FROM HEXAD 'abc-123'") - assert is_map(ast) - end - - test "parses single-modality DOCUMENT query" do - {:ok, ast} = VCLBridge.parse("SELECT DOCUMENT.* FROM HEXAD 'abc-123'") - assert is_map(ast) - end - - test "parses single-modality TEMPORAL query" do - {:ok, ast} = VCLBridge.parse("SELECT TEMPORAL.* FROM HEXAD 'abc-123'") - assert is_map(ast) - end - - test "parses single-modality PROVENANCE query" do - {:ok, ast} = VCLBridge.parse("SELECT PROVENANCE.* FROM HEXAD 'abc-123'") - assert is_map(ast) - end - - test "parses single-modality SPATIAL query" do - {:ok, ast} = VCLBridge.parse("SELECT SPATIAL.* FROM HEXAD 'abc-123'") - assert is_map(ast) - end - - test "parses multi-modality query" do - {:ok, ast} = - VCLBridge.parse("SELECT GRAPH.*, VECTOR.*, DOCUMENT.* FROM HEXAD 'abc-123'") - - assert is_map(ast) - end - - test "parses ALL modality wildcard" do - {:ok, ast} = VCLBridge.parse("SELECT * FROM HEXAD 'abc-123'") - assert is_map(ast) - end - - test "parses query with WHERE clause" do - {:ok, ast} = - VCLBridge.parse( - "SELECT DOCUMENT.* FROM HEXAD 'abc-123' WHERE DOCUMENT.title = 'Test'" - ) - - assert is_map(ast) - end - - test "parses query with LIMIT and OFFSET" do - {:ok, ast} = - VCLBridge.parse( - "SELECT GRAPH.* FROM HEXAD 'abc-123' LIMIT 10 OFFSET 5" - ) - - assert is_map(ast) - end - - test "parses query with ORDER BY" do - {:ok, ast} = - VCLBridge.parse( - "SELECT DOCUMENT.* FROM HEXAD 'abc-123' ORDER BY DOCUMENT.title ASC" - ) - - assert is_map(ast) - end - - test "parses query with aggregate in projection" do - {:ok, ast} = - VCLBridge.parse( - "SELECT GRAPH.* FROM HEXAD 'abc-123' LIMIT 5" - ) - - assert is_map(ast) - end - end - - describe "VCLBridge — mutations via parse_statement/1" do - test "parses INSERT with document data" do - {:ok, ast} = - VCLBridge.parse_statement( - "INSERT HEXAD WITH DOCUMENT(title = 'Test', body = 'Content')" - ) - - assert is_map(ast) - assert ast[:TAG] == "Mutation" or ast["TAG"] == "Mutation" - end - - test "parses UPDATE with set clause" do - {:ok, ast} = - VCLBridge.parse_statement( - "UPDATE HEXAD 'abc-123' SET DOCUMENT.title = 'Updated'" - ) - - assert is_map(ast) - end - - test "parses DELETE" do - {:ok, ast} = - VCLBridge.parse_statement("DELETE HEXAD 'abc-123'") - - assert is_map(ast) - end - end - - describe "VCLBridge.parse/1 — error cases" do - test "rejects empty string" do - assert {:error, _} = VCLBridge.parse("") - end - - test "rejects gibberish" do - assert {:error, _} = VCLBridge.parse("THIS IS NOT VCL AT ALL") - end - end - - # =========================================================================== - # Executor tests — verify routing and error handling (without Rust core) - # =========================================================================== - - describe "VCLExecutor.execute_string/2" do - test "returns parse error for invalid query" do - assert {:error, {:parse_error, _}} = VCLExecutor.execute_string("NOT VCL") - end - - test "returns explain plan when explain option set" do - result = - VCLExecutor.execute_string( - "SELECT GRAPH.* FROM HEXAD 'abc-123'", - explain: true - ) - - case result do - {:ok, plan} -> - assert is_map(plan) or is_list(plan) - - {:error, _} -> - # Parse might fail if bridge not started — acceptable in unit test - :ok - end - end - end - - describe "VCLExecutor.execute_statement/2" do - test "routes query AST through execute without crashing" do - {:ok, ast} = VCLBridge.parse_statement("SELECT GRAPH.* FROM HEXAD 'abc-123'") - - # Without Rust core running at 8080, the executor will fail to reach it. - # We verify it does not crash (BadMapError, etc.) — any {:ok, _} or {:error, _} is fine. - result = - try do - VCLExecutor.execute_statement(ast, timeout: 1_000) - rescue - # If the response is not JSON (e.g., some other HTTP server is on 8080), - # the executor may blow up with a BadMapError — that's a test-env issue, - # not a code bug. Accept it as a known limitation. - _ -> {:error, :rust_core_unavailable} - end - - assert is_tuple(result) - assert elem(result, 0) in [:ok, :error] - end - - test "routes mutation AST through execute_mutation without crashing" do - {:ok, ast} = VCLBridge.parse_statement("DELETE HEXAD 'abc-123'") - - result = - try do - VCLExecutor.execute_statement(ast, timeout: 1_000) - rescue - _ -> {:error, :rust_core_unavailable} - end - - assert is_tuple(result) - assert elem(result, 0) in [:ok, :error] - end - end - - # =========================================================================== - # Cross-modal condition classification - # =========================================================================== - - describe "condition classification" do - test "simple condition is pushed down" do - {:ok, ast} = - VCLBridge.parse( - "SELECT GRAPH.* FROM HEXAD 'abc-123' WHERE GRAPH.type = 'Person'" - ) - - # The where clause should be present and classified as pushdown - assert is_map(ast) - end - end - - # =========================================================================== - # Provenance query routing - # =========================================================================== - - describe "provenance query routing" do - test "parses provenance-only query" do - {:ok, ast} = - VCLBridge.parse("SELECT PROVENANCE.* FROM HEXAD 'abc-123'") - - assert is_map(ast) - end - end - - # =========================================================================== - # Spatial query routing - # =========================================================================== - - describe "spatial query routing" do - test "parses spatial-only query" do - {:ok, ast} = - VCLBridge.parse("SELECT SPATIAL.* FROM HEXAD 'abc-123'") - - assert is_map(ast) - end - end - - # =========================================================================== - # Security: null-byte sanitisation (P1 hardening) - # =========================================================================== - - describe "null-byte sanitisation" do - test "rejects query containing a null byte" do - # Null bytes in entity IDs truncate C strings at the Rust FFI boundary, - # allowing forged IDs. The built-in parser must reject them. - assert {:error, msg} = VCLBridge.parse("SELECT GRAPH FROM HEXAD abc\0evil") - assert msg =~ "null" - end - - test "rejects query with null byte in WHERE clause" do - assert {:error, msg} = - VCLBridge.parse("SELECT GRAPH FROM HEXAD abc-123 WHERE GRAPH.id = 'ok\0bad'") - - assert msg =~ "null" - end - - test "clean query without null bytes is accepted" do - assert {:ok, _ast} = VCLBridge.parse("SELECT GRAPH FROM HEXAD abc-123") - end - end -end diff --git a/verisimdb/elixir-orchestration/test/verisim/query/vcl_type_checker_test.exs b/verisimdb/elixir-orchestration/test/verisim/query/vcl_type_checker_test.exs deleted file mode 100644 index 3699706c..00000000 --- a/verisimdb/elixir-orchestration/test/verisim/query/vcl_type_checker_test.exs +++ /dev/null @@ -1,417 +0,0 @@ -# SPDX-License-Identifier: MPL-2.0 - -defmodule VeriSim.Query.VCLTypeCheckerTest do - @moduledoc """ - Tests for the Elixir-native VCL-UT type checker. - - Verifies that the type checker: - 1. Validates proof types are known - 2. Checks modality compatibility (proof requirements vs queried modalities) - 3. Splits multi-proof raw strings correctly - 4. Generates structured obligations with witness fields and circuits - 5. Determines correct composition strategies - 6. Rejects malformed queries with specific error reasons - """ - - use ExUnit.Case, async: true - - alias VeriSim.Query.VCLTypeChecker - - # =========================================================================== - # parse_proof_specs/1 — multi-proof splitting - # =========================================================================== - - describe "parse_proof_specs/1" do - test "splits AND-connected proofs" do - specs = VCLTypeChecker.parse_proof_specs(%{raw: "EXISTENCE(entity-001) AND PROVENANCE(entity-001)"}) - - assert length(specs) == 2 - assert Enum.at(specs, 0).proofType == "EXISTENCE" - assert Enum.at(specs, 0).contractName == "entity-001" - assert Enum.at(specs, 1).proofType == "PROVENANCE" - assert Enum.at(specs, 1).contractName == "entity-001" - end - - test "splits OR-connected proofs" do - specs = VCLTypeChecker.parse_proof_specs(%{raw: "EXISTENCE(a) OR INTEGRITY(b)"}) - - assert length(specs) == 2 - assert Enum.at(specs, 0).proofType == "EXISTENCE" - assert Enum.at(specs, 1).proofType == "INTEGRITY" - end - - test "handles single proof spec" do - specs = VCLTypeChecker.parse_proof_specs(%{raw: "INTEGRITY(my_contract)"}) - - assert length(specs) == 1 - assert Enum.at(specs, 0).proofType == "INTEGRITY" - assert Enum.at(specs, 0).contractName == "my_contract" - end - - test "handles proof without parentheses" do - specs = VCLTypeChecker.parse_proof_specs(%{raw: "EXISTENCE entity-001"}) - - assert length(specs) == 1 - assert Enum.at(specs, 0).proofType == "EXISTENCE" - assert Enum.at(specs, 0).contractName == "entity-001" - end - - test "handles three proofs" do - specs = VCLTypeChecker.parse_proof_specs( - %{raw: "EXISTENCE(a) AND PROVENANCE(b) AND INTEGRITY(c)"} - ) - - assert length(specs) == 3 - assert Enum.map(specs, & &1.proofType) == ["EXISTENCE", "PROVENANCE", "INTEGRITY"] - end - - test "passes through structured specs unchanged" do - input = %{proofType: "EXISTENCE", contractName: "entity-001"} - specs = VCLTypeChecker.parse_proof_specs(input) - - assert length(specs) == 1 - assert Enum.at(specs, 0) == input - end - - test "handles nil" do - assert VCLTypeChecker.parse_proof_specs(nil) == [] - end - - test "handles list of specs" do - input = [ - %{proofType: "EXISTENCE", contractName: "a"}, - %{proofType: "INTEGRITY", contractName: "b"} - ] - specs = VCLTypeChecker.parse_proof_specs(input) - - assert length(specs) == 2 - end - - test "handles list containing raw specs" do - input = [%{raw: "EXISTENCE(a) AND PROVENANCE(b)"}] - specs = VCLTypeChecker.parse_proof_specs(input) - - assert length(specs) == 2 - end - end - - # =========================================================================== - # typecheck/1 — full type checking - # =========================================================================== - - describe "typecheck/1 — valid queries" do - test "accepts simple EXISTENCE proof" do - ast = %{ - modalities: [:graph], - proof: [%{proofType: "EXISTENCE", contractName: "entity-001"}] - } - - assert {:ok, info} = VCLTypeChecker.typecheck(ast) - assert length(info.proof_obligations) == 1 - assert Enum.at(info.proof_obligations, 0).type == :existence - assert info.composition_strategy == :independent - end - - test "accepts multi-proof conjunction" do - ast = %{ - modalities: [:graph, :provenance], - proof: [ - %{proofType: "EXISTENCE", contractName: "entity-001"}, - %{proofType: "PROVENANCE", contractName: "entity-001"} - ] - } - - assert {:ok, info} = VCLTypeChecker.typecheck(ast) - assert length(info.proof_obligations) == 2 - assert info.composition_strategy == :conjunction - end - - test "accepts PROVENANCE proof with provenance modality" do - ast = %{ - modalities: [:provenance], - proof: [%{proofType: "PROVENANCE", contractName: "entity-001"}] - } - - assert {:ok, info} = VCLTypeChecker.typecheck(ast) - assert Enum.at(info.proof_obligations, 0).type == :provenance - end - - test "accepts any proof with :all modalities" do - ast = %{ - modalities: [:all], - proof: [%{proofType: "INTEGRITY", contractName: "contract-001"}] - } - - assert {:ok, info} = VCLTypeChecker.typecheck(ast) - assert Enum.at(info.proof_obligations, 0).type == :integrity - end - - test "accepts CONSISTENCY proof" do - ast = %{ - modalities: [:graph, :vector], - proof: [%{proofType: "CONSISTENCY", contractName: "entity-001"}] - } - - assert {:ok, info} = VCLTypeChecker.typecheck(ast) - assert Enum.at(info.proof_obligations, 0).type == :consistency - end - - test "accepts FRESHNESS proof with temporal modality" do - ast = %{ - modalities: [:temporal], - proof: [%{proofType: "FRESHNESS", contractName: "entity-001"}] - } - - assert {:ok, info} = VCLTypeChecker.typecheck(ast) - assert Enum.at(info.proof_obligations, 0).type == :freshness - end - - test "accepts raw proof spec string" do - ast = %{ - modalities: [:graph], - proof: %{raw: "EXISTENCE(entity-001)"} - } - - assert {:ok, info} = VCLTypeChecker.typecheck(ast) - assert length(info.proof_obligations) == 1 - end - - test "accepts raw multi-proof string" do - ast = %{ - modalities: [:graph, :provenance], - proof: %{raw: "EXISTENCE(entity-001) AND PROVENANCE(entity-001)"} - } - - assert {:ok, info} = VCLTypeChecker.typecheck(ast) - assert length(info.proof_obligations) == 2 - end - end - - describe "typecheck/1 — invalid queries" do - test "rejects unknown proof type" do - ast = %{ - modalities: [:graph], - proof: [%{proofType: "BOGUS", contractName: "entity-001"}] - } - - assert {:error, {:unknown_proof_type, msg}} = VCLTypeChecker.typecheck(ast) - assert msg =~ "BOGUS" - end - - test "rejects INTEGRITY proof without contract" do - ast = %{ - modalities: [:semantic], - proof: [%{proofType: "INTEGRITY"}] - } - - assert {:error, {:missing_contract, _}} = VCLTypeChecker.typecheck(ast) - end - - test "rejects CITATION proof without contract" do - ast = %{ - modalities: [:document], - proof: [%{proofType: "CITATION"}] - } - - assert {:error, {:missing_contract, _}} = VCLTypeChecker.typecheck(ast) - end - - test "rejects INTEGRITY proof with wrong modality" do - ast = %{ - modalities: [:graph], - proof: [%{proofType: "INTEGRITY", contractName: "contract-001"}] - } - - assert {:error, {:modality_mismatch, msg}} = VCLTypeChecker.typecheck(ast) - assert msg =~ "semantic" - end - - test "rejects PROVENANCE proof without provenance modality" do - ast = %{ - modalities: [:graph], - proof: [%{proofType: "PROVENANCE", contractName: "entity-001"}] - } - - assert {:error, {:modality_mismatch, msg}} = VCLTypeChecker.typecheck(ast) - assert msg =~ "provenance" - end - - test "rejects FRESHNESS proof without temporal modality" do - ast = %{ - modalities: [:graph], - proof: [%{proofType: "FRESHNESS", contractName: "entity-001"}] - } - - assert {:error, {:modality_mismatch, msg}} = VCLTypeChecker.typecheck(ast) - assert msg =~ "temporal" - end - - test "rejects query with no proof specs" do - ast = %{ - modalities: [:graph], - proof: [] - } - - assert {:error, {:missing_proof, _}} = VCLTypeChecker.typecheck(ast) - end - - test "rejects query with nil proof" do - ast = %{ - modalities: [:graph], - proof: nil - } - - assert {:error, {:missing_proof, _}} = VCLTypeChecker.typecheck(ast) - end - end - - # =========================================================================== - # Obligation structure - # =========================================================================== - - describe "obligation structure" do - test "obligations include witness fields" do - ast = %{ - modalities: [:graph], - proof: [%{proofType: "EXISTENCE", contractName: "entity-001"}] - } - - {:ok, info} = VCLTypeChecker.typecheck(ast) - obligation = Enum.at(info.proof_obligations, 0) - - assert is_list(obligation.witness_fields) - assert "octad_id" in obligation.witness_fields - assert "timestamp" in obligation.witness_fields - end - - test "obligations include circuit names" do - ast = %{ - modalities: [:graph], - proof: [%{proofType: "EXISTENCE", contractName: "entity-001"}] - } - - {:ok, info} = VCLTypeChecker.typecheck(ast) - obligation = Enum.at(info.proof_obligations, 0) - - assert obligation.circuit == "existence-proof-v1" - end - - test "obligations include time estimates" do - ast = %{ - modalities: [:graph, :provenance], - proof: [ - %{proofType: "EXISTENCE", contractName: "entity-001"}, - %{proofType: "PROVENANCE", contractName: "entity-001"} - ] - } - - {:ok, info} = VCLTypeChecker.typecheck(ast) - - assert info.total_estimated_ms > 0 - assert info.total_estimated_ms == - Enum.at(info.proof_obligations, 0).estimated_time_ms + - Enum.at(info.proof_obligations, 1).estimated_time_ms - end - - test "provenance circuit is provenance-proof-v1" do - ast = %{ - modalities: [:provenance], - proof: [%{proofType: "PROVENANCE", contractName: "entity-001"}] - } - - {:ok, info} = VCLTypeChecker.typecheck(ast) - assert Enum.at(info.proof_obligations, 0).circuit == "provenance-proof-v1" - end - - test "integrity circuit is integrity-proof-v1" do - ast = %{ - modalities: [:semantic], - proof: [%{proofType: "INTEGRITY", contractName: "contract-001"}] - } - - {:ok, info} = VCLTypeChecker.typecheck(ast) - assert Enum.at(info.proof_obligations, 0).circuit == "integrity-proof-v1" - end - end - - # =========================================================================== - # Composition strategy - # =========================================================================== - - describe "composition strategy" do - test "single proof is independent" do - ast = %{ - modalities: [:graph], - proof: [%{proofType: "EXISTENCE", contractName: "entity-001"}] - } - - {:ok, info} = VCLTypeChecker.typecheck(ast) - assert info.composition_strategy == :independent - assert info.is_parallelizable == true - end - - test "provenance + citation is sequential" do - ast = %{ - modalities: [:all], - proof: [ - %{proofType: "PROVENANCE", contractName: "entity-001"}, - %{proofType: "CITATION", contractName: "ref-001"} - ] - } - - {:ok, info} = VCLTypeChecker.typecheck(ast) - assert info.composition_strategy == :sequential - assert info.is_parallelizable == false - end - - test "existence + access is conjunction (parallelizable)" do - ast = %{ - modalities: [:graph], - proof: [ - %{proofType: "EXISTENCE", contractName: "entity-001"}, - %{proofType: "ACCESS", contractName: "entity-001"} - ] - } - - {:ok, info} = VCLTypeChecker.typecheck(ast) - assert info.composition_strategy == :conjunction - end - end - - # =========================================================================== - # All 11 proof types - # =========================================================================== - - describe "all proof types" do - @tag timeout: :infinity - - test "all known proof types are accepted with correct modalities" do - test_cases = [ - {:existence, [:graph], "entity-001"}, - {:integrity, [:semantic], "contract-001"}, - {:consistency, [:graph, :vector], "entity-001"}, - {:provenance, [:provenance], "entity-001"}, - {:freshness, [:temporal], "entity-001"}, - {:access, [:graph], "entity-001"}, - {:citation, [:document], "ref-001"}, - {:custom, [:graph], "my-circuit"}, - {:zkp, [:graph], "claim-001"}, - {:proven, [:semantic], "cert-001"}, - {:sanctify, [:semantic], "contract-001"} - ] - - for {type, modalities, contract} <- test_cases do - proof_type_str = type |> Atom.to_string() |> String.upcase() - ast = %{ - modalities: modalities, - proof: [%{proofType: proof_type_str, contractName: contract}] - } - - result = VCLTypeChecker.typecheck(ast) - assert {:ok, info} = result, - "#{proof_type_str} should be accepted, got: #{inspect(result)}" - assert Enum.at(info.proof_obligations, 0).type == type - end - end - end -end diff --git a/verisimdb/elixir-orchestration/test/verisim/telemetry_test.exs b/verisimdb/elixir-orchestration/test/verisim/telemetry_test.exs deleted file mode 100644 index 43c9122e..00000000 --- a/verisimdb/elixir-orchestration/test/verisim/telemetry_test.exs +++ /dev/null @@ -1,265 +0,0 @@ -# SPDX-License-Identifier: MPL-2.0 - -defmodule VeriSim.TelemetryTest do - @moduledoc """ - Tests for the telemetry collector and reporter. - - Verifies that: - 1. Collector aggregates counters and distributions correctly - 2. Reporter produces structured product insights - 3. Privacy guarantees are maintained (no PII, aggregate only) - 4. Telemetry opt-in gating works - """ - - use ExUnit.Case, async: false - - alias VeriSim.Telemetry.{Collector, Reporter} - - setup do - # Ensure collector is started and reset before each test. - case GenServer.whereis(Collector) do - nil -> - {:ok, _pid} = Collector.start_link() - :ok - _pid -> - Collector.reset() - :ok - end - - # Temporarily enable telemetry for tests. - prev = Application.get_env(:verisim, :telemetry_enabled, false) - Application.put_env(:verisim, :telemetry_enabled, true) - - on_exit(fn -> - Application.put_env(:verisim, :telemetry_enabled, prev) - Collector.reset() - end) - - :ok - end - - # ── Collector Tests ───────────────────────────────────────────────────── - - describe "Collector" do - test "increment/1 increases a counter by 1" do - Collector.increment(:test_counter) - Collector.increment(:test_counter) - Collector.increment(:test_counter) - - snapshot = Collector.snapshot() - assert snapshot[:test_counter] == 3 - end - - test "increment/2 increases a counter by specified amount" do - Collector.increment(:batch_counter, 10) - Collector.increment(:batch_counter, 5) - - snapshot = Collector.snapshot() - assert snapshot[:batch_counter] == 15 - end - - test "increment_map/3 tracks sub-key counters" do - Collector.increment_map(:modality_usage, :graph) - Collector.increment_map(:modality_usage, :graph) - Collector.increment_map(:modality_usage, :vector) - Collector.increment_map(:modality_usage, :vector, 3) - - snapshot = Collector.snapshot() - assert snapshot[{:modality_usage, :graph}] == 2 - assert snapshot[{:modality_usage, :vector}] == 4 - end - - test "record_distribution/2 tracks count, sum, min, max" do - Collector.record_distribution(:query_duration, 10.0) - Collector.record_distribution(:query_duration, 20.0) - Collector.record_distribution(:query_duration, 5.0) - - # Give GenServer time to process casts. - Process.sleep(50) - - snapshot = Collector.snapshot() - assert snapshot[{:query_duration, :count}] == 3 - # Sum stored as integer * 1000 for precision - assert snapshot[{:query_duration, :sum}] == 35_000 - assert snapshot[{:query_duration, :min}] == 5_000 - assert snapshot[{:query_duration, :max}] == 20_000 - end - - test "reset/0 clears all metrics" do - Collector.increment(:counter_to_reset, 42) - assert Collector.snapshot()[:counter_to_reset] == 42 - - Collector.reset() - assert Collector.snapshot() == %{} - end - - test "snapshot/0 returns empty map when no metrics collected" do - Collector.reset() - assert Collector.snapshot() == %{} - end - - test "enabled?/0 respects application config" do - assert Collector.enabled?() == true - - Application.put_env(:verisim, :telemetry_enabled, false) - # Also need to clear env var override - prev_env = System.get_env("VERISIM_TELEMETRY") - System.delete_env("VERISIM_TELEMETRY") - - refute Collector.enabled?() - - Application.put_env(:verisim, :telemetry_enabled, true) - if prev_env, do: System.put_env("VERISIM_TELEMETRY", prev_env) - end - end - - # ── Reporter Tests ────────────────────────────────────────────────────── - - describe "Reporter" do - test "report/0 returns structured map with all sections" do - report = Reporter.report() - - assert is_map(report.meta) - assert is_map(report.modality_heatmap) - assert is_map(report.query_patterns) - assert is_map(report.performance) - assert is_map(report.drift) - assert is_map(report.federation) - assert is_map(report.proof_types) - assert is_map(report.entities) - end - - test "report/0 includes privacy notice" do - report = Reporter.report() - assert String.contains?(report.meta.privacy_notice, "aggregate metrics only") - assert String.contains?(report.meta.privacy_notice, "No query content") - end - - test "modality_heatmap/0 shows all 8 octad modalities" do - heatmap = Reporter.modality_heatmap() - - assert Map.has_key?(heatmap.counts, :graph) - assert Map.has_key?(heatmap.counts, :vector) - assert Map.has_key?(heatmap.counts, :tensor) - assert Map.has_key?(heatmap.counts, :semantic) - assert Map.has_key?(heatmap.counts, :document) - assert Map.has_key?(heatmap.counts, :temporal) - assert Map.has_key?(heatmap.counts, :provenance) - assert Map.has_key?(heatmap.counts, :spatial) - end - - test "modality_heatmap/0 calculates percentages" do - Collector.increment_map(:modality_usage, :graph, 3) - Collector.increment_map(:modality_usage, :vector, 7) - - heatmap = Reporter.modality_heatmap() - - assert heatmap.total_modality_queries == 10 - assert heatmap.percentages[:graph] == 30.0 - assert heatmap.percentages[:vector] == 70.0 - assert heatmap.most_used == :vector - end - - test "query_patterns/0 tracks statement types" do - Collector.increment(:query_count, 5) - Collector.increment(:query_error_count, 1) - Collector.increment_map(:query_pattern, "SELECT", 3) - Collector.increment_map(:query_pattern, "INSERT", 2) - - patterns = Reporter.query_patterns() - - assert patterns.total_queries == 5 - assert patterns.error_count == 1 - assert patterns.error_rate == 20.0 - assert patterns.by_type["SELECT"] == 3 - assert patterns.by_type["INSERT"] == 2 - end - - test "performance_summary/0 reports duration statistics" do - Collector.record_distribution(:query_duration, 10.0) - Collector.record_distribution(:query_duration, 30.0) - Process.sleep(50) - - perf = Reporter.performance_summary() - - assert perf.query_count == 2 - assert perf.avg_duration_ms == 20.0 - assert perf.min_duration_ms == 10.0 - assert perf.max_duration_ms == 30.0 - end - - test "drift_report/0 tracks drift and normalisation" do - Collector.increment(:drift_detected_count, 10) - Collector.increment(:normalise_count, 8) - Collector.increment(:normalise_success_count, 7) - Collector.increment_map(:drift_modality_breakdown, :semantic, 5) - Collector.increment_map(:drift_modality_breakdown, :graph, 3) - - drift = Reporter.drift_report() - - assert drift.drift_detected_count == 10 - assert drift.normalise_attempts == 8 - assert drift.normalise_success_count == 7 - assert drift.normalise_success_rate == 87.5 - assert drift.modality_breakdown[:semantic] == 5 - assert drift.most_drifted == :semantic - end - - test "proof_type_usage/0 tracks VCL-UT adoption" do - Collector.increment_map(:proof_type_usage, "EXISTENCE", 3) - Collector.increment_map(:proof_type_usage, "INTEGRITY", 2) - - proofs = Reporter.proof_type_usage() - - assert proofs.total_proofs == 5 - assert proofs.vcl_dt_active == true - assert proofs.by_type["EXISTENCE"] == 3 - end - - test "entity_summary/0 tracks creates and deletes" do - Collector.increment(:entity_created_count, 100) - Collector.increment(:entity_deleted_count, 5) - - entities = Reporter.entity_summary() - - assert entities.created == 100 - assert entities.deleted == 5 - end - - test "report_json/0 returns valid JSON string" do - json = Reporter.report_json() - assert {:ok, decoded} = Jason.decode(json) - assert is_map(decoded["meta"]) - assert is_map(decoded["modality_heatmap"]) - end - end - - # ── Privacy Tests ─────────────────────────────────────────────────────── - - describe "Privacy" do - test "report never contains query content" do - Collector.increment(:query_count, 1) - Collector.increment_map(:query_pattern, "SELECT", 1) - - json = Reporter.report_json() - - # The report should never contain actual query strings. - refute String.contains?(json, "SELECT GRAPH") - refute String.contains?(json, "INSERT INTO") - refute String.contains?(json, "entity-") - end - - test "snapshot keys are only aggregate identifiers" do - Collector.increment(:query_count) - Collector.increment_map(:modality_usage, :graph) - - snapshot = Collector.snapshot() - - # All keys should be atoms or {atom, atom/string} tuples — never raw strings. - Enum.each(snapshot, fn {key, _value} -> - assert is_atom(key) or is_tuple(key), - "Expected atom or tuple key, got: #{inspect(key)}" - end) - end - end -end diff --git a/verisimdb/elixir-orchestration/test/verisim_test.exs b/verisimdb/elixir-orchestration/test/verisim_test.exs deleted file mode 100644 index 4c08088b..00000000 --- a/verisimdb/elixir-orchestration/test/verisim_test.exs +++ /dev/null @@ -1,10 +0,0 @@ -# SPDX-License-Identifier: MPL-2.0 - -defmodule VeriSimTest do - use ExUnit.Case - doctest VeriSim - - test "version returns string" do - assert is_binary(VeriSim.version()) - end -end diff --git a/verisimdb/examples/README.adoc b/verisimdb/examples/README.adoc deleted file mode 100644 index 2a4c3094..00000000 --- a/verisimdb/examples/README.adoc +++ /dev/null @@ -1,59 +0,0 @@ -// SPDX-License-Identifier: CC-BY-SA-4.0 - -= VeriSimDB Examples -:toc: - -== Purpose - -This directory contains **example and demo data only**. - -It exists to help developers understand VeriSimDB's data model, test their -client SDK integrations, and explore the octad concept. The data is fictional -(academic papers, authors, venues) and is intended for learning. - -== IMPORTANT: This Is Not a Data Store - -[CAUTION] -==== -**Do NOT store real application data in this directory or in this repo.** - -If your project uses VeriSimDB, you must run your **own dedicated instance** -with its own port, data directory, and container volume. See the instance -policy in `0-AI-MANIFEST.a2ml` and `.claude/CLAUDE.md` for full details. - -* Copy the client SDK from `connectors/clients//` into your project -* Add a `verisimdb.Containerfile` to your project's container stack -* Pick a unique port (not 8080) and a named volume for your data -* Never point your app at this repo's dev server -==== - -== Files - -[cols="1,3"] -|=== -| File | Description - -| `sample-data/seed.json` -| 8 fictional octad entities (papers, authors, venues) demonstrating all 8 - modalities. Use for SDK testing, VCL query practice, and understanding the - hexad/octad data shape. -|=== - -== Loading Example Data - -[source,bash] ----- -# Start a LOCAL, EPHEMERAL VeriSimDB instance for testing -cargo run --manifest-path rust-core/verisim-api/Cargo.toml & - -# Load the seed data -curl -X POST http://localhost:8080/api/v1/hexads/bulk \ - -H "Content-Type: application/json" \ - -d @examples/sample-data/seed.json - -# Query it -curl http://localhost:8080/api/v1/hexads | jq . ----- - -This creates an **in-memory** instance that loses all data when stopped. -That is intentional — example data is disposable. diff --git a/verisimdb/examples/SafeDOMExample.res b/verisimdb/examples/SafeDOMExample.res deleted file mode 100644 index e5c90460..00000000 --- a/verisimdb/examples/SafeDOMExample.res +++ /dev/null @@ -1,109 +0,0 @@ -// SPDX-License-Identifier: MPL-2.0 -// Example: Using SafeDOM for formally verified DOM mounting - -open SafeDOM - -// Example 1: Basic mounting with error handling -let mountApp = () => { - mountSafe( - "#app", - "

Hello, World!

Mounted safely with proofs.

", - ~onSuccess=el => { - Console.log("✓ App mounted successfully!") - Console.log("Element:", el) - }, - ~onError=err => { - Console.error("✗ Mount failed:", err) - } - ) -} - -// Example 2: Wait for DOM ready before mounting -let mountWhenDOMReady = () => { - mountWhenReady( - "#app", - "

App Title

", - ~onSuccess=_ => Console.log("✓ Mounted after DOM ready"), - ~onError=err => Console.error("✗ Failed:", err) - ) -} - -// Example 3: Batch mounting (atomic - all or nothing) -let mountMultiple = () => { - let specs = [ - {selector: "#header", html: "

Site Title

"}, - {selector: "#nav", html: ""}, - {selector: "#main", html: "

Content here

"}, - {selector: "#footer", html: "
© 2026
"} - ] - - switch mountBatch(specs) { - | Ok(elements) => { - Console.log(`✓ Successfully mounted ${Array.length(elements)} elements`) - elements->Array.forEach(el => Console.log(" -", el)) - } - | Error(err) => { - Console.error("✗ Batch mount failed:", err) - Console.error(" (None were mounted - atomic operation)") - } - } -} - -// Example 4: Explicit validation before mounting -let mountWithValidation = () => { - // Validate selector first - switch ProvenSelector.validate("#my-app") { - | Error(e) => Console.error(`Invalid selector: ${e}`) - | Ok(validSelector) => { - // Validate HTML - switch ProvenHTML.validate("
Content
") { - | Error(e) => Console.error(`Invalid HTML: ${e}`) - | Ok(validHtml) => { - // Now mount with proven safety - switch mount(validSelector, validHtml) { - | Mounted(el) => Console.log("✓ Mounted with validated inputs:", el) - | MountPointNotFound(s) => Console.error(`✗ Element not found: ${s}`) - | InvalidSelector(_) => Console.error("Impossible - already validated") - | InvalidHTML(_) => Console.error("Impossible - already validated") - } - } - } - } -} - -// Example 5: Integration with TEA -module MyApp = { - type model = {message: string} - type msg = NoOp - - let init = () => {message: "Hello from TEA"} - let update = (model, _msg) => model - let view = model => `

${model.message}

` -} - -let mountTEAApp = () => { - let model = MyApp.init() - let html = MyApp.view(model) - - mountWhenReady( - "#tea-app", - html, - ~onSuccess=el => { - Console.log("✓ TEA app mounted") - // Set up event handlers, subscriptions here - }, - ~onError=err => Console.error(`✗ TEA mount failed: ${err}`) - ) -} - -// Entry point -let main = () => { - Console.log("SafeDOM Examples") - Console.log("================\n") - - // Choose which example to run - mountWhenDOMReady() // Run on DOM ready -} - -// Auto-execute when module loads -main() diff --git a/verisimdb/examples/load-sample-data.sh b/verisimdb/examples/load-sample-data.sh deleted file mode 100755 index db4f64cd..00000000 --- a/verisimdb/examples/load-sample-data.sh +++ /dev/null @@ -1,168 +0,0 @@ -#!/usr/bin/env bash -# SPDX-License-Identifier: MPL-2.0 -# -# load-sample-data.sh — Load VeriSimDB sample data and run example VCL queries. -# -# This script loads 50 sample hexad entities from seed.json into a running -# VeriSimDB instance, then executes all VCL example queries from the -# vcl-queries/ directory. It is intended for demonstration and testing. -# -# The sample data tells the story of a research institution ecosystem: -# 10 academic papers, 10 researchers, 10 organisations, 10 datasets, -# and 10 events — all cross-linked via graph edges and with 10 entities -# containing intentional cross-modal drift for drift detection testing. -# -# Usage: -# ./load-sample-data.sh [--api-url URL] -# -# Prerequisites: -# - VeriSimDB rust-core running on port 8080 (or specify --api-url) -# - jq installed (for pretty-printing responses) -# - curl installed -# -# Examples: -# ./load-sample-data.sh -# ./load-sample-data.sh --api-url http://192.168.1.10:8080/api/v1 - -set -euo pipefail - -# --------------------------------------------------------------------------- -# Argument parsing -# --------------------------------------------------------------------------- - -API_URL="http://localhost:8080/api/v1" - -while [[ $# -gt 0 ]]; do - case "$1" in - --api-url) - API_URL="$2" - shift 2 - ;; - *) - # Positional fallback for bare URL argument - API_URL="$1" - shift - ;; - esac -done - -SCRIPT_DIR="$(cd "$(dirname "$0")" && pwd)" -DATA_FILE="${SCRIPT_DIR}/sample-data/seed.json" - -echo "=== VeriSimDB Sample Data Loader ===" -echo "API URL: ${API_URL}" -echo "" - -# --------------------------------------------------------------------------- -# Prerequisite checks -# --------------------------------------------------------------------------- - -command -v curl >/dev/null 2>&1 || { echo "ERROR: curl is required but not installed."; exit 1; } -command -v jq >/dev/null 2>&1 || { echo "ERROR: jq is required but not installed."; exit 1; } - -if [[ ! -f "$DATA_FILE" ]]; then - echo "ERROR: Sample data file not found at ${DATA_FILE}" - echo "Expected: examples/sample-data/seed.json" - exit 1 -fi - -# --------------------------------------------------------------------------- -# Server health check -# --------------------------------------------------------------------------- - -echo "Checking server health..." -HEALTH=$(curl -sf "${API_URL}/health" 2>/dev/null || echo '{"status":"unreachable"}') -STATUS=$(echo "$HEALTH" | jq -r '.status // "unknown"') - -if [ "$STATUS" != "ok" ] && [ "$STATUS" != "healthy" ]; then - echo "ERROR: Server not healthy at ${API_URL}. Status: ${STATUS}" - echo "" - echo "Start VeriSimDB first:" - echo " cd rust-core && cargo run" - echo "" - echo "Or specify a different URL:" - echo " ./load-sample-data.sh --api-url http://HOST:PORT/api/v1" - exit 1 -fi - -echo "Server healthy." -echo "" - -# --------------------------------------------------------------------------- -# Load entities -# --------------------------------------------------------------------------- - -echo "Loading 50 sample entities from ${DATA_FILE}..." -echo "" - -LOADED=0 -FAILED=0 - -# Process each entity from the JSON array -jq -c '.[]' "$DATA_FILE" | while read -r entity; do - ID=$(echo "$entity" | jq -r '.id') - CATEGORY=$(echo "$entity" | jq -r '.category') - TITLE=$(echo "$entity" | jq -r '.document.title // "(untitled)"') - - RESPONSE=$(curl -sf -X POST "${API_URL}/hexads" \ - -H "Content-Type: application/json" \ - -d "$entity" 2>/dev/null || echo "FAILED") - - if [ "$RESPONSE" = "FAILED" ]; then - echo " FAIL: ${ID} [${CATEGORY}] ${TITLE}" - FAILED=$((FAILED + 1)) - else - echo " OK: ${ID} [${CATEGORY}] ${TITLE}" - LOADED=$((LOADED + 1)) - fi -done - -echo "" -echo "Loading complete. Check output above for individual results." -echo "" - -# --------------------------------------------------------------------------- -# Run VCL example queries -# --------------------------------------------------------------------------- - -echo "=== Running VCL Example Queries ===" -echo "" - -VCL_DIR="${SCRIPT_DIR}/vcl-queries" - -if [[ ! -d "$VCL_DIR" ]]; then - echo "WARNING: VCL queries directory not found at ${VCL_DIR}" - echo "Skipping query execution." - exit 0 -fi - -for vcl_file in "${VCL_DIR}"/*.vcl; do - [[ -f "$vcl_file" ]] || continue - - BASENAME=$(basename "$vcl_file") - - # Extract the query: strip comment lines (--) and collapse whitespace - QUERY=$(grep -v '^--' "$vcl_file" | tr '\n' ' ' | sed 's/ */ /g' | xargs) - - if [[ -z "$QUERY" ]]; then - echo "--- ${BASENAME} --- (empty query, skipping)" - echo "" - continue - fi - - echo "--- ${BASENAME} ---" - echo "Query: ${QUERY}" - echo "" - - # Escape double quotes in the query for JSON payload - ESCAPED_QUERY=$(echo "$QUERY" | sed 's/"/\\"/g') - - RESULT=$(curl -sf -X POST "${API_URL}/vcl/execute" \ - -H "Content-Type: application/json" \ - -d "{\"query\": \"${ESCAPED_QUERY}\"}" 2>/dev/null || echo '{"error": "query execution failed or server unreachable"}') - - echo "$RESULT" | jq '.' 2>/dev/null || echo "$RESULT" - echo "" -done - -echo "=== Done ===" diff --git a/verisimdb/examples/sample-data/seed.json b/verisimdb/examples/sample-data/seed.json deleted file mode 100644 index fd5baa53..00000000 --- a/verisimdb/examples/sample-data/seed.json +++ /dev/null @@ -1,2210 +0,0 @@ -[ - { - "id": "entity-001", - "category": "paper", - "document": { - "title": "VeriSimDB: A Multimodal Database with Drift Detection", - "content": "This paper introduces VeriSimDB, a novel database system that unifies eight distinct data modalities — document, graph, vector, semantic, temporal, provenance, spatial, and tensor — into a single octad abstraction. We present a drift detection mechanism that identifies cross-modal inconsistencies in real time, enabling self-healing data pipelines. Our evaluation on the PolyBench benchmark demonstrates sub-millisecond drift scoring with 97.3% precision.", - "tags": ["multimodal", "drift detection", "database", "octad", "VeriSimDB"] - }, - "graph": { - "edges": [ - {"target": "entity-011", "relation": "authored_by"}, - {"target": "entity-012", "relation": "authored_by"}, - {"target": "entity-031", "relation": "evaluated_on"}, - {"target": "entity-041", "relation": "presented_at"}, - {"target": "entity-021", "relation": "affiliated_with"} - ] - }, - "vector": { - "embedding": [0.91, 0.23, 0.67, 0.44, 0.85, 0.12, 0.56, 0.78] - }, - "semantic": { - "types": ["https://schema.org/ScholarlyArticle", "https://schema.org/TechArticle"], - "properties": { - "keywords": ["multimodal databases", "drift detection"], - "doi": "10.1234/verisimdb-2025", - "pageCount": 14 - } - }, - "temporal": { - "created": "2025-03-15T09:00:00Z", - "modified": "2026-02-10T14:22:00Z", - "version": 3 - }, - "provenance": { - "origin": "Imperial College London — Data Systems Lab", - "actor": "entity-011", - "chain": ["draft-v1", "peer-review-revision", "camera-ready"] - }, - "spatial": { - "lat": 51.4988, - "lon": -0.1749, - "label": "Imperial College London, South Kensington" - }, - "intentional_drift": false - }, - { - "id": "entity-002", - "category": "paper", - "document": { - "title": "Formal Verification of Cross-Modal Consistency in Polystore Systems", - "content": "We present a dependent-type framework for verifying consistency invariants across heterogeneous data modalities. By encoding modal relationships as type-level constraints in Idris2, we can statically guarantee that graph edges, vector embeddings, and semantic annotations remain coherent after arbitrary transformation pipelines. The framework is evaluated against VeriSimDB's octad model.", - "tags": ["formal verification", "dependent types", "Idris2", "consistency", "polystore"] - }, - "graph": { - "edges": [ - {"target": "entity-013", "relation": "authored_by"}, - {"target": "entity-014", "relation": "authored_by"}, - {"target": "entity-032", "relation": "evaluated_on"}, - {"target": "entity-042", "relation": "presented_at"}, - {"target": "entity-022", "relation": "affiliated_with"} - ] - }, - "vector": { - "embedding": [0.88, 0.31, 0.72, 0.39, 0.81, 0.15, 0.61, 0.73] - }, - "semantic": { - "types": ["https://schema.org/ScholarlyArticle"], - "properties": { - "keywords": ["formal verification", "dependent types", "cross-modal consistency"], - "doi": "10.1234/formal-polystore-2025", - "pageCount": 12 - } - }, - "temporal": { - "created": "2025-06-01T10:30:00Z", - "modified": "2026-01-20T11:15:00Z", - "version": 2 - }, - "provenance": { - "origin": "University of Edinburgh — Laboratory for Foundations of Computer Science", - "actor": "entity-013", - "chain": ["initial-submission", "revision-1"] - }, - "spatial": { - "lat": 55.9445, - "lon": -3.1872, - "label": "University of Edinburgh, Informatics Forum" - }, - "intentional_drift": false - }, - { - "id": "entity-003", - "category": "paper", - "document": { - "title": "Octad Embeddings: Unified Vector Representations for Multi-Modal Data", - "content": "We propose a novel embedding technique that produces a single dense vector from all eight modalities of a octad record. Unlike concatenation or late-fusion approaches, our method uses attention-based cross-modal alignment to produce compact 768-dimensional embeddings that capture inter-modal relationships. Evaluation on the PolyBench and ModalMix benchmarks shows 12% improvement over baseline approaches.", - "tags": ["embeddings", "vector", "attention", "multi-modal", "fusion"] - }, - "graph": { - "edges": [ - {"target": "entity-011", "relation": "authored_by"}, - {"target": "entity-015", "relation": "authored_by"}, - {"target": "entity-031", "relation": "evaluated_on"}, - {"target": "entity-033", "relation": "evaluated_on"}, - {"target": "entity-043", "relation": "presented_at"} - ] - }, - "vector": { - "embedding": [0.85, 0.19, 0.74, 0.52, 0.79, 0.08, 0.63, 0.81] - }, - "semantic": { - "types": ["https://schema.org/ScholarlyArticle"], - "properties": { - "keywords": ["octad embeddings", "cross-modal fusion", "vector representations"], - "doi": "10.1234/octad-embed-2025", - "pageCount": 10 - } - }, - "temporal": { - "created": "2025-07-20T08:45:00Z", - "modified": "2026-02-01T16:30:00Z", - "version": 4 - }, - "provenance": { - "origin": "Imperial College London — Data Systems Lab", - "actor": "entity-011", - "chain": ["draft-v1", "internal-review", "conference-submission", "camera-ready"] - }, - "spatial": { - "lat": 51.4988, - "lon": -0.1749, - "label": "Imperial College London, South Kensington" - }, - "intentional_drift": true, - "drift_note": "Vector embedding is from draft-v1 and was never re-computed after the camera-ready text changes; the document content diverged from the embedding in version 3." - }, - { - "id": "entity-004", - "category": "paper", - "document": { - "title": "Provenance Graphs for Reproducible Data Science Pipelines", - "content": "Tracking data lineage through complex transformation pipelines is critical for scientific reproducibility. We introduce ProvenanceDAG, a directed acyclic graph representation that captures every transformation applied to a data entity across heterogeneous storage backends. Our system integrates with VeriSimDB's provenance modality and the W3C PROV-O ontology to provide end-to-end lineage tracking with sub-second query response times.", - "tags": ["provenance", "reproducibility", "data lineage", "DAG", "PROV-O"] - }, - "graph": { - "edges": [ - {"target": "entity-016", "relation": "authored_by"}, - {"target": "entity-012", "relation": "authored_by"}, - {"target": "entity-034", "relation": "evaluated_on"}, - {"target": "entity-044", "relation": "presented_at"}, - {"target": "entity-023", "relation": "affiliated_with"} - ] - }, - "vector": { - "embedding": [0.72, 0.41, 0.58, 0.33, 0.90, 0.22, 0.47, 0.66] - }, - "semantic": { - "types": ["https://schema.org/ScholarlyArticle"], - "properties": { - "keywords": ["provenance", "reproducibility", "lineage graphs"], - "doi": "10.1234/provdag-2025", - "pageCount": 11 - } - }, - "temporal": { - "created": "2025-04-10T13:00:00Z", - "modified": "2026-01-05T09:45:00Z", - "version": 2 - }, - "provenance": { - "origin": "University College London — Information Studies", - "actor": "entity-016", - "chain": ["initial-draft", "final-version"] - }, - "spatial": { - "lat": 51.5246, - "lon": -0.1340, - "label": "University College London, Gower Street" - }, - "intentional_drift": false - }, - { - "id": "entity-005", - "category": "paper", - "document": { - "title": "Semantic Drift in Knowledge Graphs: Detection and Repair Strategies", - "content": "Knowledge graphs evolve over time, and their semantic annotations can drift away from the entities they describe. We formalise the notion of semantic drift using information-theoretic measures and propose three repair strategies: anchor-based realignment, consensus-driven correction, and proof-carrying patches. Experiments on DBpedia and Wikidata subsets show our detection achieves 94% recall.", - "tags": ["semantic drift", "knowledge graphs", "repair", "information theory"] - }, - "graph": { - "edges": [ - {"target": "entity-013", "relation": "authored_by"}, - {"target": "entity-017", "relation": "authored_by"}, - {"target": "entity-035", "relation": "evaluated_on"}, - {"target": "entity-041", "relation": "presented_at"}, - {"target": "entity-024", "relation": "affiliated_with"} - ] - }, - "vector": { - "embedding": [0.79, 0.35, 0.62, 0.48, 0.83, 0.18, 0.55, 0.71] - }, - "semantic": { - "types": ["https://schema.org/ScholarlyArticle", "https://schema.org/Article"], - "properties": { - "keywords": ["semantic drift", "knowledge graphs", "repair"], - "doi": "10.1234/sem-drift-2025", - "pageCount": 13 - } - }, - "temporal": { - "created": "2025-05-22T11:00:00Z", - "modified": "2026-02-15T10:00:00Z", - "version": 3 - }, - "provenance": { - "origin": "University of Oxford — Department of Computer Science", - "actor": "entity-013", - "chain": ["draft", "revision", "accepted"] - }, - "spatial": { - "lat": 51.7598, - "lon": -1.2584, - "label": "University of Oxford, Wolfson Building" - }, - "intentional_drift": true, - "drift_note": "Semantic types list includes 'https://schema.org/Article' which is a supertype of ScholarlyArticle; the graph edges reference entity-024 (Oxford) but the provenance says Edinburgh LFCS — affiliation is inconsistent between modalities." - }, - { - "id": "entity-006", - "category": "paper", - "document": { - "title": "Tensor Decomposition for Anomaly Detection in Streaming Octad Data", - "content": "Real-time anomaly detection in multimodal data streams requires efficient tensor representations. We present HexTensor, a streaming tensor decomposition framework that processes VeriSimDB octad records as 8-mode tensors. By applying Tucker decomposition incrementally, we detect anomalous entities within 50ms of ingestion. Our approach scales to 100K entities per second on commodity hardware.", - "tags": ["tensor", "anomaly detection", "streaming", "Tucker decomposition"] - }, - "graph": { - "edges": [ - {"target": "entity-018", "relation": "authored_by"}, - {"target": "entity-011", "relation": "authored_by"}, - {"target": "entity-036", "relation": "evaluated_on"}, - {"target": "entity-042", "relation": "presented_at"}, - {"target": "entity-025", "relation": "affiliated_with"} - ] - }, - "vector": { - "embedding": [0.82, 0.27, 0.69, 0.51, 0.77, 0.14, 0.59, 0.84] - }, - "semantic": { - "types": ["https://schema.org/ScholarlyArticle"], - "properties": { - "keywords": ["tensor decomposition", "anomaly detection", "streaming"], - "doi": "10.1234/hextensor-2026", - "pageCount": 9 - } - }, - "temporal": { - "created": "2025-09-01T14:00:00Z", - "modified": "2026-02-20T08:30:00Z", - "version": 2 - }, - "provenance": { - "origin": "King's College London — Informatics Department", - "actor": "entity-018", - "chain": ["preprint", "camera-ready"] - }, - "spatial": { - "lat": 51.5115, - "lon": -0.1160, - "label": "King's College London, Strand Campus" - }, - "intentional_drift": false - }, - { - "id": "entity-007", - "category": "paper", - "document": { - "title": "R-Tree Indexing Strategies for Spatial-Temporal Octad Queries", - "content": "Spatial-temporal queries over multimodal data require specialised indexing. We evaluate four R-tree variants — R*-tree, Hilbert R-tree, STR-packed R-tree, and our novel Octad R-tree — for combined spatial and temporal range queries on VeriSimDB datasets. The Octad R-tree achieves 3.2x throughput improvement by co-locating spatial and temporal dimensions in the same tree nodes.", - "tags": ["R-tree", "spatial-temporal", "indexing", "query optimisation"] - }, - "graph": { - "edges": [ - {"target": "entity-019", "relation": "authored_by"}, - {"target": "entity-015", "relation": "authored_by"}, - {"target": "entity-037", "relation": "evaluated_on"}, - {"target": "entity-045", "relation": "presented_at"}, - {"target": "entity-026", "relation": "affiliated_with"} - ] - }, - "vector": { - "embedding": [0.76, 0.44, 0.55, 0.38, 0.88, 0.20, 0.50, 0.69] - }, - "semantic": { - "types": ["https://schema.org/ScholarlyArticle"], - "properties": { - "keywords": ["R-tree", "spatial indexing", "temporal indexing"], - "doi": "10.1234/octad-rtree-2025", - "pageCount": 10 - } - }, - "temporal": { - "created": "2025-08-15T10:00:00Z", - "modified": "2026-01-30T17:00:00Z", - "version": 5 - }, - "provenance": { - "origin": "Queen Mary University of London — School of EECS", - "actor": "entity-019", - "chain": ["draft-v1", "internal-review", "submitted", "revised", "accepted"] - }, - "spatial": { - "lat": 51.5229, - "lon": -0.0408, - "label": "Queen Mary University of London, Mile End" - }, - "intentional_drift": true, - "drift_note": "Temporal version is 5 but provenance chain has exactly 5 entries — this is correct. However, the spatial coordinates point to QMUL but the paper's affiliated_with edge points to entity-026 (Turing Institute), not QMUL. The affiliation and spatial location are inconsistent." - }, - { - "id": "entity-008", - "category": "paper", - "document": { - "title": "VCL: A Query Language for Verified Multimodal Databases", - "content": "Existing query languages are designed for single-modality databases. We present VCL (VeriSim Consonance Language), a declarative language that enables cross-modal queries with optional proof obligations. VCL extends SQL syntax with PROOF clauses for existence, consistency, and provenance verification. We describe the formal semantics, the query planner architecture, and benchmark results showing acceptable overhead for proof-carrying queries.", - "tags": ["VCL", "query language", "proof-carrying", "multimodal", "SQL extension"] - }, - "graph": { - "edges": [ - {"target": "entity-011", "relation": "authored_by"}, - {"target": "entity-014", "relation": "authored_by"}, - {"target": "entity-020", "relation": "authored_by"}, - {"target": "entity-038", "relation": "evaluated_on"}, - {"target": "entity-043", "relation": "presented_at"} - ] - }, - "vector": { - "embedding": [0.93, 0.20, 0.70, 0.46, 0.87, 0.10, 0.58, 0.80] - }, - "semantic": { - "types": ["https://schema.org/ScholarlyArticle", "https://schema.org/TechArticle"], - "properties": { - "keywords": ["query language", "VCL", "proof-carrying queries"], - "doi": "10.1234/vcl-lang-2026", - "pageCount": 16 - } - }, - "temporal": { - "created": "2025-11-01T09:00:00Z", - "modified": "2026-02-25T12:00:00Z", - "version": 3 - }, - "provenance": { - "origin": "Imperial College London — Data Systems Lab", - "actor": "entity-011", - "chain": ["design-doc", "implementation-paper", "camera-ready"] - }, - "spatial": { - "lat": 51.4988, - "lon": -0.1749, - "label": "Imperial College London, South Kensington" - }, - "intentional_drift": false - }, - { - "id": "entity-009", - "category": "paper", - "document": { - "title": "Neurosymbolic Approaches to Database Query Optimisation", - "content": "Traditional query optimisers rely on cost models that struggle with cross-modal queries. We present NeuroPlanner, a neurosymbolic query optimiser that combines learned cost estimation with symbolic reasoning about modality constraints. NeuroPlanner uses a graph neural network to predict query execution costs and a constraint solver to ensure plans respect modal consistency invariants. Evaluation shows 40% latency reduction on VeriSimDB workloads.", - "tags": ["neurosymbolic", "query optimisation", "GNN", "constraint solving"] - }, - "graph": { - "edges": [ - {"target": "entity-015", "relation": "authored_by"}, - {"target": "entity-016", "relation": "authored_by"}, - {"target": "entity-039", "relation": "evaluated_on"}, - {"target": "entity-044", "relation": "presented_at"}, - {"target": "entity-027", "relation": "affiliated_with"} - ] - }, - "vector": { - "embedding": [0.87, 0.29, 0.65, 0.42, 0.82, 0.16, 0.54, 0.76] - }, - "semantic": { - "types": ["https://schema.org/ScholarlyArticle"], - "properties": { - "keywords": ["neurosymbolic", "query optimisation", "GNN"], - "doi": "10.1234/neuroplanner-2026", - "pageCount": 12 - } - }, - "temporal": { - "created": "2025-10-05T15:00:00Z", - "modified": "2026-02-18T13:45:00Z", - "version": 2 - }, - "provenance": { - "origin": "University of Cambridge — Computer Laboratory", - "actor": "entity-015", - "chain": ["workshop-paper", "full-paper"] - }, - "spatial": { - "lat": 52.2107, - "lon": 0.0917, - "label": "University of Cambridge, William Gates Building" - }, - "intentional_drift": true, - "drift_note": "The vector embedding was computed from the workshop-paper text (version 1) but the document content is the full-paper (version 2). The embedding no longer represents the current document." - }, - { - "id": "entity-010", - "category": "paper", - "document": { - "title": "Benchmark Suite for Multimodal Database Systems", - "content": "We introduce PolyBench, a comprehensive benchmark suite for evaluating multimodal database systems. PolyBench provides standardised workloads across eight modalities including document retrieval, graph traversal, vector similarity, semantic reasoning, temporal queries, provenance tracing, spatial range queries, and tensor operations. We evaluate five systems: VeriSimDB, ArangoDB, NebulaGraph, Milvus, and a baseline PostgreSQL configuration.", - "tags": ["benchmark", "PolyBench", "multimodal", "evaluation"] - }, - "graph": { - "edges": [ - {"target": "entity-012", "relation": "authored_by"}, - {"target": "entity-020", "relation": "authored_by"}, - {"target": "entity-031", "relation": "includes_dataset"}, - {"target": "entity-045", "relation": "presented_at"}, - {"target": "entity-028", "relation": "affiliated_with"} - ] - }, - "vector": { - "embedding": [0.80, 0.36, 0.60, 0.45, 0.84, 0.19, 0.52, 0.74] - }, - "semantic": { - "types": ["https://schema.org/ScholarlyArticle", "https://schema.org/Dataset"], - "properties": { - "keywords": ["benchmark", "PolyBench", "evaluation suite"], - "doi": "10.1234/polybench-2026", - "pageCount": 18 - } - }, - "temporal": { - "created": "2025-12-01T08:00:00Z", - "modified": "2026-02-22T10:30:00Z", - "version": 4 - }, - "provenance": { - "origin": "The Alan Turing Institute", - "actor": "entity-012", - "chain": ["design-phase", "implementation", "evaluation", "publication"] - }, - "spatial": { - "lat": 51.5299, - "lon": -0.1278, - "label": "The Alan Turing Institute, British Library" - }, - "intentional_drift": true, - "drift_note": "Temporal version is 4 but provenance chain has 4 entries — those match. However, semantic types include 'https://schema.org/Dataset' which is wrong: this is a paper about a benchmark, not a dataset itself. The semantic annotation has drifted to describe the benchmark artefact rather than the paper." - }, - { - "id": "entity-011", - "category": "person", - "document": { - "title": "Dr. Amara Okonkwo", - "content": "Dr. Amara Okonkwo is a Senior Lecturer in Data Systems at Imperial College London. Her research focuses on multimodal database architectures, cross-modal drift detection, and query language design. She is the principal investigator of the VeriSimDB project and lead author of the VCL specification. Prior to Imperial, she held a postdoctoral position at ETH Zurich.", - "tags": ["researcher", "databases", "Imperial College", "VeriSimDB lead"] - }, - "graph": { - "edges": [ - {"target": "entity-001", "relation": "author_of"}, - {"target": "entity-003", "relation": "author_of"}, - {"target": "entity-006", "relation": "author_of"}, - {"target": "entity-008", "relation": "author_of"}, - {"target": "entity-021", "relation": "employed_by"} - ] - }, - "vector": { - "embedding": [0.90, 0.22, 0.68, 0.43, 0.86, 0.11, 0.57, 0.79] - }, - "semantic": { - "types": ["https://schema.org/Person", "https://schema.org/Researcher"], - "properties": { - "orcid": "0000-0001-2345-6789", - "h-index": 28, - "affiliation": "Imperial College London" - } - }, - "temporal": { - "created": "2025-01-10T10:00:00Z", - "modified": "2026-02-26T09:00:00Z", - "version": 5 - }, - "provenance": { - "origin": "Imperial College staff directory", - "actor": "system-admin", - "chain": ["initial-import", "orcid-link", "publication-update", "h-index-refresh", "profile-update"] - }, - "spatial": { - "lat": 51.4988, - "lon": -0.1749, - "label": "Imperial College London, South Kensington" - }, - "intentional_drift": false - }, - { - "id": "entity-012", - "category": "person", - "document": { - "title": "Prof. Marcus Chen-Ramirez", - "content": "Prof. Marcus Chen-Ramirez is a Professor of Information Systems at The Alan Turing Institute, with a joint appointment at University College London. His research interests include data provenance, scientific reproducibility, and benchmark design for novel database systems. He co-leads the PolyBench initiative and serves on the steering committee of the VLDB Endowment.", - "tags": ["researcher", "provenance", "benchmarks", "Turing Institute"] - }, - "graph": { - "edges": [ - {"target": "entity-001", "relation": "author_of"}, - {"target": "entity-004", "relation": "author_of"}, - {"target": "entity-010", "relation": "author_of"}, - {"target": "entity-028", "relation": "employed_by"}, - {"target": "entity-023", "relation": "affiliated_with"} - ] - }, - "vector": { - "embedding": [0.78, 0.38, 0.59, 0.35, 0.89, 0.21, 0.48, 0.67] - }, - "semantic": { - "types": ["https://schema.org/Person", "https://schema.org/Researcher"], - "properties": { - "orcid": "0000-0002-3456-7890", - "h-index": 42, - "affiliation": "The Alan Turing Institute / UCL" - } - }, - "temporal": { - "created": "2025-01-10T10:00:00Z", - "modified": "2026-02-20T15:00:00Z", - "version": 3 - }, - "provenance": { - "origin": "Turing Institute staff directory", - "actor": "system-admin", - "chain": ["initial-import", "publication-update", "h-index-refresh"] - }, - "spatial": { - "lat": 51.5299, - "lon": -0.1278, - "label": "The Alan Turing Institute, British Library" - }, - "intentional_drift": false - }, - { - "id": "entity-013", - "category": "person", - "document": { - "title": "Dr. Fiona MacLeod", - "content": "Dr. Fiona MacLeod is a Reader in Formal Methods at the University of Edinburgh. She specialises in dependent type theory, formal verification of distributed systems, and semantic drift analysis. She is a core contributor to the Idris2 ecosystem and co-author of VeriSimDB's formal verification framework.", - "tags": ["researcher", "formal methods", "Idris2", "Edinburgh"] - }, - "graph": { - "edges": [ - {"target": "entity-002", "relation": "author_of"}, - {"target": "entity-005", "relation": "author_of"}, - {"target": "entity-022", "relation": "employed_by"}, - {"target": "entity-024", "relation": "visiting_researcher"} - ] - }, - "vector": { - "embedding": [0.86, 0.30, 0.71, 0.40, 0.80, 0.14, 0.60, 0.72] - }, - "semantic": { - "types": ["https://schema.org/Person", "https://schema.org/Researcher"], - "properties": { - "orcid": "0000-0003-4567-8901", - "h-index": 19, - "affiliation": "University of Edinburgh" - } - }, - "temporal": { - "created": "2025-02-15T11:30:00Z", - "modified": "2026-01-10T14:00:00Z", - "version": 2 - }, - "provenance": { - "origin": "University of Edinburgh staff page", - "actor": "system-admin", - "chain": ["initial-import", "orcid-link"] - }, - "spatial": { - "lat": 55.9445, - "lon": -3.1872, - "label": "University of Edinburgh, Informatics Forum" - }, - "intentional_drift": false - }, - { - "id": "entity-014", - "category": "person", - "document": { - "title": "Dr. Rajesh Kapoor", - "content": "Dr. Rajesh Kapoor is a Lecturer in Programming Languages at the University of Edinburgh. His work bridges type theory and database systems, focusing on proof-carrying queries and verified query compilation. He co-designed the VCL-UT (Usage-Tracked) extension and the PROOF clause semantics.", - "tags": ["researcher", "type theory", "VCL", "query compilation"] - }, - "graph": { - "edges": [ - {"target": "entity-002", "relation": "author_of"}, - {"target": "entity-008", "relation": "author_of"}, - {"target": "entity-022", "relation": "employed_by"} - ] - }, - "vector": { - "embedding": [0.84, 0.33, 0.66, 0.41, 0.78, 0.17, 0.56, 0.70] - }, - "semantic": { - "types": ["https://schema.org/Person", "https://schema.org/Researcher"], - "properties": { - "orcid": "0000-0004-5678-9012", - "h-index": 15, - "affiliation": "University of Edinburgh" - } - }, - "temporal": { - "created": "2025-02-15T11:30:00Z", - "modified": "2026-02-12T10:00:00Z", - "version": 3 - }, - "provenance": { - "origin": "University of Edinburgh staff page", - "actor": "system-admin", - "chain": ["initial-import", "publication-sync", "profile-update"] - }, - "spatial": { - "lat": 55.9445, - "lon": -3.1872, - "label": "University of Edinburgh, Informatics Forum" - }, - "intentional_drift": true, - "drift_note": "The temporal version is 3 with 3 provenance entries (consistent), but the vector embedding was copied from entity-013 (Fiona MacLeod) during a bulk import error. The embedding represents a different person entirely." - }, - { - "id": "entity-015", - "category": "person", - "document": { - "title": "Dr. Lena Vasquez", - "content": "Dr. Lena Vasquez is a Research Fellow at the University of Cambridge. Her expertise spans machine learning for database systems, embedding techniques, and neurosymbolic AI. She is the architect of the NeuroPlanner query optimiser and co-author of the octad embedding methodology.", - "tags": ["researcher", "ML", "embeddings", "Cambridge", "neurosymbolic"] - }, - "graph": { - "edges": [ - {"target": "entity-003", "relation": "author_of"}, - {"target": "entity-007", "relation": "author_of"}, - {"target": "entity-009", "relation": "author_of"}, - {"target": "entity-027", "relation": "employed_by"} - ] - }, - "vector": { - "embedding": [0.83, 0.25, 0.73, 0.50, 0.81, 0.09, 0.62, 0.82] - }, - "semantic": { - "types": ["https://schema.org/Person", "https://schema.org/Researcher"], - "properties": { - "orcid": "0000-0005-6789-0123", - "h-index": 22, - "affiliation": "University of Cambridge" - } - }, - "temporal": { - "created": "2025-03-01T09:00:00Z", - "modified": "2026-02-15T11:30:00Z", - "version": 3 - }, - "provenance": { - "origin": "Cambridge Computer Laboratory directory", - "actor": "system-admin", - "chain": ["initial-import", "publication-update", "h-index-refresh"] - }, - "spatial": { - "lat": 52.2107, - "lon": 0.0917, - "label": "University of Cambridge, William Gates Building" - }, - "intentional_drift": false - }, - { - "id": "entity-016", - "category": "person", - "document": { - "title": "Dr. James Okafor", - "content": "Dr. James Okafor is a Senior Research Fellow at University College London, specialising in data provenance, scientific workflows, and information governance. He led the ProvenanceDAG project and works closely with the Turing Institute on reproducibility standards.", - "tags": ["researcher", "provenance", "UCL", "reproducibility"] - }, - "graph": { - "edges": [ - {"target": "entity-004", "relation": "author_of"}, - {"target": "entity-009", "relation": "author_of"}, - {"target": "entity-023", "relation": "employed_by"}, - {"target": "entity-028", "relation": "affiliated_with"} - ] - }, - "vector": { - "embedding": [0.74, 0.40, 0.57, 0.34, 0.91, 0.23, 0.46, 0.65] - }, - "semantic": { - "types": ["https://schema.org/Person", "https://schema.org/Researcher"], - "properties": { - "orcid": "0000-0006-7890-1234", - "h-index": 31, - "affiliation": "University College London" - } - }, - "temporal": { - "created": "2025-01-20T14:00:00Z", - "modified": "2026-01-25T16:00:00Z", - "version": 2 - }, - "provenance": { - "origin": "UCL staff directory", - "actor": "system-admin", - "chain": ["initial-import", "publication-update"] - }, - "spatial": { - "lat": 51.5246, - "lon": -0.1340, - "label": "University College London, Gower Street" - }, - "intentional_drift": false - }, - { - "id": "entity-017", - "category": "person", - "document": { - "title": "Dr. Sofia Petridis", - "content": "Dr. Sofia Petridis is a Postdoctoral Researcher at the University of Oxford, focusing on knowledge graph evolution, ontology alignment, and semantic web technologies. She contributed the semantic drift detection algorithms to the VeriSimDB project.", - "tags": ["researcher", "knowledge graphs", "ontology", "Oxford"] - }, - "graph": { - "edges": [ - {"target": "entity-005", "relation": "author_of"}, - {"target": "entity-024", "relation": "employed_by"} - ] - }, - "vector": { - "embedding": [0.77, 0.37, 0.64, 0.47, 0.82, 0.19, 0.53, 0.70] - }, - "semantic": { - "types": ["https://schema.org/Person", "https://schema.org/Researcher"], - "properties": { - "orcid": "0000-0007-8901-2345", - "h-index": 11, - "affiliation": "University of Oxford" - } - }, - "temporal": { - "created": "2025-04-01T10:00:00Z", - "modified": "2026-02-01T09:30:00Z", - "version": 2 - }, - "provenance": { - "origin": "Oxford CS department page", - "actor": "system-admin", - "chain": ["initial-import", "profile-update"] - }, - "spatial": { - "lat": 51.7598, - "lon": -1.2584, - "label": "University of Oxford, Wolfson Building" - }, - "intentional_drift": false - }, - { - "id": "entity-018", - "category": "person", - "document": { - "title": "Dr. Kwame Asante", - "content": "Dr. Kwame Asante is a Lecturer in Machine Learning at King's College London. His research covers tensor methods, streaming algorithms, and real-time anomaly detection. He designed the HexTensor decomposition framework and contributes to VeriSimDB's tensor modality implementation.", - "tags": ["researcher", "tensors", "streaming", "King's College"] - }, - "graph": { - "edges": [ - {"target": "entity-006", "relation": "author_of"}, - {"target": "entity-025", "relation": "employed_by"} - ] - }, - "vector": { - "embedding": [0.81, 0.28, 0.70, 0.49, 0.76, 0.13, 0.58, 0.83] - }, - "semantic": { - "types": ["https://schema.org/Person", "https://schema.org/Researcher"], - "properties": { - "orcid": "0000-0008-9012-3456", - "h-index": 14, - "affiliation": "King's College London" - } - }, - "temporal": { - "created": "2025-05-10T08:30:00Z", - "modified": "2026-02-18T14:00:00Z", - "version": 2 - }, - "provenance": { - "origin": "KCL staff directory", - "actor": "system-admin", - "chain": ["initial-import", "publication-update"] - }, - "spatial": { - "lat": 51.5115, - "lon": -0.1160, - "label": "King's College London, Strand Campus" - }, - "intentional_drift": false - }, - { - "id": "entity-019", - "category": "person", - "document": { - "title": "Dr. Yuki Tanaka", - "content": "Dr. Yuki Tanaka is a Research Associate at Queen Mary University of London, specialising in spatial indexing, computational geometry, and location-aware databases. She designed the Octad R-tree variant and contributed spatial query optimisations to VeriSimDB.", - "tags": ["researcher", "spatial", "R-tree", "QMUL"] - }, - "graph": { - "edges": [ - {"target": "entity-007", "relation": "author_of"}, - {"target": "entity-026", "relation": "employed_by"} - ] - }, - "vector": { - "embedding": [0.75, 0.43, 0.56, 0.37, 0.87, 0.21, 0.49, 0.68] - }, - "semantic": { - "types": ["https://schema.org/Person", "https://schema.org/Researcher"], - "properties": { - "orcid": "0000-0009-0123-4567", - "h-index": 9, - "affiliation": "Queen Mary University of London" - } - }, - "temporal": { - "created": "2025-06-15T13:00:00Z", - "modified": "2026-01-28T12:00:00Z", - "version": 2 - }, - "provenance": { - "origin": "QMUL EECS directory", - "actor": "system-admin", - "chain": ["initial-import", "publication-update"] - }, - "spatial": { - "lat": 51.5229, - "lon": -0.0408, - "label": "Queen Mary University of London, Mile End" - }, - "intentional_drift": true, - "drift_note": "The document says 'Research Associate' but the semantic types only list 'Person' and 'Researcher' with no specific role type. Meanwhile, the h-index of 9 was last updated 6 months ago and is now stale — the temporal.modified date was bumped for a different reason (publication list) but h-index was not refreshed." - }, - { - "id": "entity-020", - "category": "person", - "document": { - "title": "Dr. Nia Williams", - "content": "Dr. Nia Williams is a Research Software Engineer at Imperial College London, bridging the gap between database research and production systems. She co-designed VCL's runtime, the benchmark harness for PolyBench, and maintains VeriSimDB's Rust core. She advocates for reproducible research infrastructure.", - "tags": ["RSE", "Rust", "VCL", "benchmark", "Imperial"] - }, - "graph": { - "edges": [ - {"target": "entity-008", "relation": "author_of"}, - {"target": "entity-010", "relation": "author_of"}, - {"target": "entity-021", "relation": "employed_by"} - ] - }, - "vector": { - "embedding": [0.89, 0.24, 0.67, 0.44, 0.85, 0.12, 0.55, 0.77] - }, - "semantic": { - "types": ["https://schema.org/Person", "https://schema.org/Researcher"], - "properties": { - "orcid": "0000-0010-1234-5678", - "h-index": 8, - "affiliation": "Imperial College London" - } - }, - "temporal": { - "created": "2025-03-20T16:00:00Z", - "modified": "2026-02-25T18:00:00Z", - "version": 4 - }, - "provenance": { - "origin": "Imperial College RSE team page", - "actor": "system-admin", - "chain": ["initial-import", "orcid-link", "publication-update", "role-update"] - }, - "spatial": { - "lat": 51.4988, - "lon": -0.1749, - "label": "Imperial College London, South Kensington" - }, - "intentional_drift": false - }, - { - "id": "entity-021", - "category": "organisation", - "document": { - "title": "Imperial College London", - "content": "Imperial College London is a world-leading research university in South Kensington, London. The Department of Computing hosts the Data Systems Lab where VeriSimDB was conceived and developed. Imperial is ranked in the top 10 globally for computer science and engineering.", - "tags": ["university", "London", "Russell Group", "data systems"] - }, - "graph": { - "edges": [ - {"target": "entity-011", "relation": "employs"}, - {"target": "entity-020", "relation": "employs"}, - {"target": "entity-028", "relation": "partner_of"}, - {"target": "entity-029", "relation": "funded_by"} - ] - }, - "vector": { - "embedding": [0.60, 0.50, 0.45, 0.30, 0.70, 0.35, 0.40, 0.55] - }, - "semantic": { - "types": ["https://schema.org/EducationalOrganization", "https://schema.org/CollegeOrUniversity"], - "properties": { - "founded": 1907, - "country": "United Kingdom", - "ror": "https://ror.org/041kmwe10" - } - }, - "temporal": { - "created": "2025-01-01T00:00:00Z", - "modified": "2026-01-15T10:00:00Z", - "version": 2 - }, - "provenance": { - "origin": "ROR registry import", - "actor": "system-admin", - "chain": ["ror-import", "manual-enrichment"] - }, - "spatial": { - "lat": 51.4988, - "lon": -0.1749, - "label": "Imperial College London, South Kensington, London SW7 2AZ" - }, - "intentional_drift": false - }, - { - "id": "entity-022", - "category": "organisation", - "document": { - "title": "University of Edinburgh", - "content": "The University of Edinburgh is one of the oldest and most prestigious universities in the UK. The School of Informatics is Europe's largest informatics research centre, and the Laboratory for Foundations of Computer Science (LFCS) is a world leader in type theory and formal methods.", - "tags": ["university", "Edinburgh", "informatics", "formal methods"] - }, - "graph": { - "edges": [ - {"target": "entity-013", "relation": "employs"}, - {"target": "entity-014", "relation": "employs"}, - {"target": "entity-028", "relation": "partner_of"} - ] - }, - "vector": { - "embedding": [0.58, 0.52, 0.47, 0.32, 0.68, 0.37, 0.42, 0.53] - }, - "semantic": { - "types": ["https://schema.org/EducationalOrganization", "https://schema.org/CollegeOrUniversity"], - "properties": { - "founded": 1583, - "country": "United Kingdom", - "ror": "https://ror.org/01nrxwf90" - } - }, - "temporal": { - "created": "2025-01-01T00:00:00Z", - "modified": "2026-01-15T10:00:00Z", - "version": 2 - }, - "provenance": { - "origin": "ROR registry import", - "actor": "system-admin", - "chain": ["ror-import", "manual-enrichment"] - }, - "spatial": { - "lat": 55.9445, - "lon": -3.1872, - "label": "University of Edinburgh, Informatics Forum, Edinburgh EH8 9AB" - }, - "intentional_drift": false - }, - { - "id": "entity-023", - "category": "organisation", - "document": { - "title": "University College London", - "content": "UCL is a leading multidisciplinary university in central London. The Department of Information Studies and the Department of Computer Science collaborate on data management, provenance, and digital scholarship research. UCL is a founding member of the Russell Group.", - "tags": ["university", "London", "UCL", "information studies"] - }, - "graph": { - "edges": [ - {"target": "entity-016", "relation": "employs"}, - {"target": "entity-012", "relation": "affiliated_researcher"}, - {"target": "entity-028", "relation": "partner_of"} - ] - }, - "vector": { - "embedding": [0.62, 0.48, 0.44, 0.31, 0.72, 0.33, 0.41, 0.57] - }, - "semantic": { - "types": ["https://schema.org/EducationalOrganization", "https://schema.org/CollegeOrUniversity"], - "properties": { - "founded": 1826, - "country": "United Kingdom", - "ror": "https://ror.org/02jx3x895" - } - }, - "temporal": { - "created": "2025-01-01T00:00:00Z", - "modified": "2026-01-15T10:00:00Z", - "version": 2 - }, - "provenance": { - "origin": "ROR registry import", - "actor": "system-admin", - "chain": ["ror-import", "manual-enrichment"] - }, - "spatial": { - "lat": 51.5246, - "lon": -0.1340, - "label": "University College London, Gower Street, London WC1E 6BT" - }, - "intentional_drift": false - }, - { - "id": "entity-024", - "category": "organisation", - "document": { - "title": "University of Oxford — Department of Computer Science", - "content": "Oxford's Department of Computer Science conducts world-class research in algorithms, verification, machine learning, and knowledge representation. The Knowledge Representation and Reasoning group works on ontology alignment and semantic web technologies relevant to VeriSimDB's semantic modality.", - "tags": ["university", "Oxford", "knowledge representation", "verification"] - }, - "graph": { - "edges": [ - {"target": "entity-017", "relation": "employs"}, - {"target": "entity-013", "relation": "hosts_visitor"}, - {"target": "entity-030", "relation": "funded_by"} - ] - }, - "vector": { - "embedding": [0.56, 0.54, 0.49, 0.34, 0.66, 0.39, 0.44, 0.51] - }, - "semantic": { - "types": ["https://schema.org/EducationalOrganization", "https://schema.org/CollegeOrUniversity"], - "properties": { - "founded": 1096, - "country": "United Kingdom", - "ror": "https://ror.org/052gg0110" - } - }, - "temporal": { - "created": "2025-01-01T00:00:00Z", - "modified": "2026-01-15T10:00:00Z", - "version": 2 - }, - "provenance": { - "origin": "ROR registry import", - "actor": "system-admin", - "chain": ["ror-import", "manual-enrichment"] - }, - "spatial": { - "lat": 51.7598, - "lon": -1.2584, - "label": "University of Oxford, Wolfson Building, Parks Road, Oxford OX1 3QD" - }, - "intentional_drift": false - }, - { - "id": "entity-025", - "category": "organisation", - "document": { - "title": "King's College London — Department of Informatics", - "content": "King's College London's Department of Informatics focuses on AI, data science, and cybersecurity. The Streaming Analytics group, led by Dr. Kwame Asante, develops real-time data processing frameworks including the HexTensor system for VeriSimDB.", - "tags": ["university", "London", "KCL", "streaming", "AI"] - }, - "graph": { - "edges": [ - {"target": "entity-018", "relation": "employs"}, - {"target": "entity-028", "relation": "partner_of"} - ] - }, - "vector": { - "embedding": [0.64, 0.46, 0.43, 0.29, 0.74, 0.31, 0.39, 0.59] - }, - "semantic": { - "types": ["https://schema.org/EducationalOrganization", "https://schema.org/CollegeOrUniversity"], - "properties": { - "founded": 1829, - "country": "United Kingdom", - "ror": "https://ror.org/0220mzb33" - } - }, - "temporal": { - "created": "2025-01-01T00:00:00Z", - "modified": "2026-01-15T10:00:00Z", - "version": 2 - }, - "provenance": { - "origin": "ROR registry import", - "actor": "system-admin", - "chain": ["ror-import", "manual-enrichment"] - }, - "spatial": { - "lat": 51.5115, - "lon": -0.1160, - "label": "King's College London, Strand, London WC2R 2LS" - }, - "intentional_drift": false - }, - { - "id": "entity-026", - "category": "organisation", - "document": { - "title": "The Alan Turing Institute — Research Engineering Group", - "content": "The Alan Turing Institute is the UK's national institute for data science and artificial intelligence, headquartered at the British Library. The Research Engineering Group builds production-quality tools for research, including the PolyBench benchmark infrastructure used to evaluate VeriSimDB.", - "tags": ["research institute", "Turing", "data science", "AI", "REG"] - }, - "graph": { - "edges": [ - {"target": "entity-021", "relation": "partner_of"}, - {"target": "entity-022", "relation": "partner_of"}, - {"target": "entity-023", "relation": "partner_of"}, - {"target": "entity-029", "relation": "funded_by"} - ] - }, - "vector": { - "embedding": [0.66, 0.42, 0.48, 0.36, 0.76, 0.28, 0.45, 0.61] - }, - "semantic": { - "types": ["https://schema.org/ResearchOrganization", "https://schema.org/GovernmentOrganization"], - "properties": { - "founded": 2015, - "country": "United Kingdom", - "ror": "https://ror.org/01bd7kw46" - } - }, - "temporal": { - "created": "2025-01-01T00:00:00Z", - "modified": "2026-02-01T08:00:00Z", - "version": 3 - }, - "provenance": { - "origin": "Turing Institute website", - "actor": "system-admin", - "chain": ["initial-import", "partner-update", "funding-update"] - }, - "spatial": { - "lat": 51.5299, - "lon": -0.1278, - "label": "The Alan Turing Institute, British Library, 96 Euston Road, London NW1 2DB" - }, - "intentional_drift": true, - "drift_note": "The semantic types include 'GovernmentOrganization' which is incorrect — the Turing Institute is an independent research body (a charity), not a government organisation. It receives government funding via EPSRC but is not itself a government entity. The semantic annotation has conflated funding source with organisational type." - }, - { - "id": "entity-027", - "category": "organisation", - "document": { - "title": "University of Cambridge — Computer Laboratory", - "content": "The University of Cambridge Computer Laboratory is one of the oldest and most distinguished computer science departments in the world. Research groups span programming languages, systems, AI, and security. Dr. Lena Vasquez's work on neurosymbolic query optimisation is conducted here.", - "tags": ["university", "Cambridge", "computer laboratory", "PL"] - }, - "graph": { - "edges": [ - {"target": "entity-015", "relation": "employs"}, - {"target": "entity-028", "relation": "partner_of"} - ] - }, - "vector": { - "embedding": [0.55, 0.56, 0.50, 0.33, 0.65, 0.40, 0.46, 0.50] - }, - "semantic": { - "types": ["https://schema.org/EducationalOrganization", "https://schema.org/CollegeOrUniversity"], - "properties": { - "founded": 1209, - "country": "United Kingdom", - "ror": "https://ror.org/013meh722" - } - }, - "temporal": { - "created": "2025-01-01T00:00:00Z", - "modified": "2026-01-15T10:00:00Z", - "version": 2 - }, - "provenance": { - "origin": "ROR registry import", - "actor": "system-admin", - "chain": ["ror-import", "manual-enrichment"] - }, - "spatial": { - "lat": 52.2107, - "lon": 0.0917, - "label": "University of Cambridge, William Gates Building, Cambridge CB3 0FD" - }, - "intentional_drift": false - }, - { - "id": "entity-028", - "category": "organisation", - "document": { - "title": "Queen Mary University of London — School of Electronic Engineering and Computer Science", - "content": "QMUL's School of EECS is a research-intensive department in East London. The Spatial Computing group develops novel indexing techniques for geographic and location-aware data. Dr. Yuki Tanaka's Octad R-tree work is based here.", - "tags": ["university", "London", "QMUL", "EECS", "spatial"] - }, - "graph": { - "edges": [ - {"target": "entity-019", "relation": "employs"}, - {"target": "entity-026", "relation": "partner_of"} - ] - }, - "vector": { - "embedding": [0.61, 0.47, 0.46, 0.28, 0.71, 0.34, 0.38, 0.56] - }, - "semantic": { - "types": ["https://schema.org/EducationalOrganization", "https://schema.org/CollegeOrUniversity"], - "properties": { - "founded": 1885, - "country": "United Kingdom", - "ror": "https://ror.org/026zzn846" - } - }, - "temporal": { - "created": "2025-01-01T00:00:00Z", - "modified": "2026-01-15T10:00:00Z", - "version": 2 - }, - "provenance": { - "origin": "ROR registry import", - "actor": "system-admin", - "chain": ["ror-import", "manual-enrichment"] - }, - "spatial": { - "lat": 51.5229, - "lon": -0.0408, - "label": "Queen Mary University of London, Mile End Road, London E1 4NS" - }, - "intentional_drift": false - }, - { - "id": "entity-029", - "category": "organisation", - "document": { - "title": "EPSRC — Engineering and Physical Sciences Research Council", - "content": "EPSRC is the UK's main funding body for engineering and physical sciences research. It funds the VeriSimDB project through grant EP/X012345/1 ('Verified Multimodal Data Management') and supports the Turing Institute's data science programme.", - "tags": ["funder", "EPSRC", "UKRI", "research council"] - }, - "graph": { - "edges": [ - {"target": "entity-021", "relation": "funds"}, - {"target": "entity-026", "relation": "funds"}, - {"target": "entity-030", "relation": "parent_of"} - ] - }, - "vector": { - "embedding": [0.45, 0.60, 0.38, 0.25, 0.55, 0.45, 0.35, 0.42] - }, - "semantic": { - "types": ["https://schema.org/FundingAgency", "https://schema.org/GovernmentOrganization"], - "properties": { - "country": "United Kingdom", - "parent": "UKRI", - "fundref": "https://doi.org/10.13039/501100000266" - } - }, - "temporal": { - "created": "2025-01-01T00:00:00Z", - "modified": "2025-12-01T10:00:00Z", - "version": 1 - }, - "provenance": { - "origin": "CrossRef Funder Registry", - "actor": "system-admin", - "chain": ["fundref-import"] - }, - "spatial": { - "lat": 51.7520, - "lon": -1.2577, - "label": "EPSRC, Polaris House, Swindon (HQ); grant managed from Oxford" - }, - "intentional_drift": true, - "drift_note": "The spatial coordinates (51.7520, -1.2577) point to Oxford, not Swindon where EPSRC HQ actually is (51.5584, -1.7817). The label mentions Swindon but the coordinates are wrong — a classic spatial-document inconsistency." - }, - { - "id": "entity-030", - "category": "organisation", - "document": { - "title": "UKRI — UK Research and Innovation", - "content": "UKRI is the umbrella body for the UK's seven research councils, including EPSRC, and Research England and Innovate UK. UKRI coordinates national research strategy and infrastructure, including the Digital Research Infrastructure programme that supports projects like VeriSimDB.", - "tags": ["funder", "UKRI", "research infrastructure", "government"] - }, - "graph": { - "edges": [ - {"target": "entity-029", "relation": "parent_of"}, - {"target": "entity-024", "relation": "funds"} - ] - }, - "vector": { - "embedding": [0.43, 0.62, 0.36, 0.23, 0.53, 0.47, 0.33, 0.40] - }, - "semantic": { - "types": ["https://schema.org/GovernmentOrganization", "https://schema.org/FundingAgency"], - "properties": { - "country": "United Kingdom", - "established": 2018, - "fundref": "https://doi.org/10.13039/100014013" - } - }, - "temporal": { - "created": "2025-01-01T00:00:00Z", - "modified": "2025-12-01T10:00:00Z", - "version": 1 - }, - "provenance": { - "origin": "CrossRef Funder Registry", - "actor": "system-admin", - "chain": ["fundref-import"] - }, - "spatial": { - "lat": 51.7520, - "lon": -1.2577, - "label": "UKRI, Polaris House, Swindon SN2 1FL" - }, - "intentional_drift": false - }, - { - "id": "entity-031", - "category": "dataset", - "document": { - "title": "PolyBench v1.0 — Multimodal Database Benchmark Dataset", - "content": "PolyBench v1.0 is a benchmark dataset containing 100,000 synthetic octad records across five domains: academic publications, social networks, geospatial entities, financial transactions, and IoT sensor readings. Each record contains all eight VeriSimDB modalities with controlled levels of cross-modal drift for evaluation purposes.", - "tags": ["benchmark", "dataset", "PolyBench", "synthetic", "multimodal"] - }, - "graph": { - "edges": [ - {"target": "entity-001", "relation": "used_in"}, - {"target": "entity-003", "relation": "used_in"}, - {"target": "entity-010", "relation": "described_in"}, - {"target": "entity-012", "relation": "created_by"}, - {"target": "entity-020", "relation": "maintained_by"} - ] - }, - "vector": { - "embedding": [0.70, 0.45, 0.55, 0.40, 0.75, 0.30, 0.50, 0.65] - }, - "semantic": { - "types": ["https://schema.org/Dataset", "https://schema.org/CreativeWork"], - "properties": { - "size": "2.3 GB", - "records": 100000, - "format": "JSON octad", - "license": "CC-BY-4.0" - } - }, - "temporal": { - "created": "2025-06-01T00:00:00Z", - "modified": "2026-02-01T12:00:00Z", - "version": 3 - }, - "provenance": { - "origin": "Imperial College London — Data Systems Lab", - "actor": "entity-020", - "chain": ["initial-generation", "drift-injection", "v1.0-release"] - }, - "spatial": { - "lat": 51.4988, - "lon": -0.1749, - "label": "Imperial College London (generation site)" - }, - "intentional_drift": false - }, - { - "id": "entity-032", - "category": "dataset", - "document": { - "title": "IdrisProofs Corpus — Verified Database Invariants", - "content": "A curated collection of 500 Idris2 proof terms establishing cross-modal consistency invariants for the VeriSimDB octad model. Each proof covers a specific invariant (e.g., 'graph edge targets must exist as entities', 'vector dimensionality matches schema declaration'). Used to evaluate the formal verification framework in entity-002.", - "tags": ["dataset", "Idris2", "proofs", "verification", "invariants"] - }, - "graph": { - "edges": [ - {"target": "entity-002", "relation": "used_in"}, - {"target": "entity-013", "relation": "created_by"}, - {"target": "entity-022", "relation": "hosted_at"} - ] - }, - "vector": { - "embedding": [0.82, 0.35, 0.68, 0.38, 0.79, 0.16, 0.58, 0.72] - }, - "semantic": { - "types": ["https://schema.org/Dataset"], - "properties": { - "size": "45 MB", - "records": 500, - "format": "Idris2 source files", - "license": "MIT" - } - }, - "temporal": { - "created": "2025-07-01T00:00:00Z", - "modified": "2026-01-15T09:00:00Z", - "version": 2 - }, - "provenance": { - "origin": "University of Edinburgh — LFCS", - "actor": "entity-013", - "chain": ["initial-curation", "v2-expansion"] - }, - "spatial": { - "lat": 55.9445, - "lon": -3.1872, - "label": "University of Edinburgh (curation site)" - }, - "intentional_drift": false - }, - { - "id": "entity-033", - "category": "dataset", - "document": { - "title": "ModalMix — Cross-Modal Embedding Evaluation Set", - "content": "ModalMix is an evaluation dataset of 10,000 entity pairs with human-annotated similarity scores across different modality combinations. It enables evaluation of cross-modal embedding quality: does the vector similarity between two entities correlate with their semantic similarity, graph proximity, and document relevance?", - "tags": ["dataset", "embeddings", "evaluation", "similarity", "human-annotated"] - }, - "graph": { - "edges": [ - {"target": "entity-003", "relation": "used_in"}, - {"target": "entity-015", "relation": "created_by"}, - {"target": "entity-027", "relation": "hosted_at"} - ] - }, - "vector": { - "embedding": [0.78, 0.32, 0.64, 0.48, 0.73, 0.15, 0.56, 0.75] - }, - "semantic": { - "types": ["https://schema.org/Dataset"], - "properties": { - "size": "890 MB", - "records": 10000, - "format": "Parquet", - "license": "CC-BY-SA-4.0" - } - }, - "temporal": { - "created": "2025-08-15T00:00:00Z", - "modified": "2026-01-20T14:00:00Z", - "version": 2 - }, - "provenance": { - "origin": "University of Cambridge — Computer Laboratory", - "actor": "entity-015", - "chain": ["annotation-campaign", "quality-review"] - }, - "spatial": { - "lat": 52.2107, - "lon": 0.0917, - "label": "University of Cambridge (annotation site)" - }, - "intentional_drift": true, - "drift_note": "The provenance chain has 2 entries but the temporal version is also 2 (consistent). However, the document says '10,000 entity pairs' while the semantic properties say 'records: 10000'. These are entity PAIRS not individual records — the actual dataset has 20,000 rows. The semantic record count is wrong." - }, - { - "id": "entity-034", - "category": "dataset", - "document": { - "title": "W3C PROV-O Test Suite — Provenance Ontology Conformance Data", - "content": "The official W3C PROV-O conformance test suite adapted for VeriSimDB's provenance modality. Contains 200 provenance graphs in PROV-N notation covering all 16 PROV-O relations, with corresponding VeriSimDB octad representations. Used to validate ProvenanceDAG's compliance with the W3C standard.", - "tags": ["dataset", "W3C", "PROV-O", "provenance", "conformance"] - }, - "graph": { - "edges": [ - {"target": "entity-004", "relation": "used_in"}, - {"target": "entity-016", "relation": "adapted_by"}, - {"target": "entity-023", "relation": "hosted_at"} - ] - }, - "vector": { - "embedding": [0.68, 0.42, 0.53, 0.35, 0.88, 0.25, 0.44, 0.62] - }, - "semantic": { - "types": ["https://schema.org/Dataset"], - "properties": { - "size": "12 MB", - "records": 200, - "format": "PROV-N + JSON octad", - "license": "W3C Software License" - } - }, - "temporal": { - "created": "2025-05-01T00:00:00Z", - "modified": "2025-11-30T16:00:00Z", - "version": 1 - }, - "provenance": { - "origin": "W3C Provenance Working Group (adapted by UCL)", - "actor": "entity-016", - "chain": ["w3c-import"] - }, - "spatial": { - "lat": 51.5246, - "lon": -0.1340, - "label": "University College London (adaptation site)" - }, - "intentional_drift": false - }, - { - "id": "entity-035", - "category": "dataset", - "document": { - "title": "DBpedia-Wikidata Drift Corpus", - "content": "A parallel corpus of 50,000 entities present in both DBpedia and Wikidata, annotated with semantic drift labels. Each entity pair is tagged with drift type (ontology mismatch, temporal lag, factual disagreement, or no drift) and severity score. Used to evaluate semantic drift detection in entity-005.", - "tags": ["dataset", "DBpedia", "Wikidata", "drift", "parallel corpus"] - }, - "graph": { - "edges": [ - {"target": "entity-005", "relation": "used_in"}, - {"target": "entity-017", "relation": "created_by"}, - {"target": "entity-024", "relation": "hosted_at"} - ] - }, - "vector": { - "embedding": [0.73, 0.39, 0.60, 0.44, 0.80, 0.20, 0.51, 0.68] - }, - "semantic": { - "types": ["https://schema.org/Dataset"], - "properties": { - "size": "4.7 GB", - "records": 50000, - "format": "N-Triples + JSON", - "license": "CC-BY-3.0" - } - }, - "temporal": { - "created": "2025-06-20T00:00:00Z", - "modified": "2026-01-10T11:00:00Z", - "version": 3 - }, - "provenance": { - "origin": "University of Oxford — CS Department", - "actor": "entity-017", - "chain": ["dbpedia-extract", "wikidata-extract", "alignment-annotation"] - }, - "spatial": { - "lat": 51.7598, - "lon": -1.2584, - "label": "University of Oxford (curation site)" - }, - "intentional_drift": false - }, - { - "id": "entity-036", - "category": "dataset", - "document": { - "title": "StreamHex — Synthetic Streaming Octad Data Generator Output", - "content": "StreamHex is a 1-million-record dataset generated by a synthetic streaming data generator. Each record simulates a real-time octad arriving from IoT sensors, social media feeds, or financial tickers. The dataset includes injected anomalies at known positions for evaluating the HexTensor anomaly detection system.", - "tags": ["dataset", "streaming", "synthetic", "anomaly", "IoT"] - }, - "graph": { - "edges": [ - {"target": "entity-006", "relation": "used_in"}, - {"target": "entity-018", "relation": "created_by"}, - {"target": "entity-025", "relation": "hosted_at"} - ] - }, - "vector": { - "embedding": [0.71, 0.34, 0.62, 0.50, 0.72, 0.18, 0.54, 0.77] - }, - "semantic": { - "types": ["https://schema.org/Dataset"], - "properties": { - "size": "18 GB", - "records": 1000000, - "format": "Apache Parquet (streaming chunks)", - "license": "CC-BY-4.0" - } - }, - "temporal": { - "created": "2025-09-15T00:00:00Z", - "modified": "2026-02-15T10:00:00Z", - "version": 2 - }, - "provenance": { - "origin": "King's College London — Informatics Department", - "actor": "entity-018", - "chain": ["generation-run-1", "anomaly-injection"] - }, - "spatial": { - "lat": 51.5115, - "lon": -0.1160, - "label": "King's College London (generation site)" - }, - "intentional_drift": false - }, - { - "id": "entity-037", - "category": "dataset", - "document": { - "title": "GeoHex London — Spatial Octad Dataset of London Points of Interest", - "content": "A curated dataset of 25,000 London-area points of interest encoded as VeriSimDB octads. Each POI includes spatial coordinates, OpenStreetMap tags (semantic), textual descriptions (document), connectivity to nearby POIs (graph), and temporal opening/closing history. Used to benchmark spatial indexing strategies.", - "tags": ["dataset", "spatial", "London", "POI", "OpenStreetMap"] - }, - "graph": { - "edges": [ - {"target": "entity-007", "relation": "used_in"}, - {"target": "entity-019", "relation": "created_by"}, - {"target": "entity-028", "relation": "hosted_at"} - ] - }, - "vector": { - "embedding": [0.65, 0.48, 0.50, 0.36, 0.85, 0.26, 0.43, 0.60] - }, - "semantic": { - "types": ["https://schema.org/Dataset"], - "properties": { - "size": "1.2 GB", - "records": 25000, - "format": "GeoJSON + JSON octad", - "license": "ODbL-1.0" - } - }, - "temporal": { - "created": "2025-09-01T00:00:00Z", - "modified": "2026-01-25T15:00:00Z", - "version": 2 - }, - "provenance": { - "origin": "OpenStreetMap extract (London bounding box), enriched at QMUL", - "actor": "entity-019", - "chain": ["osm-extract", "octad-enrichment"] - }, - "spatial": { - "lat": 51.5074, - "lon": -0.1278, - "label": "Central London (coverage area centroid)" - }, - "intentional_drift": false - }, - { - "id": "entity-038", - "category": "dataset", - "document": { - "title": "VCL Test Suite — Query Language Conformance Tests", - "content": "A comprehensive test suite of 1,200 VCL queries covering all language features: basic SELECT, cross-modal joins, PROOF clauses (existence, consistency, provenance), temporal range filters, spatial radius queries, and error handling. Each query has expected results for deterministic evaluation of VCL implementations.", - "tags": ["dataset", "VCL", "test suite", "conformance", "queries"] - }, - "graph": { - "edges": [ - {"target": "entity-008", "relation": "used_in"}, - {"target": "entity-020", "relation": "created_by"}, - {"target": "entity-014", "relation": "reviewed_by"}, - {"target": "entity-021", "relation": "hosted_at"} - ] - }, - "vector": { - "embedding": [0.88, 0.26, 0.66, 0.43, 0.84, 0.13, 0.55, 0.78] - }, - "semantic": { - "types": ["https://schema.org/Dataset", "https://schema.org/SoftwareSourceCode"], - "properties": { - "size": "8 MB", - "records": 1200, - "format": "VCL + JSON expected results", - "license": "PMPL-1.0-or-later" - } - }, - "temporal": { - "created": "2025-12-01T00:00:00Z", - "modified": "2026-02-24T17:00:00Z", - "version": 5 - }, - "provenance": { - "origin": "Imperial College London — Data Systems Lab", - "actor": "entity-020", - "chain": ["initial-suite", "proof-clause-tests", "spatial-tests", "temporal-tests", "error-handling-tests"] - }, - "spatial": { - "lat": 51.4988, - "lon": -0.1749, - "label": "Imperial College London (development site)" - }, - "intentional_drift": false - }, - { - "id": "entity-039", - "category": "dataset", - "document": { - "title": "NeuroPlan Workload Traces — Query Optimiser Training Data", - "content": "A collection of 50,000 VCL query execution traces with measured latencies, resource consumption, and plan choices. Each trace captures the query plan selected by VeriSimDB's default optimiser alongside the optimal plan found by exhaustive search. Used to train the NeuroPlanner's GNN cost model.", - "tags": ["dataset", "query traces", "optimisation", "training data", "GNN"] - }, - "graph": { - "edges": [ - {"target": "entity-009", "relation": "used_in"}, - {"target": "entity-015", "relation": "created_by"}, - {"target": "entity-027", "relation": "hosted_at"} - ] - }, - "vector": { - "embedding": [0.84, 0.30, 0.63, 0.41, 0.80, 0.17, 0.53, 0.74] - }, - "semantic": { - "types": ["https://schema.org/Dataset"], - "properties": { - "size": "3.6 GB", - "records": 50000, - "format": "Apache Arrow", - "license": "CC-BY-4.0" - } - }, - "temporal": { - "created": "2025-11-01T00:00:00Z", - "modified": "2026-02-10T13:00:00Z", - "version": 2 - }, - "provenance": { - "origin": "University of Cambridge — Computer Laboratory", - "actor": "entity-015", - "chain": ["workload-capture", "annotation"] - }, - "spatial": { - "lat": 52.2107, - "lon": 0.0917, - "label": "University of Cambridge (capture site)" - }, - "intentional_drift": false - }, - { - "id": "entity-040", - "category": "dataset", - "document": { - "title": "DriftSeed — Controlled Drift Injection Toolkit Output", - "content": "DriftSeed is a meta-dataset of 5,000 octad records generated with precisely controlled cross-modal drift. Each record documents the type of drift injected (vector staleness, semantic type mismatch, temporal-provenance desync, spatial-document inconsistency), its severity, and ground-truth labels. Essential for calibrating VeriSimDB's drift detection thresholds.", - "tags": ["dataset", "drift", "controlled", "ground truth", "calibration"] - }, - "graph": { - "edges": [ - {"target": "entity-001", "relation": "used_in"}, - {"target": "entity-005", "relation": "used_in"}, - {"target": "entity-011", "relation": "created_by"}, - {"target": "entity-021", "relation": "hosted_at"} - ] - }, - "vector": { - "embedding": [0.76, 0.38, 0.58, 0.42, 0.82, 0.22, 0.49, 0.67] - }, - "semantic": { - "types": ["https://schema.org/Dataset"], - "properties": { - "size": "120 MB", - "records": 5000, - "format": "JSON octad with drift labels", - "license": "PMPL-1.0-or-later" - } - }, - "temporal": { - "created": "2025-04-01T00:00:00Z", - "modified": "2026-02-20T09:00:00Z", - "version": 4 - }, - "provenance": { - "origin": "Imperial College London — Data Systems Lab", - "actor": "entity-011", - "chain": ["baseline-generation", "drift-injection-v1", "drift-injection-v2", "threshold-calibration"] - }, - "spatial": { - "lat": 51.4988, - "lon": -0.1749, - "label": "Imperial College London (generation site)" - }, - "intentional_drift": false - }, - { - "id": "entity-041", - "category": "event", - "document": { - "title": "VLDB 2025 — 51st International Conference on Very Large Data Bases", - "content": "VLDB 2025 was held in London, UK, bringing together over 1,500 researchers and practitioners in data management. The conference featured the first public presentation of VeriSimDB and the semantic drift detection paper. The industrial track included demonstrations of multimodal database systems.", - "tags": ["conference", "VLDB", "2025", "London", "databases"] - }, - "graph": { - "edges": [ - {"target": "entity-001", "relation": "featured_paper"}, - {"target": "entity-005", "relation": "featured_paper"}, - {"target": "entity-026", "relation": "co-located_with"} - ] - }, - "vector": { - "embedding": [0.50, 0.55, 0.40, 0.35, 0.60, 0.50, 0.38, 0.48] - }, - "semantic": { - "types": ["https://schema.org/Event", "https://schema.org/BusinessEvent"], - "properties": { - "startDate": "2025-08-25", - "endDate": "2025-08-29", - "attendees": 1500, - "url": "https://vldb.org/2025/" - } - }, - "temporal": { - "created": "2025-01-15T00:00:00Z", - "modified": "2025-09-15T10:00:00Z", - "version": 3 - }, - "provenance": { - "origin": "VLDB Endowment", - "actor": "system-admin", - "chain": ["cfp-announcement", "programme-published", "post-event-update"] - }, - "spatial": { - "lat": 51.5032, - "lon": -0.0195, - "label": "ExCeL London, Royal Victoria Dock, London E16 1XL" - }, - "intentional_drift": false - }, - { - "id": "entity-042", - "category": "event", - "document": { - "title": "SIGMOD 2026 — ACM International Conference on Management of Data", - "content": "SIGMOD 2026 is scheduled for Berlin, Germany. The formal verification paper (entity-002) and the HexTensor paper (entity-006) are accepted for presentation. The conference includes a new track on verified data management systems, reflecting growing interest in provably correct databases.", - "tags": ["conference", "SIGMOD", "2026", "Berlin", "data management"] - }, - "graph": { - "edges": [ - {"target": "entity-002", "relation": "featured_paper"}, - {"target": "entity-006", "relation": "featured_paper"} - ] - }, - "vector": { - "embedding": [0.48, 0.57, 0.42, 0.33, 0.58, 0.52, 0.36, 0.46] - }, - "semantic": { - "types": ["https://schema.org/Event", "https://schema.org/BusinessEvent"], - "properties": { - "startDate": "2026-06-14", - "endDate": "2026-06-19", - "attendees": 1200, - "url": "https://sigmod2026.org/" - } - }, - "temporal": { - "created": "2025-06-01T00:00:00Z", - "modified": "2026-02-15T12:00:00Z", - "version": 4 - }, - "provenance": { - "origin": "ACM SIGMOD", - "actor": "system-admin", - "chain": ["cfp-announcement", "submission-deadline", "acceptance-notification", "programme-draft"] - }, - "spatial": { - "lat": 52.5200, - "lon": 13.4050, - "label": "Berlin Conference Centre, Berlin, Germany" - }, - "intentional_drift": false - }, - { - "id": "entity-043", - "category": "event", - "document": { - "title": "ICDE 2026 — 42nd IEEE International Conference on Data Engineering", - "content": "ICDE 2026 takes place in Utrecht, Netherlands. The octad embeddings paper (entity-003) and the VCL language paper (entity-008) are both accepted. A tutorial on 'Building Multimodal Databases from Scratch' by Dr. Okonkwo is a highlight of the programme.", - "tags": ["conference", "ICDE", "2026", "Utrecht", "data engineering"] - }, - "graph": { - "edges": [ - {"target": "entity-003", "relation": "featured_paper"}, - {"target": "entity-008", "relation": "featured_paper"}, - {"target": "entity-011", "relation": "tutorial_by"} - ] - }, - "vector": { - "embedding": [0.52, 0.53, 0.41, 0.37, 0.62, 0.48, 0.37, 0.50] - }, - "semantic": { - "types": ["https://schema.org/Event", "https://schema.org/BusinessEvent"], - "properties": { - "startDate": "2026-04-20", - "endDate": "2026-04-24", - "attendees": 800, - "url": "https://icde2026.org/" - } - }, - "temporal": { - "created": "2025-07-01T00:00:00Z", - "modified": "2026-02-20T09:00:00Z", - "version": 3 - }, - "provenance": { - "origin": "IEEE TCDE", - "actor": "system-admin", - "chain": ["cfp-announcement", "acceptance-notification", "programme-published"] - }, - "spatial": { - "lat": 52.0907, - "lon": 5.1214, - "label": "TivoliVredenburg, Utrecht, Netherlands" - }, - "intentional_drift": false - }, - { - "id": "entity-044", - "category": "event", - "document": { - "title": "Workshop on Provenance and Annotation of Data (IPAW 2025)", - "content": "The International Provenance and Annotation Workshop (IPAW) 2025 was co-located with eScience 2025 in Edinburgh. Dr. Okafor presented the ProvenanceDAG work (entity-004) and Dr. Vasquez presented the neurosymbolic optimiser (entity-009). The workshop attracted 120 attendees from the provenance and reproducibility communities.", - "tags": ["workshop", "provenance", "IPAW", "2025", "Edinburgh"] - }, - "graph": { - "edges": [ - {"target": "entity-004", "relation": "featured_paper"}, - {"target": "entity-009", "relation": "featured_paper"}, - {"target": "entity-016", "relation": "presented_by"}, - {"target": "entity-015", "relation": "presented_by"} - ] - }, - "vector": { - "embedding": [0.55, 0.50, 0.45, 0.38, 0.65, 0.42, 0.40, 0.52] - }, - "semantic": { - "types": ["https://schema.org/Event"], - "properties": { - "startDate": "2025-10-13", - "endDate": "2025-10-14", - "attendees": 120, - "url": "https://ipaw2025.org/" - } - }, - "temporal": { - "created": "2025-03-01T00:00:00Z", - "modified": "2025-11-01T10:00:00Z", - "version": 3 - }, - "provenance": { - "origin": "IPAW Steering Committee", - "actor": "system-admin", - "chain": ["cfp-announcement", "programme-published", "post-event-summary"] - }, - "spatial": { - "lat": 55.9445, - "lon": -3.1872, - "label": "University of Edinburgh, Informatics Forum, Edinburgh" - }, - "intentional_drift": false - }, - { - "id": "entity-045", - "category": "event", - "document": { - "title": "ACM SIGSPATIAL 2025 — 33rd International Conference on Advances in Geographic Information Systems", - "content": "SIGSPATIAL 2025 was held in Atlanta, USA. The R-tree indexing paper (entity-007) and the PolyBench benchmark paper (entity-010) were presented. The conference included a panel on 'Spatial Data in Multimodal Databases' with participants from the VeriSimDB team.", - "tags": ["conference", "SIGSPATIAL", "2025", "Atlanta", "spatial", "GIS"] - }, - "graph": { - "edges": [ - {"target": "entity-007", "relation": "featured_paper"}, - {"target": "entity-010", "relation": "featured_paper"}, - {"target": "entity-019", "relation": "presented_by"} - ] - }, - "vector": { - "embedding": [0.53, 0.51, 0.43, 0.36, 0.63, 0.46, 0.39, 0.49] - }, - "semantic": { - "types": ["https://schema.org/Event", "https://schema.org/BusinessEvent"], - "properties": { - "startDate": "2025-11-04", - "endDate": "2025-11-07", - "attendees": 600, - "url": "https://sigspatial2025.org/" - } - }, - "temporal": { - "created": "2025-04-01T00:00:00Z", - "modified": "2025-12-01T14:00:00Z", - "version": 3 - }, - "provenance": { - "origin": "ACM SIGSPATIAL", - "actor": "system-admin", - "chain": ["cfp-announcement", "programme-published", "post-event-summary"] - }, - "spatial": { - "lat": 33.7490, - "lon": -84.3880, - "label": "Georgia World Congress Center, Atlanta, GA, USA" - }, - "intentional_drift": false - }, - { - "id": "entity-046", - "category": "event", - "document": { - "title": "Turing Multimodal Data Workshop 2026", - "content": "An invitation-only workshop at the Alan Turing Institute bringing together the VeriSimDB research team to plan the next phase of development. Topics included VCL-UT finalisation, drift detection threshold calibration, and the roadmap for VeriSimDB v0.2.0. Attended by all 10 researchers in this dataset.", - "tags": ["workshop", "Turing", "VeriSimDB", "planning", "2026"] - }, - "graph": { - "edges": [ - {"target": "entity-011", "relation": "attended_by"}, - {"target": "entity-012", "relation": "attended_by"}, - {"target": "entity-013", "relation": "attended_by"}, - {"target": "entity-014", "relation": "attended_by"}, - {"target": "entity-015", "relation": "attended_by"}, - {"target": "entity-026", "relation": "hosted_by"} - ] - }, - "vector": { - "embedding": [0.58, 0.48, 0.44, 0.39, 0.67, 0.40, 0.42, 0.54] - }, - "semantic": { - "types": ["https://schema.org/Event"], - "properties": { - "startDate": "2026-01-20", - "endDate": "2026-01-21", - "attendees": 25, - "url": null - } - }, - "temporal": { - "created": "2025-11-15T00:00:00Z", - "modified": "2026-01-25T16:00:00Z", - "version": 2 - }, - "provenance": { - "origin": "Alan Turing Institute Events Office", - "actor": "entity-012", - "chain": ["invitation-sent", "post-workshop-notes"] - }, - "spatial": { - "lat": 51.5299, - "lon": -0.1278, - "label": "The Alan Turing Institute, British Library, London" - }, - "intentional_drift": false - }, - { - "id": "entity-047", - "category": "event", - "document": { - "title": "London Data Engineering Meetup — February 2026", - "content": "A monthly meetup for data engineers in London. The February 2026 edition featured a talk by Dr. Nia Williams on 'VeriSimDB: From Research Prototype to Production System', covering the Rust core architecture, Elixir orchestration layer, and deployment strategies. Over 80 attendees, with lively Q&A about drift detection in production.", - "tags": ["meetup", "London", "data engineering", "VeriSimDB", "production"] - }, - "graph": { - "edges": [ - {"target": "entity-020", "relation": "presented_by"}, - {"target": "entity-021", "relation": "sponsored_by"} - ] - }, - "vector": { - "embedding": [0.56, 0.46, 0.47, 0.41, 0.64, 0.38, 0.41, 0.55] - }, - "semantic": { - "types": ["https://schema.org/Event", "https://schema.org/SocialEvent"], - "properties": { - "startDate": "2026-02-18", - "endDate": "2026-02-18", - "attendees": 80, - "url": "https://meetup.com/london-data-engineering/events/feb-2026" - } - }, - "temporal": { - "created": "2026-01-15T00:00:00Z", - "modified": "2026-02-20T10:00:00Z", - "version": 2 - }, - "provenance": { - "origin": "Meetup.com event page", - "actor": "system-admin", - "chain": ["event-created", "post-event-update"] - }, - "spatial": { - "lat": 51.5177, - "lon": -0.0869, - "label": "Skills Matter CodeNode, 10 South Place, London EC2M 7EB" - }, - "intentional_drift": false - }, - { - "id": "entity-048", - "category": "event", - "document": { - "title": "POPL 2026 — 53rd ACM SIGPLAN Symposium on Principles of Programming Languages", - "content": "POPL 2026 was held in Denver, USA. While not a database conference, it featured a pearl paper on dependent types for database query verification, co-authored by Dr. MacLeod and Dr. Kapoor. This work underpins the PROOF clause semantics in VCL-UT.", - "tags": ["conference", "POPL", "2026", "Denver", "programming languages", "types"] - }, - "graph": { - "edges": [ - {"target": "entity-013", "relation": "presented_by"}, - {"target": "entity-014", "relation": "presented_by"}, - {"target": "entity-002", "relation": "builds_on"} - ] - }, - "vector": { - "embedding": [0.60, 0.45, 0.50, 0.42, 0.58, 0.35, 0.48, 0.56] - }, - "semantic": { - "types": ["https://schema.org/Event", "https://schema.org/BusinessEvent"], - "properties": { - "startDate": "2026-01-19", - "endDate": "2026-01-25", - "attendees": 700, - "url": "https://popl26.sigplan.org/" - } - }, - "temporal": { - "created": "2025-05-01T00:00:00Z", - "modified": "2026-02-01T08:00:00Z", - "version": 3 - }, - "provenance": { - "origin": "ACM SIGPLAN", - "actor": "system-admin", - "chain": ["cfp-announcement", "acceptance-notification", "post-event-summary"] - }, - "spatial": { - "lat": 39.7392, - "lon": -104.9903, - "label": "Colorado Convention Center, Denver, CO, USA" - }, - "intentional_drift": false - }, - { - "id": "entity-049", - "category": "event", - "document": { - "title": "EDBT 2026 — 29th International Conference on Extending Database Technology", - "content": "EDBT 2026 is scheduled for Barcelona, Spain. A demonstration paper on the VeriSimDB playground — an interactive web UI for exploring octad data — has been accepted. The demo showcases real-time drift visualisation and VCL query editing.", - "tags": ["conference", "EDBT", "2026", "Barcelona", "demo", "VeriSimDB"] - }, - "graph": { - "edges": [ - {"target": "entity-011", "relation": "presented_by"}, - {"target": "entity-020", "relation": "presented_by"} - ] - }, - "vector": { - "embedding": [0.54, 0.49, 0.46, 0.40, 0.61, 0.44, 0.40, 0.53] - }, - "semantic": { - "types": ["https://schema.org/Event", "https://schema.org/BusinessEvent"], - "properties": { - "startDate": "2026-03-24", - "endDate": "2026-03-27", - "attendees": 500, - "url": "https://edbt2026.org/" - } - }, - "temporal": { - "created": "2025-08-01T00:00:00Z", - "modified": "2026-02-22T11:00:00Z", - "version": 4 - }, - "provenance": { - "origin": "OpenProceedings.org", - "actor": "system-admin", - "chain": ["cfp-announcement", "demo-submission", "acceptance", "programme-update"] - }, - "spatial": { - "lat": 41.3874, - "lon": 2.1686, - "label": "Centre de Convencions Internacional de Barcelona, Spain" - }, - "intentional_drift": false - }, - { - "id": "entity-050", - "category": "event", - "document": { - "title": "VeriSimDB Community Hackathon 2026", - "content": "A two-day community hackathon at Imperial College London, open to anyone interested in contributing to VeriSimDB. Teams worked on VCL parser improvements, new drift detection heuristics, spatial index benchmarks, and documentation. The event produced 23 pull requests, 15 of which were merged.", - "tags": ["hackathon", "community", "VeriSimDB", "open source", "Imperial"] - }, - "graph": { - "edges": [ - {"target": "entity-011", "relation": "organised_by"}, - {"target": "entity-020", "relation": "organised_by"}, - {"target": "entity-021", "relation": "hosted_by"}, - {"target": "entity-029", "relation": "funded_by"} - ] - }, - "vector": { - "embedding": [0.57, 0.44, 0.48, 0.43, 0.66, 0.36, 0.43, 0.57] - }, - "semantic": { - "types": ["https://schema.org/Event", "https://schema.org/Hackathon"], - "properties": { - "startDate": "2026-02-08", - "endDate": "2026-02-09", - "attendees": 45, - "url": "https://verisimdb.dev/hackathon-2026" - } - }, - "temporal": { - "created": "2025-12-01T00:00:00Z", - "modified": "2026-02-12T18:00:00Z", - "version": 3 - }, - "provenance": { - "origin": "VeriSimDB project governance", - "actor": "entity-011", - "chain": ["planning", "registration-open", "post-event-report"] - }, - "spatial": { - "lat": 51.4988, - "lon": -0.1749, - "label": "Imperial College London, Huxley Building, South Kensington" - }, - "intentional_drift": false - } -] diff --git a/verisimdb/examples/smoke-test.sh b/verisimdb/examples/smoke-test.sh deleted file mode 100755 index d92af010..00000000 --- a/verisimdb/examples/smoke-test.sh +++ /dev/null @@ -1,170 +0,0 @@ -#!/usr/bin/env bash -# SPDX-License-Identifier: MPL-2.0 -# -# smoke-test.sh — VeriSimDB smoke test. -# -# Validates that the VeriSimDB server is running and can handle basic -# operations: health check, hexad CRUD, VCL query execution, drift -# endpoint, and orchestration layer health. Returns exit 0 on success, -# non-zero on failure. -# -# Suitable for CI integration — the exit code reflects the test outcome. -# Individual test results are printed to stdout. -# -# Usage: -# ./smoke-test.sh [--api-url URL] -# -# Examples: -# ./smoke-test.sh -# ./smoke-test.sh --api-url http://192.168.1.10:8080/api/v1 - -set -euo pipefail - -# --------------------------------------------------------------------------- -# Argument parsing -# --------------------------------------------------------------------------- - -API_URL="http://localhost:8080/api/v1" - -while [[ $# -gt 0 ]]; do - case "$1" in - --api-url) - API_URL="$2" - shift 2 - ;; - *) - API_URL="$1" - shift - ;; - esac -done - -FAILURES=0 - -# --------------------------------------------------------------------------- -# Helper functions -# --------------------------------------------------------------------------- - -pass() { echo " PASS: $1"; } -fail() { echo " FAIL: $1"; FAILURES=$((FAILURES + 1)); } - -# --------------------------------------------------------------------------- -# Tests -# --------------------------------------------------------------------------- - -echo "=== VeriSimDB Smoke Test ===" -echo "API URL: ${API_URL}" -echo "Timestamp: $(date -u +%Y-%m-%dT%H:%M:%SZ)" -echo "" - -# 1. Health check — verify the Rust core is responding -echo "1. Health check" -if curl -sf "${API_URL}/health" | jq -e '.status' >/dev/null 2>&1; then - pass "Server is healthy" -else - fail "Server health check failed — is VeriSimDB running?" -fi - -# 2. Create a hexad — test the write path -echo "2. Create hexad" -CREATE_PAYLOAD='{ - "id": "smoke-test-001", - "category": "test", - "document": { - "title": "Smoke Test Entity", - "content": "This entity is created by the VeriSimDB smoke test and should be deleted after the test completes.", - "tags": ["smoke-test", "ephemeral"] - }, - "graph": { - "edges": [] - }, - "vector": { - "embedding": [0.1, 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8] - }, - "semantic": { - "types": ["https://schema.org/Thing"], - "properties": {"test": true} - }, - "temporal": { - "created": "2026-02-28T00:00:00Z", - "modified": "2026-02-28T00:00:00Z", - "version": 1 - }, - "provenance": { - "origin": "smoke-test.sh", - "actor": "ci-runner", - "chain": ["created-by-smoke-test"] - }, - "spatial": { - "lat": 51.5074, - "lon": -0.1278, - "label": "Central London (test)" - } -}' - -CREATE_RESP=$(curl -sf -X POST "${API_URL}/hexads" \ - -H "Content-Type: application/json" \ - -d "$CREATE_PAYLOAD" 2>/dev/null || echo "FAILED") - -if [ "$CREATE_RESP" != "FAILED" ]; then - pass "Hexad created" -else - fail "Hexad creation failed" -fi - -# 3. Read the hexad back — test the read path -echo "3. Read hexad" -if curl -sf "${API_URL}/hexads/smoke-test-001" | jq -e '.id' >/dev/null 2>&1; then - pass "Hexad readable" -else - fail "Hexad read failed" -fi - -# 4. Execute a VCL query — test the query engine -echo "4. VCL query" -VCL_RESP=$(curl -sf -X POST "${API_URL}/vcl/execute" \ - -H "Content-Type: application/json" \ - -d '{"query":"SELECT * FROM hexads LIMIT 5"}' 2>/dev/null || echo "FAILED") - -if [ "$VCL_RESP" != "FAILED" ]; then - pass "VCL query executed" -else - fail "VCL query failed" -fi - -# 5. Check drift endpoint — test drift scoring -echo "5. Drift check" -if curl -sf "${API_URL}/drift/entity/smoke-test-001" >/dev/null 2>&1; then - pass "Drift endpoint responds" -else - fail "Drift endpoint failed" -fi - -# 6. Telemetry / orchestration layer — test the Elixir layer (may not be -# running in all environments, so a failure here is noted but expected) -ORCH_URL="${API_URL/8080/4080}" -ORCH_URL="${ORCH_URL/\/api\/v1/}" -echo "6. Telemetry (orchestration layer)" -if curl -sf "${ORCH_URL}/health" | jq -e '.status' >/dev/null 2>&1; then - pass "Orchestration layer healthy" -else - fail "Orchestration layer unreachable (expected if Elixir not running)" -fi - -# 7. Delete test entity — clean up after ourselves -echo "7. Cleanup" -curl -sf -X DELETE "${API_URL}/hexads/smoke-test-001" >/dev/null 2>&1 || true -pass "Cleanup attempted" - -# --------------------------------------------------------------------------- -# Summary -# --------------------------------------------------------------------------- - -echo "" -if [ $FAILURES -eq 0 ]; then - echo "=== ALL PASSED ===" - exit 0 -else - echo "=== ${FAILURES} FAILURE(S) ===" - exit 1 -fi diff --git a/verisimdb/examples/vcl-queries/01-basic-search.vcl b/verisimdb/examples/vcl-queries/01-basic-search.vcl deleted file mode 100644 index ba2ec3df..00000000 --- a/verisimdb/examples/vcl-queries/01-basic-search.vcl +++ /dev/null @@ -1,18 +0,0 @@ --- SPDX-License-Identifier: MPL-2.0 --- --- 01-basic-search.vcl --- Basic text search: find entities mentioning "drift detection" --- --- This searches the document modality's full-text index (Tantivy). --- VeriSimDB indexes the 'content' and 'title' fields of each hexad's --- document modality using Tantivy, a Rust-native full-text search engine. --- The CONTAINS clause performs a phrase search by default; use OR/AND --- operators for boolean queries. --- --- Expected results: entity-001 (VeriSimDB paper), entity-005 (semantic --- drift paper), entity-040 (DriftSeed dataset), entity-047 (meetup), --- and others mentioning drift detection in their document content. - -SELECT document, semantic FROM hexads -WHERE document CONTAINS 'drift detection' -LIMIT 10 diff --git a/verisimdb/examples/vcl-queries/02-vector-similarity.vcl b/verisimdb/examples/vcl-queries/02-vector-similarity.vcl deleted file mode 100644 index 978febcd..00000000 --- a/verisimdb/examples/vcl-queries/02-vector-similarity.vcl +++ /dev/null @@ -1,19 +0,0 @@ --- SPDX-License-Identifier: MPL-2.0 --- --- 02-vector-similarity.vcl --- Vector similarity: find entities with embeddings close to entity-001 --- --- Uses HNSW (Hierarchical Navigable Small World) approximate nearest- --- neighbour search over the vector modality. VeriSimDB stores embeddings --- in an HNSW index for sub-millisecond similarity queries even at --- million-entity scale. --- --- The SIMILAR TO ENTITY clause looks up the target entity's embedding --- and uses it as the query vector. You can also supply a raw vector --- with SIMILAR TO VECTOR([0.91, 0.23, ...]). --- --- Expected results: entity-008 (VCL paper, same lead author), entity-003 --- (hexad embeddings, same lab), and other entities with nearby embeddings. - -SEARCH VECTOR SIMILAR TO ENTITY 'entity-001' -LIMIT 5 diff --git a/verisimdb/examples/vcl-queries/03-cross-modal-join.vcl b/verisimdb/examples/vcl-queries/03-cross-modal-join.vcl deleted file mode 100644 index 299c6e44..00000000 --- a/verisimdb/examples/vcl-queries/03-cross-modal-join.vcl +++ /dev/null @@ -1,21 +0,0 @@ --- SPDX-License-Identifier: MPL-2.0 --- --- 03-cross-modal-join.vcl --- Cross-modal join: graph edges + vector similarity combined --- --- This query demonstrates VeriSimDB's ability to combine constraints --- from different modalities in a single query. Here we find entities --- that are connected to entity-011 (Dr. Amara Okonkwo) via graph edges --- AND have embeddings within a cosine similarity threshold of 0.7. --- --- The query planner resolves this by first traversing graph edges to --- find connected entities, then filtering by vector similarity. The --- execution order is optimised based on estimated cardinality. --- --- Expected results: entity-001, entity-003, entity-006, entity-008 --- (papers authored by Dr. Okonkwo with similar embeddings). - -SELECT graph, vector, document FROM hexads -WHERE graph.edge_to = 'entity-011' -AND vector.similarity > 0.7 -LIMIT 20 diff --git a/verisimdb/examples/vcl-queries/04-drift-detection.vcl b/verisimdb/examples/vcl-queries/04-drift-detection.vcl deleted file mode 100644 index 21fb0ee0..00000000 --- a/verisimdb/examples/vcl-queries/04-drift-detection.vcl +++ /dev/null @@ -1,25 +0,0 @@ --- SPDX-License-Identifier: MPL-2.0 --- --- 04-drift-detection.vcl --- Drift detection: find all entities with cross-modal drift above 0.3 --- --- VeriSimDB continuously computes a drift_score for each hexad entity. --- The drift score measures cross-modal inconsistency: how much the --- different modality representations of the same entity have diverged --- from each other. A score of 0.0 means perfect consistency; 1.0 means --- the modalities are completely unrelated. --- --- Common drift causes: --- - Vector embedding computed from an older version of the document --- - Semantic types that no longer match the document content --- - Temporal version count mismatching provenance chain length --- - Spatial coordinates inconsistent with document-described location --- --- Expected results: entities with intentional_drift=true in the sample --- data should appear here, including entity-003 (stale embedding), --- entity-005 (affiliation mismatch), entity-009 (embedding from wrong --- version), entity-029 (spatial coordinates wrong), etc. - -SELECT id, drift_score, document.title FROM hexads -WHERE drift_score > 0.3 -ORDER BY drift_score DESC diff --git a/verisimdb/examples/vcl-queries/05-provenance-chain.vcl b/verisimdb/examples/vcl-queries/05-provenance-chain.vcl deleted file mode 100644 index f46ac162..00000000 --- a/verisimdb/examples/vcl-queries/05-provenance-chain.vcl +++ /dev/null @@ -1,21 +0,0 @@ --- SPDX-License-Identifier: MPL-2.0 --- --- 05-provenance-chain.vcl --- Provenance chain: trace the full origin and transformation history --- --- The provenance modality records the complete lineage of each entity: --- where it originated, who created or modified it, and the ordered --- sequence of transformations it underwent. This is modelled after the --- W3C PROV-O ontology. --- --- This query retrieves the provenance chain for entity-031 (the --- PolyBench benchmark dataset), showing how it was generated, had --- drift injected, and was released as v1.0. --- --- Expected result: --- origin: "Imperial College London — Data Systems Lab" --- actor: "entity-020" (Dr. Nia Williams) --- chain: ["initial-generation", "drift-injection", "v1.0-release"] - -SELECT provenance FROM hexads -WHERE id = 'entity-031' diff --git a/verisimdb/examples/vcl-queries/06-temporal-range.vcl b/verisimdb/examples/vcl-queries/06-temporal-range.vcl deleted file mode 100644 index c378699d..00000000 --- a/verisimdb/examples/vcl-queries/06-temporal-range.vcl +++ /dev/null @@ -1,21 +0,0 @@ --- SPDX-License-Identifier: MPL-2.0 --- --- 06-temporal-range.vcl --- Temporal range: find entities modified in the last 30 days --- --- The temporal modality tracks creation time, last modification time, --- and version number for each hexad entity. This enables time-bounded --- queries that find recently changed data — useful for incremental --- synchronisation, change feeds, and audit trails. --- --- The temporal.modified field is an ISO 8601 timestamp. VeriSimDB --- indexes it in a B-tree for efficient range queries. --- --- Expected results: many entities in the sample data were modified in --- February 2026, so this query (relative to 2026-02-28) should return --- entity-001, entity-008, entity-011, and others recently updated. - -SELECT temporal, document.title FROM hexads -WHERE temporal.modified > '2026-01-28T00:00:00Z' -ORDER BY temporal.modified DESC -LIMIT 20 diff --git a/verisimdb/examples/vcl-queries/07-spatial-radius.vcl b/verisimdb/examples/vcl-queries/07-spatial-radius.vcl deleted file mode 100644 index 304ad534..00000000 --- a/verisimdb/examples/vcl-queries/07-spatial-radius.vcl +++ /dev/null @@ -1,23 +0,0 @@ --- SPDX-License-Identifier: MPL-2.0 --- --- 07-spatial-radius.vcl --- Spatial radius: find entities within 5km of central London --- --- The spatial modality stores latitude/longitude coordinates and an --- optional label for each hexad entity. VeriSimDB indexes spatial data --- using an R-tree (specifically the Hexad R-tree variant described in --- entity-007) for efficient radius and bounding-box queries. --- --- The WITHIN RADIUS clause takes three arguments: --- 1. Latitude of the centre point (decimal degrees) --- 2. Longitude of the centre point (decimal degrees) --- 3. Radius in metres --- --- Central London coordinates (51.5074, -0.1278) with a 5km radius --- should capture entities at Imperial College, UCL, KCL, the Turing --- Institute, and the Skills Matter venue, but exclude Edinburgh, --- Oxford, Cambridge, and international locations. - -SELECT spatial, document.title FROM hexads -WHERE spatial WITHIN RADIUS(51.5074, -0.1278, 5000) -LIMIT 15 diff --git a/verisimdb/examples/vcl-queries/08-proof-existence.vcl b/verisimdb/examples/vcl-queries/08-proof-existence.vcl deleted file mode 100644 index 40d01a47..00000000 --- a/verisimdb/examples/vcl-queries/08-proof-existence.vcl +++ /dev/null @@ -1,24 +0,0 @@ --- SPDX-License-Identifier: MPL-2.0 --- --- 08-proof-existence.vcl --- Proof-carrying query: verify entity existence across all modalities --- --- VCL-UT (Dependent Type) extends VCL with PROOF clauses that attach --- machine-checkable proof certificates to query results. This is --- VeriSimDB's distinguishing feature: not just returning data, but --- returning data with formal guarantees. --- --- PROOF EXISTENCE(entity-id) verifies that: --- 1. The entity exists in the primary store --- 2. All eight modality representations are present and non-null --- 3. The entity ID is consistent across all modality indices --- --- The proof certificate is returned as a separate field in the result --- and can be independently verified by any Idris2-compatible checker. --- --- Expected result: entity-001 exists with all modalities populated, --- so the proof should succeed and the certificate should be attached. - -SELECT * FROM hexads -WHERE id = 'entity-001' -PROOF EXISTENCE(entity-001) diff --git a/verisimdb/examples/vcl-queries/09-proof-consistency.vcl b/verisimdb/examples/vcl-queries/09-proof-consistency.vcl deleted file mode 100644 index 0f70522b..00000000 --- a/verisimdb/examples/vcl-queries/09-proof-consistency.vcl +++ /dev/null @@ -1,24 +0,0 @@ --- SPDX-License-Identifier: MPL-2.0 --- --- 09-proof-consistency.vcl --- Proof of consistency: verify graph-semantic coherence for entity-005 --- --- PROOF CONSISTENCY(entity-id) verifies that two or more modalities --- of the same entity are mutually consistent. Specifically, it checks: --- - Graph edge targets exist as entities (referential integrity) --- - Semantic type annotations are valid schema.org types --- - Graph relationships are compatible with semantic types --- (e.g., an entity typed as 'Person' should have 'author_of' edges, --- not 'employs' edges) --- --- Entity-005 has intentional drift: its semantic types and graph edges --- contain an affiliation inconsistency (Oxford in graph, Edinburgh in --- provenance). The PROOF CONSISTENCY clause should FAIL for this entity --- and return a detailed explanation of the inconsistency. --- --- Expected result: proof failure with diagnostic information about --- the affiliation mismatch between graph and provenance modalities. - -SELECT graph, semantic FROM hexads -WHERE id = 'entity-005' -PROOF CONSISTENCY(entity-005) diff --git a/verisimdb/examples/vcl-queries/10-multi-modal-pipeline.vcl b/verisimdb/examples/vcl-queries/10-multi-modal-pipeline.vcl deleted file mode 100644 index eff7eef2..00000000 --- a/verisimdb/examples/vcl-queries/10-multi-modal-pipeline.vcl +++ /dev/null @@ -1,32 +0,0 @@ --- SPDX-License-Identifier: MPL-2.0 --- --- 10-multi-modal-pipeline.vcl --- Multi-modal pipeline: query all 8 octad modalities for a single entity --- --- This is the "full octad query" — VeriSimDB's most comprehensive query --- type. It retrieves all eight modality representations of an entity in --- a single result, demonstrating the database's unique ability to unify --- heterogeneous data models under one roof. --- --- The eight modalities are: --- 1. graph — edges connecting this entity to others --- 2. vector — dense embedding for similarity search --- 3. tensor — multi-dimensional array representation --- 4. semantic — schema.org types and structured properties --- 5. document — full-text title, content, and tags --- 6. temporal — creation time, modification time, version --- 7. provenance — origin, actor, and transformation chain --- 8. spatial — latitude, longitude, and location label --- --- The PROOF clause combines two checks: --- - EXISTENCE: all modalities are present and populated --- - PROVENANCE: the provenance chain is well-formed and each step --- references a valid actor --- --- Expected result: complete octad data for entity-001 (the VeriSimDB --- paper) with both proof certificates attached. - -SELECT graph, vector, tensor, semantic, document, temporal, provenance, spatial -FROM hexads -WHERE id = 'entity-001' -PROOF EXISTENCE(entity-001) AND PROVENANCE(entity-001) diff --git a/verisimdb/examples/web-project-deno.json b/verisimdb/examples/web-project-deno.json deleted file mode 100644 index 5ddd3bd7..00000000 --- a/verisimdb/examples/web-project-deno.json +++ /dev/null @@ -1,20 +0,0 @@ -{ - "// NOTE": "Example deno.json for ReScript web projects", - "tasks": { - "build": "deno run -A npm:rescript", - "clean": "deno run -A npm:rescript clean", - "watch": "deno run -A npm:rescript -w", - "serve": "deno run -A jsr:@std/http/file-server .", - "test": "deno test --allow-all" - }, - "imports": { - "rescript": "^12.0.0", - "@rescript/core": "npm:@rescript/core@^1.6.0", - "safe-dom/": "https://raw.githubusercontent.com/hyperpolymath/rescript-dom-mounter/main/src/", - "proven/": "../proven/bindings/rescript/src/" - }, - "compilerOptions": { - "allowJs": true, - "checkJs": false - } -} diff --git a/verisimdb/ffi/zig/build.zig b/verisimdb/ffi/zig/build.zig deleted file mode 100644 index 4a2e049a..00000000 --- a/verisimdb/ffi/zig/build.zig +++ /dev/null @@ -1,94 +0,0 @@ -// {{PROJECT}} FFI Build Configuration -// SPDX-License-Identifier: MPL-2.0 - -const std = @import("std"); - -pub fn build(b: *std.Build) void { - const target = b.standardTargetOptions(.{}); - const optimize = b.standardOptimizeOption(.{}); - - // Shared library (.so, .dylib, .dll) - const lib = b.addSharedLibrary(.{ - .name = "{{project}}", - .root_source_file = b.path("src/main.zig"), - .target = target, - .optimize = optimize, - }); - - // Set version - lib.version = .{ .major = 0, .minor = 1, .patch = 0 }; - - // Static library (.a) - const lib_static = b.addStaticLibrary(.{ - .name = "{{project}}", - .root_source_file = b.path("src/main.zig"), - .target = target, - .optimize = optimize, - }); - - // Install artifacts - b.installArtifact(lib); - b.installArtifact(lib_static); - - // Generate header file for C compatibility - const header = b.addInstallHeader( - b.path("include/{{project}}.h"), - "{{project}}.h", - ); - b.getInstallStep().dependOn(&header.step); - - // Unit tests - const lib_tests = b.addTest(.{ - .root_source_file = b.path("src/main.zig"), - .target = target, - .optimize = optimize, - }); - - const run_lib_tests = b.addRunArtifact(lib_tests); - - const test_step = b.step("test", "Run library tests"); - test_step.dependOn(&run_lib_tests.step); - - // Integration tests - const integration_tests = b.addTest(.{ - .root_source_file = b.path("test/integration_test.zig"), - .target = target, - .optimize = optimize, - }); - - integration_tests.linkLibrary(lib); - - const run_integration_tests = b.addRunArtifact(integration_tests); - - const integration_test_step = b.step("test-integration", "Run integration tests"); - integration_test_step.dependOn(&run_integration_tests.step); - - // Documentation - const docs = b.addTest(.{ - .root_source_file = b.path("src/main.zig"), - .target = target, - .optimize = .Debug, - }); - - const docs_step = b.step("docs", "Generate documentation"); - docs_step.dependOn(&b.addInstallDirectory(.{ - .source_dir = docs.getEmittedDocs(), - .install_dir = .prefix, - .install_subdir = "docs", - }).step); - - // Benchmark (if needed) - const bench = b.addExecutable(.{ - .name = "{{project}}-bench", - .root_source_file = b.path("bench/bench.zig"), - .target = target, - .optimize = .ReleaseFast, - }); - - bench.linkLibrary(lib); - - const run_bench = b.addRunArtifact(bench); - - const bench_step = b.step("bench", "Run benchmarks"); - bench_step.dependOn(&run_bench.step); -} diff --git a/verisimdb/ffi/zig/src/main.zig b/verisimdb/ffi/zig/src/main.zig deleted file mode 100644 index 6b233bc7..00000000 --- a/verisimdb/ffi/zig/src/main.zig +++ /dev/null @@ -1,274 +0,0 @@ -// {{PROJECT}} FFI Implementation -// -// This module implements the C-compatible FFI declared in src/abi/Foreign.idr -// All types and layouts must match the Idris2 ABI definitions. -// -// SPDX-License-Identifier: MPL-2.0 - -const std = @import("std"); - -// Version information (keep in sync with project) -const VERSION = "0.1.0"; -const BUILD_INFO = "{{PROJECT}} built with Zig " ++ @import("builtin").zig_version_string; - -/// Thread-local error storage -threadlocal var last_error: ?[]const u8 = null; - -/// Set the last error message -fn setError(msg: []const u8) void { - last_error = msg; -} - -/// Clear the last error -fn clearError() void { - last_error = null; -} - -//============================================================================== -// Core Types (must match src/abi/Types.idr) -//============================================================================== - -/// Result codes (must match Idris2 Result type) -pub const Result = enum(c_int) { - ok = 0, - @"error" = 1, - invalid_param = 2, - out_of_memory = 3, - null_pointer = 4, -}; - -/// Library handle (opaque to prevent direct access) -pub const Handle = opaque { - // Internal state hidden from C - allocator: std.mem.Allocator, - initialized: bool, - // Add your fields here -}; - -//============================================================================== -// Library Lifecycle -//============================================================================== - -/// Initialize the library -/// Returns a handle, or null on failure -export fn {{project}}_init() ?*Handle { - const allocator = std.heap.c_allocator; - - const handle = allocator.create(Handle) catch { - setError("Failed to allocate handle"); - return null; - }; - - // Initialize handle - handle.* = .{ - .allocator = allocator, - .initialized = true, - }; - - clearError(); - return handle; -} - -/// Free the library handle -export fn {{project}}_free(handle: ?*Handle) void { - const h = handle orelse return; - const allocator = h.allocator; - - // Clean up resources - h.initialized = false; - - allocator.destroy(h); - clearError(); -} - -//============================================================================== -// Core Operations -//============================================================================== - -/// Process data (example operation) -export fn {{project}}_process(handle: ?*Handle, input: u32) Result { - const h = handle orelse { - setError("Null handle"); - return .null_pointer; - }; - - if (!h.initialized) { - setError("Handle not initialized"); - return .@"error"; - } - - // Example processing logic - _ = input; - - clearError(); - return .ok; -} - -//============================================================================== -// String Operations -//============================================================================== - -/// Get a string result (example) -/// Caller must free the returned string -export fn {{project}}_get_string(handle: ?*Handle) ?[*:0]const u8 { - const h = handle orelse { - setError("Null handle"); - return null; - }; - - if (!h.initialized) { - setError("Handle not initialized"); - return null; - } - - // Example: allocate and return a string - const result = h.allocator.dupeZ(u8, "Example result") catch { - setError("Failed to allocate string"); - return null; - }; - - clearError(); - return result.ptr; -} - -/// Free a string allocated by the library -export fn {{project}}_free_string(str: ?[*:0]const u8) void { - const s = str orelse return; - const allocator = std.heap.c_allocator; - - const slice = std.mem.span(s); - allocator.free(slice); -} - -//============================================================================== -// Array/Buffer Operations -//============================================================================== - -/// Process an array of data -export fn {{project}}_process_array( - handle: ?*Handle, - buffer: ?[*]const u8, - len: u32, -) Result { - const h = handle orelse { - setError("Null handle"); - return .null_pointer; - }; - - const buf = buffer orelse { - setError("Null buffer"); - return .null_pointer; - }; - - if (!h.initialized) { - setError("Handle not initialized"); - return .@"error"; - } - - // Access the buffer - const data = buf[0..len]; - _ = data; - - // Process data here - - clearError(); - return .ok; -} - -//============================================================================== -// Error Handling -//============================================================================== - -/// Get the last error message -/// Returns null if no error -export fn {{project}}_last_error() ?[*:0]const u8 { - const err = last_error orelse return null; - - // Return C string (static storage, no need to free) - const allocator = std.heap.c_allocator; - const c_str = allocator.dupeZ(u8, err) catch return null; - return c_str.ptr; -} - -//============================================================================== -// Version Information -//============================================================================== - -/// Get the library version -export fn {{project}}_version() [*:0]const u8 { - return VERSION.ptr; -} - -/// Get build information -export fn {{project}}_build_info() [*:0]const u8 { - return BUILD_INFO.ptr; -} - -//============================================================================== -// Callback Support -//============================================================================== - -/// Callback function type (C ABI) -pub const Callback = *const fn (u64, u32) callconv(.C) u32; - -/// Register a callback -export fn {{project}}_register_callback( - handle: ?*Handle, - callback: ?Callback, -) Result { - const h = handle orelse { - setError("Null handle"); - return .null_pointer; - }; - - const cb = callback orelse { - setError("Null callback"); - return .null_pointer; - }; - - if (!h.initialized) { - setError("Handle not initialized"); - return .@"error"; - } - - // Store callback for later use - _ = cb; - - clearError(); - return .ok; -} - -//============================================================================== -// Utility Functions -//============================================================================== - -/// Check if handle is initialized -export fn {{project}}_is_initialized(handle: ?*Handle) u32 { - const h = handle orelse return 0; - return if (h.initialized) 1 else 0; -} - -//============================================================================== -// Tests -//============================================================================== - -test "lifecycle" { - const handle = {{project}}_init() orelse return error.InitFailed; - defer {{project}}_free(handle); - - try std.testing.expect({{project}}_is_initialized(handle) == 1); -} - -test "error handling" { - const result = {{project}}_process(null, 0); - try std.testing.expectEqual(Result.null_pointer, result); - - const err = {{project}}_last_error(); - try std.testing.expect(err != null); -} - -test "version" { - const ver = {{project}}_version(); - const ver_str = std.mem.span(ver); - try std.testing.expectEqualStrings(VERSION, ver_str); -} diff --git a/verisimdb/ffi/zig/test/integration_test.zig b/verisimdb/ffi/zig/test/integration_test.zig deleted file mode 100644 index 03419949..00000000 --- a/verisimdb/ffi/zig/test/integration_test.zig +++ /dev/null @@ -1,182 +0,0 @@ -// {{PROJECT}} Integration Tests -// SPDX-License-Identifier: MPL-2.0 -// -// These tests verify that the Zig FFI correctly implements the Idris2 ABI - -const std = @import("std"); -const testing = std.testing; - -// Import FFI functions -extern fn {{project}}_init() ?*opaque {}; -extern fn {{project}}_free(?*opaque {}) void; -extern fn {{project}}_process(?*opaque {}, u32) c_int; -extern fn {{project}}_get_string(?*opaque {}) ?[*:0]const u8; -extern fn {{project}}_free_string(?[*:0]const u8) void; -extern fn {{project}}_last_error() ?[*:0]const u8; -extern fn {{project}}_version() [*:0]const u8; -extern fn {{project}}_is_initialized(?*opaque {}) u32; - -//============================================================================== -// Lifecycle Tests -//============================================================================== - -test "create and destroy handle" { - const handle = {{project}}_init() orelse return error.InitFailed; - defer {{project}}_free(handle); - - try testing.expect(handle != null); -} - -test "handle is initialized" { - const handle = {{project}}_init() orelse return error.InitFailed; - defer {{project}}_free(handle); - - const initialized = {{project}}_is_initialized(handle); - try testing.expectEqual(@as(u32, 1), initialized); -} - -test "null handle is not initialized" { - const initialized = {{project}}_is_initialized(null); - try testing.expectEqual(@as(u32, 0), initialized); -} - -//============================================================================== -// Operation Tests -//============================================================================== - -test "process with valid handle" { - const handle = {{project}}_init() orelse return error.InitFailed; - defer {{project}}_free(handle); - - const result = {{project}}_process(handle, 42); - try testing.expectEqual(@as(c_int, 0), result); // 0 = ok -} - -test "process with null handle returns error" { - const result = {{project}}_process(null, 42); - try testing.expectEqual(@as(c_int, 4), result); // 4 = null_pointer -} - -//============================================================================== -// String Tests -//============================================================================== - -test "get string result" { - const handle = {{project}}_init() orelse return error.InitFailed; - defer {{project}}_free(handle); - - const str = {{project}}_get_string(handle); - defer if (str) |s| {{project}}_free_string(s); - - try testing.expect(str != null); -} - -test "get string with null handle" { - const str = {{project}}_get_string(null); - try testing.expect(str == null); -} - -//============================================================================== -// Error Handling Tests -//============================================================================== - -test "last error after null handle operation" { - _ = {{project}}_process(null, 0); - - const err = {{project}}_last_error(); - try testing.expect(err != null); - - if (err) |e| { - const err_str = std.mem.span(e); - try testing.expect(err_str.len > 0); - } -} - -test "no error after successful operation" { - const handle = {{project}}_init() orelse return error.InitFailed; - defer {{project}}_free(handle); - - _ = {{project}}_process(handle, 0); - - // Error should be cleared after successful operation - // (This depends on implementation) -} - -//============================================================================== -// Version Tests -//============================================================================== - -test "version string is not empty" { - const ver = {{project}}_version(); - const ver_str = std.mem.span(ver); - - try testing.expect(ver_str.len > 0); -} - -test "version string is semantic version format" { - const ver = {{project}}_version(); - const ver_str = std.mem.span(ver); - - // Should be in format X.Y.Z - try testing.expect(std.mem.count(u8, ver_str, ".") >= 1); -} - -//============================================================================== -// Memory Safety Tests -//============================================================================== - -test "multiple handles are independent" { - const h1 = {{project}}_init() orelse return error.InitFailed; - defer {{project}}_free(h1); - - const h2 = {{project}}_init() orelse return error.InitFailed; - defer {{project}}_free(h2); - - try testing.expect(h1 != h2); - - // Operations on h1 should not affect h2 - _ = {{project}}_process(h1, 1); - _ = {{project}}_process(h2, 2); -} - -test "double free is safe" { - const handle = {{project}}_init() orelse return error.InitFailed; - - {{project}}_free(handle); - {{project}}_free(handle); // Should not crash -} - -test "free null is safe" { - {{project}}_free(null); // Should not crash -} - -//============================================================================== -// Thread Safety Tests (if applicable) -//============================================================================== - -test "concurrent operations" { - const handle = {{project}}_init() orelse return error.InitFailed; - defer {{project}}_free(handle); - - const ThreadContext = struct { - h: *opaque {}, - id: u32, - }; - - const thread_fn = struct { - fn run(ctx: ThreadContext) void { - _ = {{project}}_process(ctx.h, ctx.id); - } - }.run; - - var threads: [4]std.Thread = undefined; - for (&threads, 0..) |*thread, i| { - thread.* = try std.Thread.spawn(.{}, thread_fn, .{ - ThreadContext{ .h = handle, .id = @intCast(i) }, - }); - } - - for (threads) |thread| { - thread.join(); - } -} diff --git a/verisimdb/fuzz/Cargo.toml b/verisimdb/fuzz/Cargo.toml deleted file mode 100644 index 0e5c04eb..00000000 --- a/verisimdb/fuzz/Cargo.toml +++ /dev/null @@ -1,28 +0,0 @@ -# SPDX-License-Identifier: MPL-2.0 -[package] -name = "verisimdb-fuzz" -version = "0.0.0" -publish = false -edition = "2021" - -[package.metadata] -cargo-fuzz = true - -[dependencies] -libfuzzer-sys = "0.4" - -[dependencies.verisim-octad] -path = "../rust-core/verisim-octad" - -# Prevent this from interfering with workspaces -[workspace] -members = ["."] - -[[bin]] -name = "fuzz_octad_id" -path = "fuzz_targets/fuzz_octad_id.rs" -test = false -doc = false - -[profile.release] -debug = 1 diff --git a/verisimdb/fuzz/fuzz_targets/fuzz_octad_id.rs b/verisimdb/fuzz/fuzz_targets/fuzz_octad_id.rs deleted file mode 100644 index 8894618c..00000000 --- a/verisimdb/fuzz/fuzz_targets/fuzz_octad_id.rs +++ /dev/null @@ -1,19 +0,0 @@ -// SPDX-License-Identifier: MPL-2.0 -// Fuzz target for Octad UUID parsing and validation - -#![no_main] - -use libfuzzer_sys::fuzz_target; - -fuzz_target!(|data: &[u8]| { - // Try to parse data as a UUID string - if let Ok(s) = std::str::from_utf8(data) { - // Test UUID parsing doesn't panic - let _ = uuid::Uuid::parse_str(s); - - // Test that malformed UUIDs are handled gracefully - if s.len() < 128 { - let _ = s.trim(); - } - } -}); diff --git a/verisimdb/historiographic-custodian.html b/verisimdb/historiographic-custodian.html deleted file mode 100644 index cf6667cd..00000000 --- a/verisimdb/historiographic-custodian.html +++ /dev/null @@ -1,258 +0,0 @@ -import React, { useState, useEffect } from 'react'; -import { - Shield, - AlertTriangle, - CheckCircle, - FileText, - Clock, - Users, - Search, - ChevronRight, - Fingerprint, - Info, - History -} from 'lucide-react'; - -const App = () => { - const [selectedIssue, setSelectedIssue] = useState(null); - const [isSigning, setIsSigning] = useState(false); - const [auditLog, setAuditLog] = useState([ - { id: 1, action: "Policy Updated", target: "0x882A...", actor: "did:verisim:custodian_02", time: "2h ago" }, - { id: 2, action: "Manual Repair", target: "0x441F...", actor: "did:verisim:custodian_01", time: "5h ago" } - ]); - - const pendingIssues = [ - { - id: "0x12AB...990F", - type: "Formal Drift (Contract Breach)", - modality: "Graph / Semantic", - severity: "High", - contract: "CitationContract", - breach: "Invariant 'claim_validity' failed.", - cause: "Reference 0x990F... has status 'retracted'.", - implication: "Approving this update will logically invalidate the authority of the parent Hexad.", - timestamp: "12 mins ago" - }, - { - id: "0xCC21...110E", - type: "Topological Mutation", - modality: "Graph", - severity: "Medium", - contract: "TaxonomyContract", - breach: "Edge-addition rate exceeds Poisson threshold (λ=0.05).", - cause: "Rapid batch update of 450 nodes detected.", - implication: "Possible systematic bias or archival error in metadata ingestion.", - timestamp: "45 mins ago" - } - ]; - - const handleSign = () => { - setIsSigning(true); - // Simulate sactify-php + proven ZKP generation - setTimeout(() => { - setAuditLog([ - { - id: Date.now(), - action: "Formal Repair Signed", - target: selectedIssue.id, - actor: "did:verisim:custodian_01", - time: "Just now" - }, - ...auditLog - ]); - setIsSigning(false); - setSelectedIssue(null); - }, 2000); - }; - - return ( -
- {/* Header */} - - -
- {/* Left: Pending Queue */} -
-
-

- - Pending Reviews -

- - {pendingIssues.length} New - -
- -
- {pendingIssues.map((issue) => ( - - ))} -
-
- - {/* Center: Detailed Evidence (Proven Explainability) */} -
- {selectedIssue ? ( -
-
-

Explainability Trace

-

Formal Verification Engine (Proven v1.1)

-
- -
-
-

Violation Details

-
- -
-

{selectedIssue.breach}

-

{selectedIssue.cause}

-
-
-
- -
-

Logical Implication

-
- -

{selectedIssue.implication}

-
-
- -
-

Proposed Repair

-
-

Revert citation mapping to state 0x4F2... and append retraction proof to the Temporal Ledger.

-
- - COHERENCE RESTORED AFTER REPAIR -
-
-
-
- -
- -
-
- ) : ( -
- -

Select a pending issue to review formal drift evidence and authorize repairs.

-
- )} -
- - {/* Right: Audit Log & Stats */} -
-
-

- - Recent Activity -

-
- {auditLog.map(log => ( -
-

{log.action}

-

Target: {log.target}

-
- {log.actor.split(':').pop()} - {log.time} -
-
- ))} -
-
- -
-

- - Epistemic Health -

-
-
-
- Formal Coherence - 98.2% -
-
-
-
-
-
-
- Quorums Reached - 24 / 24 -
-
-
-
-
-
-
-
-
-
- ); -}; - -const Activity = ({ className }) => ( - - - -); - -export default App; diff --git a/verisimdb/kraft-comparison.adoc b/verisimdb/kraft-comparison.adoc deleted file mode 100644 index dfb69b36..00000000 --- a/verisimdb/kraft-comparison.adoc +++ /dev/null @@ -1,69 +0,0 @@ -// SPDX-License-Identifier: CC-BY-SA-4.0 - -= KRaft vs VeriSimDB Comparison - -Comparison of Apache Kafka's KRaft (Kafka Raft) metadata management with VeriSimDB's federated registry approach. - -[cols="1,2,2",options="header"] -|=== -|Feature |KRaft (Kafka) |VeriSimDB (Proposed) - -|**The "Brain"** -|Metadata Quorum (3-5 nodes) -|Federated Registry Quorum - -|**Consistency** -|Strong (via Raft) -|Strong (via Paxos/Raft Integration) - -|**Data Format** -|Event-sourced Log -|Merkle Search Tree (MST) - -|**Scalability** -|Millions of partitions -|Millions of Octads - -|**Addressing** -|Broker/Topic Mapping -|UUID/Modality Mapping -|=== - -== Key Similarities - -Both KRaft and VeriSimDB: - -- Use consensus protocols for strong consistency -- Replicate metadata across multiple nodes -- Support horizontal scalability -- Provide fault-tolerant coordination - -== Key Differences - -**KRaft:** - -- Designed for streaming platform coordination -- Event log as primary data structure -- Broker-centric addressing -- Partition assignment management - -**VeriSimDB:** - -- Designed for federated knowledge stores -- Merkle Search Tree for verifiable state -- UUID-centric addressing (content-addressable) -- Modality store routing - -== Integration Opportunities - -VeriSimDB's federated registry could potentially: - -1. Use KRaft as a consensus backend (alternative to custom Raft implementation) -2. Leverage Kafka's mature Raft implementation -3. Benefit from Kafka's operational tooling -4. Reuse broker coordination patterns for store coordination - -== References - -- link:docs/technical-specification-kraft-metadata-log.adoc[VeriSimDB KRaft Integration Spec] -- https://cwiki.apache.org/confluence/display/KAFKA/KIP-500%3A+Replace+ZooKeeper+with+a+Self-Managed+Metadata+Quorum[KIP-500: Replace ZooKeeper with a Self-Managed Metadata Quorum] diff --git a/verisimdb/lib/verisim/adaptive_learner.ex b/verisimdb/lib/verisim/adaptive_learner.ex deleted file mode 100644 index e6a5bc0a..00000000 --- a/verisimdb/lib/verisim/adaptive_learner.ex +++ /dev/null @@ -1,516 +0,0 @@ -# SPDX-License-Identifier: MPL-2.0 - -defmodule VeriSim.AdaptiveLearner do - @moduledoc """ - Adaptive learning for VeriSimDB optimization. - - Uses feedback loops (observe → decide → act → measure) to learn: - - Cache TTL policies (hit rate optimization) - - Normalization thresholds (reduce false positives) - - Drift tolerance (balance repair cost vs consistency) - - Query plan selection (learn from execution times) - - Architecture: - - GenServer per learning domain - - Periodic measurement and adjustment - - Convergence detection (stop learning when stable) - - Rollback on regression - """ - - use GenServer - require Logger - - @type learning_domain :: - :cache_ttl - | :normalization_threshold - | :drift_tolerance - | :query_plan - - @type observation :: %{ - timestamp: DateTime.t(), - metrics: map(), - action_taken: atom(), - result: map() - } - - defmodule State do - @type t :: %__MODULE__{ - domain: atom(), - observations: [map()], - current_policy: map(), - baseline_policy: map(), - learning_rate: float(), - converged: boolean(), - last_adjustment: DateTime.t() | nil - } - - defstruct domain: nil, - observations: [], - current_policy: %{}, - baseline_policy: %{}, - learning_rate: 0.1, - converged: false, - last_adjustment: nil - end - - # ============================================================================ - # Client API - # ============================================================================ - - @doc """ - Start adaptive learner for a domain. - """ - def start_link(domain, initial_policy \\ %{}) do - GenServer.start_link(__MODULE__, {domain, initial_policy}, name: via_tuple(domain)) - end - - @doc """ - Record observation for learning. - """ - def observe(domain, metrics) do - GenServer.cast(via_tuple(domain), {:observe, metrics}) - end - - @doc """ - Get current learned policy. - """ - def get_policy(domain) do - GenServer.call(via_tuple(domain), :get_policy) - end - - @doc """ - Force policy adjustment (manual trigger). - """ - def adjust_policy(domain) do - GenServer.call(via_tuple(domain), :adjust_policy) - end - - @doc """ - Reset to baseline policy (rollback learning). - """ - def reset_policy(domain) do - GenServer.call(via_tuple(domain), :reset_policy) - end - - # ============================================================================ - # Server Callbacks - # ============================================================================ - - @impl true - def init({domain, initial_policy}) do - state = %State{ - domain: domain, - current_policy: initial_policy, - baseline_policy: initial_policy, - observations: [] - } - - # Schedule periodic learning - schedule_learning() - - Logger.info("Adaptive learner started for #{domain}") - {:ok, state} - end - - @impl true - def handle_cast({:observe, metrics}, state) do - observation = %{ - timestamp: DateTime.utc_now(), - metrics: metrics - } - - new_state = %{state | observations: [observation | state.observations] |> Enum.take(1000)} - - {:noreply, new_state} - end - - @impl true - def handle_call(:get_policy, _from, state) do - {:reply, state.current_policy, state} - end - - @impl true - def handle_call(:adjust_policy, _from, state) do - case adjust_policy_for_domain(state) do - {:ok, new_policy} -> - new_state = %{state | current_policy: new_policy, last_adjustment: DateTime.utc_now()} - {:reply, {:ok, new_policy}, new_state} - - {:error, reason} -> - {:reply, {:error, reason}, state} - end - end - - @impl true - def handle_call(:reset_policy, _from, state) do - Logger.warn("Resetting policy for #{state.domain} to baseline") - - new_state = %{ - state - | current_policy: state.baseline_policy, - converged: false, - last_adjustment: DateTime.utc_now() - } - - {:reply, :ok, new_state} - end - - @impl true - def handle_info(:learn, state) do - new_state = - if should_learn?(state) do - case adjust_policy_for_domain(state) do - {:ok, new_policy} -> - %{state | current_policy: new_policy, last_adjustment: DateTime.utc_now()} - - {:error, reason} -> - Logger.error("Learning failed for #{state.domain}: #{inspect(reason)}") - state - end - else - state - end - - schedule_learning() - {:noreply, new_state} - end - - # ============================================================================ - # Learning Algorithms (Domain-Specific) - # ============================================================================ - - defp adjust_policy_for_domain(%State{domain: :cache_ttl} = state) do - learn_cache_ttl(state) - end - - defp adjust_policy_for_domain(%State{domain: :normalization_threshold} = state) do - learn_normalization_threshold(state) - end - - defp adjust_policy_for_domain(%State{domain: :drift_tolerance} = state) do - learn_drift_tolerance(state) - end - - defp adjust_policy_for_domain(%State{domain: :query_plan} = state) do - learn_query_plan_selection(state) - end - - # ============================================================================ - # Cache TTL Learning - # ============================================================================ - - defp learn_cache_ttl(state) do - recent_obs = Enum.take(state.observations, 100) - - if length(recent_obs) < 10 do - {:error, :insufficient_data} - else - hit_rate = calculate_hit_rate(recent_obs) - current_ttl = Map.get(state.current_policy, :ttl_seconds, 300) - - new_ttl = - cond do - # Low hit rate → increase TTL (cache longer) - hit_rate < 0.4 -> - increase = current_ttl * state.learning_rate - min(current_ttl + increase, 3600) - - # Very high hit rate → decrease TTL (fresher data) - hit_rate > 0.9 -> - decrease = current_ttl * state.learning_rate - max(current_ttl - decrease, 30) - - # Hit rate in acceptable range → maintain - true -> - current_ttl - end - |> round() - - Logger.info("Cache TTL adjusted: #{current_ttl}s → #{new_ttl}s (hit_rate: #{hit_rate})") - - new_policy = Map.put(state.current_policy, :ttl_seconds, new_ttl) - {:ok, new_policy} - end - end - - defp calculate_hit_rate(observations) do - totals = - Enum.reduce(observations, %{hits: 0, misses: 0}, fn obs, acc -> - %{ - hits: acc.hits + Map.get(obs.metrics, :cache_hits, 0), - misses: acc.misses + Map.get(obs.metrics, :cache_misses, 0) - } - end) - - total_requests = totals.hits + totals.misses - - if total_requests > 0 do - totals.hits / total_requests - else - 0.0 - end - end - - # ============================================================================ - # Normalization Threshold Learning - # ============================================================================ - - defp learn_normalization_threshold(state) do - recent_obs = Enum.take(state.observations, 100) - - if length(recent_obs) < 10 do - {:error, :insufficient_data} - else - false_positive_rate = calculate_false_positive_rate(recent_obs) - current_threshold = Map.get(state.current_policy, :deviance_threshold, 0.1) - - new_threshold = - cond do - # Too many false positives → relax threshold - false_positive_rate > 0.2 -> - increase = current_threshold * state.learning_rate - min(current_threshold + increase, 0.5) - - # Too few alerts (may be missing real issues) → tighten - false_positive_rate < 0.05 -> - decrease = current_threshold * state.learning_rate - max(current_threshold - decrease, 0.01) - - true -> - current_threshold - end - - Logger.info( - "Normalization threshold adjusted: #{current_threshold} → #{new_threshold} (FP rate: #{false_positive_rate})" - ) - - new_policy = Map.put(state.current_policy, :deviance_threshold, new_threshold) - {:ok, new_policy} - end - end - - defp calculate_false_positive_rate(observations) do - alerts = - Enum.filter(observations, fn obs -> - Map.get(obs.metrics, :deviance_alert, false) - end) - - false_positives = - Enum.count(alerts, fn obs -> - Map.get(obs.metrics, :was_false_positive, false) - end) - - if length(alerts) > 0 do - false_positives / length(alerts) - else - 0.0 - end - end - - # ============================================================================ - # Drift Tolerance Learning - # ============================================================================ - - defp learn_drift_tolerance(state) do - recent_obs = Enum.take(state.observations, 100) - - if length(recent_obs) < 10 do - {:error, :insufficient_data} - else - repair_cost = calculate_average_repair_cost(recent_obs) - inconsistency_impact = calculate_inconsistency_impact(recent_obs) - - current_tolerance = Map.get(state.current_policy, :drift_tolerance_threshold, 0.05) - - new_tolerance = - cond do - # Repairs are expensive, inconsistency impact is low → tolerate more drift - repair_cost > 1000 and inconsistency_impact < 0.1 -> - increase = current_tolerance * state.learning_rate - min(current_tolerance + increase, 0.2) - - # Repairs are cheap, inconsistency impact is high → tighten tolerance - repair_cost < 100 and inconsistency_impact > 0.5 -> - decrease = current_tolerance * state.learning_rate - max(current_tolerance - decrease, 0.01) - - true -> - current_tolerance - end - - Logger.info( - "Drift tolerance adjusted: #{current_tolerance} → #{new_tolerance} (repair_cost: #{repair_cost}, impact: #{inconsistency_impact})" - ) - - new_policy = Map.put(state.current_policy, :drift_tolerance_threshold, new_tolerance) - {:ok, new_policy} - end - end - - defp calculate_average_repair_cost(observations) do - repairs = Enum.filter(observations, fn obs -> Map.has_key?(obs.metrics, :repair_cost_ms) end) - - if length(repairs) > 0 do - Enum.sum(Enum.map(repairs, fn obs -> obs.metrics.repair_cost_ms end)) / length(repairs) - else - 0.0 - end - end - - defp calculate_inconsistency_impact(observations) do - inconsistencies = - Enum.count(observations, fn obs -> - Map.get(obs.metrics, :user_affected_by_inconsistency, false) - end) - - if length(observations) > 0 do - inconsistencies / length(observations) - else - 0.0 - end - end - - # ============================================================================ - # Query Plan Selection Learning - # ============================================================================ - - defp learn_query_plan_selection(state) do - recent_obs = Enum.take(state.observations, 200) - - if length(recent_obs) < 20 do - {:error, :insufficient_data} - else - # Group observations by query pattern - patterns = group_by_query_pattern(recent_obs) - - # For each pattern, find best-performing plan - learned_plans = - patterns - |> Enum.map(fn {pattern, observations} -> - best_plan = find_best_plan(observations) - {pattern, best_plan} - end) - |> Map.new() - - current_plans = Map.get(state.current_policy, :preferred_plans, %{}) - - # Merge learned plans with existing (learned plans take precedence) - new_plans = Map.merge(current_plans, learned_plans) - - Logger.info("Query plan selection updated: learned #{map_size(learned_plans)} patterns") - - new_policy = Map.put(state.current_policy, :preferred_plans, new_plans) - {:ok, new_policy} - end - end - - defp group_by_query_pattern(observations) do - Enum.group_by(observations, fn obs -> - # Extract query pattern (e.g., "SELECT-graph-WHERE-LIMIT") - Map.get(obs.metrics, :query_pattern, :unknown) - end) - end - - defp find_best_plan(observations) do - # Find plan with lowest average execution time - observations - |> Enum.group_by(fn obs -> Map.get(obs.metrics, :plan_id, :unknown) end) - |> Enum.map(fn {plan_id, obs_for_plan} -> - avg_time = - Enum.sum(Enum.map(obs_for_plan, fn o -> Map.get(o.metrics, :execution_time_ms, 0) end)) / - length(obs_for_plan) - - {plan_id, avg_time} - end) - |> Enum.min_by(fn {_plan_id, avg_time} -> avg_time end, fn -> {:unknown, :infinity} end) - |> elem(0) - end - - # ============================================================================ - # Convergence Detection - # ============================================================================ - - defp should_learn?(state) do - cond do - state.converged -> - false - - length(state.observations) < 10 -> - false - - recently_adjusted?(state) -> - false - - true -> - # Check if policy has stabilized - if policy_stable?(state) do - Logger.info("Learning converged for #{state.domain}") - false - else - true - end - end - end - - defp recently_adjusted?(state) do - case state.last_adjustment do - nil -> - false - - last_time -> - seconds_since = DateTime.diff(DateTime.utc_now(), last_time, :second) - # Don't adjust more than once per hour - seconds_since < 3600 - end - end - - defp policy_stable?(state) do - # Check last 10 observations for stability - recent = Enum.take(state.observations, 10) - - if length(recent) < 10 do - false - else - # Calculate variance in key metric - variance = calculate_metric_variance(recent, state.domain) - # Stable if variance < 5% - variance < 0.05 - end - end - - defp calculate_metric_variance(observations, domain) do - metric_key = - case domain do - :cache_ttl -> :cache_hit_rate - :normalization_threshold -> :false_positive_rate - :drift_tolerance -> :inconsistency_impact - :query_plan -> :avg_execution_time - end - - values = - Enum.map(observations, fn obs -> - Map.get(obs.metrics, metric_key, 0.0) - end) - - if length(values) > 1 do - mean = Enum.sum(values) / length(values) - variance = Enum.sum(Enum.map(values, fn v -> (v - mean) ** 2 end)) / length(values) - :math.sqrt(variance) / mean - else - 1.0 - end - end - - # ============================================================================ - # Private Helpers - # ============================================================================ - - defp via_tuple(domain) do - {:via, Registry, {VeriSim.LearnerRegistry, domain}} - end - - defp schedule_learning do - # Learn every 10 minutes - Process.send_after(self(), :learn, 600_000) - end -end diff --git a/verisimdb/lib/verisim/circuit_breaker.ex b/verisimdb/lib/verisim/circuit_breaker.ex deleted file mode 100644 index 46df5897..00000000 --- a/verisimdb/lib/verisim/circuit_breaker.ex +++ /dev/null @@ -1,301 +0,0 @@ -# SPDX-License-Identifier: MPL-2.0 - -defmodule VeriSim.CircuitBreaker do - @moduledoc """ - Circuit breaker implementation to prevent cascading failures. - - States: - - :closed - Normal operation, requests pass through - - :open - Too many failures, requests fail fast - - :half_open - Testing if service recovered - - Transitions: - - :closed → :open when failure_count >= threshold - - :open → :half_open after timeout expires - - :half_open → :closed on successful request - - :half_open → :open on failed request - """ - - use GenServer - require Logger - - @type state :: :closed | :open | :half_open - - defmodule State do - @type t :: %__MODULE__{ - store_id: String.t(), - state: :closed | :open | :half_open, - failure_count: non_neg_integer(), - failure_threshold: pos_integer(), - timeout_ms: pos_integer(), - last_failure_time: DateTime.t() | nil, - success_count: non_neg_integer(), - total_requests: non_neg_integer() - } - - defstruct store_id: nil, - state: :closed, - failure_count: 0, - failure_threshold: 5, - timeout_ms: 60_000, - # 1 minute - last_failure_time: nil, - success_count: 0, - total_requests: 0 - end - - # ============================================================================ - # Client API - # ============================================================================ - - @doc """ - Start circuit breaker for a store. - """ - def start_link(store_id, opts \\ []) do - GenServer.start_link(__MODULE__, {store_id, opts}, name: via_tuple(store_id)) - end - - @doc """ - Call function with circuit breaker protection. - - Returns: - - {:ok, result} if function succeeds - - {:error, {:circuit_open, msg}} if circuit is open - - {:error, reason} if function fails - """ - def call_with_breaker(store_id, func, opts \\ []) do - ensure_started(store_id, opts) - - case get_state(store_id) do - :open -> - Logger.warn("Circuit breaker open for store #{store_id}, failing fast") - {:error, {:circuit_open, "Circuit breaker is open for store #{store_id}"}} - - :half_open -> - # Try one request to test recovery - execute_and_record(store_id, func) - - :closed -> - # Normal operation - execute_and_record(store_id, func) - end - end - - @doc """ - Get current circuit breaker state for a store. - """ - def get_state(store_id) do - GenServer.call(via_tuple(store_id), :get_state) - end - - @doc """ - Get circuit breaker statistics. - """ - def get_stats(store_id) do - GenServer.call(via_tuple(store_id), :get_stats) - end - - @doc """ - Manually reset circuit breaker (force close). - """ - def reset(store_id) do - GenServer.call(via_tuple(store_id), :reset) - end - - # ============================================================================ - # Server Callbacks - # ============================================================================ - - @impl true - def init({store_id, opts}) do - state = %State{ - store_id: store_id, - state: :closed, - failure_count: 0, - failure_threshold: Keyword.get(opts, :failure_threshold, 5), - timeout_ms: Keyword.get(opts, :timeout_ms, 60_000) - } - - # Schedule periodic health check - schedule_health_check() - - Logger.info("Circuit breaker started for store #{store_id}") - {:ok, state} - end - - @impl true - def handle_call(:get_state, _from, state) do - {:reply, state.state, state} - end - - @impl true - def handle_call(:get_stats, _from, state) do - stats = %{ - state: state.state, - failure_count: state.failure_count, - success_count: state.success_count, - total_requests: state.total_requests, - failure_rate: - if state.total_requests > 0 do - state.failure_count / state.total_requests * 100 - else - 0.0 - end, - last_failure_time: state.last_failure_time - } - - {:reply, stats, state} - end - - @impl true - def handle_call(:reset, _from, state) do - Logger.info("Circuit breaker reset for store #{state.store_id}") - - new_state = %{ - state - | state: :closed, - failure_count: 0, - success_count: 0, - total_requests: 0, - last_failure_time: nil - } - - {:reply, :ok, new_state} - end - - @impl true - def handle_call({:record_success, result}, _from, state) do - new_state = %{ - state - | success_count: state.success_count + 1, - total_requests: state.total_requests + 1 - } - - new_state = - case state.state do - :half_open -> - # Success in half_open → transition to closed - Logger.info("Circuit breaker closed for store #{state.store_id} (successful recovery)") - %{new_state | state: :closed, failure_count: 0} - - :closed -> - # Success in closed → reset failure count - %{new_state | failure_count: 0} - - :open -> - # Should not happen, but handle gracefully - new_state - end - - {:reply, {:ok, result}, new_state} - end - - @impl true - def handle_call({:record_failure, error}, _from, state) do - new_state = %{ - state - | failure_count: state.failure_count + 1, - total_requests: state.total_requests + 1, - last_failure_time: DateTime.utc_now() - } - - new_state = - case state.state do - :half_open -> - # Failure in half_open → back to open - Logger.warn("Circuit breaker reopened for store #{state.store_id} (recovery failed)") - %{new_state | state: :open} - - :closed when new_state.failure_count >= state.failure_threshold -> - # Too many failures → open circuit - Logger.error( - "Circuit breaker opened for store #{state.store_id} (#{new_state.failure_count} failures)" - ) - - %{new_state | state: :open} - - :closed -> - # Still below threshold - new_state - - :open -> - # Already open - new_state - end - - {:reply, {:error, error}, new_state} - end - - @impl true - def handle_info(:health_check, state) do - new_state = - case state.state do - :open -> - # Check if timeout expired - if time_since_last_failure(state) >= state.timeout_ms do - Logger.info( - "Circuit breaker transitioning to half_open for store #{state.store_id}" - ) - - %{state | state: :half_open} - else - state - end - - _ -> - state - end - - schedule_health_check() - {:noreply, new_state} - end - - # ============================================================================ - # Private Helpers - # ============================================================================ - - defp via_tuple(store_id) do - {:via, Registry, {VeriSim.CircuitBreakerRegistry, store_id}} - end - - defp ensure_started(store_id, opts) do - case Registry.lookup(VeriSim.CircuitBreakerRegistry, store_id) do - [] -> - # Start circuit breaker if not exists - case DynamicSupervisor.start_child( - VeriSim.CircuitBreakerSupervisor, - {__MODULE__, {store_id, opts}} - ) do - {:ok, _pid} -> :ok - {:error, {:already_started, _pid}} -> :ok - error -> error - end - - [_] -> - :ok - end - end - - defp execute_and_record(store_id, func) do - case func.() do - {:ok, result} -> - GenServer.call(via_tuple(store_id), {:record_success, result}) - - {:error, reason} = error -> - GenServer.call(via_tuple(store_id), {:record_failure, reason}) - error - end - end - - defp time_since_last_failure(state) do - case state.last_failure_time do - nil -> :infinity - time -> DateTime.diff(DateTime.utc_now(), time, :millisecond) - end - end - - defp schedule_health_check do - # Check every 10 seconds - Process.send_after(self(), :health_check, 10_000) - end -end diff --git a/verisimdb/lib/verisim/error_recovery.ex b/verisimdb/lib/verisim/error_recovery.ex deleted file mode 100644 index a8c2fd76..00000000 --- a/verisimdb/lib/verisim/error_recovery.ex +++ /dev/null @@ -1,414 +0,0 @@ -# SPDX-License-Identifier: MPL-2.0 - -defmodule VeriSim.ErrorRecovery do - @moduledoc """ - Error recovery strategies for VCL queries. - - Implements: - 1. Retry with exponential backoff - 2. Circuit breaker pattern - 3. Fallback to cached results - 4. Partial results handling - 5. Compensating transactions - - All recovery attempts are logged to verisim-temporal for audit trail. - """ - - require Logger - alias VeriSim.QueryCache - alias VeriSim.CircuitBreaker - alias VeriSim.Temporal - - @type recovery_strategy :: - :retry - | :cache_fallback - | :partial_results - | :fail_fast - | :compensate - - @type recovery_opts :: [ - max_retries: pos_integer(), - base_delay_ms: pos_integer(), - strategy: recovery_strategy(), - min_quorum: pos_integer() - ] - - # ============================================================================ - # Retry with Exponential Backoff - # ============================================================================ - - @doc """ - Execute function with automatic retry on recoverable errors. - - ## Options - * `:max_retries` - Maximum retry attempts (default: 3) - * `:base_delay_ms` - Base delay in milliseconds (default: 100) - * `:max_delay_ms` - Maximum delay cap (default: 10_000) - - ## Examples - - iex> ErrorRecovery.retry_with_backoff(fn -> - ...> execute_query(query) - ...> end, max_retries: 5) - {:ok, result} - """ - def retry_with_backoff(func, opts \\ []) do - max_retries = Keyword.get(opts, :max_retries, 3) - base_delay_ms = Keyword.get(opts, :base_delay_ms, 100) - max_delay_ms = Keyword.get(opts, :max_delay_ms, 10_000) - - do_retry(func, max_retries, base_delay_ms, max_delay_ms, 0, []) - end - - defp do_retry(func, max_retries, base_delay_ms, max_delay_ms, attempt, errors) - when attempt < max_retries do - case func.() do - {:ok, result} -> - if attempt > 0 do - Logger.info("Retry succeeded on attempt #{attempt + 1}") - end - - {:ok, result} - - {:error, error} when is_recoverable?(error) -> - delay_ms = calculate_backoff(attempt, base_delay_ms, max_delay_ms) - - Logger.debug( - "Retry attempt #{attempt + 1}/#{max_retries} after #{delay_ms}ms for error: #{inspect(error)}" - ) - - Process.sleep(delay_ms) - do_retry(func, max_retries, base_delay_ms, max_delay_ms, attempt + 1, [error | errors]) - - {:error, error} -> - Logger.error("Non-recoverable error, failing immediately: #{inspect(error)}") - {:error, error} - end - end - - defp do_retry(_func, _max_retries, _base_delay_ms, _max_delay_ms, attempt, errors) do - Logger.error("Max retries exceeded (#{attempt} attempts), errors: #{inspect(Enum.reverse(errors))}") - {:error, {:max_retries_exceeded, Enum.reverse(errors)}} - end - - defp calculate_backoff(attempt, base_delay_ms, max_delay_ms) do - # Exponential backoff: delay = base * 2^attempt - delay = base_delay_ms * :math.pow(2, attempt) |> round() - - # Add jitter (±25%) - jitter = delay * (0.75 + :rand.uniform() * 0.5) |> round() - - # Cap at max_delay_ms - min(jitter, max_delay_ms) - end - - @doc """ - Check if an error is recoverable (worth retrying). - """ - def is_recoverable?({:store_unavailable, _}), do: true - def is_recoverable?({:network_error, _}), do: true - def is_recoverable?({:timeout, _}), do: true - def is_recoverable?({:temporary_failure, _}), do: true - def is_recoverable?({:resource_exhausted, _}), do: true - def is_recoverable?(_), do: false - - # ============================================================================ - # Circuit Breaker Integration - # ============================================================================ - - @doc """ - Execute function with circuit breaker protection. - - If the store has failed too many times, the circuit breaker opens - and requests fail fast without attempting execution. - """ - def execute_with_circuit_breaker(store_id, func, opts \\ []) do - case CircuitBreaker.call_with_breaker(store_id, func, opts) do - {:ok, result} -> - {:ok, result} - - {:error, {:circuit_open, _msg}} = error -> - Logger.warn("Circuit breaker open for store #{store_id}, failing fast") - error - - {:error, _reason} = error -> - error - end - end - - # ============================================================================ - # Cache Fallback - # ============================================================================ - - @doc """ - Execute query with fallback to cached results if execution fails. - - Even stale cache results are returned if the store is unavailable. - """ - def execute_with_cache_fallback(query, execute_func) do - cache_key = QueryCache.query_result_key(query.ast) - - case execute_func.() do - {:ok, result} -> - # Success, cache for next time - QueryCache.put(cache_key, result, ttl: 300) - {:ok, result} - - {:error, {:store_unavailable, _}} = error -> - # Try cache even if expired - case QueryCache.get(cache_key) do - {:ok, cached} -> - Logger.warn( - "Using stale cached result due to store unavailability (cache age: #{cache_age(cached)}s)" - ) - - {:ok, %{cached | stale: true, warning: "Data may be outdated due to store unavailability"}} - - {:error, :not_found} -> - Logger.error("No cached result available for fallback") - error - end - - {:error, _} = error -> - error - end - end - - defp cache_age(cached_entry) do - DateTime.diff(DateTime.utc_now(), cached_entry.created_at, :second) - end - - # ============================================================================ - # Partial Results Handling - # ============================================================================ - - @doc """ - Execute federated query with partial results handling. - - Returns partial results if enough stores succeed (quorum-based). - - ## Options - * `:min_quorum` - Minimum successful stores required (default: 1) - * `:timeout_ms` - Per-store timeout (default: 30_000) - """ - def execute_with_partial_results(query, stores, execute_on_store_func, opts \\ []) do - min_quorum = Keyword.get(opts, :min_quorum, 1) - timeout_ms = Keyword.get(opts, :timeout_ms, 30_000) - - # Execute on all stores in parallel - results = - stores - |> Task.async_stream( - fn store -> - {store, execute_on_store_func.(store, query)} - end, - timeout: timeout_ms, - max_concurrency: length(stores) - ) - |> Enum.to_list() - - # Separate successes and failures - {succeeded, failed} = categorize_results(results) - - cond do - Enum.empty?(failed) -> - # All succeeded - Logger.info("Federation query succeeded on all #{length(succeeded)} stores") - {:ok, combine_results(succeeded)} - - Enum.empty?(succeeded) -> - # All failed - Logger.error("Federation query failed on all #{length(stores)} stores") - {:error, {:all_stores_failed, failed}} - - length(succeeded) >= min_quorum -> - # Enough succeeded for partial results - Logger.warn( - "Partial results from federation: succeeded=#{length(succeeded)}, failed=#{length(failed)}" - ) - - {:partial, combine_results(succeeded), - %{ - failed_stores: Enum.map(failed, fn {store, _} -> store end), - warning: "Not all stores responded" - }} - - true -> - # Not enough succeeded - Logger.error( - "Insufficient quorum: needed #{min_quorum}, got #{length(succeeded)}" - ) - - {:error, {:insufficient_quorum, length(succeeded), min_quorum}} - end - end - - defp categorize_results(results) do - Enum.reduce(results, {[], []}, fn - {:ok, {store, {:ok, result}}}, {succ, fail} -> - {[{store, result} | succ], fail} - - {:ok, {store, {:error, reason}}}, {succ, fail} -> - {succ, [{store, reason} | fail]} - - {:exit, reason}, {succ, fail} -> - {succ, [{:unknown_store, reason} | fail]} - end) - end - - defp combine_results(store_results) do - store_results - |> Enum.flat_map(fn {_store, result} -> result.data end) - end - - # ============================================================================ - # Compensating Transactions - # ============================================================================ - - @doc """ - Execute mutation with compensating transaction support. - - If any step fails, all previous steps are rolled back using - their compensation functions. - """ - def execute_with_compensation(mutation_id, steps) do - saga = %{ - id: mutation_id, - steps: [], - completed_steps: [] - } - - try do - saga = execute_saga_steps(saga, steps) - commit_saga(saga) - {:ok, saga.completed_steps} - rescue - error -> - Logger.error("Mutation #{mutation_id} failed, rolling back: #{inspect(error)}") - rollback_saga(saga) - {:error, {:mutation_failed, error}} - end - end - - defp execute_saga_steps(saga, []), do: saga - - defp execute_saga_steps(saga, [{name, forward_func, compensate_func} | rest]) do - Logger.debug("Executing saga step: #{name}") - - case forward_func.() do - {:ok, result} -> - updated_saga = %{ - saga - | completed_steps: [{name, result, compensate_func} | saga.completed_steps] - } - - execute_saga_steps(updated_saga, rest) - - {:error, reason} -> - raise "Saga step #{name} failed: #{inspect(reason)}" - end - end - - defp commit_saga(saga) do - Logger.info("Saga #{saga.id} committed successfully (#{length(saga.completed_steps)} steps)") - - # Log to temporal for audit trail - Temporal.append_audit_log("saga_commits", %{ - saga_id: saga.id, - steps: Enum.map(saga.completed_steps, fn {name, _, _} -> name end), - timestamp: DateTime.utc_now() - }) - - saga - end - - defp rollback_saga(saga) do - Logger.warn("Rolling back saga #{saga.id} (#{length(saga.completed_steps)} steps)") - - # Execute compensation functions in reverse order - saga.completed_steps - |> Enum.reverse() - |> Enum.each(fn {name, result, compensate_func} -> - Logger.debug("Compensating step: #{name}") - - case compensate_func.(result) do - :ok -> - Logger.debug("Compensated step #{name} successfully") - - {:error, reason} -> - Logger.error("Failed to compensate step #{name}: #{inspect(reason)}") - end - end) - - # Log to temporal for audit trail - Temporal.append_audit_log("saga_rollbacks", %{ - saga_id: saga.id, - steps: Enum.map(saga.completed_steps, fn {name, _, _} -> name end), - timestamp: DateTime.utc_now() - }) - - :ok - end - - # ============================================================================ - # Recovery Strategy Selection - # ============================================================================ - - @doc """ - Automatically select recovery strategy based on error type. - """ - def execute_with_auto_recovery(query, execute_func, opts \\ []) do - case execute_func.() do - {:ok, result} -> - {:ok, result} - - {:error, {:store_unavailable, _} = error} -> - Logger.info("Store unavailable, trying cache fallback") - execute_with_cache_fallback(query, execute_func) - - {:error, {:network_error, _} = error} -> - Logger.info("Network error, retrying with backoff") - retry_with_backoff(execute_func, opts) - - {:error, {:timeout, _} = error} -> - Logger.info("Timeout, trying cache fallback") - execute_with_cache_fallback(query, execute_func) - - {:error, {:permission_denied, _} = error} -> - Logger.error("Permission denied, failing immediately") - {:error, error} - - {:error, error} -> - if is_recoverable?(error) do - Logger.info("Recoverable error, retrying: #{inspect(error)}") - retry_with_backoff(execute_func, opts) - else - Logger.error("Non-recoverable error: #{inspect(error)}") - {:error, error} - end - end - end - - # ============================================================================ - # Error Logging - # ============================================================================ - - @doc """ - Log error to audit trail in verisim-temporal. - """ - def log_error(error, context) do - entry = %{ - timestamp: DateTime.utc_now(), - error: inspect(error), - error_recoverable: is_recoverable?(error), - query_id: context[:query_id], - user_id: context[:user_id], - octad_ids: context[:octad_ids], - recovery_attempted: context[:recovery_attempted], - recovery_successful: context[:recovery_successful], - retry_count: context[:retry_count] - } - - Temporal.append_audit_log("errors", entry) - end -end diff --git a/verisimdb/lib/verisim/query_cache.ex b/verisimdb/lib/verisim/query_cache.ex deleted file mode 100644 index 5f9534f8..00000000 --- a/verisimdb/lib/verisim/query_cache.ex +++ /dev/null @@ -1,577 +0,0 @@ -# SPDX-License-Identifier: MPL-2.0 - -defmodule VeriSim.QueryCache do - @moduledoc """ - Multi-layer query caching for VeriSimDB. - - Caches: - 1. Query results (expensive ZKP-verified queries) - 2. Parsed ASTs (avoid re-parsing identical queries) - 3. Execution plans (avoid re-planning) - 4. ZKP proofs (expensive to generate) - 5. Store metadata (indexes, capabilities) - 6. Registry lookups (UUID → store mappings) - 7. Temporal versions (frequently accessed historical states) - - Cache Layers: - - L1: In-memory ETS (hot data, <1ms access) - - L2: Distributed cache across nodes (warm data, <10ms) - - L3: Persistent cache in verisim-temporal (cold data, <100ms) - - Cache Policies: - - TTL-based expiration - - Drift-aware invalidation - - Per-modality policies (Vector can be stale, Semantic cannot) - - LRU eviction when memory limit reached - """ - - use GenServer - require Logger - - @type cache_key :: String.t() - @type cache_value :: any() - @type cache_layer :: :l1 | :l2 | :l3 - @type cache_policy :: :strict | :relaxed | :aggressive - - @type cache_config :: %{ - max_memory_mb: integer(), - default_ttl_seconds: integer(), - enable_l2: boolean(), - enable_l3: boolean(), - policy: cache_policy(), - modality_policies: %{String.t() => cache_policy()} - } - - # Default configuration - @default_config %{ - max_memory_mb: 1024, # 1GB L1 cache - default_ttl_seconds: 300, # 5 minutes - enable_l2: true, - enable_l3: true, - policy: :relaxed, - modality_policies: %{ - "VECTOR" => :aggressive, # Vector results rarely change - "GRAPH" => :relaxed, # Graph can be cached briefly - "DOCUMENT" => :aggressive, # Document rarely changes - "SEMANTIC" => :strict, # ZKP proofs must be fresh - "TENSOR" => :relaxed, - "TEMPORAL" => :aggressive # Historical data immutable - } - } - - # Cache entry structure - defmodule CacheEntry do - @type t :: %__MODULE__{ - key: String.t(), - value: any(), - created_at: DateTime.t(), - expires_at: DateTime.t(), - access_count: integer(), - last_accessed: DateTime.t(), - size_bytes: integer(), - layer: :l1 | :l2 | :l3, - tags: [String.t()] # For invalidation - } - - defstruct [ - :key, - :value, - :created_at, - :expires_at, - access_count: 0, - :last_accessed, - :size_bytes, - layer: :l1, - tags: [] - ] - end - - # === Public API === - - @doc """ - Get cached value, checking all layers (L1 → L2 → L3). - Returns {:ok, value} or {:error, :not_found}. - """ - def get(key, opts \\ []) do - start_time = System.monotonic_time(:microsecond) - - result = case get_from_l1(key) do - {:ok, entry} -> - record_hit(:l1, key, start_time) - {:ok, entry.value} - - {:error, :not_found} -> - # Try L2 - case get_config().enable_l2 && get_from_l2(key) do - {:ok, entry} -> - # Promote to L1 - put_in_l1(key, entry) - record_hit(:l2, key, start_time) - {:ok, entry.value} - - _ -> - # Try L3 - case get_config().enable_l3 && get_from_l3(key) do - {:ok, entry} -> - # Promote to L1 and L2 - put_in_l1(key, entry) - put_in_l2(key, entry) - record_hit(:l3, key, start_time) - {:ok, entry.value} - - _ -> - record_miss(key, start_time) - {:error, :not_found} - end - end - end - - result - end - - @doc """ - Put value in cache with optional TTL and tags. - """ - def put(key, value, opts \\ []) do - ttl = Keyword.get(opts, :ttl, get_config().default_ttl_seconds) - tags = Keyword.get(opts, :tags, []) - layer = Keyword.get(opts, :layer, :l1) - - entry = %CacheEntry{ - key: key, - value: value, - created_at: DateTime.utc_now(), - expires_at: DateTime.add(DateTime.utc_now(), ttl, :second), - last_accessed: DateTime.utc_now(), - size_bytes: estimate_size(value), - layer: layer, - tags: tags - } - - case layer do - :l1 -> put_in_l1(key, entry) - :l2 -> put_in_l2(key, entry) - :l3 -> put_in_l3(key, entry) - :all -> - put_in_l1(key, entry) - put_in_l2(key, entry) - put_in_l3(key, entry) - end - - :ok - end - - @doc """ - Invalidate cache entry by key. - """ - def invalidate(key) do - GenServer.call(__MODULE__, {:invalidate, key}) - end - - @doc """ - Invalidate all cache entries with specific tag. - Used for drift-aware invalidation. - - Examples: - invalidate_by_tag("octad:abc-123") # Invalidate all queries for octad - invalidate_by_tag("modality:GRAPH") # Invalidate all graph queries - invalidate_by_tag("federation:/universities/*") # Invalidate federation - """ - def invalidate_by_tag(tag) do - GenServer.call(__MODULE__, {:invalidate_by_tag, tag}) - end - - @doc """ - Clear entire cache (all layers). - """ - def clear_all do - GenServer.call(__MODULE__, :clear_all) - end - - @doc """ - Get cache statistics. - """ - def stats do - GenServer.call(__MODULE__, :stats) - end - - # === Cache Key Generation === - - @doc """ - Generate cache key for query result. - Includes query hash + modalities + source + conditions. - """ - def query_result_key(query_ast) do - query_hash = :crypto.hash(:blake3, :erlang.term_to_binary(query_ast)) - |> Base.encode16(case: :lower) - - "query:result:#{query_hash}" - end - - @doc """ - Generate cache key for parsed AST. - """ - def parsed_ast_key(raw_query) do - query_hash = :crypto.hash(:blake3, raw_query) |> Base.encode16(case: :lower) - "query:ast:#{query_hash}" - end - - @doc """ - Generate cache key for execution plan. - """ - def execution_plan_key(query_ast, optimization_mode) do - query_hash = :crypto.hash(:blake3, :erlang.term_to_binary(query_ast)) - |> Base.encode16(case: :lower) - - "query:plan:#{optimization_mode}:#{query_hash}" - end - - @doc """ - Generate cache key for ZKP proof. - """ - def zkp_proof_key(contract_name, data_hash) do - "zkp:proof:#{contract_name}:#{data_hash}" - end - - @doc """ - Generate cache key for registry lookup. - """ - def registry_key(octad_id) do - "registry:#{octad_id}" - end - - @doc """ - Generate cache key for temporal version. - """ - def temporal_version_key(octad_id, timestamp) do - ts_str = DateTime.to_iso8601(timestamp) - "temporal:#{octad_id}:#{ts_str}" - end - - # === Cache Policy Enforcement === - - @doc """ - Check if value should be cached based on modality policy. - """ - def should_cache?(modality, query_type) do - policy = get_policy_for_modality(modality) - - case {policy, query_type} do - {:strict, :dependent_type} -> true # Always cache verified queries - {:strict, :slipstream} -> false # Never cache unverified - {:relaxed, _} -> true # Cache both - {:aggressive, _} -> true # Cache everything - end - end - - @doc """ - Get TTL for modality based on policy. - """ - def get_ttl_for_modality(modality) do - policy = get_policy_for_modality(modality) - - case policy do - :strict -> 60 # 1 minute (short-lived) - :relaxed -> 300 # 5 minutes (default) - :aggressive -> 3600 # 1 hour (long-lived) - end - end - - defp get_policy_for_modality(modality) do - config = get_config() - Map.get(config.modality_policies, modality, config.policy) - end - - # === GenServer Implementation === - - def start_link(opts) do - GenServer.start_link(__MODULE__, opts, name: __MODULE__) - end - - @impl true - def init(_opts) do - # Create ETS tables for L1 cache - :ets.new(:cache_l1, [:set, :public, :named_table, read_concurrency: true]) - :ets.new(:cache_l2, [:set, :public, :named_table, read_concurrency: true]) - :ets.new(:cache_stats, [:set, :public, :named_table]) - :ets.new(:cache_tags, [:bag, :public, :named_table]) # key → tags mapping - - # Schedule periodic cleanup - schedule_cleanup() - - state = %{ - config: @default_config, - hits: %{l1: 0, l2: 0, l3: 0}, - misses: 0, - evictions: 0, - current_memory_bytes: 0 - } - - {:ok, state} - end - - @impl true - def handle_call({:invalidate, key}, _from, state) do - invalidate_key(key) - {:reply, :ok, state} - end - - @impl true - def handle_call({:invalidate_by_tag, tag}, _from, state) do - # Find all keys with this tag - keys = :ets.match(:cache_tags, {:"$1", tag}) - |> Enum.map(fn [key] -> key end) - - # Invalidate each - Enum.each(keys, &invalidate_key/1) - - Logger.info("Invalidated #{length(keys)} cache entries with tag: #{tag}") - - {:reply, {:ok, length(keys)}, state} - end - - @impl true - def handle_call(:clear_all, _from, state) do - :ets.delete_all_objects(:cache_l1) - :ets.delete_all_objects(:cache_tags) - - if state.config.enable_l2, do: clear_l2() - if state.config.enable_l3, do: clear_l3() - - new_state = %{state | current_memory_bytes: 0} - {:reply, :ok, new_state} - end - - @impl true - def handle_call(:stats, _from, state) do - total_requests = state.hits.l1 + state.hits.l2 + state.hits.l3 + state.misses - hit_rate = if total_requests > 0 do - (state.hits.l1 + state.hits.l2 + state.hits.l3) / total_requests * 100 - else - 0.0 - end - - l1_count = :ets.info(:cache_l1, :size) - - stats = %{ - hit_rate: Float.round(hit_rate, 2), - hits: state.hits, - misses: state.misses, - evictions: state.evictions, - l1_entries: l1_count, - memory_mb: Float.round(state.current_memory_bytes / 1_000_000, 2), - memory_limit_mb: state.config.max_memory_mb - } - - {:reply, stats, state} - end - - @impl true - def handle_info(:cleanup, state) do - # Remove expired entries - now = DateTime.utc_now() - - expired_keys = :ets.match(:cache_l1, {:"$1", %{expires_at: :"$2"}}) - |> Enum.filter(fn [_key, expires_at] -> - DateTime.compare(expires_at, now) == :lt - end) - |> Enum.map(fn [key, _] -> key end) - - Enum.each(expired_keys, &invalidate_key/1) - - # Check memory limit and evict if needed - new_state = if state.current_memory_bytes > state.config.max_memory_mb * 1_000_000 do - evict_lru_entries(state) - else - state - end - - # Schedule next cleanup - schedule_cleanup() - - {:noreply, new_state} - end - - # === Private Helpers === - - defp get_from_l1(key) do - case :ets.lookup(:cache_l1, key) do - [{^key, entry}] -> - # Check if expired - if DateTime.compare(entry.expires_at, DateTime.utc_now()) == :gt do - # Update access count and time - updated_entry = %{entry | - access_count: entry.access_count + 1, - last_accessed: DateTime.utc_now() - } - :ets.insert(:cache_l1, {key, updated_entry}) - {:ok, updated_entry} - else - :ets.delete(:cache_l1, key) - {:error, :expired} - end - - [] -> - {:error, :not_found} - end - end - - defp put_in_l1(key, entry) do - :ets.insert(:cache_l1, {key, entry}) - - # Store tag mappings - Enum.each(entry.tags, fn tag -> - :ets.insert(:cache_tags, {key, tag}) - end) - - :ok - end - - defp get_from_l2(key) do - case :ets.lookup(:cache_l2, key) do - [{^key, entry}] -> - if DateTime.compare(entry.expires_at, DateTime.utc_now()) == :gt do - updated_entry = %{entry | - access_count: entry.access_count + 1, - last_accessed: DateTime.utc_now() - } - :ets.insert(:cache_l2, {key, updated_entry}) - {:ok, updated_entry} - else - :ets.delete(:cache_l2, key) - {:error, :expired} - end - [] -> - {:error, :not_found} - end - end - - defp put_in_l2(key, entry) do - # L2 entries get 3x the TTL of L1 - extended_entry = %{entry | - layer: :l2, - expires_at: DateTime.add(entry.expires_at, entry.expires_at |> DateTime.diff(entry.created_at, :second) |> Kernel.*(2), :second) - } - :ets.insert(:cache_l2, {key, extended_entry}) - :ok - end - - defp get_from_l3(key) do - path = l3_cache_path(key) - case File.read(path) do - {:ok, content} -> - case :erlang.binary_to_term(content) do - %CacheEntry{} = entry -> - if DateTime.compare(entry.expires_at, DateTime.utc_now()) == :gt do - {:ok, entry} - else - File.rm(path) - {:error, :expired} - end - _ -> - {:error, :not_found} - end - {:error, _} -> - {:error, :not_found} - end - end - - defp put_in_l3(key, entry) do - path = l3_cache_path(key) - File.mkdir_p!(Path.dirname(path)) - l3_entry = %{entry | layer: :l3} - File.write!(path, :erlang.term_to_binary(l3_entry)) - :ok - end - - defp clear_l2 do - :ets.delete_all_objects(:cache_l2) - :ok - end - - defp clear_l3 do - l3_dir = l3_cache_dir() - if File.exists?(l3_dir) do - File.rm_rf!(l3_dir) - File.mkdir_p!(l3_dir) - end - :ok - end - - defp invalidate_key(key) do - :ets.delete(:cache_l1, key) - :ets.delete(:cache_l2, key) - :ets.match_delete(:cache_tags, {key, :_}) - - # Invalidate L3 - path = l3_cache_path(key) - File.rm(path) - - :ok - end - - defp evict_lru_entries(state) do - # Get all entries sorted by last_accessed - entries = :ets.tab2list(:cache_l1) - |> Enum.sort_by(fn {_key, entry} -> entry.last_accessed end) - - # Evict oldest 10% - evict_count = div(length(entries), 10) - to_evict = Enum.take(entries, evict_count) - - Enum.each(to_evict, fn {key, _entry} -> - invalidate_key(key) - end) - - Logger.info("Evicted #{evict_count} LRU cache entries") - - %{state | evictions: state.evictions + evict_count} - end - - defp record_hit(layer, _key, start_time) do - duration = System.monotonic_time(:microsecond) - start_time - Logger.debug("Cache hit (#{layer}): #{duration}μs") - GenServer.cast(__MODULE__, {:record_hit, layer}) - end - - defp record_miss(_key, start_time) do - duration = System.monotonic_time(:microsecond) - start_time - Logger.debug("Cache miss: #{duration}μs") - GenServer.cast(__MODULE__, :record_miss) - end - - @impl true - def handle_cast({:record_hit, layer}, state) do - new_hits = Map.update!(state.hits, layer, &(&1 + 1)) - {:noreply, %{state | hits: new_hits}} - end - - @impl true - def handle_cast(:record_miss, state) do - {:noreply, %{state | misses: state.misses + 1}} - end - - defp estimate_size(value) do - # Rough estimate of term size in bytes - :erlang.external_size(value) - end - - defp schedule_cleanup do - # Run cleanup every 60 seconds - Process.send_after(self(), :cleanup, 60_000) - end - - defp l3_cache_dir do - Path.join(System.tmp_dir!(), "verisimdb_cache_l3") - end - - defp l3_cache_path(key) do - safe_key = key |> :erlang.phash2() |> Integer.to_string() - Path.join(l3_cache_dir(), "#{safe_key}.cache") - end - - defp get_config do - # TODO: Make this configurable - @default_config - end -end diff --git a/verisimdb/lib/verisim/query_planner_bidirectional.ex b/verisimdb/lib/verisim/query_planner_bidirectional.ex deleted file mode 100644 index 21257ac9..00000000 --- a/verisimdb/lib/verisim/query_planner_bidirectional.ex +++ /dev/null @@ -1,331 +0,0 @@ -# SPDX-License-Identifier: MPL-2.0 - -defmodule VeriSim.QueryPlanner.Bidirectional do - @moduledoc """ - Bidirectional query optimization using both forward and backward propagation. - - Forward Propagation (Top-Down): - - Push predicates down from SELECT to WHERE - - Example: If LIMIT 10, tell stores "only need 10 results" - - Backward Propagation (Bottom-Up): - - Pull constraints up from stores to query plan - - Example: If store has index on 'year', reorganize to use it - - Combined: - - Forward pass: Optimize predicates - - Backward pass: Rewrite based on store capabilities - - Forward pass: Execute optimized plan - """ - - alias VeriSim.QueryPlannerConfig - - @doc """ - Optimize query using bidirectional propagation. - - Steps: - 1. Forward pass: Push down predicates (LIMIT, WHERE conditions) - 2. Backward pass: Pull up store capabilities (indexes, partitions) - 3. Generate final plan incorporating both - """ - def optimize_bidirectional(query_ast) do - # Phase 1: Forward propagation - forward_plan = forward_propagate(query_ast) - - # Phase 2: Backward propagation (gather store capabilities) - store_hints = backward_propagate(forward_plan) - - # Phase 3: Merge and finalize - finalize_plan(forward_plan, store_hints) - end - - # === Forward Propagation (Top-Down) === - - defp forward_propagate(query_ast) do - query_ast - |> push_limit_down() - |> push_predicates_down() - |> push_projections_down() - |> eliminate_redundant_operations() - end - - defp push_limit_down(query_ast) do - # If query has LIMIT, tell each store to limit results early - case query_ast.limit do - nil -> query_ast - limit_value -> - # Push LIMIT to each modality operation - updated_operations = Enum.map(query_ast.operations, fn op -> - %{op | early_limit: limit_value * 2} # 2x buffer for joins - end) - %{query_ast | operations: updated_operations} - end - end - - defp push_predicates_down(query_ast) do - # Push WHERE conditions to the stores that can evaluate them - case query_ast.where do - nil -> query_ast - conditions -> - # Decompose conditions by modality - condition_map = decompose_conditions_by_modality(conditions) - - # Assign to operations - updated_operations = Enum.map(query_ast.operations, fn op -> - modality_conditions = Map.get(condition_map, op.modality, []) - %{op | pushed_predicates: modality_conditions} - end) - - %{query_ast | operations: updated_operations} - end - end - - defp push_projections_down(query_ast) do - # If SELECT only requests specific fields, tell stores to return subset - # Example: SELECT GRAPH(nodes, edges) → only fetch nodes and edges - case query_ast.projections do - nil -> query_ast - projections -> - updated_operations = Enum.map(query_ast.operations, fn op -> - fields = Map.get(projections, op.modality, :all) - %{op | projection: fields} - end) - %{query_ast | operations: updated_operations} - end - end - - defp eliminate_redundant_operations(query_ast) do - # Remove operations that can't possibly return results - # Example: If SEMANTIC condition fails, no need to query other modalities - query_ast - end - - # === Backward Propagation (Bottom-Up) === - - defp backward_propagate(forward_plan) do - # Query each store for capabilities and hints - forward_plan.operations - |> Enum.map(&gather_store_hints/1) - |> aggregate_hints() - end - - defp gather_store_hints(operation) do - store_id = operation.store_id - modality = operation.modality - - # Ask store: "What indexes/optimizations do you have?" - hints = case modality do - "GRAPH" -> gather_graph_hints(store_id, operation) - "VECTOR" -> gather_vector_hints(store_id, operation) - "DOCUMENT" -> gather_document_hints(store_id, operation) - "SEMANTIC" -> gather_semantic_hints(store_id, operation) - "TENSOR" -> gather_tensor_hints(store_id, operation) - "TEMPORAL" -> gather_temporal_hints(store_id, operation) - end - - %{operation: operation, hints: hints} - end - - defp gather_graph_hints(store_id, operation) do - # Ask Oxigraph: "What indexes exist for this query?" - # Example response: [:edge_type_index, :node_label_index] - case OxigraphClient.get_available_indexes(store_id, operation.condition) do - {:ok, indexes} -> - %{ - available_indexes: indexes, - supports_parallel: true, - estimated_cardinality: OxigraphClient.estimate_result_count(store_id, operation.condition), - suggestion: suggest_graph_optimization(indexes, operation) - } - {:error, _} -> - %{available_indexes: [], supports_parallel: false} - end - end - - defp gather_vector_hints(store_id, operation) do - # Ask Milvus: "What's the best way to run this similarity search?" - case MilvusClient.get_index_info(store_id) do - {:ok, index_info} -> - %{ - index_type: index_info.type, # HNSW, IVF, etc. - dimension: index_info.dimension, - metric: index_info.metric, - supports_filtering: index_info.supports_filtering, - suggestion: suggest_vector_optimization(index_info, operation) - } - {:error, _} -> - %{index_type: :unknown} - end - end - - defp gather_document_hints(store_id, operation) do - # Ask Tantivy: "Do you have inverted index for this query?" - case TantivyClient.get_schema_info(store_id) do - {:ok, schema} -> - %{ - indexed_fields: schema.indexed_fields, - supports_phrase_query: schema.supports_phrase_query, - has_stored_fields: schema.has_stored_fields, - suggestion: suggest_document_optimization(schema, operation) - } - {:error, _} -> - %{indexed_fields: []} - end - end - - defp gather_semantic_hints(_store_id, operation) do - # Semantic operations are ZKP-based, limited optimization - %{ - contract_name: operation.condition.contract_name, - verification_cost: ProvenLibrary.estimate_verification_cost(operation.condition), - suggestion: :execute_late # ZKP verification is expensive, do last - } - end - - defp gather_tensor_hints(_store_id, _operation) do - # Tensor operations depend on shape/dtype - %{supports_gpu: BurnClient.has_gpu?()} - end - - defp gather_temporal_hints(store_id, operation) do - # Ask verisim-temporal: "Is this version cached?" - case VeriSimTemporal.check_cache(store_id, operation.condition) do - {:ok, :cached} -> %{suggestion: :execute_early, cached: true} - {:ok, :not_cached} -> %{suggestion: :normal, cached: false} - {:error, _} -> %{cached: false} - end - end - - # === Suggestion Heuristics === - - defp suggest_graph_optimization(indexes, operation) do - cond do - :edge_type_index in indexes and operation.condition.edge_type != nil -> - {:use_index, :edge_type_index} - :node_label_index in indexes -> - {:use_index, :node_label_index} - true -> - :full_scan - end - end - - defp suggest_vector_optimization(index_info, operation) do - cond do - index_info.type == :hnsw and operation.condition.threshold > 0.9 -> - {:use_hnsw, :high_precision} - index_info.supports_filtering and operation.condition.filter != nil -> - {:use_filtering, :pre_filter} - true -> - :standard_ann - end - end - - defp suggest_document_optimization(schema, operation) do - query_fields = extract_fulltext_fields(operation.condition) - - if Enum.all?(query_fields, &(&1 in schema.indexed_fields)) do - {:use_inverted_index, query_fields} - else - :full_scan - end - end - - # === Finalization (Merge Forward + Backward) === - - defp finalize_plan(forward_plan, store_hints) do - # Reorder operations based on backward hints - operations_with_hints = Enum.zip(forward_plan.operations, store_hints) - - # Sort by: - # 1. Cached operations first - # 2. Indexed operations second - # 3. Full scans last - sorted_operations = Enum.sort_by(operations_with_hints, fn {_op, hints} -> - priority = case hints.hints.suggestion do - :execute_early -> 1 - {:use_index, _} -> 2 - {:use_hnsw, _} -> 2 - :execute_late -> 99 - _ -> 50 - end - priority - end) - - final_operations = Enum.map(sorted_operations, fn {op, hints} -> - %{op | optimization_hint: hints.hints.suggestion} - end) - - %{forward_plan | operations: final_operations, optimization: :bidirectional} - end - - # === Helper Functions === - - defp decompose_conditions_by_modality(conditions) when is_list(conditions) do - Enum.reduce(conditions, %{}, fn condition, acc -> - modality = classify_condition_modality(condition) - Map.update(acc, modality, [condition], &[condition | &1]) - end) - end - - defp decompose_conditions_by_modality(conditions) when is_map(conditions) do - # Single condition wrapped in a map - modality = classify_condition_modality(conditions) - %{modality => [conditions]} - end - - defp decompose_conditions_by_modality(_), do: %{} - - defp classify_condition_modality(condition) when is_map(condition) do - cond do - Map.has_key?(condition, :embedding) or Map.has_key?(condition, :similar_to) -> - "VECTOR" - Map.has_key?(condition, :edge_type) or Map.has_key?(condition, :traversal) -> - "GRAPH" - Map.has_key?(condition, :fulltext) or Map.has_key?(condition, :contains) -> - "DOCUMENT" - Map.has_key?(condition, :proof) or Map.has_key?(condition, :contract) -> - "SEMANTIC" - Map.has_key?(condition, :shape) or Map.has_key?(condition, :tensor_op) -> - "TENSOR" - Map.has_key?(condition, :version) or Map.has_key?(condition, :as_of) -> - "TEMPORAL" - true -> - "GRAPH" # Default modality - end - end - - defp classify_condition_modality(condition) when is_binary(condition) do - upper = String.upcase(condition) - cond do - String.contains?(upper, "SIMILAR") or String.contains?(upper, "EMBEDDING") -> "VECTOR" - String.contains?(upper, "CITES") or String.contains?(upper, "EDGE") or String.contains?(upper, ")-[") -> "GRAPH" - String.contains?(upper, "FULLTEXT") or String.contains?(upper, "CONTAINS") -> "DOCUMENT" - String.contains?(upper, "PROOF") or String.contains?(upper, "VERIFY") -> "SEMANTIC" - String.contains?(upper, "TENSOR") or String.contains?(upper, "SHAPE") -> "TENSOR" - String.contains?(upper, "VERSION") or String.contains?(upper, "AS OF") -> "TEMPORAL" - true -> "GRAPH" - end - end - - defp classify_condition_modality(_), do: "GRAPH" - - defp aggregate_hints(hints_list) do - hints_list - end - - defp extract_fulltext_fields(condition) when is_map(condition) do - case Map.get(condition, :fulltext) do - nil -> - case Map.get(condition, :fields) do - nil -> ["title", "body"] # Default searchable fields - fields when is_list(fields) -> fields - field when is_binary(field) -> [field] - _ -> ["title", "body"] - end - %{fields: fields} when is_list(fields) -> fields - _ -> ["title", "body"] - end - end - - defp extract_fulltext_fields(_condition), do: ["title", "body"] -end diff --git a/verisimdb/lib/verisim/query_planner_config.ex b/verisimdb/lib/verisim/query_planner_config.ex deleted file mode 100644 index 1300ccd6..00000000 --- a/verisimdb/lib/verisim/query_planner_config.ex +++ /dev/null @@ -1,274 +0,0 @@ -# SPDX-License-Identifier: MPL-2.0 - -defmodule VeriSim.QueryPlannerConfig do - @moduledoc """ - Configuration for query planner tuning. - - Supports three optimization modes: - - :conservative - Worst-case estimates, prioritize correctness - - :balanced - Historical averages, good for most workloads - - :aggressive - Optimistic estimates, prioritize speed - - Can be set globally or per-modality. - """ - - use GenServer - - @type optimization_mode :: :conservative | :balanced | :aggressive - - @type config :: %{ - global_mode: optimization_mode(), - modality_overrides: %{String.t() => optimization_mode()}, - statistics_weight: float(), # How much to trust historical data (0.0-1.0) - enable_adaptive: boolean(), # Auto-tune based on query patterns - } - - # Default configuration - @default_config %{ - global_mode: :balanced, - modality_overrides: %{ - # Vector searches are predictable → aggressive - "VECTOR" => :aggressive, - # Graph traversal is unpredictable → conservative - "GRAPH" => :conservative, - # Semantic ZKP verification is expensive → conservative - "SEMANTIC" => :conservative, - }, - statistics_weight: 0.7, # 70% historical, 30% estimates - enable_adaptive: true, - } - - # === API === - - @doc """ - Get optimization mode for a specific modality. - Falls back to global mode if no override. - """ - def get_mode_for_modality(modality) do - config = get_config() - Map.get(config.modality_overrides, modality, config.global_mode) - end - - @doc """ - Set global optimization mode. - """ - def set_global_mode(mode) when mode in [:conservative, :balanced, :aggressive] do - GenServer.call(__MODULE__, {:set_global_mode, mode}) - end - - @doc """ - Set optimization mode for specific modality. - """ - def set_modality_mode(modality, mode) when mode in [:conservative, :balanced, :aggressive] do - GenServer.call(__MODULE__, {:set_modality_mode, modality, mode}) - end - - @doc """ - Enable/disable adaptive tuning. - When enabled, system auto-adjusts based on query performance. - """ - def set_adaptive(enabled) when is_boolean(enabled) do - GenServer.call(__MODULE__, {:set_adaptive, enabled}) - end - - # === Selectivity Multipliers === - - @doc """ - Get selectivity multiplier based on optimization mode. - - Conservative: Multiply by 2.0 (assume more results) - Balanced: Multiply by 1.0 (use estimate as-is) - Aggressive: Multiply by 0.5 (assume fewer results) - """ - def selectivity_multiplier(mode) do - case mode do - :conservative -> 2.0 - :balanced -> 1.0 - :aggressive -> 0.5 - end - end - - @doc """ - Get cost multiplier based on optimization mode. - - Conservative: Add safety buffer (1.5x) - Balanced: Use estimate as-is (1.0x) - Aggressive: Assume best case (0.8x) - """ - def cost_multiplier(mode) do - case mode do - :conservative -> 1.5 - :balanced -> 1.0 - :aggressive -> 0.8 - end - end - - # === Adaptive Tuning === - - @doc """ - Record actual query performance for adaptive tuning. - If estimates were consistently wrong, adjust mode automatically. - """ - def record_performance(modality, estimated_cost, actual_cost, estimated_selectivity, actual_selectivity) do - GenServer.cast(__MODULE__, {:record_performance, %{ - modality: modality, - estimated_cost: estimated_cost, - actual_cost: actual_cost, - estimated_selectivity: estimated_selectivity, - actual_selectivity: actual_selectivity, - timestamp: DateTime.utc_now(), - }}) - end - - # === GenServer Implementation === - - def start_link(_opts) do - GenServer.start_link(__MODULE__, @default_config, name: __MODULE__) - end - - @impl true - def init(config) do - # Load from persistent storage if available - config = load_config_from_storage() || config - {:ok, %{config: config, performance_history: []}} - end - - @impl true - def handle_call({:set_global_mode, mode}, _from, state) do - new_config = %{state.config | global_mode: mode} - persist_config(new_config) - {:reply, :ok, %{state | config: new_config}} - end - - @impl true - def handle_call({:set_modality_mode, modality, mode}, _from, state) do - new_overrides = Map.put(state.config.modality_overrides, modality, mode) - new_config = %{state.config | modality_overrides: new_overrides} - persist_config(new_config) - {:reply, :ok, %{state | config: new_config}} - end - - @impl true - def handle_call({:set_adaptive, enabled}, _from, state) do - new_config = %{state.config | enable_adaptive: enabled} - persist_config(new_config) - {:reply, :ok, %{state | config: new_config}} - end - - @impl true - def handle_call(:get_config, _from, state) do - {:reply, state.config, state} - end - - @impl true - def handle_cast({:record_performance, perf}, state) do - new_history = [perf | state.performance_history] |> Enum.take(1000) # Keep last 1000 - - # Adaptive tuning: adjust modes if estimates are consistently wrong - new_state = if state.config.enable_adaptive do - maybe_auto_tune(%{state | performance_history: new_history}) - else - %{state | performance_history: new_history} - end - - {:noreply, new_state} - end - - # === Private Helpers === - - defp get_config do - GenServer.call(__MODULE__, :get_config) - end - - defp maybe_auto_tune(state) do - # Analyze last 50 queries for each modality - recent = Enum.take(state.performance_history, 50) - - state.config.modality_overrides - |> Enum.reduce(state, fn {modality, current_mode}, acc -> - modality_perfs = Enum.filter(recent, &(&1.modality == modality)) - - if length(modality_perfs) >= 10 do - # Calculate average error - avg_cost_error = calculate_average_error(modality_perfs, :cost) - avg_selectivity_error = calculate_average_error(modality_perfs, :selectivity) - - # Adjust mode based on error patterns - new_mode = case {avg_cost_error, avg_selectivity_error} do - # Consistently underestimating → more conservative - {error, _} when error < -0.3 -> shift_mode(current_mode, :more_conservative) - # Consistently overestimating → more aggressive - {error, _} when error > 0.3 -> shift_mode(current_mode, :more_aggressive) - # Good estimates → keep current - _ -> current_mode - end - - if new_mode != current_mode do - Logger.info("Adaptive tuning: #{modality} mode changed from #{current_mode} to #{new_mode}") - update_modality_mode(acc, modality, new_mode) - else - acc - end - else - acc - end - end) - end - - defp calculate_average_error(performances, :cost) do - performances - |> Enum.map(fn p -> (p.estimated_cost - p.actual_cost) / p.actual_cost end) - |> Enum.sum() - |> Kernel./(length(performances)) - end - - defp calculate_average_error(performances, :selectivity) do - performances - |> Enum.map(fn p -> (p.estimated_selectivity - p.actual_selectivity) / p.actual_selectivity end) - |> Enum.sum() - |> Kernel./(length(performances)) - end - - defp shift_mode(:conservative, :more_conservative), do: :conservative - defp shift_mode(:conservative, :more_aggressive), do: :balanced - defp shift_mode(:balanced, :more_conservative), do: :conservative - defp shift_mode(:balanced, :more_aggressive), do: :aggressive - defp shift_mode(:aggressive, :more_conservative), do: :balanced - defp shift_mode(:aggressive, :more_aggressive), do: :aggressive - - defp update_modality_mode(state, modality, new_mode) do - new_overrides = Map.put(state.config.modality_overrides, modality, new_mode) - new_config = %{state.config | modality_overrides: new_overrides} - persist_config(new_config) - %{state | config: new_config} - end - - defp load_config_from_storage do - path = config_storage_path() - case File.read(path) do - {:ok, content} -> - try do - :erlang.binary_to_term(content) - rescue - _ -> nil - end - {:error, _} -> nil - end - end - - defp persist_config(config) do - path = config_storage_path() - File.mkdir_p!(Path.dirname(path)) - File.write!(path, :erlang.term_to_binary(config)) - :ok - rescue - e -> - require Logger - Logger.warning("Failed to persist config: #{inspect(e)}") - :ok - end - - defp config_storage_path do - Path.join([System.tmp_dir!(), "verisimdb", "query_planner_config.bin"]) - end -end diff --git a/verisimdb/lib/verisim/query_router_cached.ex b/verisimdb/lib/verisim/query_router_cached.ex deleted file mode 100644 index 7c3dd03a..00000000 --- a/verisimdb/lib/verisim/query_router_cached.ex +++ /dev/null @@ -1,334 +0,0 @@ -# SPDX-License-Identifier: MPL-2.0 - -defmodule VeriSim.QueryRouter.Cached do - @moduledoc """ - Query router with integrated caching. - - Wraps VeriSim.QueryRouter to add caching at multiple levels: - 1. Parsed AST cache (avoid re-parsing) - 2. Execution plan cache (avoid re-planning) - 3. Query result cache (avoid re-execution) - 4. ZKP proof cache (avoid re-generation) - - Drift-aware: Automatically invalidates cache when drift detected. - """ - - alias VeriSim.QueryCache - alias VeriSim.QueryRouter - alias VeriSim.DriftMonitor - require Logger - - @doc """ - Execute query with caching. - - Cache strategy: - - Dependent-type queries: Cache AST, plan, results, and proofs - - Slipstream queries: Cache AST and plan only (results too volatile) - - Invalidate on drift detection - """ - def execute_with_cache(raw_query, opts \\ []) do - use_dependent_types = Keyword.get(opts, :use_dependent_types, false) - force_fresh = Keyword.get(opts, :force_fresh, false) - - # Step 1: Try to get parsed AST from cache - {ast, ast_cache_hit?} = if force_fresh do - {parse_query(raw_query), false} - else - case get_cached_ast(raw_query) do - {:ok, cached_ast} -> {cached_ast, true} - {:error, :not_found} -> - ast = parse_query(raw_query) - cache_ast(raw_query, ast) - {ast, false} - end - end - - Logger.debug("AST cache #{if ast_cache_hit?, do: "HIT", else: "MISS"}") - - # Step 2: Try to get execution plan from cache - {plan, plan_cache_hit?} = if force_fresh do - {generate_plan(ast), false} - else - case get_cached_plan(ast) do - {:ok, cached_plan} -> {cached_plan, true} - {:error, :not_found} -> - plan = generate_plan(ast) - cache_plan(ast, plan) - {plan, false} - end - end - - Logger.debug("Plan cache #{if plan_cache_hit?, do: "HIT", else: "MISS"}") - - # Step 3: Check if we should cache query results - should_cache_result? = should_cache_query_result?(ast, use_dependent_types) - - # Step 4: Try to get query results from cache - if should_cache_result? and not force_fresh do - case get_cached_result(ast) do - {:ok, cached_result} -> - Logger.info("Query result cache HIT") - {:ok, %{cached_result | cache_hit: true}} - - {:error, :not_found} -> - # Execute and cache - execute_and_cache(ast, plan, use_dependent_types) - end - else - # Don't use result cache - execute_query(ast, plan, use_dependent_types) - end - end - - # === Cache Retrieval === - - defp get_cached_ast(raw_query) do - key = QueryCache.parsed_ast_key(raw_query) - QueryCache.get(key) - end - - defp get_cached_plan(ast) do - optimization_mode = VeriSim.QueryPlannerConfig.get_mode_for_modality( - List.first(ast.modalities) - ) - key = QueryCache.execution_plan_key(ast, optimization_mode) - QueryCache.get(key) - end - - defp get_cached_result(ast) do - key = QueryCache.query_result_key(ast) - QueryCache.get(key) - end - - # === Cache Storage === - - defp cache_ast(raw_query, ast) do - key = QueryCache.parsed_ast_key(raw_query) - tags = ["ast"] - - QueryCache.put(key, ast, - ttl: 3600, # ASTs don't change, cache for 1 hour - tags: tags, - layer: :l1 - ) - end - - defp cache_plan(ast, plan) do - optimization_mode = VeriSim.QueryPlannerConfig.get_mode_for_modality( - List.first(ast.modalities) - ) - key = QueryCache.execution_plan_key(ast, optimization_mode) - tags = ["plan"] ++ extract_tags_from_ast(ast) - - QueryCache.put(key, plan, - ttl: 600, # Plans can change with statistics, cache for 10 minutes - tags: tags, - layer: :l1 - ) - end - - defp cache_result(ast, result) do - key = QueryCache.query_result_key(ast) - tags = extract_tags_from_ast(ast) - - # TTL based on modality policy - ttl = get_result_ttl(ast) - - QueryCache.put(key, result, - ttl: ttl, - tags: tags, - layer: :all # Store in all layers (L1, L2, L3) - ) - end - - defp cache_zkp_proof(contract_name, data_hash, proof) do - key = QueryCache.zkp_proof_key(contract_name, data_hash) - tags = ["zkp", "contract:#{contract_name}"] - - QueryCache.put(key, proof, - ttl: 1800, # ZKP proofs valid for 30 minutes - tags: tags, - layer: :all - ) - end - - # === Query Execution === - - defp execute_and_cache(ast, plan, use_dependent_types) do - case execute_query(ast, plan, use_dependent_types) do - {:ok, result} -> - # Cache the result - cache_result(ast, result) - - # If ZKP proof included, cache it separately - if result.proof do - data_hash = compute_data_hash(result.data) - cache_zkp_proof(result.proof.contract_name, data_hash, result.proof) - end - - {:ok, %{result | cache_hit: false}} - - error -> - error - end - end - - defp execute_query(ast, plan, use_dependent_types) do - # Delegate to actual QueryRouter - QueryRouter.handle_query(%{ - ast: ast, - plan: plan, - typed: use_dependent_types - }) - end - - # === Cache Policy Decisions === - - defp should_cache_query_result?(ast, use_dependent_types) do - # Check each modality's cache policy - ast.modalities - |> Enum.all?(fn modality -> - modality_str = modality_to_string(modality) - query_type = if use_dependent_types, do: :dependent_type, else: :slipstream - - QueryCache.should_cache?(modality_str, query_type) - end) - end - - defp get_result_ttl(ast) do - # Use the most conservative TTL among all modalities - ast.modalities - |> Enum.map(fn modality -> - modality_str = modality_to_string(modality) - QueryCache.get_ttl_for_modality(modality_str) - end) - |> Enum.min() - end - - # === Drift-Aware Invalidation === - - @doc """ - Invalidate cache when drift is detected for a octad. - Should be called by DriftMonitor when drift detected. - """ - def invalidate_on_drift(octad_id) do - Logger.info("Invalidating cache due to drift for octad: #{octad_id}") - - # Invalidate all queries involving this octad - QueryCache.invalidate_by_tag("octad:#{octad_id}") - - :ok - end - - @doc """ - Invalidate cache when drift is detected across a federation. - """ - def invalidate_federation_cache(federation_pattern) do - Logger.info("Invalidating cache for federation: #{federation_pattern}") - - QueryCache.invalidate_by_tag("federation:#{federation_pattern}") - - :ok - end - - @doc """ - Invalidate cache for specific modality. - Used when modality store is updated. - """ - def invalidate_modality_cache(modality) do - Logger.info("Invalidating cache for modality: #{modality}") - - QueryCache.invalidate_by_tag("modality:#{modality}") - - :ok - end - - # === Cache Warming === - - @doc """ - Pre-warm cache with common queries. - Should be called at startup or after major data changes. - """ - def warm_cache(common_queries) do - Logger.info("Warming cache with #{length(common_queries)} queries") - - common_queries - |> Task.async_stream( - fn query -> - execute_with_cache(query, use_dependent_types: false) - end, - max_concurrency: 10, - timeout: 30_000 - ) - |> Stream.run() - - Logger.info("Cache warming complete") - end - - # === Helper Functions === - - defp parse_query(raw_query) do - modalities = extract_modalities(raw_query) - source = extract_source(raw_query) - %{raw: raw_query, modalities: modalities, source: source} - end - - defp extract_modalities(query) do - upper = String.upcase(query) - modalities = [] - modalities = if String.contains?(upper, "GRAPH"), do: ["GRAPH" | modalities], else: modalities - modalities = if String.contains?(upper, "VECTOR"), do: ["VECTOR" | modalities], else: modalities - modalities = if String.contains?(upper, "TENSOR"), do: ["TENSOR" | modalities], else: modalities - modalities = if String.contains?(upper, "SEMANTIC"), do: ["SEMANTIC" | modalities], else: modalities - modalities = if String.contains?(upper, "DOCUMENT"), do: ["DOCUMENT" | modalities], else: modalities - modalities = if String.contains?(upper, "TEMPORAL"), do: ["TEMPORAL" | modalities], else: modalities - if modalities == [], do: ["GRAPH"], else: Enum.reverse(modalities) - end - - defp extract_source(query) do - cond do - String.contains?(query, "FEDERATION") -> - case Regex.run(~r/FEDERATION\s+(\S+)/, query) do - [_, pattern] -> {:federation, pattern, %{}} - _ -> {:octad, "unknown"} - end - true -> - {:octad, "unknown"} - end - end - - defp generate_plan(ast) do - # Call query planner - VeriSim.QueryPlanner.plan_query(ast) - end - - defp extract_tags_from_ast(ast) do - tags = [] - - # Add octad tags - octad_tags = case ast.source do - {:octad, id} -> ["octad:#{id}"] - {:federation, pattern, _} -> ["federation:#{pattern}"] - {:store, store_id} -> ["store:#{store_id}"] - _ -> [] - end - - # Add modality tags - modality_tags = Enum.map(ast.modalities, fn mod -> - "modality:#{modality_to_string(mod)}" - end) - - tags ++ octad_tags ++ modality_tags - end - - defp modality_to_string(modality) when is_binary(modality), do: modality - defp modality_to_string(modality) when is_atom(modality) do - modality |> Atom.to_string() |> String.upcase() - end - defp modality_to_string(modality), do: "#{modality}" - - defp compute_data_hash(data) do - :crypto.hash(:blake3, :erlang.term_to_binary(data)) - |> Base.encode16(case: :lower) - end -end diff --git a/verisimdb/llm-warmup-dev.md b/verisimdb/llm-warmup-dev.md deleted file mode 100644 index 90b48f34..00000000 --- a/verisimdb/llm-warmup-dev.md +++ /dev/null @@ -1,153 +0,0 @@ -# VeriSimDB — LLM Context (Developer) - -## Identity - -VeriSimDB (Veridical Simulacrum Database) — 8-modality entity consistency -engine with drift detection, self-normalisation, and formally verified queries. -Part of the nextgen-databases monorepo. License: PMPL-1.0-or-later. -Author: Jonathan D.A. Jewell. - -## Architecture (Marr's Three Levels) - -**Computational**: Maintain cross-modal consistency across 8 representations. -**Algorithmic**: Octad entities, drift detection with thresholds, OTP supervision. -**Implementational**: Rust stores + Elixir coordination + VCL queries. - -``` -Elixir OTP: EntityServer, DriftMonitor, QueryRouter, SchemaRegistry - ↓ HTTP -Rust Core: graph, vector, tensor, semantic, document, temporal, - provenance, spatial, octad, drift, normalizer, api -``` - -## Rust Workspace (`Cargo.toml`) - -10 library crates + 1 binary crate (verisim-api). Workspace at root. -Build: `OPENSSL_NO_VENDOR=1 cargo build --release` -oxrocksdb-sys eliminated (redb pure-Rust backend). protoc eliminated -(pre-generated). No C++ deps. - -## Elixir Layer (`elixir-orchestration/`) - -OTP app with supervision tree. GenServer per entity. -Hypatia integration: ScanIngester, PatternQuery, DispatchBridge (37 tests). -Built-in VCL parser (no external runtime needed). -Product telemetry: opt-in ETS collector + JSON reporter. - -## Key Subsystems - -### Drift Detection -6 drift types: semantic_vector, graph_document, temporal_consistency, -tensor, schema, quality. Configurable thresholds gate normalisation. - -### Self-Normalisation -1. Identify most authoritative modality -2. Regenerate drifted modalities -3. Validate consistency -4. Atomic update - -### Federation -10 adapters: MongoDB, Redis, Neo4j, ClickHouse, SurrealDB, SQLite, -DuckDB, VectorDB, InfluxDB, ObjectStorage. -Integration tests: 105 tests across 7 adapters (need test-infra stack). - -### VCL (VeriSim Consonance Language) -Type system: VCL-UT. 11 proof types. Multi-proof parsing. -Modality compatibility validation. ReScript playground wired to backend. - -### Hypatia Pipeline -``` -panic-attack assail → ScanIngester → octads → PatternQuery → DispatchBridge → gitbot-fleet -``` -954 patterns tracked across 298 repos. Fleet dispatch logged (JSONL), -live execution needs GitHub PAT. - -## Client SDKs (`connectors/clients/`) - -6 SDKs: Rust, V, Elixir, ReScript, Julia, Gleam. -Shared: JSON Schema, OpenAPI, protobuf (`connectors/shared/`). - -## ABI/FFI (`src/abi/`, `ffi/zig/`) - -Idris2 ABI definitions (formal proofs). Zig FFI C-ABI bridge. -Generated C headers in `generated/abi/`. - -## Container Stack - -Podman (never Docker). Chainguard base images. -- `container/Containerfile` — main build -- `container/compose.toml` — selur-compose (3 services) -- `container/.gatekeeper.yaml` — svalinn edge gateway -- `container/manifest.toml` — cerro-torre signing -- `container/ct-build.sh` — build/sign/verify pipeline - -Test infra: `connectors/test-infra/compose.toml` — 7 databases. - -## Instance Policy (CRITICAL) - -This repo = source code + examples ONLY. Each consumer runs own instance: -- IDApTIK: port 8090, volume idaptik-verisimdb-data -- Burble: port 8091, volume burble-verisimdb-data -- Hypatia: port 8092, volume hypatia-verisimdb-data - -Never store app data here. Never point at localhost:8080. - -## Commands - -```bash -just build / build-all / build-elixir / build-abi / build-ffi -just test / test-elixir / test-integration / test-all -just serve / serve-otp -just fmt / lint / fmt-elixir -just container-build / container-run / deploy / deploy-stop -just panic-scan / hypatia-scan / license-check / check-scm -just doctor / heal / tour / help-me -``` - -## Language Policy - -Allowed: Rust, Elixir, ReScript, VCL, Idris2, Zig. -Banned: Python, Go, Node.js. - -## Code Patterns - -### Octad creation (Rust) -```rust -let octad = OctadBuilder::new() - .with_document("Title", "Body") - .with_embedding(vec![0.1, 0.2]) - .with_types(vec!["http://example.org/Doc"]) - .build(); -``` - -### Entity server (Elixir) -```elixir -{:ok, _} = VeriSim.EntityServer.start_link("eid") -{:ok, state} = VeriSim.EntityServer.get("eid") -``` - -## Known Issues - -All 25 historical issues resolved. See KNOWN-ISSUES.adoc. - -## File Map - -| Path | What | -|------|------| -| `Cargo.toml` | Workspace root | -| `rust-core/verisim-*/` | 10 modality crates | -| `elixir-orchestration/lib/verisim/` | OTP modules | -| `elixir-orchestration/lib/verisim/hypatia/` | Hypatia integration | -| `connectors/clients/` | 6 SDKs | -| `connectors/test-infra/` | Integration test databases | -| `container/` | Containerfiles + compose | -| `playground/` | VCL playground (ReScript) | -| `src/abi/` | Idris2 ABI | -| `ffi/zig/` | Zig FFI | -| `spec/` | Grammar EBNF | -| `verification/` | Verification gateway | -| `docs/` | Documentation | -| `verisimdb-data/` | Git-backed flat-file data | -| `.machine_readable/` | STATE.scm, META.scm, ECOSYSTEM.scm | -| `.claude/CLAUDE.md` | Full AI instructions | -| `0-AI-MANIFEST.a2ml` | Universal AI entry point | diff --git a/verisimdb/llm-warmup-user.md b/verisimdb/llm-warmup-user.md deleted file mode 100644 index 9b10451a..00000000 --- a/verisimdb/llm-warmup-user.md +++ /dev/null @@ -1,69 +0,0 @@ -# VeriSimDB — LLM Context (User) - -## What It Is - -VeriSimDB is an 8-modality database engine. Every entity is stored across -Graph, Vector, Tensor, Semantic, Document, Temporal, Provenance, and Spatial -representations simultaneously (the "octad"). Drift between modalities is -detected and self-healed. - -## Architecture - -- **Rust core** (`rust-core/`): 10 crates for modality stores + API server -- **Elixir/OTP** (`elixir-orchestration/`): GenServer per entity, drift - monitoring, query routing, schema registry -- **VCL**: Custom query language (not SQL) -- **Federation**: 10 adapters (MongoDB, Redis, Neo4j, ClickHouse, SurrealDB, - SQLite, DuckDB, VectorDB, InfluxDB, ObjectStorage) -- **ABI/FFI**: Idris2 formal spec + Zig C-ABI bridge - -## Quick Commands - -```bash -just build # Build Rust (release) -just build-elixir # Build Elixir layer -just build-all # Both -just serve # API server on :8080 -just serve-otp # Full OTP orchestrator -just test-all # All tests -just doctor # Check prerequisites -``` - -## Core Concepts - -- **Octad**: One entity, 8 synchronized representations -- **Drift**: Divergence between modalities (semantic-vector, graph-document, etc.) -- **Self-normalisation**: When drift exceeds threshold, most authoritative - modality regenerates the others atomically -- **Federation**: Coordinates existing databases (MongoDB, Redis, etc.) as - modality backends - -## Prerequisites - -Rust (nightly), Elixir 1.17+, Erlang/OTP 27+, Zig 0.14+, -openssl-devel, pkg-config, just. Optional: Idris2, Podman. - -## Container - -```bash -just container-build # Podman build -just container-run # Run on :8080 -``` - -## Instance Policy - -This repo is source code + examples only. Each consuming project -(IDApTIK, Burble, Hypatia) runs its own VeriSimDB instance on a -dedicated port with its own data volume. - -## Key Paths - -| Path | Purpose | -|------|---------| -| `rust-core/` | Rust modality crates | -| `elixir-orchestration/` | OTP supervision tree | -| `connectors/clients/` | 6 SDKs (Rust, V, Elixir, ReScript, Julia, Gleam) | -| `connectors/test-infra/` | 7-database integration test stack | -| `container/` | Containerfile + compose | -| `playground/` | VCL playground (ReScript) | -| `.claude/CLAUDE.md` | Full AI context | diff --git a/verisimdb/opsm.toml b/verisimdb/opsm.toml deleted file mode 100644 index 1a894291..00000000 --- a/verisimdb/opsm.toml +++ /dev/null @@ -1,15 +0,0 @@ -# SPDX-License-Identifier: MPL-2.0 -[opsm] -name = "verisimdb" -role = "dogfood-wave-1" -policy = "strict" - -[paths] -use_pathroot = true - -[install] -mode = "auto" -allow_untrusted = false - -[telemetry] -enabled = false diff --git a/verisimdb/osv-scanner.toml b/verisimdb/osv-scanner.toml deleted file mode 100644 index 4a91084d..00000000 --- a/verisimdb/osv-scanner.toml +++ /dev/null @@ -1,4 +0,0 @@ -[[IgnoredVulns]] -id = "RUSTSEC-2025-0134" -ignoreUntil = 2026-12-31 -reason = "Transitive via axum-server TLS helper (rustls-pemfile 2.2.0); tracked for migration to rustls-pki-types APIs upstream." diff --git a/verisimdb/playground/build.mjs b/verisimdb/playground/build.mjs deleted file mode 100644 index 5771ab09..00000000 --- a/verisimdb/playground/build.mjs +++ /dev/null @@ -1,15 +0,0 @@ -// SPDX-License-Identifier: MPL-2.0 -// Build script — bundles ReScript output for production PWA deployment. -import * as esbuild from "esbuild"; - -await esbuild.build({ - entryPoints: ["src/App.res.mjs"], - bundle: true, - minify: true, - outfile: "public/app.js", - format: "esm", - target: "es2022", - sourcemap: true, -}); - -console.log("Build complete: public/app.js"); diff --git a/verisimdb/playground/deno.json b/verisimdb/playground/deno.json deleted file mode 100644 index e736ae51..00000000 --- a/verisimdb/playground/deno.json +++ /dev/null @@ -1,17 +0,0 @@ -{ - "name": "vcl-playground", - "version": "0.1.0", - "tasks": { - "res:build": "rescript", - "res:dev": "rescript -w", - "dev": "concurrently \"npm:res:dev\" \"npx serve public -l 3000\"", - "build": "rescript && deno run -A build.mjs", - "clean": "rescript clean" - }, - "imports": { - "@rescript/core": "npm:@rescript/core@^1.0.0", - "rescript": "^12.0.0", - "concurrently": "npm:concurrently@^9.0.0", - "esbuild": "npm:esbuild@^0.24.0" - } -} \ No newline at end of file diff --git a/verisimdb/playground/public/index.html b/verisimdb/playground/public/index.html deleted file mode 100644 index 6498c21b..00000000 --- a/verisimdb/playground/public/index.html +++ /dev/null @@ -1,99 +0,0 @@ - - - - - - - - - VCL Playground — VeriSimDB - - - - - - - -
-

VCL Playground

-
-
- -
-
-
- - VCL -
-
-
- -
-
-
- Editor - 0 chars -
-
- -
-
- - - - - - -
-
Ready
-
- -
-
- Output - table -
-
-Welcome to the VCL Playground! - -VCL (VeriSim Consonance Language) is the native query interface for VeriSimDB, -the octad multimodal database with drift detection and self-normalisation. - -Octad modalities: GRAPH, VECTOR, TENSOR, SEMANTIC, DOCUMENT, TEMPORAL, PROVENANCE, SPATIAL - -When verisim-api is running on localhost:8080, queries execute against the real backend. -Otherwise, demo mode simulates responses offline. - -Toggle VCL-UT mode for dependent-type features including: - - PROOF clause (11 types: EXISTENCE, INTEGRITY, CONSISTENCY, PROVENANCE, FRESHNESS, etc.) - - THRESHOLD constraints - - Multi-proof composition: PROOF A AND B AND C - - Formal verification of query results - -Try an example query or start typing in the editor. -
-
-
- -
- Mode: VCL - playground.verisimdb.org - Demo mode (offline) -
- - - - - diff --git a/verisimdb/playground/public/manifest.json b/verisimdb/playground/public/manifest.json deleted file mode 100644 index 4caea5df..00000000 --- a/verisimdb/playground/public/manifest.json +++ /dev/null @@ -1,21 +0,0 @@ -{ - "name": "VCL Playground", - "short_name": "VCL Play", - "description": "Interactive VCL query editor for VeriSimDB with VCL-UT mode", - "start_url": "/", - "display": "standalone", - "background_color": "#0f172a", - "theme_color": "#3b82f6", - "icons": [ - { - "src": "icon-192.png", - "sizes": "192x192", - "type": "image/png" - }, - { - "src": "icon-512.png", - "sizes": "512x512", - "type": "image/png" - } - ] -} diff --git a/verisimdb/playground/public/style.css b/verisimdb/playground/public/style.css deleted file mode 100644 index 769de76e..00000000 --- a/verisimdb/playground/public/style.css +++ /dev/null @@ -1,370 +0,0 @@ -/* SPDX-License-Identifier: MPL-2.0 */ -/* VCL Playground — Dark theme matching VeriSimDB branding */ - -:root { - --bg-primary: #0f172a; - --bg-secondary: #1e293b; - --bg-editor: #0d1117; - --bg-output: #1a1f2e; - --text-primary: #e2e8f0; - --text-secondary: #94a3b8; - --text-muted: #64748b; - --accent: #3b82f6; - --accent-hover: #2563eb; - --success: #22c55e; - --warning: #f59e0b; - --error: #ef4444; - --border: #334155; - --keyword: #93c5fd; - --modality: #86efac; - --string: #fde68a; - --number: #67e8f9; - --comment: #64748b; - --proof: #c4b5fd; - --font-mono: "JetBrains Mono", "Fira Code", "Cascadia Code", monospace; - --font-sans: "Inter", system-ui, -apple-system, sans-serif; -} - -* { - margin: 0; - padding: 0; - box-sizing: border-box; -} - -body { - font-family: var(--font-sans); - background: var(--bg-primary); - color: var(--text-primary); - min-height: 100vh; - display: flex; - flex-direction: column; -} - -/* Header */ -.header { - display: flex; - align-items: center; - justify-content: space-between; - padding: 0.75rem 1.5rem; - background: var(--bg-secondary); - border-bottom: 1px solid var(--border); - flex-shrink: 0; -} - -.header h1 { - font-size: 1.25rem; - font-weight: 600; - display: flex; - align-items: center; - gap: 0.5rem; -} - -.header h1 .logo { - color: var(--accent); -} - -.header-controls { - display: flex; - align-items: center; - gap: 1rem; -} - -/* VCL-UT Toggle */ -.mode-toggle { - display: flex; - align-items: center; - gap: 0.5rem; - font-size: 0.875rem; - color: var(--text-secondary); -} - -.mode-toggle label { - cursor: pointer; - user-select: none; -} - -.toggle-switch { - position: relative; - width: 44px; - height: 24px; - background: var(--border); - border-radius: 12px; - cursor: pointer; - transition: background 0.2s; -} - -.toggle-switch.active { - background: var(--accent); -} - -.toggle-switch .toggle-knob { - position: absolute; - top: 2px; - left: 2px; - width: 20px; - height: 20px; - background: white; - border-radius: 50%; - transition: transform 0.2s; -} - -.toggle-switch.active .toggle-knob { - transform: translateX(20px); -} - -.mode-badge { - font-size: 0.75rem; - padding: 0.2rem 0.5rem; - border-radius: 4px; - font-weight: 600; - letter-spacing: 0.05em; -} - -.mode-badge.vcl { - background: rgba(59, 130, 246, 0.2); - color: var(--accent); - border: 1px solid rgba(59, 130, 246, 0.3); -} - -.mode-badge.vcl-dt { - background: rgba(196, 181, 253, 0.2); - color: var(--proof); - border: 1px solid rgba(196, 181, 253, 0.3); -} - -/* Main layout */ -.main { - display: flex; - flex: 1; - overflow: hidden; -} - -/* Editor panel */ -.editor-panel { - flex: 1; - display: flex; - flex-direction: column; - border-right: 1px solid var(--border); - min-width: 0; -} - -.panel-header { - display: flex; - align-items: center; - justify-content: space-between; - padding: 0.5rem 1rem; - background: var(--bg-secondary); - border-bottom: 1px solid var(--border); - font-size: 0.875rem; - color: var(--text-secondary); -} - -.editor-area { - flex: 1; - position: relative; - overflow: hidden; -} - -#editor { - width: 100%; - height: 100%; - background: var(--bg-editor); - color: var(--text-primary); - font-family: var(--font-mono); - font-size: 14px; - line-height: 1.6; - padding: 1rem; - border: none; - outline: none; - resize: none; - tab-size: 2; - white-space: pre; - overflow: auto; -} - -#editor::placeholder { - color: var(--text-muted); -} - -/* Syntax highlight overlay */ -.highlight-overlay { - position: absolute; - top: 0; - left: 0; - width: 100%; - height: 100%; - padding: 1rem; - font-family: var(--font-mono); - font-size: 14px; - line-height: 1.6; - white-space: pre-wrap; - word-wrap: break-word; - pointer-events: none; - overflow: hidden; - color: transparent; -} - -.highlight-overlay .kw { color: var(--keyword); font-weight: 600; } -.highlight-overlay .mod { color: var(--modality); font-weight: 600; } -.highlight-overlay .str { color: var(--string); } -.highlight-overlay .num { color: var(--number); } -.highlight-overlay .proof { color: var(--proof); font-weight: 600; } -.highlight-overlay .cmt { color: var(--comment); font-style: italic; } - -/* Toolbar */ -.toolbar { - display: flex; - align-items: center; - gap: 0.5rem; - padding: 0.5rem 1rem; - background: var(--bg-secondary); - border-top: 1px solid var(--border); -} - -.btn { - padding: 0.4rem 1rem; - border: 1px solid var(--border); - border-radius: 6px; - background: var(--bg-primary); - color: var(--text-primary); - font-family: var(--font-sans); - font-size: 0.875rem; - cursor: pointer; - transition: all 0.15s; - display: flex; - align-items: center; - gap: 0.4rem; -} - -.btn:hover { - background: var(--border); -} - -.btn-primary { - background: var(--accent); - border-color: var(--accent); - color: white; - font-weight: 600; -} - -.btn-primary:hover { - background: var(--accent-hover); -} - -.btn kbd { - font-size: 0.75rem; - padding: 0.1rem 0.3rem; - background: rgba(0, 0, 0, 0.2); - border-radius: 3px; -} - -/* Output panel */ -.output-panel { - flex: 1; - display: flex; - flex-direction: column; - min-width: 0; -} - -#output { - flex: 1; - padding: 1rem; - background: var(--bg-output); - font-family: var(--font-mono); - font-size: 13px; - line-height: 1.6; - overflow: auto; - white-space: pre-wrap; - word-wrap: break-word; -} - -.output-info { color: var(--text-secondary); } -.output-success { color: var(--success); } -.output-error { color: var(--error); } -.output-warning { color: var(--warning); } -.output-timing { color: var(--text-muted); font-size: 0.8em; } - -/* Lint diagnostics */ -.lint-bar { - padding: 0.4rem 1rem; - background: var(--bg-secondary); - border-top: 1px solid var(--border); - font-size: 0.8rem; - color: var(--text-muted); - display: flex; - align-items: center; - gap: 0.5rem; -} - -.lint-error { color: var(--error); } -.lint-warning { color: var(--warning); } -.lint-hint { color: var(--accent); } - -/* Example queries sidebar */ -.examples-drawer { - padding: 1rem; - background: var(--bg-secondary); - border-top: 1px solid var(--border); - max-height: 200px; - overflow-y: auto; -} - -.example-query { - padding: 0.5rem 0.75rem; - margin-bottom: 0.25rem; - border-radius: 4px; - cursor: pointer; - font-family: var(--font-mono); - font-size: 0.8rem; - color: var(--text-secondary); - transition: background 0.15s; -} - -.example-query:hover { - background: var(--bg-primary); - color: var(--text-primary); -} - -.example-query .example-label { - font-family: var(--font-sans); - font-size: 0.75rem; - color: var(--text-muted); - margin-bottom: 0.2rem; -} - -/* Status bar */ -.status-bar { - display: flex; - align-items: center; - justify-content: space-between; - padding: 0.3rem 1rem; - background: var(--accent); - color: white; - font-size: 0.75rem; - flex-shrink: 0; -} - -.status-bar.vcl-dt { - background: #7c3aed; -} - -/* Responsive */ -@media (max-width: 768px) { - .main { - flex-direction: column; - } - - .editor-panel { - border-right: none; - border-bottom: 1px solid var(--border); - max-height: 50vh; - } - - .header-controls { - gap: 0.5rem; - } - - .header h1 { - font-size: 1rem; - } -} diff --git a/verisimdb/playground/public/sw.js b/verisimdb/playground/public/sw.js deleted file mode 100644 index f778a452..00000000 --- a/verisimdb/playground/public/sw.js +++ /dev/null @@ -1,31 +0,0 @@ -// SPDX-License-Identifier: MPL-2.0 -// Service Worker for VCL Playground PWA — offline support. - -const CACHE_NAME = "vcl-playground-v1"; -const ASSETS = ["/", "/index.html", "/app.js", "/style.css", "/manifest.json"]; - -self.addEventListener("install", (event) => { - event.waitUntil( - caches.open(CACHE_NAME).then((cache) => cache.addAll(ASSETS)) - ); -}); - -self.addEventListener("activate", (event) => { - event.waitUntil( - caches.keys().then((names) => - Promise.all( - names - .filter((name) => name !== CACHE_NAME) - .map((name) => caches.delete(name)) - ) - ) - ); -}); - -self.addEventListener("fetch", (event) => { - event.respondWith( - caches.match(event.request).then((cached) => { - return cached || fetch(event.request); - }) - ); -}); diff --git a/verisimdb/playground/rescript.json b/verisimdb/playground/rescript.json deleted file mode 100644 index f7f1d551..00000000 --- a/verisimdb/playground/rescript.json +++ /dev/null @@ -1,18 +0,0 @@ -{ - "name": "vcl-playground", - "sources": [ - { - "dir": "src", - "subdirs": true - } - ], - "package-specs": [ - { - "module": "esmodule", - "in-source": true - } - ], - "suffix": ".res.mjs", - "bs-dependencies": ["@rescript/core"], - "bsc-flags": ["-open RescriptCore"] -} diff --git a/verisimdb/playground/src/ApiClient.res b/verisimdb/playground/src/ApiClient.res deleted file mode 100644 index 263f5d77..00000000 --- a/verisimdb/playground/src/ApiClient.res +++ /dev/null @@ -1,278 +0,0 @@ -// SPDX-License-Identifier: MPL-2.0 -// VCL Playground API client — connects to a real verisim-api backend. -// Falls back gracefully to demo mode when the backend is unreachable. - -/// Response shape from POST /api/v1/vcl/execute. -type vclResponse = { - success: bool, - statement_type: string, - row_count: int, - data: JSON.t, - message: option, -} - -/// Connection state for the backend. -type connectionState = - | Disconnected - | Connecting - | Connected(string) - | Failed(string) - -// === Fetch API bindings (no external package needed) === - -type response - -@val external fetch: (string, {..}) => promise = "fetch" -@val external fetchGet: string => promise = "fetch" -@send external responseJson: response => promise = "json" -@get external responseOk: response => bool = "ok" -@get external responseStatus: response => int = "status" - -/// Backend URL — defaults to localhost:8080 (verisim-api default port). -/// Override by setting window.__VERISIM_API_URL__ before the script loads. -@val @scope("window") external apiUrlOverride: Nullable.t = "__VERISIM_API_URL__" - -let getBaseUrl = (): string => { - switch Nullable.toOption(apiUrlOverride) { - | Some(url) => url - | None => "http://localhost:8080" - } -} - -/// Check backend health by hitting GET /api/v1/health. -let checkHealth = async (): result => { - let url = getBaseUrl() ++ "/api/v1/health" - try { - let response = await fetchGet(url) - if responseOk(response) { - Ok(getBaseUrl()) - } else { - Error("Backend returned " ++ Int.toString(responseStatus(response))) - } - } catch { - | _ => Error("Backend unreachable at " ++ url) - } -} - -/// Execute a VCL query against the real backend. -/// Returns Ok(vclResponse) on success, Error(string) on failure. -let executeQuery = async (query: string): result => { - let url = getBaseUrl() ++ "/api/v1/vcl/execute" - let bodyDict = Dict.make() - Dict.set(bodyDict, "query", JSON.Encode.string(query)) - let body = JSON.Encode.object(bodyDict) - - try { - let response = await fetch( - url, - { - "method": "POST", - "headers": {"Content-Type": "application/json"}, - "body": JSON.stringify(body), - }, - ) - - let json = await responseJson(response) - - if responseOk(response) { - // Parse the response fields. - switch JSON.Classify.classify(json) { - | JSON.Classify.Object(obj) => { - let success = switch Dict.get(obj, "success") { - | Some(v) => - switch JSON.Classify.classify(v) { - | JSON.Classify.Bool(b) => b - | _ => false - } - | None => false - } - let statementType = switch Dict.get(obj, "statement_type") { - | Some(v) => - switch JSON.Classify.classify(v) { - | JSON.Classify.String(s) => s - | _ => "UNKNOWN" - } - | None => "UNKNOWN" - } - let rowCount = switch Dict.get(obj, "row_count") { - | Some(v) => - switch JSON.Classify.classify(v) { - | JSON.Classify.Number(n) => Float.toInt(n) - | _ => 0 - } - | None => 0 - } - let data = switch Dict.get(obj, "data") { - | Some(v) => v - | None => JSON.Encode.null - } - let message = switch Dict.get(obj, "message") { - | Some(v) => - switch JSON.Classify.classify(v) { - | JSON.Classify.String(s) => Some(s) - | _ => None - } - | None => None - } - - Ok({ - success, - statement_type: statementType, - row_count: rowCount, - data, - message, - }) - } - | _ => Error("Unexpected response format") - } - } else { - // Parse error message from response body. - switch JSON.Classify.classify(json) { - | JSON.Classify.Object(obj) => - switch Dict.get(obj, "error") { - | Some(v) => - switch JSON.Classify.classify(v) { - | JSON.Classify.String(s) => Error(s) - | _ => - Error("Backend error (status " ++ Int.toString(responseStatus(response)) ++ ")") - } - | None => - Error("Backend error (status " ++ Int.toString(responseStatus(response)) ++ ")") - } - | _ => Error("Backend error (status " ++ Int.toString(responseStatus(response)) ++ ")") - } - } - } catch { - | exn => - let msg = switch exn { - | Exn.Error(e) => - switch Exn.message(e) { - | Some(m) => m - | None => "Network error" - } - | _ => "Network error" - } - Error(msg) - } -} - -// === Response conversion helpers (defined before toExecuteResult) === - -/// Convert a JSON value to a display string for table cells. -let jsonToString = (value: JSON.t): string => { - switch JSON.Classify.classify(value) { - | JSON.Classify.String(s) => s - | JSON.Classify.Number(n) => - if Float.mod(n, 1.0) == 0.0 { - Int.toString(Float.toInt(n)) - } else { - Float.toFixed(n, ~digits=3) - } - | JSON.Classify.Bool(b) => if b { "true" } else { "false" } - | JSON.Classify.Null => "null" - | _ => JSON.stringify(value) - } -} - -/// Format an EXPLAIN response into readable text. -let formatExplainResponse = (data: JSON.t): string => { - let text = ref("=== EXPLAIN OUTPUT (from backend) ===\n\n") - switch JSON.Classify.classify(data) { - | JSON.Classify.Object(obj) => { - switch Dict.get(obj, "query") { - | Some(q) => - switch JSON.Classify.classify(q) { - | JSON.Classify.String(s) => text := text.contents ++ "Query: " ++ s ++ "\n\n" - | _ => () - } - | None => () - } - switch Dict.get(obj, "plan") { - | Some(plan) => text := text.contents ++ "Plan:\n" ++ JSON.stringify(plan, ~space=2) ++ "\n" - | None => () - } - } - | _ => text := text.contents ++ JSON.stringify(data, ~space=2) ++ "\n" - } - text.contents -} - -/// Convert a VCL response with a JSON data array into a table result. -let formatAsTable = (response: vclResponse): DemoExecutor.executeResult => { - switch JSON.Classify.classify(response.data) { - | JSON.Classify.Array(items) => - if Array.length(items) == 0 { - DemoExecutor.Success({ - columns: ["result"], - rows: [], - timing_ms: 0.0, - row_count: 0, - }) - } else { - // Extract columns from the first item's keys. - let firstItem = items[0] - let columns = switch firstItem { - | Some(item) => - switch JSON.Classify.classify(item) { - | JSON.Classify.Object(obj) => Dict.keysToArray(obj) - | _ => ["value"] - } - | None => ["value"] - } - - // Extract rows. - let rows = items->Array.map(item => - columns->Array.map(col => - switch JSON.Classify.classify(item) { - | JSON.Classify.Object(obj) => - switch Dict.get(obj, col) { - | Some(v) => jsonToString(v) - | None => "null" - } - | _ => jsonToString(item) - } - ) - ) - - DemoExecutor.Success({ - columns, - rows, - timing_ms: 0.0, - row_count: response.row_count, - }) - } - | JSON.Classify.Object(_) => - // Single object result (e.g., COUNT, SHOW STATUS) — render as key-value pairs. - let text = JSON.stringify(response.data, ~space=2) - switch response.message { - | Some(msg) => DemoExecutor.ExplainResult(msg ++ "\n\n" ++ text) - | None => - DemoExecutor.ExplainResult(response.statement_type ++ " result:\n\n" ++ text) - } - | JSON.Classify.Null => - switch response.message { - | Some(msg) => DemoExecutor.ExplainResult(msg) - | None => DemoExecutor.ExplainResult(response.statement_type ++ " completed successfully.") - } - | _ => DemoExecutor.ExplainResult(JSON.stringify(response.data, ~space=2)) - } -} - -/// Convert a VCL API response into a DemoExecutor-compatible result. -/// This bridges the real backend response format to the existing rendering code. -let toExecuteResult = (response: vclResponse): DemoExecutor.executeResult => { - if !response.success { - DemoExecutor.Error( - switch response.message { - | Some(msg) => msg - | None => "Query failed" - }, - ) - } else if response.statement_type == "EXPLAIN" { - // EXPLAIN returns structured JSON — format it as readable text. - DemoExecutor.ExplainResult(formatExplainResponse(response.data)) - } else { - // Convert JSON data array to columns + rows table format. - formatAsTable(response) - } -} diff --git a/verisimdb/playground/src/App.res b/verisimdb/playground/src/App.res deleted file mode 100644 index 70fdfd97..00000000 --- a/verisimdb/playground/src/App.res +++ /dev/null @@ -1,349 +0,0 @@ -// SPDX-License-Identifier: MPL-2.0 -// VCL Playground — main application entry point. -// Wires up the editor, VCL-UT toggle, linter, formatter, and query executor. -// Tries the real verisim-api backend first, falls back to demo mode. - -// === DOM helpers === - -@val external document: {..} = "document" - -let getElementById = (id: string): {..} => document["getElementById"](id) -let addEventListener = (el: {..}, event: string, handler: {..} => unit): unit => - el["addEventListener"](event, handler) - -// === State === - -let vclDtMode = ref(false) -let backendConnected = ref(false) -let queryInFlight = ref(false) - -// === Initialization === - -let rec init = () => { - let editor = getElementById("editor") - let output = getElementById("output") - let lintBar = getElementById("lint-bar") - let charCount = getElementById("char-count") - let modeBadge = getElementById("mode-badge") - let statusMode = getElementById("status-mode") - let statusBar = getElementById("status-bar") - let statusConnection = getElementById("status-connection") - let toggle = getElementById("vcl-dt-toggle") - - // === Check backend connectivity === - checkBackend(statusConnection)->ignore - - // === VCL-UT Toggle === - let updateMode = () => { - if vclDtMode.contents { - toggle["classList"]["add"]("active")->ignore - modeBadge["className"] = "mode-badge vcl-dt" - modeBadge["textContent"] = "VCL-UT" - statusMode["textContent"] = "Mode: VCL-UT (Dependent Types)" - statusBar["classList"]["add"]("vcl-dt")->ignore - } else { - toggle["classList"]["remove"]("active")->ignore - modeBadge["className"] = "mode-badge vcl" - modeBadge["textContent"] = "VCL" - statusMode["textContent"] = "Mode: VCL" - statusBar["classList"]["remove"]("vcl-dt")->ignore - } - } - - addEventListener(toggle, "click", _ => { - vclDtMode := !vclDtMode.contents - updateMode()->ignore - // Re-lint current query - let query = editor["value"] - if String.trim(query) !== "" { - runLint(query, lintBar) - } - }) - - // Keyboard accessibility for toggle - addEventListener(toggle, "keydown", e => { - let key: string = e["key"] - if key == " " || key == "Enter" { - e["preventDefault"]() - vclDtMode := !vclDtMode.contents - updateMode() - } - }) - - // === Editor events === - addEventListener(editor, "input", _ => { - let query: string = editor["value"] - let len = String.length(query) - charCount["textContent"] = `${Int.toString(len)} chars` - - // Live lint - if String.trim(query) !== "" { - runLint(query, lintBar) - } else { - lintBar["textContent"] = "Ready" - lintBar["className"] = "lint-bar" - } - }) - - // Ctrl+Enter to run - addEventListener(editor, "keydown", e => { - let key: string = e["key"] - let ctrlKey: bool = e["ctrlKey"] - let metaKey: bool = e["metaKey"] - if key == "Enter" && (ctrlKey || metaKey) { - e["preventDefault"]() - runQuery(editor, output) - } - // Tab inserts spaces - if key == "Tab" { - e["preventDefault"]() - // Insert 2 spaces at cursor - let start: int = editor["selectionStart"] - let endd: int = editor["selectionEnd"] - let value: string = editor["value"] - editor["value"] = - String.slice(value, ~start=0, ~end=start) ++ " " ++ String.sliceToEnd(value, ~start=endd) - editor["selectionStart"] = start + 2 - editor["selectionEnd"] = start + 2 - } - }) - - // === Button handlers === - addEventListener(getElementById("run-btn"), "click", _ => { - runQuery(editor, output) - }) - - addEventListener(getElementById("explain-btn"), "click", _ => { - let query: string = editor["value"] - if String.trim(query) !== "" { - let explainQuery = if String.includes(String.toUpperCase(query), "EXPLAIN") { - query - } else { - "EXPLAIN " ++ query - } - executeAndDisplay(explainQuery, output) - } - }) - - addEventListener(getElementById("lint-btn"), "click", _ => { - let query: string = editor["value"] - if String.trim(query) !== "" { - let diagnostics = Linter.lint(query, ~vclDt=vclDtMode.contents) - if Array.length(diagnostics) == 0 { - output["innerHTML"] = `No lint issues found.` - } else { - let html = - diagnostics - ->Array.map(d => { - let cls = switch d.severity { - | Linter.Error => "output-error" - | Linter.Warning => "output-warning" - | Linter.Hint => "output-info" - } - `[${d.code}] ${Linter.severityToString(d.severity)}: ${d.message}` - }) - ->Array.join("\n") - output["innerHTML"] = html - } - } - }) - - addEventListener(getElementById("format-btn"), "click", _ => { - let query: string = editor["value"] - if String.trim(query) !== "" { - editor["value"] = Formatter.formatVcl(query) - // Trigger input event to update char count - let inputEvent = document["createEvent"]("Event") - inputEvent["initEvent"]("input", true, true)->ignore - editor["dispatchEvent"](inputEvent)->ignore - } - }) - - addEventListener(getElementById("clear-btn"), "click", _ => { - editor["value"] = "" - output["innerHTML"] = `Output cleared.` - lintBar["textContent"] = "Ready" - charCount["textContent"] = "0 chars" - }) - - addEventListener(getElementById("examples-btn"), "click", _ => { - let exs = Examples.forMode(vclDtMode.contents) - let html = - exs - ->Array.map(ex => { - let escaped = String.replaceAll(String.replaceAll(ex.query, "<", "<"), ">", ">") - let dtBadge = if ex.vclDt { - ` DT` - } else { - "" - } - `
-
${ex.label}${dtBadge}
- ${escaped} -
` - }) - ->Array.join("") - - output["innerHTML"] = `
${html}
` - - // Add click handlers to examples - let exampleEls = output["querySelectorAll"](".example-query") - let len: int = exampleEls["length"] - let i = ref(0) - while i.contents < len { - let el = exampleEls[i.contents]->Option.getExn - addEventListener(el, "click", _ => { - let q: string = el["getAttribute"]("data-query") - editor["value"] = q - let inputEvent = document["createEvent"]("Event") - inputEvent["initEvent"]("input", true, true)->ignore - editor["dispatchEvent"](inputEvent)->ignore - }) - i := i.contents + 1 - } - }) -} - -// === Check backend health === - -and checkBackend = async (statusEl: {..}): unit => { - statusEl["textContent"] = "Connecting..." - let result = await ApiClient.checkHealth() - switch result { - | Ok(url) => - backendConnected := true - statusEl["textContent"] = "Connected to " ++ url - statusEl["style"]["color"] = "var(--accent, #4ade80)" - | Error(_) => - backendConnected := false - statusEl["textContent"] = "Demo mode (offline)" - statusEl["style"]["color"] = "" - } -} - -// === Query execution === - -and runQuery = (editor: {..}, output: {..}) => { - let query: string = editor["value"] - if String.trim(query) !== "" { - executeAndDisplay(query, output) - } -} - -and executeAndDisplay = (query: string, output: {..}) => { - if !queryInFlight.contents { - if backendConnected.contents { - // Execute against real backend (async). - queryInFlight := true - output["innerHTML"] = `Executing query...` - executeOnBackend(query, output)->ignore - } else { - // Fall back to demo executor (synchronous). - let result = DemoExecutor.execute(query, ~vclDt=vclDtMode.contents) - renderResult(result, output) - } - } -} - -and executeOnBackend = async (query: string, output: {..}): unit => { - let startTime = Date.now() - let response = await ApiClient.executeQuery(query) - let elapsed = Date.now() -. startTime - queryInFlight := false - - switch response { - | Ok(apiResponse) => { - let result = ApiClient.toExecuteResult(apiResponse) - // Inject real timing into success results. - let timedResult = switch result { - | DemoExecutor.Success(data) => - DemoExecutor.Success({...data, timing_ms: elapsed, row_count: apiResponse.row_count}) - | other => other - } - renderResult(timedResult, output) - } - | Error(msg) => - // Backend failed — try demo mode as fallback. - output["innerHTML"] = - `Backend error: ${msg}\n` ++ - `Falling back to demo mode...` - let _ = setTimeout(() => { - let result = DemoExecutor.execute(query, ~vclDt=vclDtMode.contents) - renderResult(result, output) - }, 300) - } -} - -and renderResult = (result: DemoExecutor.executeResult, output: {..}) => { - switch result { - | DemoExecutor.Success(data) => { - // Render as table - let headerHtml = data.columns->Array.map(c => `
`)->Array.join("") - let rowsHtml = - data.rows - ->Array.map(row => { - let cells = row->Array.map(cell => ``)->Array.join("") - `${cells}` - }) - ->Array.join("\n") - - let tableStyle = "border-collapse:collapse;width:100%;font-size:0.85rem;" - let cellStyle = "border:1px solid var(--border);padding:0.3rem 0.6rem;text-align:left;" - let headerStyle = - cellStyle ++ "background:var(--bg-secondary);font-weight:600;color:var(--accent);" - - // Inline styles since we're injecting HTML - let styledTable = String.replaceAll( - String.replaceAll( - `
${c}${cell}
${headerHtml}${rowsHtml}
`, - "", - ``, - ), - "", - ``, - ) - - let source = if backendConnected.contents { "live" } else { "demo" } - - output["innerHTML"] = - styledTable ++ - `\n(${Int.toString(data.row_count)} rows, ${Float.toFixed(data.timing_ms, ~digits=1)}ms, ${source})` - } - | DemoExecutor.ExplainResult(text) => { - let escaped = String.replaceAll(String.replaceAll(text, "<", "<"), ">", ">") - output["innerHTML"] = `
${escaped}
` - } - | DemoExecutor.Error(msg) => { - output["innerHTML"] = `ERROR: ${msg}` - } - } -} - -// === Lint helper === - -and runLint = (query: string, lintBar: {..}) => { - let diagnostics = Linter.lint(query, ~vclDt=vclDtMode.contents) - let errors = diagnostics->Array.filter(d => d.severity == Linter.Error)->Array.length - let warnings = diagnostics->Array.filter(d => d.severity == Linter.Warning)->Array.length - let hints = diagnostics->Array.filter(d => d.severity == Linter.Hint)->Array.length - - if errors > 0 { - lintBar["innerHTML"] = - `${Int.toString(errors)} error(s), ${Int.toString(warnings)} warning(s), ${Int.toString(hints)} hint(s)` - } else if warnings > 0 { - lintBar["innerHTML"] = - `${Int.toString(warnings)} warning(s), ${Int.toString(hints)} hint(s)` - } else if hints > 0 { - lintBar["innerHTML"] = `${Int.toString(hints)} hint(s)` - } else { - lintBar["innerHTML"] = `No issues` - } -} - -// === setTimeout binding === -@val external setTimeout: (unit => unit, int) => int = "setTimeout" - -// === Boot === - -// Wait for DOM -addEventListener(document, "DOMContentLoaded", _ => init()) diff --git a/verisimdb/playground/src/DemoExecutor.res b/verisimdb/playground/src/DemoExecutor.res deleted file mode 100644 index e68891eb..00000000 --- a/verisimdb/playground/src/DemoExecutor.res +++ /dev/null @@ -1,98 +0,0 @@ -// SPDX-License-Identifier: MPL-2.0 -// Demo query executor — simulates VeriSimDB responses offline. -// In production, this would call the real verisim-api endpoint. - -type queryResult = { - columns: array, - rows: array>, - timing_ms: float, - row_count: int, -} - -type executeResult = - | Success(queryResult) - | ExplainResult(string) - | Error(string) - -/// Generate demo data based on the query modalities. -let execute = (query: string, ~vclDt: bool=false): executeResult => { - let upper = String.toUpperCase(query) - let startTime = Date.now() - - // EXPLAIN mode - if String.includes(upper, "EXPLAIN") { - let modalities = VclKeywords.modalities->Array.filter(m => String.includes(upper, m)) - let plan = ref("=== EXPLAIN OUTPUT ===\n\n") - plan := plan.contents ++ "Strategy: " ++ (if Array.length(modalities) >= 2 { "Parallel" } else { "Sequential" }) ++ "\n\n" - - modalities->Array.forEachWithIndex((m, i) => { - let cost = switch m { - | "TEMPORAL" => "30.0" - | "VECTOR" => "50.0" - | "DOCUMENT" => "80.0" - | "GRAPH" => "150.0" - | "TENSOR" => "200.0" - | "SEMANTIC" => "300.0" - | _ => "100.0" - } - plan := plan.contents ++ `Step ${Int.toString(i + 1)}: ${m} query\n` - plan := plan.contents ++ ` Estimated cost: ${cost}ms\n` - plan := plan.contents ++ ` Estimated rows: 100\n` - plan := plan.contents ++ ` Selectivity: 0.5\n\n` - }) - - if vclDt && String.includes(upper, "PROOF") { - plan := plan.contents ++ "Proof verification: ENABLED\n" - plan := plan.contents ++ "ZKP scheme: PLONK\n" - plan := plan.contents ++ "Circuit compilation: deferred\n" - } - - let elapsed = Date.now() -. startTime - plan := plan.contents ++ `\nPlan generated in ${Float.toFixed(elapsed, ~digits=1)}ms\n` - ExplainResult(plan.contents) - } - // DELETE/UPDATE — always deny in demo mode - else if String.includes(upper, "DELETE") || String.includes(upper, "UPDATE") { - Error("Write operations are disabled in demo mode") - } - // SELECT queries — generate demo data - else if String.includes(upper, "SELECT") { - let modalities = VclKeywords.modalities->Array.filter(m => String.includes(upper, m)) - if Array.length(modalities) == 0 { - Error("No modalities specified in SELECT clause") - } else { - let columns = ["id"]->Array.concat( - modalities->Array.map(m => String.toLowerCase(m) ++ "_data") - ) - let rowCount = if String.includes(upper, "LIMIT") { 5 } else { 10 } - let rows = Array.fromInitializer(~length=rowCount, i => { - let id = `hexad-${Int.toString(1000 + i)}` - let modalityData = modalities->Array.map(m => - switch m { - | "GRAPH" => `{edges: ${Int.toString(3 + i)}, type: "Entity"}` - | "VECTOR" => `[${Float.toFixed(Float.fromInt(i) *. 0.1, ~digits=2)}, 0.50, 0.30]` - | "TENSOR" => `shape=[3,3], dtype=f32` - | "SEMANTIC" => if vclDt { `{proof: "verified", scheme: "PLONK"}` } else { `{types: ["Thing"]}` } - | "DOCUMENT" => `"Sample document ${Int.toString(i + 1)}"` - | "TEMPORAL" => `{version: ${Int.toString(i + 1)}, ts: "2026-02-28"}` - | "PROVENANCE" => `{source: "scan-v1", actor: "hypatia", chain_length: ${Int.toString(i + 1)}}` - | "SPATIAL" => `{lat: ${Float.toFixed(51.5 +. Float.fromInt(i) *. 0.01, ~digits=4)}, lon: -0.1278}` - | _ => "null" - } - ) - [id]->Array.concat(modalityData) - }) - - let elapsed = Date.now() -. startTime +. 15.0 // simulate some latency - - Success({ - columns, - rows, - timing_ms: elapsed, - row_count: rowCount, - }) - } - } else { - Error("Unrecognized query — VCL queries must start with SELECT, EXPLAIN, INSERT, UPDATE, or DELETE") - } -} diff --git a/verisimdb/playground/src/Examples.res b/verisimdb/playground/src/Examples.res deleted file mode 100644 index 717167df..00000000 --- a/verisimdb/playground/src/Examples.res +++ /dev/null @@ -1,107 +0,0 @@ -// SPDX-License-Identifier: MPL-2.0 -// Example VCL queries for the playground. -// Covers all 8 octad modalities, real backend queries, and VCL-UT proof types. - -type example = { - label: string, - query: string, - vclDt: bool, -} - -let examples = [ - // --- Standard VCL examples --- - { - label: "List all hexads", - query: "SELECT * FROM hexads LIMIT 10", - vclDt: false, - }, - { - label: "Full-text search", - query: "SEARCH TEXT 'multimodal database' LIMIT 10", - vclDt: false, - }, - { - label: "Vector similarity search", - query: "SEARCH VECTOR [0.1, 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8] LIMIT 5", - vclDt: false, - }, - { - label: "Graph traversal", - query: "SEARCH RELATED 'entity-1' BY 'relates_to'", - vclDt: false, - }, - { - label: "Insert a hexad", - query: "INSERT INTO hexads (title, body)\nVALUES ('My Entity', 'A multimodal entity in VeriSimDB')", - vclDt: false, - }, - { - label: "Show server status", - query: "SHOW STATUS", - vclDt: false, - }, - { - label: "Show drift metrics", - query: "SHOW DRIFT", - vclDt: false, - }, - { - label: "Explain a query", - query: "EXPLAIN SELECT * FROM hexads WHERE id = 'my-entity' LIMIT 1", - vclDt: false, - }, - { - label: "Count hexads", - query: "COUNT hexads", - vclDt: false, - }, - { - label: "Multi-modality query (demo)", - query: "SELECT GRAPH, VECTOR, DOCUMENT, PROVENANCE\nFROM HEXAD\nWHERE name CONTAINS 'example'\nORDER BY score DESC\nLIMIT 20", - vclDt: false, - }, - { - label: "Temporal query (demo)", - query: "SELECT TEMPORAL, PROVENANCE\nFROM HEXAD\nAT TIME '2026-02-28T00:00:00Z'\nWHERE id = 'entity-123'\nLIMIT 1", - vclDt: false, - }, - { - label: "Federation query (demo)", - query: "SELECT GRAPH\nFROM FEDERATION STORE 'remote-cluster-1'\nHEXAD\nWHERE region = 'eu-west'\nLIMIT 25", - vclDt: false, - }, - // --- VCL-UT examples --- - { - label: "Proof of existence (VCL-UT)", - query: "SELECT SEMANTIC\nFROM HEXAD\nPROOF EXISTENCE\nTHRESHOLD 0.95\nWHERE type = 'Certificate'\nLIMIT 10", - vclDt: true, - }, - { - label: "Integrity proof (VCL-UT)", - query: "SELECT SEMANTIC, DOCUMENT\nFROM HEXAD\nPROOF INTEGRITY\nTHRESHOLD 0.99\nWHERE classification = 'audit-trail'\nLIMIT 5", - vclDt: true, - }, - { - label: "Consistency check (VCL-UT)", - query: "SELECT GRAPH, SEMANTIC\nFROM HEXAD\nPROOF CONSISTENCY\nTHRESHOLD 0.9\nWHERE DRIFT THRESHOLD 0.1\nLIMIT 20", - vclDt: true, - }, - { - label: "Provenance proof (VCL-UT)", - query: "SELECT PROVENANCE, SEMANTIC\nFROM HEXAD\nPROOF PROVENANCE\nTHRESHOLD 0.95\nWHERE source = 'verified-origin'\nLIMIT 10", - vclDt: true, - }, - { - label: "Freshness proof (VCL-UT)", - query: "SELECT TEMPORAL, SEMANTIC\nFROM HEXAD\nPROOF FRESHNESS\nTHRESHOLD 0.99\nWHERE age_ms < 86400000\nLIMIT 10", - vclDt: true, - }, - { - label: "Multi-proof composition (VCL-UT)", - query: "SELECT SEMANTIC, PROVENANCE, TEMPORAL\nFROM HEXAD\nPROOF EXISTENCE AND INTEGRITY AND FRESHNESS\nTHRESHOLD 0.95\nWHERE type = 'critical-entity'\nLIMIT 5", - vclDt: true, - }, -] - -let forMode = (vclDt: bool): array => - examples->Array.filter(e => !e.vclDt || vclDt) diff --git a/verisimdb/playground/src/Formatter.res b/verisimdb/playground/src/Formatter.res deleted file mode 100644 index 34e67abc..00000000 --- a/verisimdb/playground/src/Formatter.res +++ /dev/null @@ -1,67 +0,0 @@ -// SPDX-License-Identifier: MPL-2.0 -// VCL formatter — canonical formatting for queries. - -/// Clause-starting keywords that get their own line. -let clauseStarters = [ - "SELECT", "FROM", "WHERE", "ORDER", "GROUP", "HAVING", "LIMIT", - "OFFSET", "JOIN", "ON", "WITH", "SET", "INTO", "VALUES", - "TRAVERSE", "PROOF", "EXPLAIN", -] - -let formatVcl = (query: string): string => { - let upper = String.toUpperCase - let tokens = - Js.String2.splitByRe(query, %re("/(\s+|'[^']*'|\"[^\"]*\")/")) - ->Array.filterMap(t => t) - ->Array.filter(t => String.trim(t) !== "") - - let result = ref("") - let isFirst = ref(true) - - tokens->Array.forEach(token => { - let trimmed = String.trim(token) - if trimmed === "" { - // Whitespace — will be normalized - if !(String.endsWith(result.contents, " ") || String.endsWith(result.contents, "\n")) { - result := result.contents ++ " " - } - } else if String.startsWith(trimmed, "'") || String.startsWith(trimmed, "\"") { - // String literal — preserve as-is - result := result.contents ++ trimmed - } else { - let word = upper(trimmed) - let formatted = if VclKeywords.isKeyword(word) || VclKeywords.isModality(word) { - word - } else { - trimmed - } - - if clauseStarters->Array.includes(word) && !isFirst.contents { - // Remove trailing space - if String.endsWith(result.contents, " ") { - result := String.slice(result.contents, ~start=0, ~end=String.length(result.contents) - 1) - } - // Check EXPLAIN + SELECT same line - if word === "SELECT" && String.endsWith(String.trim(result.contents), "EXPLAIN") { - result := result.contents ++ " " ++ formatted - } else { - result := result.contents ++ "\n" ++ formatted - } - } else if word === "AND" || word === "OR" { - if String.endsWith(result.contents, " ") { - result := String.slice(result.contents, ~start=0, ~end=String.length(result.contents) - 1) - } - result := result.contents ++ "\n " ++ formatted - } else { - if !(String.endsWith(result.contents, " ") || String.endsWith(result.contents, "\n") || result.contents === "") { - result := result.contents ++ " " - } - result := result.contents ++ formatted - } - - isFirst := false - } - }) - - String.trim(result.contents) -} diff --git a/verisimdb/playground/src/Highlighter.res b/verisimdb/playground/src/Highlighter.res deleted file mode 100644 index 9d0e8697..00000000 --- a/verisimdb/playground/src/Highlighter.res +++ /dev/null @@ -1,83 +0,0 @@ -// SPDX-License-Identifier: MPL-2.0 -// VCL syntax highlighting for the playground editor. -// Produces HTML spans with CSS classes for keyword colouring. - -let highlightVcl = (text: string, ~vclDt: bool=false): string => { - let result = ref("") - let chars = String.split(text, "") - let len = Array.length(chars) - let i = ref(0) - - while i.contents < len { - let ch = chars[i.contents]->Option.getOr("") - - // String literals - if ch == "'" || ch == "\"" { - let quote = ch - let start = i.contents - i := i.contents + 1 - while i.contents < len && chars[i.contents]->Option.getOr("") != quote { - if chars[i.contents]->Option.getOr("") == "\\" { - i := i.contents + 1 - } - i := i.contents + 1 - } - if i.contents < len { - i := i.contents + 1 - } - let slice = String.slice(text, ~start, ~end=i.contents) - result := result.contents ++ `${slice}` - } - // Comments (-- single line) - else if ch == "-" && i.contents + 1 < len && chars[i.contents + 1]->Option.getOr("") == "-" { - let start = i.contents - while i.contents < len && chars[i.contents]->Option.getOr("") != "\n" { - i := i.contents + 1 - } - let slice = String.slice(text, ~start, ~end=i.contents) - result := result.contents ++ `${slice}` - } - // Words - else if Js.Re.test_(%re("/[a-zA-Z_]/"), ch) { - let start = i.contents - while i.contents < len && Js.Re.test_(%re("/[a-zA-Z0-9_]/"), chars[i.contents]->Option.getOr("")) { - i := i.contents + 1 - } - let word = String.slice(text, ~start, ~end=i.contents) - let upper = String.toUpperCase(word) - - if VclKeywords.isModality(upper) { - result := result.contents ++ `${word}` - } else if VclKeywords.isProofType(upper) && vclDt { - result := result.contents ++ `${word}` - } else if VclKeywords.isKeyword(upper) { - result := result.contents ++ `${word}` - } else { - result := result.contents ++ word - } - } - // Numbers - else if Js.Re.test_(%re("/[0-9]/"), ch) { - let start = i.contents - while i.contents < len && Js.Re.test_(%re("/[0-9.]/"), chars[i.contents]->Option.getOr("")) { - i := i.contents + 1 - } - let num = String.slice(text, ~start, ~end=i.contents) - result := result.contents ++ `${num}` - } - // Everything else - else { - // HTML-escape < > & - let escaped = switch ch { - | "<" => "<" - | ">" => ">" - | "&" => "&" - | c => c - } - result := result.contents ++ escaped - i := i.contents + 1 - } - } - - result.contents -} diff --git a/verisimdb/playground/src/Linter.res b/verisimdb/playground/src/Linter.res deleted file mode 100644 index d38e1a41..00000000 --- a/verisimdb/playground/src/Linter.res +++ /dev/null @@ -1,140 +0,0 @@ -// SPDX-License-Identifier: MPL-2.0 -// VCL client-side linter — mirrors the Rust linter rules (VCL001–VCL011). - -type severity = Hint | Warning | Error - -type diagnostic = { - code: string, - severity: severity, - message: string, -} - -let severityToString = s => - switch s { - | Hint => "hint" - | Warning => "warning" - | Error => "error" - } - -let lint = (query: string, ~vclDt: bool=false): array => { - let diagnostics = [] - let upper = String.toUpperCase(query) - let tokens = - Js.String2.splitByRe(String.trim(upper), %re("/\s+/")) - ->Array.filterMap(x => x) - - let has = tok => tokens->Array.includes(tok) - - let isSelect = has("SELECT") - let isDelete = has("DELETE") - let isUpdate = has("UPDATE") - let isExplain = has("EXPLAIN") - - // VCL001: Missing LIMIT - if isSelect && !has("LIMIT") && !isExplain { - diagnostics->Array.push({ - code: "VCL001", - severity: Warning, - message: "Query lacks LIMIT clause — may return unbounded results", - }) - } - - // VCL002: SELECT all modalities - if isSelect { - let count = - VclKeywords.modalities->Array.filter(m => has(m))->Array.length - if count >= 6 { - diagnostics->Array.push({ - code: "VCL002", - severity: Hint, - message: "Query selects " ++ Int.toString(count) ++ " of 8 modalities — consider selecting only what you need", - }) - } - } - - // VCL003: Semantic without PROOF - if has("SEMANTIC") && !has("PROOF") && isSelect { - diagnostics->Array.push({ - code: "VCL003", - severity: if vclDt { Error } else { Warning }, - message: "Semantic modality accessed without PROOF clause", - }) - } - - // VCL004: TRAVERSE without DEPTH - if has("TRAVERSE") && !has("DEPTH") { - diagnostics->Array.push({ - code: "VCL004", - severity: Error, - message: "TRAVERSE without DEPTH limit — may explore entire graph", - }) - } - - // VCL005: DRIFT without THRESHOLD - if (has("DRIFT") || has("CONSISTENCY")) && !has("THRESHOLD") { - diagnostics->Array.push({ - code: "VCL005", - severity: Hint, - message: "DRIFT/CONSISTENCY check without THRESHOLD — using implicit default", - }) - } - - // VCL006: ORDER BY without LIMIT - if has("ORDER") && !has("LIMIT") && isSelect { - diagnostics->Array.push({ - code: "VCL006", - severity: Warning, - message: "ORDER BY without LIMIT — sorting potentially unbounded result set", - }) - } - - // VCL007: Dangerous write without WHERE - if (isDelete || isUpdate) && !has("WHERE") { - diagnostics->Array.push({ - code: "VCL007", - severity: Error, - message: "DELETE/UPDATE without WHERE clause — affects all entities", - }) - } - - // VCL010: Multi-modality without EXPLAIN - if isSelect && !isExplain { - let count = - VclKeywords.modalities->Array.filter(m => has(m))->Array.length - if count >= 3 { - diagnostics->Array.push({ - code: "VCL010", - severity: Hint, - message: "Multi-modality query — consider running EXPLAIN first", - }) - } - } - - // VCL011: FEDERATION without STORE - if has("FEDERATION") && !has("STORE") { - diagnostics->Array.push({ - code: "VCL011", - severity: Warning, - message: "FEDERATION query without STORE — will query all federated instances", - }) - } - - // VCL-UT specific: PROOF required for all semantic access - if vclDt && isSelect && has("SEMANTIC") && !has("PROOF") { - // Already covered by VCL003 with Error severity - ignore() - } - - // Sort: errors first - diagnostics->Array.sort((a, b) => { - let severityOrder = s => - switch s { - | Error => 0 - | Warning => 1 - | Hint => 2 - } - Float.fromInt(severityOrder(a.severity) - severityOrder(b.severity)) - }) - - diagnostics -} diff --git a/verisimdb/playground/src/VclKeywords.res b/verisimdb/playground/src/VclKeywords.res deleted file mode 100644 index f3cd4750..00000000 --- a/verisimdb/playground/src/VclKeywords.res +++ /dev/null @@ -1,38 +0,0 @@ -// SPDX-License-Identifier: MPL-2.0 -// VCL keyword definitions shared across syntax highlighting, completion, and linting. -// Updated for the octad architecture (8 modalities) and 11 proof types. - -let keywords = [ - "SELECT", "FROM", "WHERE", "PROOF", "LIMIT", "OFFSET", "ORDER", "BY", - "GROUP", "HAVING", "AS", "AND", "OR", "NOT", "IN", "BETWEEN", "LIKE", - "EXISTS", "CONTAINS", "SIMILAR", "TO", "TRAVERSE", "DEPTH", "THRESHOLD", - "DRIFT", "CONSISTENCY", "AT", "TIME", "EXPLAIN", "INSERT", "UPDATE", - "DELETE", "SET", "INTO", "VALUES", "CREATE", "DROP", "ALTER", "JOIN", - "ON", "WITH", "FEDERATION", "STORE", "HEXAD", "ALL", "ASC", "DESC", - "COUNT", "SUM", "AVG", "MIN", "MAX", "DISTINCT", "ANALYZE", - "SHOW", "STATUS", "SEARCH", "TEXT", "RELATED", "WITHIN", "RADIUS", - "BOUNDS", "NEAREST", "REFLECT", -] - -/// Octad modalities — 8 stores that form the core of each entity. -let modalities = [ - "GRAPH", "VECTOR", "TENSOR", "SEMANTIC", "DOCUMENT", "TEMPORAL", - "PROVENANCE", "SPATIAL", -] - -/// All 11 proof types supported by the VCL-UT type checker. -let proofTypes = [ - "EXISTENCE", "CONSISTENCY", "INTEGRITY", "PROVENANCE", - "FRESHNESS", "ACCESS", "CITATION", "CUSTOM", - "ZKP", "PROVEN", "SANCTIFY", -] - -/// VCL-UT specific keywords (only active in VCL-UT mode). -let vclDtKeywords = [ - "PROOF", "THRESHOLD", "VERIFY", "CERTIFY", "ATTEST", - "WITNESS", "CIRCUIT", "COMMITMENT", -] - -let isKeyword = word => keywords->Array.includes(String.toUpperCase(word)) -let isModality = word => modalities->Array.includes(String.toUpperCase(word)) -let isProofType = word => proofTypes->Array.includes(String.toUpperCase(word)) diff --git a/verisimdb/proven-coherence.md b/verisimdb/proven-coherence.md deleted file mode 100644 index 91a801b9..00000000 --- a/verisimdb/proven-coherence.md +++ /dev/null @@ -1,258 +0,0 @@ -import React, { useState, useEffect } from 'react'; -import { - Shield, - AlertTriangle, - CheckCircle, - FileText, - Clock, - Users, - Search, - ChevronRight, - Fingerprint, - Info, - History -} from 'lucide-react'; - -const App = () => { - const [selectedIssue, setSelectedIssue] = useState(null); - const [isSigning, setIsSigning] = useState(false); - const [auditLog, setAuditLog] = useState([ - { id: 1, action: "Policy Updated", target: "0x882A...", actor: "did:verisim:custodian_02", time: "2h ago" }, - { id: 2, action: "Manual Repair", target: "0x441F...", actor: "did:verisim:custodian_01", time: "5h ago" } - ]); - - const pendingIssues = [ - { - id: "0x12AB...990F", - type: "Formal Drift (Contract Breach)", - modality: "Graph / Semantic", - severity: "High", - contract: "CitationContract", - breach: "Invariant 'claim_validity' failed.", - cause: "Reference 0x990F... has status 'retracted'.", - implication: "Approving this update will logically invalidate the authority of the parent Octad.", - timestamp: "12 mins ago" - }, - { - id: "0xCC21...110E", - type: "Topological Mutation", - modality: "Graph", - severity: "Medium", - contract: "TaxonomyContract", - breach: "Edge-addition rate exceeds Poisson threshold (λ=0.05).", - cause: "Rapid batch update of 450 nodes detected.", - implication: "Possible systematic bias or archival error in metadata ingestion.", - timestamp: "45 mins ago" - } - ]; - - const handleSign = () => { - setIsSigning(true); - // Simulate sactify-php + proven ZKP generation - setTimeout(() => { - setAuditLog([ - { - id: Date.now(), - action: "Formal Repair Signed", - target: selectedIssue.id, - actor: "did:verisim:custodian_01", - time: "Just now" - }, - ...auditLog - ]); - setIsSigning(false); - setSelectedIssue(null); - }, 2000); - }; - - return ( -
- {/* Header */} - - -
- {/* Left: Pending Queue */} -
-
-

- - Pending Reviews -

- - {pendingIssues.length} New - -
- -
- {pendingIssues.map((issue) => ( - - ))} -
-
- - {/* Center: Detailed Evidence (Proven Explainability) */} -
- {selectedIssue ? ( -
-
-

Explainability Trace

-

Formal Verification Engine (Proven v1.1)

-
- -
-
-

Violation Details

-
- -
-

{selectedIssue.breach}

-

{selectedIssue.cause}

-
-
-
- -
-

Logical Implication

-
- -

{selectedIssue.implication}

-
-
- -
-

Proposed Repair

-
-

Revert citation mapping to state 0x4F2... and append retraction proof to the Temporal Ledger.

-
- - COHERENCE RESTORED AFTER REPAIR -
-
-
-
- -
- -
-
- ) : ( -
- -

Select a pending issue to review formal drift evidence and authorize repairs.

-
- )} -
- - {/* Right: Audit Log & Stats */} -
-
-

- - Recent Activity -

-
- {auditLog.map(log => ( -
-

{log.action}

-

Target: {log.target}

-
- {log.actor.split(':').pop()} - {log.time} -
-
- ))} -
-
- -
-

- - Epistemic Health -

-
-
-
- Formal Coherence - 98.2% -
-
-
-
-
-
-
- Quorums Reached - 24 / 24 -
-
-
-
-
-
-
-
-
-
- ); -}; - -const Activity = ({ className }) => ( - - - -); - -export default App; diff --git a/verisimdb/references.bib b/verisimdb/references.bib deleted file mode 100644 index 6a0040e3..00000000 --- a/verisimdb/references.bib +++ /dev/null @@ -1,147 +0,0 @@ -% SPDX-License-Identifier: MPL-2.0 -% Copyright (c) 2026 Jonathan D.A. Jewell (hyperpolymath) -% -% References for VeriSimDB: multi-model database with VCL -% (8 query modalities), vector/graph/document storage -% - -@article{codd1970, - author = {Edgar F. Codd}, - title = {A Relational Model of Data for Large Shared Data Banks}, - journal = {Communications of the ACM}, - volume = {13}, - number = {6}, - pages = {377--387}, - year = {1970}, - note = {Introduces the relational model. VeriSimDB extends Codd's - foundational abstractions to multi-model storage while - preserving relational algebraic reasoning where applicable.} -} - -@article{stonebraker2005, - author = {Michael Stonebraker and Ugur \c{C}etintemel}, - title = {``One Size Fits All'': An Idea Whose Time Has Come and Gone}, - journal = {Proceedings of the International Conference on Data Engineering - (ICDE)}, - pages = {2--11}, - year = {2005}, - publisher = {IEEE}, - note = {Argues that specialised database engines outperform general-purpose - ones. VeriSimDB's 8 query modalities address this by providing - specialised access paths within a unified system.} -} - -@article{johnson2021, - author = {Jeff Johnson and Matthijs Douze and Herv\'{e} J\'{e}gou}, - title = {Billion-Scale Similarity Search with {GPUs}}, - journal = {IEEE Transactions on Big Data}, - volume = {7}, - number = {3}, - pages = {535--547}, - year = {2021}, - note = {Describes FAISS and approximate nearest neighbor techniques at - scale. VeriSimDB's vector storage modality builds on these - indexing strategies for similarity queries.} -} - -@book{pierce2002, - author = {Benjamin C. Pierce}, - title = {Types and Programming Languages}, - publisher = {MIT Press}, - year = {2002}, - note = {Reference for VeriSimDB's type system design. VCL's 8 query - modalities are type-checked to ensure well-formedness of - cross-model queries.} -} - -@article{plotkin1981, - author = {Gordon D. Plotkin}, - title = {A Structural Approach to Operational Semantics}, - journal = {Technical Report DAIMI FN-19, Computer Science Department, - Aarhus University}, - year = {1981}, - note = {Establishes structural operational semantics (SOS). VCL's - formal semantics use SOS-style rules to define the meaning - of each query modality.} -} - -@inproceedings{lu2019, - author = {Yi Lu and Anil Shanbhag and Alekh Jindal and Samuel Madden}, - title = {Multi-Model Databases: A New Journey to Handle the Variety - of Data}, - booktitle = {ACM Computing Surveys}, - volume = {52}, - number = {3}, - pages = {1--38}, - year = {2019}, - note = {Comprehensive survey of multi-model database architectures. - VeriSimDB's unified storage layer draws on the taxonomy and - design patterns catalogued here.} -} - -@inproceedings{malkov2020, - author = {Yury A. Malkov and Dmitry A. Yashunin}, - title = {Efficient and Robust Approximate Nearest Neighbor Search - Using Hierarchical Navigable Small World Graphs}, - journal = {IEEE Transactions on Pattern Analysis and Machine Intelligence}, - volume = {42}, - number = {4}, - pages = {824--836}, - year = {2020}, - note = {Introduces HNSW, a graph-based index for approximate nearest - neighbor search. VeriSimDB uses HNSW-derived structures - for its vector similarity modality.} -} - -@inproceedings{angles2018, - author = {Renzo Angles and Marcelo Arenas and Pablo Barcel\'{o} - and Aidan Hogan and Juan Reutter and Domagoj Vrgo\v{c}}, - title = {Foundations of Modern Query Languages for Graph Databases}, - journal = {ACM Computing Surveys}, - volume = {50}, - number = {5}, - pages = {1--40}, - year = {2018}, - note = {Surveys formal foundations of graph query languages. VeriSimDB's - graph modality builds on the path pattern semantics and - regular path query algebra described herein.} -} - -@article{stonebraker2018, - author = {Michael Stonebraker}, - title = {The Case for Polystores}, - journal = {ACM SIGMOD Record}, - volume = {47}, - number = {3}, - pages = {8--9}, - year = {2018}, - note = {Argues for polystore architectures integrating heterogeneous - data stores. VeriSimDB takes the alternative approach of - a single multi-model engine with polystore-level flexibility.} -} - -@inproceedings{ozsu2007, - author = {M. Tamer \"{O}zsu and Patrick Valduriez}, - title = {Principles of Distributed Database Systems}, - publisher = {Springer}, - edition = {3rd}, - year = {2011}, - note = {Comprehensive treatment of distributed data management. - VeriSimDB's distributed query execution across modalities - draws on the query decomposition and optimization techniques - presented here.} -} - -@article{hellerstein2007, - author = {Joseph M. Hellerstein and Michael Stonebraker and James Hamilton}, - title = {Architecture of a Database System}, - journal = {Foundations and Trends in Databases}, - volume = {1}, - number = {2}, - pages = {141--259}, - year = {2007}, - note = {Detailed walkthrough of database system architecture. VeriSimDB's - storage manager, query processor, and transaction subsystem - follow the architectural patterns described here, extended - for multi-model operation.} -} diff --git a/verisimdb/rescript.json b/verisimdb/rescript.json deleted file mode 100644 index f4c88c28..00000000 --- a/verisimdb/rescript.json +++ /dev/null @@ -1,13 +0,0 @@ -{ - "name": "verisimdb-registry", - "sources": [ - { "dir": "src", "subdirs": true } - ], - "package-specs": [ - { "module": "es6", "in-source": true } - ], - "suffix": ".res.mjs", - "bs-dependencies": [ - "@rescript/core" - ] -} diff --git a/verisimdb/rust-core/fuzz/Cargo.toml b/verisimdb/rust-core/fuzz/Cargo.toml deleted file mode 100644 index 86be4290..00000000 --- a/verisimdb/rust-core/fuzz/Cargo.toml +++ /dev/null @@ -1,22 +0,0 @@ -# SPDX-License-Identifier: MPL-2.0 - -[package] -name = "verisimdb-fuzz" -version = "0.0.0" -publish = false -edition = "2021" - -[package.metadata] -cargo-fuzz = true - -[dependencies] -libfuzzer-sys = "0.4" - -[dependencies.verisim-api] -path = "../verisim-api" - -[[bin]] -name = "fuzz_vcl_parser" -path = "fuzz_targets/fuzz_vcl_parser.rs" -test = false -doc = false diff --git a/verisimdb/rust-core/fuzz/fuzz_targets/fuzz_vcl_parser.rs b/verisimdb/rust-core/fuzz/fuzz_targets/fuzz_vcl_parser.rs deleted file mode 100644 index 05242456..00000000 --- a/verisimdb/rust-core/fuzz/fuzz_targets/fuzz_vcl_parser.rs +++ /dev/null @@ -1,25 +0,0 @@ -// SPDX-License-Identifier: MPL-2.0 -// -// Fuzz target for the VCL parser. -// Run with: cargo +nightly fuzz run fuzz_vcl_parser -// -// This fuzzer feeds arbitrary byte strings to the VCL parser to find -// panics, hangs, or memory safety issues. The parser should gracefully -// reject invalid input without crashing. - -#![no_main] - -use libfuzzer_sys::fuzz_target; - -fuzz_target!(|data: &[u8]| { - // Only attempt to parse valid UTF-8 strings — the VCL parser - // operates on &str, not raw bytes. - if let Ok(input) = std::str::from_utf8(data) { - // Limit input size to prevent timeouts on extremely long strings - if input.len() <= 4096 { - // The parser should never panic on any valid UTF-8 input. - // We don't care about the result — only that it doesn't crash. - let _ = verisim_api::vcl::parse(input); - } - } -}); diff --git a/verisimdb/rust-core/verisim-api/Cargo.toml b/verisimdb/rust-core/verisim-api/Cargo.toml deleted file mode 100644 index 18956ab2..00000000 --- a/verisimdb/rust-core/verisim-api/Cargo.toml +++ /dev/null @@ -1,60 +0,0 @@ -# SPDX-License-Identifier: MPL-2.0 - -[package] -name = "verisim-api" -description = "HTTP API server for VeriSimDB" -version.workspace = true -edition.workspace = true -authors.workspace = true -license.workspace = true - -[dependencies] -verisim-octad = { path = "../verisim-octad" } -verisim-normalizer = { path = "../verisim-normalizer" } -verisim-drift = { path = "../verisim-drift" } -verisim-graph = { path = "../verisim-graph" } -verisim-vector = { path = "../verisim-vector" } -verisim-document = { path = "../verisim-document" } -verisim-tensor = { path = "../verisim-tensor" } -verisim-semantic = { path = "../verisim-semantic" } -verisim-temporal = { path = "../verisim-temporal" } -verisim-provenance = { path = "../verisim-provenance" } -verisim-spatial = { path = "../verisim-spatial" } -verisim-planner = { path = "../verisim-planner" } - -axum.workspace = true -tokio.workspace = true -tower.workspace = true -hyper.workspace = true -serde.workspace = true -serde_json.workspace = true -chrono.workspace = true -thiserror.workspace = true -tracing.workspace = true -tracing-subscriber.workspace = true -prometheus.workspace = true -reqwest.workspace = true -async-graphql.workspace = true -async-graphql-axum.workspace = true -tonic.workspace = true -tonic-prost.workspace = true -prost.workspace = true -prost-types.workspace = true -sha2.workspace = true -axum-server.workspace = true -rustls.workspace = true -hex = "0.4" - -[features] -default = [] -# Enable persistent storage backends (redb for graph, file-backed Tantivy for documents, WAL). -# Requires VERISIM_PERSISTENCE_DIR environment variable at runtime. -persistent = ["verisim-graph/redb-backend"] - -# Build-dependencies removed: protobuf code is pre-generated at src/proto/verisim.rs. -# To regenerate after changing proto/verisim.proto, run: -# protoc --prost_out=src/proto proto/verisim.proto -# Or use tonic-build manually. - -[dev-dependencies] -proptest.workspace = true diff --git a/verisimdb/rust-core/verisim-api/build.rs b/verisimdb/rust-core/verisim-api/build.rs deleted file mode 100644 index d2fa725b..00000000 --- a/verisimdb/rust-core/verisim-api/build.rs +++ /dev/null @@ -1,15 +0,0 @@ -// SPDX-License-Identifier: MPL-2.0 -// -// Build script for verisim-api. -// -// Protobuf code is pre-generated at src/proto/verisim.rs and committed to -// the repository. This eliminates protoc as a build-time dependency. -// -// To regenerate after changing proto/verisim.proto, install protoc and run: -// -// cd rust-core/verisim-api -// protoc --prost_out=src/proto proto/verisim.proto - -fn main() { - // No-op: proto code is pre-generated. -} diff --git a/verisimdb/rust-core/verisim-api/proto/verisim.proto b/verisimdb/rust-core/verisim-api/proto/verisim.proto deleted file mode 100644 index 1ef65bce..00000000 --- a/verisimdb/rust-core/verisim-api/proto/verisim.proto +++ /dev/null @@ -1,178 +0,0 @@ -// SPDX-License-Identifier: MPL-2.0 -// VeriSimDB gRPC service definitions. - -syntax = "proto3"; - -package verisim; - -// ============================================================================ -// Planner Service -// ============================================================================ - -service VeriSimPlanner { - // Optimize a logical plan into a physical plan. - rpc OptimizePlan(LogicalPlanRequest) returns (PhysicalPlanResponse); - // Generate EXPLAIN output for a logical plan. - rpc ExplainPlan(LogicalPlanRequest) returns (ExplainResponse); - // Get current planner configuration. - rpc GetConfig(Empty) returns (PlannerConfigResponse); - // Update planner configuration. - rpc SetConfig(PlannerConfigRequest) returns (PlannerConfigResponse); - // Get per-modality statistics snapshot. - rpc GetStats(Empty) returns (StatsResponse); -} - -// ============================================================================ -// Octad Service -// ============================================================================ - -service VeriSimOctad { - // Create a new octad. - rpc Create(OctadCreateRequest) returns (OctadResponse); - // Get a octad by ID. - rpc Get(OctadIdRequest) returns (OctadResponse); - // Update an existing octad. - rpc Update(OctadUpdateRequest) returns (OctadResponse); - // Delete a octad. - rpc Delete(OctadIdRequest) returns (Empty); - // Full-text search. - rpc SearchText(TextSearchRequest) returns (SearchResponse); - // Vector similarity search. - rpc SearchVector(VectorSearchRequest) returns (SearchResponse); -} - -// ============================================================================ -// Common Messages -// ============================================================================ - -message Empty {} - -// ============================================================================ -// Planner Messages -// ============================================================================ - -// Logical plan passed as JSON string (mirrors the REST API). -// This avoids duplicating the full plan type hierarchy in protobuf. -message LogicalPlanRequest { - string plan_json = 1; -} - -message PhysicalPlanResponse { - repeated PlanStepMsg steps = 1; - string strategy = 2; - double total_time_ms = 3; - uint64 total_estimated_rows = 4; - repeated string notes = 5; -} - -message PlanStepMsg { - int32 step = 1; - string operation = 2; - string modality = 3; - double time_ms = 4; - uint64 estimated_rows = 5; - double selectivity = 6; - string optimization_hint = 7; -} - -message ExplainResponse { - repeated PlanStepMsg steps = 1; - repeated ModalityCostMsg cost_breakdown = 2; - repeated PerformanceHintMsg performance_hints = 3; - double total_cost_ms = 4; - string strategy = 5; - string text_output = 6; -} - -message ModalityCostMsg { - string modality = 1; - double time_ms = 2; - double percentage = 3; -} - -message PerformanceHintMsg { - string severity = 1; - string message = 2; -} - -message PlannerConfigRequest { - string global_mode = 1; - double statistics_weight = 2; - bool enable_adaptive = 3; - int32 parallel_threshold = 4; -} - -message PlannerConfigResponse { - string global_mode = 1; - double statistics_weight = 2; - bool enable_adaptive = 3; - int32 parallel_threshold = 4; -} - -message StoreStatsMsg { - string modality = 1; - uint64 total_rows = 2; - double avg_latency_ms = 3; - uint64 avg_rows_returned = 4; - uint64 query_count = 5; -} - -message StatsResponse { - repeated StoreStatsMsg stores = 1; -} - -// ============================================================================ -// Octad Messages -// ============================================================================ - -message OctadCreateRequest { - string title = 1; - string body = 2; - repeated float embedding = 3; - repeated string types = 4; -} - -message OctadIdRequest { - string id = 1; -} - -message OctadUpdateRequest { - string id = 1; - string title = 2; - string body = 3; - repeated float embedding = 4; - repeated string types = 5; -} - -message OctadResponse { - string id = 1; - string created_at = 2; - string modified_at = 3; - uint64 version = 4; - bool has_graph = 5; - bool has_vector = 6; - bool has_tensor = 7; - bool has_semantic = 8; - bool has_document = 9; - uint64 version_count = 10; -} - -message TextSearchRequest { - string query = 1; - int32 limit = 2; -} - -message VectorSearchRequest { - repeated float vector = 1; - int32 k = 2; -} - -message SearchResultMsg { - string id = 1; - float score = 2; - string title = 3; -} - -message SearchResponse { - repeated SearchResultMsg results = 1; -} diff --git a/verisimdb/rust-core/verisim-api/src/a2ml.rs b/verisimdb/rust-core/verisim-api/src/a2ml.rs deleted file mode 100644 index a8c8bbf2..00000000 --- a/verisimdb/rust-core/verisim-api/src/a2ml.rs +++ /dev/null @@ -1,360 +0,0 @@ -// SPDX-License-Identifier: MPL-2.0 -// SPDX-FileCopyrightText: 2026 Jonathan D.A. Jewell (hyperpolymath) -//! A2ML (Annotated Attribute Markup Language) response helpers for verisim-api. -//! -//! The hyperpolymath no-JSON-emit rule requires that all tool and service -//! outputs use A2ML format rather than JSON. This module provides: -//! -//! - [`a2ml_response`] — wrap an A2ML body in an Axum `Response` -//! - [`a2ml_error`] — format an error response -//! - A2ML serialisation for the proof-attempts response types -//! -//! ## A2ML format used here -//! -//! ```text -//! # SPDX-License-Identifier: MPL-2.0 -//! @(key="value", ...): -//! @(key="value"):@end -//! @end -//! ``` -//! -//! Attribute values are always quoted strings. Numeric values are -//! formatted as decimal strings. There is no `null` — absent optionals -//! are omitted from the attribute list entirely. - -use axum::{ - http::{header, StatusCode}, - response::{IntoResponse, Response}, -}; - -// ── Content-type constant ───────────────────────────────────────────────────── - -/// Content-Type for A2ML responses. -pub const A2ML_CONTENT_TYPE: &str = "text/a2ml; charset=utf-8"; - -// ── Core response helpers ───────────────────────────────────────────────────── - -/// Wrap an A2ML body string in an Axum [`Response`] with the correct -/// `Content-Type`. -pub fn a2ml_response(status: StatusCode, body: String) -> Response { - ( - status, - [(header::CONTENT_TYPE, A2ML_CONTENT_TYPE)], - body, - ) - .into_response() -} - -/// Convenience alias so callers can `use crate::a2ml::String` implicitly. -pub type String = std::string::String; - -/// Format a standardised A2ML error block. -/// -/// ```text -/// @error(code="clickhouse_error", http-status="502"):@end -/// ``` -pub fn a2ml_error(code: &str, http_status: u16) -> String { - format!("@error(code=\"{code}\", http-status=\"{http_status}\"):@end\n") -} - -/// Format an A2ML error block with an additional detail message. -/// -/// ```text -/// @error(code="serialisation_failed", http-status="400", detail="..."):@end -/// ``` -pub fn a2ml_error_detail(code: &str, http_status: u16, detail: &str) -> String { - // Escape double-quotes inside `detail` so the A2ML stays well-formed. - let escaped = detail.replace('"', "\\\""); - format!("@error(code=\"{code}\", http-status=\"{http_status}\", detail=\"{escaped}\"):@end\n") -} - -// ── proof-attempts serialisation ────────────────────────────────────────────── - -/// A single proof-attempt row as returned by the list endpoint. -pub struct ProofAttemptRowA2ml<'a> { - pub attempt_id: &'a str, - pub obligation_id: &'a str, - pub repo: &'a str, - pub file: &'a str, - pub claim: &'a str, - pub obligation_class: &'a str, - pub prover_used: &'a str, - pub outcome: &'a str, - pub duration_ms: u64, - pub confidence: f64, - pub parent_attempt_id: Option<&'a str>, - pub strategy_tag: &'a str, - pub started_at: &'a str, - pub completed_at: &'a str, -} - -impl<'a> ProofAttemptRowA2ml<'a> { - /// Render the row as an inline A2ML `@row(...):@end` element. - pub fn to_a2ml(&self, indent: &str) -> String { - let mut s = format!( - "{indent}@row(attempt-id=\"{}\", obligation-id=\"{}\", repo=\"{}\", file=\"{}\", \ - claim=\"{}\", class=\"{}\", prover=\"{}\", outcome=\"{}\", \ - duration-ms=\"{}\", confidence=\"{:.4}\", strategy-tag=\"{}\", \ - started-at=\"{}\", completed-at=\"{}\"", - self.attempt_id, - self.obligation_id, - self.repo, - self.file, - self.claim, - self.obligation_class, - self.prover_used, - self.outcome, - self.duration_ms, - self.confidence, - self.strategy_tag, - self.started_at, - self.completed_at, - ); - if let Some(parent) = self.parent_attempt_id { - s.push_str(&format!(", parent-attempt-id=\"{parent}\"")); - } - s.push_str("):@end\n"); - s - } -} - -/// Serialise a list of raw ClickHouse JSONEachRow rows (as `serde_json::Value`) -/// into a complete A2ML `@proof-attempts` block. -/// -/// Rows that are missing required fields are silently skipped. -pub fn proof_attempts_to_a2ml(rows: &[serde_json::Value]) -> String { - let mut out = String::new(); - out.push_str("@proof-attempts():\n"); - for v in rows { - let Some(attempt_id) = v["attempt_id"].as_str() else { continue }; - let Some(obligation_id) = v["obligation_id"].as_str() else { continue }; - let Some(repo) = v["repo"].as_str() else { continue }; - let Some(file) = v["file"].as_str() else { continue }; - let Some(claim) = v["claim"].as_str() else { continue }; - let Some(obligation_class) = v["obligation_class"].as_str() else { continue }; - let Some(prover_used) = v["prover_used"].as_str() else { continue }; - let Some(outcome) = v["outcome"].as_str() else { continue }; - let duration_ms = v["duration_ms"].as_u64().unwrap_or(0); - let confidence = v["confidence"].as_f64().unwrap_or(0.0); - let parent_attempt_id = v["parent_attempt_id"].as_str(); - let strategy_tag = v["strategy_tag"].as_str().unwrap_or(""); - let started_at = v["started_at"].as_str().unwrap_or(""); - let completed_at = v["completed_at"].as_str().unwrap_or(""); - - let row = ProofAttemptRowA2ml { - attempt_id, - obligation_id, - repo, - file, - claim, - obligation_class, - prover_used, - outcome, - duration_ms, - confidence, - parent_attempt_id, - strategy_tag, - started_at, - completed_at, - }; - out.push_str(&row.to_a2ml(" ")); - } - out.push_str("@end\n"); - out -} - -/// Serialise the `inserted` acknowledgement for a single proof attempt. -/// -/// ```text -/// @inserted(status="ok", attempt-id="abc123"):@end -/// ``` -pub fn inserted_to_a2ml(attempt_id: &str) -> String { - format!("@inserted(status=\"ok\", attempt-id=\"{attempt_id}\"):@end\n") -} - -// ── strategy serialisation ──────────────────────────────────────────────────── - -/// A single prover recommendation. -pub struct RecommendationA2ml { - pub prover: String, - pub success_rate: f64, - pub avg_duration_ms: f64, - pub total_attempts: u64, -} - -/// Serialise strategy recommendations into an A2ML block. -/// -/// ```text -/// @strategy-recommendations(): -/// @recommendation(prover="echidna", success-rate="0.9500", ...):@end -/// @end -/// ``` -pub fn strategy_to_a2ml(recommendations: &[RecommendationA2ml]) -> String { - let mut out = String::from("@strategy-recommendations():\n"); - for r in recommendations { - out.push_str(&format!( - " @recommendation(prover=\"{}\", success-rate=\"{:.4}\", \ - avg-duration-ms=\"{:.2}\", total-attempts=\"{}\"):@end\n", - r.prover, r.success_rate, r.avg_duration_ms, r.total_attempts, - )); - } - out.push_str("@end\n"); - out -} - -/// Parse ClickHouse JSONEachRow recommendations into typed structs. -/// Malformed lines are silently skipped. -pub fn parse_recommendations(text: &str) -> Vec { - text.lines() - .filter(|l| !l.trim().is_empty()) - .filter_map(|line| { - let v: serde_json::Value = serde_json::from_str(line).ok()?; - Some(RecommendationA2ml { - prover: v["prover_used"].as_str()?.to_string(), - success_rate: v["success_rate"].as_f64().unwrap_or(0.0), - avg_duration_ms: v["avg_duration_ms"].as_f64().unwrap_or(0.0), - total_attempts: v["total_attempts"].as_u64().unwrap_or(0), - }) - }) - .collect() -} - -// ── certificates serialisation ──────────────────────────────────────────────── - -/// A single certificate row. -pub struct CertRowA2ml { - pub prover_used: String, - pub status: String, - pub success_rate: f64, - pub total_attempts: u64, -} - -/// Serialise certificate rows into an A2ML block. -/// -/// ```text -/// @certificates(): -/// @cert(prover-used="echidna", status="PROVEN", ...):@end -/// @end -/// ``` -pub fn certificates_to_a2ml(rows: &[CertRowA2ml]) -> String { - let mut out = String::from("@certificates():\n"); - for r in rows { - out.push_str(&format!( - " @cert(prover-used=\"{}\", status=\"{}\", \ - success-rate=\"{:.4}\", total-attempts=\"{}\"):@end\n", - r.prover_used, r.status, r.success_rate, r.total_attempts, - )); - } - out.push_str("@end\n"); - out -} - -/// Parse ClickHouse JSONEachRow certificate rows into typed structs. -/// Malformed lines are silently skipped. -pub fn parse_certs(text: &str) -> Vec { - text.lines() - .filter(|l| !l.trim().is_empty()) - .filter_map(|line| { - let v: serde_json::Value = serde_json::from_str(line).ok()?; - Some(CertRowA2ml { - prover_used: v["prover_used"].as_str()?.to_string(), - status: v["status"].as_str().unwrap_or("pending").to_string(), - success_rate: v["success_rate"].as_f64().unwrap_or(0.0), - total_attempts: v["total_attempts"].as_u64().unwrap_or(0), - }) - }) - .collect() -} - -// ── tests ───────────────────────────────────────────────────────────────────── - -#[cfg(test)] -mod tests { - use super::*; - - #[test] - fn error_format() { - let s = a2ml_error("clickhouse_error", 502); - assert!(s.contains("@error(")); - assert!(s.contains("code=\"clickhouse_error\"")); - assert!(s.contains("http-status=\"502\"")); - assert!(s.ends_with(":@end\n")); - } - - #[test] - fn error_detail_escapes_quotes() { - let s = a2ml_error_detail("bad_input", 400, r#"has "quotes""#); - assert!(s.contains(r#"detail="has \"quotes\"""#)); - } - - #[test] - fn inserted_format() { - let s = inserted_to_a2ml("abc-123"); - assert_eq!(s, "@inserted(status=\"ok\", attempt-id=\"abc-123\"):@end\n"); - } - - #[test] - fn proof_attempts_to_a2ml_empty() { - let s = proof_attempts_to_a2ml(&[]); - assert_eq!(s, "@proof-attempts():\n@end\n"); - } - - #[test] - fn proof_attempts_to_a2ml_skips_malformed() { - // Row missing 'attempt_id' field should be skipped - let rows = vec![serde_json::json!({"obligation_id": "x"})]; - let s = proof_attempts_to_a2ml(&rows); - // Only the wrapper tags, no @row - assert!(!s.contains("@row(")); - } - - #[test] - fn strategy_to_a2ml_format() { - let recs = vec![ - RecommendationA2ml { - prover: "echidna".to_string(), - success_rate: 0.95, - avg_duration_ms: 120.5, - total_attempts: 100, - }, - ]; - let s = strategy_to_a2ml(&recs); - assert!(s.starts_with("@strategy-recommendations():\n")); - assert!(s.contains("prover=\"echidna\"")); - assert!(s.contains("success-rate=\"0.9500\"")); - assert!(s.contains("avg-duration-ms=\"120.50\"")); - assert!(s.contains("total-attempts=\"100\"")); - assert!(s.ends_with("@end\n")); - } - - #[test] - fn certificates_to_a2ml_format() { - let rows = vec![CertRowA2ml { - prover_used: "idris2".to_string(), - status: "PROVEN".to_string(), - success_rate: 0.80, - total_attempts: 50, - }]; - let s = certificates_to_a2ml(&rows); - assert!(s.contains("prover-used=\"idris2\"")); - assert!(s.contains("status=\"PROVEN\"")); - assert!(s.contains("success-rate=\"0.8000\"")); - } - - #[test] - fn parse_recommendations_skips_malformed() { - let text = "{\"prover_used\":\"echidna\",\"success_rate\":0.9,\"avg_duration_ms\":100.0,\"total_attempts\":10}\n{bad json}\n"; - let recs = parse_recommendations(text); - assert_eq!(recs.len(), 1); - assert_eq!(recs[0].prover, "echidna"); - } - - #[test] - fn parse_certs_skips_malformed() { - let text = "{\"prover_used\":\"lean4\",\"status\":\"PROVEN\",\"success_rate\":0.75,\"total_attempts\":8}\n"; - let certs = parse_certs(text); - assert_eq!(certs.len(), 1); - assert_eq!(certs[0].prover_used, "lean4"); - assert_eq!(certs[0].status, "PROVEN"); - } -} diff --git a/verisimdb/rust-core/verisim-api/src/auth.rs b/verisimdb/rust-core/verisim-api/src/auth.rs deleted file mode 100644 index 6e134850..00000000 --- a/verisimdb/rust-core/verisim-api/src/auth.rs +++ /dev/null @@ -1,778 +0,0 @@ -// SPDX-License-Identifier: MPL-2.0 -// Copyright (c) 2026 Jonathan D.A. Jewell (hyperpolymath) -// -//! Authentication and rate limiting for VeriSimDB API. -//! -//! Supports two authentication methods: -//! - **API Key**: Passed via `X-API-Key` header -//! - **JWT Bearer Token**: Passed via `Authorization: Bearer ` header -//! -//! Rate limiting is per-client (identified by API key or IP address). - -use axum::{ - extract::{Request, State}, - http::{header, StatusCode}, - middleware::Next, - response::{IntoResponse, Response}, - Json, -}; -use serde::{Deserialize, Serialize}; -use sha2::{Digest, Sha256}; -use std::collections::HashMap; -use std::sync::{Arc, Mutex}; -use std::time::{Duration, Instant}; -use tracing::{info, warn}; - -/// Authentication configuration. -#[derive(Debug, Clone, Serialize, Deserialize)] -pub struct AuthConfig { - /// Whether authentication is enabled. When disabled, all requests pass through. - pub enabled: bool, - /// Whether to allow unauthenticated access to health/ready/metrics endpoints. - pub allow_public_health: bool, - /// Maximum requests per minute per client (0 = unlimited). - pub rate_limit_per_minute: u32, - /// JWT secret for HMAC-SHA256 verification (if using JWT). - #[serde(skip_serializing)] - pub jwt_secret: Option, -} - -impl Default for AuthConfig { - fn default() -> Self { - Self { - enabled: false, - allow_public_health: true, - rate_limit_per_minute: 0, - jwt_secret: None, - } - } -} - -/// Client identity extracted from an authenticated request. -#[derive(Debug, Clone)] -pub struct ClientIdentity { - /// The client identifier (API key hash or JWT subject). - pub id: String, - /// Role assigned to this client. - pub role: ClientRole, -} - -/// Role-based access level. -#[derive(Debug, Clone, Copy, PartialEq, Eq, Serialize, Deserialize)] -pub enum ClientRole { - /// Read-only access to all endpoints. - Reader, - /// Read and write access to octad CRUD and queries. - Writer, - /// Full access including admin operations (config, normalizer triggers). - Admin, -} - -/// Registered API key with associated metadata. -#[derive(Debug, Clone, Serialize, Deserialize)] -pub struct ApiKeyEntry { - /// SHA-256 hash of the API key (never store plaintext). - pub key_hash: String, - /// Human-readable label for this key. - pub label: String, - /// Role granted to holders of this key. - pub role: ClientRole, - /// Whether this key is currently active. - pub active: bool, -} - -/// In-memory API key registry. -/// -/// In production, this would be backed by a persistent store. For now, -/// keys are registered at startup or via admin endpoints. -#[derive(Debug, Clone)] -pub struct ApiKeyRegistry { - /// Map from key hash → entry. - keys: Arc>>, -} - -impl ApiKeyRegistry { - /// Create a new empty registry. - pub fn new() -> Self { - Self { - keys: Arc::new(Mutex::new(HashMap::new())), - } - } - - /// Register a new API key. The key is hashed before storage. - pub fn register(&self, plaintext_key: &str, label: &str, role: ClientRole) { - let hash = hash_key(plaintext_key); - let entry = ApiKeyEntry { - key_hash: hash.clone(), - label: label.to_string(), - role, - active: true, - }; - let mut keys = self.keys.lock().expect("key registry lock"); - keys.insert(hash, entry); - } - - /// Validate an API key. Returns the entry if the key is valid and active. - pub fn validate(&self, plaintext_key: &str) -> Option { - let hash = hash_key(plaintext_key); - let keys = self.keys.lock().expect("key registry lock"); - keys.get(&hash) - .filter(|entry| entry.active) - .cloned() - } - - /// Revoke an API key by its hash. - pub fn revoke(&self, key_hash: &str) -> bool { - let mut keys = self.keys.lock().expect("key registry lock"); - if let Some(entry) = keys.get_mut(key_hash) { - entry.active = false; - true - } else { - false - } - } - - /// List all registered keys (without plaintext). - pub fn list(&self) -> Vec { - let keys = self.keys.lock().expect("key registry lock"); - keys.values().cloned().collect() - } -} - -impl Default for ApiKeyRegistry { - fn default() -> Self { - Self::new() - } -} - -/// Per-client rate limiter using a sliding window counter. -#[derive(Debug, Clone)] -pub struct RateLimiter { - /// Map from client ID → (request timestamps in current window). - windows: Arc>>>, - /// Maximum requests per window. - max_requests: u32, - /// Window duration. - window: Duration, -} - -impl RateLimiter { - /// Create a new rate limiter. - pub fn new(max_requests_per_minute: u32) -> Self { - Self { - windows: Arc::new(Mutex::new(HashMap::new())), - max_requests: max_requests_per_minute, - window: Duration::from_secs(60), - } - } - - /// Check if a client is allowed to make a request. Returns true if allowed. - pub fn check(&self, client_id: &str) -> bool { - if self.max_requests == 0 { - return true; // Unlimited. - } - - let now = Instant::now(); - let mut windows = self.windows.lock().expect("rate limiter lock"); - let timestamps = windows.entry(client_id.to_string()).or_default(); - - // Remove expired timestamps. - timestamps.retain(|t| now.duration_since(*t) < self.window); - - if timestamps.len() as u32 >= self.max_requests { - false - } else { - timestamps.push(now); - true - } - } - - /// Get the number of remaining requests for a client in the current window. - pub fn remaining(&self, client_id: &str) -> u32 { - if self.max_requests == 0 { - return u32::MAX; - } - - let now = Instant::now(); - let mut windows = self.windows.lock().expect("rate limiter lock"); - let timestamps = windows.entry(client_id.to_string()).or_default(); - timestamps.retain(|t| now.duration_since(*t) < self.window); - - self.max_requests.saturating_sub(timestamps.len() as u32) - } -} - -/// Shared authentication state, added to `AppState`. -#[derive(Debug, Clone)] -pub struct AuthState { - pub config: AuthConfig, - pub key_registry: ApiKeyRegistry, - pub rate_limiter: RateLimiter, - /// RBAC state for fine-grained authorization checks. - pub rbac: crate::rbac::RbacState, -} - -impl AuthState { - /// Create auth state from config. - pub fn new(config: AuthConfig) -> Self { - let rate_limiter = RateLimiter::new(config.rate_limit_per_minute); - Self { - config, - key_registry: ApiKeyRegistry::new(), - rate_limiter, - rbac: crate::rbac::RbacState::default(), - } - } - - /// Create auth state from config with a custom RBAC policy. - pub fn with_rbac(config: AuthConfig, rbac: crate::rbac::RbacState) -> Self { - let rate_limiter = RateLimiter::new(config.rate_limit_per_minute); - Self { - config, - key_registry: ApiKeyRegistry::new(), - rate_limiter, - rbac, - } - } -} - -impl Default for AuthState { - fn default() -> Self { - Self::new(AuthConfig::default()) - } -} - -/// Authentication error response. -#[derive(Debug, Serialize)] -pub struct AuthError { - pub error: String, - pub code: u16, -} - -/// Axum middleware that performs authentication and rate limiting. -/// -/// This middleware: -/// 1. Checks if auth is enabled (passes through if disabled) -/// 2. Allows public health endpoints if configured -/// 3. Extracts API key from `X-API-Key` header or JWT from `Authorization: Bearer` -/// 4. Validates the credential against the key registry -/// 5. Checks rate limits for the identified client -pub async fn auth_middleware( - State(auth): State, - request: Request, - next: Next, -) -> Response { - // If auth is disabled, pass through. - if !auth.config.enabled { - return next.run(request).await; - } - - let path = request.uri().path().to_string(); - - // Allow public health endpoints without auth. - if auth.config.allow_public_health - && (path == "/health" || path == "/ready" || path == "/metrics") - { - return next.run(request).await; - } - - // Extract credential. - let identity = match extract_identity(&request, &auth) { - Ok(id) => id, - Err(response) => return response, - }; - - // Rate limit check. - if !auth.rate_limiter.check(&identity.id) { - warn!(client = %identity.id, "Rate limit exceeded"); - let remaining = auth.rate_limiter.remaining(&identity.id); - return ( - StatusCode::TOO_MANY_REQUESTS, - [ - (header::HeaderName::from_static("x-ratelimit-remaining"), - remaining.to_string()), - (header::RETRY_AFTER, "60".to_string()), - ], - Json(AuthError { - error: "Rate limit exceeded".to_string(), - code: 429, - }), - ) - .into_response(); - } - - // RBAC authorization check. - let method = request.method().clone(); - if let Err(authz_err) = crate::rbac::check_authorization( - &identity, - &path, - &method, - &auth.rbac, - ) { - warn!( - client = %identity.id, - role = ?identity.role, - path = %path, - "Authorization denied: {}", - authz_err.error - ); - return authz_err.into_response(); - } - - next.run(request).await -} - -/// Extract client identity from request headers. -fn extract_identity(request: &Request, auth: &AuthState) -> Result { - // Try X-API-Key header first. - if let Some(api_key) = request - .headers() - .get("x-api-key") - .and_then(|v| v.to_str().ok()) - { - if let Some(entry) = auth.key_registry.validate(api_key) { - info!(label = %entry.label, role = ?entry.role, "API key authenticated"); - return Ok(ClientIdentity { - id: entry.key_hash.clone(), - role: entry.role, - }); - } - return Err(( - StatusCode::UNAUTHORIZED, - Json(AuthError { - error: "Invalid API key".to_string(), - code: 401, - }), - ) - .into_response()); - } - - // Try Authorization: Bearer header. - if let Some(auth_header) = request - .headers() - .get(header::AUTHORIZATION) - .and_then(|v| v.to_str().ok()) - { - if let Some(token) = auth_header.strip_prefix("Bearer ") { - match validate_jwt(token, &auth.config) { - Ok(identity) => { - info!(subject = %identity.id, role = ?identity.role, "JWT authenticated"); - return Ok(identity); - } - Err(msg) => { - return Err(( - StatusCode::UNAUTHORIZED, - Json(AuthError { - error: msg, - code: 401, - }), - ) - .into_response()); - } - } - } - } - - // No credentials provided. - Err(( - StatusCode::UNAUTHORIZED, - [(header::WWW_AUTHENTICATE, "Bearer, ApiKey")], - Json(AuthError { - error: "Authentication required. Provide X-API-Key header or Authorization: Bearer ".to_string(), - code: 401, - }), - ) - .into_response()) -} - -/// Validate a JWT token (HMAC-SHA256). -/// -/// VeriSimDB uses a minimal JWT implementation: we only verify the signature -/// and extract the `sub` (subject) and `role` claims. Expiration is checked -/// via the `exp` claim. -fn validate_jwt(token: &str, config: &AuthConfig) -> Result { - let secret = config - .jwt_secret - .as_deref() - .ok_or_else(|| "JWT authentication not configured".to_string())?; - - let parts: Vec<&str> = token.split('.').collect(); - if parts.len() != 3 { - return Err("Invalid JWT format".to_string()); - } - - // Decode the header and payload. - let payload_bytes = base64url_decode(parts[1]) - .map_err(|_| "Invalid JWT payload encoding".to_string())?; - - // Verify HMAC-SHA256 signature. - let signing_input = format!("{}.{}", parts[0], parts[1]); - let expected_sig = hmac_sha256(signing_input.as_bytes(), secret.as_bytes()); - let actual_sig = base64url_decode(parts[2]) - .map_err(|_| "Invalid JWT signature encoding".to_string())?; - - if expected_sig != actual_sig { - return Err("Invalid JWT signature".to_string()); - } - - // Parse claims. - let claims: serde_json::Value = serde_json::from_slice(&payload_bytes) - .map_err(|_| "Invalid JWT payload".to_string())?; - - // Check expiration. - if let Some(exp) = claims.get("exp").and_then(|v| v.as_i64()) { - let now = std::time::SystemTime::now() - .duration_since(std::time::UNIX_EPOCH) - .expect("TODO: handle error") - .as_secs() as i64; - if now > exp { - return Err("JWT token expired".to_string()); - } - } - - let subject = claims - .get("sub") - .and_then(|v| v.as_str()) - .unwrap_or("unknown") - .to_string(); - - let role = claims - .get("role") - .and_then(|v| v.as_str()) - .map(|r| match r { - "admin" => ClientRole::Admin, - "writer" => ClientRole::Writer, - _ => ClientRole::Reader, - }) - .unwrap_or(ClientRole::Reader); - - Ok(ClientIdentity { - id: subject, - role, - }) -} - -/// Hash an API key with SHA-256 for storage. -fn hash_key(key: &str) -> String { - let mut hasher = Sha256::new(); - hasher.update(key.as_bytes()); - let hash = hasher.finalize(); - hex::encode(hash) -} - -/// Compute HMAC-SHA256. -fn hmac_sha256(data: &[u8], key: &[u8]) -> Vec { - // HMAC: H((key XOR opad) || H((key XOR ipad) || message)) - let block_size = 64; - let mut key_block = vec![0u8; block_size]; - - if key.len() > block_size { - let mut hasher = Sha256::new(); - hasher.update(key); - let hashed = hasher.finalize(); - key_block[..hashed.len()].copy_from_slice(&hashed); - } else { - key_block[..key.len()].copy_from_slice(key); - } - - let mut ipad = vec![0x36u8; block_size]; - let mut opad = vec![0x5cu8; block_size]; - for i in 0..block_size { - ipad[i] ^= key_block[i]; - opad[i] ^= key_block[i]; - } - - // Inner hash. - let mut inner_hasher = Sha256::new(); - inner_hasher.update(&ipad); - inner_hasher.update(data); - let inner_hash = inner_hasher.finalize(); - - // Outer hash. - let mut outer_hasher = Sha256::new(); - outer_hasher.update(&opad); - outer_hasher.update(&inner_hash); - outer_hasher.finalize().to_vec() -} - -/// Base64url decode (RFC 4648 without padding). -fn base64url_decode(input: &str) -> Result, &'static str> { - // Add padding if needed. - let padded = match input.len() % 4 { - 2 => format!("{input}=="), - 3 => format!("{input}="), - 0 => input.to_string(), - _ => return Err("invalid base64url length"), - }; - - // Replace URL-safe characters with standard base64. - let standard = padded.replace('-', "+").replace('_', "/"); - - // Decode using a simple base64 decoder. - base64_decode(&standard) -} - -/// Simple base64 decoder. -fn base64_decode(input: &str) -> Result, &'static str> { - const CHARSET: &[u8] = b"ABCDEFGHIJKLMNOPQRSTUVWXYZabcdefghijklmnopqrstuvwxyz0123456789+/"; - - fn char_to_val(c: u8) -> Result { - if c == b'=' { - return Ok(0); - } - CHARSET - .iter() - .position(|&x| x == c) - .map(|p| p as u8) - .ok_or("invalid base64 character") - } - - let bytes = input.as_bytes(); - if bytes.len() % 4 != 0 { - return Err("invalid base64 length"); - } - - let mut result = Vec::with_capacity(bytes.len() * 3 / 4); - - for chunk in bytes.chunks(4) { - let a = char_to_val(chunk[0])?; - let b = char_to_val(chunk[1])?; - let c = char_to_val(chunk[2])?; - let d = char_to_val(chunk[3])?; - - result.push((a << 2) | (b >> 4)); - if chunk[2] != b'=' { - result.push((b << 4) | (c >> 2)); - } - if chunk[3] != b'=' { - result.push((c << 6) | d); - } - } - - Ok(result) -} - -#[cfg(test)] -mod tests { - use super::*; - - #[test] - fn test_hash_key_deterministic() { - let hash1 = hash_key("test-key-123"); - let hash2 = hash_key("test-key-123"); - assert_eq!(hash1, hash2); - } - - #[test] - fn test_hash_key_different_keys() { - let hash1 = hash_key("key-a"); - let hash2 = hash_key("key-b"); - assert_ne!(hash1, hash2); - } - - #[test] - fn test_api_key_registry_register_and_validate() { - let registry = ApiKeyRegistry::new(); - registry.register("my-secret-key", "Test Key", ClientRole::Writer); - - let entry = registry.validate("my-secret-key"); - assert!(entry.is_some()); - let entry = entry.expect("TODO: handle error"); - assert_eq!(entry.label, "Test Key"); - assert_eq!(entry.role, ClientRole::Writer); - assert!(entry.active); - } - - #[test] - fn test_api_key_registry_invalid_key() { - let registry = ApiKeyRegistry::new(); - registry.register("correct-key", "Valid", ClientRole::Reader); - - let result = registry.validate("wrong-key"); - assert!(result.is_none()); - } - - #[test] - fn test_api_key_registry_revoke() { - let registry = ApiKeyRegistry::new(); - registry.register("revokable-key", "Temp", ClientRole::Admin); - - let hash = hash_key("revokable-key"); - assert!(registry.revoke(&hash)); - - let result = registry.validate("revokable-key"); - assert!(result.is_none()); - } - - #[test] - fn test_rate_limiter_allows_within_limit() { - let limiter = RateLimiter::new(10); - for _ in 0..10 { - assert!(limiter.check("client-1")); - } - } - - #[test] - fn test_rate_limiter_blocks_over_limit() { - let limiter = RateLimiter::new(3); - assert!(limiter.check("client-x")); - assert!(limiter.check("client-x")); - assert!(limiter.check("client-x")); - assert!(!limiter.check("client-x")); // 4th should be blocked - } - - #[test] - fn test_rate_limiter_unlimited() { - let limiter = RateLimiter::new(0); - for _ in 0..1000 { - assert!(limiter.check("anyone")); - } - } - - #[test] - fn test_rate_limiter_per_client() { - let limiter = RateLimiter::new(2); - assert!(limiter.check("alice")); - assert!(limiter.check("alice")); - assert!(!limiter.check("alice")); // Alice blocked - - // Bob should still have quota - assert!(limiter.check("bob")); - assert!(limiter.check("bob")); - } - - #[test] - fn test_rate_limiter_remaining() { - let limiter = RateLimiter::new(5); - assert_eq!(limiter.remaining("test"), 5); - limiter.check("test"); - assert_eq!(limiter.remaining("test"), 4); - } - - #[test] - fn test_hmac_sha256_known_vector() { - // RFC 4231 Test Case 2 - let key = b"Jefe"; - let data = b"what do ya want for nothing?"; - let mac = hmac_sha256(data, key); - let hex = hex::encode(&mac); - assert_eq!( - hex, - "5bdcc146bf60754e6a042426089575c75a003f089d2739839dec58b964ec3843" - ); - } - - #[test] - fn test_base64url_decode() { - // "hello" in base64url is "aGVsbG8" - let decoded = base64url_decode("aGVsbG8").expect("TODO: handle error"); - assert_eq!(decoded, b"hello"); - } - - #[test] - fn test_auth_config_default() { - let config = AuthConfig::default(); - assert!(!config.enabled); - assert!(config.allow_public_health); - assert_eq!(config.rate_limit_per_minute, 0); - } - - #[test] - fn test_jwt_validation() { - // Create a minimal JWT for testing. - let config = AuthConfig { - enabled: true, - allow_public_health: true, - rate_limit_per_minute: 0, - jwt_secret: Some("test-secret".to_string()), - }; - - // Create JWT: header.payload.signature - let header = base64url_encode(b"{\"alg\":\"HS256\",\"typ\":\"JWT\"}"); - // Set exp far in the future. - let payload = base64url_encode( - b"{\"sub\":\"test-user\",\"role\":\"admin\",\"exp\":9999999999}", - ); - let signing_input = format!("{header}.{payload}"); - let sig = hmac_sha256(signing_input.as_bytes(), b"test-secret"); - let sig_encoded = base64url_encode(&sig); - let token = format!("{header}.{payload}.{sig_encoded}"); - - let result = validate_jwt(&token, &config); - assert!(result.is_ok()); - let identity = result.expect("TODO: handle error"); - assert_eq!(identity.id, "test-user"); - assert_eq!(identity.role, ClientRole::Admin); - } - - #[test] - fn test_jwt_expired() { - let config = AuthConfig { - enabled: true, - allow_public_health: true, - rate_limit_per_minute: 0, - jwt_secret: Some("test-secret".to_string()), - }; - - let header = base64url_encode(b"{\"alg\":\"HS256\",\"typ\":\"JWT\"}"); - let payload = base64url_encode(b"{\"sub\":\"expired-user\",\"exp\":1}"); - let signing_input = format!("{header}.{payload}"); - let sig = hmac_sha256(signing_input.as_bytes(), b"test-secret"); - let sig_encoded = base64url_encode(&sig); - let token = format!("{header}.{payload}.{sig_encoded}"); - - let result = validate_jwt(&token, &config); - assert!(result.is_err()); - assert!(result.unwrap_err().contains("expired")); - } - - #[test] - fn test_jwt_invalid_signature() { - let config = AuthConfig { - enabled: true, - allow_public_health: true, - rate_limit_per_minute: 0, - jwt_secret: Some("correct-secret".to_string()), - }; - - let header = base64url_encode(b"{\"alg\":\"HS256\",\"typ\":\"JWT\"}"); - let payload = base64url_encode(b"{\"sub\":\"hacker\",\"exp\":9999999999}"); - let signing_input = format!("{header}.{payload}"); - // Sign with WRONG secret. - let sig = hmac_sha256(signing_input.as_bytes(), b"wrong-secret"); - let sig_encoded = base64url_encode(&sig); - let token = format!("{header}.{payload}.{sig_encoded}"); - - let result = validate_jwt(&token, &config); - assert!(result.is_err()); - assert!(result.unwrap_err().contains("signature")); - } - - /// Base64url encode for test helpers. - fn base64url_encode(input: &[u8]) -> String { - const CHARSET: &[u8] = - b"ABCDEFGHIJKLMNOPQRSTUVWXYZabcdefghijklmnopqrstuvwxyz0123456789+/"; - - let mut result = String::new(); - let mut i = 0; - while i < input.len() { - let a = input[i]; - let b = if i + 1 < input.len() { input[i + 1] } else { 0 }; - let c = if i + 2 < input.len() { input[i + 2] } else { 0 }; - - result.push(CHARSET[(a >> 2) as usize] as char); - result.push(CHARSET[((a & 0x03) << 4 | b >> 4) as usize] as char); - - if i + 1 < input.len() { - result.push(CHARSET[((b & 0x0f) << 2 | c >> 6) as usize] as char); - } - if i + 2 < input.len() { - result.push(CHARSET[(c & 0x3f) as usize] as char); - } - i += 3; - } - - // Convert to URL-safe base64 (no padding). - result.replace('+', "-").replace('/', "_") - } -} diff --git a/verisimdb/rust-core/verisim-api/src/federation.rs b/verisimdb/rust-core/verisim-api/src/federation.rs deleted file mode 100644 index d252ee9a..00000000 --- a/verisimdb/rust-core/verisim-api/src/federation.rs +++ /dev/null @@ -1,660 +0,0 @@ -// SPDX-License-Identifier: MPL-2.0 -//! Federation Protocol for VeriSimDB -//! -//! Enables cross-instance querying with drift-aware consistency policies. -//! Each VeriSimDB instance can federate with others to form a distributed -//! knowledge network while maintaining local autonomy. -//! -//! ## Protocol -//! -//! Federation uses a pull-based HTTP protocol where the coordinator -//! (the instance receiving the query) fans out requests to peer stores -//! and aggregates results according to the configured drift policy. -//! -//! ## Drift Policies -//! -//! - **Strict**: Only return results from stores with drift score < threshold. -//! - **Repair**: Return results and trigger normalization on drifted stores. -//! - **Tolerate**: Return all results, annotating drifted ones. -//! - **Latest**: Return only the most recent version from each store. - -use axum::{ - extract::{Query, State}, - http::{HeaderMap, StatusCode}, - routing::{get, post}, - Json, Router, -}; -use serde::{Deserialize, Serialize}; -use sha2::{Sha256, Digest}; -use std::collections::HashMap; -use std::sync::{Arc, RwLock}; -use tracing::{error, info, warn, instrument}; - -// --------------------------------------------------------------------------- -// Types -// --------------------------------------------------------------------------- - -/// Drift policy for federated queries. -#[derive(Debug, Clone, Copy, Serialize, Deserialize, PartialEq, Eq)] -#[serde(rename_all = "lowercase")] -pub enum DriftPolicy { - Strict, - Repair, - Tolerate, - Latest, -} - -impl Default for DriftPolicy { - fn default() -> Self { - DriftPolicy::Tolerate - } -} - -/// A registered peer store in the federation. -#[derive(Debug, Clone, Serialize, Deserialize)] -pub struct PeerStore { - /// Unique store identifier. - pub store_id: String, - /// HTTP endpoint URL (e.g., "https://store-2.verisimdb.example.com/api/v1"). - pub endpoint: String, - /// Modalities this store supports. - pub modalities: Vec, - /// Trust level (0.0 - 1.0). - pub trust_level: f64, - /// Last health check timestamp (RFC 3339). - pub last_seen: Option, - /// Average response time in milliseconds. - pub response_time_ms: Option, - /// SHA-256 hash of the peer's secret (not serialized to clients). - #[serde(skip)] - pub secret_hash: Option, -} - -/// A federation query request. -#[derive(Debug, Serialize, Deserialize)] -pub struct FederationQueryRequest { - /// Pattern to match stores (e.g., "/universities/*", or a specific store ID). - pub pattern: String, - /// Modalities to query. - pub modalities: Vec, - /// Drift policy. - #[serde(default)] - pub drift_policy: DriftPolicy, - /// Maximum results per store. - pub limit: Option, - /// Optional text query. - pub text_query: Option, - /// Optional vector query. - pub vector_query: Option>, -} - -/// A single result from a federated query. -#[derive(Debug, Serialize, Deserialize)] -pub struct FederationResult { - /// Which store provided this result. - pub source_store: String, - /// Octad ID. - pub octad_id: String, - /// Relevance score. - pub score: f64, - /// Whether the source store has drift issues. - pub drifted: bool, - /// Result data (modality-dependent). - pub data: serde_json::Value, -} - -/// Response for federation queries. -#[derive(Debug, Serialize, Deserialize)] -pub struct FederationQueryResponse { - /// Query results aggregated from all matching stores. - pub results: Vec, - /// Stores that were queried. - pub stores_queried: Vec, - /// Stores that failed or were excluded by drift policy. - pub stores_excluded: Vec, - /// The drift policy applied. - pub drift_policy: DriftPolicy, -} - -/// Registration request to join the federation. -#[derive(Debug, Serialize, Deserialize)] -pub struct RegisterRequest { - pub store_id: String, - pub endpoint: String, - pub modalities: Vec, - /// Pre-shared key for this peer (sent once at registration, stored as SHA-256 hash). - pub secret: Option, -} - -/// Health check from a peer. -#[derive(Debug, Serialize, Deserialize)] -pub struct HeartbeatRequest { - pub store_id: String, - pub drift_scores: HashMap, -} - -/// Federation registry query parameters. -#[derive(Debug, Deserialize)] -pub struct PeerQueryParams { - pub modality: Option, -} - -// --------------------------------------------------------------------------- -// Federation State -// --------------------------------------------------------------------------- - -/// Shared federation state — registry of known peer stores. -#[derive(Clone)] -pub struct FederationState { - pub peers: Arc>>, - pub self_store_id: String, - pub self_endpoint: String, - /// Drift threshold for Strict policy. - pub strict_drift_threshold: f64, - /// Pre-shared keys for federation peers (store_id → key). - /// Parsed from `VERISIM_FEDERATION_KEYS` env var (comma-separated `store_id:key`). - /// When empty, federation registration is disabled (closed by default). - federation_keys: Arc>, -} - -impl FederationState { - pub fn new(self_store_id: String, self_endpoint: String) -> Self { - let keys = Self::load_federation_keys(); - Self { - peers: Arc::new(RwLock::new(HashMap::new())), - self_store_id, - self_endpoint, - strict_drift_threshold: 0.3, - federation_keys: Arc::new(keys), - } - } - - /// Load federation keys from `VERISIM_FEDERATION_KEYS` env var. - /// Format: `store_id1:key1,store_id2:key2,...` - fn load_federation_keys() -> HashMap { - match std::env::var("VERISIM_FEDERATION_KEYS") { - Ok(val) if !val.is_empty() => { - val.split(',') - .filter_map(|pair| { - let parts: Vec<&str> = pair.splitn(2, ':').collect(); - if parts.len() == 2 { - Some((parts[0].trim().to_string(), parts[1].trim().to_string())) - } else { - warn!(entry = %pair, "Invalid federation key entry (expected store_id:key)"); - None - } - }) - .collect() - } - _ => HashMap::new(), - } - } - - /// Check if federation registration is enabled (requires keys to be configured). - pub fn registration_enabled(&self) -> bool { - !self.federation_keys.is_empty() - } - - /// Validate a PSK for a given store_id. - fn validate_psk(&self, store_id: &str, provided_secret: &str) -> bool { - if let Some(expected_key) = self.federation_keys.get(store_id) { - expected_key == provided_secret - } else { - false - } - } - - /// Validate the X-Federation-PSK header for an existing peer. - fn validate_peer_header(&self, store_id: &str, headers: &HeaderMap) -> Result<(), StatusCode> { - let psk = headers - .get("X-Federation-PSK") - .and_then(|v| v.to_str().ok()) - .unwrap_or(""); - - if psk.is_empty() { - return Err(StatusCode::UNAUTHORIZED); - } - - // Check against stored hash - let peers = self.peers.read().map_err(|_| { - error!("Federation peers RwLock poisoned"); - StatusCode::INTERNAL_SERVER_ERROR - })?; - - if let Some(peer) = peers.get(store_id) { - if let Some(ref stored_hash) = peer.secret_hash { - let provided_hash = sha256_hex(psk); - if stored_hash == &provided_hash { - return Ok(()); - } - } - } - - Err(StatusCode::UNAUTHORIZED) - } -} - -/// Compute SHA-256 hex digest of a string. -fn sha256_hex(input: &str) -> String { - let mut hasher = Sha256::new(); - hasher.update(input.as_bytes()); - format!("{:x}", hasher.finalize()) -} - -// --------------------------------------------------------------------------- -// Router -// --------------------------------------------------------------------------- - -/// Build federation API routes. -pub fn federation_router(state: FederationState) -> Router { - Router::new() - .route("/federation/peers", get(list_peers)) - .route("/federation/register", post(register_peer)) - .route("/federation/heartbeat", post(heartbeat)) - .route("/federation/query", post(federation_query)) - .route("/federation/deregister/{store_id}", post(deregister_peer)) - .with_state(state) -} - -// --------------------------------------------------------------------------- -// Handlers -// --------------------------------------------------------------------------- - -/// List all known peer stores. -#[instrument(skip(state))] -async fn list_peers( - State(state): State, - Query(params): Query, -) -> Result>, StatusCode> { - let peers = state.peers.read().map_err(|_| { - tracing::error!("Federation peers RwLock poisoned"); - StatusCode::INTERNAL_SERVER_ERROR - })?; - - let result: Vec = peers - .values() - .filter(|p| { - if let Some(ref modality) = params.modality { - p.modalities.iter().any(|m| m == modality) - } else { - true - } - }) - .cloned() - .collect(); - - Ok(Json(result)) -} - -/// Register a new peer store in the federation. -/// -/// Requires `VERISIM_FEDERATION_KEYS` to be set with the store_id:key pair. -/// Federation registration is disabled (403) when no keys are configured. -#[instrument(skip(state))] -async fn register_peer( - State(state): State, - Json(request): Json, -) -> Result<(StatusCode, Json), StatusCode> { - // Reject if federation registration is disabled - if !state.registration_enabled() { - warn!("Federation registration rejected: VERISIM_FEDERATION_KEYS not configured"); - return Err(StatusCode::FORBIDDEN); - } - - // Validate store_id format (alphanumeric + dash + underscore, max 128) - if request.store_id.is_empty() - || request.store_id.len() > 128 - || !request.store_id.chars().all(|c| c.is_alphanumeric() || c == '-' || c == '_' || c == '/') - { - return Err(StatusCode::BAD_REQUEST); - } - - // Validate PSK - let secret = request.secret.as_deref().unwrap_or(""); - if !state.validate_psk(&request.store_id, secret) { - warn!(store_id = %request.store_id, "Federation registration rejected: invalid PSK"); - return Err(StatusCode::UNAUTHORIZED); - } - - let secret_hash = Some(sha256_hex(secret)); - - let peer = PeerStore { - store_id: request.store_id.clone(), - endpoint: request.endpoint, - modalities: request.modalities, - trust_level: 1.0, - last_seen: Some(chrono::Utc::now().to_rfc3339()), - response_time_ms: None, - secret_hash, - }; - - info!(store_id = %request.store_id, "Registered peer store"); - - state - .peers - .write() - .map_err(|_| { - error!("Federation peers RwLock poisoned"); - StatusCode::INTERNAL_SERVER_ERROR - })? - .insert(request.store_id, peer.clone()); - - Ok((StatusCode::CREATED, Json(peer))) -} - -/// Receive a heartbeat from a peer. -/// -/// Requires `X-Federation-PSK` header matching the stored peer secret. -#[instrument(skip(state, headers))] -async fn heartbeat( - State(state): State, - headers: HeaderMap, - Json(request): Json, -) -> Result { - state.validate_peer_header(&request.store_id, &headers)?; - - let mut peers = state.peers.write().map_err(|_| { - error!("Federation peers RwLock poisoned"); - StatusCode::INTERNAL_SERVER_ERROR - })?; - - if let Some(peer) = peers.get_mut(&request.store_id) { - peer.last_seen = Some(chrono::Utc::now().to_rfc3339()); - Ok(StatusCode::OK) - } else { - Ok(StatusCode::NOT_FOUND) - } -} - -/// Remove a peer from the federation. -#[instrument(skip(state))] -async fn deregister_peer( - State(state): State, - axum::extract::Path(store_id): axum::extract::Path, -) -> Result { - let removed = state - .peers - .write() - .map_err(|_| { - tracing::error!("Federation peers RwLock poisoned"); - StatusCode::INTERNAL_SERVER_ERROR - })? - .remove(&store_id); - - if removed.is_some() { - info!(store_id = %store_id, "Deregistered peer store"); - Ok(StatusCode::OK) - } else { - Ok(StatusCode::NOT_FOUND) - } -} - -/// Execute a federated query across matching peer stores. -#[instrument(skip(state))] -async fn federation_query( - State(state): State, - Json(request): Json, -) -> Result, StatusCode> { - let limit = request.limit.unwrap_or(100).min(1000); - - // Collect matching stores and apply drift policy (hold lock briefly) - let (stores_to_query, stores_excluded) = { - let peers = state.peers.read().map_err(|_| { - tracing::error!("Federation peers RwLock poisoned"); - StatusCode::INTERNAL_SERVER_ERROR - })?; - - let matching: Vec = peers - .values() - .filter(|p| pattern_matches(&request.pattern, &p.store_id)) - .filter(|p| { - request - .modalities - .iter() - .all(|m| p.modalities.iter().any(|pm| pm == m)) - }) - .cloned() - .collect(); - - let mut included = Vec::new(); - let mut excluded = Vec::new(); - - for store in matching { - // Skip self to prevent infinite recursion - if store.store_id == state.self_store_id { - continue; - } - - let include = match request.drift_policy { - DriftPolicy::Strict => { - store.trust_level >= (1.0 - state.strict_drift_threshold) - } - DriftPolicy::Repair | DriftPolicy::Tolerate | DriftPolicy::Latest => true, - }; - - if include { - included.push(store); - } else { - info!( - store_id = %store.store_id, - trust = store.trust_level, - "Excluded store due to Strict drift policy" - ); - excluded.push(store.store_id.clone()); - } - } - - (included, excluded) - }; - // RwLock dropped here — safe to make async HTTP calls - - let stores_queried: Vec = stores_to_query - .iter() - .map(|s| s.store_id.clone()) - .collect(); - - let client = reqwest::Client::new(); - let text_query = request.text_query.clone(); - let vector_query = request.vector_query.clone(); - let drift_policy = request.drift_policy; - - // Fan out parallel queries - let mut handles = Vec::new(); - for store in stores_to_query { - let client = client.clone(); - let text_q = text_query.clone(); - let vector_q = vector_query.clone(); - - let handle = tokio::spawn(async move { - let timeout = std::time::Duration::from_secs(10); - match tokio::time::timeout( - timeout, - query_single_peer(&client, &store, text_q.as_deref(), vector_q.as_deref(), limit), - ) - .await - { - Ok(Ok(results)) => results, - Ok(Err(e)) => { - warn!(store_id = %store.store_id, error = %e, "Peer query failed"); - Vec::new() - } - Err(_) => { - warn!(store_id = %store.store_id, "Peer query timed out after 10s"); - Vec::new() - } - } - }); - handles.push(handle); - } - - // Collect results from all peers - let mut all_results: Vec = Vec::new(); - for handle in handles { - match handle.await { - Ok(mut results) => all_results.append(&mut results), - Err(e) => warn!(error = %e, "Peer query task panicked"), - } - } - - // Sort by score descending and apply global limit - all_results.sort_by(|a, b| b.score.partial_cmp(&a.score).unwrap_or(std::cmp::Ordering::Equal)); - all_results.truncate(limit); - - Ok(Json(FederationQueryResponse { - results: all_results, - stores_queried, - stores_excluded, - drift_policy, - })) -} - -/// Query a single peer store via HTTP. -async fn query_single_peer( - client: &reqwest::Client, - store: &PeerStore, - text_query: Option<&str>, - vector_query: Option<&[f32]>, - limit: usize, -) -> Result, String> { - let endpoint = &store.endpoint; - let store_id = &store.store_id; - - let response_items: Vec = if let Some(q) = text_query { - // Text search - let url = format!("{}/search/text", endpoint); - let resp = client - .get(&url) - .query(&[("q", q), ("limit", &limit.to_string())]) - .send() - .await - .map_err(|e| format!("HTTP request to {} failed: {}", store_id, e))?; - - if !resp.status().is_success() { - return Err(format!("Peer {} returned status {}", store_id, resp.status())); - } - - resp.json() - .await - .map_err(|e| format!("Failed to parse response from {}: {}", store_id, e))? - } else if let Some(vec) = vector_query { - // Vector search - let url = format!("{}/search/vector", endpoint); - let body = serde_json::json!({ "vector": vec, "k": limit }); - let resp = client - .post(&url) - .json(&body) - .send() - .await - .map_err(|e| format!("HTTP request to {} failed: {}", store_id, e))?; - - if !resp.status().is_success() { - return Err(format!("Peer {} returned status {}", store_id, resp.status())); - } - - resp.json() - .await - .map_err(|e| format!("Failed to parse response from {}: {}", store_id, e))? - } else { - // No specific query — list octads from the peer's /octads endpoint - let url = format!("{}/octads", endpoint); - let resp = client - .get(&url) - .query(&[("limit", &limit.to_string())]) - .send() - .await - .map_err(|e| format!("HTTP request to {} failed: {}", store_id, e))?; - - if !resp.status().is_success() { - return Err(format!("Peer {} returned status {}", store_id, resp.status())); - } - - resp.json() - .await - .map_err(|e| format!("Failed to parse response from {}: {}", store_id, e))? - }; - - // Map response items to FederationResult - let results = response_items - .into_iter() - .map(|item| FederationResult { - source_store: store_id.clone(), - octad_id: item["id"].as_str().unwrap_or("unknown").to_string(), - score: item["score"].as_f64().unwrap_or(0.0), - drifted: false, - data: item, - }) - .collect(); - - Ok(results) -} - -// --------------------------------------------------------------------------- -// Helpers -// --------------------------------------------------------------------------- - -/// Match a federation pattern against a store ID. -/// Supports glob-style wildcards: "/universities/*" matches "/universities/oxford". -fn pattern_matches(pattern: &str, store_id: &str) -> bool { - if pattern == "*" { - return true; - } - - if let Some(prefix) = pattern.strip_suffix("/*") { - store_id.starts_with(prefix) - } else { - pattern == store_id - } -} - -#[cfg(test)] -mod tests { - use super::*; - - #[test] - fn test_pattern_matching() { - assert!(pattern_matches("*", "any-store")); - assert!(pattern_matches("/universities/*", "/universities/oxford")); - assert!(pattern_matches("/universities/*", "/universities/cambridge")); - assert!(!pattern_matches("/universities/*", "/hospitals/nhs")); - assert!(pattern_matches("store-1", "store-1")); - assert!(!pattern_matches("store-1", "store-2")); - } - - #[test] - fn test_federation_state() { - let state = FederationState::new("self".to_string(), "http://localhost:8080".to_string()); - - let peer = PeerStore { - store_id: "peer-1".to_string(), - endpoint: "http://peer-1:8080".to_string(), - modalities: vec!["graph".to_string(), "vector".to_string()], - trust_level: 0.95, - last_seen: None, - response_time_ms: None, - secret_hash: None, - }; - - state - .peers - .write() - .expect("TODO: handle error") - .insert("peer-1".to_string(), peer); - - let peers = state.peers.read().expect("TODO: handle error"); - assert_eq!(peers.len(), 1); - assert_eq!(peers["peer-1"].trust_level, 0.95); - } - - #[test] - fn test_registration_disabled_by_default() { - let state = FederationState::new("self".to_string(), "http://localhost:8080".to_string()); - // Without VERISIM_FEDERATION_KEYS, registration should be disabled - assert!(!state.registration_enabled()); - } - - #[test] - fn test_sha256_hex() { - let hash = sha256_hex("test-secret"); - assert_eq!(hash.len(), 64); // SHA-256 produces 64 hex chars - } -} diff --git a/verisimdb/rust-core/verisim-api/src/graphql.rs b/verisimdb/rust-core/verisim-api/src/graphql.rs deleted file mode 100644 index 8e7e5caa..00000000 --- a/verisimdb/rust-core/verisim-api/src/graphql.rs +++ /dev/null @@ -1,555 +0,0 @@ -// SPDX-License-Identifier: MPL-2.0 -//! GraphQL API for VeriSimDB. -//! -//! Exposes planner, octad, search, drift, and normalizer operations -//! via a GraphQL schema at `/graphql`. - -use async_graphql::{ - Context, EmptySubscription, InputObject, Object, Schema, SimpleObject, - http::GraphiQLSource, -}; -use async_graphql_axum::{GraphQLRequest, GraphQLResponse}; -use tracing::error; -use axum::{ - extract::State as AxumState, - response::{Html, IntoResponse}, - routing::{get, post}, - Router, -}; - -use verisim_planner::{ - ExplainOutput as PlannerExplainOutput, - LogicalPlan, -}; - -use crate::AppState; - -// ============================================================================ -// GraphQL Output Types -// ============================================================================ - -/// Health check result. -#[derive(SimpleObject)] -struct Health { - status: String, - version: String, - uptime_seconds: u64, -} - -/// Octad summary. -#[derive(SimpleObject)] -struct Octad { - id: String, - created_at: String, - modified_at: String, - version: u64, - has_graph: bool, - has_vector: bool, - has_tensor: bool, - has_semantic: bool, - has_document: bool, - version_count: u64, -} - -/// Search result entry. -#[derive(SimpleObject)] -struct SearchResult { - id: String, - score: f32, - title: Option, -} - -/// Drift status for a single drift type. -#[derive(SimpleObject)] -struct DriftStatus { - drift_type: String, - current_score: f64, - moving_average: f64, - max_score: f64, - measurement_count: u64, -} - -/// A single step in a physical plan. -#[derive(SimpleObject)] -struct PlanStep { - step: i32, - operation: String, - modality: String, - time_ms: f64, - estimated_rows: u64, - selectivity: f64, - optimization_hint: Option, -} - -/// Optimized physical plan. -#[derive(SimpleObject)] -struct PhysicalPlan { - steps: Vec, - strategy: String, - total_time_ms: f64, - total_estimated_rows: u64, - notes: Vec, -} - -/// Cost breakdown for a modality. -#[derive(SimpleObject)] -struct ModalityCost { - modality: String, - time_ms: f64, - percentage: f64, -} - -/// Performance hint. -#[derive(SimpleObject)] -struct PerformanceHint { - severity: String, - message: String, -} - -/// EXPLAIN output. -#[derive(SimpleObject)] -struct ExplainOutput { - steps: Vec, - cost_breakdown: Vec, - performance_hints: Vec, - total_cost_ms: f64, - strategy: String, - text_output: String, -} - -/// Planner configuration. -#[derive(SimpleObject, Clone)] -struct PlannerConfigOutput { - global_mode: String, - statistics_weight: f64, - enable_adaptive: bool, - parallel_threshold: i32, -} - -/// Store statistics for a modality. -#[derive(SimpleObject)] -struct StoreStats { - modality: String, - total_rows: u64, - avg_latency_ms: f64, - avg_rows_returned: u64, - query_count: u64, -} - -/// Full statistics snapshot. -#[derive(SimpleObject)] -struct PlannerStats { - stores: Vec, -} - -// ============================================================================ -// GraphQL Input Types -// ============================================================================ - -/// Octad creation/update input. -#[derive(InputObject)] -struct OctadInput { - title: Option, - body: Option, - embedding: Option>, - types: Option>, - tensor_shape: Option>, - tensor_data: Option>, -} - -/// Planner configuration input. -#[derive(InputObject)] -struct PlannerConfigInput { - global_mode: Option, - statistics_weight: Option, - enable_adaptive: Option, - parallel_threshold: Option, -} - -// ============================================================================ -// Query Root -// ============================================================================ - -pub struct QueryRoot; - -#[Object] -impl QueryRoot { - /// Health check. - async fn health(&self, ctx: &Context<'_>) -> async_graphql::Result { - let state = ctx.data::()?; - Ok(Health { - status: "healthy".to_string(), - version: env!("CARGO_PKG_VERSION").to_string(), - uptime_seconds: state.start_time.elapsed().as_secs(), - }) - } - - /// Get a octad by ID. - async fn octad(&self, ctx: &Context<'_>, id: String) -> async_graphql::Result> { - let state = ctx.data::()?; - let octad_id = verisim_octad::OctadId::new(&id); - - use verisim_octad::OctadStore; - match state.octad_store.get(&octad_id).await { - Ok(Some(h)) => Ok(Some(Octad { - id: h.id.to_string(), - created_at: h.status.created_at.to_rfc3339(), - modified_at: h.status.modified_at.to_rfc3339(), - version: h.status.version, - has_graph: h.graph_node.is_some(), - has_vector: h.embedding.is_some(), - has_tensor: h.tensor.is_some(), - has_semantic: h.semantic.is_some(), - has_document: h.document.is_some(), - version_count: h.version_count, - })), - Ok(None) => Ok(None), - Err(e) => { - error!(error = %e, "GraphQL octad query failed"); - Err(async_graphql::Error::new("Internal server error")) - } - } - } - - /// Search by text. - async fn search_text( - &self, - ctx: &Context<'_>, - query: String, - limit: Option, - ) -> async_graphql::Result> { - let state = ctx.data::()?; - let limit = limit.unwrap_or(10) as usize; - - use verisim_octad::OctadStore; - let octads = state - .octad_store - .search_text(&query, limit) - .await - .map_err(|e| { - error!(error = %e, "GraphQL text search failed"); - async_graphql::Error::new("Internal server error") - })?; - - Ok(octads - .iter() - .enumerate() - .map(|(i, h)| SearchResult { - id: h.id.to_string(), - score: 1.0 - (i as f32 * 0.1), - title: h.document.as_ref().map(|d| d.title.clone()), - }) - .collect()) - } - - /// Get drift status for all drift types. - async fn drift_status(&self, ctx: &Context<'_>) -> async_graphql::Result> { - let state = ctx.data::()?; - let all_metrics = state.drift_detector.all_metrics() - .map_err(|e| { - error!(error = %e, "GraphQL drift status query failed"); - async_graphql::Error::new("Internal server error") - })?; - - Ok(all_metrics - .iter() - .map(|(dt, m)| DriftStatus { - drift_type: dt.to_string(), - current_score: m.current_score, - moving_average: m.moving_average, - max_score: m.max_score, - measurement_count: m.measurement_count, - }) - .collect()) - } - - /// Get current planner configuration. - async fn planner_config(&self, ctx: &Context<'_>) -> async_graphql::Result { - let state = ctx.data::()?; - let planner = state.planner.lock().map_err(|_| { error!("Planner lock poisoned in GraphQL"); async_graphql::Error::new("Internal server error") })?; - let cfg = planner.config(); - Ok(PlannerConfigOutput { - global_mode: format!("{:?}", cfg.global_mode), - statistics_weight: cfg.statistics_weight, - enable_adaptive: cfg.enable_adaptive, - parallel_threshold: cfg.parallel_threshold as i32, - }) - } - - /// Get planner statistics. - async fn planner_stats(&self, ctx: &Context<'_>) -> async_graphql::Result { - let state = ctx.data::()?; - let planner = state.planner.lock().map_err(|_| { error!("Planner lock poisoned in GraphQL"); async_graphql::Error::new("Internal server error") })?; - - let stores: Vec = verisim_planner::Modality::ALL - .iter() - .filter_map(|m| { - planner.stats().get(*m).map(|s| StoreStats { - modality: m.to_string(), - total_rows: s.total_rows, - avg_latency_ms: s.avg_latency_ms, - avg_rows_returned: s.avg_rows_returned, - query_count: s.query_count, - }) - }) - .collect(); - - Ok(PlannerStats { stores }) - } - - /// EXPLAIN a logical plan. - async fn explain_plan( - &self, - ctx: &Context<'_>, - plan_json: String, - ) -> async_graphql::Result { - let state = ctx.data::()?; - let logical: LogicalPlan = serde_json::from_str(&plan_json) - .map_err(|e| async_graphql::Error::new(format!("Invalid plan JSON: {}", e)))?; - - let planner = state.planner.lock().map_err(|_| { - error!("Planner lock poisoned in GraphQL explain"); - async_graphql::Error::new("Internal server error") - })?; - let explain = planner - .explain(&logical) - .map_err(|e| async_graphql::Error::new(format!("Plan error: {}", e)))?; - - Ok(convert_explain(&explain)) - } -} - -// ============================================================================ -// Mutation Root -// ============================================================================ - -pub struct MutationRoot; - -#[Object] -impl MutationRoot { - /// Create a new octad. - async fn create_octad( - &self, - ctx: &Context<'_>, - input: OctadInput, - ) -> async_graphql::Result { - let state = ctx.data::()?; - - let mut octad_input = verisim_octad::OctadInput::default(); - if let Some(title) = &input.title { - octad_input.document = Some(verisim_octad::OctadDocumentInput { - title: title.clone(), - body: input.body.clone().unwrap_or_default(), - fields: std::collections::HashMap::new(), - }); - } - if let Some(embedding) = &input.embedding { - octad_input.vector = Some(verisim_octad::OctadVectorInput { - embedding: embedding.clone(), - model: None, - }); - } - if let Some(types) = &input.types { - octad_input.semantic = Some(verisim_octad::OctadSemanticInput { - types: types.clone(), - properties: std::collections::HashMap::new(), - }); - } - - use verisim_octad::OctadStore; - let h = state - .octad_store - .create(octad_input) - .await - .map_err(|e| { - error!(error = %e, "GraphQL octad creation failed"); - async_graphql::Error::new("Internal server error") - })?; - - Ok(Octad { - id: h.id.to_string(), - created_at: h.status.created_at.to_rfc3339(), - modified_at: h.status.modified_at.to_rfc3339(), - version: h.status.version, - has_graph: h.graph_node.is_some(), - has_vector: h.embedding.is_some(), - has_tensor: h.tensor.is_some(), - has_semantic: h.semantic.is_some(), - has_document: h.document.is_some(), - version_count: h.version_count, - }) - } - - /// Delete a octad. - async fn delete_octad(&self, ctx: &Context<'_>, id: String) -> async_graphql::Result { - let state = ctx.data::()?; - let octad_id = verisim_octad::OctadId::new(&id); - - use verisim_octad::OctadStore; - state - .octad_store - .delete(&octad_id) - .await - .map_err(|e| { - error!(error = %e, "GraphQL octad deletion failed"); - async_graphql::Error::new("Internal server error") - })?; - - Ok(true) - } - - /// Optimize a logical plan into a physical plan. - async fn optimize_plan( - &self, - ctx: &Context<'_>, - plan_json: String, - ) -> async_graphql::Result { - let state = ctx.data::()?; - let logical: LogicalPlan = serde_json::from_str(&plan_json) - .map_err(|e| async_graphql::Error::new(format!("Invalid plan JSON: {}", e)))?; - - let planner = state.planner.lock().map_err(|_| { error!("Planner lock poisoned in GraphQL"); async_graphql::Error::new("Internal server error") })?; - let physical = planner - .optimize(&logical) - .map_err(|e| async_graphql::Error::new(e.to_string()))?; - - Ok(convert_physical_plan(&physical)) - } - - /// Update planner configuration. - async fn update_planner_config( - &self, - ctx: &Context<'_>, - input: PlannerConfigInput, - ) -> async_graphql::Result { - let state = ctx.data::()?; - let mut planner = state.planner.lock().map_err(|_| { error!("Planner lock poisoned in GraphQL"); async_graphql::Error::new("Internal server error") })?; - - let mut cfg = planner.config().clone(); - if let Some(mode) = &input.global_mode { - cfg.global_mode = match mode.to_lowercase().as_str() { - "conservative" => verisim_planner::OptimizationMode::Conservative, - "aggressive" => verisim_planner::OptimizationMode::Aggressive, - _ => verisim_planner::OptimizationMode::Balanced, - }; - } - if let Some(w) = input.statistics_weight { - cfg.statistics_weight = w; - } - if let Some(a) = input.enable_adaptive { - cfg.enable_adaptive = a; - } - if let Some(t) = input.parallel_threshold { - cfg.parallel_threshold = t as usize; - } - planner.set_config(cfg); - - let cfg = planner.config(); - Ok(PlannerConfigOutput { - global_mode: format!("{:?}", cfg.global_mode), - statistics_weight: cfg.statistics_weight, - enable_adaptive: cfg.enable_adaptive, - parallel_threshold: cfg.parallel_threshold as i32, - }) - } -} - -// ============================================================================ -// Schema Construction -// ============================================================================ - -pub type VeriSimSchema = Schema; - -/// Build the GraphQL schema with AppState as context data. -pub fn build_schema(state: AppState) -> VeriSimSchema { - Schema::build(QueryRoot, MutationRoot, EmptySubscription) - .data(state) - .finish() -} - -/// GraphQL request handler. -async fn graphql_handler( - AxumState(schema): AxumState, - req: GraphQLRequest, -) -> GraphQLResponse { - schema.execute(req.into_inner()).await.into() -} - -/// GraphiQL playground handler. -async fn graphiql_handler() -> impl IntoResponse { - Html(GraphiQLSource::build().endpoint("/graphql").finish()) -} - -/// Build the GraphQL axum router. -pub fn graphql_router(state: AppState) -> Router { - let schema = build_schema(state); - - Router::new() - .route("/graphql", post(graphql_handler)) - .route("/graphiql", get(graphiql_handler)) - .with_state(schema) -} - -// ============================================================================ -// Conversion Helpers -// ============================================================================ - -fn convert_physical_plan(p: &verisim_planner::PhysicalPlan) -> PhysicalPlan { - PhysicalPlan { - steps: p - .steps - .iter() - .map(|s| PlanStep { - step: s.step as i32, - operation: s.operation.clone(), - modality: s.modality.to_string(), - time_ms: s.cost.time_ms, - estimated_rows: s.cost.estimated_rows, - selectivity: s.cost.selectivity, - optimization_hint: s.optimization_hint.clone(), - }) - .collect(), - strategy: format!("{:?}", p.strategy), - total_time_ms: p.total_cost.time_ms, - total_estimated_rows: p.total_cost.estimated_rows, - notes: p.notes.clone(), - } -} - -fn convert_explain(e: &PlannerExplainOutput) -> ExplainOutput { - ExplainOutput { - steps: e - .steps - .iter() - .map(|s| PlanStep { - step: s.step as i32, - operation: s.operation.clone(), - modality: s.modality.to_string(), - time_ms: s.estimated_cost_ms, - estimated_rows: s.estimated_rows, - selectivity: s.estimated_selectivity, - optimization_hint: s.optimization_hint.clone(), - }) - .collect(), - cost_breakdown: e - .cost_breakdown - .iter() - .map(|cb| ModalityCost { - modality: cb.modality.to_string(), - time_ms: cb.time_ms, - percentage: cb.percentage, - }) - .collect(), - performance_hints: e - .performance_hints - .iter() - .map(|h| PerformanceHint { - severity: h.severity.clone(), - message: h.message.clone(), - }) - .collect(), - total_cost_ms: e.total_cost_ms, - strategy: e.strategy.clone(), - text_output: e.text_output.clone(), - } -} diff --git a/verisimdb/rust-core/verisim-api/src/groove.rs b/verisimdb/rust-core/verisim-api/src/groove.rs deleted file mode 100644 index 4ed1e0da..00000000 --- a/verisimdb/rust-core/verisim-api/src/groove.rs +++ /dev/null @@ -1,890 +0,0 @@ -// SPDX-License-Identifier: MPL-2.0 -// Copyright (c) 2026 Jonathan D.A. Jewell (hyperpolymath) -// -//! Groove Protocol connection lifecycle for VeriSimDB. -//! -//! Implements the connect/disconnect lifecycle from the Groove Protocol spec -//! (section 4). Allows groove-aware systems (Gossamer, Burble, PanLL, etc.) -//! to establish typed connections with VeriSimDB for octad storage, drift -//! detection, and provenance services. -//! -//! Endpoints: -//! GET /.well-known/groove — Capability manifest -//! POST /.well-known/groove/connect — Establish connection (spec 4.2) -//! POST /.well-known/groove/disconnect — Tear down connection (spec 4.5) -//! GET /.well-known/groove/heartbeat — Heartbeat keepalive (spec 4.3) -//! GET /.well-known/groove/status — Current connection states -//! -//! Connection state machine: -//! DISCOVERED -> NEGOTIATING -> CONNECTED -> ACTIVE -> DISCONNECTING -> DISCONNECTED -//! | | -//! REJECTED DEGRADED -> RECONNECTING -> ACTIVE - -use axum::{ - extract::{Query, State}, - http::StatusCode, - response::IntoResponse, - routing::{get, post}, - Json, Router, -}; -use serde::{Deserialize, Serialize}; -use std::collections::HashMap; -use std::sync::{Arc, Mutex}; -use std::time::{Instant, SystemTime, UNIX_EPOCH}; -use tracing::{info, warn}; - -/// Groove connection state, mirroring the spec state machine. -#[derive(Debug, Clone, Copy, PartialEq, Eq, Serialize, Deserialize)] -#[serde(rename_all = "lowercase")] -pub enum ConnectionState { - /// Peer discovered but not yet negotiated. - Discovered, - /// Capability negotiation in progress. - Negotiating, - /// Connection established, capabilities available. - Connected, - /// Active data exchange (heartbeats received). - Active, - /// Heartbeats missed, capabilities may be unreliable. - Degraded, - /// Attempting to re-establish after degradation. - Reconnecting, - /// Graceful shutdown in progress. - Disconnecting, - /// Connection closed. - Disconnected, - /// Peer rejected (incompatible capabilities or security). - Rejected, -} - -/// Information about a single groove connection. -#[derive(Debug, Clone, Serialize)] -pub struct ConnectionInfo { - /// Peer's self-reported service ID. - pub peer_id: String, - /// Current lifecycle state. - pub state: ConnectionState, - /// Unix timestamp (milliseconds) when the connection was established. - pub connected_at_ms: u64, - /// Unix timestamp (milliseconds) of the last heartbeat. - pub last_heartbeat_ms: u64, - /// Monotonic instant of last heartbeat (not serialised, used for timeout). - #[serde(skip)] - pub last_heartbeat_instant: Instant, - /// Capabilities that matched between provider and consumer. - pub matched_capabilities: Vec, -} - -/// Shared state for all groove connections. -#[derive(Debug, Clone)] -pub struct GrooveState { - /// Active connections keyed by session ID. - connections: Arc>>, -} - -impl GrooveState { - /// Create a new empty groove state. - pub fn new() -> Self { - Self { - connections: Arc::new(Mutex::new(HashMap::new())), - } - } -} - -impl Default for GrooveState { - fn default() -> Self { - Self::new() - } -} - -/// Static capability manifest for VeriSimDB. -/// -/// Declares what VeriSimDB offers to groove consumers (octad storage, drift -/// detection, provenance, spatial search, etc.) and what it consumes from -/// groove partners (integrity verification, scanning). -fn manifest() -> serde_json::Value { - serde_json::json!({ - "groove_version": "1", - "service_id": "verisimdb", - "service_version": env!("CARGO_PKG_VERSION"), - "capabilities": { - "octad-storage": { - "type": "octad-storage", - "description": "8-modality entity storage with drift detection and self-normalisation", - "protocol": "http", - "endpoint": "/api/v1/octads", - "requires_auth": false, - "panel_compatible": true - }, - "drift-detection": { - "type": "drift-detection", - "description": "Cross-modal drift measurement and alerting", - "protocol": "http", - "endpoint": "/api/v1/drift/status", - "requires_auth": false, - "panel_compatible": true - }, - "provenance": { - "type": "provenance", - "description": "Hash-chain provenance tracking for entity lineage", - "protocol": "http", - "endpoint": "/api/v1/provenance", - "requires_auth": false, - "panel_compatible": true - }, - "vector-search": { - "type": "vector-search", - "description": "HNSW similarity search over entity embeddings", - "protocol": "http", - "endpoint": "/api/v1/search/vector", - "requires_auth": false, - "panel_compatible": false - }, - "spatial-search": { - "type": "spatial-search", - "description": "R-tree geospatial queries (radius, bounds, nearest)", - "protocol": "http", - "endpoint": "/api/v1/spatial/search", - "requires_auth": false, - "panel_compatible": false - }, - "vcl": { - "type": "vcl", - "description": "VeriSim Consonance Language — type-safe multi-modal queries", - "protocol": "http", - "endpoint": "/api/v1/vcl/execute", - "requires_auth": false, - "panel_compatible": true - }, - "feedback": { - "type": "feedback", - "description": "Groove-routed feedback receiver — stores feedback targeted at VeriSimDB", - "protocol": "http", - "endpoint": "/.well-known/groove/feedback", - "requires_auth": false, - "panel_compatible": false - }, - "health-mesh": { - "type": "health-mesh", - "description": "Inter-service health mesh — monitors peer status via groove probing", - "protocol": "http", - "endpoint": "/.well-known/groove/mesh", - "requires_auth": false, - "panel_compatible": true - } - }, - "consumes": ["integrity", "scanning"], - "endpoints": { - "api": "http://localhost:8080/api/v1", - "health": "http://localhost:8080/health", - "graphql": "http://localhost:8080/graphql" - }, - "health": "/health", - "heartbeat": { - "interval_ms": 5000, - "timeout_ms": 15000 - } - }) -} - -/// Our offered capability IDs (for matching against consumer "consumes" list). -const OFFERED_CAPABILITIES: &[&str] = &[ - "octad-storage", - "drift-detection", - "provenance", - "vector-search", - "spatial-search", - "vcl", -]; - -/// Heartbeat timeout: 15 seconds (3 missed heartbeats at 5s interval, per spec 4.3). -const HEARTBEAT_TIMEOUT_MS: u64 = 15_000; - -// --- Request/Response types --- - -/// Body for POST /.well-known/groove/connect -#[derive(Debug, Deserialize)] -pub struct ConnectRequest { - /// The peer's service ID. - pub service_id: Option, - /// The peer's service version. - pub service_version: Option, - /// Capabilities the peer consumes (we check if we offer them). - pub consumes: Option>, - /// Full peer manifest (opaque, stored for reference). - #[serde(flatten)] - pub extra: HashMap, -} - -/// Response for POST /.well-known/groove/connect -#[derive(Debug, Serialize)] -pub struct ConnectResponse { - pub ok: bool, - pub session_id: Option, - pub provider: &'static str, - pub state: &'static str, - #[serde(skip_serializing_if = "Option::is_none")] - pub error: Option, - #[serde(skip_serializing_if = "Option::is_none")] - pub matched_capabilities: Option>, -} - -/// Body for POST /.well-known/groove/disconnect -#[derive(Debug, Deserialize)] -pub struct DisconnectRequest { - pub session_id: String, -} - -/// Query params for GET /.well-known/groove/heartbeat -#[derive(Debug, Deserialize)] -pub struct HeartbeatQuery { - pub session_id: String, -} - -// --- Handlers --- - -/// GET /.well-known/groove — Return the capability manifest. -async fn groove_manifest_handler() -> impl IntoResponse { - Json(manifest()) -} - -/// POST /.well-known/groove/connect — Establish a groove connection. -/// -/// The consumer sends its manifest. VeriSimDB checks structural compatibility -/// (does the consumer consume something we offer?) and returns a session ID -/// if compatible. Per spec section 4.2. -async fn groove_connect_handler( - State(groove): State, - Json(req): Json, -) -> impl IntoResponse { - let peer_id = req.service_id.unwrap_or_else(|| "unknown".to_string()); - let peer_consumes = req.consumes.unwrap_or_default(); - - // Structural compatibility check: does the peer consume anything we offer? - let matched: Vec = peer_consumes - .iter() - .filter(|cap| OFFERED_CAPABILITIES.contains(&cap.as_str())) - .cloned() - .collect(); - - if matched.is_empty() && !peer_consumes.is_empty() { - info!( - peer_id = %peer_id, - "Groove connection rejected: no capability match" - ); - return ( - StatusCode::CONFLICT, - Json(ConnectResponse { - ok: false, - session_id: None, - provider: "verisimdb", - state: "rejected", - error: Some("no matching capabilities".to_string()), - matched_capabilities: None, - }), - ); - } - - let session_id = generate_session_id(); - let now_ms = unix_now_ms(); - - let conn_info = ConnectionInfo { - peer_id: peer_id.clone(), - state: ConnectionState::Connected, - connected_at_ms: now_ms, - last_heartbeat_ms: now_ms, - last_heartbeat_instant: Instant::now(), - matched_capabilities: matched.clone(), - }; - - { - let mut connections = groove.connections.lock().expect("groove lock poisoned"); - connections.insert(session_id.clone(), conn_info); - } - - info!( - peer_id = %peer_id, - session_id = %session_id, - capabilities = ?matched, - "Groove connection established" - ); - - ( - StatusCode::OK, - Json(ConnectResponse { - ok: true, - session_id: Some(session_id), - provider: "verisimdb", - state: "connected", - error: None, - matched_capabilities: Some(matched), - }), - ) -} - -/// POST /.well-known/groove/disconnect — Tear down a groove connection. -/// -/// Consumes the linear connection handle. Per spec section 4.5. -async fn groove_disconnect_handler( - State(groove): State, - Json(req): Json, -) -> impl IntoResponse { - let mut connections = groove.connections.lock().expect("groove lock poisoned"); - - match connections.remove(&req.session_id) { - Some(info) => { - info!( - peer_id = %info.peer_id, - session_id = %req.session_id, - "Groove connection disconnected" - ); - ( - StatusCode::OK, - Json(serde_json::json!({"ok": true, "state": "disconnected"})), - ) - } - None => ( - StatusCode::NOT_FOUND, - Json(serde_json::json!({"ok": false, "error": "session not found"})), - ), - } -} - -/// GET /.well-known/groove/heartbeat — Heartbeat from connected peer. -/// -/// Per spec section 4.3. Returns 204 No Content on success. -async fn groove_heartbeat_handler( - State(groove): State, - Query(params): Query, -) -> impl IntoResponse { - let mut connections = groove.connections.lock().expect("groove lock poisoned"); - - match connections.get_mut(¶ms.session_id) { - Some(info) => { - info.last_heartbeat_ms = unix_now_ms(); - info.last_heartbeat_instant = Instant::now(); - // Promote from connected/degraded to active on heartbeat. - if info.state == ConnectionState::Connected || info.state == ConnectionState::Degraded { - info.state = ConnectionState::Active; - } - StatusCode::NO_CONTENT.into_response() - } - None => ( - StatusCode::NOT_FOUND, - Json(serde_json::json!({"ok": false, "error": "session not found"})), - ) - .into_response(), - } -} - -/// GET /.well-known/groove/status — Current connection state for all peers. -/// -/// Also performs heartbeat timeout checks: connections that have not sent a -/// heartbeat within 15 seconds are transitioned to DEGRADED; connections -/// already DEGRADED that miss another timeout are removed. -async fn groove_status_handler(State(groove): State) -> impl IntoResponse { - let mut connections = groove.connections.lock().expect("groove lock poisoned"); - let now = Instant::now(); - - // Check heartbeat timeouts and update states. - let mut to_remove = Vec::new(); - for (session_id, info) in connections.iter_mut() { - let elapsed_ms = now.duration_since(info.last_heartbeat_instant).as_millis() as u64; - - if elapsed_ms > HEARTBEAT_TIMEOUT_MS { - if info.state == ConnectionState::Degraded { - // Already degraded and still no heartbeat — remove. - warn!( - peer_id = %info.peer_id, - session_id = %session_id, - "Groove peer timed out, removing" - ); - to_remove.push(session_id.clone()); - } else if info.state == ConnectionState::Active - || info.state == ConnectionState::Connected - { - // Transition to degraded. - warn!( - peer_id = %info.peer_id, - session_id = %session_id, - elapsed_ms = elapsed_ms, - "Groove peer degraded (no heartbeat)" - ); - info.state = ConnectionState::Degraded; - } - } - } - - for session_id in &to_remove { - connections.remove(session_id); - } - - // Build response. - let status: HashMap<&String, _> = connections - .iter() - .map(|(id, info)| { - ( - id, - serde_json::json!({ - "peer_id": info.peer_id, - "state": info.state, - "connected_at_ms": info.connected_at_ms, - "last_heartbeat_ms": info.last_heartbeat_ms, - "matched_capabilities": info.matched_capabilities, - }), - ) - }) - .collect(); - - Json(serde_json::json!({ - "service_id": "verisimdb", - "active_connections": connections.len(), - "connections": status, - })) -} - -// --- Health Mesh --- - -/// Cached health state of groove peers, updated by the mesh monitor. -#[derive(Debug, Clone, Serialize)] -pub struct PeerHealth { - /// Peer's self-reported service ID. - pub service_id: String, - /// Port the peer was discovered on. - pub port: u16, - /// "up", "degraded", or "down". - pub status: String, - /// Unix timestamp (milliseconds) of last successful probe. - pub last_seen_ms: u64, -} - -/// Shared mesh state: list of peer health entries. -#[derive(Debug, Clone)] -pub struct MeshState { - peers: Arc>>, - last_probe_ms: Arc>, -} - -impl MeshState { - /// Create an empty mesh state. - pub fn new() -> Self { - Self { - peers: Arc::new(Mutex::new(Vec::new())), - last_probe_ms: Arc::new(Mutex::new(0)), - } - } -} - -impl Default for MeshState { - fn default() -> Self { - Self::new() - } -} - -/// Known ports to probe for groove peers (excluding our own port). -const MESH_PROBE_PORTS: &[u16] = &[6473, 8000, 8081, 8091, 8092]; - -/// Probe all known groove peers and update the mesh state. -/// -/// Called periodically by the mesh monitor background task. -fn probe_mesh_peers(mesh: &MeshState) { - let now_ms = unix_now_ms(); - let mut results = Vec::new(); - - for &port in MESH_PROBE_PORTS { - let addr_str = format!("127.0.0.1:{}", port); - let addr: std::net::SocketAddr = match addr_str.parse() { - Ok(a) => a, - Err(_) => continue, - }; - - match std::net::TcpStream::connect_timeout(&addr, std::time::Duration::from_millis(500)) { - Ok(mut stream) => { - use std::io::{Read, Write}; - stream.set_read_timeout(Some(std::time::Duration::from_millis(500))).ok(); - stream.set_write_timeout(Some(std::time::Duration::from_millis(500))).ok(); - - let request = format!( - "GET /.well-known/groove/status HTTP/1.0\r\nHost: {}\r\nConnection: close\r\n\r\n", - addr_str - ); - - if stream.write_all(request.as_bytes()).is_ok() { - let mut buf = vec![0u8; 4096]; - let service_id = match stream.read(&mut buf) { - Ok(n) if n > 0 => { - let resp = String::from_utf8_lossy(&buf[..n]); - extract_service_id(&resp) - } - _ => "unknown".to_string(), - }; - results.push(PeerHealth { - service_id, - port, - status: "up".to_string(), - last_seen_ms: now_ms, - }); - } - } - Err(_) => { - // Peer unreachable — don't include in results. - } - } - } - - if let Ok(mut peers) = mesh.peers.lock() { - *peers = results; - } - if let Ok(mut ts) = mesh.last_probe_ms.lock() { - *ts = now_ms; - } -} - -/// Extract service_id from a groove status HTTP response body. -fn extract_service_id(response: &str) -> String { - // Find body after headers. - let body = if let Some(idx) = response.find("\r\n\r\n") { - &response[idx + 4..] - } else { - response - }; - - if let Ok(v) = serde_json::from_str::(body) { - if let Some(id) = v.get("service").and_then(|s| s.as_str()) { - return id.to_string(); - } - if let Some(id) = v.get("service_id").and_then(|s| s.as_str()) { - return id.to_string(); - } - } - - "unknown".to_string() -} - -/// Spawn the mesh health monitor as a background tokio task. -/// -/// Probes peers every 30 seconds. The MeshState is shared with the -/// HTTP handler via Arc. -pub fn spawn_mesh_monitor(mesh: MeshState) { - tokio::spawn(async move { - loop { - // Run probe on a blocking thread to avoid blocking the async runtime. - let mesh_clone = mesh.clone(); - let _ = tokio::task::spawn_blocking(move || { - probe_mesh_peers(&mesh_clone); - }) - .await; - - tokio::time::sleep(std::time::Duration::from_secs(30)).await; - } - }); -} - -/// GET /.well-known/groove/mesh — Return the cached mesh health view. -async fn groove_mesh_handler(State(mesh): State) -> impl IntoResponse { - let peers = mesh.peers.lock().unwrap_or_else(|e| e.into_inner()).clone(); - let last_probe = *mesh.last_probe_ms.lock().unwrap_or_else(|e| e.into_inner()); - - Json(serde_json::json!({ - "service_id": "verisimdb", - "timestamp_ms": unix_now_ms(), - "last_probe_ms": last_probe, - "peer_count": peers.len(), - "peers": peers, - })) -} - -// --- Feedback --- - -/// Shared feedback store: timestamped feedback entries. -#[derive(Debug, Clone)] -pub struct FeedbackStore { - entries: Arc>>, -} - -/// A single feedback entry received via the Groove mesh. -#[derive(Debug, Clone, Serialize)] -pub struct FeedbackEntry { - pub id: String, - pub timestamp_ms: u64, - pub source_service: String, - pub target_service: String, - pub category: String, - pub message: String, - pub metadata: serde_json::Value, -} - -/// Maximum stored feedback entries. -const MAX_FEEDBACK_ENTRIES: usize = 10_000; - -impl FeedbackStore { - /// Create an empty feedback store. - pub fn new() -> Self { - Self { - entries: Arc::new(Mutex::new(Vec::new())), - } - } -} - -impl Default for FeedbackStore { - fn default() -> Self { - Self::new() - } -} - -/// Body for POST /.well-known/groove/feedback. -#[derive(Debug, Deserialize)] -pub struct FeedbackRequest { - #[serde(default = "default_feedback_type")] - pub r#type: String, - #[serde(default = "default_verisimdb")] - pub target_service: String, - #[serde(default = "default_other")] - pub category: String, - #[serde(default)] - pub message: String, - #[serde(default)] - pub metadata: serde_json::Value, - #[serde(default = "default_unknown")] - pub source_service: String, -} - -fn default_feedback_type() -> String { "feedback".to_string() } -fn default_verisimdb() -> String { "verisimdb".to_string() } -fn default_other() -> String { "other".to_string() } -fn default_unknown() -> String { "unknown".to_string() } - -/// POST /.well-known/groove/feedback — Receive feedback from the Groove mesh. -async fn groove_feedback_handler( - State(store): State, - Json(req): Json, -) -> impl IntoResponse { - let valid_categories = ["bug", "feature", "ux", "performance", "other"]; - if !valid_categories.contains(&req.category.as_str()) { - return ( - StatusCode::BAD_REQUEST, - Json(serde_json::json!({ - "ok": false, - "error": format!("invalid category: {}", req.category), - })), - ); - } - - let now_ms = unix_now_ms(); - let id = format!("groove-feedback-{now_ms}"); - - let entry = FeedbackEntry { - id: id.clone(), - timestamp_ms: now_ms, - source_service: req.source_service, - target_service: req.target_service.clone(), - category: req.category, - message: req.message, - metadata: req.metadata, - }; - - { - let mut entries = store.entries.lock().expect("feedback lock poisoned"); - entries.push(entry); - if entries.len() > MAX_FEEDBACK_ENTRIES { - entries.remove(0); - } - } - - info!(id = %id, "Groove feedback accepted"); - - ( - StatusCode::OK, - Json(serde_json::json!({ - "ok": true, - "routed_to": req.target_service, - "id": id, - })), - ) -} - -/// GET /.well-known/groove/feedback — List stored feedback entries. -async fn groove_feedback_list_handler( - State(store): State, -) -> impl IntoResponse { - let entries = store.entries.lock().unwrap_or_else(|e| e.into_inner()).clone(); - Json(serde_json::json!({ - "count": entries.len(), - "entries": entries, - })) -} - -/// Build the groove sub-router. -/// -/// Mounted at `/.well-known/groove` in the main application router. -/// Uses its own `GrooveState` (not `AppState`) to keep the connection -/// tracker lightweight and independent of the database layer. -/// -/// Includes health mesh monitoring and feedback-o-tron endpoints. -pub fn groove_router() -> Router { - let groove_state = GrooveState::new(); - let mesh_state = MeshState::new(); - let feedback_store = FeedbackStore::new(); - - // Spawn the background mesh monitor task. - spawn_mesh_monitor(mesh_state.clone()); - - // Connection lifecycle sub-router (uses GrooveState). - let connection_router = Router::new() - .route( - "/.well-known/groove", - get(groove_manifest_handler), - ) - .route( - "/.well-known/groove/connect", - post(groove_connect_handler), - ) - .route( - "/.well-known/groove/disconnect", - post(groove_disconnect_handler), - ) - .route( - "/.well-known/groove/heartbeat", - get(groove_heartbeat_handler), - ) - .route( - "/.well-known/groove/status", - get(groove_status_handler), - ) - .with_state(groove_state); - - // Health mesh sub-router (uses MeshState). - let mesh_router = Router::new() - .route( - "/.well-known/groove/mesh", - get(groove_mesh_handler), - ) - .with_state(mesh_state); - - // Feedback sub-router (uses FeedbackStore). - let feedback_router = Router::new() - .route( - "/.well-known/groove/feedback", - get(groove_feedback_list_handler).post(groove_feedback_handler), - ) - .with_state(feedback_store); - - connection_router - .merge(mesh_router) - .merge(feedback_router) -} - -// --- Helpers --- - -/// Generate a random hex session ID (32 hex characters = 16 bytes). -fn generate_session_id() -> String { - use std::fmt::Write; - let mut bytes = [0u8; 16]; - // Use getrandom via std (available since Rust 1.36+, backed by OS entropy). - // Fallback: timestamp + counter if getrandom is somehow unavailable. - #[cfg(unix)] - { - use std::io::Read; - if let Ok(mut f) = std::fs::File::open("/dev/urandom") { - let _ = f.read_exact(&mut bytes); - } - } - #[cfg(not(unix))] - { - // Fallback: use system time as entropy source (not cryptographically strong, - // but groove session IDs are not security-critical). - let t = SystemTime::now() - .duration_since(UNIX_EPOCH) - .unwrap_or_default() - .as_nanos(); - bytes[..8].copy_from_slice(&t.to_le_bytes()[..8]); - bytes[8..].copy_from_slice(&(t.wrapping_mul(6364136223846793005)).to_le_bytes()[..8]); - } - - let mut s = String::with_capacity(32); - for b in &bytes { - let _ = write!(s, "{:02x}", b); - } - s -} - -/// Current Unix time in milliseconds. -fn unix_now_ms() -> u64 { - SystemTime::now() - .duration_since(UNIX_EPOCH) - .unwrap_or_default() - .as_millis() as u64 -} - -#[cfg(test)] -mod tests { - use super::*; - - #[test] - fn test_manifest_has_required_fields() { - let m = manifest(); - assert_eq!(m["service_id"], "verisimdb"); - assert_eq!(m["groove_version"], "1"); - assert!(m["capabilities"]["octad-storage"].is_object()); - assert!(m["capabilities"]["drift-detection"].is_object()); - assert!(m["capabilities"]["provenance"].is_object()); - } - - #[test] - fn test_session_id_generation() { - let id1 = generate_session_id(); - let id2 = generate_session_id(); - assert_eq!(id1.len(), 32); - assert_eq!(id2.len(), 32); - // Should be different (probabilistically). - assert_ne!(id1, id2); - } - - #[test] - fn test_groove_state_connect_disconnect() { - let state = GrooveState::new(); - let session_id = "test-session-123".to_string(); - let now_ms = unix_now_ms(); - - // Insert a connection. - { - let mut conns = state.connections.lock().expect("TODO: handle error"); - conns.insert( - session_id.clone(), - ConnectionInfo { - peer_id: "burble".to_string(), - state: ConnectionState::Connected, - connected_at_ms: now_ms, - last_heartbeat_ms: now_ms, - last_heartbeat_instant: Instant::now(), - matched_capabilities: vec!["octad-storage".to_string()], - }, - ); - assert_eq!(conns.len(), 1); - } - - // Remove it. - { - let mut conns = state.connections.lock().expect("TODO: handle error"); - let removed = conns.remove(&session_id); - assert!(removed.is_some()); - assert_eq!(conns.len(), 0); - } - } - - #[test] - fn test_capability_matching() { - let peer_consumes = vec![ - "octad-storage".to_string(), - "unknown-cap".to_string(), - ]; - - let matched: Vec = peer_consumes - .iter() - .filter(|cap| OFFERED_CAPABILITIES.contains(&cap.as_str())) - .cloned() - .collect(); - - assert_eq!(matched, vec!["octad-storage".to_string()]); - } -} diff --git a/verisimdb/rust-core/verisim-api/src/grpc.rs b/verisimdb/rust-core/verisim-api/src/grpc.rs deleted file mode 100644 index 91f4d5cf..00000000 --- a/verisimdb/rust-core/verisim-api/src/grpc.rs +++ /dev/null @@ -1,444 +0,0 @@ -// SPDX-License-Identifier: MPL-2.0 -//! gRPC API for VeriSimDB. -//! -//! Exposes planner and octad operations via gRPC on a separate port (50051). - -use tonic::{Request, Response, Status}; -use tracing::error; - -use verisim_planner::LogicalPlan; - -use crate::AppState; - -// Pre-generated protobuf types (from proto/verisim.proto via prost-build). -// Using a committed file eliminates the protoc build dependency. -#[path = "proto/verisim.rs"] -pub mod proto; - -use proto::veri_sim_planner_server::{VeriSimPlanner, VeriSimPlannerServer}; -use proto::veri_sim_octad_server::{VeriSimOctad, VeriSimOctadServer}; - -// ============================================================================ -// Planner gRPC Service -// ============================================================================ - -pub struct PlannerService { - state: AppState, -} - -impl PlannerService { - pub fn new(state: AppState) -> Self { - Self { state } - } -} - -#[tonic::async_trait] -impl VeriSimPlanner for PlannerService { - async fn optimize_plan( - &self, - request: Request, - ) -> Result, Status> { - let req = request.into_inner(); - let logical: LogicalPlan = serde_json::from_str(&req.plan_json) - .map_err(|e| Status::invalid_argument(format!("Invalid plan JSON: {}", e)))?; - - let planner = self.state.planner.lock() - .map_err(|_| { error!("Planner lock poisoned in gRPC"); Status::internal("Internal server error") })?; - let physical = planner - .optimize(&logical) - .map_err(|e| Status::invalid_argument(e.to_string()))?; - - let steps: Vec = physical - .steps - .iter() - .map(|s| proto::PlanStepMsg { - step: s.step as i32, - operation: s.operation.clone(), - modality: s.modality.to_string(), - time_ms: s.cost.time_ms, - estimated_rows: s.cost.estimated_rows, - selectivity: s.cost.selectivity, - optimization_hint: s.optimization_hint.clone().unwrap_or_default(), - }) - .collect(); - - Ok(Response::new(proto::PhysicalPlanResponse { - steps, - strategy: format!("{:?}", physical.strategy), - total_time_ms: physical.total_cost.time_ms, - total_estimated_rows: physical.total_cost.estimated_rows, - notes: physical.notes.clone(), - })) - } - - async fn explain_plan( - &self, - request: Request, - ) -> Result, Status> { - let req = request.into_inner(); - let logical: LogicalPlan = serde_json::from_str(&req.plan_json) - .map_err(|e| Status::invalid_argument(format!("Invalid plan JSON: {}", e)))?; - - let planner = self.state.planner.lock() - .map_err(|_| { error!("Planner lock poisoned in gRPC"); Status::internal("Internal server error") })?; - let explain = planner - .explain(&logical) - .map_err(|e| Status::invalid_argument(e.to_string()))?; - - let steps: Vec = explain - .steps - .iter() - .map(|s| proto::PlanStepMsg { - step: s.step as i32, - operation: s.operation.clone(), - modality: s.modality.to_string(), - time_ms: s.estimated_cost_ms, - estimated_rows: s.estimated_rows, - selectivity: s.estimated_selectivity, - optimization_hint: s.optimization_hint.clone().unwrap_or_default(), - }) - .collect(); - - let cost_breakdown: Vec = explain - .cost_breakdown - .iter() - .map(|cb| proto::ModalityCostMsg { - modality: cb.modality.to_string(), - time_ms: cb.time_ms, - percentage: cb.percentage, - }) - .collect(); - - let hints: Vec = explain - .performance_hints - .iter() - .map(|h| proto::PerformanceHintMsg { - severity: h.severity.clone(), - message: h.message.clone(), - }) - .collect(); - - Ok(Response::new(proto::ExplainResponse { - steps, - cost_breakdown, - performance_hints: hints, - total_cost_ms: explain.total_cost_ms, - strategy: explain.strategy.clone(), - text_output: explain.text_output.clone(), - })) - } - - async fn get_config( - &self, - _request: Request, - ) -> Result, Status> { - let planner = self.state.planner.lock() - .map_err(|_| { error!("Planner lock poisoned in gRPC"); Status::internal("Internal server error") })?; - let cfg = planner.config(); - - Ok(Response::new(proto::PlannerConfigResponse { - global_mode: format!("{:?}", cfg.global_mode), - statistics_weight: cfg.statistics_weight, - enable_adaptive: cfg.enable_adaptive, - parallel_threshold: cfg.parallel_threshold as i32, - })) - } - - async fn set_config( - &self, - request: Request, - ) -> Result, Status> { - let req = request.into_inner(); - let mut planner = self.state.planner.lock() - .map_err(|_| { error!("Planner lock poisoned in gRPC"); Status::internal("Internal server error") })?; - - let mut cfg = planner.config().clone(); - if !req.global_mode.is_empty() { - cfg.global_mode = match req.global_mode.to_lowercase().as_str() { - "conservative" => verisim_planner::OptimizationMode::Conservative, - "aggressive" => verisim_planner::OptimizationMode::Aggressive, - _ => verisim_planner::OptimizationMode::Balanced, - }; - } - if req.statistics_weight > 0.0 { - cfg.statistics_weight = req.statistics_weight; - } - cfg.enable_adaptive = req.enable_adaptive; - if req.parallel_threshold > 0 { - cfg.parallel_threshold = req.parallel_threshold as usize; - } - planner.set_config(cfg); - - let cfg = planner.config(); - Ok(Response::new(proto::PlannerConfigResponse { - global_mode: format!("{:?}", cfg.global_mode), - statistics_weight: cfg.statistics_weight, - enable_adaptive: cfg.enable_adaptive, - parallel_threshold: cfg.parallel_threshold as i32, - })) - } - - async fn get_stats( - &self, - _request: Request, - ) -> Result, Status> { - let planner = self.state.planner.lock() - .map_err(|_| { error!("Planner lock poisoned in gRPC"); Status::internal("Internal server error") })?; - - let stores: Vec = verisim_planner::Modality::ALL - .iter() - .filter_map(|m| { - planner.stats().get(*m).map(|s| proto::StoreStatsMsg { - modality: m.to_string(), - total_rows: s.total_rows, - avg_latency_ms: s.avg_latency_ms, - avg_rows_returned: s.avg_rows_returned, - query_count: s.query_count, - }) - }) - .collect(); - - Ok(Response::new(proto::StatsResponse { stores })) - } -} - -// ============================================================================ -// Octad gRPC Service -// ============================================================================ - -pub struct OctadService { - state: AppState, -} - -impl OctadService { - pub fn new(state: AppState) -> Self { - Self { state } - } -} - -#[tonic::async_trait] -impl VeriSimOctad for OctadService { - async fn create( - &self, - request: Request, - ) -> Result, Status> { - let req = request.into_inner(); - let mut input = verisim_octad::OctadInput::default(); - - if !req.title.is_empty() { - input.document = Some(verisim_octad::OctadDocumentInput { - title: req.title, - body: req.body, - fields: std::collections::HashMap::new(), - }); - } - if !req.embedding.is_empty() { - input.vector = Some(verisim_octad::OctadVectorInput { - embedding: req.embedding, - model: None, - }); - } - if !req.types.is_empty() { - input.semantic = Some(verisim_octad::OctadSemanticInput { - types: req.types, - properties: std::collections::HashMap::new(), - }); - } - - use verisim_octad::OctadStore; - let h = self - .state - .octad_store - .create(input) - .await - .map_err(|e| { - error!(error = %e, "gRPC octad creation failed"); - Status::internal("Internal server error") - })?; - - Ok(Response::new(octad_to_proto(&h))) - } - - async fn get( - &self, - request: Request, - ) -> Result, Status> { - let id = request.into_inner().id; - let octad_id = verisim_octad::OctadId::new(&id); - - use verisim_octad::OctadStore; - let h = self - .state - .octad_store - .get(&octad_id) - .await - .map_err(|e| { - error!(error = %e, "gRPC octad get failed"); - Status::internal("Internal server error") - })? - .ok_or_else(|| Status::not_found(format!("Octad {} not found", id)))?; - - Ok(Response::new(octad_to_proto(&h))) - } - - async fn update( - &self, - request: Request, - ) -> Result, Status> { - let req = request.into_inner(); - let octad_id = verisim_octad::OctadId::new(&req.id); - - let mut input = verisim_octad::OctadInput::default(); - if !req.title.is_empty() { - input.document = Some(verisim_octad::OctadDocumentInput { - title: req.title, - body: req.body, - fields: std::collections::HashMap::new(), - }); - } - if !req.embedding.is_empty() { - input.vector = Some(verisim_octad::OctadVectorInput { - embedding: req.embedding, - model: None, - }); - } - if !req.types.is_empty() { - input.semantic = Some(verisim_octad::OctadSemanticInput { - types: req.types, - properties: std::collections::HashMap::new(), - }); - } - - use verisim_octad::OctadStore; - let h = self - .state - .octad_store - .update(&octad_id, input) - .await - .map_err(|e| { - error!(error = %e, "gRPC octad update failed"); - Status::internal("Internal server error") - })?; - - Ok(Response::new(octad_to_proto(&h))) - } - - async fn delete( - &self, - request: Request, - ) -> Result, Status> { - let id = request.into_inner().id; - let octad_id = verisim_octad::OctadId::new(&id); - - use verisim_octad::OctadStore; - self.state - .octad_store - .delete(&octad_id) - .await - .map_err(|e| { - error!(error = %e, "gRPC octad deletion failed"); - Status::internal("Internal server error") - })?; - - Ok(Response::new(proto::Empty {})) - } - - async fn search_text( - &self, - request: Request, - ) -> Result, Status> { - let req = request.into_inner(); - let limit = if req.limit > 0 { req.limit as usize } else { 10 }; - - use verisim_octad::OctadStore; - let octads = self - .state - .octad_store - .search_text(&req.query, limit) - .await - .map_err(|e| { - error!(error = %e, "gRPC text search failed"); - Status::internal("Internal server error") - })?; - - let results: Vec = octads - .iter() - .enumerate() - .map(|(i, h)| proto::SearchResultMsg { - id: h.id.to_string(), - score: 1.0 - (i as f32 * 0.1), - title: h - .document - .as_ref() - .map(|d| d.title.clone()) - .unwrap_or_default(), - }) - .collect(); - - Ok(Response::new(proto::SearchResponse { results })) - } - - async fn search_vector( - &self, - request: Request, - ) -> Result, Status> { - let req = request.into_inner(); - let k = if req.k > 0 { req.k as usize } else { 10 }; - - use verisim_octad::OctadStore; - let octads = self - .state - .octad_store - .search_similar(&req.vector, k) - .await - .map_err(|e| { - error!(error = %e, "gRPC vector search failed"); - Status::internal("Internal server error") - })?; - - let results: Vec = octads - .iter() - .enumerate() - .map(|(i, h)| proto::SearchResultMsg { - id: h.id.to_string(), - score: 1.0 - (i as f32 * 0.1), - title: h - .document - .as_ref() - .map(|d| d.title.clone()) - .unwrap_or_default(), - }) - .collect(); - - Ok(Response::new(proto::SearchResponse { results })) - } -} - -// ============================================================================ -// Helpers -// ============================================================================ - -fn octad_to_proto(h: &verisim_octad::Octad) -> proto::OctadResponse { - proto::OctadResponse { - id: h.id.to_string(), - created_at: h.status.created_at.to_rfc3339(), - modified_at: h.status.modified_at.to_rfc3339(), - version: h.status.version, - has_graph: h.graph_node.is_some(), - has_vector: h.embedding.is_some(), - has_tensor: h.tensor.is_some(), - has_semantic: h.semantic.is_some(), - has_document: h.document.is_some(), - version_count: h.version_count, - } -} - -/// Build gRPC server routes. Returns a tonic Router that can be served. -pub fn build_grpc_router(state: AppState) -> tonic::transport::server::Router { - let planner_svc = VeriSimPlannerServer::new(PlannerService::new(state.clone())); - let octad_svc = VeriSimOctadServer::new(OctadService::new(state)); - - tonic::transport::Server::builder() - .add_service(planner_svc) - .add_service(octad_svc) -} diff --git a/verisimdb/rust-core/verisim-api/src/lib.rs b/verisimdb/rust-core/verisim-api/src/lib.rs deleted file mode 100644 index 3f827aa3..00000000 --- a/verisimdb/rust-core/verisim-api/src/lib.rs +++ /dev/null @@ -1,2744 +0,0 @@ -// SPDX-License-Identifier: MPL-2.0 -//! VeriSim API -//! -//! HTTP API server for VeriSimDB. -//! Exposes all database functionality via REST endpoints. - -#![forbid(unsafe_code)] -pub mod a2ml; -pub mod auth; -pub mod federation; -pub mod graphql; -pub mod groove; -pub mod grpc; -pub mod proof_attempts; -pub mod rbac; -pub mod transaction; -pub mod vcl; - -use axum::{ - extract::{Path, Query, State}, - http::StatusCode, - middleware as axum_middleware, - response::IntoResponse, - routing::{delete, get, post, put}, - Json, Router, -}; -use serde::{Deserialize, Serialize}; -use std::sync::Arc; -use thiserror::Error; -use tokio::net::TcpListener; -use tracing::{error, info, instrument}; - -use std::sync::Mutex; - -use verisim_document::TantivyDocumentStore; -use verisim_drift::{DriftDetector, DriftMetrics, DriftThresholds, DriftType}; -#[cfg(not(feature = "persistent"))] -use verisim_graph::SimpleGraphStore; -#[cfg(feature = "persistent")] -use verisim_graph::RedbGraphStore; -use verisim_planner::{ - CacheConfig, ExplainOutput, ExplainAnalyzeOutput, LogicalPlan, ParamValue, - PhysicalPlan, PlanCache, Planner, PlannerConfig, PreparedId, PreparedStatement, - Profiler, SlowQueryLog, SlowQuerySummary, StatisticsCollector, -}; -use verisim_octad::{ - BoundingBox, Coordinates, OctadConfig, OctadDocumentInput, OctadGraphInput, - OctadId, OctadInput, OctadProvenanceInput, OctadSemanticInput, OctadSnapshot, - OctadSpatialInput, OctadStore, OctadTemporalInput, OctadTensorInput, - OctadVectorInput, InMemoryOctadStore, ProvenanceStore, SpatialStore, -}; -use verisim_provenance::InMemoryProvenanceStore; -use verisim_spatial::InMemorySpatialStore; -use verisim_normalizer::{create_default_normalizer, Normalizer, NormalizerStatus}; -use verisim_semantic::InMemorySemanticStore; -use verisim_semantic::zkp_bridge::{self as zkp_api, PrivacyLevel, ZkpProofRequest as ZkpBridgeRequest}; -use verisim_semantic::circuit_registry::CircuitRegistry; -use verisim_temporal::InMemoryVersionStore; -use verisim_tensor::InMemoryTensorStore; -use verisim_vector::{DistanceMetric, BruteForceVectorStore}; - -/// Type alias for our concrete OctadStore implementation (octad: 8 modality stores). -/// -/// When the `persistent` feature is enabled, the graph store uses redb (pure Rust, -/// ACID, single-file B-tree) and the document store uses file-backed Tantivy. -/// WAL is enabled for crash recovery. Requires `VERISIM_PERSISTENCE_DIR` at runtime. -#[cfg(not(feature = "persistent"))] -pub type ConcreteOctadStore = InMemoryOctadStore< - SimpleGraphStore, - BruteForceVectorStore, - TantivyDocumentStore, - InMemoryTensorStore, - InMemorySemanticStore, - InMemoryVersionStore, - InMemoryProvenanceStore, - InMemorySpatialStore, ->; - -/// Persistent variant: redb graph store, file-backed Tantivy, WAL enabled. -#[cfg(feature = "persistent")] -pub type ConcreteOctadStore = InMemoryOctadStore< - RedbGraphStore, - BruteForceVectorStore, - TantivyDocumentStore, - InMemoryTensorStore, - InMemorySemanticStore, - InMemoryVersionStore, - InMemoryProvenanceStore, - InMemorySpatialStore, ->; - -/// API errors -#[derive(Error, Debug)] -pub enum ApiError { - #[error("Not found: {0}")] - NotFound(String), - - #[error("Bad request: {0}")] - BadRequest(String), - - #[error("Internal error: {0}")] - Internal(String), - - #[error("Serialization error: {0}")] - Serialization(String), -} - -impl IntoResponse for ApiError { - fn into_response(self) -> axum::response::Response { - let (status, client_message) = match &self { - ApiError::NotFound(msg) => (StatusCode::NOT_FOUND, msg.clone()), - ApiError::BadRequest(msg) => (StatusCode::BAD_REQUEST, msg.clone()), - ApiError::Internal(msg) => { - error!(error = %msg, "Internal server error"); - (StatusCode::INTERNAL_SERVER_ERROR, "Internal server error".to_string()) - } - ApiError::Serialization(msg) => { - error!(error = %msg, "Serialization error"); - (StatusCode::INTERNAL_SERVER_ERROR, "Internal server error".to_string()) - } - }; - - let body = Json(ErrorResponse { - error: client_message, - code: status.as_u16(), - }); - - (status, body).into_response() - } -} - -/// Error response body -#[derive(Debug, Serialize, Deserialize)] -pub struct ErrorResponse { - pub error: String, - pub code: u16, -} - -/// API configuration -#[derive(Debug, Clone, Serialize, Deserialize)] -pub struct ApiConfig { - /// Host to bind to - pub host: String, - /// Port to bind to (HTTP) - pub port: u16, - /// Port for gRPC server (default: 50051). Set to 0 to disable gRPC. - pub grpc_port: u16, - /// Enable CORS - pub enable_cors: bool, - /// API version prefix - pub version_prefix: String, - /// Vector dimension for embeddings - pub vector_dimension: usize, - /// Persistence directory for the `persistent` feature. - /// Overrides `VERISIM_PERSISTENCE_DIR` env var when set. - pub persistence_dir: Option, - /// Maximum request body size in bytes (default: 10MB) - pub max_body_size: usize, - /// Request timeout in seconds (default: 30) - pub request_timeout_secs: u64, - /// Maximum concurrent connections (default: 1024) - pub max_connections: usize, -} - -impl Default for ApiConfig { - fn default() -> Self { - Self { - host: "[::1]".to_string(), - port: 8080, - grpc_port: 50051, - enable_cors: true, - version_prefix: "/api/v1".to_string(), - vector_dimension: 384, - persistence_dir: None, - max_body_size: 10 * 1024 * 1024, // 10MB - request_timeout_secs: 30, - max_connections: 1024, - } - } -} - -/// Maximum number of results allowed in any search/list endpoint. -const MAX_RESULT_LIMIT: usize = 1000; - -/// Validate and cap a limit parameter. -fn validate_limit(limit: usize) -> usize { - limit.min(MAX_RESULT_LIMIT) -} - -/// Validate a octad ID: max 128 chars, alphanumeric + dash + underscore only. -fn validate_octad_id(id: &str) -> Result<(), ApiError> { - if id.is_empty() { - return Err(ApiError::BadRequest("Octad ID must not be empty".to_string())); - } - if id.len() > 128 { - return Err(ApiError::BadRequest("Octad ID must be at most 128 characters".to_string())); - } - if !id.chars().all(|c| c.is_alphanumeric() || c == '-' || c == '_') { - return Err(ApiError::BadRequest( - "Octad ID must contain only alphanumeric characters, dashes, and underscores".to_string(), - )); - } - Ok(()) -} - -/// Validate that all vector components are finite (no NaN/Inf). -fn validate_vector(v: &[f32]) -> Result<(), ApiError> { - if !v.iter().all(|x| x.is_finite()) { - return Err(ApiError::BadRequest( - "Vector contains NaN or Inf values".to_string(), - )); - } - Ok(()) -} - -/// Health check response -#[derive(Debug, Serialize, Deserialize)] -pub struct HealthResponse { - pub status: String, - pub version: String, - pub uptime_seconds: u64, - #[serde(skip_serializing_if = "Option::is_none")] - pub degraded_reason: Option, -} - -/// Octad create/update request -#[derive(Debug, Serialize, Deserialize)] -pub struct OctadRequest { - /// Document title - pub title: Option, - /// Document body - pub body: Option, - /// Vector embedding - pub embedding: Option>, - /// Semantic types - pub types: Option>, - /// Relationships (predicate, target_id) - pub relationships: Option>, - /// Tensor data - pub tensor: Option, - /// Temporal observation (real-world clock, distinct from ingestion time) - pub temporal: Option, - /// Provenance event - pub provenance: Option, - /// Spatial coordinates - pub spatial: Option, - /// Metadata - pub metadata: Option>, -} - -/// Temporal observation in request — caller-supplied real-world time -#[derive(Debug, Serialize, Deserialize)] -pub struct TemporalRequest { - /// RFC 3339 timestamp of the underlying real-world event (e.g. an email's - /// `Date:` header). Stored on the octad as `OctadStatus.observed_at`. - pub observed_at: String, -} - -/// Provenance event data in request -#[derive(Debug, Serialize, Deserialize)] -pub struct ProvenanceRequest { - /// Event type: created, modified, imported, normalized, drift_repaired, deleted, merged - pub event_type: String, - /// Who or what caused this event - pub actor: String, - /// Optional source identifier - pub source: Option, - /// Human-readable description - pub description: String, -} - -/// Spatial coordinates in request -#[derive(Debug, Serialize, Deserialize)] -pub struct SpatialRequest { - /// Latitude (WGS84, -90 to 90) - pub latitude: f64, - /// Longitude (WGS84, -180 to 180) - pub longitude: f64, - /// Altitude in metres (optional) - pub altitude: Option, - /// Geometry type (defaults to Point) - pub geometry_type: Option, - /// SRID (defaults to 4326) - pub srid: Option, - /// Spatial properties - pub properties: Option>, -} - -impl OctadRequest { - /// Convert to OctadInput. Fallible because the temporal field is supplied - /// as RFC 3339 text and may fail to parse — surface that as 400 Bad Request - /// rather than swallowing it as a missing observation. - fn to_octad_input(&self) -> Result { - let mut input = OctadInput::default(); - - if let (Some(title), Some(body)) = (&self.title, &self.body) { - input.document = Some(OctadDocumentInput { - title: title.clone(), - body: body.clone(), - fields: std::collections::HashMap::new(), - }); - } else if let Some(title) = &self.title { - input.document = Some(OctadDocumentInput { - title: title.clone(), - body: String::new(), - fields: std::collections::HashMap::new(), - }); - } - - if let Some(embedding) = &self.embedding { - input.vector = Some(OctadVectorInput { - embedding: embedding.clone(), - model: None, - }); - } - - if let Some(types) = &self.types { - input.semantic = Some(OctadSemanticInput { - types: types.clone(), - properties: std::collections::HashMap::new(), - }); - } - - if let Some(relationships) = &self.relationships { - input.graph = Some(OctadGraphInput { - relationships: relationships.clone(), - }); - } - - if let Some(tensor) = &self.tensor { - input.tensor = Some(OctadTensorInput { - shape: tensor.shape.clone(), - data: tensor.data.clone(), - }); - } - - if let Some(temporal) = &self.temporal { - let observed_at = chrono::DateTime::parse_from_rfc3339(&temporal.observed_at) - .map_err(|e| { - ApiError::BadRequest(format!( - "temporal.observed_at must be RFC 3339: {e}" - )) - })? - .with_timezone(&chrono::Utc); - input.temporal = Some(OctadTemporalInput { observed_at }); - } - - if let Some(provenance) = &self.provenance { - input.provenance = Some(OctadProvenanceInput { - event_type: provenance.event_type.clone(), - actor: provenance.actor.clone(), - source: provenance.source.clone(), - description: provenance.description.clone(), - }); - } - - if let Some(spatial) = &self.spatial { - input.spatial = Some(OctadSpatialInput { - latitude: spatial.latitude, - longitude: spatial.longitude, - altitude: spatial.altitude, - geometry_type: spatial.geometry_type.clone(), - srid: spatial.srid, - properties: spatial.properties.clone().unwrap_or_default(), - }); - } - - if let Some(metadata) = &self.metadata { - input.metadata = metadata.clone(); - } - - Ok(input) - } -} - -/// Tensor data in request -#[derive(Debug, Serialize, Deserialize)] -pub struct TensorRequest { - pub shape: Vec, - pub data: Vec, -} - -/// Octad response -#[derive(Debug, Serialize, Deserialize)] -pub struct OctadResponse { - pub id: String, - pub status: OctadStatusResponse, - pub has_graph: bool, - pub has_vector: bool, - pub has_tensor: bool, - pub has_semantic: bool, - pub has_document: bool, - pub has_provenance: bool, - pub has_spatial: bool, - pub version_count: u64, - pub provenance_chain_length: u64, - /// Populated only when the caller passes `?include=types` on a GET. Listed - /// in the order the entity declares them. Cheap to surface but opt-in so - /// the default response shape stays small. - #[serde(default, skip_serializing_if = "Option::is_none")] - pub semantic_types: Option>, - /// Populated only when the caller passes `?include=embedding` on a GET. - /// The stored vector is returned verbatim (same dimension, same - /// component values) so the Vector shape is byte-exactly verifiable. - #[serde(default, skip_serializing_if = "Option::is_none")] - pub embedding: Option>, -} - -/// Per-request flags controlling which expensive or large fields the GET -/// /octads/{id} response should include. Driven by `?include=…` (a -/// comma-separated list). -#[derive(Debug, Default, Clone, Copy)] -pub struct IncludeFlags { - /// Surface `semantic_types` (the entity's declared type IRIs) - pub types: bool, - /// Surface the stored vector bytes (Phase 1 Step 3 — Task #12) - pub embedding: bool, -} - -impl IncludeFlags { - /// Parse a `?include=` value (comma-separated). Unknown tokens are ignored - /// rather than errored — the API's forward-compatibility contract is that - /// a future client may request includes the current server doesn't know - /// about. - pub fn parse(raw: Option<&str>) -> Self { - let mut flags = Self::default(); - if let Some(s) = raw { - for token in s.split(',').map(|t| t.trim()).filter(|t| !t.is_empty()) { - match token { - "types" => flags.types = true, - "embedding" => flags.embedding = true, - _ => {} - } - } - } - flags - } -} - -/// Query parameters for `GET /octads/{id}` -#[derive(Debug, Deserialize)] -pub struct OctadGetQuery { - /// Comma-separated list of optional includes — currently `types`, - /// `embedding`. Unknown tokens are ignored. - pub include: Option, -} - -/// Status response -#[derive(Debug, Serialize, Deserialize)] -pub struct OctadStatusResponse { - pub created_at: String, - pub modified_at: String, - /// RFC 3339 territory clock — the real-world time the entity was observed, - /// distinct from `created_at` (the database ingestion time). Omitted from - /// the JSON when the entity has no caller-supplied observation time. - #[serde(skip_serializing_if = "Option::is_none")] - pub observed_at: Option, - pub version: u64, -} - -impl From<&verisim_octad::Octad> for OctadResponse { - fn from(h: &verisim_octad::Octad) -> Self { - Self::with_includes(h, IncludeFlags::default()) - } -} - -impl OctadResponse { - /// Build a response from an octad, honouring per-request include flags - /// for opt-in fields (semantic types, vector bytes, …). - pub fn with_includes(h: &verisim_octad::Octad, flags: IncludeFlags) -> Self { - let semantic_types = if flags.types { - Some( - h.semantic - .as_ref() - .map(|s| s.types.clone()) - .unwrap_or_default(), - ) - } else { - None - }; - let embedding = if flags.embedding { - h.embedding.as_ref().map(|e| e.vector.clone()) - } else { - None - }; - Self { - id: h.id.to_string(), - status: OctadStatusResponse { - created_at: h.status.created_at.to_rfc3339(), - modified_at: h.status.modified_at.to_rfc3339(), - observed_at: h.status.observed_at.map(|t| t.to_rfc3339()), - version: h.status.version, - }, - has_graph: h.graph_node.is_some(), - has_vector: h.embedding.is_some(), - has_tensor: h.tensor.is_some(), - has_semantic: h.semantic.is_some(), - has_document: h.document.is_some(), - has_provenance: h.provenance_chain_length > 0, - has_spatial: h.spatial_data.is_some(), - version_count: h.version_count, - provenance_chain_length: h.provenance_chain_length, - semantic_types, - embedding, - } - } -} - -/// Pagination query parameters for list endpoints -#[derive(Debug, Deserialize)] -pub struct ListQuery { - /// Maximum number of results (default 100, max 1000) - pub limit: Option, - /// Offset for pagination (default 0) - pub offset: Option, -} - -/// Search query parameters -#[derive(Debug, Deserialize)] -pub struct SearchQuery { - /// Text query for document search - pub q: Option, - /// Number of results - pub limit: Option, -} - -/// Vector search request -#[derive(Debug, Serialize, Deserialize)] -pub struct VectorSearchRequest { - /// Query vector - pub vector: Vec, - /// Number of results - pub k: Option, -} - -/// Search result -#[derive(Debug, Serialize, Deserialize)] -pub struct SearchResultResponse { - pub id: String, - pub score: f32, - pub title: Option, -} - -/// Drift status response -#[derive(Debug, Serialize, Deserialize)] -pub struct DriftStatusResponse { - pub drift_type: String, - pub current_score: f64, - pub moving_average: f64, - pub max_score: f64, - pub measurement_count: u64, -} - -impl DriftStatusResponse { - fn from_metrics(drift_type: DriftType, metrics: &DriftMetrics) -> Self { - Self { - drift_type: drift_type.to_string(), - current_score: metrics.current_score, - moving_average: metrics.moving_average, - max_score: metrics.max_score, - measurement_count: metrics.measurement_count, - } - } -} - -/// Application state -#[derive(Clone)] -pub struct AppState { - pub start_time: std::time::Instant, - pub octad_store: Arc, - pub drift_detector: Arc, - pub normalizer: Arc, - pub planner: Arc>, - pub plan_cache: Arc, - pub slow_query_log: Arc, - pub transaction_manager: Arc, - pub circuit_registry: Arc, - pub federation: federation::FederationState, - pub auth: auth::AuthState, - pub config: ApiConfig, -} - -impl AppState { - /// Create new application state with default configuration (async version). - /// - /// With the `persistent` feature enabled, reads `VERISIM_PERSISTENCE_DIR` - /// to determine where to store data on disk. Defaults to `/var/lib/verisimdb` - /// if the variable is unset. - pub async fn new_async(config: ApiConfig) -> Result { - let octad_config = OctadConfig { - vector_dimension: config.vector_dimension, - ..Default::default() - }; - - // --- In-memory stores (default, no `persistent` feature) --- - #[cfg(not(feature = "persistent"))] - let (graph, document) = { - let g = Arc::new( - SimpleGraphStore::in_memory().map_err(|e| ApiError::Internal(e.to_string()))?, - ); - let d = Arc::new( - TantivyDocumentStore::in_memory() - .map_err(|e| ApiError::Internal(e.to_string()))?, - ); - (g, d) - }; - - // --- Persistent stores (with `persistent` feature) --- - #[cfg(feature = "persistent")] - let persist_dir = config - .persistence_dir - .clone() - .or_else(|| std::env::var("VERISIM_PERSISTENCE_DIR").ok()) - .unwrap_or_else(|| "/var/lib/verisimdb".to_string()); - - #[cfg(feature = "persistent")] - let (graph, document) = { - std::fs::create_dir_all(&persist_dir) - .map_err(|e| ApiError::Internal(format!("create persistence dir: {e}")))?; - - info!(dir = %persist_dir, "Persistent storage enabled"); - - let g = Arc::new( - RedbGraphStore::persistent(format!("{}/graph.redb", persist_dir)) - .map_err(|e| ApiError::Internal(e.to_string()))?, - ); - let d = Arc::new( - TantivyDocumentStore::persistent(format!("{}/documents", persist_dir)) - .map_err(|e| ApiError::Internal(e.to_string()))?, - ); - (g, d) - }; - - let vector = Arc::new(BruteForceVectorStore::new( - config.vector_dimension, - DistanceMetric::Cosine, - )); - let tensor = Arc::new(InMemoryTensorStore::new()); - let semantic = Arc::new(InMemorySemanticStore::new()); - let temporal = Arc::new(InMemoryVersionStore::new()); - let provenance = Arc::new(InMemoryProvenanceStore::new()); - let spatial = Arc::new(InMemorySpatialStore::new()); - - let octad_store_inner = InMemoryOctadStore::new( - octad_config, - graph, - vector, - document, - tensor, - semantic, - temporal, - provenance, - spatial, - ); - - // Enable WAL for crash recovery when persistent. - #[cfg(feature = "persistent")] - let octad_store_inner = octad_store_inner - .with_wal( - format!("{}/wal", persist_dir), - verisim_octad::SyncMode::Fsync, - ) - .map_err(|e| ApiError::Internal(format!("WAL init: {e}")))?; - - // Replay WAL to recover octad status registry after crash. - // Modality data is already in redb (loaded by persistent store constructors). - // This rebuilds the in-memory octad registry from WAL entries. - #[cfg(feature = "persistent")] - { - let wal_dir = format!("{}/wal", persist_dir); - match octad_store_inner.replay_wal(&wal_dir).await { - Ok(0) => info!("WAL replay: clean start (no entries to replay)"), - Ok(n) => info!(recovered = n, "WAL replay: recovered {} entities", n), - Err(e) => tracing::warn!("WAL replay failed (non-fatal): {e}"), - } - } - - let octad_store = Arc::new(octad_store_inner); - - let drift_detector = Arc::new(DriftDetector::new(DriftThresholds::default())); - let normalizer = Arc::new(create_default_normalizer(drift_detector.clone()).await); - - let planner = Arc::new(Mutex::new(Planner::new(PlannerConfig::default()))); - let plan_cache = Arc::new(PlanCache::new(CacheConfig::default())); - let slow_query_log = Arc::new(SlowQueryLog::new(Default::default())); - let transaction_manager = Arc::new( - transaction::TransactionManager::new(transaction::TransactionConfig::default()), - ); - - let self_endpoint = format!("http://{}:{}{}", config.host, config.port, config.version_prefix); - let federation = federation::FederationState::new( - "self".to_string(), - self_endpoint, - ); - - let auth = auth::AuthState::default(); - let circuit_registry = Arc::new(CircuitRegistry::new()); - - Ok(Self { - start_time: std::time::Instant::now(), - octad_store, - drift_detector, - normalizer, - planner, - plan_cache, - slow_query_log, - transaction_manager, - circuit_registry, - federation, - auth, - config, - }) - } -} - -/// Build the API router -pub fn build_router(state: AppState) -> Router { - let federation_routes = federation::federation_router(state.federation.clone()); - let auth_state = state.auth.clone(); - - Router::new() - // Health endpoints - .route("/health", get(health_handler)) - .route("/ready", get(ready_handler)) - .route("/metrics", get(metrics_handler)) - // Octad CRUD - .route("/octads", get(list_octads_handler).post(create_octad_handler)) - .route("/octads/{id}", get(get_octad_handler)) - .route("/octads/{id}", put(update_octad_handler)) - .route("/octads/{id}", delete(delete_octad_handler)) - // Search endpoints - .route("/search/text", get(text_search_handler)) - .route("/search/vector", post(vector_search_handler)) - .route("/search/related/{id}", get(related_search_handler)) - // Drift and normalization - .route("/drift/status", get(drift_status_handler)) - .route("/drift/entity/{id}", get(entity_drift_handler)) - .route("/normalizer/status", get(normalizer_status_handler)) - .route("/normalizer/trigger/{id}", post(trigger_normalization_handler)) - // Meta-query store (homoiconicity: queries as octads) - .route("/queries", post(store_query_handler)) - .route("/queries/similar", post(similar_queries_handler)) - .route("/queries/{id}/optimize", put(optimize_query_handler)) - // Query planner - .route("/query/plan", post(query_plan_handler)) - .route("/query/explain", post(query_explain_handler)) - .route("/planner/config", get(get_planner_config_handler)) - .route("/planner/config", put(put_planner_config_handler)) - .route("/planner/stats", get(planner_stats_handler)) - // EXPLAIN ANALYZE - .route("/query/explain-analyze", post(query_explain_analyze_handler)) - // Prepared statements - .route("/prepared", post(prepared_create_handler)) - .route("/prepared/{id}", get(prepared_get_handler)) - .route("/prepared/{id}/execute", post(prepared_execute_handler)) - .route("/prepared/stats", get(prepared_stats_handler)) - // Slow query log - .route("/planner/slow-queries", get(slow_queries_handler)) - // Transaction endpoints - .route("/transactions/begin", post(transaction_begin_handler)) - .route("/transactions/{id}/commit", post(transaction_commit_handler)) - .route("/transactions/{id}/rollback", post(transaction_rollback_handler)) - .route("/transactions/{id}", get(transaction_status_handler)) - // ZKP proof endpoints - .route("/proofs/generate", post(proof_generate_handler)) - .route("/proofs/verify", post(proof_verify_handler)) - .route("/proofs/generate-with-circuit", post(proof_generate_with_circuit_handler)) - // Provenance endpoints - .route("/provenance/{id}", get(provenance_get_chain_handler)) - .route("/provenance/{id}/record", post(provenance_record_handler)) - .route("/provenance/{id}/verify", get(provenance_verify_handler)) - // Spatial search endpoints - .route("/spatial/search/radius", post(spatial_radius_search_handler)) - .route("/spatial/search/bounds", post(spatial_bounds_search_handler)) - .route("/spatial/search/nearest", post(spatial_nearest_handler)) - // VCL text query endpoint (used by verisim-repl) - .route("/vcl/execute", post(vcl::vcl_execute_handler)) - // Proof-attempts pipeline (ClickHouse-backed, versioned under /api/v1/) - .route( - "/api/v1/proof_attempts", - get(proof_attempts::list_proof_attempts).post(proof_attempts::insert_proof_attempt), - ) - .route( - "/api/v1/proof_attempts/strategy", - get(proof_attempts::strategy), - ) - .route( - "/api/v1/proof_attempts/certificates", - get(proof_attempts::certificates), - ) - // Authentication middleware layer - .layer(axum_middleware::from_fn_with_state( - auth_state, - auth::auth_middleware, - )) - .with_state(state.clone()) - // GraphQL endpoint - .merge(graphql::graphql_router(state)) - // Federation endpoints (separate state) - .merge(federation_routes) - // Groove Protocol connection lifecycle (separate state — lightweight, - // independent of the database layer). Per spec section 4. - .merge(groove::groove_router()) -} - -/// Health check handler — verifies drift detector status and reports degraded when critical -#[instrument(skip(state))] -async fn health_handler(State(state): State) -> (StatusCode, Json) { - let uptime = state.start_time.elapsed().as_secs(); - let version = env!("CARGO_PKG_VERSION").to_string(); - - // Check drift detector health - match state.drift_detector.health_check() { - Ok(health) => { - use verisim_drift::HealthStatus; - let (status_str, reason) = match health.status { - HealthStatus::Critical => ( - "degraded", - Some(format!( - "Critical drift on {:?}: score {:.3}", - health.worst_drift_type, health.worst_score - )), - ), - HealthStatus::Degraded => ( - "degraded", - Some(format!( - "Degraded drift on {:?}: score {:.3}", - health.worst_drift_type, health.worst_score - )), - ), - HealthStatus::Warning => ("healthy", None), - HealthStatus::Healthy => ("healthy", None), - }; - - ( - StatusCode::OK, - Json(HealthResponse { - status: status_str.to_string(), - version, - uptime_seconds: uptime, - degraded_reason: reason, - }), - ) - } - Err(_) => ( - StatusCode::OK, - Json(HealthResponse { - status: "degraded".to_string(), - version, - uptime_seconds: uptime, - degraded_reason: Some("Drift detector unavailable".to_string()), - }), - ), - } -} - -/// Prometheus metrics handler — exposes drift and query metrics for scraping -#[instrument(skip(state))] -async fn metrics_handler( - State(state): State, -) -> Result<(StatusCode, [(axum::http::header::HeaderName, &'static str); 1], String), ApiError> { - use prometheus::{Encoder, TextEncoder, GaugeVec, Opts, Registry}; - - let registry = Registry::new(); - - // Drift gauges - let drift_gauge = GaugeVec::new( - Opts::new("verisimdb_drift_score", "Current drift score by type"), - &["drift_type"], - ) - .map_err(|e| ApiError::Internal(e.to_string()))?; - - let drift_avg_gauge = GaugeVec::new( - Opts::new("verisimdb_drift_moving_average", "Drift moving average by type"), - &["drift_type"], - ) - .map_err(|e| ApiError::Internal(e.to_string()))?; - - let drift_count_gauge = GaugeVec::new( - Opts::new("verisimdb_drift_measurement_count", "Drift measurement count by type"), - &["drift_type"], - ) - .map_err(|e| ApiError::Internal(e.to_string()))?; - - registry.register(Box::new(drift_gauge.clone())).map_err(|e| ApiError::Internal(e.to_string()))?; - registry.register(Box::new(drift_avg_gauge.clone())).map_err(|e| ApiError::Internal(e.to_string()))?; - registry.register(Box::new(drift_count_gauge.clone())).map_err(|e| ApiError::Internal(e.to_string()))?; - - // Populate drift metrics - if let Ok(all_metrics) = state.drift_detector.all_metrics() { - for (drift_type, metrics) in &all_metrics { - let label = drift_type.to_string(); - drift_gauge.with_label_values(&[&label]).set(metrics.current_score); - drift_avg_gauge.with_label_values(&[&label]).set(metrics.moving_average); - drift_count_gauge.with_label_values(&[&label]).set(metrics.measurement_count as f64); - } - } - - // Uptime gauge - let uptime = prometheus::Gauge::new("verisimdb_uptime_seconds", "Server uptime in seconds") - .map_err(|e| ApiError::Internal(e.to_string()))?; - uptime.set(state.start_time.elapsed().as_secs() as f64); - registry.register(Box::new(uptime)).map_err(|e| ApiError::Internal(e.to_string()))?; - - // Encode - let encoder = TextEncoder::new(); - let mut buffer = Vec::new(); - encoder.encode(®istry.gather(), &mut buffer) - .map_err(|e| ApiError::Internal(e.to_string()))?; - - let output = String::from_utf8(buffer) - .map_err(|e| ApiError::Internal(e.to_string()))?; - - Ok(( - StatusCode::OK, - [(axum::http::header::CONTENT_TYPE, "text/plain; version=0.0.4; charset=utf-8")], - output, - )) -} - -/// Readiness check handler — checks octad store accessibility and drift detector health -#[instrument(skip(state))] -async fn ready_handler(State(state): State) -> StatusCode { - // Check octad store is accessible (try a list with limit 0) - if state.octad_store.list(1, 0).await.is_err() { - return StatusCode::SERVICE_UNAVAILABLE; - } - - // Check drift detector is responsive - if state.drift_detector.health_check().is_err() { - return StatusCode::SERVICE_UNAVAILABLE; - } - - StatusCode::OK -} - -/// List octads handler with pagination -#[instrument(skip(state))] -async fn list_octads_handler( - State(state): State, - Query(params): Query, -) -> Result>, ApiError> { - let limit = validate_limit(params.limit.unwrap_or(100)); - let offset = params.offset.unwrap_or(0); - - let octads = state - .octad_store - .list(limit, offset) - .await - .map_err(|e| ApiError::Internal(e.to_string()))?; - - let responses: Vec = octads.iter().map(OctadResponse::from).collect(); - Ok(Json(responses)) -} - -/// Create octad handler -#[instrument(skip(state, request))] -async fn create_octad_handler( - State(state): State, - Json(request): Json, -) -> Result<(StatusCode, Json), ApiError> { - let input = request.to_octad_input()?; - - let octad = state - .octad_store - .create(input) - .await - .map_err(|e| ApiError::Internal(e.to_string()))?; - - Ok((StatusCode::CREATED, Json(OctadResponse::from(&octad)))) -} - -/// Get octad handler -#[instrument(skip(state))] -async fn get_octad_handler( - State(state): State, - Path(id): Path, - Query(query): Query, -) -> Result, ApiError> { - validate_octad_id(&id)?; - let octad_id = OctadId::new(&id); - - let octad = state - .octad_store - .get(&octad_id) - .await - .map_err(|e| ApiError::Internal(e.to_string()))? - .ok_or_else(|| ApiError::NotFound(format!("Octad {} not found", id)))?; - - let flags = IncludeFlags::parse(query.include.as_deref()); - Ok(Json(OctadResponse::with_includes(&octad, flags))) -} - -/// Update octad handler -#[instrument(skip(state, request))] -async fn update_octad_handler( - State(state): State, - Path(id): Path, - Json(request): Json, -) -> Result, ApiError> { - validate_octad_id(&id)?; - let octad_id = OctadId::new(&id); - let input = request.to_octad_input()?; - - let octad = state - .octad_store - .update(&octad_id, input) - .await - .map_err(|e| match e { - verisim_octad::OctadError::NotFound(_) => { - ApiError::NotFound(format!("Octad {} not found", id)) - } - _ => ApiError::Internal(e.to_string()), - })?; - - Ok(Json(OctadResponse::from(&octad))) -} - -/// Delete octad handler -#[instrument(skip(state))] -async fn delete_octad_handler( - State(state): State, - Path(id): Path, -) -> Result { - validate_octad_id(&id)?; - let octad_id = OctadId::new(&id); - - state - .octad_store - .delete(&octad_id) - .await - .map_err(|e| match e { - verisim_octad::OctadError::NotFound(_) => { - ApiError::NotFound(format!("Octad {} not found", id)) - } - _ => ApiError::Internal(e.to_string()), - })?; - - Ok(StatusCode::NO_CONTENT) -} - -/// Text search handler -#[instrument(skip(state))] -async fn text_search_handler( - State(state): State, - Query(query): Query, -) -> Result>, ApiError> { - let q = match query.q { - Some(q) if !q.is_empty() => q, - _ => return Err(ApiError::BadRequest("Query parameter 'q' must not be empty".to_string())), - }; - let limit = validate_limit(query.limit.unwrap_or(10)); - - let octads = state - .octad_store - .search_text(&q, limit) - .await - .map_err(|e| ApiError::Internal(e.to_string()))?; - - let results: Vec = octads - .iter() - .enumerate() - .map(|(i, h)| SearchResultResponse { - id: h.id.to_string(), - score: 1.0 - (i as f32 * 0.1), // Approximate score based on ranking - title: h.document.as_ref().map(|d| d.title.clone()), - }) - .collect(); - - Ok(Json(results)) -} - -/// Vector search handler -#[instrument(skip(state, request))] -async fn vector_search_handler( - State(state): State, - Json(request): Json, -) -> Result>, ApiError> { - let k = validate_limit(request.k.unwrap_or(10)); - - if request.vector.len() != state.config.vector_dimension { - return Err(ApiError::BadRequest(format!( - "Vector dimension mismatch: expected {}, got {}", - state.config.vector_dimension, - request.vector.len() - ))); - } - validate_vector(&request.vector)?; - - let octads = state - .octad_store - .search_similar(&request.vector, k) - .await - .map_err(|e| ApiError::Internal(e.to_string()))?; - - let results: Vec = octads - .iter() - .enumerate() - .map(|(i, h)| SearchResultResponse { - id: h.id.to_string(), - score: 1.0 - (i as f32 * 0.1), // Approximate score based on ranking - title: h.document.as_ref().map(|d| d.title.clone()), - }) - .collect(); - - Ok(Json(results)) -} - -/// Related entities search handler -#[instrument(skip(state))] -async fn related_search_handler( - State(state): State, - Path(id): Path, - Query(query): Query, -) -> Result>, ApiError> { - validate_octad_id(&id)?; - let octad_id = OctadId::new(&id); - let predicate = query.predicate.unwrap_or_else(|| "related".to_string()); - - let octads = state - .octad_store - .query_related(&octad_id, &predicate) - .await - .map_err(|e| ApiError::Internal(e.to_string()))?; - - let responses: Vec = octads.iter().map(OctadResponse::from).collect(); - - Ok(Json(responses)) -} - -/// Query parameters for related search -#[derive(Debug, Deserialize)] -pub struct RelatedQuery { - pub predicate: Option, -} - -/// Drift status handler -#[instrument(skip(state))] -async fn drift_status_handler( - State(state): State, -) -> Result>, ApiError> { - let all_metrics = state.drift_detector.all_metrics() - .map_err(|e| ApiError::Internal(e.to_string()))?; - - let responses: Vec = all_metrics - .iter() - .map(|(drift_type, metrics)| DriftStatusResponse::from_metrics(*drift_type, metrics)) - .collect(); - - Ok(Json(responses)) -} - -/// Entity drift response -#[derive(Debug, Serialize, Deserialize)] -pub struct EntityDriftResponse { - pub entity_id: String, - pub score: f64, - pub drift_type: String, - pub status: String, -} - -/// Entity drift handler — get drift info for a single entity -#[instrument(skip(state))] -async fn entity_drift_handler( - State(state): State, - Path(id): Path, -) -> Result, ApiError> { - validate_octad_id(&id)?; - let octad_id = OctadId::new(&id); - - // Verify octad exists - let _octad = state - .octad_store - .get(&octad_id) - .await - .map_err(|e| ApiError::Internal(e.to_string()))? - .ok_or_else(|| ApiError::NotFound(format!("Octad {} not found", id)))?; - - // Get aggregate health from drift detector - let all_metrics = state.drift_detector.all_metrics() - .map_err(|e| ApiError::Internal(e.to_string()))?; - let (worst_type, worst_score) = all_metrics - .iter() - .max_by(|a, b| a.1.current_score.partial_cmp(&b.1.current_score).unwrap_or(std::cmp::Ordering::Equal)) - .map(|(dt, m)| (dt.to_string(), m.current_score)) - .unwrap_or_else(|| ("none".to_string(), 0.0)); - - let status = if worst_score >= 0.7 { - "critical" - } else if worst_score >= 0.3 { - "warning" - } else { - "healthy" - }; - - Ok(Json(EntityDriftResponse { - entity_id: id, - score: worst_score, - drift_type: worst_type, - status: status.to_string(), - })) -} - -/// Normalizer status handler -#[instrument(skip(state))] -async fn normalizer_status_handler( - State(state): State, -) -> Result, ApiError> { - let status = state.normalizer.status().await; - Ok(Json(status)) -} - -/// Trigger normalization handler -#[instrument(skip(state))] -async fn trigger_normalization_handler( - State(state): State, - Path(id): Path, -) -> Result { - validate_octad_id(&id)?; - let octad_id = OctadId::new(&id); - - // Check if octad exists - let _octad = state - .octad_store - .get(&octad_id) - .await - .map_err(|e| ApiError::Internal(e.to_string()))? - .ok_or_else(|| ApiError::NotFound(format!("Octad {} not found", id)))?; - - // In a full implementation, this would trigger actual normalization - // For now, we just verify the octad exists and return accepted - info!(id = %id, "Normalization triggered for octad"); - - Ok(StatusCode::ACCEPTED) -} - -// --- Query Planner Handlers --- - -/// Query plan handler — optimize a logical plan into a physical plan -#[instrument(skip(state, plan))] -async fn query_plan_handler( - State(state): State, - Json(plan): Json, -) -> Result, ApiError> { - let planner = state.planner.lock().map_err(|_| ApiError::Internal("Planner lock poisoned".to_string()))?; - let physical = planner - .optimize(&plan) - .map_err(|e| ApiError::BadRequest(e.to_string()))?; - Ok(Json(physical)) -} - -/// Query explain handler — generate EXPLAIN output for a logical plan -#[instrument(skip(state, plan))] -async fn query_explain_handler( - State(state): State, - Json(plan): Json, -) -> Result, ApiError> { - let planner = state.planner.lock().map_err(|_| ApiError::Internal("Planner lock poisoned".to_string()))?; - let explain = planner - .explain(&plan) - .map_err(|e| ApiError::BadRequest(e.to_string()))?; - Ok(Json(explain)) -} - -/// Get planner configuration -#[instrument(skip(state))] -async fn get_planner_config_handler( - State(state): State, -) -> Result, ApiError> { - let planner = state.planner.lock().map_err(|_| ApiError::Internal("Planner lock poisoned".to_string()))?; - Ok(Json(planner.config().clone())) -} - -/// Update planner configuration -#[instrument(skip(state, config))] -async fn put_planner_config_handler( - State(state): State, - Json(config): Json, -) -> Result, ApiError> { - let mut planner = state.planner.lock().map_err(|_| ApiError::Internal("Planner lock poisoned".to_string()))?; - planner.set_config(config); - Ok(Json(planner.config().clone())) -} - -/// Planner statistics snapshot -#[instrument(skip(state))] -async fn planner_stats_handler( - State(state): State, -) -> Result, ApiError> { - let planner = state.planner.lock().map_err(|_| ApiError::Internal("Planner lock poisoned".to_string()))?; - Ok(Json(planner.stats().clone())) -} - -// --- Meta-Query Store (Homoiconicity) --- - -/// Store query request body -#[derive(Debug, Serialize, Deserialize)] -pub struct StoreQueryRequest { - /// The VCL query text - pub query: String, - /// Optional embedding for the query - pub embedding: Option>, - /// Optional cost vector from the planner - pub cost_vector: Option>, - /// Optional proof obligations - pub proof_obligations: Option>, -} - -/// Store a VCL query as a octad (homoiconicity) -#[instrument(skip(state, request))] -async fn store_query_handler( - State(state): State, - Json(request): Json, -) -> Result<(StatusCode, Json), ApiError> { - use verisim_octad::QueryOctadBuilder; - - let mut builder = QueryOctadBuilder::new(&request.query); - - if let Some(embedding) = request.embedding { - validate_vector(&embedding)?; - builder = builder.with_embedding(embedding); - } - - if let Some(costs) = request.cost_vector { - builder = builder.with_cost_vector(costs); - } - - if let Some(obligations) = request.proof_obligations { - builder = builder.with_proof_obligations(obligations); - } - - builder = builder.with_metadata("stored_at", chrono::Utc::now().to_rfc3339()); - - let (_query_id, input) = builder.build(); - - let octad = state - .octad_store - .create(input) - .await - .map_err(|e| ApiError::Internal(e.to_string()))?; - - info!(octad_id = %octad.id, "Stored query as octad"); - - Ok((StatusCode::CREATED, Json(OctadResponse::from(&octad)))) -} - -/// Find similar past queries by vector similarity -#[instrument(skip(state, request))] -async fn similar_queries_handler( - State(state): State, - Json(request): Json, -) -> Result>, ApiError> { - let k = validate_limit(request.k.unwrap_or(10)); - - if request.vector.len() != state.config.vector_dimension { - return Err(ApiError::BadRequest(format!( - "Vector dimension mismatch: expected {}, got {}", - state.config.vector_dimension, - request.vector.len() - ))); - } - validate_vector(&request.vector)?; - - // Search for similar octads (which includes query-octads) - let octads = state - .octad_store - .search_similar(&request.vector, k) - .await - .map_err(|e| ApiError::Internal(e.to_string()))?; - - // Filter to only query octads (those with "vcl_query" type in document fields) - let results: Vec = octads - .iter() - .filter(|h| { - h.document - .as_ref() - .map(|d| d.title.starts_with("VCL Query:")) - .unwrap_or(false) - }) - .enumerate() - .map(|(i, h)| SearchResultResponse { - id: h.id.to_string(), - score: 1.0 - (i as f32 * 0.1), - title: h.document.as_ref().map(|d| d.title.clone()), - }) - .collect(); - - Ok(Json(results)) -} - -/// Optimize a stored query — re-plan and update its tensor modality with new costs. -/// This is reflection: the system modifying its own queries based on learned costs. -/// Accepts an optional LogicalPlan body; if absent, uses existing tensor as baseline. -#[instrument(skip(state))] -async fn optimize_query_handler( - State(state): State, - Path(id): Path, - body: Option>, -) -> Result, ApiError> { - validate_octad_id(&id)?; - let octad_id = OctadId::new(&id); - - // Get the existing query octad - let octad = state - .octad_store - .get(&octad_id) - .await - .map_err(|e| ApiError::Internal(e.to_string()))? - .ok_or_else(|| ApiError::NotFound(format!("Query octad {} not found", id)))?; - - // Compute cost vector from the planner - let cost_vector = if let Some(Json(logical_plan)) = body { - // If a logical plan was provided, run the planner on it - let planner = state.planner.lock().map_err(|_| { - ApiError::Internal("Planner lock poisoned".to_string()) - })?; - - match planner.explain(&logical_plan) { - Ok(explain) => { - explain - .steps - .iter() - .map(|s| s.estimated_cost_ms) - .collect::>() - } - Err(_) => vec![1.0, 0.0, 0.0], - } - } else { - // No plan provided — use existing tensor data scaled by a learning factor, - // or default to [1.0] if no tensor exists - octad - .tensor - .as_ref() - .map(|t| t.data.iter().map(|v| v * 0.95).collect()) - .unwrap_or_else(|| vec![1.0]) - }; - - // Update the octad with new tensor data (cost vector) - let mut update_input = OctadInput::default(); - update_input.tensor = Some(OctadTensorInput { - shape: vec![1, cost_vector.len()], - data: cost_vector, - }); - update_input - .metadata - .insert("optimized_at".to_string(), chrono::Utc::now().to_rfc3339()); - - let updated = state - .octad_store - .update(&octad_id, update_input) - .await - .map_err(|e| ApiError::Internal(e.to_string()))?; - - info!(octad_id = %id, "Optimized query octad with new cost vector"); - - Ok(Json(OctadResponse::from(&updated))) -} - -// --- EXPLAIN ANALYZE Handler --- - -/// EXPLAIN ANALYZE request — execute a plan and return actual timings -#[derive(Debug, Serialize, Deserialize)] -pub struct ExplainAnalyzeRequest { - /// The logical plan to analyze - pub plan: LogicalPlan, - /// Simulated step execution times (milliseconds) for profiling. - /// In a real execution engine, these would be measured; here they - /// can be provided for testing/simulation purposes. - pub simulated_timings: Option>, -} - -/// EXPLAIN ANALYZE handler — produces plan estimates with simulated actual timings -#[instrument(skip(state, request))] -async fn query_explain_analyze_handler( - State(state): State, - Json(request): Json, -) -> Result, ApiError> { - let mut planner = state.planner.lock().map_err(|_| ApiError::Internal("Planner lock poisoned".to_string()))?; - let explain = planner - .explain(&request.plan) - .map_err(|e| ApiError::BadRequest(e.to_string()))?; - let physical = planner - .optimize(&request.plan) - .map_err(|e| ApiError::BadRequest(e.to_string()))?; - - let plan_id = format!("analyze-{}", chrono::Utc::now().timestamp_millis()); - let mut profiler = Profiler::new(&plan_id, &physical); - - // Record simulated or default step timings - let now = chrono::Utc::now(); - for (i, step) in physical.steps.iter().enumerate() { - let actual_ms = request.simulated_timings - .as_ref() - .and_then(|t| t.get(i).copied()) - .unwrap_or(step.cost.time_ms * 1.1); // Default: 10% slower than estimate - profiler.record_step(i, actual_ms, step.cost.estimated_rows, now, now); - } - - let profile = profiler.finish(planner.stats_mut()); - let output = explain.with_profile(&profile); - - Ok(Json(output)) -} - -// --- Prepared Statements Handlers --- - -/// Request to create a prepared statement -#[derive(Debug, Serialize, Deserialize)] -pub struct PreparedCreateRequest { - /// The VCL query text - pub query: String, - /// The logical plan for the query - pub plan: LogicalPlan, -} - -/// Create a prepared statement -#[instrument(skip(state, request))] -async fn prepared_create_handler( - State(state): State, - Json(request): Json, -) -> Result<(StatusCode, Json), ApiError> { - let id = state.plan_cache.prepare(&request.query, request.plan).await; - - let stmt = state.plan_cache.get(&id).await - .ok_or_else(|| ApiError::Internal("Failed to retrieve prepared statement after creation".to_string()))?; - - Ok((StatusCode::CREATED, Json(stmt))) -} - -/// Get a prepared statement by ID -#[instrument(skip(state))] -async fn prepared_get_handler( - State(state): State, - Path(id): Path, -) -> Result, ApiError> { - let prep_id = PreparedId::new(&id); - let stmt = state.plan_cache.get(&prep_id).await - .ok_or_else(|| ApiError::NotFound(format!("Prepared statement '{}' not found", id)))?; - Ok(Json(stmt)) -} - -/// Execute a prepared statement with parameters -#[derive(Debug, Serialize, Deserialize)] -pub struct PreparedExecuteRequest { - /// Parameter bindings - pub params: std::collections::HashMap, -} - -/// Execute a prepared statement -#[instrument(skip(state, request))] -async fn prepared_execute_handler( - State(state): State, - Path(id): Path, - Json(request): Json, -) -> Result, ApiError> { - let prep_id = PreparedId::new(&id); - - let stmt = state.plan_cache - .execute_prepared(&prep_id, &request.params) - .await - .map_err(|e| ApiError::BadRequest(e.to_string()))?; - - // Use cached physical plan if available, otherwise optimize - let physical = if let Some(cached) = stmt.cached_physical_plan { - cached - } else { - let planner = state.planner.lock().map_err(|_| ApiError::Internal("Planner lock poisoned".to_string()))?; - planner.optimize(&stmt.logical_plan).map_err(|e| ApiError::Internal(e.to_string()))? - }; - - // Cache the physical plan for future use - state.plan_cache.cache_plan(&prep_id, physical.clone()).await; - - Ok(Json(physical)) -} - -/// Get prepared statement cache statistics -#[instrument(skip(state))] -async fn prepared_stats_handler( - State(state): State, -) -> Result, ApiError> { - let stats = state.plan_cache.stats_async().await; - Ok(Json(stats)) -} - -/// Get slow query log summary -#[instrument(skip(state))] -async fn slow_queries_handler( - State(state): State, -) -> Result, ApiError> { - let summary = state.slow_query_log.summary(); - Ok(Json(summary)) -} - -// --- Transaction Handlers --- - -/// Begin a new transaction -#[instrument(skip(state))] -async fn transaction_begin_handler( - State(state): State, -) -> Result<(StatusCode, Json), ApiError> { - let txn_id = state.transaction_manager - .begin() - .await - .map_err(|e| ApiError::Internal(e.to_string()))?; - - let status = state.transaction_manager - .status(&txn_id) - .await - .map_err(|e| ApiError::Internal(e.to_string()))?; - - Ok((StatusCode::CREATED, Json(status))) -} - -/// Commit a transaction -#[instrument(skip(state))] -async fn transaction_commit_handler( - State(state): State, - Path(id): Path, -) -> Result, ApiError> { - let txn_id = transaction::TransactionId::from_str(&id); - - let _ops = state.transaction_manager - .commit(&txn_id) - .await - .map_err(|e| match e { - transaction::TransactionError::NotFound(_) => ApiError::NotFound(e.to_string()), - _ => ApiError::BadRequest(e.to_string()), - })?; - - // In a full implementation, ops would be applied to the octad store here - let status = state.transaction_manager - .status(&txn_id) - .await - .map_err(|e| ApiError::Internal(e.to_string()))?; - - Ok(Json(status)) -} - -/// Rollback a transaction -#[instrument(skip(state))] -async fn transaction_rollback_handler( - State(state): State, - Path(id): Path, -) -> Result, ApiError> { - let txn_id = transaction::TransactionId::from_str(&id); - - let _discarded = state.transaction_manager - .rollback(&txn_id) - .await - .map_err(|e| match e { - transaction::TransactionError::NotFound(_) => ApiError::NotFound(e.to_string()), - _ => ApiError::BadRequest(e.to_string()), - })?; - - let status = state.transaction_manager - .status(&txn_id) - .await - .map_err(|e| ApiError::Internal(e.to_string()))?; - - Ok(Json(status)) -} - -/// Get transaction status -#[instrument(skip(state))] -async fn transaction_status_handler( - State(state): State, - Path(id): Path, -) -> Result, ApiError> { - let txn_id = transaction::TransactionId::from_str(&id); - - let status = state.transaction_manager - .status(&txn_id) - .await - .map_err(|e| match e { - transaction::TransactionError::NotFound(_) => ApiError::NotFound(e.to_string()), - _ => ApiError::Internal(e.to_string()), - })?; - - Ok(Json(status)) -} - -// --- ZKP Proof Handlers --- - -/// API request for proof generation -#[derive(Debug, Serialize, Deserialize)] -pub struct ProofGenerateRequest { - /// The claim to prove (base64-encoded or plain text) - pub claim: String, - /// Privacy level: "public", "private", or "zero_knowledge" - pub privacy_level: Option, - /// Optional membership set for Merkle inclusion proofs - pub membership_set: Option>, - /// Index of the claim in the membership set - pub membership_index: Option, -} - -/// API request for proof verification -#[derive(Debug, Serialize, Deserialize)] -pub struct ProofVerifyRequest { - /// The proof to verify (serialized) - pub proof: zkp_api::ZkpProof, - /// The claim the proof is for - pub claim: String, -} - -/// API request for circuit-based proof generation -#[derive(Debug, Serialize, Deserialize)] -pub struct ProofWithCircuitRequest { - /// The claim to prove - pub claim: String, - /// Privacy level - pub privacy_level: Option, - /// Circuit name to verify against - pub circuit_name: String, - /// Witness data (private inputs) - pub witness: Option>, - /// Public inputs - pub public_inputs: Option>, -} - -/// API response for proof operations -#[derive(Debug, Serialize, Deserialize)] -pub struct ProofResponse { - pub success: bool, - #[serde(skip_serializing_if = "Option::is_none")] - pub proof: Option, - #[serde(skip_serializing_if = "Option::is_none")] - pub verified: Option, - #[serde(skip_serializing_if = "Option::is_none")] - pub error: Option, -} - -fn parse_privacy_level(s: &str) -> Result { - match s.to_lowercase().as_str() { - "public" => Ok(PrivacyLevel::Public), - "private" => Ok(PrivacyLevel::Private), - "zero_knowledge" | "zeroknowledge" | "zk" => Ok(PrivacyLevel::ZeroKnowledge), - other => Err(ApiError::BadRequest(format!( - "Unknown privacy level: '{}'. Use 'public', 'private', or 'zero_knowledge'", - other - ))), - } -} - -/// Generate a privacy-aware ZKP proof -#[instrument(skip(_state, request))] -async fn proof_generate_handler( - State(_state): State, - Json(request): Json, -) -> Result, ApiError> { - let privacy_level = match &request.privacy_level { - Some(level) => parse_privacy_level(level)?, - None => PrivacyLevel::Public, - }; - - let membership_set = request.membership_set.as_ref().map(|set| { - set.iter().map(|s| s.as_bytes().to_vec()).collect::>() - }); - - let bridge_request = ZkpBridgeRequest { - claim: request.claim.as_bytes().to_vec(), - privacy_level, - circuit_name: None, - witness: None, - public_inputs: None, - membership_set, - membership_index: request.membership_index, - }; - - match zkp_api::generate_zkp(&bridge_request) { - Ok(proof) => Ok(Json(ProofResponse { - success: true, - proof: Some(proof), - verified: None, - error: None, - })), - Err(e) => Ok(Json(ProofResponse { - success: false, - proof: None, - verified: None, - error: Some(e.to_string()), - })), - } -} - -/// Verify a previously generated ZKP proof -#[instrument(skip(_state, request))] -async fn proof_verify_handler( - State(_state): State, - Json(request): Json, -) -> Result, ApiError> { - let verified = zkp_api::verify_zkp(&request.proof, request.claim.as_bytes()); - - Ok(Json(ProofResponse { - success: true, - proof: None, - verified: Some(verified), - error: None, - })) -} - -/// Generate a proof with circuit verification -#[instrument(skip(state, request))] -async fn proof_generate_with_circuit_handler( - State(state): State, - Json(request): Json, -) -> Result, ApiError> { - let privacy_level = match &request.privacy_level { - Some(level) => parse_privacy_level(level)?, - None => PrivacyLevel::Public, - }; - - let bridge_request = ZkpBridgeRequest { - claim: request.claim.as_bytes().to_vec(), - privacy_level, - circuit_name: Some(request.circuit_name), - witness: request.witness, - public_inputs: request.public_inputs, - membership_set: None, - membership_index: None, - }; - - match zkp_api::generate_zkp_with_circuit(&bridge_request, &state.circuit_registry) { - Ok(proof) => Ok(Json(ProofResponse { - success: true, - proof: Some(proof), - verified: None, - error: None, - })), - Err(e) => Ok(Json(ProofResponse { - success: false, - proof: None, - verified: None, - error: Some(e.to_string()), - })), - } -} - -/// Start the API server (HTTP + gRPC) with graceful shutdown and hardening. -/// -/// HTTP server on `config.port` (default 8080) — for external clients. -/// gRPC server on `config.grpc_port` (default 50051) — for internal/federation. -/// Set `grpc_port = 0` to disable gRPC. -/// -/// Hardening: -/// - Request body size limit (`max_body_size`) -/// - Request timeout (`request_timeout_secs`) -/// - Graceful shutdown on SIGINT/SIGTERM with WAL flush -pub async fn serve(config: ApiConfig) -> Result<(), std::io::Error> { - let state = AppState::new_async(config.clone()) - .await - .map_err(|e| std::io::Error::other(e.to_string()))?; - - let octad_store = state.octad_store.clone(); - - // Build HTTP router with hardening middleware - let app = build_router(state.clone()) - .layer(axum::extract::DefaultBodyLimit::max(config.max_body_size)); - - // Start HTTP server - let http_addr = format!("{}:{}", config.host, config.port); - info!(addr = %http_addr, "Starting VeriSimDB HTTP server"); - let listener = TcpListener::bind(&http_addr).await?; - - let http_server = axum::serve(listener, app) - .with_graceful_shutdown(shutdown_signal()); - - // Start gRPC server (if enabled) - if config.grpc_port > 0 { - let grpc_addr = format!("{}:{}", config.host, config.grpc_port); - info!(addr = %grpc_addr, "Starting VeriSimDB gRPC server"); - - let grpc_router = grpc::build_grpc_router(state); - let grpc_addr_parsed: std::net::SocketAddr = grpc_addr - .parse() - .map_err(|e: std::net::AddrParseError| std::io::Error::other(e.to_string()))?; - - // Run both servers concurrently — when either stops (shutdown signal), both stop - tokio::select! { - result = http_server => { - if let Err(e) = result { - tracing::error!("HTTP server error: {e}"); - } - } - result = grpc_router.serve(grpc_addr_parsed) => { - if let Err(e) = result { - tracing::error!("gRPC server error: {e}"); - } - } - } - } else { - // HTTP only (gRPC disabled) - http_server.await?; - } - - // Clean shutdown: flush WAL - info!("VeriSimDB: servers stopped, flushing WAL..."); - if let Err(e) = octad_store.graceful_shutdown().await { - tracing::warn!("Graceful shutdown error (non-fatal): {e}"); - } - - Ok(()) -} - -/// Wait for a shutdown signal (Ctrl+C or SIGTERM). -async fn shutdown_signal() { - let ctrl_c = async { - tokio::signal::ctrl_c() - .await - .expect("Failed to install Ctrl+C handler"); - }; - - #[cfg(unix)] - let terminate = async { - tokio::signal::unix::signal(tokio::signal::unix::SignalKind::terminate()) - .expect("Failed to install SIGTERM handler") - .recv() - .await; - }; - - #[cfg(not(unix))] - let terminate = std::future::pending::<()>(); - - tokio::select! { - _ = ctrl_c => info!("Received Ctrl+C, initiating graceful shutdown"), - _ = terminate => info!("Received SIGTERM, initiating graceful shutdown"), - } -} - -/// Start the API server with TLS (HTTPS) and graceful shutdown. -pub async fn serve_tls( - config: ApiConfig, - cert_path: &str, - key_path: &str, -) -> Result<(), std::io::Error> { - use axum_server::tls_rustls::RustlsConfig; - - let state = AppState::new_async(config.clone()) - .await - .map_err(|e| std::io::Error::other(e.to_string()))?; - - let octad_store = state.octad_store.clone(); - let app = build_router(state); - - let addr = format!("{}:{}", config.host, config.port); - info!(addr = %addr, cert = %cert_path, "Starting VeriSimDB API server with TLS"); - - let tls_config = RustlsConfig::from_pem_file(cert_path, key_path) - .await - .map_err(|e| std::io::Error::other(e.to_string()))?; - - let addr: std::net::SocketAddr = addr - .parse() - .map_err(|e: std::net::AddrParseError| std::io::Error::other(e.to_string()))?; - - let handle = axum_server::Handle::new(); - let shutdown_handle = handle.clone(); - - // Spawn shutdown listener - tokio::spawn(async move { - shutdown_signal().await; - shutdown_handle.graceful_shutdown(Some(std::time::Duration::from_secs(10))); - }); - - axum_server::bind_rustls(addr, tls_config) - .handle(handle) - .serve(app.into_make_service()) - .await?; - - // After server stops, perform clean shutdown - info!("VeriSimDB TLS: server stopped, flushing WAL..."); - if let Err(e) = octad_store.graceful_shutdown().await { - tracing::warn!("Graceful shutdown error (non-fatal): {e}"); - } - - Ok(()) -} - -// --------------------------------------------------------------------------- -// Provenance endpoint handlers -// --------------------------------------------------------------------------- - -/// Provenance chain response -#[derive(Debug, Serialize, Deserialize)] -pub struct ProvenanceChainResponse { - pub entity_id: String, - pub chain_length: usize, - pub chain_valid: bool, - pub records: Vec, -} - -/// A single provenance record in the response -#[derive(Debug, Serialize, Deserialize)] -pub struct ProvenanceRecordResponse { - pub event_type: String, - pub actor: String, - pub timestamp: String, - pub source: Option, - pub description: String, - pub content_hash: String, -} - -/// GET /provenance/{id} — retrieve the full provenance chain for an entity -#[instrument(skip(state))] -async fn provenance_get_chain_handler( - State(state): State, - Path(id): Path, -) -> Result, ApiError> { - validate_octad_id(&id)?; - - // Check entity exists - let octad_id = OctadId::new(&id); - let exists = state - .octad_store - .status(&octad_id) - .await - .map_err(|e| ApiError::Internal(e.to_string()))?; - if exists.is_none() { - return Err(ApiError::NotFound(format!("Entity {} not found", id))); - } - - let chain = state - .octad_store - .provenance_store() - .get_chain(&id) - .await - .map_err(|e| ApiError::Internal(e.to_string()))?; - - let chain_valid = state - .octad_store - .provenance_store() - .verify_chain(&id) - .await - .unwrap_or(false); - - let records: Vec = chain - .records - .iter() - .map(|r| ProvenanceRecordResponse { - event_type: format!("{:?}", r.event_type), - actor: r.actor.clone(), - timestamp: r.timestamp.to_rfc3339(), - source: r.source.clone(), - description: r.description.clone(), - content_hash: r.content_hash.clone(), - }) - .collect(); - - Ok(Json(ProvenanceChainResponse { - entity_id: id, - chain_length: records.len(), - chain_valid, - records, - })) -} - -/// POST /provenance/{id}/record — record a new provenance event -#[instrument(skip(state, body))] -async fn provenance_record_handler( - State(state): State, - Path(id): Path, - Json(body): Json, -) -> Result, ApiError> { - validate_octad_id(&id)?; - - let octad_id = OctadId::new(&id); - let input = OctadInput { - provenance: Some(OctadProvenanceInput { - event_type: body.event_type, - actor: body.actor, - source: body.source, - description: body.description, - }), - ..Default::default() - }; - - let octad = state - .octad_store - .update(&octad_id, input) - .await - .map_err(|e| ApiError::Internal(e.to_string()))?; - - Ok(Json(serde_json::json!({ - "entity_id": id, - "chain_length": octad.provenance_chain_length, - "recorded": true, - }))) -} - -/// GET /provenance/{id}/verify — verify provenance chain integrity -#[instrument(skip(state))] -async fn provenance_verify_handler( - State(state): State, - Path(id): Path, -) -> Result, ApiError> { - validate_octad_id(&id)?; - - let octad_id = OctadId::new(&id); - let status = state - .octad_store - .status(&octad_id) - .await - .map_err(|e| ApiError::Internal(e.to_string()))? - .ok_or_else(|| ApiError::NotFound(format!("Entity {} not found", id)))?; - - Ok(Json(serde_json::json!({ - "entity_id": id, - "has_provenance": status.modality_status.provenance, - "chain_valid": true, - }))) -} - -// --------------------------------------------------------------------------- -// Spatial endpoint handlers -// --------------------------------------------------------------------------- - -/// Radius search request -#[derive(Debug, Deserialize)] -pub struct RadiusSearchRequest { - pub latitude: f64, - pub longitude: f64, - pub radius_km: f64, - pub limit: Option, -} - -/// Bounding box search request -#[derive(Debug, Deserialize)] -pub struct BoundsSearchRequest { - pub min_lat: f64, - pub min_lon: f64, - pub max_lat: f64, - pub max_lon: f64, - pub limit: Option, -} - -/// K-nearest search request -#[derive(Debug, Deserialize)] -pub struct NearestSearchRequest { - pub latitude: f64, - pub longitude: f64, - pub k: Option, -} - -/// Spatial search result response -#[derive(Debug, Serialize)] -pub struct SpatialSearchResultResponse { - pub entity_id: String, - pub latitude: f64, - pub longitude: f64, - pub distance_km: f64, -} - -/// POST /spatial/search/radius — find entities within a given radius -#[instrument(skip_all)] -async fn spatial_radius_search_handler( - State(state): State, - Json(body): Json, -) -> Result>, ApiError> { - let limit = validate_limit(body.limit.unwrap_or(100)); - - if !(-90.0..=90.0).contains(&body.latitude) || !(-180.0..=180.0).contains(&body.longitude) { - return Err(ApiError::BadRequest("Invalid coordinates".to_string())); - } - if body.radius_km <= 0.0 { - return Err(ApiError::BadRequest("Radius must be positive".to_string())); - } - - let center = Coordinates { - latitude: body.latitude, - longitude: body.longitude, - altitude: None, - }; - - let results = state - .octad_store - .spatial_store() - .search_radius(¢er, body.radius_km, limit) - .await - .map_err(|e| ApiError::Internal(e.to_string()))?; - - let response = results - .into_iter() - .map(|r| SpatialSearchResultResponse { - entity_id: r.entity_id, - latitude: r.data.coordinates.latitude, - longitude: r.data.coordinates.longitude, - distance_km: r.distance_km, - }) - .collect(); - - Ok(Json(response)) -} - -/// POST /spatial/search/bounds — find entities within a bounding box -#[instrument(skip_all)] -async fn spatial_bounds_search_handler( - State(state): State, - Json(body): Json, -) -> Result>, ApiError> { - let limit = validate_limit(body.limit.unwrap_or(100)); - - if body.min_lat > body.max_lat || body.min_lon > body.max_lon { - return Err(ApiError::BadRequest( - "min values must be less than max values".to_string(), - )); - } - - let bounds = BoundingBox { - min_lat: body.min_lat, - min_lon: body.min_lon, - max_lat: body.max_lat, - max_lon: body.max_lon, - }; - - let results = state - .octad_store - .spatial_store() - .search_within(&bounds, limit) - .await - .map_err(|e| ApiError::Internal(e.to_string()))?; - - let response = results - .into_iter() - .map(|r| SpatialSearchResultResponse { - entity_id: r.entity_id, - latitude: r.data.coordinates.latitude, - longitude: r.data.coordinates.longitude, - distance_km: r.distance_km, - }) - .collect(); - - Ok(Json(response)) -} - -/// POST /spatial/search/nearest — find k nearest entities to a point -#[instrument(skip_all)] -async fn spatial_nearest_handler( - State(state): State, - Json(body): Json, -) -> Result>, ApiError> { - if !(-90.0..=90.0).contains(&body.latitude) || !(-180.0..=180.0).contains(&body.longitude) { - return Err(ApiError::BadRequest("Invalid coordinates".to_string())); - } - - let k = body.k.unwrap_or(10).min(MAX_RESULT_LIMIT); - - let point = Coordinates { - latitude: body.latitude, - longitude: body.longitude, - altitude: None, - }; - - let results = state - .octad_store - .spatial_store() - .nearest(&point, k) - .await - .map_err(|e| ApiError::Internal(e.to_string()))?; - - let response = results - .into_iter() - .map(|r| SpatialSearchResultResponse { - entity_id: r.entity_id, - latitude: r.data.coordinates.latitude, - longitude: r.data.coordinates.longitude, - distance_km: r.distance_km, - }) - .collect(); - - Ok(Json(response)) -} - -#[cfg(test)] -mod tests { - use super::*; - use axum::body::Body; - use axum::http::{Request, StatusCode}; - use tower::ServiceExt; - - async fn create_test_state() -> AppState { - let mut config = ApiConfig { - vector_dimension: 3, - ..Default::default() - }; - - // When the `persistent` feature is enabled, each test gets a unique temp directory - // to avoid redb lock contention between parallel tests. - #[cfg(feature = "persistent")] - { - use std::sync::atomic::{AtomicU64, Ordering}; - static COUNTER: AtomicU64 = AtomicU64::new(0); - let id = COUNTER.fetch_add(1, Ordering::Relaxed); - let tmp = std::env::temp_dir().join(format!( - "verisimdb-test-{}-{}", - std::process::id(), - id, - )); - config.persistence_dir = Some(tmp.to_string_lossy().into_owned()); - } - - AppState::new_async(config).await.expect("TODO: handle error") - } - - #[tokio::test] - async fn test_health_endpoint() { - let state = create_test_state().await; - let app = build_router(state); - - let response = app - .oneshot( - Request::builder() - .uri("/health") - .body(Body::empty()) - .expect("TODO: handle error"), - ) - .await - .expect("TODO: handle error"); - - assert_eq!(response.status(), StatusCode::OK); - } - - #[tokio::test] - async fn test_ready_endpoint() { - let state = create_test_state().await; - let app = build_router(state); - - let response = app - .oneshot( - Request::builder() - .uri("/ready") - .body(Body::empty()) - .expect("TODO: handle error"), - ) - .await - .expect("TODO: handle error"); - - assert_eq!(response.status(), StatusCode::OK); - } - - #[tokio::test] - async fn test_create_and_get_octad() { - let state = create_test_state().await; - let app = build_router(state); - - // Create a octad - let create_request = OctadRequest { - title: Some("Test Document".to_string()), - body: Some("Test body content".to_string()), - embedding: Some(vec![0.1, 0.2, 0.3]), - types: None, - relationships: None, - tensor: None, - temporal: None, - metadata: None, - provenance: None, - spatial: None, - }; - - let response = app - .clone() - .oneshot( - Request::builder() - .method("POST") - .uri("/octads") - .header("content-type", "application/json") - .body(Body::from(serde_json::to_string(&create_request).expect("TODO: handle error"))) - .expect("TODO: handle error"), - ) - .await - .expect("TODO: handle error"); - - assert_eq!(response.status(), StatusCode::CREATED); - - // Parse response to get ID - let body = axum::body::to_bytes(response.into_body(), 1024 * 1024) - .await - .expect("TODO: handle error"); - let created: OctadResponse = serde_json::from_slice(&body).expect("TODO: handle error"); - - // Get the octad - let response = app - .oneshot( - Request::builder() - .uri(format!("/octads/{}", created.id)) - .body(Body::empty()) - .expect("TODO: handle error"), - ) - .await - .expect("TODO: handle error"); - - assert_eq!(response.status(), StatusCode::OK); - } - - #[tokio::test] - async fn test_include_embedding_byte_exact_round_trip() { - // Phase 1 gap closure: GET /octads/{id}?include=embedding must - // surface the stored Vec verbatim — same dimension, same - // component values, no truncation, no reorder. This is what makes - // the Vector shape byte-exactly veridical. - let state = create_test_state().await; - let app = build_router(state); - - let original: Vec = vec![0.1, 0.2, 0.3]; - let create_request = OctadRequest { - title: Some("Vec round-trip".to_string()), - body: Some("Body".to_string()), - embedding: Some(original.clone()), - types: None, - relationships: None, - tensor: None, - temporal: None, - metadata: None, - provenance: None, - spatial: None, - }; - let response = app - .clone() - .oneshot( - Request::builder() - .method("POST") - .uri("/octads") - .header("content-type", "application/json") - .body(Body::from( - serde_json::to_string(&create_request).expect("serialize create request"), - )) - .expect("build POST request"), - ) - .await - .expect("oneshot create"); - assert_eq!(response.status(), StatusCode::CREATED); - let body = axum::body::to_bytes(response.into_body(), 1024 * 1024) - .await - .expect("read create body"); - let created: OctadResponse = - serde_json::from_slice(&body).expect("decode create response"); - assert!( - created.embedding.is_none(), - "POST response must not surface embedding by default" - ); - - // GET without include — embedding must be omitted - let response = app - .clone() - .oneshot( - Request::builder() - .uri(format!("/octads/{}", created.id)) - .body(Body::empty()) - .expect("build GET request"), - ) - .await - .expect("oneshot get"); - let body = axum::body::to_bytes(response.into_body(), 1024 * 1024) - .await - .expect("read get body"); - let got: OctadResponse = serde_json::from_slice(&body).expect("decode get response"); - assert!( - got.embedding.is_none(), - "default GET must not surface embedding" - ); - - // GET with ?include=embedding — bytes must match exactly - let response = app - .oneshot( - Request::builder() - .uri(format!("/octads/{}?include=embedding", created.id)) - .body(Body::empty()) - .expect("build GET request"), - ) - .await - .expect("oneshot get with include=embedding"); - let body = axum::body::to_bytes(response.into_body(), 1024 * 1024) - .await - .expect("read get body"); - let got: OctadResponse = - serde_json::from_slice(&body).expect("decode get response with embedding"); - let returned = got - .embedding - .as_ref() - .expect("?include=embedding must surface embedding"); - assert_eq!( - returned.len(), - original.len(), - "byte-exact round-trip — same dimension" - ); - for (i, (a, b)) in original.iter().zip(returned.iter()).enumerate() { - assert_eq!( - a.to_bits(), - b.to_bits(), - "byte-exact round-trip — component {i} differs ({a} vs {b})" - ); - } - } - - #[tokio::test] - async fn test_include_types_round_trip() { - // Phase 1 gap closure: GET /octads/{id}?include=types must surface the - // entity's declared semantic types as semantic_types: Vec. The - // default GET (no include) must still omit the field so legacy clients - // see no schema change. - let state = create_test_state().await; - let app = build_router(state); - - let create_request = OctadRequest { - title: Some("Typed entity".to_string()), - body: Some("Body".to_string()), - embedding: Some(vec![0.1, 0.2, 0.3]), - types: Some(vec![ - "https://example.org/Email".to_string(), - "https://schema.org/Message".to_string(), - ]), - relationships: None, - tensor: None, - temporal: None, - metadata: None, - provenance: None, - spatial: None, - }; - let response = app - .clone() - .oneshot( - Request::builder() - .method("POST") - .uri("/octads") - .header("content-type", "application/json") - .body(Body::from( - serde_json::to_string(&create_request).expect("serialize create request"), - )) - .expect("build POST request"), - ) - .await - .expect("oneshot create"); - assert_eq!(response.status(), StatusCode::CREATED); - - let body = axum::body::to_bytes(response.into_body(), 1024 * 1024) - .await - .expect("read create body"); - let created: OctadResponse = - serde_json::from_slice(&body).expect("decode create response"); - assert!( - created.semantic_types.is_none(), - "default response (POST) must not surface semantic_types" - ); - - // GET without include — types still omitted - let response = app - .clone() - .oneshot( - Request::builder() - .uri(format!("/octads/{}", created.id)) - .body(Body::empty()) - .expect("build GET request"), - ) - .await - .expect("oneshot get"); - assert_eq!(response.status(), StatusCode::OK); - let body = axum::body::to_bytes(response.into_body(), 1024 * 1024) - .await - .expect("read get body"); - let got: OctadResponse = serde_json::from_slice(&body).expect("decode get response"); - assert!( - got.semantic_types.is_none(), - "default GET must not surface semantic_types" - ); - - // GET with ?include=types — types surfaced verbatim - let response = app - .oneshot( - Request::builder() - .uri(format!("/octads/{}?include=types", created.id)) - .body(Body::empty()) - .expect("build GET request"), - ) - .await - .expect("oneshot get with include"); - assert_eq!(response.status(), StatusCode::OK); - let body = axum::body::to_bytes(response.into_body(), 1024 * 1024) - .await - .expect("read get body"); - let got: OctadResponse = - serde_json::from_slice(&body).expect("decode get response with include"); - let types = got - .semantic_types - .as_ref() - .expect("?include=types must surface semantic_types"); - assert_eq!( - types, - &vec![ - "https://example.org/Email".to_string(), - "https://schema.org/Message".to_string(), - ], - "semantic_types must round-trip in declaration order" - ); - } - - #[tokio::test] - async fn test_temporal_observed_at_round_trip() { - // Phase 1 gap closure: caller-supplied real-world time (e.g. an email's - // Date: header) must round-trip through POST/GET as RFC 3339 in the - // OctadStatusResponse.observed_at field. - let state = create_test_state().await; - let app = build_router(state); - - let create_request = OctadRequest { - title: Some("Email subject".to_string()), - body: Some("Email body".to_string()), - embedding: Some(vec![0.1, 0.2, 0.3]), - types: None, - relationships: None, - tensor: None, - temporal: Some(TemporalRequest { - observed_at: "2026-04-27T15:30:00Z".to_string(), - }), - metadata: None, - provenance: None, - spatial: None, - }; - - let response = app - .clone() - .oneshot( - Request::builder() - .method("POST") - .uri("/octads") - .header("content-type", "application/json") - .body(Body::from( - serde_json::to_string(&create_request).expect("serialize create request"), - )) - .expect("build POST request"), - ) - .await - .expect("oneshot create"); - assert_eq!(response.status(), StatusCode::CREATED); - - let body = axum::body::to_bytes(response.into_body(), 1024 * 1024) - .await - .expect("read create body"); - let created: OctadResponse = serde_json::from_slice(&body).expect("decode response"); - - let observed = created - .status - .observed_at - .as_ref() - .expect("create response surfaces observed_at"); - assert!( - observed.starts_with("2026-04-27T15:30:00"), - "observed_at must reflect caller-supplied time, got {observed}" - ); - - // Bad RFC 3339 must surface as 400, not be silently dropped - let bad_request = OctadRequest { - title: Some("Bad".to_string()), - body: Some("Bad".to_string()), - embedding: Some(vec![0.1, 0.2, 0.3]), - types: None, - relationships: None, - tensor: None, - temporal: Some(TemporalRequest { - observed_at: "not-a-date".to_string(), - }), - metadata: None, - provenance: None, - spatial: None, - }; - let bad_response = app - .oneshot( - Request::builder() - .method("POST") - .uri("/octads") - .header("content-type", "application/json") - .body(Body::from( - serde_json::to_string(&bad_request).expect("serialize bad request"), - )) - .expect("build bad POST request"), - ) - .await - .expect("oneshot bad create"); - assert_eq!(bad_response.status(), StatusCode::BAD_REQUEST); - } - - #[tokio::test] - async fn test_text_search() { - let state = create_test_state().await; - let app = build_router(state); - - // Create a octad - let create_request = OctadRequest { - title: Some("Rust Programming".to_string()), - body: Some("Rust is a systems programming language".to_string()), - embedding: Some(vec![0.1, 0.2, 0.3]), - types: None, - relationships: None, - tensor: None, - temporal: None, - metadata: None, - provenance: None, - spatial: None, - }; - - let _ = app - .clone() - .oneshot( - Request::builder() - .method("POST") - .uri("/octads") - .header("content-type", "application/json") - .body(Body::from(serde_json::to_string(&create_request).expect("TODO: handle error"))) - .expect("TODO: handle error"), - ) - .await - .expect("TODO: handle error"); - - // Search for it - let response = app - .oneshot( - Request::builder() - .uri("/search/text?q=Rust&limit=10") - .body(Body::empty()) - .expect("TODO: handle error"), - ) - .await - .expect("TODO: handle error"); - - assert_eq!(response.status(), StatusCode::OK); - } - - #[tokio::test] - async fn test_drift_status() { - let state = create_test_state().await; - let app = build_router(state); - - let response = app - .oneshot( - Request::builder() - .uri("/drift/status") - .body(Body::empty()) - .expect("TODO: handle error"), - ) - .await - .expect("TODO: handle error"); - - assert_eq!(response.status(), StatusCode::OK); - } -} diff --git a/verisimdb/rust-core/verisim-api/src/main.rs b/verisimdb/rust-core/verisim-api/src/main.rs deleted file mode 100644 index a7889a09..00000000 --- a/verisimdb/rust-core/verisim-api/src/main.rs +++ /dev/null @@ -1,110 +0,0 @@ -// SPDX-License-Identifier: MPL-2.0 -//! VeriSimDB API server binary -//! -//! Starts the HTTP API server for VeriSimDB. -//! Defaults to IPv6-only ([::]). Set VERISIM_ENABLE_IPV4=true for dual-stack. -//! Set VERISIM_TLS_CERT and VERISIM_TLS_KEY for HTTPS mode. - -use verisim_api::ApiConfig; - -#[tokio::main] -async fn main() -> Result<(), Box> { - // Install ring as the default crypto provider (pure Rust, no OpenSSL/aws-lc-sys) - rustls::crypto::ring::default_provider() - .install_default() - .expect("failed to install ring crypto provider"); - - // Initialize tracing with structured JSON output - let json_logging = std::env::var("VERISIM_LOG_FORMAT") - .map(|v| v == "json") - .unwrap_or(true); // JSON by default in production - - let env_filter = tracing_subscriber::EnvFilter::try_from_default_env() - .unwrap_or_else(|_| tracing_subscriber::EnvFilter::new("info")); - - if json_logging { - tracing_subscriber::fmt() - .json() - .with_env_filter(env_filter) - .init(); - } else { - tracing_subscriber::fmt() - .with_env_filter(env_filter) - .init(); - } - - // IPv6-only by default; VERISIM_ENABLE_IPV4=true for dual-stack (0.0.0.0) - let default_host = if std::env::var("VERISIM_ENABLE_IPV4") - .map(|v| v == "true" || v == "1") - .unwrap_or(false) - { - "0.0.0.0".to_string() - } else { - "[::]".to_string() - }; - - let persist_dir = std::env::var("VERISIM_PERSISTENCE_DIR").ok(); - - let config = ApiConfig { - host: std::env::var("VERISIM_HOST").unwrap_or(default_host), - port: std::env::var("VERISIM_PORT") - .ok() - .and_then(|v| v.parse().ok()) - .unwrap_or(8080), - enable_cors: std::env::var("VERISIM_ENABLE_CORS") - .map(|v| v != "false" && v != "0") - .unwrap_or(true), - version_prefix: std::env::var("VERISIM_API_PREFIX") - .unwrap_or_else(|_| "/api/v1".to_string()), - vector_dimension: std::env::var("VERISIM_VECTOR_DIM") - .ok() - .and_then(|v| v.parse().ok()) - .unwrap_or(384), - persistence_dir: persist_dir.clone(), - grpc_port: std::env::var("VERISIM_GRPC_PORT") - .ok() - .and_then(|v| v.parse().ok()) - .unwrap_or(50051), - max_body_size: std::env::var("VERISIM_MAX_BODY_SIZE") - .ok() - .and_then(|v| v.parse().ok()) - .unwrap_or(10 * 1024 * 1024), - request_timeout_secs: std::env::var("VERISIM_TIMEOUT_SECS") - .ok() - .and_then(|v| v.parse().ok()) - .unwrap_or(30), - max_connections: std::env::var("VERISIM_MAX_CONNECTIONS") - .ok() - .and_then(|v| v.parse().ok()) - .unwrap_or(1024), - }; - - let storage_mode = if cfg!(feature = "persistent") { "persistent" } else { "in-memory" }; - - tracing::info!( - host = %config.host, - port = %config.port, - storage = %storage_mode, - persistence_dir = ?persist_dir, - "Starting VeriSimDB API server" - ); - - // Check for TLS configuration - let tls_cert = std::env::var("VERISIM_TLS_CERT").ok(); - let tls_key = std::env::var("VERISIM_TLS_KEY").ok(); - - match (tls_cert, tls_key) { - (Some(cert_path), Some(key_path)) => { - tracing::info!(cert = %cert_path, "Starting with TLS enabled"); - verisim_api::serve_tls(config, &cert_path, &key_path).await?; - } - (Some(_), None) | (None, Some(_)) => { - return Err("Both VERISIM_TLS_CERT and VERISIM_TLS_KEY must be set for TLS".into()); - } - (None, None) => { - verisim_api::serve(config).await?; - } - } - - Ok(()) -} diff --git a/verisimdb/rust-core/verisim-api/src/proof_attempts.rs b/verisimdb/rust-core/verisim-api/src/proof_attempts.rs deleted file mode 100644 index 67d5c4d2..00000000 --- a/verisimdb/rust-core/verisim-api/src/proof_attempts.rs +++ /dev/null @@ -1,338 +0,0 @@ -// SPDX-License-Identifier: MPL-2.0 -// SPDX-FileCopyrightText: 2026 Jonathan D.A. Jewell (hyperpolymath) -//! Proof-attempts API handlers. -//! -//! Three REST endpoints that bridge the proof-attempts pipeline to ClickHouse: -//! -//! | Method | Path | Purpose | -//! |--------|-------------------------------------------|-----------------------------------| -//! | GET | /proof_attempts?limit=N&offset=M | List attempt rows (paged) | -//! | POST | /proof_attempts | Insert a single attempt row | -//! | GET | /proof_attempts/strategy?class=X&limit=N | Recommend best provers for class | -//! | GET | /proof_attempts/certificates?class=X | PROVEN/pending cert status | -//! -//! Responses are emitted as A2ML (text/a2ml) — never JSON. ClickHouse is -//! still spoken to over its JSONEachRow wire format internally, but that is -//! parsing only (inbound from ClickHouse, never emitted to callers). -//! -//! The ClickHouse URL is read from `VERISIM_CLICKHOUSE_URL` at request time. -//! Optional write auth: if `VERISIM_PROOF_ATTEMPTS_TOKEN` is set, callers must -//! provide either `X-Proof-Attempts-Token: ` or -//! `Authorization: Bearer ` on `POST /proof_attempts`. - -use axum::{ - extract::Query, - http::{HeaderMap, StatusCode}, - response::Response, -}; -use reqwest::Client; -use serde::{Deserialize, Serialize}; -use std::time::Duration; -use tracing::{debug, warn}; - -use crate::a2ml::{ - a2ml_error, a2ml_error_detail, a2ml_response, - certificates_to_a2ml, inserted_to_a2ml, - parse_certs, parse_recommendations, proof_attempts_to_a2ml, - strategy_to_a2ml, -}; - -// ── Shared HTTP client (one per module, constructed lazily) ────────────────── - -fn ch_client() -> &'static Client { - use std::sync::OnceLock; - static CLIENT: OnceLock = OnceLock::new(); - CLIENT.get_or_init(|| { - Client::builder() - .timeout(Duration::from_secs(10)) - .build() - .expect("failed to build ClickHouse HTTP client") - }) -} - -fn ch_url() -> String { - std::env::var("VERISIM_CLICKHOUSE_URL") - .unwrap_or_else(|_| "http://localhost:8123".to_string()) -} - -fn required_insert_token() -> Option { - std::env::var("VERISIM_PROOF_ATTEMPTS_TOKEN") - .ok() - .map(|s| s.trim().to_string()) - .filter(|s| !s.is_empty()) -} - -fn provided_insert_token(headers: &HeaderMap) -> Option { - if let Some(v) = headers - .get("x-proof-attempts-token") - .and_then(|v| v.to_str().ok()) - { - return Some(v.trim().to_string()); - } - - headers - .get("authorization") - .and_then(|v| v.to_str().ok()) - .and_then(|v| v.strip_prefix("Bearer ")) - .map(|v| v.trim().to_string()) -} - -// ── Inbound proof-attempt row (matches echidnabot VeriSimWriter schema) ─────── - -/// A single proof attempt submitted by echidnabot. -#[derive(Debug, Deserialize, Serialize)] -pub struct ProofAttemptRow { - pub attempt_id: String, - pub obligation_id: String, - pub repo: String, - pub file: String, - pub claim: String, - pub obligation_class: String, - pub prover_used: String, - pub outcome: String, - pub duration_ms: u64, - pub confidence: f64, - pub parent_attempt_id: Option, - pub strategy_tag: String, - pub started_at: String, - pub completed_at: String, - pub prover_output: String, - pub error_message: Option, -} - -// ── Query params ───────────────────────────────────────────────────────────── - -#[derive(Debug, Deserialize)] -pub struct ListParams { - pub limit: Option, - pub offset: Option, -} - -#[derive(Debug, Deserialize)] -pub struct StrategyParams { - pub class: String, - pub limit: Option, -} - -#[derive(Debug, Deserialize)] -pub struct ClassParam { - pub class: String, -} - -// ── Handlers ───────────────────────────────────────────────────────────────── - -/// GET /proof_attempts?limit=N&offset=M -/// -/// Returns up to `limit` (default 1000, max 20000) recent proof-attempt rows -/// from ClickHouse as an A2ML document for retraining the Julia ML models. -pub async fn list_proof_attempts( - Query(params): Query, -) -> Response { - let limit = params.limit.unwrap_or(1000).clamp(1, 20000); - let offset = params.offset.unwrap_or(0); - - let sql = format!( - "SELECT attempt_id, obligation_id, repo, file, claim, obligation_class, \ - prover_used, outcome, duration_ms, confidence, parent_attempt_id, \ - strategy_tag, started_at, completed_at \ - FROM verisim.proof_attempts \ - ORDER BY started_at DESC \ - LIMIT {limit} \ - OFFSET {offset} \ - FORMAT JSONEachRow" - ); - - match ch_client() - .post(ch_url()) - .header("Content-Type", "text/plain") - .body(sql) - .send() - .await - { - Ok(resp) if resp.status().is_success() => { - let text = resp.text().await.unwrap_or_default(); - // Parse ClickHouse JSONEachRow internally (never emitted as JSON) - let rows: Vec = text - .lines() - .filter(|l| !l.trim().is_empty()) - .filter_map(|l| serde_json::from_str(l).ok()) - .collect(); - a2ml_response(StatusCode::OK, proof_attempts_to_a2ml(&rows)) - } - Ok(resp) => { - let status = resp.status().as_u16(); - let body = resp.text().await.unwrap_or_default(); - warn!("proof_attempts list: ClickHouse {status}: {body}"); - a2ml_response(StatusCode::BAD_GATEWAY, a2ml_error("clickhouse_error", status)) - } - Err(e) => { - warn!("proof_attempts list: unreachable: {e}"); - a2ml_response(StatusCode::SERVICE_UNAVAILABLE, a2ml_error("clickhouse_unreachable", 503)) - } - } -} - -/// POST /proof_attempts -/// -/// Inserts a single attempt row into `verisim.proof_attempts` via the -/// ClickHouse HTTP INSERT … FORMAT JSONEachRow endpoint. -pub async fn insert_proof_attempt( - headers: HeaderMap, - axum::Json(row): axum::Json, -) -> Response { - if let Some(expected) = required_insert_token() { - match provided_insert_token(&headers) { - Some(provided) if provided == expected => {} - Some(_) => { - return a2ml_response( - StatusCode::UNAUTHORIZED, - a2ml_error_detail("invalid_proof_attempts_token", 401, "Token mismatch"), - ); - } - None => { - return a2ml_response( - StatusCode::UNAUTHORIZED, - a2ml_error_detail( - "proof_attempts_auth_required", - 401, - "Set X-Proof-Attempts-Token or Authorization: Bearer", - ), - ); - } - } - } - - let url = format!( - "{}/?query=INSERT+INTO+verisim.proof_attempts+FORMAT+JSONEachRow", - ch_url() - ); - - // Serialise the inbound row to JSONEachRow for the ClickHouse wire format. - // This is internal I/O to ClickHouse, not an outbound response. - let body = match serde_json::to_string(&row) { - Ok(s) => s, - Err(e) => { - warn!("proof_attempts: failed to serialise row: {e}"); - return a2ml_response( - StatusCode::BAD_REQUEST, - a2ml_error_detail("serialisation_failed", 400, &e.to_string()), - ); - } - }; - - debug!(attempt_id = %row.attempt_id, prover = %row.prover_used, "inserting proof attempt"); - - match ch_client() - .post(&url) - .header("Content-Type", "application/json") - .body(body) - .send() - .await - { - Ok(resp) if resp.status().is_success() => { - a2ml_response(StatusCode::CREATED, inserted_to_a2ml(&row.attempt_id)) - } - Ok(resp) => { - let status = resp.status().as_u16(); - let detail = resp.text().await.unwrap_or_default(); - warn!("proof_attempts: ClickHouse returned {status}: {detail}"); - a2ml_response( - StatusCode::BAD_GATEWAY, - a2ml_error_detail("clickhouse_error", status, &detail), - ) - } - Err(e) => { - warn!("proof_attempts: ClickHouse unreachable: {e}"); - a2ml_response( - StatusCode::SERVICE_UNAVAILABLE, - a2ml_error_detail("clickhouse_unreachable", 503, &e.to_string()), - ) - } - } -} - -/// GET /proof_attempts/strategy?class=X&limit=N -/// -/// Returns up to `limit` (default 5) prover recommendations for the given -/// obligation class, ordered by success rate descending. Queries the -/// `mv_proven_certificates` view which pre-computes success_rate and -/// avg_duration_ms so no Rust arithmetic is needed. -pub async fn strategy( - Query(params): Query, -) -> Response { - let limit = params.limit.unwrap_or(5).min(50); - let class = params.class.replace('\'', "\\'"); // minimal SQL escape - - let sql = format!( - "SELECT prover_used, success_rate, avg_duration_ms, total_attempts \ - FROM verisim.mv_proven_certificates \ - WHERE obligation_class = '{class}' \ - ORDER BY success_rate DESC, avg_duration_ms ASC \ - LIMIT {limit} \ - FORMAT JSONEachRow" - ); - - match ch_client() - .post(ch_url()) - .header("Content-Type", "text/plain") - .body(sql) - .send() - .await - { - Ok(resp) if resp.status().is_success() => { - let text = resp.text().await.unwrap_or_default(); - let recommendations = parse_recommendations(&text); - a2ml_response(StatusCode::OK, strategy_to_a2ml(&recommendations)) - } - Ok(resp) => { - let status = resp.status().as_u16(); - let body = resp.text().await.unwrap_or_default(); - warn!("proof_attempts/strategy: ClickHouse {status}: {body}"); - a2ml_response(StatusCode::BAD_GATEWAY, a2ml_error("clickhouse_error", status)) - } - Err(e) => { - warn!("proof_attempts/strategy: unreachable: {e}"); - a2ml_response(StatusCode::SERVICE_UNAVAILABLE, a2ml_error("clickhouse_unreachable", 503)) - } - } -} - -/// GET /proof_attempts/certificates?class=X -/// -/// Returns PROVEN / pending status for every (class, prover) pair from -/// `mv_proven_certificates`. -pub async fn certificates( - Query(params): Query, -) -> Response { - let class = params.class.replace('\'', "\\'"); - - let sql = format!( - "SELECT prover_used, status, success_rate, total_attempts \ - FROM verisim.mv_proven_certificates \ - WHERE obligation_class = '{class}' \ - FORMAT JSONEachRow" - ); - - match ch_client() - .post(ch_url()) - .header("Content-Type", "text/plain") - .body(sql) - .send() - .await - { - Ok(resp) if resp.status().is_success() => { - let text = resp.text().await.unwrap_or_default(); - let rows = parse_certs(&text); - a2ml_response(StatusCode::OK, certificates_to_a2ml(&rows)) - } - Ok(resp) => { - let status = resp.status().as_u16(); - warn!("proof_attempts/certificates: ClickHouse {status}"); - a2ml_response(StatusCode::BAD_GATEWAY, a2ml_error("clickhouse_error", status)) - } - Err(e) => { - warn!("proof_attempts/certificates: unreachable: {e}"); - a2ml_response(StatusCode::SERVICE_UNAVAILABLE, a2ml_error("clickhouse_unreachable", 503)) - } - } -} diff --git a/verisimdb/rust-core/verisim-api/src/proto/verisim.rs b/verisimdb/rust-core/verisim-api/src/proto/verisim.rs deleted file mode 100644 index 92de8dca..00000000 --- a/verisimdb/rust-core/verisim-api/src/proto/verisim.rs +++ /dev/null @@ -1,1445 +0,0 @@ -// SPDX-License-Identifier: MPL-2.0 -// This file is @generated by prost-build from proto/verisim.proto. -// Do NOT edit manually — regenerate with: cargo build -p verisim-api (requires protoc) -// Pre-generated to eliminate the protoc build dependency. -#[derive(Clone, Copy, PartialEq, Eq, Hash, ::prost::Message)] -pub struct Empty {} -/// Logical plan passed as JSON string (mirrors the REST API). -/// This avoids duplicating the full plan type hierarchy in protobuf. -#[derive(Clone, PartialEq, Eq, Hash, ::prost::Message)] -pub struct LogicalPlanRequest { - #[prost(string, tag = "1")] - pub plan_json: ::prost::alloc::string::String, -} -#[derive(Clone, PartialEq, ::prost::Message)] -pub struct PhysicalPlanResponse { - #[prost(message, repeated, tag = "1")] - pub steps: ::prost::alloc::vec::Vec, - #[prost(string, tag = "2")] - pub strategy: ::prost::alloc::string::String, - #[prost(double, tag = "3")] - pub total_time_ms: f64, - #[prost(uint64, tag = "4")] - pub total_estimated_rows: u64, - #[prost(string, repeated, tag = "5")] - pub notes: ::prost::alloc::vec::Vec<::prost::alloc::string::String>, -} -#[derive(Clone, PartialEq, ::prost::Message)] -pub struct PlanStepMsg { - #[prost(int32, tag = "1")] - pub step: i32, - #[prost(string, tag = "2")] - pub operation: ::prost::alloc::string::String, - #[prost(string, tag = "3")] - pub modality: ::prost::alloc::string::String, - #[prost(double, tag = "4")] - pub time_ms: f64, - #[prost(uint64, tag = "5")] - pub estimated_rows: u64, - #[prost(double, tag = "6")] - pub selectivity: f64, - #[prost(string, tag = "7")] - pub optimization_hint: ::prost::alloc::string::String, -} -#[derive(Clone, PartialEq, ::prost::Message)] -pub struct ExplainResponse { - #[prost(message, repeated, tag = "1")] - pub steps: ::prost::alloc::vec::Vec, - #[prost(message, repeated, tag = "2")] - pub cost_breakdown: ::prost::alloc::vec::Vec, - #[prost(message, repeated, tag = "3")] - pub performance_hints: ::prost::alloc::vec::Vec, - #[prost(double, tag = "4")] - pub total_cost_ms: f64, - #[prost(string, tag = "5")] - pub strategy: ::prost::alloc::string::String, - #[prost(string, tag = "6")] - pub text_output: ::prost::alloc::string::String, -} -#[derive(Clone, PartialEq, ::prost::Message)] -pub struct ModalityCostMsg { - #[prost(string, tag = "1")] - pub modality: ::prost::alloc::string::String, - #[prost(double, tag = "2")] - pub time_ms: f64, - #[prost(double, tag = "3")] - pub percentage: f64, -} -#[derive(Clone, PartialEq, Eq, Hash, ::prost::Message)] -pub struct PerformanceHintMsg { - #[prost(string, tag = "1")] - pub severity: ::prost::alloc::string::String, - #[prost(string, tag = "2")] - pub message: ::prost::alloc::string::String, -} -#[derive(Clone, PartialEq, ::prost::Message)] -pub struct PlannerConfigRequest { - #[prost(string, tag = "1")] - pub global_mode: ::prost::alloc::string::String, - #[prost(double, tag = "2")] - pub statistics_weight: f64, - #[prost(bool, tag = "3")] - pub enable_adaptive: bool, - #[prost(int32, tag = "4")] - pub parallel_threshold: i32, -} -#[derive(Clone, PartialEq, ::prost::Message)] -pub struct PlannerConfigResponse { - #[prost(string, tag = "1")] - pub global_mode: ::prost::alloc::string::String, - #[prost(double, tag = "2")] - pub statistics_weight: f64, - #[prost(bool, tag = "3")] - pub enable_adaptive: bool, - #[prost(int32, tag = "4")] - pub parallel_threshold: i32, -} -#[derive(Clone, PartialEq, ::prost::Message)] -pub struct StoreStatsMsg { - #[prost(string, tag = "1")] - pub modality: ::prost::alloc::string::String, - #[prost(uint64, tag = "2")] - pub total_rows: u64, - #[prost(double, tag = "3")] - pub avg_latency_ms: f64, - #[prost(uint64, tag = "4")] - pub avg_rows_returned: u64, - #[prost(uint64, tag = "5")] - pub query_count: u64, -} -#[derive(Clone, PartialEq, ::prost::Message)] -pub struct StatsResponse { - #[prost(message, repeated, tag = "1")] - pub stores: ::prost::alloc::vec::Vec, -} -#[derive(Clone, PartialEq, ::prost::Message)] -pub struct OctadCreateRequest { - #[prost(string, tag = "1")] - pub title: ::prost::alloc::string::String, - #[prost(string, tag = "2")] - pub body: ::prost::alloc::string::String, - #[prost(float, repeated, tag = "3")] - pub embedding: ::prost::alloc::vec::Vec, - #[prost(string, repeated, tag = "4")] - pub types: ::prost::alloc::vec::Vec<::prost::alloc::string::String>, -} -#[derive(Clone, PartialEq, Eq, Hash, ::prost::Message)] -pub struct OctadIdRequest { - #[prost(string, tag = "1")] - pub id: ::prost::alloc::string::String, -} -#[derive(Clone, PartialEq, ::prost::Message)] -pub struct OctadUpdateRequest { - #[prost(string, tag = "1")] - pub id: ::prost::alloc::string::String, - #[prost(string, tag = "2")] - pub title: ::prost::alloc::string::String, - #[prost(string, tag = "3")] - pub body: ::prost::alloc::string::String, - #[prost(float, repeated, tag = "4")] - pub embedding: ::prost::alloc::vec::Vec, - #[prost(string, repeated, tag = "5")] - pub types: ::prost::alloc::vec::Vec<::prost::alloc::string::String>, -} -#[derive(Clone, PartialEq, Eq, Hash, ::prost::Message)] -pub struct OctadResponse { - #[prost(string, tag = "1")] - pub id: ::prost::alloc::string::String, - #[prost(string, tag = "2")] - pub created_at: ::prost::alloc::string::String, - #[prost(string, tag = "3")] - pub modified_at: ::prost::alloc::string::String, - #[prost(uint64, tag = "4")] - pub version: u64, - #[prost(bool, tag = "5")] - pub has_graph: bool, - #[prost(bool, tag = "6")] - pub has_vector: bool, - #[prost(bool, tag = "7")] - pub has_tensor: bool, - #[prost(bool, tag = "8")] - pub has_semantic: bool, - #[prost(bool, tag = "9")] - pub has_document: bool, - #[prost(uint64, tag = "10")] - pub version_count: u64, -} -#[derive(Clone, PartialEq, Eq, Hash, ::prost::Message)] -pub struct TextSearchRequest { - #[prost(string, tag = "1")] - pub query: ::prost::alloc::string::String, - #[prost(int32, tag = "2")] - pub limit: i32, -} -#[derive(Clone, PartialEq, ::prost::Message)] -pub struct VectorSearchRequest { - #[prost(float, repeated, tag = "1")] - pub vector: ::prost::alloc::vec::Vec, - #[prost(int32, tag = "2")] - pub k: i32, -} -#[derive(Clone, PartialEq, ::prost::Message)] -pub struct SearchResultMsg { - #[prost(string, tag = "1")] - pub id: ::prost::alloc::string::String, - #[prost(float, tag = "2")] - pub score: f32, - #[prost(string, tag = "3")] - pub title: ::prost::alloc::string::String, -} -#[derive(Clone, PartialEq, ::prost::Message)] -pub struct SearchResponse { - #[prost(message, repeated, tag = "1")] - pub results: ::prost::alloc::vec::Vec, -} -/// Generated client implementations. -pub mod veri_sim_planner_client { - #![allow( - unused_variables, - dead_code, - missing_docs, - clippy::wildcard_imports, - clippy::let_unit_value, - )] - use tonic::codegen::*; - use tonic::codegen::http::Uri; - #[derive(Debug, Clone)] - pub struct VeriSimPlannerClient { - inner: tonic::client::Grpc, - } - impl VeriSimPlannerClient { - /// Attempt to create a new client by connecting to a given endpoint. - pub async fn connect(dst: D) -> Result - where - D: TryInto, - D::Error: Into, - { - let conn = tonic::transport::Endpoint::new(dst)?.connect().await?; - Ok(Self::new(conn)) - } - } - impl VeriSimPlannerClient - where - T: tonic::client::GrpcService, - T::Error: Into, - T::ResponseBody: Body + std::marker::Send + 'static, - ::Error: Into + std::marker::Send, - { - pub fn new(inner: T) -> Self { - let inner = tonic::client::Grpc::new(inner); - Self { inner } - } - pub fn with_origin(inner: T, origin: Uri) -> Self { - let inner = tonic::client::Grpc::with_origin(inner, origin); - Self { inner } - } - pub fn with_interceptor( - inner: T, - interceptor: F, - ) -> VeriSimPlannerClient> - where - F: tonic::service::Interceptor, - T::ResponseBody: Default, - T: tonic::codegen::Service< - http::Request, - Response = http::Response< - >::ResponseBody, - >, - >, - , - >>::Error: Into + std::marker::Send + std::marker::Sync, - { - VeriSimPlannerClient::new(InterceptedService::new(inner, interceptor)) - } - /// Compress requests with the given encoding. - /// - /// This requires the server to support it otherwise it might respond with an - /// error. - #[must_use] - pub fn send_compressed(mut self, encoding: CompressionEncoding) -> Self { - self.inner = self.inner.send_compressed(encoding); - self - } - /// Enable decompressing responses. - #[must_use] - pub fn accept_compressed(mut self, encoding: CompressionEncoding) -> Self { - self.inner = self.inner.accept_compressed(encoding); - self - } - /// Limits the maximum size of a decoded message. - /// - /// Default: `4MB` - #[must_use] - pub fn max_decoding_message_size(mut self, limit: usize) -> Self { - self.inner = self.inner.max_decoding_message_size(limit); - self - } - /// Limits the maximum size of an encoded message. - /// - /// Default: `usize::MAX` - #[must_use] - pub fn max_encoding_message_size(mut self, limit: usize) -> Self { - self.inner = self.inner.max_encoding_message_size(limit); - self - } - /// Optimize a logical plan into a physical plan. - pub async fn optimize_plan( - &mut self, - request: impl tonic::IntoRequest, - ) -> std::result::Result< - tonic::Response, - tonic::Status, - > { - self.inner - .ready() - .await - .map_err(|e| { - tonic::Status::unknown( - format!("Service was not ready: {}", e.into()), - ) - })?; - let codec = tonic_prost::ProstCodec::default(); - let path = http::uri::PathAndQuery::from_static( - "/verisim.VeriSimPlanner/OptimizePlan", - ); - let mut req = request.into_request(); - req.extensions_mut() - .insert(GrpcMethod::new("verisim.VeriSimPlanner", "OptimizePlan")); - self.inner.unary(req, path, codec).await - } - /// Generate EXPLAIN output for a logical plan. - pub async fn explain_plan( - &mut self, - request: impl tonic::IntoRequest, - ) -> std::result::Result< - tonic::Response, - tonic::Status, - > { - self.inner - .ready() - .await - .map_err(|e| { - tonic::Status::unknown( - format!("Service was not ready: {}", e.into()), - ) - })?; - let codec = tonic_prost::ProstCodec::default(); - let path = http::uri::PathAndQuery::from_static( - "/verisim.VeriSimPlanner/ExplainPlan", - ); - let mut req = request.into_request(); - req.extensions_mut() - .insert(GrpcMethod::new("verisim.VeriSimPlanner", "ExplainPlan")); - self.inner.unary(req, path, codec).await - } - /// Get current planner configuration. - pub async fn get_config( - &mut self, - request: impl tonic::IntoRequest, - ) -> std::result::Result< - tonic::Response, - tonic::Status, - > { - self.inner - .ready() - .await - .map_err(|e| { - tonic::Status::unknown( - format!("Service was not ready: {}", e.into()), - ) - })?; - let codec = tonic_prost::ProstCodec::default(); - let path = http::uri::PathAndQuery::from_static( - "/verisim.VeriSimPlanner/GetConfig", - ); - let mut req = request.into_request(); - req.extensions_mut() - .insert(GrpcMethod::new("verisim.VeriSimPlanner", "GetConfig")); - self.inner.unary(req, path, codec).await - } - /// Update planner configuration. - pub async fn set_config( - &mut self, - request: impl tonic::IntoRequest, - ) -> std::result::Result< - tonic::Response, - tonic::Status, - > { - self.inner - .ready() - .await - .map_err(|e| { - tonic::Status::unknown( - format!("Service was not ready: {}", e.into()), - ) - })?; - let codec = tonic_prost::ProstCodec::default(); - let path = http::uri::PathAndQuery::from_static( - "/verisim.VeriSimPlanner/SetConfig", - ); - let mut req = request.into_request(); - req.extensions_mut() - .insert(GrpcMethod::new("verisim.VeriSimPlanner", "SetConfig")); - self.inner.unary(req, path, codec).await - } - /// Get per-modality statistics snapshot. - pub async fn get_stats( - &mut self, - request: impl tonic::IntoRequest, - ) -> std::result::Result, tonic::Status> { - self.inner - .ready() - .await - .map_err(|e| { - tonic::Status::unknown( - format!("Service was not ready: {}", e.into()), - ) - })?; - let codec = tonic_prost::ProstCodec::default(); - let path = http::uri::PathAndQuery::from_static( - "/verisim.VeriSimPlanner/GetStats", - ); - let mut req = request.into_request(); - req.extensions_mut() - .insert(GrpcMethod::new("verisim.VeriSimPlanner", "GetStats")); - self.inner.unary(req, path, codec).await - } - } -} -/// Generated server implementations. -pub mod veri_sim_planner_server { - #![allow( - unused_variables, - dead_code, - missing_docs, - clippy::wildcard_imports, - clippy::let_unit_value, - )] - use tonic::codegen::*; - /// Generated trait containing gRPC methods that should be implemented for use with VeriSimPlannerServer. - #[async_trait] - pub trait VeriSimPlanner: std::marker::Send + std::marker::Sync + 'static { - /// Optimize a logical plan into a physical plan. - async fn optimize_plan( - &self, - request: tonic::Request, - ) -> std::result::Result< - tonic::Response, - tonic::Status, - >; - /// Generate EXPLAIN output for a logical plan. - async fn explain_plan( - &self, - request: tonic::Request, - ) -> std::result::Result, tonic::Status>; - /// Get current planner configuration. - async fn get_config( - &self, - request: tonic::Request, - ) -> std::result::Result< - tonic::Response, - tonic::Status, - >; - /// Update planner configuration. - async fn set_config( - &self, - request: tonic::Request, - ) -> std::result::Result< - tonic::Response, - tonic::Status, - >; - /// Get per-modality statistics snapshot. - async fn get_stats( - &self, - request: tonic::Request, - ) -> std::result::Result, tonic::Status>; - } - #[derive(Debug)] - pub struct VeriSimPlannerServer { - inner: Arc, - accept_compression_encodings: EnabledCompressionEncodings, - send_compression_encodings: EnabledCompressionEncodings, - max_decoding_message_size: Option, - max_encoding_message_size: Option, - } - impl VeriSimPlannerServer { - pub fn new(inner: T) -> Self { - Self::from_arc(Arc::new(inner)) - } - pub fn from_arc(inner: Arc) -> Self { - Self { - inner, - accept_compression_encodings: Default::default(), - send_compression_encodings: Default::default(), - max_decoding_message_size: None, - max_encoding_message_size: None, - } - } - pub fn with_interceptor( - inner: T, - interceptor: F, - ) -> InterceptedService - where - F: tonic::service::Interceptor, - { - InterceptedService::new(Self::new(inner), interceptor) - } - /// Enable decompressing requests with the given encoding. - #[must_use] - pub fn accept_compressed(mut self, encoding: CompressionEncoding) -> Self { - self.accept_compression_encodings.enable(encoding); - self - } - /// Compress responses with the given encoding, if the client supports it. - #[must_use] - pub fn send_compressed(mut self, encoding: CompressionEncoding) -> Self { - self.send_compression_encodings.enable(encoding); - self - } - /// Limits the maximum size of a decoded message. - /// - /// Default: `4MB` - #[must_use] - pub fn max_decoding_message_size(mut self, limit: usize) -> Self { - self.max_decoding_message_size = Some(limit); - self - } - /// Limits the maximum size of an encoded message. - /// - /// Default: `usize::MAX` - #[must_use] - pub fn max_encoding_message_size(mut self, limit: usize) -> Self { - self.max_encoding_message_size = Some(limit); - self - } - } - impl tonic::codegen::Service> for VeriSimPlannerServer - where - T: VeriSimPlanner, - B: Body + std::marker::Send + 'static, - B::Error: Into + std::marker::Send + 'static, - { - type Response = http::Response; - type Error = std::convert::Infallible; - type Future = BoxFuture; - fn poll_ready( - &mut self, - _cx: &mut Context<'_>, - ) -> Poll> { - Poll::Ready(Ok(())) - } - fn call(&mut self, req: http::Request) -> Self::Future { - match req.uri().path() { - "/verisim.VeriSimPlanner/OptimizePlan" => { - #[allow(non_camel_case_types)] - struct OptimizePlanSvc(pub Arc); - impl< - T: VeriSimPlanner, - > tonic::server::UnaryService - for OptimizePlanSvc { - type Response = super::PhysicalPlanResponse; - type Future = BoxFuture< - tonic::Response, - tonic::Status, - >; - fn call( - &mut self, - request: tonic::Request, - ) -> Self::Future { - let inner = Arc::clone(&self.0); - let fut = async move { - ::optimize_plan(&inner, request).await - }; - Box::pin(fut) - } - } - let accept_compression_encodings = self.accept_compression_encodings; - let send_compression_encodings = self.send_compression_encodings; - let max_decoding_message_size = self.max_decoding_message_size; - let max_encoding_message_size = self.max_encoding_message_size; - let inner = self.inner.clone(); - let fut = async move { - let method = OptimizePlanSvc(inner); - let codec = tonic_prost::ProstCodec::default(); - let mut grpc = tonic::server::Grpc::new(codec) - .apply_compression_config( - accept_compression_encodings, - send_compression_encodings, - ) - .apply_max_message_size_config( - max_decoding_message_size, - max_encoding_message_size, - ); - let res = grpc.unary(method, req).await; - Ok(res) - }; - Box::pin(fut) - } - "/verisim.VeriSimPlanner/ExplainPlan" => { - #[allow(non_camel_case_types)] - struct ExplainPlanSvc(pub Arc); - impl< - T: VeriSimPlanner, - > tonic::server::UnaryService - for ExplainPlanSvc { - type Response = super::ExplainResponse; - type Future = BoxFuture< - tonic::Response, - tonic::Status, - >; - fn call( - &mut self, - request: tonic::Request, - ) -> Self::Future { - let inner = Arc::clone(&self.0); - let fut = async move { - ::explain_plan(&inner, request).await - }; - Box::pin(fut) - } - } - let accept_compression_encodings = self.accept_compression_encodings; - let send_compression_encodings = self.send_compression_encodings; - let max_decoding_message_size = self.max_decoding_message_size; - let max_encoding_message_size = self.max_encoding_message_size; - let inner = self.inner.clone(); - let fut = async move { - let method = ExplainPlanSvc(inner); - let codec = tonic_prost::ProstCodec::default(); - let mut grpc = tonic::server::Grpc::new(codec) - .apply_compression_config( - accept_compression_encodings, - send_compression_encodings, - ) - .apply_max_message_size_config( - max_decoding_message_size, - max_encoding_message_size, - ); - let res = grpc.unary(method, req).await; - Ok(res) - }; - Box::pin(fut) - } - "/verisim.VeriSimPlanner/GetConfig" => { - #[allow(non_camel_case_types)] - struct GetConfigSvc(pub Arc); - impl tonic::server::UnaryService - for GetConfigSvc { - type Response = super::PlannerConfigResponse; - type Future = BoxFuture< - tonic::Response, - tonic::Status, - >; - fn call( - &mut self, - request: tonic::Request, - ) -> Self::Future { - let inner = Arc::clone(&self.0); - let fut = async move { - ::get_config(&inner, request).await - }; - Box::pin(fut) - } - } - let accept_compression_encodings = self.accept_compression_encodings; - let send_compression_encodings = self.send_compression_encodings; - let max_decoding_message_size = self.max_decoding_message_size; - let max_encoding_message_size = self.max_encoding_message_size; - let inner = self.inner.clone(); - let fut = async move { - let method = GetConfigSvc(inner); - let codec = tonic_prost::ProstCodec::default(); - let mut grpc = tonic::server::Grpc::new(codec) - .apply_compression_config( - accept_compression_encodings, - send_compression_encodings, - ) - .apply_max_message_size_config( - max_decoding_message_size, - max_encoding_message_size, - ); - let res = grpc.unary(method, req).await; - Ok(res) - }; - Box::pin(fut) - } - "/verisim.VeriSimPlanner/SetConfig" => { - #[allow(non_camel_case_types)] - struct SetConfigSvc(pub Arc); - impl< - T: VeriSimPlanner, - > tonic::server::UnaryService - for SetConfigSvc { - type Response = super::PlannerConfigResponse; - type Future = BoxFuture< - tonic::Response, - tonic::Status, - >; - fn call( - &mut self, - request: tonic::Request, - ) -> Self::Future { - let inner = Arc::clone(&self.0); - let fut = async move { - ::set_config(&inner, request).await - }; - Box::pin(fut) - } - } - let accept_compression_encodings = self.accept_compression_encodings; - let send_compression_encodings = self.send_compression_encodings; - let max_decoding_message_size = self.max_decoding_message_size; - let max_encoding_message_size = self.max_encoding_message_size; - let inner = self.inner.clone(); - let fut = async move { - let method = SetConfigSvc(inner); - let codec = tonic_prost::ProstCodec::default(); - let mut grpc = tonic::server::Grpc::new(codec) - .apply_compression_config( - accept_compression_encodings, - send_compression_encodings, - ) - .apply_max_message_size_config( - max_decoding_message_size, - max_encoding_message_size, - ); - let res = grpc.unary(method, req).await; - Ok(res) - }; - Box::pin(fut) - } - "/verisim.VeriSimPlanner/GetStats" => { - #[allow(non_camel_case_types)] - struct GetStatsSvc(pub Arc); - impl tonic::server::UnaryService - for GetStatsSvc { - type Response = super::StatsResponse; - type Future = BoxFuture< - tonic::Response, - tonic::Status, - >; - fn call( - &mut self, - request: tonic::Request, - ) -> Self::Future { - let inner = Arc::clone(&self.0); - let fut = async move { - ::get_stats(&inner, request).await - }; - Box::pin(fut) - } - } - let accept_compression_encodings = self.accept_compression_encodings; - let send_compression_encodings = self.send_compression_encodings; - let max_decoding_message_size = self.max_decoding_message_size; - let max_encoding_message_size = self.max_encoding_message_size; - let inner = self.inner.clone(); - let fut = async move { - let method = GetStatsSvc(inner); - let codec = tonic_prost::ProstCodec::default(); - let mut grpc = tonic::server::Grpc::new(codec) - .apply_compression_config( - accept_compression_encodings, - send_compression_encodings, - ) - .apply_max_message_size_config( - max_decoding_message_size, - max_encoding_message_size, - ); - let res = grpc.unary(method, req).await; - Ok(res) - }; - Box::pin(fut) - } - _ => { - Box::pin(async move { - let mut response = http::Response::new( - tonic::body::Body::default(), - ); - let headers = response.headers_mut(); - headers - .insert( - tonic::Status::GRPC_STATUS, - (tonic::Code::Unimplemented as i32).into(), - ); - headers - .insert( - http::header::CONTENT_TYPE, - tonic::metadata::GRPC_CONTENT_TYPE, - ); - Ok(response) - }) - } - } - } - } - impl Clone for VeriSimPlannerServer { - fn clone(&self) -> Self { - let inner = self.inner.clone(); - Self { - inner, - accept_compression_encodings: self.accept_compression_encodings, - send_compression_encodings: self.send_compression_encodings, - max_decoding_message_size: self.max_decoding_message_size, - max_encoding_message_size: self.max_encoding_message_size, - } - } - } - /// Generated gRPC service name - pub const SERVICE_NAME: &str = "verisim.VeriSimPlanner"; - impl tonic::server::NamedService for VeriSimPlannerServer { - const NAME: &'static str = SERVICE_NAME; - } -} -/// Generated client implementations. -pub mod veri_sim_octad_client { - #![allow( - unused_variables, - dead_code, - missing_docs, - clippy::wildcard_imports, - clippy::let_unit_value, - )] - use tonic::codegen::*; - use tonic::codegen::http::Uri; - #[derive(Debug, Clone)] - pub struct VeriSimOctadClient { - inner: tonic::client::Grpc, - } - impl VeriSimOctadClient { - /// Attempt to create a new client by connecting to a given endpoint. - pub async fn connect(dst: D) -> Result - where - D: TryInto, - D::Error: Into, - { - let conn = tonic::transport::Endpoint::new(dst)?.connect().await?; - Ok(Self::new(conn)) - } - } - impl VeriSimOctadClient - where - T: tonic::client::GrpcService, - T::Error: Into, - T::ResponseBody: Body + std::marker::Send + 'static, - ::Error: Into + std::marker::Send, - { - pub fn new(inner: T) -> Self { - let inner = tonic::client::Grpc::new(inner); - Self { inner } - } - pub fn with_origin(inner: T, origin: Uri) -> Self { - let inner = tonic::client::Grpc::with_origin(inner, origin); - Self { inner } - } - pub fn with_interceptor( - inner: T, - interceptor: F, - ) -> VeriSimOctadClient> - where - F: tonic::service::Interceptor, - T::ResponseBody: Default, - T: tonic::codegen::Service< - http::Request, - Response = http::Response< - >::ResponseBody, - >, - >, - , - >>::Error: Into + std::marker::Send + std::marker::Sync, - { - VeriSimOctadClient::new(InterceptedService::new(inner, interceptor)) - } - /// Compress requests with the given encoding. - /// - /// This requires the server to support it otherwise it might respond with an - /// error. - #[must_use] - pub fn send_compressed(mut self, encoding: CompressionEncoding) -> Self { - self.inner = self.inner.send_compressed(encoding); - self - } - /// Enable decompressing responses. - #[must_use] - pub fn accept_compressed(mut self, encoding: CompressionEncoding) -> Self { - self.inner = self.inner.accept_compressed(encoding); - self - } - /// Limits the maximum size of a decoded message. - /// - /// Default: `4MB` - #[must_use] - pub fn max_decoding_message_size(mut self, limit: usize) -> Self { - self.inner = self.inner.max_decoding_message_size(limit); - self - } - /// Limits the maximum size of an encoded message. - /// - /// Default: `usize::MAX` - #[must_use] - pub fn max_encoding_message_size(mut self, limit: usize) -> Self { - self.inner = self.inner.max_encoding_message_size(limit); - self - } - /// Create a new octad. - pub async fn create( - &mut self, - request: impl tonic::IntoRequest, - ) -> std::result::Result, tonic::Status> { - self.inner - .ready() - .await - .map_err(|e| { - tonic::Status::unknown( - format!("Service was not ready: {}", e.into()), - ) - })?; - let codec = tonic_prost::ProstCodec::default(); - let path = http::uri::PathAndQuery::from_static( - "/verisim.VeriSimOctad/Create", - ); - let mut req = request.into_request(); - req.extensions_mut() - .insert(GrpcMethod::new("verisim.VeriSimOctad", "Create")); - self.inner.unary(req, path, codec).await - } - /// Get a octad by ID. - pub async fn get( - &mut self, - request: impl tonic::IntoRequest, - ) -> std::result::Result, tonic::Status> { - self.inner - .ready() - .await - .map_err(|e| { - tonic::Status::unknown( - format!("Service was not ready: {}", e.into()), - ) - })?; - let codec = tonic_prost::ProstCodec::default(); - let path = http::uri::PathAndQuery::from_static("/verisim.VeriSimOctad/Get"); - let mut req = request.into_request(); - req.extensions_mut().insert(GrpcMethod::new("verisim.VeriSimOctad", "Get")); - self.inner.unary(req, path, codec).await - } - /// Update an existing octad. - pub async fn update( - &mut self, - request: impl tonic::IntoRequest, - ) -> std::result::Result, tonic::Status> { - self.inner - .ready() - .await - .map_err(|e| { - tonic::Status::unknown( - format!("Service was not ready: {}", e.into()), - ) - })?; - let codec = tonic_prost::ProstCodec::default(); - let path = http::uri::PathAndQuery::from_static( - "/verisim.VeriSimOctad/Update", - ); - let mut req = request.into_request(); - req.extensions_mut() - .insert(GrpcMethod::new("verisim.VeriSimOctad", "Update")); - self.inner.unary(req, path, codec).await - } - /// Delete a octad. - pub async fn delete( - &mut self, - request: impl tonic::IntoRequest, - ) -> std::result::Result, tonic::Status> { - self.inner - .ready() - .await - .map_err(|e| { - tonic::Status::unknown( - format!("Service was not ready: {}", e.into()), - ) - })?; - let codec = tonic_prost::ProstCodec::default(); - let path = http::uri::PathAndQuery::from_static( - "/verisim.VeriSimOctad/Delete", - ); - let mut req = request.into_request(); - req.extensions_mut() - .insert(GrpcMethod::new("verisim.VeriSimOctad", "Delete")); - self.inner.unary(req, path, codec).await - } - /// Full-text search. - pub async fn search_text( - &mut self, - request: impl tonic::IntoRequest, - ) -> std::result::Result, tonic::Status> { - self.inner - .ready() - .await - .map_err(|e| { - tonic::Status::unknown( - format!("Service was not ready: {}", e.into()), - ) - })?; - let codec = tonic_prost::ProstCodec::default(); - let path = http::uri::PathAndQuery::from_static( - "/verisim.VeriSimOctad/SearchText", - ); - let mut req = request.into_request(); - req.extensions_mut() - .insert(GrpcMethod::new("verisim.VeriSimOctad", "SearchText")); - self.inner.unary(req, path, codec).await - } - /// Vector similarity search. - pub async fn search_vector( - &mut self, - request: impl tonic::IntoRequest, - ) -> std::result::Result, tonic::Status> { - self.inner - .ready() - .await - .map_err(|e| { - tonic::Status::unknown( - format!("Service was not ready: {}", e.into()), - ) - })?; - let codec = tonic_prost::ProstCodec::default(); - let path = http::uri::PathAndQuery::from_static( - "/verisim.VeriSimOctad/SearchVector", - ); - let mut req = request.into_request(); - req.extensions_mut() - .insert(GrpcMethod::new("verisim.VeriSimOctad", "SearchVector")); - self.inner.unary(req, path, codec).await - } - } -} -/// Generated server implementations. -pub mod veri_sim_octad_server { - #![allow( - unused_variables, - dead_code, - missing_docs, - clippy::wildcard_imports, - clippy::let_unit_value, - )] - use tonic::codegen::*; - /// Generated trait containing gRPC methods that should be implemented for use with VeriSimOctadServer. - #[async_trait] - pub trait VeriSimOctad: std::marker::Send + std::marker::Sync + 'static { - /// Create a new octad. - async fn create( - &self, - request: tonic::Request, - ) -> std::result::Result, tonic::Status>; - /// Get a octad by ID. - async fn get( - &self, - request: tonic::Request, - ) -> std::result::Result, tonic::Status>; - /// Update an existing octad. - async fn update( - &self, - request: tonic::Request, - ) -> std::result::Result, tonic::Status>; - /// Delete a octad. - async fn delete( - &self, - request: tonic::Request, - ) -> std::result::Result, tonic::Status>; - /// Full-text search. - async fn search_text( - &self, - request: tonic::Request, - ) -> std::result::Result, tonic::Status>; - /// Vector similarity search. - async fn search_vector( - &self, - request: tonic::Request, - ) -> std::result::Result, tonic::Status>; - } - #[derive(Debug)] - pub struct VeriSimOctadServer { - inner: Arc, - accept_compression_encodings: EnabledCompressionEncodings, - send_compression_encodings: EnabledCompressionEncodings, - max_decoding_message_size: Option, - max_encoding_message_size: Option, - } - impl VeriSimOctadServer { - pub fn new(inner: T) -> Self { - Self::from_arc(Arc::new(inner)) - } - pub fn from_arc(inner: Arc) -> Self { - Self { - inner, - accept_compression_encodings: Default::default(), - send_compression_encodings: Default::default(), - max_decoding_message_size: None, - max_encoding_message_size: None, - } - } - pub fn with_interceptor( - inner: T, - interceptor: F, - ) -> InterceptedService - where - F: tonic::service::Interceptor, - { - InterceptedService::new(Self::new(inner), interceptor) - } - /// Enable decompressing requests with the given encoding. - #[must_use] - pub fn accept_compressed(mut self, encoding: CompressionEncoding) -> Self { - self.accept_compression_encodings.enable(encoding); - self - } - /// Compress responses with the given encoding, if the client supports it. - #[must_use] - pub fn send_compressed(mut self, encoding: CompressionEncoding) -> Self { - self.send_compression_encodings.enable(encoding); - self - } - /// Limits the maximum size of a decoded message. - /// - /// Default: `4MB` - #[must_use] - pub fn max_decoding_message_size(mut self, limit: usize) -> Self { - self.max_decoding_message_size = Some(limit); - self - } - /// Limits the maximum size of an encoded message. - /// - /// Default: `usize::MAX` - #[must_use] - pub fn max_encoding_message_size(mut self, limit: usize) -> Self { - self.max_encoding_message_size = Some(limit); - self - } - } - impl tonic::codegen::Service> for VeriSimOctadServer - where - T: VeriSimOctad, - B: Body + std::marker::Send + 'static, - B::Error: Into + std::marker::Send + 'static, - { - type Response = http::Response; - type Error = std::convert::Infallible; - type Future = BoxFuture; - fn poll_ready( - &mut self, - _cx: &mut Context<'_>, - ) -> Poll> { - Poll::Ready(Ok(())) - } - fn call(&mut self, req: http::Request) -> Self::Future { - match req.uri().path() { - "/verisim.VeriSimOctad/Create" => { - #[allow(non_camel_case_types)] - struct CreateSvc(pub Arc); - impl< - T: VeriSimOctad, - > tonic::server::UnaryService - for CreateSvc { - type Response = super::OctadResponse; - type Future = BoxFuture< - tonic::Response, - tonic::Status, - >; - fn call( - &mut self, - request: tonic::Request, - ) -> Self::Future { - let inner = Arc::clone(&self.0); - let fut = async move { - ::create(&inner, request).await - }; - Box::pin(fut) - } - } - let accept_compression_encodings = self.accept_compression_encodings; - let send_compression_encodings = self.send_compression_encodings; - let max_decoding_message_size = self.max_decoding_message_size; - let max_encoding_message_size = self.max_encoding_message_size; - let inner = self.inner.clone(); - let fut = async move { - let method = CreateSvc(inner); - let codec = tonic_prost::ProstCodec::default(); - let mut grpc = tonic::server::Grpc::new(codec) - .apply_compression_config( - accept_compression_encodings, - send_compression_encodings, - ) - .apply_max_message_size_config( - max_decoding_message_size, - max_encoding_message_size, - ); - let res = grpc.unary(method, req).await; - Ok(res) - }; - Box::pin(fut) - } - "/verisim.VeriSimOctad/Get" => { - #[allow(non_camel_case_types)] - struct GetSvc(pub Arc); - impl< - T: VeriSimOctad, - > tonic::server::UnaryService for GetSvc { - type Response = super::OctadResponse; - type Future = BoxFuture< - tonic::Response, - tonic::Status, - >; - fn call( - &mut self, - request: tonic::Request, - ) -> Self::Future { - let inner = Arc::clone(&self.0); - let fut = async move { - ::get(&inner, request).await - }; - Box::pin(fut) - } - } - let accept_compression_encodings = self.accept_compression_encodings; - let send_compression_encodings = self.send_compression_encodings; - let max_decoding_message_size = self.max_decoding_message_size; - let max_encoding_message_size = self.max_encoding_message_size; - let inner = self.inner.clone(); - let fut = async move { - let method = GetSvc(inner); - let codec = tonic_prost::ProstCodec::default(); - let mut grpc = tonic::server::Grpc::new(codec) - .apply_compression_config( - accept_compression_encodings, - send_compression_encodings, - ) - .apply_max_message_size_config( - max_decoding_message_size, - max_encoding_message_size, - ); - let res = grpc.unary(method, req).await; - Ok(res) - }; - Box::pin(fut) - } - "/verisim.VeriSimOctad/Update" => { - #[allow(non_camel_case_types)] - struct UpdateSvc(pub Arc); - impl< - T: VeriSimOctad, - > tonic::server::UnaryService - for UpdateSvc { - type Response = super::OctadResponse; - type Future = BoxFuture< - tonic::Response, - tonic::Status, - >; - fn call( - &mut self, - request: tonic::Request, - ) -> Self::Future { - let inner = Arc::clone(&self.0); - let fut = async move { - ::update(&inner, request).await - }; - Box::pin(fut) - } - } - let accept_compression_encodings = self.accept_compression_encodings; - let send_compression_encodings = self.send_compression_encodings; - let max_decoding_message_size = self.max_decoding_message_size; - let max_encoding_message_size = self.max_encoding_message_size; - let inner = self.inner.clone(); - let fut = async move { - let method = UpdateSvc(inner); - let codec = tonic_prost::ProstCodec::default(); - let mut grpc = tonic::server::Grpc::new(codec) - .apply_compression_config( - accept_compression_encodings, - send_compression_encodings, - ) - .apply_max_message_size_config( - max_decoding_message_size, - max_encoding_message_size, - ); - let res = grpc.unary(method, req).await; - Ok(res) - }; - Box::pin(fut) - } - "/verisim.VeriSimOctad/Delete" => { - #[allow(non_camel_case_types)] - struct DeleteSvc(pub Arc); - impl< - T: VeriSimOctad, - > tonic::server::UnaryService - for DeleteSvc { - type Response = super::Empty; - type Future = BoxFuture< - tonic::Response, - tonic::Status, - >; - fn call( - &mut self, - request: tonic::Request, - ) -> Self::Future { - let inner = Arc::clone(&self.0); - let fut = async move { - ::delete(&inner, request).await - }; - Box::pin(fut) - } - } - let accept_compression_encodings = self.accept_compression_encodings; - let send_compression_encodings = self.send_compression_encodings; - let max_decoding_message_size = self.max_decoding_message_size; - let max_encoding_message_size = self.max_encoding_message_size; - let inner = self.inner.clone(); - let fut = async move { - let method = DeleteSvc(inner); - let codec = tonic_prost::ProstCodec::default(); - let mut grpc = tonic::server::Grpc::new(codec) - .apply_compression_config( - accept_compression_encodings, - send_compression_encodings, - ) - .apply_max_message_size_config( - max_decoding_message_size, - max_encoding_message_size, - ); - let res = grpc.unary(method, req).await; - Ok(res) - }; - Box::pin(fut) - } - "/verisim.VeriSimOctad/SearchText" => { - #[allow(non_camel_case_types)] - struct SearchTextSvc(pub Arc); - impl< - T: VeriSimOctad, - > tonic::server::UnaryService - for SearchTextSvc { - type Response = super::SearchResponse; - type Future = BoxFuture< - tonic::Response, - tonic::Status, - >; - fn call( - &mut self, - request: tonic::Request, - ) -> Self::Future { - let inner = Arc::clone(&self.0); - let fut = async move { - ::search_text(&inner, request).await - }; - Box::pin(fut) - } - } - let accept_compression_encodings = self.accept_compression_encodings; - let send_compression_encodings = self.send_compression_encodings; - let max_decoding_message_size = self.max_decoding_message_size; - let max_encoding_message_size = self.max_encoding_message_size; - let inner = self.inner.clone(); - let fut = async move { - let method = SearchTextSvc(inner); - let codec = tonic_prost::ProstCodec::default(); - let mut grpc = tonic::server::Grpc::new(codec) - .apply_compression_config( - accept_compression_encodings, - send_compression_encodings, - ) - .apply_max_message_size_config( - max_decoding_message_size, - max_encoding_message_size, - ); - let res = grpc.unary(method, req).await; - Ok(res) - }; - Box::pin(fut) - } - "/verisim.VeriSimOctad/SearchVector" => { - #[allow(non_camel_case_types)] - struct SearchVectorSvc(pub Arc); - impl< - T: VeriSimOctad, - > tonic::server::UnaryService - for SearchVectorSvc { - type Response = super::SearchResponse; - type Future = BoxFuture< - tonic::Response, - tonic::Status, - >; - fn call( - &mut self, - request: tonic::Request, - ) -> Self::Future { - let inner = Arc::clone(&self.0); - let fut = async move { - ::search_vector(&inner, request).await - }; - Box::pin(fut) - } - } - let accept_compression_encodings = self.accept_compression_encodings; - let send_compression_encodings = self.send_compression_encodings; - let max_decoding_message_size = self.max_decoding_message_size; - let max_encoding_message_size = self.max_encoding_message_size; - let inner = self.inner.clone(); - let fut = async move { - let method = SearchVectorSvc(inner); - let codec = tonic_prost::ProstCodec::default(); - let mut grpc = tonic::server::Grpc::new(codec) - .apply_compression_config( - accept_compression_encodings, - send_compression_encodings, - ) - .apply_max_message_size_config( - max_decoding_message_size, - max_encoding_message_size, - ); - let res = grpc.unary(method, req).await; - Ok(res) - }; - Box::pin(fut) - } - _ => { - Box::pin(async move { - let mut response = http::Response::new( - tonic::body::Body::default(), - ); - let headers = response.headers_mut(); - headers - .insert( - tonic::Status::GRPC_STATUS, - (tonic::Code::Unimplemented as i32).into(), - ); - headers - .insert( - http::header::CONTENT_TYPE, - tonic::metadata::GRPC_CONTENT_TYPE, - ); - Ok(response) - }) - } - } - } - } - impl Clone for VeriSimOctadServer { - fn clone(&self) -> Self { - let inner = self.inner.clone(); - Self { - inner, - accept_compression_encodings: self.accept_compression_encodings, - send_compression_encodings: self.send_compression_encodings, - max_decoding_message_size: self.max_decoding_message_size, - max_encoding_message_size: self.max_encoding_message_size, - } - } - } - /// Generated gRPC service name - pub const SERVICE_NAME: &str = "verisim.VeriSimOctad"; - impl tonic::server::NamedService for VeriSimOctadServer { - const NAME: &'static str = SERVICE_NAME; - } -} diff --git a/verisimdb/rust-core/verisim-api/src/rbac.rs b/verisimdb/rust-core/verisim-api/src/rbac.rs deleted file mode 100644 index 49a9d4ed..00000000 --- a/verisimdb/rust-core/verisim-api/src/rbac.rs +++ /dev/null @@ -1,1084 +0,0 @@ -// SPDX-License-Identifier: MPL-2.0 -// Copyright (c) 2026 Jonathan D.A. Jewell (hyperpolymath) -// -//! Role-Based Access Control (RBAC) for VeriSimDB API. -//! -//! Provides fine-grained permission checking beyond the basic [`ClientRole`] -//! authentication layer. Supports: -//! -//! - **Global permissions**: Apply to all resources (e.g., Admin can do everything). -//! - **Per-modality permissions**: Grant read/write/execute per VeriSimDB modality -//! (graph, vector, tensor, semantic, document, temporal). -//! - **Entity-level ACLs**: Override permissions for specific octad entities. -//! - **Audit logging**: Every access decision is recorded with timestamps. -//! -//! # Integration with auth middleware -//! -//! After the [`auth_middleware`](crate::auth::auth_middleware) extracts a -//! [`ClientIdentity`](crate::auth::ClientIdentity), call -//! [`check_authorization`] to verify that the client's role permits the -//! requested operation on the target resource. - -use crate::auth::{ClientIdentity, ClientRole}; -use axum::http::{Method, StatusCode}; -use axum::response::IntoResponse; -use axum::Json; -use serde::{Deserialize, Serialize}; -use std::collections::HashMap; -use std::sync::{Arc, Mutex}; -use std::time::SystemTime; -use tracing::{info, warn}; - -// --------------------------------------------------------------------------- -// Permission types -// --------------------------------------------------------------------------- - -/// Individual permission that can be granted to a role or entity ACL. -#[derive(Debug, Clone, Copy, PartialEq, Eq, Hash, Serialize, Deserialize)] -pub enum Permission { - /// Read data (GET requests, search, list). - Read, - /// Write data (POST, PUT, DELETE on octads and related resources). - Write, - /// Administrative operations (config changes, normalizer triggers, key management). - Admin, - /// Execute queries (VCL execution, query planning, EXPLAIN). - Execute, -} - -impl std::fmt::Display for Permission { - fn fmt(&self, f: &mut std::fmt::Formatter<'_>) -> std::fmt::Result { - match self { - Permission::Read => write!(f, "read"), - Permission::Write => write!(f, "write"), - Permission::Admin => write!(f, "admin"), - Permission::Execute => write!(f, "execute"), - } - } -} - -/// Permissions scoped to a specific VeriSimDB modality. -/// -/// For example, a role might have `Read` + `Execute` on the `vector` modality -/// but only `Read` on `graph`. -#[derive(Debug, Clone, Serialize, Deserialize)] -pub struct ModalityPermission { - /// The modality name (e.g., "graph", "vector", "tensor", "semantic", - /// "document", "temporal"). - pub modality: String, - /// Permissions granted for this modality. - pub permissions: Vec, -} - -// --------------------------------------------------------------------------- -// Role definitions -// --------------------------------------------------------------------------- - -/// A named role definition with global and per-modality permissions. -/// -/// Roles are composable: a client's effective permissions are the union of -/// global permissions and any modality-specific grants. -#[derive(Debug, Clone, Serialize, Deserialize)] -pub struct RoleDefinition { - /// Human-readable role name (e.g., "reader", "writer", "admin", - /// "vector-analyst"). - pub name: String, - /// Permissions that apply to every modality and every resource. - pub global_permissions: Vec, - /// Per-modality permission overrides. Key is the modality name. - pub modality_permissions: HashMap>, -} - -impl RoleDefinition { - /// Check whether this role grants a specific permission globally. - pub fn has_global_permission(&self, permission: Permission) -> bool { - self.global_permissions.contains(&permission) - } - - /// Check whether this role grants a specific permission for a modality. - /// - /// Returns `true` if the permission is granted either globally or - /// specifically for the named modality. - pub fn has_modality_permission(&self, modality: &str, permission: Permission) -> bool { - if self.has_global_permission(permission) { - return true; - } - self.modality_permissions - .get(modality) - .map(|perms| perms.contains(&permission)) - .unwrap_or(false) - } -} - -// --------------------------------------------------------------------------- -// RBAC policy -// --------------------------------------------------------------------------- - -/// The complete RBAC policy for the VeriSimDB instance. -/// -/// Contains all role definitions and entity-level ACL overrides. -#[derive(Debug, Clone, Serialize, Deserialize)] -pub struct RbacPolicy { - /// Named role definitions. Key is the role name (lowercase). - pub roles: HashMap, - /// Entity-level ACL overrides. Outer key is the entity (octad) ID, - /// inner tuples are `(client_id, Vec)` pairs. - pub entity_acls: HashMap)>>, -} - -impl Default for RbacPolicy { - /// Construct the default policy with built-in reader/writer/admin roles. - fn default() -> Self { - let mut roles = HashMap::new(); - - // Reader: read + execute globally. - roles.insert( - "reader".to_string(), - RoleDefinition { - name: "reader".to_string(), - global_permissions: vec![Permission::Read, Permission::Execute], - modality_permissions: HashMap::new(), - }, - ); - - // Writer: read + write + execute globally. - roles.insert( - "writer".to_string(), - RoleDefinition { - name: "writer".to_string(), - global_permissions: vec![Permission::Read, Permission::Write, Permission::Execute], - modality_permissions: HashMap::new(), - }, - ); - - // Admin: every permission globally. - roles.insert( - "admin".to_string(), - RoleDefinition { - name: "admin".to_string(), - global_permissions: vec![ - Permission::Read, - Permission::Write, - Permission::Admin, - Permission::Execute, - ], - modality_permissions: HashMap::new(), - }, - ); - - Self { - roles, - entity_acls: HashMap::new(), - } - } -} - -impl RbacPolicy { - /// Create a new empty policy (no roles, no ACLs). - pub fn new() -> Self { - Self { - roles: HashMap::new(), - entity_acls: HashMap::new(), - } - } - - /// Look up the [`RoleDefinition`] for a [`ClientRole`]. - pub fn role_for(&self, client_role: ClientRole) -> Option<&RoleDefinition> { - let key = match client_role { - ClientRole::Reader => "reader", - ClientRole::Writer => "writer", - ClientRole::Admin => "admin", - }; - self.roles.get(key) - } - - /// Register or replace a custom role definition. - pub fn set_role(&mut self, role: RoleDefinition) { - self.roles.insert(role.name.clone(), role); - } - - /// Add an entity-level ACL entry granting `permissions` to `client_id` - /// on `entity_id`. - pub fn add_entity_acl( - &mut self, - entity_id: &str, - client_id: &str, - permissions: Vec, - ) { - self.entity_acls - .entry(entity_id.to_string()) - .or_default() - .push((client_id.to_string(), permissions)); - } - - /// Check whether `client_id` has `permission` on `entity_id` via an - /// entity-level ACL entry. - pub fn check_entity_acl( - &self, - entity_id: &str, - client_id: &str, - permission: Permission, - ) -> bool { - self.entity_acls - .get(entity_id) - .map(|acls| { - acls.iter().any(|(cid, perms)| { - cid == client_id && perms.contains(&permission) - }) - }) - .unwrap_or(false) - } -} - -// --------------------------------------------------------------------------- -// Audit log -// --------------------------------------------------------------------------- - -/// Outcome of an authorization decision. -#[derive(Debug, Clone, Copy, PartialEq, Eq, Serialize, Deserialize)] -pub enum AccessDecision { - /// Access was permitted. - Allowed, - /// Access was denied. - Denied, -} - -impl std::fmt::Display for AccessDecision { - fn fmt(&self, f: &mut std::fmt::Formatter<'_>) -> std::fmt::Result { - match self { - AccessDecision::Allowed => write!(f, "ALLOWED"), - AccessDecision::Denied => write!(f, "DENIED"), - } - } -} - -/// A single entry in the authorization audit log. -#[derive(Debug, Clone, Serialize, Deserialize)] -pub struct AuditEntry { - /// Timestamp of the decision (seconds since UNIX epoch). - pub timestamp: u64, - /// Client identifier (API key hash or JWT subject). - pub client_id: String, - /// Role of the client at the time of the decision. - pub client_role: String, - /// The resource path that was accessed. - pub resource_path: String, - /// The HTTP method used. - pub method: String, - /// The permission that was required. - pub required_permission: Permission, - /// The decision outcome. - pub decision: AccessDecision, - /// Optional reason string for denials. - pub reason: Option, -} - -/// Thread-safe audit log that records all authorization decisions. -#[derive(Debug, Clone)] -pub struct AuditLog { - entries: Arc>>, - /// Maximum number of entries retained (ring buffer behaviour). - max_entries: usize, -} - -impl AuditLog { - /// Create a new audit log with the given capacity. - pub fn new(max_entries: usize) -> Self { - Self { - entries: Arc::new(Mutex::new(Vec::with_capacity(max_entries.min(4096)))), - max_entries, - } - } - - /// Record an authorization decision. - pub fn record(&self, entry: AuditEntry) { - let mut entries = self.entries.lock().expect("audit log lock"); - if entries.len() >= self.max_entries { - // Drop the oldest entry to stay within capacity. - entries.remove(0); - } - entries.push(entry); - } - - /// Retrieve a snapshot of all recorded entries. - pub fn entries(&self) -> Vec { - let entries = self.entries.lock().expect("audit log lock"); - entries.clone() - } - - /// Return the number of recorded entries. - pub fn len(&self) -> usize { - let entries = self.entries.lock().expect("audit log lock"); - entries.len() - } - - /// Check if the audit log is empty. - pub fn is_empty(&self) -> bool { - self.len() == 0 - } - - /// Clear all entries. - pub fn clear(&self) { - let mut entries = self.entries.lock().expect("audit log lock"); - entries.clear(); - } -} - -impl Default for AuditLog { - fn default() -> Self { - Self::new(10_000) - } -} - -// --------------------------------------------------------------------------- -// RBAC state (shared across the application) -// --------------------------------------------------------------------------- - -/// Shared RBAC state that is attached to the Axum application state. -#[derive(Debug, Clone)] -pub struct RbacState { - /// The active RBAC policy. - pub policy: Arc>, - /// Authorization audit log. - pub audit_log: AuditLog, -} - -impl RbacState { - /// Create RBAC state from a policy. - pub fn new(policy: RbacPolicy) -> Self { - Self { - policy: Arc::new(Mutex::new(policy)), - audit_log: AuditLog::default(), - } - } -} - -impl Default for RbacState { - fn default() -> Self { - Self::new(RbacPolicy::default()) - } -} - -// --------------------------------------------------------------------------- -// Authorization error -// --------------------------------------------------------------------------- - -/// Error returned when an authorization check fails. -#[derive(Debug, Clone, Serialize, Deserialize)] -pub struct AuthzError { - /// Human-readable error message. - pub error: String, - /// HTTP status code (always 403 for authorization failures). - pub code: u16, - /// The permission that was required but not granted. - pub required_permission: String, -} - -impl IntoResponse for AuthzError { - fn into_response(self) -> axum::response::Response { - (StatusCode::FORBIDDEN, Json(self)).into_response() - } -} - -// --------------------------------------------------------------------------- -// Permission derivation from HTTP method + path -// --------------------------------------------------------------------------- - -/// Derive the required [`Permission`] from an HTTP method and resource path. -/// -/// The mapping follows REST conventions: -/// - `GET` / `HEAD` / `OPTIONS` -> [`Permission::Read`] -/// - `POST` to query/plan/explain endpoints -> [`Permission::Execute`] -/// - `POST` / `PUT` / `PATCH` / `DELETE` -> [`Permission::Write`] -/// - Admin endpoints (`/normalizer/trigger`, `/planner/config` PUT) -> -/// [`Permission::Admin`] -pub fn required_permission(method: &Method, path: &str) -> Permission { - // Admin endpoints (explicitly listed). - if is_admin_path(method, path) { - return Permission::Admin; - } - - // Query execution endpoints. - if is_execute_path(method, path) { - return Permission::Execute; - } - - // Read-only methods. - if matches!(*method, Method::GET | Method::HEAD | Method::OPTIONS) { - return Permission::Read; - } - - // All other mutating methods. - Permission::Write -} - -/// Check if a (method, path) combination targets an admin-only resource. -fn is_admin_path(method: &Method, path: &str) -> bool { - // Normalizer trigger is admin-only. - if path.starts_with("/normalizer/trigger") && *method == Method::POST { - return true; - } - // Planner config mutation is admin-only. - if path.starts_with("/planner/config") && *method == Method::PUT { - return true; - } - false -} - -/// Check if a (method, path) combination targets a query execution resource. -fn is_execute_path(method: &Method, path: &str) -> bool { - if *method != Method::POST { - return false; - } - path.starts_with("/query/plan") - || path.starts_with("/query/explain") - || path.starts_with("/queries/similar") - || path.starts_with("/search/") -} - -/// Extract the modality name from a resource path, if applicable. -/// -/// Returns `None` when the path does not target a specific modality. -/// For octad-level paths we return `None` because octads span all modalities. -pub fn modality_from_path(path: &str) -> Option<&str> { - if path.starts_with("/search/vector") { - return Some("vector"); - } - if path.starts_with("/search/text") { - return Some("document"); - } - if path.starts_with("/search/related") { - return Some("graph"); - } - if path.starts_with("/drift") { - return Some("temporal"); - } - None -} - -/// Extract the entity (octad) ID from a resource path, if applicable. -pub fn entity_from_path(path: &str) -> Option<&str> { - // Matches /octads/{id} and sub-paths. - if let Some(rest) = path.strip_prefix("/octads/") { - let id = rest.split('/').next().unwrap_or(rest); - if !id.is_empty() { - return Some(id); - } - } - // Matches /drift/entity/{id}. - if let Some(rest) = path.strip_prefix("/drift/entity/") { - let id = rest.split('/').next().unwrap_or(rest); - if !id.is_empty() { - return Some(id); - } - } - // Matches /normalizer/trigger/{id}. - if let Some(rest) = path.strip_prefix("/normalizer/trigger/") { - let id = rest.split('/').next().unwrap_or(rest); - if !id.is_empty() { - return Some(id); - } - } - None -} - -// --------------------------------------------------------------------------- -// Core authorization check -// --------------------------------------------------------------------------- - -/// Check whether a client is authorized to access a resource. -/// -/// This is the primary entry point for RBAC enforcement. It: -/// -/// 1. Determines the required [`Permission`] from the HTTP method and path. -/// 2. Looks up the client's [`RoleDefinition`] in the active policy. -/// 3. Checks entity-level ACLs (which can grant access that the role alone -/// would not). -/// 4. Checks modality-specific permissions when the path targets a specific -/// modality. -/// 5. Falls back to global role permissions. -/// 6. Records the decision in the audit log. -/// -/// Returns `Ok(())` when access is granted, or `Err(AuthzError)` when denied. -pub fn check_access( - identity: &ClientIdentity, - resource_path: &str, - method: &Method, - rbac: &RbacState, -) -> Result<(), AuthzError> { - let permission = required_permission(method, resource_path); - let policy = rbac.policy.lock().expect("rbac policy lock"); - - let now_secs = SystemTime::now() - .duration_since(SystemTime::UNIX_EPOCH) - .unwrap_or_default() - .as_secs(); - - let role_name = match identity.role { - ClientRole::Reader => "reader", - ClientRole::Writer => "writer", - ClientRole::Admin => "admin", - }; - - // Helper to record + return a decision. - let record = |decision: AccessDecision, reason: Option| { - let entry = AuditEntry { - timestamp: now_secs, - client_id: identity.id.clone(), - client_role: role_name.to_string(), - resource_path: resource_path.to_string(), - method: method.to_string(), - required_permission: permission, - decision, - reason: reason.clone(), - }; - rbac.audit_log.record(entry); - }; - - // --- 1. Entity-level ACL check (overrides role) --- - if let Some(entity_id) = entity_from_path(resource_path) { - if policy.check_entity_acl(entity_id, &identity.id, permission) { - info!( - client = %identity.id, - role = %role_name, - path = %resource_path, - permission = %permission, - "Access ALLOWED via entity ACL" - ); - record(AccessDecision::Allowed, Some("entity ACL grant".to_string())); - return Ok(()); - } - } - - // --- 2. Look up role definition --- - let role_def = match policy.role_for(identity.role) { - Some(rd) => rd, - None => { - let reason = format!("No role definition found for '{}'", role_name); - warn!( - client = %identity.id, - role = %role_name, - path = %resource_path, - "Access DENIED: {}", reason - ); - record(AccessDecision::Denied, Some(reason.clone())); - return Err(AuthzError { - error: reason, - code: 403, - required_permission: permission.to_string(), - }); - } - }; - - // --- 3. Modality-specific check --- - if let Some(modality) = modality_from_path(resource_path) { - if role_def.has_modality_permission(modality, permission) { - info!( - client = %identity.id, - role = %role_name, - path = %resource_path, - modality = %modality, - permission = %permission, - "Access ALLOWED via modality permission" - ); - record(AccessDecision::Allowed, Some(format!("modality '{}' grant", modality))); - return Ok(()); - } - // If the role has a modality_permissions entry for this modality - // but the specific permission is missing, deny rather than falling - // through to global (explicit modality config is restrictive). - if role_def.modality_permissions.contains_key(modality) { - let reason = format!( - "Role '{}' lacks '{}' permission on modality '{}'", - role_name, permission, modality - ); - warn!( - client = %identity.id, - role = %role_name, - path = %resource_path, - "Access DENIED: {}", reason - ); - record(AccessDecision::Denied, Some(reason.clone())); - return Err(AuthzError { - error: reason, - code: 403, - required_permission: permission.to_string(), - }); - } - } - - // --- 4. Global permission check --- - if role_def.has_global_permission(permission) { - info!( - client = %identity.id, - role = %role_name, - path = %resource_path, - permission = %permission, - "Access ALLOWED via global permission" - ); - record(AccessDecision::Allowed, Some("global role grant".to_string())); - return Ok(()); - } - - // --- 5. Denied --- - let reason = format!( - "Role '{}' does not have '{}' permission", - role_name, permission - ); - warn!( - client = %identity.id, - role = %role_name, - path = %resource_path, - "Access DENIED: {}", reason - ); - record(AccessDecision::Denied, Some(reason.clone())); - Err(AuthzError { - error: reason, - code: 403, - required_permission: permission.to_string(), - }) -} - -/// Convenience wrapper intended for use by the authentication middleware. -/// -/// After [`extract_identity`](crate::auth) succeeds, the middleware calls -/// this function to perform authorization. If the check fails, the returned -/// error is directly convertible to an Axum response via [`IntoResponse`]. -pub fn check_authorization( - identity: &ClientIdentity, - resource_path: &str, - method: &Method, - rbac: &RbacState, -) -> Result<(), AuthzError> { - check_access(identity, resource_path, method, rbac) -} - -// --------------------------------------------------------------------------- -// Tests -// --------------------------------------------------------------------------- - -#[cfg(test)] -mod tests { - use super::*; - use crate::auth::{ClientIdentity, ClientRole}; - - /// Helper: build an [`RbacState`] with the default policy. - fn default_rbac() -> RbacState { - RbacState::default() - } - - /// Helper: build a [`ClientIdentity`] with the given role. - fn identity(id: &str, role: ClientRole) -> ClientIdentity { - ClientIdentity { - id: id.to_string(), - role, - } - } - - // ------------------------------------------------------------------ - // Test 1: Admin has all permissions - // ------------------------------------------------------------------ - #[test] - fn test_admin_has_all_permissions() { - let rbac = default_rbac(); - let admin = identity("admin-user", ClientRole::Admin); - - // Read - assert!(check_access(&admin, "/octads", &Method::GET, &rbac).is_ok()); - // Write - assert!(check_access(&admin, "/octads", &Method::POST, &rbac).is_ok()); - // Execute - assert!(check_access(&admin, "/query/plan", &Method::POST, &rbac).is_ok()); - // Admin - assert!(check_access( - &admin, - "/normalizer/trigger/entity-1", - &Method::POST, - &rbac - ).is_ok()); - assert!(check_access( - &admin, - "/planner/config", - &Method::PUT, - &rbac - ).is_ok()); - } - - // ------------------------------------------------------------------ - // Test 2: Reader denied write access - // ------------------------------------------------------------------ - #[test] - fn test_reader_denied_write_access() { - let rbac = default_rbac(); - let reader = identity("reader-user", ClientRole::Reader); - - // Write should be denied. - let result = check_access(&reader, "/octads", &Method::POST, &rbac); - assert!(result.is_err()); - let err = result.unwrap_err(); - assert_eq!(err.code, 403); - assert_eq!(err.required_permission, "write"); - - // PUT should be denied. - let result = check_access(&reader, "/octads/some-id", &Method::PUT, &rbac); - assert!(result.is_err()); - - // DELETE should be denied. - let result = check_access(&reader, "/octads/some-id", &Method::DELETE, &rbac); - assert!(result.is_err()); - } - - // ------------------------------------------------------------------ - // Test 3: Reader allowed read and execute - // ------------------------------------------------------------------ - #[test] - fn test_reader_allowed_read_and_execute() { - let rbac = default_rbac(); - let reader = identity("reader-user", ClientRole::Reader); - - // Read is allowed. - assert!(check_access(&reader, "/octads", &Method::GET, &rbac).is_ok()); - assert!(check_access(&reader, "/octads/abc", &Method::GET, &rbac).is_ok()); - assert!(check_access(&reader, "/drift/status", &Method::GET, &rbac).is_ok()); - - // Execute is allowed (query endpoints). - assert!(check_access(&reader, "/query/plan", &Method::POST, &rbac).is_ok()); - assert!(check_access(&reader, "/query/explain", &Method::POST, &rbac).is_ok()); - assert!(check_access(&reader, "/search/vector", &Method::POST, &rbac).is_ok()); - } - - // ------------------------------------------------------------------ - // Test 4: Writer allowed write, denied admin - // ------------------------------------------------------------------ - #[test] - fn test_writer_allowed_write_denied_admin() { - let rbac = default_rbac(); - let writer = identity("writer-user", ClientRole::Writer); - - // Write is allowed. - assert!(check_access(&writer, "/octads", &Method::POST, &rbac).is_ok()); - assert!(check_access(&writer, "/octads/abc", &Method::PUT, &rbac).is_ok()); - assert!(check_access(&writer, "/octads/abc", &Method::DELETE, &rbac).is_ok()); - - // Read and execute also allowed. - assert!(check_access(&writer, "/octads", &Method::GET, &rbac).is_ok()); - assert!(check_access(&writer, "/query/plan", &Method::POST, &rbac).is_ok()); - - // Admin is denied. - let result = check_access( - &writer, - "/normalizer/trigger/entity-1", - &Method::POST, - &rbac, - ); - assert!(result.is_err()); - let err = result.unwrap_err(); - assert_eq!(err.required_permission, "admin"); - - let result = check_access(&writer, "/planner/config", &Method::PUT, &rbac); - assert!(result.is_err()); - } - - // ------------------------------------------------------------------ - // Test 5: Per-modality permissions work - // ------------------------------------------------------------------ - #[test] - fn test_per_modality_permissions() { - let mut policy = RbacPolicy::new(); - - // Create a custom role that can only read vectors and write documents. - let mut modality_perms = HashMap::new(); - modality_perms.insert( - "vector".to_string(), - vec![Permission::Read, Permission::Execute], - ); - modality_perms.insert( - "document".to_string(), - vec![Permission::Read, Permission::Write], - ); - // Explicitly empty graph permissions — should deny graph access. - modality_perms.insert("graph".to_string(), vec![Permission::Read]); - - policy.set_role(RoleDefinition { - name: "reader".to_string(), - global_permissions: vec![], // No global perms — all modality-specific. - modality_permissions: modality_perms, - }); - - let rbac = RbacState::new(policy); - let user = identity("modal-user", ClientRole::Reader); - - // Vector search (Execute on vector) → allowed. - assert!(check_access(&user, "/search/vector", &Method::POST, &rbac).is_ok()); - - // Text search (requires execute on document modality) → denied - // because document modality only has Read and Write. - let result = check_access(&user, "/search/text", &Method::POST, &rbac); - assert!(result.is_err()); - - // Graph related search (requires execute on graph modality) → denied - // because graph modality only has Read. - let result = check_access(&user, "/search/related/xyz", &Method::POST, &rbac); - assert!(result.is_err()); - } - - // ------------------------------------------------------------------ - // Test 6: Entity-level ACLs work - // ------------------------------------------------------------------ - #[test] - fn test_entity_level_acls() { - let mut policy = RbacPolicy::default(); - - // Grant the reader "writer-user" write access to a specific entity. - policy.add_entity_acl("secret-octad", "special-reader", vec![Permission::Write]); - - let rbac = RbacState::new(policy); - let reader = identity("special-reader", ClientRole::Reader); - - // Normally a reader cannot write. - let result = check_access( - &reader, - "/octads/other-octad", - &Method::PUT, - &rbac, - ); - assert!(result.is_err()); - - // But the entity ACL grants write on "secret-octad". - assert!(check_access( - &reader, - "/octads/secret-octad", - &Method::PUT, - &rbac, - ).is_ok()); - } - - // ------------------------------------------------------------------ - // Test 7: Audit log records decisions - // ------------------------------------------------------------------ - #[test] - fn test_audit_log_records_decisions() { - let rbac = default_rbac(); - let admin = identity("audit-admin", ClientRole::Admin); - let reader = identity("audit-reader", ClientRole::Reader); - - // Successful access. - let _ = check_access(&admin, "/octads", &Method::GET, &rbac); - // Denied access. - let _ = check_access(&reader, "/octads", &Method::POST, &rbac); - - let entries = rbac.audit_log.entries(); - assert!(entries.len() >= 2, "Expected at least 2 audit entries, got {}", entries.len()); - - // First entry: admin allowed read. - let allowed_entry = entries.iter().find(|e| e.decision == AccessDecision::Allowed); - assert!(allowed_entry.is_some(), "Expected an ALLOWED entry"); - let allowed = allowed_entry.expect("TODO: handle error"); - assert_eq!(allowed.client_id, "audit-admin"); - assert_eq!(allowed.client_role, "admin"); - assert_eq!(allowed.required_permission, Permission::Read); - - // Second entry: reader denied write. - let denied_entry = entries.iter().find(|e| e.decision == AccessDecision::Denied); - assert!(denied_entry.is_some(), "Expected a DENIED entry"); - let denied = denied_entry.expect("TODO: handle error"); - assert_eq!(denied.client_id, "audit-reader"); - assert_eq!(denied.client_role, "reader"); - assert_eq!(denied.required_permission, Permission::Write); - assert!(denied.reason.is_some()); - } - - // ------------------------------------------------------------------ - // Test 8: Audit log capacity limit - // ------------------------------------------------------------------ - #[test] - fn test_audit_log_capacity() { - let log = AuditLog::new(3); - assert!(log.is_empty()); - - for i in 0..5 { - log.record(AuditEntry { - timestamp: i as u64, - client_id: format!("client-{}", i), - client_role: "reader".to_string(), - resource_path: "/test".to_string(), - method: "GET".to_string(), - required_permission: Permission::Read, - decision: AccessDecision::Allowed, - reason: None, - }); - } - - // Should only retain the last 3 entries. - assert_eq!(log.len(), 3); - let entries = log.entries(); - assert_eq!(entries[0].client_id, "client-2"); - assert_eq!(entries[1].client_id, "client-3"); - assert_eq!(entries[2].client_id, "client-4"); - } - - // ------------------------------------------------------------------ - // Test 9: Required permission derivation - // ------------------------------------------------------------------ - #[test] - fn test_required_permission_derivation() { - // Read - assert_eq!(required_permission(&Method::GET, "/octads"), Permission::Read); - assert_eq!(required_permission(&Method::HEAD, "/octads"), Permission::Read); - assert_eq!(required_permission(&Method::OPTIONS, "/anything"), Permission::Read); - - // Write - assert_eq!(required_permission(&Method::POST, "/octads"), Permission::Write); - assert_eq!(required_permission(&Method::PUT, "/octads/abc"), Permission::Write); - assert_eq!(required_permission(&Method::DELETE, "/octads/abc"), Permission::Write); - - // Execute - assert_eq!(required_permission(&Method::POST, "/query/plan"), Permission::Execute); - assert_eq!(required_permission(&Method::POST, "/query/explain"), Permission::Execute); - assert_eq!(required_permission(&Method::POST, "/search/vector"), Permission::Execute); - assert_eq!(required_permission(&Method::POST, "/search/text?q=foo"), Permission::Execute); - assert_eq!(required_permission(&Method::POST, "/queries/similar"), Permission::Execute); - - // Admin - assert_eq!( - required_permission(&Method::POST, "/normalizer/trigger/abc"), - Permission::Admin - ); - assert_eq!( - required_permission(&Method::PUT, "/planner/config"), - Permission::Admin - ); - } - - // ------------------------------------------------------------------ - // Test 10: Entity and modality extraction from paths - // ------------------------------------------------------------------ - #[test] - fn test_entity_and_modality_extraction() { - // Entity extraction - assert_eq!(entity_from_path("/octads/my-entity"), Some("my-entity")); - assert_eq!(entity_from_path("/octads/abc/sub"), Some("abc")); - assert_eq!(entity_from_path("/drift/entity/drift-id"), Some("drift-id")); - assert_eq!(entity_from_path("/normalizer/trigger/norm-id"), Some("norm-id")); - assert_eq!(entity_from_path("/octads"), None); - assert_eq!(entity_from_path("/search/text"), None); - - // Modality extraction - assert_eq!(modality_from_path("/search/vector"), Some("vector")); - assert_eq!(modality_from_path("/search/text"), Some("document")); - assert_eq!(modality_from_path("/search/related/xyz"), Some("graph")); - assert_eq!(modality_from_path("/drift/status"), Some("temporal")); - assert_eq!(modality_from_path("/octads"), None); - assert_eq!(modality_from_path("/query/plan"), None); - } - - // ------------------------------------------------------------------ - // Test 11: Custom role definition - // ------------------------------------------------------------------ - #[test] - fn test_custom_role_definition() { - let mut policy = RbacPolicy::new(); - - let mut modality_perms = HashMap::new(); - modality_perms.insert( - "vector".to_string(), - vec![Permission::Read, Permission::Write, Permission::Execute], - ); - - policy.set_role(RoleDefinition { - name: "writer".to_string(), - global_permissions: vec![Permission::Read], - modality_permissions: modality_perms, - }); - - let rbac = RbacState::new(policy); - let writer = identity("custom-writer", ClientRole::Writer); - - // Global read → allowed. - assert!(check_access(&writer, "/octads", &Method::GET, &rbac).is_ok()); - - // Vector write → allowed (modality grant). - assert!(check_access(&writer, "/search/vector", &Method::POST, &rbac).is_ok()); - - // Octad write → denied (no global write, no modality for octad paths). - let result = check_access(&writer, "/octads", &Method::POST, &rbac); - assert!(result.is_err()); - } - - // ------------------------------------------------------------------ - // Test 12: Missing role definition handled gracefully - // ------------------------------------------------------------------ - #[test] - fn test_missing_role_definition() { - // Policy with no roles at all. - let policy = RbacPolicy::new(); - let rbac = RbacState::new(policy); - let reader = identity("orphan-user", ClientRole::Reader); - - let result = check_access(&reader, "/octads", &Method::GET, &rbac); - assert!(result.is_err()); - let err = result.unwrap_err(); - assert!(err.error.contains("No role definition found")); - } - - // ------------------------------------------------------------------ - // Test 13: check_authorization wrapper works identically - // ------------------------------------------------------------------ - #[test] - fn test_check_authorization_wrapper() { - let rbac = default_rbac(); - let admin = identity("wrapper-admin", ClientRole::Admin); - let reader = identity("wrapper-reader", ClientRole::Reader); - - // Admin should pass. - assert!(check_authorization(&admin, "/octads", &Method::POST, &rbac).is_ok()); - - // Reader write should fail. - assert!(check_authorization(&reader, "/octads", &Method::POST, &rbac).is_err()); - } - - // ------------------------------------------------------------------ - // Test 14: Entity ACL does not leak to other entities - // ------------------------------------------------------------------ - #[test] - fn test_entity_acl_isolation() { - let mut policy = RbacPolicy::default(); - policy.add_entity_acl("entity-alpha", "reader-x", vec![Permission::Write]); - - let rbac = RbacState::new(policy); - let reader = identity("reader-x", ClientRole::Reader); - - // Write to entity-alpha → allowed via ACL. - assert!(check_access(&reader, "/octads/entity-alpha", &Method::PUT, &rbac).is_ok()); - - // Write to entity-beta → denied (no ACL for this entity). - let result = check_access(&reader, "/octads/entity-beta", &Method::PUT, &rbac); - assert!(result.is_err()); - - // Write to entity-alpha by a DIFFERENT client → denied. - let other_reader = identity("reader-y", ClientRole::Reader); - let result = check_access( - &other_reader, - "/octads/entity-alpha", - &Method::PUT, - &rbac, - ); - assert!(result.is_err()); - } - - // ------------------------------------------------------------------ - // Test 15: Audit log clear - // ------------------------------------------------------------------ - #[test] - fn test_audit_log_clear() { - let rbac = default_rbac(); - let admin = identity("clear-admin", ClientRole::Admin); - - let _ = check_access(&admin, "/octads", &Method::GET, &rbac); - assert!(!rbac.audit_log.is_empty()); - - rbac.audit_log.clear(); - assert!(rbac.audit_log.is_empty()); - assert_eq!(rbac.audit_log.len(), 0); - } -} diff --git a/verisimdb/rust-core/verisim-api/src/transaction.rs b/verisimdb/rust-core/verisim-api/src/transaction.rs deleted file mode 100644 index 785e5974..00000000 --- a/verisimdb/rust-core/verisim-api/src/transaction.rs +++ /dev/null @@ -1,466 +0,0 @@ -// SPDX-License-Identifier: MPL-2.0 -// Copyright (c) 2026 Jonathan D.A. Jewell (hyperpolymath) -//! ACID Transaction Manager for VeriSimDB. -//! -//! Coordinates multi-modality transactions using a write-ahead log (WAL) -//! for durability and crash recovery. Transactions buffer operations and -//! apply them atomically on commit, or discard them on rollback. -//! -//! # Usage -//! -//! ```ignore -//! let txn_id = manager.begin().await; -//! manager.buffer_operation(&txn_id, op).await?; -//! manager.commit(&txn_id).await?; // or rollback(&txn_id) -//! ``` - -use std::collections::HashMap; -use std::sync::Arc; - -use chrono::Utc; -use serde::{Deserialize, Serialize}; -use sha2::{Digest, Sha256}; -use tokio::sync::RwLock; -use tracing::{info, warn}; - -/// Unique identifier for a transaction. -#[derive(Debug, Clone, Hash, Eq, PartialEq, Serialize, Deserialize)] -pub struct TransactionId(String); - -impl TransactionId { - /// Generate a new unique transaction ID. - pub fn new() -> Self { - let now = Utc::now(); - let mut hasher = Sha256::new(); - hasher.update(now.timestamp_nanos_opt().unwrap_or(0).to_le_bytes()); - hasher.update(std::process::id().to_le_bytes()); - let hash = hasher.finalize(); - Self(hex::encode(&hash[..16])) - } - - pub fn as_str(&self) -> &str { - &self.0 - } - - /// Create a TransactionId from an existing string (for API lookups). - pub fn from_str(s: &str) -> Self { - Self(s.to_string()) - } -} - -impl std::fmt::Display for TransactionId { - fn fmt(&self, f: &mut std::fmt::Formatter<'_>) -> std::fmt::Result { - write!(f, "txn_{}", &self.0[..12]) - } -} - -/// State of a transaction. -#[derive(Debug, Clone, Copy, PartialEq, Eq, Serialize, Deserialize)] -pub enum TransactionState { - /// Transaction is active and accepting operations. - Active, - /// Transaction has been committed. - Committed, - /// Transaction has been rolled back. - RolledBack, -} - -/// A buffered operation within a transaction. -#[derive(Debug, Clone, Serialize, Deserialize)] -pub struct BufferedOperation { - /// Target entity ID - pub entity_id: String, - /// Operation type - pub operation: OperationType, - /// Serialized payload (JSON) - pub payload: Vec, - /// Timestamp of the operation - pub timestamp: String, -} - -/// Types of operations that can be buffered in a transaction. -#[derive(Debug, Clone, Serialize, Deserialize)] -pub enum OperationType { - /// Create a new octad - Create, - /// Update an existing octad - Update, - /// Delete a octad - Delete, -} - -/// An active transaction with its buffered operations. -#[derive(Debug, Clone, Serialize, Deserialize)] -pub struct Transaction { - /// Transaction identifier - pub id: TransactionId, - /// Current state - pub state: TransactionState, - /// Buffered operations (applied on commit) - pub operations: Vec, - /// When the transaction was started - pub started_at: String, - /// When the transaction was completed (committed or rolled back) - pub completed_at: Option, -} - -/// Transaction status response for the API. -#[derive(Debug, Serialize, Deserialize)] -pub struct TransactionStatus { - pub id: String, - pub state: TransactionState, - pub operation_count: usize, - pub started_at: String, - pub completed_at: Option, -} - -impl From<&Transaction> for TransactionStatus { - fn from(txn: &Transaction) -> Self { - Self { - id: txn.id.0.clone(), - state: txn.state, - operation_count: txn.operations.len(), - started_at: txn.started_at.clone(), - completed_at: txn.completed_at.clone(), - } - } -} - -/// Errors from the transaction manager. -#[derive(Debug, Clone, Serialize, Deserialize)] -pub enum TransactionError { - /// Transaction not found. - NotFound(String), - /// Transaction is not in an active state. - NotActive(String), - /// Transaction has already been committed. - AlreadyCommitted(String), - /// Transaction has already been rolled back. - AlreadyRolledBack(String), - /// Maximum concurrent transactions exceeded. - TooManyTransactions, -} - -impl std::fmt::Display for TransactionError { - fn fmt(&self, f: &mut std::fmt::Formatter<'_>) -> std::fmt::Result { - match self { - Self::NotFound(id) => write!(f, "transaction not found: {}", id), - Self::NotActive(id) => write!(f, "transaction not active: {}", id), - Self::AlreadyCommitted(id) => write!(f, "transaction already committed: {}", id), - Self::AlreadyRolledBack(id) => write!(f, "transaction already rolled back: {}", id), - Self::TooManyTransactions => write!(f, "maximum concurrent transactions exceeded"), - } - } -} - -impl std::error::Error for TransactionError {} - -/// Configuration for the transaction manager. -#[derive(Debug, Clone, Serialize, Deserialize)] -pub struct TransactionConfig { - /// Maximum number of concurrent active transactions. - pub max_concurrent: usize, - /// Transaction timeout in seconds (auto-rollback after this). - pub timeout_seconds: u64, -} - -impl Default for TransactionConfig { - fn default() -> Self { - Self { - max_concurrent: 256, - timeout_seconds: 300, // 5 minutes - } - } -} - -/// Transaction manager — coordinates multi-modality ACID transactions. -/// -/// All operations within a transaction are buffered in memory. On commit, -/// the entire batch is written as a WAL commit record, then applied to -/// the stores. On rollback, the buffer is simply discarded. -pub struct TransactionManager { - config: TransactionConfig, - transactions: Arc>>, -} - -impl TransactionManager { - /// Create a new transaction manager. - pub fn new(config: TransactionConfig) -> Self { - Self { - config, - transactions: Arc::new(RwLock::new(HashMap::new())), - } - } - - /// Begin a new transaction. - pub async fn begin(&self) -> Result { - let mut txns = self.transactions.write().await; - - // Check concurrent transaction limit - let active_count = txns - .values() - .filter(|t| t.state == TransactionState::Active) - .count(); - if active_count >= self.config.max_concurrent { - return Err(TransactionError::TooManyTransactions); - } - - let id = TransactionId::new(); - let txn = Transaction { - id: id.clone(), - state: TransactionState::Active, - operations: Vec::new(), - started_at: Utc::now().to_rfc3339(), - completed_at: None, - }; - - info!(txn_id = %id, "Transaction started"); - txns.insert(id.clone(), txn); - Ok(id) - } - - /// Buffer an operation within a transaction. - pub async fn buffer_operation( - &self, - txn_id: &TransactionId, - operation: BufferedOperation, - ) -> Result<(), TransactionError> { - let mut txns = self.transactions.write().await; - let txn = txns - .get_mut(txn_id) - .ok_or_else(|| TransactionError::NotFound(txn_id.0.clone()))?; - - match txn.state { - TransactionState::Active => { - txn.operations.push(operation); - Ok(()) - } - TransactionState::Committed => { - Err(TransactionError::AlreadyCommitted(txn_id.0.clone())) - } - TransactionState::RolledBack => { - Err(TransactionError::AlreadyRolledBack(txn_id.0.clone())) - } - } - } - - /// Commit a transaction — marks all buffered operations as committed. - /// - /// Returns the list of buffered operations so the caller (API layer) - /// can apply them to the octad store. - pub async fn commit( - &self, - txn_id: &TransactionId, - ) -> Result, TransactionError> { - let mut txns = self.transactions.write().await; - let txn = txns - .get_mut(txn_id) - .ok_or_else(|| TransactionError::NotFound(txn_id.0.clone()))?; - - match txn.state { - TransactionState::Active => { - txn.state = TransactionState::Committed; - txn.completed_at = Some(Utc::now().to_rfc3339()); - let ops = txn.operations.clone(); - info!(txn_id = %txn_id, ops = ops.len(), "Transaction committed"); - Ok(ops) - } - TransactionState::Committed => { - Err(TransactionError::AlreadyCommitted(txn_id.0.clone())) - } - TransactionState::RolledBack => { - Err(TransactionError::AlreadyRolledBack(txn_id.0.clone())) - } - } - } - - /// Rollback a transaction — discard all buffered operations. - pub async fn rollback( - &self, - txn_id: &TransactionId, - ) -> Result { - let mut txns = self.transactions.write().await; - let txn = txns - .get_mut(txn_id) - .ok_or_else(|| TransactionError::NotFound(txn_id.0.clone()))?; - - match txn.state { - TransactionState::Active => { - let discarded = txn.operations.len(); - txn.operations.clear(); - txn.state = TransactionState::RolledBack; - txn.completed_at = Some(Utc::now().to_rfc3339()); - warn!(txn_id = %txn_id, discarded = discarded, "Transaction rolled back"); - Ok(discarded) - } - TransactionState::Committed => { - Err(TransactionError::AlreadyCommitted(txn_id.0.clone())) - } - TransactionState::RolledBack => { - Err(TransactionError::AlreadyRolledBack(txn_id.0.clone())) - } - } - } - - /// Get the status of a transaction. - pub async fn status( - &self, - txn_id: &TransactionId, - ) -> Result { - let txns = self.transactions.read().await; - let txn = txns - .get(txn_id) - .ok_or_else(|| TransactionError::NotFound(txn_id.0.clone()))?; - Ok(TransactionStatus::from(txn)) - } - - /// Clean up completed transactions older than the timeout. - pub async fn cleanup_expired(&self) -> usize { - let mut txns = self.transactions.write().await; - let now = Utc::now(); - let timeout = chrono::Duration::seconds(self.config.timeout_seconds as i64); - - let expired_ids: Vec = txns - .iter() - .filter(|(_, txn)| { - if txn.state != TransactionState::Active { - // Already completed — clean up after timeout - if let Some(ref completed) = txn.completed_at { - if let Ok(completed_at) = chrono::DateTime::parse_from_rfc3339(completed) { - return now.signed_duration_since(completed_at) > timeout; - } - } - false - } else { - // Active but expired — auto-rollback candidate - if let Ok(started) = chrono::DateTime::parse_from_rfc3339(&txn.started_at) { - now.signed_duration_since(started) > timeout - } else { - false - } - } - }) - .map(|(id, _)| id.clone()) - .collect(); - - let count = expired_ids.len(); - for id in expired_ids { - txns.remove(&id); - } - count - } -} - -#[cfg(test)] -mod tests { - use super::*; - - #[tokio::test] - async fn test_begin_commit() { - let mgr = TransactionManager::new(TransactionConfig::default()); - - let txn_id = mgr.begin().await.expect("TODO: handle error"); - let status = mgr.status(&txn_id).await.expect("TODO: handle error"); - assert_eq!(status.state, TransactionState::Active); - assert_eq!(status.operation_count, 0); - - // Buffer an operation - mgr.buffer_operation( - &txn_id, - BufferedOperation { - entity_id: "hex-001".to_string(), - operation: OperationType::Create, - payload: b"{}".to_vec(), - timestamp: Utc::now().to_rfc3339(), - }, - ) - .await - .expect("TODO: handle error"); - - let status = mgr.status(&txn_id).await.expect("TODO: handle error"); - assert_eq!(status.operation_count, 1); - - // Commit - let ops = mgr.commit(&txn_id).await.expect("TODO: handle error"); - assert_eq!(ops.len(), 1); - - let status = mgr.status(&txn_id).await.expect("TODO: handle error"); - assert_eq!(status.state, TransactionState::Committed); - } - - #[tokio::test] - async fn test_begin_rollback() { - let mgr = TransactionManager::new(TransactionConfig::default()); - - let txn_id = mgr.begin().await.expect("TODO: handle error"); - - mgr.buffer_operation( - &txn_id, - BufferedOperation { - entity_id: "hex-002".to_string(), - operation: OperationType::Update, - payload: b"{}".to_vec(), - timestamp: Utc::now().to_rfc3339(), - }, - ) - .await - .expect("TODO: handle error"); - - let discarded = mgr.rollback(&txn_id).await.expect("TODO: handle error"); - assert_eq!(discarded, 1); - - let status = mgr.status(&txn_id).await.expect("TODO: handle error"); - assert_eq!(status.state, TransactionState::RolledBack); - } - - #[tokio::test] - async fn test_double_commit_fails() { - let mgr = TransactionManager::new(TransactionConfig::default()); - let txn_id = mgr.begin().await.expect("TODO: handle error"); - mgr.commit(&txn_id).await.expect("TODO: handle error"); - let result = mgr.commit(&txn_id).await; - assert!(result.is_err()); - } - - #[tokio::test] - async fn test_operation_on_committed_fails() { - let mgr = TransactionManager::new(TransactionConfig::default()); - let txn_id = mgr.begin().await.expect("TODO: handle error"); - mgr.commit(&txn_id).await.expect("TODO: handle error"); - - let result = mgr - .buffer_operation( - &txn_id, - BufferedOperation { - entity_id: "hex-003".to_string(), - operation: OperationType::Delete, - payload: b"{}".to_vec(), - timestamp: Utc::now().to_rfc3339(), - }, - ) - .await; - assert!(result.is_err()); - } - - #[tokio::test] - async fn test_max_concurrent_transactions() { - let mgr = TransactionManager::new(TransactionConfig { - max_concurrent: 2, - timeout_seconds: 300, - }); - - let _t1 = mgr.begin().await.expect("TODO: handle error"); - let _t2 = mgr.begin().await.expect("TODO: handle error"); - let result = mgr.begin().await; - assert!(result.is_err()); - } - - #[tokio::test] - async fn test_not_found() { - let mgr = TransactionManager::new(TransactionConfig::default()); - let fake_id = TransactionId("nonexistent".to_string()); - assert!(mgr.status(&fake_id).await.is_err()); - assert!(mgr.commit(&fake_id).await.is_err()); - assert!(mgr.rollback(&fake_id).await.is_err()); - } -} diff --git a/verisimdb/rust-core/verisim-api/src/vcl.rs b/verisimdb/rust-core/verisim-api/src/vcl.rs deleted file mode 100644 index 45050666..00000000 --- a/verisimdb/rust-core/verisim-api/src/vcl.rs +++ /dev/null @@ -1,984 +0,0 @@ -// SPDX-License-Identifier: MPL-2.0 -// Copyright (c) 2026 Jonathan D.A. Jewell (hyperpolymath) -//! -//! VCL execution endpoint — accepts VCL text queries, parses them, and -//! routes to the appropriate store operations. -//! -//! This is a lightweight server-side VCL parser that handles the core -//! query operations directly against the octad store. It bridges the gap -//! between the REPL client (which sends raw VCL text) and the REST API -//! (which expects structured JSON requests). -//! -//! ## Supported VCL Statements -//! -//! - `SELECT [modalities] FROM octads [WHERE id = '...'] [LIMIT n]` -//! - `SEARCH TEXT '' [LIMIT n]` -//! - `SEARCH VECTOR [v1, v2, ...] [LIMIT n]` -//! - `SEARCH RELATED '' [BY '']` -//! - `INSERT INTO octads (fields...) VALUES (values...)` -//! - `DELETE FROM octads WHERE id = ''` -//! - `SHOW STATUS` / `SHOW DRIFT` / `SHOW NORMALIZER` -//! - `SHOW OCTADS [LIMIT n]` -//! - `COUNT octads` -//! - `EXPLAIN ` - -use axum::{extract::State, Json}; -use serde::{Deserialize, Serialize}; -use serde_json::{json, Value}; -use tracing::{info, instrument}; - -use verisim_octad::{OctadId, OctadInput, OctadDocumentInput, OctadStore}; - -use crate::{ApiError, AppState, OctadResponse}; - -/// VCL execute request — wraps a raw VCL query string. -#[derive(Debug, Deserialize)] -pub struct VclExecuteRequest { - /// The VCL query text to parse and execute. - pub query: String, -} - -/// VCL execute response — returns structured results from a query. -#[derive(Debug, Serialize)] -pub struct VclExecuteResponse { - /// Whether the query executed successfully. - pub success: bool, - /// The type of statement that was executed. - pub statement_type: String, - /// Number of rows/items in the result. - pub row_count: usize, - /// The result data (schema depends on statement type). - pub data: Value, - /// Optional message (e.g., for INSERT/DELETE confirmations). - #[serde(skip_serializing_if = "Option::is_none")] - pub message: Option, - /// VCL-UT safety level achieved (0-9), if TypeLL validation was performed. - #[serde(skip_serializing_if = "Option::is_none")] - pub safety_level: Option, - /// VCL-UT query path used: "VCL (Slipstream)", "VCL-UT", or "VCL-UT". - #[serde(skip_serializing_if = "Option::is_none")] - pub query_path: Option, - /// Proof obligations generated by TypeLL (for ECHIDNA dispatch). - #[serde(skip_serializing_if = "Option::is_none")] - pub proof_obligations: Option>, -} - -/// Execute a VCL query string against the database. -/// -/// Parses the query, determines the operation, executes it against the -/// octad store, and returns structured results. -#[instrument(skip(state, request), fields(query = %request.query))] -pub async fn vcl_execute_handler( - State(state): State, - Json(request): Json, -) -> Result, ApiError> { - let query = request.query.trim(); - - if query.is_empty() { - return Err(ApiError::BadRequest("Empty query".to_string())); - } - - // Normalize: strip trailing semicolons, collapse whitespace. - let query = query.trim_end_matches(';').trim(); - - // Parse and route the query. - let tokens = tokenize(query); - if tokens.is_empty() { - return Err(ApiError::BadRequest("Empty query after parsing".to_string())); - } - - // Pre-flight: validate with TypeLL VCL-UT checker if available. - // This is non-blocking — if TypeLL is unreachable, we proceed without it. - let typell_report = validate_with_typell(query).await; - - let mut result = match tokens[0].to_uppercase().as_str() { - "SELECT" => execute_select(&state, &tokens, query).await, - "SEARCH" => execute_search(&state, &tokens).await, - "INSERT" => execute_insert(&state, query).await, - "DELETE" => execute_delete(&state, &tokens).await, - "SHOW" => execute_show(&state, &tokens).await, - "COUNT" => execute_count(&state, &tokens).await, - "EXPLAIN" => execute_explain(&state, &tokens, query).await, - other => Err(ApiError::BadRequest(format!( - "Unknown VCL statement: '{}'. Supported: SELECT, SEARCH, INSERT, DELETE, SHOW, COUNT, EXPLAIN", - other - ))), - }?; - - // Attach TypeLL safety metadata to the response - if let Some((level, path, obligations)) = typell_report { - result.safety_level = Some(level); - result.query_path = Some(path); - if !obligations.is_empty() { - result.proof_obligations = Some(obligations); - } - } - - info!( - statement_type = %result.statement_type, - row_count = result.row_count, - safety_level = ?result.safety_level, - "VCL query executed" - ); - - Ok(Json(result)) -} - -/// Tokenize a VCL query into whitespace-separated tokens, respecting -/// quoted strings (single and double quotes). -/// Validate a VCL query against TypeLL's VCL-UT 10-level type checker. -/// -/// Calls the TypeLL server (default: localhost:7800) to verify the query's -/// type safety level. Returns (safety_level, query_path, proof_obligations) -/// if TypeLL is reachable, or None if it's unavailable. -/// -/// This is a non-blocking, best-effort validation — query execution proceeds -/// regardless. The safety metadata is attached to the response for downstream -/// consumers (PanLL, ECHIDNA) to act on. -async fn validate_with_typell(query: &str) -> Option<(u8, String, Vec)> { - let typell_url = std::env::var("TYPELL_URL") - .unwrap_or_else(|_| "http://localhost:7800".to_string()); - - let check_body = serde_json::json!({ - "expression": query, - "context": "{\"language\":\"vcl\",\"dialect\":\"vcl-ut\",\"features\":[\"dependent\",\"linear\",\"proof-carrying\"]}" - }); - - let client = reqwest::Client::builder() - .timeout(std::time::Duration::from_secs(3)) - .build() - .ok()?; - - let resp = client - .post(format!("{}/api/v1/check", typell_url)) - .json(&check_body) - .send() - .await - .ok()?; - - if !resp.status().is_success() { - return None; - } - - let body: serde_json::Value = resp.json().await.ok()?; - - // Extract VCL-UT level from features (e.g., "vcl-ut-l4") - let level = body - .get("features") - .and_then(|f| f.as_array()) - .and_then(|arr| { - arr.iter().find_map(|v| { - let s = v.as_str()?; - if let Some(stripped) = s.strip_prefix("vcl-ut-l") { - stripped.parse::().ok() - } else { - None - } - }) - }) - .unwrap_or(0); - - // Extract query path from features (e.g., "path:VCL-UT") - let path = body - .get("features") - .and_then(|f| f.as_array()) - .and_then(|arr| { - arr.iter().find_map(|v| { - let s = v.as_str()?; - s.strip_prefix("path:").map(|p| p.to_string()) - }) - }) - .unwrap_or_else(|| "VCL (Slipstream)".to_string()); - - // Extract proof obligations - let obligations: Vec = body - .get("proof_obligations") - .and_then(|p| p.as_array()) - .map(|arr| { - arr.iter() - .filter_map(|v| v.as_str().map(|s| s.to_string())) - .collect() - }) - .unwrap_or_default(); - - info!( - safety_level = level, - query_path = %path, - obligations = obligations.len(), - "TypeLL VCL-UT pre-flight check complete" - ); - - Some((level, path, obligations)) -} - -fn tokenize(input: &str) -> Vec { - let mut tokens = Vec::new(); - let mut current = String::new(); - let mut in_single_quote = false; - let mut in_double_quote = false; - - for ch in input.chars() { - match ch { - '\'' if !in_double_quote => { - in_single_quote = !in_single_quote; - current.push(ch); - } - '"' if !in_single_quote => { - in_double_quote = !in_double_quote; - current.push(ch); - } - ' ' | '\t' | '\n' if !in_single_quote && !in_double_quote => { - if !current.is_empty() { - tokens.push(std::mem::take(&mut current)); - } - } - _ => { - current.push(ch); - } - } - } - if !current.is_empty() { - tokens.push(current); - } - - tokens -} - -/// Strip surrounding quotes (single or double) from a string. -fn unquote(s: &str) -> &str { - if (s.starts_with('\'') && s.ends_with('\'')) || (s.starts_with('"') && s.ends_with('"')) { - &s[1..s.len() - 1] - } else { - s - } -} - -/// Parse a LIMIT clause from the end of the token list. -/// Returns (limit_value, index_of_limit_keyword_or_end). -fn parse_limit(tokens: &[String]) -> (usize, usize) { - for (i, token) in tokens.iter().enumerate() { - if token.to_uppercase() == "LIMIT" { - if let Some(next) = tokens.get(i + 1) { - if let Ok(n) = next.parse::() { - return (n.min(1000), i); - } - } - } - } - (100, tokens.len()) // default limit -} - -// --------------------------------------------------------------------------- -// SELECT -// --------------------------------------------------------------------------- - -/// Execute a SELECT query. -/// -/// Supported forms: -/// - `SELECT * FROM octads` — list all octads -/// - `SELECT * FROM octads WHERE id = ''` — get one octad -/// - `SELECT * FROM octads LIMIT n` — list with limit -async fn execute_select( - state: &AppState, - tokens: &[String], - _raw: &str, -) -> Result { - let (limit, _) = parse_limit(tokens); - - // Check for WHERE id = '...' - let where_id = find_where_id(tokens); - - if let Some(id) = where_id { - // Single octad lookup - let octad_id = OctadId::new(id); - let octad = state - .octad_store - .get(&octad_id) - .await - .map_err(|e| ApiError::Internal(e.to_string()))? - .ok_or_else(|| ApiError::NotFound(format!("Octad '{}' not found", id)))?; - - let response = OctadResponse::from(&octad); - Ok(VclExecuteResponse { - success: true, - statement_type: "SELECT".to_string(), - row_count: 1, - data: serde_json::to_value(vec![response]) - .map_err(|e| ApiError::Serialization(e.to_string()))?, - message: None, - safety_level: None, - query_path: None, - proof_obligations: None, - }) - } else { - // List octads - let octads = state - .octad_store - .list(limit, 0) - .await - .map_err(|e| ApiError::Internal(e.to_string()))?; - - let responses: Vec = octads.iter().map(OctadResponse::from).collect(); - let count = responses.len(); - - Ok(VclExecuteResponse { - success: true, - statement_type: "SELECT".to_string(), - row_count: count, - data: serde_json::to_value(responses) - .map_err(|e| ApiError::Serialization(e.to_string()))?, - message: None, - safety_level: None, - query_path: None, - proof_obligations: None, - }) - } -} - -/// Find `WHERE id = ''` in token list. -fn find_where_id<'a>(tokens: &'a [String]) -> Option<&'a str> { - for (i, token) in tokens.iter().enumerate() { - if token.to_uppercase() == "WHERE" { - // Expect: WHERE id = '' - if tokens.get(i + 1).map(|t| t.to_lowercase()) == Some("id".to_string()) { - if tokens.get(i + 2).map(|t| t.as_str()) == Some("=") { - if let Some(val) = tokens.get(i + 3) { - return Some(unquote(val)); - } - } - } - } - } - None -} - -// --------------------------------------------------------------------------- -// SEARCH -// --------------------------------------------------------------------------- - -/// Execute a SEARCH query. -/// -/// Supported forms: -/// - `SEARCH TEXT '' [LIMIT n]` -/// - `SEARCH VECTOR [v1, v2, ...] [LIMIT n]` -/// - `SEARCH RELATED '' [BY '']` -async fn execute_search( - state: &AppState, - tokens: &[String], -) -> Result { - if tokens.len() < 3 { - return Err(ApiError::BadRequest( - "SEARCH requires at least: SEARCH TEXT '' or SEARCH VECTOR [...]".to_string(), - )); - } - - match tokens[1].to_uppercase().as_str() { - "TEXT" => { - let query_text = unquote(&tokens[2]); - let (limit, _) = parse_limit(tokens); - - let octads = state - .octad_store - .search_text(query_text, limit) - .await - .map_err(|e| ApiError::Internal(e.to_string()))?; - - let results: Vec = octads - .iter() - .enumerate() - .map(|(i, h)| { - json!({ - "id": h.id.to_string(), - "score": 1.0 - (i as f64 * 0.1), - "title": h.document.as_ref().map(|d| d.title.clone()), - "has_graph": h.graph_node.is_some(), - "has_vector": h.embedding.is_some(), - "has_document": h.document.is_some(), - }) - }) - .collect(); - - let count = results.len(); - Ok(VclExecuteResponse { - success: true, - statement_type: "SEARCH TEXT".to_string(), - row_count: count, - data: json!(results), - message: None, - safety_level: None, - query_path: None, - proof_obligations: None, - }) - } - "VECTOR" => { - // Parse vector: [v1, v2, v3, ...] - // Tokens after VECTOR up to LIMIT are the vector components. - let (limit, limit_idx) = parse_limit(tokens); - let vector_str: String = tokens[2..limit_idx].join(" "); - let vector = parse_vector(&vector_str)?; - - if vector.len() != state.config.vector_dimension { - return Err(ApiError::BadRequest(format!( - "Vector dimension mismatch: expected {}, got {}", - state.config.vector_dimension, - vector.len() - ))); - } - - let octads = state - .octad_store - .search_similar(&vector, limit) - .await - .map_err(|e| ApiError::Internal(e.to_string()))?; - - let results: Vec = octads - .iter() - .enumerate() - .map(|(i, h)| { - json!({ - "id": h.id.to_string(), - "score": 1.0 - (i as f64 * 0.1), - "title": h.document.as_ref().map(|d| d.title.clone()), - }) - }) - .collect(); - - let count = results.len(); - Ok(VclExecuteResponse { - success: true, - statement_type: "SEARCH VECTOR".to_string(), - row_count: count, - data: json!(results), - message: None, - safety_level: None, - query_path: None, - proof_obligations: None, - }) - } - "RELATED" => { - if tokens.len() < 3 { - return Err(ApiError::BadRequest( - "SEARCH RELATED requires: SEARCH RELATED '' [BY '']".to_string(), - )); - } - let id = unquote(&tokens[2]); - let octad_id = OctadId::new(id); - - let predicate = tokens - .iter() - .position(|t| t.to_uppercase() == "BY") - .and_then(|i| tokens.get(i + 1)) - .map(|t| unquote(t)) - .unwrap_or("related"); - - let octads = state - .octad_store - .query_related(&octad_id, predicate) - .await - .map_err(|e| ApiError::Internal(e.to_string()))?; - - let responses: Vec = octads.iter().map(OctadResponse::from).collect(); - let count = responses.len(); - - Ok(VclExecuteResponse { - success: true, - statement_type: "SEARCH RELATED".to_string(), - row_count: count, - data: serde_json::to_value(responses) - .map_err(|e| ApiError::Serialization(e.to_string()))?, - message: None, - safety_level: None, - query_path: None, - proof_obligations: None, - }) - } - other => Err(ApiError::BadRequest(format!( - "Unknown SEARCH type: '{}'. Use TEXT, VECTOR, or RELATED.", - other - ))), - } -} - -/// Parse a vector from a string like `[0.1, 0.2, 0.3]` or `0.1 0.2 0.3`. -fn parse_vector(s: &str) -> Result, ApiError> { - let cleaned = s - .trim() - .trim_start_matches('[') - .trim_end_matches(']') - .replace(',', " "); - - let values: Result, _> = cleaned - .split_whitespace() - .filter(|s| !s.is_empty()) - .map(|v| v.parse::()) - .collect(); - - values.map_err(|e| ApiError::BadRequest(format!("Invalid vector: {}", e))) -} - -// --------------------------------------------------------------------------- -// INSERT -// --------------------------------------------------------------------------- - -/// Execute an INSERT statement. -/// -/// Supported form: -/// `INSERT INTO octads (title, body) VALUES ('', '<body>')` -/// -/// Also accepts simplified form: -/// `INSERT '<title>' '<body>'` -async fn execute_insert( - state: &AppState, - raw: &str, -) -> Result<VclExecuteResponse, ApiError> { - let upper = raw.to_uppercase(); - - let (title, body) = if upper.starts_with("INSERT INTO") { - // Parse: INSERT INTO octads (title, body) VALUES ('...', '...') - parse_insert_values(raw)? - } else { - // Simplified: INSERT '<title>' '<body>' - let tokens = tokenize(raw); - if tokens.len() < 3 { - return Err(ApiError::BadRequest( - "INSERT requires: INSERT INTO octads (title, body) VALUES ('<title>', '<body>')".to_string(), - )); - } - ( - unquote(&tokens[1]).to_string(), - unquote(&tokens[2]).to_string(), - ) - }; - - let mut input = OctadInput::default(); - input.document = Some(OctadDocumentInput { - title: title.clone(), - body, - fields: std::collections::HashMap::new(), - }); - - let octad = state - .octad_store - .create(input) - .await - .map_err(|e| ApiError::Internal(e.to_string()))?; - - let response = OctadResponse::from(&octad); - - Ok(VclExecuteResponse { - success: true, - statement_type: "INSERT".to_string(), - row_count: 1, - data: serde_json::to_value(vec![&response]) - .map_err(|e| ApiError::Serialization(e.to_string()))?, - message: Some(format!("Inserted octad '{}'", response.id)), - safety_level: None, - query_path: None, - proof_obligations: None, - }) -} - -/// Parse VALUES clause from INSERT INTO ... VALUES ('...', '...'). -fn parse_insert_values(raw: &str) -> Result<(String, String), ApiError> { - let upper = raw.to_uppercase(); - let values_idx = upper - .find("VALUES") - .ok_or_else(|| ApiError::BadRequest("INSERT INTO requires a VALUES clause".to_string()))?; - - let values_part = &raw[values_idx + 6..].trim(); - let values_part = values_part - .trim_start_matches('(') - .trim_end_matches(')'); - - // Split on comma, respecting quotes. - let value_tokens = tokenize(&values_part.replace(',', " ")); - - let title = value_tokens - .first() - .map(|s| unquote(s).to_string()) - .unwrap_or_default(); - let body = value_tokens - .get(1) - .map(|s| unquote(s).to_string()) - .unwrap_or_default(); - - Ok((title, body)) -} - -// --------------------------------------------------------------------------- -// DELETE -// --------------------------------------------------------------------------- - -/// Execute a DELETE statement. -/// -/// Supported form: -/// `DELETE FROM octads WHERE id = '<id>'` -async fn execute_delete( - state: &AppState, - tokens: &[String], -) -> Result<VclExecuteResponse, ApiError> { - let id = find_where_id(tokens).ok_or_else(|| { - ApiError::BadRequest( - "DELETE requires: DELETE FROM octads WHERE id = '<id>'".to_string(), - ) - })?; - - let octad_id = OctadId::new(id); - - state - .octad_store - .delete(&octad_id) - .await - .map_err(|e| match e { - verisim_octad::OctadError::NotFound(_) => { - ApiError::NotFound(format!("Octad '{}' not found", id)) - } - _ => ApiError::Internal(e.to_string()), - })?; - - Ok(VclExecuteResponse { - success: true, - statement_type: "DELETE".to_string(), - row_count: 1, - data: json!(null), - message: Some(format!("Deleted octad '{}'", id)), - safety_level: None, - query_path: None, - proof_obligations: None, - }) -} - -// --------------------------------------------------------------------------- -// SHOW -// --------------------------------------------------------------------------- - -/// Execute a SHOW query. -/// -/// Supported forms: -/// - `SHOW STATUS` — server health -/// - `SHOW DRIFT` — drift metrics -/// - `SHOW NORMALIZER` — normalizer status -/// - `SHOW OCTADS [LIMIT n]` — list octads (alias for SELECT) -async fn execute_show( - state: &AppState, - tokens: &[String], -) -> Result<VclExecuteResponse, ApiError> { - if tokens.len() < 2 { - return Err(ApiError::BadRequest( - "SHOW requires: SHOW STATUS | SHOW DRIFT | SHOW NORMALIZER | SHOW OCTADS".to_string(), - )); - } - - match tokens[1].to_uppercase().as_str() { - "STATUS" | "HEALTH" => { - let uptime = state.start_time.elapsed().as_secs(); - let version = env!("CARGO_PKG_VERSION"); - - let health = state.drift_detector.health_check(); - let (status, reason) = match health { - Ok(h) => { - use verisim_drift::HealthStatus; - match h.status { - HealthStatus::Critical | HealthStatus::Degraded => ( - "degraded", - Some(format!("{:?}: {:.3}", h.worst_drift_type, h.worst_score)), - ), - _ => ("healthy", None), - } - } - Err(_) => ("degraded", Some("Drift detector unavailable".to_string())), - }; - - Ok(VclExecuteResponse { - success: true, - statement_type: "SHOW STATUS".to_string(), - row_count: 1, - data: json!({ - "status": status, - "version": version, - "uptime_seconds": uptime, - "degraded_reason": reason, - }), - message: None, - safety_level: None, - query_path: None, - proof_obligations: None, - }) - } - "DRIFT" => { - let all_metrics = state - .drift_detector - .all_metrics() - .map_err(|e| ApiError::Internal(e.to_string()))?; - - let results: Vec<Value> = all_metrics - .iter() - .map(|(drift_type, metrics)| { - json!({ - "drift_type": drift_type.to_string(), - "current_score": metrics.current_score, - "moving_average": metrics.moving_average, - "max_score": metrics.max_score, - "measurement_count": metrics.measurement_count, - }) - }) - .collect(); - - let count = results.len(); - Ok(VclExecuteResponse { - success: true, - statement_type: "SHOW DRIFT".to_string(), - row_count: count, - data: json!(results), - message: None, - safety_level: None, - query_path: None, - proof_obligations: None, - }) - } - "NORMALIZER" => { - let status = state.normalizer.status().await; - Ok(VclExecuteResponse { - success: true, - statement_type: "SHOW NORMALIZER".to_string(), - row_count: 1, - data: serde_json::to_value(status) - .map_err(|e| ApiError::Serialization(e.to_string()))?, - message: None, - safety_level: None, - query_path: None, - proof_obligations: None, - }) - } - "OCTADS" => { - let (limit, _) = parse_limit(tokens); - let octads = state - .octad_store - .list(limit, 0) - .await - .map_err(|e| ApiError::Internal(e.to_string()))?; - - let responses: Vec<OctadResponse> = octads.iter().map(OctadResponse::from).collect(); - let count = responses.len(); - - Ok(VclExecuteResponse { - success: true, - statement_type: "SHOW OCTADS".to_string(), - row_count: count, - data: serde_json::to_value(responses) - .map_err(|e| ApiError::Serialization(e.to_string()))?, - message: None, - safety_level: None, - query_path: None, - proof_obligations: None, - }) - } - other => Err(ApiError::BadRequest(format!( - "Unknown SHOW target: '{}'. Use STATUS, DRIFT, NORMALIZER, or OCTADS.", - other - ))), - } -} - -// --------------------------------------------------------------------------- -// COUNT -// --------------------------------------------------------------------------- - -/// Execute a COUNT query. -/// -/// Supported form: -/// - `COUNT octads` — return total octad count -async fn execute_count( - state: &AppState, - _tokens: &[String], -) -> Result<VclExecuteResponse, ApiError> { - // List with a large limit to count (in a real DB this would be a COUNT query). - let octads = state - .octad_store - .list(1000, 0) - .await - .map_err(|e| ApiError::Internal(e.to_string()))?; - - let count = octads.len(); - - Ok(VclExecuteResponse { - success: true, - statement_type: "COUNT".to_string(), - row_count: 1, - data: json!({ "count": count }), - message: None, - safety_level: None, - query_path: None, - proof_obligations: None, - }) -} - -// --------------------------------------------------------------------------- -// EXPLAIN -// --------------------------------------------------------------------------- - -/// Execute an EXPLAIN query — describes what a query would do without executing. -/// -/// Supported form: -/// - `EXPLAIN <any VCL query>` -async fn execute_explain( - _state: &AppState, - tokens: &[String], - raw: &str, -) -> Result<VclExecuteResponse, ApiError> { - if tokens.len() < 2 { - return Err(ApiError::BadRequest("EXPLAIN requires a query to explain".to_string())); - } - - let inner_query = &raw[raw.to_uppercase().find("EXPLAIN").expect("TODO: handle error") + 7..].trim(); - let inner_tokens = tokenize(inner_query); - - if inner_tokens.is_empty() { - return Err(ApiError::BadRequest("EXPLAIN requires a query".to_string())); - } - - let statement_type = inner_tokens[0].to_uppercase(); - let (limit, _) = parse_limit(&inner_tokens); - let where_id = find_where_id(&inner_tokens); - - let plan = match statement_type.as_str() { - "SELECT" => { - if where_id.is_some() { - json!({ - "operation": "Point Lookup", - "target": "octad_store", - "method": "get_by_id", - "cost": "O(1)", - "estimated_rows": 1, - }) - } else { - json!({ - "operation": "Sequential Scan", - "target": "octad_store", - "method": "list", - "limit": limit, - "cost": "O(n)", - "estimated_rows": limit, - }) - } - } - "SEARCH" => { - let search_type = inner_tokens.get(1).map(|t| t.to_uppercase()).unwrap_or_default(); - match search_type.as_str() { - "TEXT" => json!({ - "operation": "Full-Text Search", - "target": "tantivy_document_store", - "method": "search_text", - "limit": limit, - "cost": "O(log n)", - "index": "tantivy_inverted_index", - }), - "VECTOR" => json!({ - "operation": "Approximate Nearest Neighbor", - "target": "hnsw_vector_store", - "method": "search_similar", - "limit": limit, - "cost": "O(log n)", - "index": "hnsw_graph", - }), - "RELATED" => json!({ - "operation": "Graph Traversal", - "target": "oxigraph_store", - "method": "query_related", - "cost": "O(degree)", - "index": "rdf_triple_index", - }), - _ => json!({"operation": "Unknown search type"}), - } - } - "INSERT" => json!({ - "operation": "Multi-Modal Insert", - "targets": ["document_store", "graph_store", "vector_store", "semantic_store", "temporal_store"], - "method": "create", - "cost": "O(1) per modality", - }), - "DELETE" => json!({ - "operation": "Multi-Modal Delete", - "targets": ["all_modality_stores"], - "method": "delete", - "cost": "O(1)", - }), - _ => json!({"operation": format!("Unrecognized: {}", statement_type)}), - }; - - Ok(VclExecuteResponse { - success: true, - statement_type: "EXPLAIN".to_string(), - row_count: 1, - data: json!({ - "query": inner_query, - "plan": plan, - }), - message: None, - safety_level: None, - query_path: None, - proof_obligations: None, - }) -} - -#[cfg(test)] -mod tests { - use super::*; - - #[test] - fn test_tokenize_simple() { - let tokens = tokenize("SELECT * FROM octads"); - assert_eq!(tokens, vec!["SELECT", "*", "FROM", "octads"]); - } - - #[test] - fn test_tokenize_quoted() { - let tokens = tokenize("SEARCH TEXT 'hello world' LIMIT 10"); - assert_eq!(tokens, vec!["SEARCH", "TEXT", "'hello world'", "LIMIT", "10"]); - } - - #[test] - fn test_unquote() { - assert_eq!(unquote("'hello'"), "hello"); - assert_eq!(unquote("\"hello\""), "hello"); - assert_eq!(unquote("hello"), "hello"); - } - - #[test] - fn test_parse_limit() { - let tokens: Vec<String> = vec!["SELECT", "*", "FROM", "octads", "LIMIT", "50"] - .into_iter() - .map(String::from) - .collect(); - let (limit, idx) = parse_limit(&tokens); - assert_eq!(limit, 50); - assert_eq!(idx, 4); - } - - #[test] - fn test_parse_limit_default() { - let tokens: Vec<String> = vec!["SELECT", "*", "FROM", "octads"] - .into_iter() - .map(String::from) - .collect(); - let (limit, idx) = parse_limit(&tokens); - assert_eq!(limit, 100); - assert_eq!(idx, 4); - } - - #[test] - fn test_find_where_id() { - let tokens: Vec<String> = vec!["SELECT", "*", "FROM", "octads", "WHERE", "id", "=", "'abc-123'"] - .into_iter() - .map(String::from) - .collect(); - assert_eq!(find_where_id(&tokens), Some("abc-123")); - } - - #[test] - fn test_parse_vector() { - let v = parse_vector("[0.1, 0.2, 0.3]").expect("TODO: handle error"); - assert_eq!(v.len(), 3); - assert!((v[0] - 0.1).abs() < 0.001); - } -} diff --git a/verisimdb/rust-core/verisim-document/Cargo.toml b/verisimdb/rust-core/verisim-document/Cargo.toml deleted file mode 100644 index a994d70a..00000000 --- a/verisimdb/rust-core/verisim-document/Cargo.toml +++ /dev/null @@ -1,21 +0,0 @@ -# SPDX-License-Identifier: MPL-2.0 - -[package] -name = "verisim-document" -description = "Document modality - full-text search via Tantivy" -version.workspace = true -edition.workspace = true -authors.workspace = true -license.workspace = true - -[dependencies] -tantivy.workspace = true -serde.workspace = true -serde_json.workspace = true -thiserror.workspace = true -tracing.workspace = true -async-trait.workspace = true -tokio.workspace = true - -[dev-dependencies] -proptest.workspace = true diff --git a/verisimdb/rust-core/verisim-document/src/lib.rs b/verisimdb/rust-core/verisim-document/src/lib.rs deleted file mode 100644 index c900b389..00000000 --- a/verisimdb/rust-core/verisim-document/src/lib.rs +++ /dev/null @@ -1,332 +0,0 @@ -// SPDX-License-Identifier: MPL-2.0 -//! VeriSim Document Modality -//! -//! Full-text search via Tantivy. -//! Implements Marr's Computational Level: "What text matches?" - -#![forbid(unsafe_code)] -use async_trait::async_trait; -use serde::{Deserialize, Serialize}; -use std::collections::HashMap; -use std::path::Path; -use std::sync::Arc; -use tantivy::collector::TopDocs; -use tantivy::query::QueryParser; -use tantivy::schema::{Field, Schema, Value, STORED, TEXT}; -use tantivy::snippet::SnippetGenerator; -use tantivy::{Index, IndexReader, IndexWriter, ReloadPolicy, TantivyDocument}; -use thiserror::Error; -use tokio::sync::RwLock; - -/// Document modality errors -#[derive(Error, Debug)] -pub enum DocumentError { - #[error("Index error: {0}")] - IndexError(String), - - #[error("Document not found: {0}")] - NotFound(String), - - #[error("Query parse error: {0}")] - QueryError(String), - - #[error("Schema error: {0}")] - SchemaError(String), - - #[error("IO error: {0}")] - IoError(#[from] std::io::Error), -} - -impl From<tantivy::TantivyError> for DocumentError { - fn from(e: tantivy::TantivyError) -> Self { - DocumentError::IndexError(e.to_string()) - } -} - -impl From<tantivy::query::QueryParserError> for DocumentError { - fn from(e: tantivy::query::QueryParserError) -> Self { - DocumentError::QueryError(e.to_string()) - } -} - -impl From<tantivy::directory::error::OpenDirectoryError> for DocumentError { - fn from(e: tantivy::directory::error::OpenDirectoryError) -> Self { - DocumentError::IoError(std::io::Error::other(e.to_string())) - } -} - -/// A document for full-text indexing -#[derive(Debug, Clone, Serialize, Deserialize)] -pub struct Document { - /// Unique identifier (matches Octad entity ID) - pub id: String, - /// Document title - pub title: String, - /// Main content body - pub body: String, - /// Additional searchable fields - pub fields: HashMap<String, String>, - /// Non-searchable metadata - pub metadata: HashMap<String, String>, -} - -impl Document { - /// Create a new document - pub fn new(id: impl Into<String>, title: impl Into<String>, body: impl Into<String>) -> Self { - Self { - id: id.into(), - title: title.into(), - body: body.into(), - fields: HashMap::new(), - metadata: HashMap::new(), - } - } - - /// Add a searchable field - pub fn with_field(mut self, key: impl Into<String>, value: impl Into<String>) -> Self { - self.fields.insert(key.into(), value.into()); - self - } - - /// Add metadata - pub fn with_metadata(mut self, key: impl Into<String>, value: impl Into<String>) -> Self { - self.metadata.insert(key.into(), value.into()); - self - } -} - -/// Search result with score and highlights -#[derive(Debug, Clone, Serialize, Deserialize)] -pub struct SearchResult { - /// Document ID - pub id: String, - /// Relevance score - pub score: f32, - /// Document title - pub title: String, - /// Snippet with highlights - pub snippet: Option<String>, -} - -/// Document store trait for cross-modal consistency -#[async_trait] -pub trait DocumentStore: Send + Sync { - /// Index a document - async fn index(&self, doc: &Document) -> Result<(), DocumentError>; - - /// Search documents - async fn search(&self, query: &str, limit: usize) -> Result<Vec<SearchResult>, DocumentError>; - - /// Get document by ID - async fn get(&self, id: &str) -> Result<Option<Document>, DocumentError>; - - /// Delete document by ID - async fn delete(&self, id: &str) -> Result<(), DocumentError>; - - /// Commit pending changes - async fn commit(&self) -> Result<(), DocumentError>; -} - -/// Schema fields for Tantivy -struct DocumentSchema { - id: Field, - title: Field, - body: Field, - schema: Schema, -} - -impl DocumentSchema { - fn new() -> Self { - let mut schema_builder = Schema::builder(); - let id = schema_builder.add_text_field("id", TEXT | STORED); - let title = schema_builder.add_text_field("title", TEXT | STORED); - let body = schema_builder.add_text_field("body", TEXT | STORED); - let schema = schema_builder.build(); - - Self { id, title, body, schema } - } -} - -/// Tantivy-backed document store -pub struct TantivyDocumentStore { - schema: DocumentSchema, - index: Index, - writer: Arc<RwLock<IndexWriter>>, - reader: IndexReader, - documents: Arc<RwLock<HashMap<String, Document>>>, -} - -impl TantivyDocumentStore { - /// Create an in-memory store - pub fn in_memory() -> Result<Self, DocumentError> { - let schema = DocumentSchema::new(); - let index = Index::create_in_ram(schema.schema.clone()); - let writer = index.writer(50_000_000)?; - let reader = index - .reader_builder() - .reload_policy(ReloadPolicy::OnCommitWithDelay) - .try_into()?; - - Ok(Self { - schema, - index, - writer: Arc::new(RwLock::new(writer)), - reader, - documents: Arc::new(RwLock::new(HashMap::new())), - }) - } - - /// Create a persistent store - pub fn persistent(path: impl AsRef<Path>) -> Result<Self, DocumentError> { - let schema = DocumentSchema::new(); - std::fs::create_dir_all(path.as_ref())?; - let dir = tantivy::directory::MmapDirectory::open(path)?; - let index = Index::open_or_create(dir, schema.schema.clone())?; - let writer = index.writer(50_000_000)?; - let reader = index - .reader_builder() - .reload_policy(ReloadPolicy::OnCommitWithDelay) - .try_into()?; - - Ok(Self { - schema, - index, - writer: Arc::new(RwLock::new(writer)), - reader, - documents: Arc::new(RwLock::new(HashMap::new())), - }) - } -} - -#[async_trait] -impl DocumentStore for TantivyDocumentStore { - async fn index(&self, doc: &Document) -> Result<(), DocumentError> { - let mut tantivy_doc = TantivyDocument::default(); - tantivy_doc.add_text(self.schema.id, &doc.id); - tantivy_doc.add_text(self.schema.title, &doc.title); - tantivy_doc.add_text(self.schema.body, &doc.body); - - // Delete existing document with same ID - let term = tantivy::Term::from_field_text(self.schema.id, &doc.id); - { - let writer = self.writer.write().await; - writer.delete_term(term); - writer.add_document(tantivy_doc)?; - } - - // Store original document - self.documents.write().await.insert(doc.id.clone(), doc.clone()); - - Ok(()) - } - - async fn search(&self, query: &str, limit: usize) -> Result<Vec<SearchResult>, DocumentError> { - let searcher = self.reader.searcher(); - let query_parser = QueryParser::for_index( - &self.index, - vec![self.schema.title, self.schema.body], - ); - - let parsed_query = query_parser.parse_query(query)?; - let top_docs = searcher.search(&parsed_query, &TopDocs::with_limit(limit).order_by_score())?; - - // Create snippet generator for body field - let snippet_generator = SnippetGenerator::create( - &searcher, - &parsed_query, - self.schema.body, - )?; - - let mut results = Vec::new(); - for (score, doc_address) in top_docs { - let retrieved_doc: TantivyDocument = searcher.doc(doc_address)?; - - let id = retrieved_doc - .get_first(self.schema.id) - .and_then(|v| v.as_str()) - .unwrap_or("") - .to_string(); - - let title = retrieved_doc - .get_first(self.schema.title) - .and_then(|v| v.as_str()) - .unwrap_or("") - .to_string(); - - // Generate snippet with highlights - let snippet = snippet_generator.snippet_from_doc(&retrieved_doc); - let snippet_html = snippet.to_html(); - let snippet_text = if snippet_html.is_empty() { - None - } else { - Some(snippet_html) - }; - - results.push(SearchResult { - id, - score, - title, - snippet: snippet_text, - }); - } - - Ok(results) - } - - async fn get(&self, id: &str) -> Result<Option<Document>, DocumentError> { - Ok(self.documents.read().await.get(id).cloned()) - } - - async fn delete(&self, id: &str) -> Result<(), DocumentError> { - let term = tantivy::Term::from_field_text(self.schema.id, id); - self.writer.write().await.delete_term(term); - self.documents.write().await.remove(id); - Ok(()) - } - - async fn commit(&self) -> Result<(), DocumentError> { - self.writer.write().await.commit()?; - self.reader.reload()?; - Ok(()) - } -} - -#[cfg(test)] -mod tests { - use super::*; - - #[tokio::test] - async fn test_index_and_search() { - let store = TantivyDocumentStore::in_memory().expect("TODO: handle error"); - - let doc1 = Document::new("d1", "Rust Programming", "Rust is a systems programming language"); - let doc2 = Document::new("d2", "Python Tutorial", "Python is great for beginners"); - - store.index(&doc1).await.expect("TODO: handle error"); - store.index(&doc2).await.expect("TODO: handle error"); - store.commit().await.expect("TODO: handle error"); - - let results = store.search("Rust", 10).await.expect("TODO: handle error"); - assert_eq!(results.len(), 1); - assert_eq!(results[0].id, "d1"); - } - - #[tokio::test] - async fn test_search_with_snippets() { - let store = TantivyDocumentStore::in_memory().expect("TODO: handle error"); - - let doc = Document::new( - "d1", - "Rust Guide", - "Rust is a systems programming language focused on safety and performance", - ); - store.index(&doc).await.expect("TODO: handle error"); - store.commit().await.expect("TODO: handle error"); - - let results = store.search("safety", 10).await.expect("TODO: handle error"); - assert_eq!(results.len(), 1); - assert!(results[0].snippet.is_some(), "Snippet should not be None"); - let snippet = results[0].snippet.as_ref().expect("TODO: handle error"); - assert!(snippet.contains("safety"), "Snippet should contain the search term"); - } -} diff --git a/verisimdb/rust-core/verisim-document/tests/property_tests.proptest-regressions b/verisimdb/rust-core/verisim-document/tests/property_tests.proptest-regressions deleted file mode 100644 index 7a7760c4..00000000 --- a/verisimdb/rust-core/verisim-document/tests/property_tests.proptest-regressions +++ /dev/null @@ -1,7 +0,0 @@ -# Seeds for failure cases proptest has generated in the past. It is -# automatically read and these particular cases re-run before any -# novel cases are generated. -# -# It is recommended to check this file in to source control so that -# everyone who runs the test benefits from these saved cases. -cc 36c667d0e33d21c6e1111c3f4c94cfd78d26ad2a41020d8ea3b7950eace8c54d # shrinks to id = "00aaaa0a", title = "a aaa", body = ".!., AaA?00A00AaA..?" diff --git a/verisimdb/rust-core/verisim-document/tests/property_tests.rs b/verisimdb/rust-core/verisim-document/tests/property_tests.rs deleted file mode 100644 index 8876a346..00000000 --- a/verisimdb/rust-core/verisim-document/tests/property_tests.rs +++ /dev/null @@ -1,224 +0,0 @@ -// SPDX-License-Identifier: MPL-2.0 -//! Property-based tests for document modality - -use proptest::prelude::*; -use verisim_document::{Document, DocumentStore, TantivyDocumentStore}; - -/// Generate arbitrary document IDs -fn arb_id() -> impl Strategy<Value = String> { - "[a-z0-9]{8,16}" -} - -/// Generate arbitrary document titles -fn arb_title() -> impl Strategy<Value = String> { - "[A-Za-z ]{5,50}" -} - -/// Generate arbitrary document bodies -fn arb_body() -> impl Strategy<Value = String> { - "[A-Za-z0-9 .,!?]{20,200}" -} - -proptest! { - #[test] - fn test_insert_then_retrieve( - id in arb_id(), - title in arb_title(), - body in arb_body() - ) { - let runtime = tokio::runtime::Runtime::new().unwrap(); - runtime.block_on(async { - let store = TantivyDocumentStore::in_memory().unwrap(); - let doc = Document::new(&id, &title, &body); - - // Index document - store.index(&doc).await.unwrap(); - store.commit().await.unwrap(); - - // Retrieve by ID - let retrieved = store.get(&id).await.unwrap(); - prop_assert!(retrieved.is_some()); - - let retrieved = retrieved.unwrap(); - prop_assert_eq!(&retrieved.id, &id); - prop_assert_eq!(&retrieved.title, &title); - prop_assert_eq!(&retrieved.body, &body); - - Ok(()) - })?; - } - - #[test] - fn test_search_returns_indexed_document( - id in arb_id(), - title in arb_title(), - body in arb_body() - ) { - let runtime = tokio::runtime::Runtime::new().unwrap(); - runtime.block_on(async { - let store = TantivyDocumentStore::in_memory().unwrap(); - let doc = Document::new(&id, &title, &body); - - store.index(&doc).await.unwrap(); - store.commit().await.unwrap(); - - // Search by ID should work via get() - let retrieved = store.get(&id).await.unwrap(); - prop_assert!(retrieved.is_some(), "Document should be retrievable by ID"); - - // Search by at least part of title - let title_words: Vec<&str> = title.split_whitespace().collect(); - if let Some(first_word) = title_words.first() { - if first_word.len() > 2 { - // Search might find it (depends on tokenization) - let _results = store.search(first_word, 10).await.unwrap(); - // Note: We don't assert results.len() > 0 because Tantivy tokenization - // may not match our expectations for all random inputs - } - } - - Ok(()) - })?; - } - - #[test] - fn test_update_document( - id in arb_id(), - title1 in arb_title(), - title2 in arb_title(), - body1 in arb_body(), - body2 in arb_body() - ) { - let runtime = tokio::runtime::Runtime::new().unwrap(); - runtime.block_on(async { - let store = TantivyDocumentStore::in_memory().unwrap(); - - // Index first version - let doc1 = Document::new(&id, &title1, &body1); - store.index(&doc1).await.unwrap(); - store.commit().await.unwrap(); - - // Update with second version (same ID, different content) - let doc2 = Document::new(&id, &title2, &body2); - store.index(&doc2).await.unwrap(); - store.commit().await.unwrap(); - - // Should only find the updated version - let retrieved = store.get(&id).await.unwrap(); - prop_assert!(retrieved.is_some()); - - let retrieved = retrieved.unwrap(); - prop_assert_eq!(&retrieved.title, &title2, "Should have updated title"); - prop_assert_eq!(&retrieved.body, &body2, "Should have updated body"); - - Ok(()) - })?; - } - - #[test] - fn test_delete_document( - id in arb_id(), - title in arb_title(), - body in arb_body() - ) { - let runtime = tokio::runtime::Runtime::new().unwrap(); - runtime.block_on(async { - let store = TantivyDocumentStore::in_memory().unwrap(); - let doc = Document::new(&id, &title, &body); - - // Index document - store.index(&doc).await.unwrap(); - store.commit().await.unwrap(); - - // Verify it exists - prop_assert!(store.get(&id).await.unwrap().is_some()); - - // Delete document - store.delete(&id).await.unwrap(); - store.commit().await.unwrap(); - - // Verify it's gone - prop_assert!(store.get(&id).await.unwrap().is_none()); - - Ok(()) - })?; - } - - #[test] - fn test_multiple_documents_search( - docs in prop::collection::vec( - (arb_id(), arb_title(), arb_body()), - 1..10 - ) - ) { - let runtime = tokio::runtime::Runtime::new().unwrap(); - runtime.block_on(async { - let store = TantivyDocumentStore::in_memory().unwrap(); - - // Index all documents - for (id, title, body) in &docs { - let doc = Document::new(id, title, body); - store.index(&doc).await.unwrap(); - } - store.commit().await.unwrap(); - - // Each document should be retrievable by ID - for (id, _title, _body) in &docs { - let retrieved = store.get(id).await.unwrap(); - prop_assert!(retrieved.is_some(), "Document {} should be retrievable", id); - } - - Ok(()) - })?; - } -} - -/// Integration test: realistic usage pattern -#[tokio::test] -async fn test_realistic_document_lifecycle() { - let store = TantivyDocumentStore::in_memory().unwrap(); - - // Create initial document - let doc = Document::new( - "paper-123", - "Machine Learning for Databases", - "This paper explores the use of machine learning techniques for optimizing database query plans." - ) - .with_field("authors", "Smith, Jones, Johnson") - .with_field("year", "2024") - .with_metadata("citation_count", "42"); - - // Index it - store.index(&doc).await.unwrap(); - store.commit().await.unwrap(); - - // Verify document exists by ID - let retrieved = store.get("paper-123").await.unwrap(); - assert!(retrieved.is_some()); - assert_eq!(retrieved.unwrap().title, "Machine Learning for Databases"); - - // Update the document - let updated_doc = Document::new( - "paper-123", - "Machine Learning for Database Query Optimization", - "This paper explores the use of deep learning techniques for optimizing database query plans and execution." - ) - .with_field("authors", "Smith, Jones, Johnson, Williams") - .with_field("year", "2024") - .with_metadata("citation_count", "45"); - - store.index(&updated_doc).await.unwrap(); - store.commit().await.unwrap(); - - // Verify update by ID (most reliable check) - let retrieved = store.get("paper-123").await.unwrap().unwrap(); - assert!(retrieved.title.contains("Query Optimization")); - assert!(retrieved.body.contains("deep learning")); - - // Delete the document - store.delete("paper-123").await.unwrap(); - store.commit().await.unwrap(); - - // Verify deletion by ID - assert!(store.get("paper-123").await.unwrap().is_none(), "Document should be deleted"); -} diff --git a/verisimdb/rust-core/verisim-drift/Cargo.toml b/verisimdb/rust-core/verisim-drift/Cargo.toml deleted file mode 100644 index 8d5c98c5..00000000 --- a/verisimdb/rust-core/verisim-drift/Cargo.toml +++ /dev/null @@ -1,21 +0,0 @@ -# SPDX-License-Identifier: MPL-2.0 - -[package] -name = "verisim-drift" -description = "Drift detection - monitors cross-modal consistency degradation" -version.workspace = true -edition.workspace = true -authors.workspace = true -license.workspace = true - -[dependencies] -serde.workspace = true -chrono.workspace = true -thiserror.workspace = true -tracing.workspace = true -async-trait.workspace = true -tokio.workspace = true -prometheus.workspace = true - -[dev-dependencies] -proptest.workspace = true diff --git a/verisimdb/rust-core/verisim-drift/src/calculator.rs b/verisimdb/rust-core/verisim-drift/src/calculator.rs deleted file mode 100644 index 9ba728d0..00000000 --- a/verisimdb/rust-core/verisim-drift/src/calculator.rs +++ /dev/null @@ -1,650 +0,0 @@ -// SPDX-License-Identifier: MPL-2.0 -//! Drift calculation algorithms -//! -//! Computes actual drift scores between modalities to detect consistency degradation. - -use crate::DriftType; - -/// Drift calculator for computing cross-modal consistency scores -pub struct DriftCalculator { - /// Minimum similarity threshold below which drift is detected - pub similarity_threshold: f64, -} - -impl Default for DriftCalculator { - fn default() -> Self { - Self { - similarity_threshold: 0.8, - } - } -} - -impl DriftCalculator { - /// Create a new drift calculator with custom threshold - pub fn new(similarity_threshold: f64) -> Self { - Self { similarity_threshold } - } - - /// Calculate semantic-vector drift score - /// - /// Measures how well the vector embedding captures the semantic meaning. - /// This is computed by comparing the embedding similarity with semantic type similarity. - /// - /// Returns a score from 0.0 (no drift) to 1.0 (maximum drift) - pub fn semantic_vector_drift( - &self, - embedding: &[f32], - semantic_types: &[String], - type_embeddings: &[(String, Vec<f32>)], - ) -> f64 { - if semantic_types.is_empty() || type_embeddings.is_empty() { - return 0.0; // No drift if no semantic types or embeddings to compare - } - - // Find type embeddings for the entity's semantic types - let relevant_embeddings: Vec<&Vec<f32>> = type_embeddings - .iter() - .filter(|(type_iri, _)| semantic_types.contains(type_iri)) - .map(|(_, emb)| emb) - .collect(); - - if relevant_embeddings.is_empty() { - return 0.0; - } - - // Compute average similarity with type embeddings - let mut total_similarity = 0.0; - for type_emb in &relevant_embeddings { - total_similarity += cosine_similarity_f32(embedding, type_emb); - } - let avg_similarity = total_similarity / relevant_embeddings.len() as f64; - - // Convert similarity to drift score (inverse relationship) - // High similarity = low drift, low similarity = high drift - let drift_score = 1.0 - avg_similarity; - drift_score.clamp(0.0, 1.0) - } - - /// Calculate graph-document drift score - /// - /// Measures consistency between graph relationships and document content. - /// Checks if entities mentioned in document have corresponding graph edges. - /// - /// Returns a score from 0.0 (no drift) to 1.0 (maximum drift) - pub fn graph_document_drift( - &self, - document_text: &str, - document_entities: &[String], - graph_relationships: &[(String, String)], // (predicate, target) - ) -> f64 { - if document_entities.is_empty() { - return 0.0; - } - - // Count how many document entities have corresponding graph relationships - let graph_targets: Vec<&String> = graph_relationships.iter().map(|(_, t)| t).collect(); - - let mut matched = 0; - for entity in document_entities { - // Check if entity appears in graph relationships - if graph_targets.iter().any(|t| t.contains(entity) || entity.contains(*t)) { - matched += 1; - } - // Also check if entity is mentioned in document - if !document_text.to_lowercase().contains(&entity.to_lowercase()) { - // Entity in list but not in document text - potential issue - } - } - - // Drift score based on coverage - let coverage = if !document_entities.is_empty() { - matched as f64 / document_entities.len() as f64 - } else { - 1.0 - }; - - // Also consider if there are graph relationships not reflected in document - let extra_graph_ratio = if !graph_relationships.is_empty() { - let unmatched_graph = graph_relationships - .iter() - .filter(|(_, t)| !document_entities.iter().any(|e| t.contains(e))) - .count(); - unmatched_graph as f64 / graph_relationships.len() as f64 - } else { - 0.0 - }; - - // Combine metrics: low coverage or high extra graph = high drift - let drift_score = (1.0 - coverage + extra_graph_ratio) / 2.0; - drift_score.clamp(0.0, 1.0) - } - - /// Calculate temporal consistency drift score - /// - /// Measures if there are inconsistencies in the version history, - /// such as conflicting updates or temporal anomalies. - /// - /// Returns a score from 0.0 (no drift) to 1.0 (maximum drift) - pub fn temporal_consistency_drift( - &self, - version_timestamps: &[i64], // Unix timestamps - version_hashes: &[u64], // Content hashes for each version - ) -> f64 { - if version_timestamps.len() < 2 { - return 0.0; // No drift with single version - } - - let mut issues = 0.0; - let total_checks = (version_timestamps.len() - 1) as f64; - - // Check for timestamp ordering issues - for window in version_timestamps.windows(2) { - if window[1] < window[0] { - issues += 1.0; // Timestamp goes backwards - } - } - - // Check for duplicate content (same hash, different version) - let mut hash_counts = std::collections::HashMap::new(); - for hash in version_hashes { - *hash_counts.entry(hash).or_insert(0) += 1; - } - let duplicates = hash_counts.values().filter(|&&c| c > 1).count(); - issues += duplicates as f64 * 0.5; // Partial penalty for duplicates - - // Check for suspiciously large time gaps (might indicate data loss) - if version_timestamps.len() >= 2 { - let mut deltas: Vec<i64> = version_timestamps - .windows(2) - .map(|w| w[1] - w[0]) - .collect(); - deltas.sort(); - if deltas.len() >= 3 { - let median = deltas[deltas.len() / 2]; - let max = *deltas.last().unwrap_or(&0); - if median > 0 && max > median * 10 { - issues += 0.5; // Large gap detected - } - } - } - - let drift_score = issues / (total_checks + 1.0); - drift_score.clamp(0.0, 1.0) - } - - /// Calculate tensor drift score - /// - /// Measures if tensor representations are consistent with expected properties. - /// - /// Returns a score from 0.0 (no drift) to 1.0 (maximum drift) - pub fn tensor_drift( - &self, - tensor_data: &[f64], - expected_shape: &[usize], - actual_shape: &[usize], - expected_stats: Option<TensorStats>, - ) -> f64 { - let mut drift_score = 0.0; - - // Check shape consistency - if expected_shape != actual_shape { - let shape_diff = expected_shape - .iter() - .zip(actual_shape.iter()) - .map(|(e, a)| (*e as f64 - *a as f64).abs() / *e as f64) - .sum::<f64>() - / expected_shape.len().max(actual_shape.len()) as f64; - drift_score += shape_diff * 0.5; - } - - // Check statistical properties if expected stats provided - if let Some(expected) = expected_stats { - let actual = TensorStats::compute(tensor_data); - - // Compare means - if expected.mean.abs() > 1e-10 { - let mean_diff = (actual.mean - expected.mean).abs() / expected.mean.abs(); - drift_score += mean_diff.min(1.0) * 0.2; - } - - // Compare std deviations - if expected.std_dev > 1e-10 { - let std_diff = (actual.std_dev - expected.std_dev).abs() / expected.std_dev; - drift_score += std_diff.min(1.0) * 0.2; - } - - // Check for NaN/Inf values (always bad) - if actual.has_nan || actual.has_inf { - drift_score += 0.3; - } - } - - drift_score.clamp(0.0, 1.0) - } - - /// Calculate provenance drift score - /// - /// Measures integrity of the provenance chain. A broken chain (hash - /// mismatch, missing parent) produces maximum drift. Staleness (long - /// gap since last event) produces moderate drift. - /// - /// Returns a score from 0.0 (no drift) to 1.0 (maximum drift) - pub fn provenance_drift( - &self, - chain_valid: bool, - chain_length: usize, - seconds_since_last_event: Option<u64>, - expected_max_gap_seconds: u64, - ) -> f64 { - let mut drift_score = 0.0; - - // Chain integrity is critical — broken chain = maximum provenance drift - if !chain_valid { - return 1.0; - } - - // Empty chain for an existing entity is suspicious - if chain_length == 0 { - drift_score += 0.5; - } - - // Staleness: if the last event is much older than expected, lineage - // tracking may have stopped - if let Some(gap) = seconds_since_last_event { - if expected_max_gap_seconds > 0 && gap > expected_max_gap_seconds { - let staleness = (gap as f64 / expected_max_gap_seconds as f64 - 1.0).min(1.0); - drift_score += staleness * 0.3; - } - } - - drift_score.clamp(0.0, 1.0) - } - - /// Calculate spatial drift score - /// - /// Measures consistency between spatial coordinates and location - /// references in the document/graph modalities. If an entity claims - /// to be "in London" but has coordinates pointing to New York, that's - /// spatial drift. - /// - /// Returns a score from 0.0 (no drift) to 1.0 (maximum drift) - pub fn spatial_drift( - &self, - has_coordinates: bool, - has_location_mentions: bool, - coordinate_matches_mentions: bool, - coordinates_valid: bool, - ) -> f64 { - let mut drift_score: f64 = 0.0; - - // Invalid coordinates (NaN, out of WGS84 range) = high drift - if has_coordinates && !coordinates_valid { - return 0.9; - } - - // Document mentions locations but no spatial data indexed - if has_location_mentions && !has_coordinates { - drift_score += 0.4; - } - - // Has coordinates but they don't match document location mentions - if has_coordinates && has_location_mentions && !coordinate_matches_mentions { - drift_score += 0.5; - } - - // Has coordinates but no location context in other modalities - if has_coordinates && !has_location_mentions { - drift_score += 0.1; // Minor — coordinates without context is OK - } - - drift_score.clamp(0.0, 1.0) - } - - /// Calculate schema drift score - /// - /// Measures if the entity violates cross-modal schema constraints. - /// - /// Returns a score from 0.0 (no drift) to 1.0 (maximum drift) - pub fn schema_drift( - &self, - required_modalities: &[&str], - present_modalities: &[&str], - schema_violations: usize, - total_constraints: usize, - ) -> f64 { - let mut drift_score = 0.0; - - // Check modality coverage - let missing_modalities = required_modalities - .iter() - .filter(|m| !present_modalities.contains(m)) - .count(); - if !required_modalities.is_empty() { - drift_score += (missing_modalities as f64 / required_modalities.len() as f64) * 0.5; - } - - // Check schema constraint violations - if total_constraints > 0 { - drift_score += (schema_violations as f64 / total_constraints as f64) * 0.5; - } - - drift_score.clamp(0.0, 1.0) - } - - /// Calculate overall quality drift score - /// - /// Aggregates all drift metrics into an overall quality score. - /// Updated for octad (8 modalities) with provenance and spatial. - /// - /// Returns a score from 0.0 (no drift) to 1.0 (maximum drift) - pub fn quality_drift( - &self, - semantic_vector: f64, - graph_document: f64, - temporal_consistency: f64, - tensor: f64, - schema: f64, - ) -> f64 { - // Weighted average of all drift scores - // Weights reflect relative importance - let weights = [ - (semantic_vector, 0.25), - (graph_document, 0.25), - (temporal_consistency, 0.20), - (tensor, 0.15), - (schema, 0.15), - ]; - - let weighted_sum: f64 = weights.iter().map(|(score, weight)| score * weight).sum(); - weighted_sum.clamp(0.0, 1.0) - } - - /// Calculate overall quality drift score including all 8 modality drift types - pub fn quality_drift_octad( - &self, - semantic_vector: f64, - graph_document: f64, - temporal_consistency: f64, - tensor: f64, - schema: f64, - provenance: f64, - spatial: f64, - ) -> f64 { - // Rebalanced weights for octad (8 modalities) - let weights = [ - (semantic_vector, 0.20), - (graph_document, 0.20), - (temporal_consistency, 0.15), - (tensor, 0.10), - (schema, 0.10), - (provenance, 0.15), - (spatial, 0.10), - ]; - - let weighted_sum: f64 = weights.iter().map(|(score, weight)| score * weight).sum(); - weighted_sum.clamp(0.0, 1.0) - } - - /// Determine drift type from individual scores - pub fn primary_drift_type( - &self, - semantic_vector: f64, - graph_document: f64, - temporal_consistency: f64, - tensor: f64, - schema: f64, - ) -> DriftType { - let scores = [ - (DriftType::SemanticVectorDrift, semantic_vector), - (DriftType::GraphDocumentDrift, graph_document), - (DriftType::TemporalConsistencyDrift, temporal_consistency), - (DriftType::TensorDrift, tensor), - (DriftType::SchemaDrift, schema), - ]; - - scores - .iter() - .max_by(|a, b| a.1.partial_cmp(&b.1).unwrap_or(std::cmp::Ordering::Equal)) - .map(|(t, _)| *t) - .unwrap_or(DriftType::QualityDrift) - } - - /// Determine drift type from all octad scores (including provenance + spatial) - pub fn primary_drift_type_octad( - &self, - semantic_vector: f64, - graph_document: f64, - temporal_consistency: f64, - tensor: f64, - schema: f64, - provenance: f64, - spatial: f64, - ) -> DriftType { - let scores = [ - (DriftType::SemanticVectorDrift, semantic_vector), - (DriftType::GraphDocumentDrift, graph_document), - (DriftType::TemporalConsistencyDrift, temporal_consistency), - (DriftType::TensorDrift, tensor), - (DriftType::SchemaDrift, schema), - (DriftType::ProvenanceDrift, provenance), - (DriftType::SpatialDrift, spatial), - ]; - - scores - .iter() - .max_by(|a, b| a.1.partial_cmp(&b.1).unwrap_or(std::cmp::Ordering::Equal)) - .map(|(t, _)| *t) - .unwrap_or(DriftType::QualityDrift) - } -} - -/// Statistics for tensor data -#[derive(Debug, Clone)] -pub struct TensorStats { - pub mean: f64, - pub std_dev: f64, - pub min: f64, - pub max: f64, - pub has_nan: bool, - pub has_inf: bool, -} - -impl TensorStats { - /// Compute statistics from tensor data - pub fn compute(data: &[f64]) -> Self { - if data.is_empty() { - return Self { - mean: 0.0, - std_dev: 0.0, - min: 0.0, - max: 0.0, - has_nan: false, - has_inf: false, - }; - } - - let has_nan = data.iter().any(|x| x.is_nan()); - let has_inf = data.iter().any(|x| x.is_infinite()); - - // Filter out NaN/Inf for statistics - let valid: Vec<f64> = data.iter().copied().filter(|x| x.is_finite()).collect(); - - if valid.is_empty() { - return Self { - mean: 0.0, - std_dev: 0.0, - min: 0.0, - max: 0.0, - has_nan, - has_inf, - }; - } - - let sum: f64 = valid.iter().sum(); - let mean = sum / valid.len() as f64; - - let variance: f64 = valid.iter().map(|x| (x - mean).powi(2)).sum::<f64>() / valid.len() as f64; - let std_dev = variance.sqrt(); - - let min = valid.iter().copied().fold(f64::INFINITY, f64::min); - let max = valid.iter().copied().fold(f64::NEG_INFINITY, f64::max); - - Self { - mean, - std_dev, - min, - max, - has_nan, - has_inf, - } - } -} - -/// Compute cosine similarity between two f32 vectors -fn cosine_similarity_f32(a: &[f32], b: &[f32]) -> f64 { - if a.len() != b.len() || a.is_empty() { - return 0.0; - } - - let dot: f64 = a.iter().zip(b.iter()).map(|(x, y)| (*x as f64) * (*y as f64)).sum(); - let norm_a: f64 = a.iter().map(|x| (*x as f64).powi(2)).sum::<f64>().sqrt(); - let norm_b: f64 = b.iter().map(|x| (*x as f64).powi(2)).sum::<f64>().sqrt(); - - if norm_a > 0.0 && norm_b > 0.0 { - dot / (norm_a * norm_b) - } else { - 0.0 - } -} - -#[cfg(test)] -mod tests { - use super::*; - - #[test] - fn test_semantic_vector_drift_no_drift() { - let calc = DriftCalculator::default(); - - // Embedding that matches type embedding - let embedding = vec![1.0, 0.0, 0.0]; - let semantic_types = vec!["https://example.org/Person".to_string()]; - let type_embeddings = vec![( - "https://example.org/Person".to_string(), - vec![1.0, 0.0, 0.0], - )]; - - let drift = calc.semantic_vector_drift(&embedding, &semantic_types, &type_embeddings); - assert!(drift < 0.1, "Expected low drift, got {}", drift); - } - - #[test] - fn test_semantic_vector_drift_high_drift() { - let calc = DriftCalculator::default(); - - // Embedding completely different from type embedding - let embedding = vec![1.0, 0.0, 0.0]; - let semantic_types = vec!["https://example.org/Person".to_string()]; - let type_embeddings = vec![( - "https://example.org/Person".to_string(), - vec![0.0, 1.0, 0.0], - )]; - - let drift = calc.semantic_vector_drift(&embedding, &semantic_types, &type_embeddings); - assert!(drift > 0.5, "Expected high drift, got {}", drift); - } - - #[test] - fn test_graph_document_drift() { - let calc = DriftCalculator::default(); - - let document_text = "Alice knows Bob and Charlie"; - let document_entities = vec!["Alice".to_string(), "Bob".to_string(), "Charlie".to_string()]; - let graph_relationships = vec![ - ("knows".to_string(), "Bob".to_string()), - ("knows".to_string(), "Charlie".to_string()), - ]; - - let drift = - calc.graph_document_drift(document_text, &document_entities, &graph_relationships); - // Should have moderate drift since Alice is not a target in relationships - assert!(drift < 0.7, "Drift score: {}", drift); - } - - #[test] - fn test_temporal_consistency_drift() { - let calc = DriftCalculator::default(); - - // Normal sequence - let timestamps = vec![1000, 2000, 3000, 4000]; - let hashes = vec![1, 2, 3, 4]; - let drift = calc.temporal_consistency_drift(×tamps, &hashes); - assert!(drift < 0.1, "Expected low drift for normal sequence, got {}", drift); - - // Out of order timestamps - let timestamps = vec![1000, 3000, 2000, 4000]; - let drift = calc.temporal_consistency_drift(×tamps, &hashes); - assert!(drift > 0.0, "Expected drift for out-of-order timestamps, got {}", drift); - } - - #[test] - fn test_tensor_drift() { - let calc = DriftCalculator::default(); - - let data = vec![1.0, 2.0, 3.0, 4.0]; - let expected_shape = vec![2, 2]; - let actual_shape = vec![2, 2]; - let expected_stats = Some(TensorStats { - mean: 2.5, - std_dev: 1.118, - min: 1.0, - max: 4.0, - has_nan: false, - has_inf: false, - }); - - let drift = calc.tensor_drift(&data, &expected_shape, &actual_shape, expected_stats); - assert!(drift < 0.3, "Expected low drift for matching tensor, got {}", drift); - - // Test with NaN values - let data_with_nan = vec![1.0, f64::NAN, 3.0, 4.0]; - let _drift = calc.tensor_drift(&data_with_nan, &expected_shape, &actual_shape, None); - // Should not panic, just handle gracefully - } - - #[test] - fn test_schema_drift() { - let calc = DriftCalculator::default(); - - let required = vec!["graph", "vector", "document"]; - let present = vec!["graph", "vector"]; - - let drift = calc.schema_drift(&required, &present, 1, 5); - assert!(drift > 0.0, "Expected drift for missing modality"); - assert!(drift < 1.0, "Drift should be bounded"); - } - - #[test] - fn test_quality_drift() { - let calc = DriftCalculator::default(); - - // All metrics good - let drift = calc.quality_drift(0.1, 0.1, 0.1, 0.1, 0.1); - assert!(drift < 0.2, "Expected low overall drift"); - - // Some metrics bad - let drift = calc.quality_drift(0.8, 0.2, 0.1, 0.1, 0.1); - assert!(drift > 0.1, "Expected higher drift with semantic-vector issues"); - } - - #[test] - fn test_tensor_stats() { - let data = vec![1.0, 2.0, 3.0, 4.0, 5.0]; - let stats = TensorStats::compute(&data); - - assert!((stats.mean - 3.0).abs() < 1e-10); - assert!((stats.min - 1.0).abs() < 1e-10); - assert!((stats.max - 5.0).abs() < 1e-10); - assert!(!stats.has_nan); - assert!(!stats.has_inf); - } -} diff --git a/verisimdb/rust-core/verisim-drift/src/lib.rs b/verisimdb/rust-core/verisim-drift/src/lib.rs deleted file mode 100644 index 33041778..00000000 --- a/verisimdb/rust-core/verisim-drift/src/lib.rs +++ /dev/null @@ -1,551 +0,0 @@ -// SPDX-License-Identifier: MPL-2.0 -//! VeriSim Drift Detection -//! -//! Monitors cross-modal consistency degradation and triggers normalization. -//! This is the "early warning system" for data quality issues. - -#![forbid(unsafe_code)] -use chrono::{DateTime, Utc}; -use prometheus::{Counter, Gauge, Registry}; -use serde::{Deserialize, Serialize}; -use std::collections::HashMap; -use std::sync::{Arc, RwLock}; -use thiserror::Error; -use tokio::sync::mpsc; - -// Drift calculation algorithms -mod calculator; -pub use calculator::{DriftCalculator, TensorStats}; - -/// Drift detection errors -#[derive(Error, Debug)] -pub enum DriftError { - #[error("Metric not found: {0}")] - MetricNotFound(String), - - #[error("Invalid threshold: {0}")] - InvalidThreshold(String), - - #[error("Channel error: {0}")] - ChannelError(String), - - #[error("Lock poisoned: internal concurrency error")] - LockPoisoned, -} - -/// Types of drift that can be detected -#[derive(Debug, Clone, Copy, Serialize, Deserialize, PartialEq, Eq, Hash)] -pub enum DriftType { - /// Vector embeddings diverge from semantic meaning - SemanticVectorDrift, - /// Graph structure doesn't match document content - GraphDocumentDrift, - /// Temporal versions become inconsistent - TemporalConsistencyDrift, - /// Tensor representations diverge - TensorDrift, - /// Cross-modal schema violations - SchemaDrift, - /// Provenance chain integrity broken or lineage inconsistent - ProvenanceDrift, - /// Spatial coordinates inconsistent with graph/document location mentions - SpatialDrift, - /// Overall data quality degradation - QualityDrift, -} - -impl std::fmt::Display for DriftType { - fn fmt(&self, f: &mut std::fmt::Formatter<'_>) -> std::fmt::Result { - match self { - DriftType::SemanticVectorDrift => write!(f, "semantic_vector_drift"), - DriftType::GraphDocumentDrift => write!(f, "graph_document_drift"), - DriftType::TemporalConsistencyDrift => write!(f, "temporal_consistency_drift"), - DriftType::TensorDrift => write!(f, "tensor_drift"), - DriftType::SchemaDrift => write!(f, "schema_drift"), - DriftType::ProvenanceDrift => write!(f, "provenance_drift"), - DriftType::SpatialDrift => write!(f, "spatial_drift"), - DriftType::QualityDrift => write!(f, "quality_drift"), - } - } -} - -/// Severity levels for drift alerts -#[derive(Debug, Clone, Copy, Serialize, Deserialize, PartialEq, Eq, PartialOrd, Ord)] -pub enum DriftSeverity { - Info, - Warning, - Critical, - Emergency, -} - -/// A detected drift event -#[derive(Debug, Clone, Serialize, Deserialize)] -pub struct DriftEvent { - /// Type of drift - pub drift_type: DriftType, - /// Severity level - pub severity: DriftSeverity, - /// Affected entity IDs - pub affected_entities: Vec<String>, - /// Drift score (0.0 - 1.0, higher = worse) - pub score: f64, - /// When detected - pub detected_at: DateTime<Utc>, - /// Description - pub description: String, - /// Suggested remediation - pub remediation: Option<String>, -} - -impl DriftEvent { - /// Create a new drift event - pub fn new(drift_type: DriftType, score: f64, description: impl Into<String>) -> Self { - let severity = if score > 0.9 { - DriftSeverity::Emergency - } else if score > 0.7 { - DriftSeverity::Critical - } else if score > 0.5 { - DriftSeverity::Warning - } else { - DriftSeverity::Info - }; - - Self { - drift_type, - severity, - affected_entities: Vec::new(), - score, - detected_at: Utc::now(), - description: description.into(), - remediation: None, - } - } - - /// Add affected entities - pub fn with_entities(mut self, entities: Vec<String>) -> Self { - self.affected_entities = entities; - self - } - - /// Add remediation suggestion - pub fn with_remediation(mut self, remediation: impl Into<String>) -> Self { - self.remediation = Some(remediation.into()); - self - } -} - -/// Threshold policy for drift detection -#[derive(Debug, Clone, Serialize, Deserialize)] -pub enum ThresholdPolicy { - /// Fixed threshold value - Fixed(f64), - /// Adaptive threshold: base + (moving_avg * sensitivity) - Adaptive { base: f64, sensitivity: f64 }, -} - -impl ThresholdPolicy { - /// Compute the effective threshold given the current moving average - pub fn effective_threshold(&self, moving_average: f64) -> f64 { - match self { - ThresholdPolicy::Fixed(v) => *v, - ThresholdPolicy::Adaptive { base, sensitivity } => { - base + (moving_average * sensitivity) - } - } - } -} - -/// Threshold configuration for drift detection -#[derive(Debug, Clone, Serialize, Deserialize)] -pub struct DriftThresholds { - /// Threshold for semantic-vector drift - pub semantic_vector: f64, - /// Threshold for graph-document drift - pub graph_document: f64, - /// Threshold for temporal consistency drift - pub temporal_consistency: f64, - /// Threshold for tensor drift - pub tensor: f64, - /// Threshold for schema drift - pub schema: f64, - /// Threshold for provenance chain integrity drift - pub provenance: f64, - /// Threshold for spatial consistency drift - pub spatial: f64, - /// Threshold for overall quality drift - pub quality: f64, - /// Optional adaptive policies per drift type (overrides fixed thresholds) - #[serde(default)] - pub adaptive_policies: HashMap<DriftType, ThresholdPolicy>, -} - -impl Default for DriftThresholds { - fn default() -> Self { - Self { - semantic_vector: 0.3, - graph_document: 0.4, - temporal_consistency: 0.2, - tensor: 0.35, - schema: 0.1, - provenance: 0.1, - spatial: 0.3, - quality: 0.25, - adaptive_policies: HashMap::new(), - } - } -} - -impl DriftThresholds { - /// Get the effective threshold for a drift type, considering adaptive policies - pub fn effective_threshold(&self, drift_type: DriftType, moving_average: f64) -> f64 { - if let Some(policy) = self.adaptive_policies.get(&drift_type) { - return policy.effective_threshold(moving_average); - } - // Fall back to fixed thresholds - match drift_type { - DriftType::SemanticVectorDrift => self.semantic_vector, - DriftType::GraphDocumentDrift => self.graph_document, - DriftType::TemporalConsistencyDrift => self.temporal_consistency, - DriftType::TensorDrift => self.tensor, - DriftType::SchemaDrift => self.schema, - DriftType::ProvenanceDrift => self.provenance, - DriftType::SpatialDrift => self.spatial, - DriftType::QualityDrift => self.quality, - } - } -} - -/// Metrics for a specific drift type -#[derive(Debug, Clone, Serialize, Deserialize)] -pub struct DriftMetrics { - /// Current drift score - pub current_score: f64, - /// Moving average - pub moving_average: f64, - /// Maximum observed score - pub max_score: f64, - /// Number of measurements - pub measurement_count: u64, - /// Last measurement time - pub last_measured: DateTime<Utc>, - /// Historical scores (last N measurements) - pub history: Vec<(DateTime<Utc>, f64)>, -} - -impl Default for DriftMetrics { - fn default() -> Self { - Self { - current_score: 0.0, - moving_average: 0.0, - max_score: 0.0, - measurement_count: 0, - last_measured: Utc::now(), - history: Vec::new(), - } - } -} - -impl DriftMetrics { - /// Record a new measurement - pub fn record(&mut self, score: f64) { - self.current_score = score; - self.measurement_count += 1; - self.last_measured = Utc::now(); - - if score > self.max_score { - self.max_score = score; - } - - // Update moving average (exponential) - let alpha = 0.1; - self.moving_average = alpha * score + (1.0 - alpha) * self.moving_average; - - // Keep last 100 measurements - self.history.push((Utc::now(), score)); - if self.history.len() > 100 { - self.history.remove(0); - } - } - - /// Check if score exceeds threshold - pub fn exceeds_threshold(&self, threshold: f64) -> bool { - self.current_score > threshold - } - - /// Get trend (positive = increasing drift) - pub fn trend(&self) -> f64 { - if self.history.len() < 2 { - return 0.0; - } - - let recent: Vec<_> = self.history.iter().rev().take(10).collect(); - let older: Vec<_> = self.history.iter().rev().skip(10).take(10).collect(); - - if older.is_empty() { - return 0.0; - } - - let recent_avg: f64 = recent.iter().map(|(_, s)| s).sum::<f64>() / recent.len() as f64; - let older_avg: f64 = older.iter().map(|(_, s)| s).sum::<f64>() / older.len() as f64; - - recent_avg - older_avg - } -} - -/// Drift detector - monitors and reports drift events -pub struct DriftDetector { - thresholds: DriftThresholds, - metrics: Arc<RwLock<HashMap<DriftType, DriftMetrics>>>, - event_sender: Option<mpsc::Sender<DriftEvent>>, - prometheus_registry: Option<Registry>, - // Prometheus metrics - drift_score_gauge: Option<HashMap<DriftType, Gauge>>, - drift_event_counter: Option<HashMap<DriftType, Counter>>, -} - -impl DriftDetector { - /// Create a new drift detector - pub fn new(thresholds: DriftThresholds) -> Self { - let mut metrics = HashMap::new(); - for drift_type in [ - DriftType::SemanticVectorDrift, - DriftType::GraphDocumentDrift, - DriftType::TemporalConsistencyDrift, - DriftType::TensorDrift, - DriftType::SchemaDrift, - DriftType::ProvenanceDrift, - DriftType::SpatialDrift, - DriftType::QualityDrift, - ] { - metrics.insert(drift_type, DriftMetrics::default()); - } - - Self { - thresholds, - metrics: Arc::new(RwLock::new(metrics)), - event_sender: None, - prometheus_registry: None, - drift_score_gauge: None, - drift_event_counter: None, - } - } - - /// Create with default thresholds - pub fn with_defaults() -> Self { - Self::new(DriftThresholds::default()) - } - - /// Set event channel for drift notifications - pub fn with_event_channel(mut self, sender: mpsc::Sender<DriftEvent>) -> Self { - self.event_sender = Some(sender); - self - } - - /// Register Prometheus metrics - pub fn with_prometheus(mut self, registry: Registry) -> Result<Self, DriftError> { - let mut gauges = HashMap::new(); - let mut counters = HashMap::new(); - - for drift_type in [ - DriftType::SemanticVectorDrift, - DriftType::GraphDocumentDrift, - DriftType::TemporalConsistencyDrift, - DriftType::TensorDrift, - DriftType::SchemaDrift, - DriftType::ProvenanceDrift, - DriftType::SpatialDrift, - DriftType::QualityDrift, - ] { - let gauge = Gauge::new( - format!("verisim_drift_score_{}", drift_type), - format!("Current drift score for {}", drift_type), - ) - .map_err(|e| DriftError::InvalidThreshold(e.to_string()))?; - registry - .register(Box::new(gauge.clone())) - .map_err(|e| DriftError::InvalidThreshold(e.to_string()))?; - gauges.insert(drift_type, gauge); - - let counter = Counter::new( - format!("verisim_drift_events_{}", drift_type), - format!("Number of drift events for {}", drift_type), - ) - .map_err(|e| DriftError::InvalidThreshold(e.to_string()))?; - registry - .register(Box::new(counter.clone())) - .map_err(|e| DriftError::InvalidThreshold(e.to_string()))?; - counters.insert(drift_type, counter); - } - - self.prometheus_registry = Some(registry); - self.drift_score_gauge = Some(gauges); - self.drift_event_counter = Some(counters); - Ok(self) - } - - /// Record a drift measurement - pub async fn record(&self, drift_type: DriftType, score: f64, entities: Vec<String>) -> Result<Option<DriftEvent>, DriftError> { - // Update metrics - { - let mut metrics = self.metrics.write().map_err(|_| DriftError::LockPoisoned)?; - if let Some(m) = metrics.get_mut(&drift_type) { - m.record(score); - } - } - - // Update Prometheus gauge - if let Some(ref gauges) = self.drift_score_gauge { - if let Some(gauge) = gauges.get(&drift_type) { - gauge.set(score); - } - } - - // Check threshold (adaptive or fixed) - let moving_avg = { - let metrics = self.metrics.read().map_err(|_| DriftError::LockPoisoned)?; - metrics - .get(&drift_type) - .map(|m| m.moving_average) - .unwrap_or(0.0) - }; - let threshold = self.thresholds.effective_threshold(drift_type, moving_avg); - - if score > threshold { - let event = DriftEvent::new( - drift_type, - score, - format!( - "{} detected with score {:.3} (threshold: {:.3})", - drift_type, score, threshold - ), - ) - .with_entities(entities); - - // Update Prometheus counter - if let Some(ref counters) = self.drift_event_counter { - if let Some(counter) = counters.get(&drift_type) { - counter.inc(); - } - } - - // Send event notification - if let Some(ref sender) = self.event_sender { - sender - .send(event.clone()) - .await - .map_err(|e| DriftError::ChannelError(e.to_string()))?; - } - - return Ok(Some(event)); - } - - Ok(None) - } - - /// Get current metrics for a drift type - pub fn get_metrics(&self, drift_type: DriftType) -> Result<Option<DriftMetrics>, DriftError> { - let metrics = self.metrics.read().map_err(|_| DriftError::LockPoisoned)?; - Ok(metrics.get(&drift_type).cloned()) - } - - /// Get all metrics - pub fn all_metrics(&self) -> Result<HashMap<DriftType, DriftMetrics>, DriftError> { - let metrics = self.metrics.read().map_err(|_| DriftError::LockPoisoned)?; - Ok(metrics.clone()) - } - - /// Check overall health - pub fn health_check(&self) -> Result<DriftHealthStatus, DriftError> { - let metrics = self.metrics.read().map_err(|_| DriftError::LockPoisoned)?; - let mut worst_score = 0.0; - let mut worst_type = DriftType::QualityDrift; - - for (drift_type, m) in metrics.iter() { - if m.current_score > worst_score { - worst_score = m.current_score; - worst_type = *drift_type; - } - } - - let status = if worst_score > 0.9 { - HealthStatus::Critical - } else if worst_score > 0.7 { - HealthStatus::Degraded - } else if worst_score > 0.5 { - HealthStatus::Warning - } else { - HealthStatus::Healthy - }; - - Ok(DriftHealthStatus { - status, - worst_drift_type: worst_type, - worst_score, - checked_at: Utc::now(), - }) - } -} - -/// Overall health status -#[derive(Debug, Clone, Copy, Serialize, Deserialize, PartialEq, Eq)] -pub enum HealthStatus { - Healthy, - Warning, - Degraded, - Critical, -} - -/// Drift health status report -#[derive(Debug, Clone, Serialize, Deserialize)] -pub struct DriftHealthStatus { - pub status: HealthStatus, - pub worst_drift_type: DriftType, - pub worst_score: f64, - pub checked_at: DateTime<Utc>, -} - -#[cfg(test)] -mod tests { - use super::*; - - #[tokio::test] - async fn test_drift_detection() { - let detector = DriftDetector::with_defaults(); - - // Record normal score - let event = detector - .record(DriftType::SemanticVectorDrift, 0.1, vec![]) - .await - .expect("TODO: handle error"); - assert!(event.is_none()); - - // Record high score (above threshold of 0.3 for semantic_vector) - // Score 0.6 triggers Warning severity (> 0.5) - let event = detector - .record(DriftType::SemanticVectorDrift, 0.6, vec!["entity1".to_string()]) - .await - .expect("TODO: handle error"); - assert!(event.is_some()); - assert_eq!(event.expect("TODO: handle error").severity, DriftSeverity::Warning); - } - - #[test] - fn test_drift_metrics() { - let mut metrics = DriftMetrics::default(); - - metrics.record(0.1); - metrics.record(0.2); - metrics.record(0.3); - - assert_eq!(metrics.current_score, 0.3); - assert_eq!(metrics.max_score, 0.3); - assert_eq!(metrics.measurement_count, 3); - } - - #[test] - fn test_health_check() { - let detector = DriftDetector::with_defaults(); - let status = detector.health_check().expect("TODO: handle error"); - assert_eq!(status.status, HealthStatus::Healthy); - } -} diff --git a/verisimdb/rust-core/verisim-graph/Cargo.toml b/verisimdb/rust-core/verisim-graph/Cargo.toml deleted file mode 100644 index 652e2d52..00000000 --- a/verisimdb/rust-core/verisim-graph/Cargo.toml +++ /dev/null @@ -1,34 +0,0 @@ -# SPDX-License-Identifier: MPL-2.0 - -[package] -name = "verisim-graph" -description = "Graph modality - RDF and property graph storage (pure Rust default, Oxigraph optional)" -version.workspace = true -edition.workspace = true -authors.workspace = true -license.workspace = true - -[features] -default = [] -# Enable the Oxigraph backend for full RDF/SPARQL support. -# Requires a C++ linker (oxigraph → oxrocksdb-sys transitive dependency). -oxigraph-backend = ["dep:oxigraph"] -# Enable the redb backend for persistent graph storage (pure Rust, no C/C++). -redb-backend = ["dep:redb", "dep:serde_json"] - -[dependencies] -serde.workspace = true -thiserror.workspace = true -async-trait.workspace = true -tokio.workspace = true - -# Optional: Oxigraph for full RDF/SPARQL (pulls in C++ via oxrocksdb-sys) -oxigraph = { workspace = true, optional = true } - -# Optional: redb for pure-Rust persistent graph storage (B-tree, ACID) -redb = { workspace = true, optional = true } -serde_json = { workspace = true, optional = true } - -[dev-dependencies] -proptest.workspace = true -tempfile = "3" diff --git a/verisimdb/rust-core/verisim-graph/src/lib.rs b/verisimdb/rust-core/verisim-graph/src/lib.rs deleted file mode 100644 index 387620f0..00000000 --- a/verisimdb/rust-core/verisim-graph/src/lib.rs +++ /dev/null @@ -1,427 +0,0 @@ -// SPDX-License-Identifier: MPL-2.0 -//! VeriSim Graph Modality -//! -//! RDF and property graph storage. Provides two backends: -//! -//! - **`SimpleGraphStore`** (default) — Pure Rust, in-memory HashMap/BTreeMap -//! store. Zero C/C++ dependencies, builds on any platform without a C++ linker. -//! -//! - **`OxiGraphStore`** (feature: `oxigraph-backend`) — Full Oxigraph RDF store -//! with SPARQL support. Requires C++ linker for the RocksDB transitive dependency. -//! -//! Implements Marr's Computational Level: "What relationships exist?" - -#![forbid(unsafe_code)] -use async_trait::async_trait; -use serde::{Deserialize, Serialize}; -use std::collections::{HashMap, HashSet}; -use std::sync::RwLock; -use thiserror::Error; - -// Re-export Oxigraph backend when feature is enabled -#[cfg(feature = "oxigraph-backend")] -mod oxigraph_backend; -#[cfg(feature = "oxigraph-backend")] -pub use oxigraph_backend::OxiGraphStore; - -// Re-export redb backend when feature is enabled -#[cfg(feature = "redb-backend")] -mod redb_backend; -#[cfg(feature = "redb-backend")] -pub use redb_backend::RedbGraphStore; - -/// Graph modality errors -#[derive(Error, Debug)] -pub enum GraphError { - #[error("Store error: {0}")] - StoreError(String), - - #[error("Parse error: {0}")] - ParseError(String), - - #[error("Entity not found: {0}")] - NotFound(String), - - #[error("Invalid IRI: {0}")] - InvalidIri(String), - - #[error("Lock poisoned")] - LockPoisoned, -} - -/// A node in the graph (entity reference) -#[derive(Debug, Clone, Serialize, Deserialize, PartialEq, Eq, Hash)] -pub struct GraphNode { - /// IRI of the node - pub iri: String, - /// Local name (last segment of IRI) - pub local_name: String, -} - -impl GraphNode { - /// Create a new graph node from an IRI - pub fn new(iri: impl Into<String>) -> Self { - let iri = iri.into(); - let local_name = iri - .rsplit_once('/') - .or_else(|| iri.rsplit_once('#')) - .map(|(_, name)| name.to_string()) - .unwrap_or_else(|| iri.clone()); - Self { iri, local_name } - } -} - -/// An edge in the graph (relationship) -#[derive(Debug, Clone, Serialize, Deserialize)] -pub struct GraphEdge { - /// Subject node - pub subject: GraphNode, - /// Predicate (relationship type) - pub predicate: GraphNode, - /// Object (target node or literal) - pub object: GraphObject, -} - -/// Object of a triple (can be node or literal) -#[derive(Debug, Clone, Serialize, Deserialize)] -pub enum GraphObject { - Node(GraphNode), - Literal { value: String, datatype: Option<String> }, -} - -/// Graph store trait for cross-modal consistency. -/// -/// All graph backends implement this trait, allowing the octad store to be -/// generic over the concrete backend. -#[async_trait] -pub trait GraphStore: Send + Sync { - /// Insert a triple - async fn insert(&self, edge: &GraphEdge) -> Result<(), GraphError>; - - /// Query outgoing edges from a node - async fn outgoing(&self, node: &GraphNode) -> Result<Vec<GraphEdge>, GraphError>; - - /// Query incoming edges to a node - async fn incoming(&self, node: &GraphNode) -> Result<Vec<GraphEdge>, GraphError>; - - /// Check if a triple exists - async fn exists(&self, edge: &GraphEdge) -> Result<bool, GraphError>; - - /// Delete a triple - async fn delete(&self, edge: &GraphEdge) -> Result<(), GraphError>; - - /// Get all nodes connected to a given node within N hops - async fn neighborhood(&self, node: &GraphNode, hops: usize) -> Result<Vec<GraphNode>, GraphError>; -} - -// ═══════════════════════════════════════════════════════════════════════════ -// SimpleGraphStore — Pure Rust in-memory graph store -// ═══════════════════════════════════════════════════════════════════════════ - -/// A canonicalised triple key for deduplication and lookup. -/// -/// Stores `(subject_iri, predicate_iri, object_key)` where `object_key` is -/// either the IRI for nodes or `"literal::<value>"` for literals. -#[derive(Debug, Clone, PartialEq, Eq, Hash)] -struct TripleKey(String, String, String); - -impl TripleKey { - fn from_edge(edge: &GraphEdge) -> Self { - let obj_key = match &edge.object { - GraphObject::Node(n) => n.iri.clone(), - GraphObject::Literal { value, .. } => format!("literal::{}", value), - }; - Self(edge.subject.iri.clone(), edge.predicate.iri.clone(), obj_key) - } -} - -/// Pure Rust in-memory graph store. -/// -/// Uses HashMap indices for O(1) subject/object lookups and a HashSet for -/// deduplication. No external dependencies — builds on any platform. -/// -/// Thread-safe via `RwLock` — concurrent reads, exclusive writes. -pub struct SimpleGraphStore { - /// All edges stored as a set of triple keys → edge data - edges: RwLock<HashMap<TripleKey, GraphEdge>>, - /// Subject index: subject IRI → set of triple keys - subject_idx: RwLock<HashMap<String, HashSet<TripleKey>>>, - /// Object index: object IRI → set of triple keys (nodes only) - object_idx: RwLock<HashMap<String, HashSet<TripleKey>>>, -} - -impl SimpleGraphStore { - /// Create a new empty in-memory graph store. - pub fn new() -> Self { - Self { - edges: RwLock::new(HashMap::new()), - subject_idx: RwLock::new(HashMap::new()), - object_idx: RwLock::new(HashMap::new()), - } - } - - /// Create a new store — mirrors OxiGraphStore::in_memory() API for drop-in - /// replacement. - pub fn in_memory() -> Result<Self, GraphError> { - Ok(Self::new()) - } -} - -impl Default for SimpleGraphStore { - fn default() -> Self { - Self::new() - } -} - -#[async_trait] -impl GraphStore for SimpleGraphStore { - async fn insert(&self, edge: &GraphEdge) -> Result<(), GraphError> { - let key = TripleKey::from_edge(edge); - - // Update subject index - self.subject_idx - .write() - .map_err(|_| GraphError::LockPoisoned)? - .entry(edge.subject.iri.clone()) - .or_default() - .insert(key.clone()); - - // Update object index (nodes only) - if let GraphObject::Node(n) = &edge.object { - self.object_idx - .write() - .map_err(|_| GraphError::LockPoisoned)? - .entry(n.iri.clone()) - .or_default() - .insert(key.clone()); - } - - // Insert the edge - self.edges - .write() - .map_err(|_| GraphError::LockPoisoned)? - .insert(key, edge.clone()); - - Ok(()) - } - - async fn outgoing(&self, node: &GraphNode) -> Result<Vec<GraphEdge>, GraphError> { - let subject_idx = self.subject_idx.read().map_err(|_| GraphError::LockPoisoned)?; - let edges = self.edges.read().map_err(|_| GraphError::LockPoisoned)?; - - let result = match subject_idx.get(&node.iri) { - Some(keys) => keys - .iter() - .filter_map(|k| edges.get(k).cloned()) - .collect(), - None => Vec::new(), - }; - - Ok(result) - } - - async fn incoming(&self, node: &GraphNode) -> Result<Vec<GraphEdge>, GraphError> { - let object_idx = self.object_idx.read().map_err(|_| GraphError::LockPoisoned)?; - let edges = self.edges.read().map_err(|_| GraphError::LockPoisoned)?; - - let result = match object_idx.get(&node.iri) { - Some(keys) => keys - .iter() - .filter_map(|k| edges.get(k).cloned()) - .collect(), - None => Vec::new(), - }; - - Ok(result) - } - - async fn exists(&self, edge: &GraphEdge) -> Result<bool, GraphError> { - let key = TripleKey::from_edge(edge); - let edges = self.edges.read().map_err(|_| GraphError::LockPoisoned)?; - Ok(edges.contains_key(&key)) - } - - async fn delete(&self, edge: &GraphEdge) -> Result<(), GraphError> { - let key = TripleKey::from_edge(edge); - - // Remove from subject index - if let Ok(mut idx) = self.subject_idx.write() { - if let Some(keys) = idx.get_mut(&edge.subject.iri) { - keys.remove(&key); - if keys.is_empty() { - idx.remove(&edge.subject.iri); - } - } - } - - // Remove from object index - if let GraphObject::Node(n) = &edge.object { - if let Ok(mut idx) = self.object_idx.write() { - if let Some(keys) = idx.get_mut(&n.iri) { - keys.remove(&key); - if keys.is_empty() { - idx.remove(&n.iri); - } - } - } - } - - // Remove the edge - self.edges - .write() - .map_err(|_| GraphError::LockPoisoned)? - .remove(&key); - - Ok(()) - } - - async fn neighborhood(&self, node: &GraphNode, hops: usize) -> Result<Vec<GraphNode>, GraphError> { - let mut visited = HashSet::new(); - let mut frontier = vec![node.clone()]; - visited.insert(node.iri.clone()); - - for _ in 0..hops { - let mut next_frontier = Vec::new(); - for current in frontier { - for edge in self.outgoing(¤t).await? { - if let GraphObject::Node(n) = edge.object { - if visited.insert(n.iri.clone()) { - next_frontier.push(n); - } - } - } - for edge in self.incoming(¤t).await? { - if visited.insert(edge.subject.iri.clone()) { - next_frontier.push(edge.subject); - } - } - } - frontier = next_frontier; - } - - Ok(visited.into_iter().map(GraphNode::new).collect()) - } -} - -// ═══════════════════════════════════════════════════════════════════════════ -// Tests -// ═══════════════════════════════════════════════════════════════════════════ - -#[cfg(test)] -mod tests { - use super::*; - - #[tokio::test] - async fn test_insert_and_query() { - let store = SimpleGraphStore::in_memory().expect("TODO: handle error"); - let edge = GraphEdge { - subject: GraphNode::new("https://example.org/Alice"), - predicate: GraphNode::new("https://example.org/knows"), - object: GraphObject::Node(GraphNode::new("https://example.org/Bob")), - }; - - store.insert(&edge).await.expect("TODO: handle error"); - assert!(store.exists(&edge).await.expect("TODO: handle error")); - - let outgoing = store.outgoing(&edge.subject).await.expect("TODO: handle error"); - assert_eq!(outgoing.len(), 1); - } - - #[tokio::test] - async fn test_incoming_edges() { - let store = SimpleGraphStore::new(); - let edge = GraphEdge { - subject: GraphNode::new("https://example.org/Alice"), - predicate: GraphNode::new("https://example.org/knows"), - object: GraphObject::Node(GraphNode::new("https://example.org/Bob")), - }; - - store.insert(&edge).await.expect("TODO: handle error"); - - let bob = GraphNode::new("https://example.org/Bob"); - let incoming = store.incoming(&bob).await.expect("TODO: handle error"); - assert_eq!(incoming.len(), 1); - } - - #[tokio::test] - async fn test_delete_edge() { - let store = SimpleGraphStore::new(); - let edge = GraphEdge { - subject: GraphNode::new("https://example.org/Alice"), - predicate: GraphNode::new("https://example.org/knows"), - object: GraphObject::Node(GraphNode::new("https://example.org/Bob")), - }; - - store.insert(&edge).await.expect("TODO: handle error"); - assert!(store.exists(&edge).await.expect("TODO: handle error")); - - store.delete(&edge).await.expect("TODO: handle error"); - assert!(!store.exists(&edge).await.expect("TODO: handle error")); - } - - #[tokio::test] - async fn test_literal_object() { - let store = SimpleGraphStore::new(); - let edge = GraphEdge { - subject: GraphNode::new("https://example.org/Alice"), - predicate: GraphNode::new("https://example.org/name"), - object: GraphObject::Literal { - value: "Alice".to_string(), - datatype: Some("http://www.w3.org/2001/XMLSchema#string".to_string()), - }, - }; - - store.insert(&edge).await.expect("TODO: handle error"); - assert!(store.exists(&edge).await.expect("TODO: handle error")); - - let outgoing = store.outgoing(&edge.subject).await.expect("TODO: handle error"); - assert_eq!(outgoing.len(), 1); - } - - #[tokio::test] - async fn test_neighborhood() { - let store = SimpleGraphStore::new(); - - // Alice -> Bob -> Carol - let e1 = GraphEdge { - subject: GraphNode::new("https://example.org/Alice"), - predicate: GraphNode::new("https://example.org/knows"), - object: GraphObject::Node(GraphNode::new("https://example.org/Bob")), - }; - let e2 = GraphEdge { - subject: GraphNode::new("https://example.org/Bob"), - predicate: GraphNode::new("https://example.org/knows"), - object: GraphObject::Node(GraphNode::new("https://example.org/Carol")), - }; - - store.insert(&e1).await.expect("TODO: handle error"); - store.insert(&e2).await.expect("TODO: handle error"); - - let alice = GraphNode::new("https://example.org/Alice"); - - // 1 hop: Alice, Bob - let neighbors_1 = store.neighborhood(&alice, 1).await.expect("TODO: handle error"); - assert_eq!(neighbors_1.len(), 2); - - // 2 hops: Alice, Bob, Carol - let neighbors_2 = store.neighborhood(&alice, 2).await.expect("TODO: handle error"); - assert_eq!(neighbors_2.len(), 3); - } - - #[tokio::test] - async fn test_deduplication() { - let store = SimpleGraphStore::new(); - let edge = GraphEdge { - subject: GraphNode::new("https://example.org/Alice"), - predicate: GraphNode::new("https://example.org/knows"), - object: GraphObject::Node(GraphNode::new("https://example.org/Bob")), - }; - - // Insert the same edge twice - store.insert(&edge).await.expect("TODO: handle error"); - store.insert(&edge).await.expect("TODO: handle error"); - - let outgoing = store.outgoing(&edge.subject).await.expect("TODO: handle error"); - assert_eq!(outgoing.len(), 1, "Duplicate edges should be deduplicated"); - } -} diff --git a/verisimdb/rust-core/verisim-graph/src/oxigraph_backend.rs b/verisimdb/rust-core/verisim-graph/src/oxigraph_backend.rs deleted file mode 100644 index 31dcd20e..00000000 --- a/verisimdb/rust-core/verisim-graph/src/oxigraph_backend.rs +++ /dev/null @@ -1,170 +0,0 @@ -// SPDX-License-Identifier: MPL-2.0 -//! Oxigraph-backed graph store. -//! -//! This module is only compiled when the `oxigraph-backend` feature is enabled. -//! It provides full RDF/SPARQL support via Oxigraph, but requires a C++ linker -//! for the transitive oxrocksdb-sys dependency. - -use async_trait::async_trait; -use oxigraph::model::{GraphName, NamedNode, Quad, Subject, Term}; -use oxigraph::store::Store; -use std::path::Path; - -use crate::{GraphEdge, GraphError, GraphNode, GraphObject, GraphStore}; - -/// Oxigraph-backed graph store with full RDF/SPARQL support. -/// -/// Requires the `oxigraph-backend` feature flag. Supports both in-memory and -/// persistent (RocksDB-backed) storage. -pub struct OxiGraphStore { - store: Store, -} - -impl OxiGraphStore { - /// Create a new in-memory store - pub fn in_memory() -> Result<Self, GraphError> { - Ok(Self { - store: Store::new().map_err(|e| GraphError::StoreError(e.to_string()))?, - }) - } - - /// Open or create a persistent store (requires filesystem + C++ linker at build time) - pub fn persistent(path: impl AsRef<Path>) -> Result<Self, GraphError> { - Ok(Self { - store: Store::open(path).map_err(|e| GraphError::StoreError(e.to_string()))?, - }) - } - - /// Convert GraphEdge to Oxigraph Quad - fn edge_to_quad(&self, edge: &GraphEdge) -> Result<Quad, GraphError> { - let subject = Subject::NamedNode( - NamedNode::new(&edge.subject.iri) - .map_err(|e| GraphError::InvalidIri(e.to_string()))?, - ); - let predicate = NamedNode::new(&edge.predicate.iri) - .map_err(|e| GraphError::InvalidIri(e.to_string()))?; - let object = match &edge.object { - GraphObject::Node(n) => Term::NamedNode( - NamedNode::new(&n.iri).map_err(|e| GraphError::InvalidIri(e.to_string()))?, - ), - GraphObject::Literal { value, datatype: _ } => { - Term::Literal(oxigraph::model::Literal::new_simple_literal(value)) - } - }; - Ok(Quad::new(subject, predicate, object, GraphName::DefaultGraph)) - } -} - -#[async_trait] -impl GraphStore for OxiGraphStore { - async fn insert(&self, edge: &GraphEdge) -> Result<(), GraphError> { - let quad = self.edge_to_quad(edge)?; - self.store.insert(&quad).map_err(|e| GraphError::StoreError(e.to_string()))?; - Ok(()) - } - - async fn outgoing(&self, node: &GraphNode) -> Result<Vec<GraphEdge>, GraphError> { - let subject = NamedNode::new(&node.iri).map_err(|e| GraphError::InvalidIri(e.to_string()))?; - let mut edges = Vec::new(); - - for quad in self.store.quads_for_pattern(Some(subject.as_ref().into()), None, None, None) { - let quad = quad.map_err(|e| GraphError::StoreError(e.to_string()))?; - let predicate = GraphNode::new(quad.predicate.as_str()); - let object = match quad.object { - Term::NamedNode(n) => GraphObject::Node(GraphNode::new(n.as_str())), - Term::Literal(l) => GraphObject::Literal { - value: l.value().to_string(), - datatype: Some(l.datatype().as_str().to_string()), - }, - _ => continue, - }; - edges.push(GraphEdge { - subject: node.clone(), - predicate, - object, - }); - } - Ok(edges) - } - - async fn incoming(&self, node: &GraphNode) -> Result<Vec<GraphEdge>, GraphError> { - let object = NamedNode::new(&node.iri).map_err(|e| GraphError::InvalidIri(e.to_string()))?; - let mut edges = Vec::new(); - - for quad in self.store.quads_for_pattern(None, None, Some(object.as_ref().into()), None) { - let quad = quad.map_err(|e| GraphError::StoreError(e.to_string()))?; - let subject = match quad.subject { - Subject::NamedNode(n) => GraphNode::new(n.as_str()), - _ => continue, - }; - let predicate = GraphNode::new(quad.predicate.as_str()); - edges.push(GraphEdge { - subject, - predicate, - object: GraphObject::Node(node.clone()), - }); - } - Ok(edges) - } - - async fn exists(&self, edge: &GraphEdge) -> Result<bool, GraphError> { - let quad = self.edge_to_quad(edge)?; - self.store.contains(&quad).map_err(|e| GraphError::StoreError(e.to_string())) - } - - async fn delete(&self, edge: &GraphEdge) -> Result<(), GraphError> { - let quad = self.edge_to_quad(edge)?; - self.store.remove(&quad).map_err(|e| GraphError::StoreError(e.to_string()))?; - Ok(()) - } - - async fn neighborhood(&self, node: &GraphNode, hops: usize) -> Result<Vec<GraphNode>, GraphError> { - use std::collections::HashSet; - - let mut visited = HashSet::new(); - let mut frontier = vec![node.clone()]; - visited.insert(node.iri.clone()); - - for _ in 0..hops { - let mut next_frontier = Vec::new(); - for current in frontier { - for edge in self.outgoing(¤t).await? { - if let GraphObject::Node(n) = edge.object { - if visited.insert(n.iri.clone()) { - next_frontier.push(n); - } - } - } - for edge in self.incoming(¤t).await? { - if visited.insert(edge.subject.iri.clone()) { - next_frontier.push(edge.subject); - } - } - } - frontier = next_frontier; - } - - Ok(visited.into_iter().map(GraphNode::new).collect()) - } -} - -#[cfg(test)] -mod tests { - use super::*; - - #[tokio::test] - async fn test_oxigraph_insert_and_query() { - let store = OxiGraphStore::in_memory().expect("TODO: handle error"); - let edge = GraphEdge { - subject: GraphNode::new("https://example.org/Alice"), - predicate: GraphNode::new("https://example.org/knows"), - object: GraphObject::Node(GraphNode::new("https://example.org/Bob")), - }; - - store.insert(&edge).await.expect("TODO: handle error"); - assert!(store.exists(&edge).await.expect("TODO: handle error")); - - let outgoing = store.outgoing(&edge.subject).await.expect("TODO: handle error"); - assert_eq!(outgoing.len(), 1); - } -} diff --git a/verisimdb/rust-core/verisim-graph/src/redb_backend.rs b/verisimdb/rust-core/verisim-graph/src/redb_backend.rs deleted file mode 100644 index 78077b17..00000000 --- a/verisimdb/rust-core/verisim-graph/src/redb_backend.rs +++ /dev/null @@ -1,602 +0,0 @@ -// SPDX-License-Identifier: MPL-2.0 -// Copyright (c) 2026 Jonathan D.A. Jewell (hyperpolymath) <j.d.a.jewell@open.ac.uk> -// -// redb-backed persistent graph store. -// -// This module is only compiled when the `redb-backend` feature is enabled. -// It provides durable graph triple storage using redb (pure Rust, B-tree, ACID, -// single-file database). No C/C++ dependencies — builds on any platform with a -// Rust toolchain. -// -// # Storage Design -// -// Three redb tables store the graph: -// -// 1. **`triples`** — Primary triple store. -// Key: `"{subject}\0{predicate}\0{object_key}"` (null-separated composite key) -// Value: JSON-serialised `GraphEdge` -// -// 2. **`subject_idx`** — Subject index for outgoing edge lookups. -// Key: `"{subject_iri}\0{triple_key}"` (composite for prefix scanning) -// Value: empty (`&[u8]` — presence in index is sufficient) -// -// 3. **`object_idx`** — Object index for incoming edge lookups (node objects only). -// Key: `"{object_iri}\0{triple_key}"` (composite for prefix scanning) -// Value: empty -// -// This design uses redb's efficient `range()` with prefix scanning for -// O(log n) subject/object lookups rather than MultimapTable, which avoids -// the complexity of value deduplication. - -use std::path::{Path, PathBuf}; -use std::sync::Arc; - -use async_trait::async_trait; -use redb::{Database, ReadableDatabase, TableDefinition}; -use serde_json; - -use crate::{GraphEdge, GraphError, GraphNode, GraphObject, GraphStore}; - -/// Primary triple store: composite triple key → serialised GraphEdge. -const TRIPLES: TableDefinition<&[u8], &[u8]> = TableDefinition::new("triples"); - -/// Subject index: `"{subject_iri}\0{triple_key}"` → empty value. -const SUBJECT_IDX: TableDefinition<&[u8], &[u8]> = TableDefinition::new("subject_idx"); - -/// Object index: `"{object_iri}\0{triple_key}"` → empty value. -const OBJECT_IDX: TableDefinition<&[u8], &[u8]> = TableDefinition::new("object_idx"); - -/// Separator byte for composite keys (null byte — not valid in IRIs). -const SEP: u8 = 0x00; - -/// A persistent graph store backed by redb. -/// -/// Provides the same `GraphStore` interface as `SimpleGraphStore` (in-memory) -/// and `OxiGraphStore` (Oxigraph), but with durable on-disk storage via a -/// pure-Rust B-tree database. No C/C++ dependencies. -/// -/// Thread-safe: `Database` is `Send + Sync` with internal locking. -/// -/// # Example -/// -/// ```rust,ignore -/// use verisim_graph::{RedbGraphStore, GraphStore, GraphEdge, GraphNode, GraphObject}; -/// -/// let store = RedbGraphStore::persistent("/tmp/graph-test.redb").unwrap(); -/// -/// let edge = GraphEdge { -/// subject: GraphNode::new("https://example.org/Alice"), -/// predicate: GraphNode::new("https://example.org/knows"), -/// object: GraphObject::Node(GraphNode::new("https://example.org/Bob")), -/// }; -/// -/// store.insert(&edge).await.unwrap(); -/// let outgoing = store.outgoing(&edge.subject).await.unwrap(); -/// assert_eq!(outgoing.len(), 1); -/// ``` -pub struct RedbGraphStore { - db: Arc<Database>, - #[allow(dead_code)] - path: PathBuf, -} - -impl RedbGraphStore { - /// Open or create a persistent graph store at the given path. - pub fn persistent(path: impl AsRef<Path>) -> Result<Self, GraphError> { - let path = path.as_ref().to_path_buf(); - - if let Some(parent) = path.parent() { - std::fs::create_dir_all(parent) - .map_err(|e| GraphError::StoreError(format!("create dirs: {e}")))?; - } - - let db = Database::create(&path) - .map_err(|e| GraphError::StoreError(format!("open redb: {e}")))?; - - Ok(Self { - db: Arc::new(db), - path, - }) - } - - /// Build a composite triple key from an edge: `"{subject}\0{predicate}\0{object_key}"`. - fn triple_key(edge: &GraphEdge) -> Vec<u8> { - let obj_key = match &edge.object { - GraphObject::Node(n) => n.iri.as_bytes().to_vec(), - GraphObject::Literal { value, .. } => { - let mut k = b"literal::".to_vec(); - k.extend_from_slice(value.as_bytes()); - k - } - }; - let mut key = Vec::with_capacity( - edge.subject.iri.len() + edge.predicate.iri.len() + obj_key.len() + 2, - ); - key.extend_from_slice(edge.subject.iri.as_bytes()); - key.push(SEP); - key.extend_from_slice(edge.predicate.iri.as_bytes()); - key.push(SEP); - key.extend_from_slice(&obj_key); - key - } - - /// Build a subject-index key: `"{subject_iri}\0{triple_key}"`. - fn subject_index_key(subject_iri: &str, triple_key: &[u8]) -> Vec<u8> { - let mut key = Vec::with_capacity(subject_iri.len() + 1 + triple_key.len()); - key.extend_from_slice(subject_iri.as_bytes()); - key.push(SEP); - key.extend_from_slice(triple_key); - key - } - - /// Build an object-index key: `"{object_iri}\0{triple_key}"`. - fn object_index_key(object_iri: &str, triple_key: &[u8]) -> Vec<u8> { - let mut key = Vec::with_capacity(object_iri.len() + 1 + triple_key.len()); - key.extend_from_slice(object_iri.as_bytes()); - key.push(SEP); - key.extend_from_slice(triple_key); - key - } - - /// Build a prefix for scanning all entries for a given IRI: `"{iri}\0"`. - fn iri_prefix(iri: &str) -> Vec<u8> { - let mut prefix = Vec::with_capacity(iri.len() + 1); - prefix.extend_from_slice(iri.as_bytes()); - prefix.push(SEP); - prefix - } - - /// Deserialise a GraphEdge from JSON bytes. - fn deserialise_edge(bytes: &[u8]) -> Result<GraphEdge, GraphError> { - serde_json::from_slice(bytes) - .map_err(|e| GraphError::StoreError(format!("deserialise edge: {e}"))) - } - - /// Serialise a GraphEdge to JSON bytes. - fn serialise_edge(edge: &GraphEdge) -> Result<Vec<u8>, GraphError> { - serde_json::to_vec(edge) - .map_err(|e| GraphError::StoreError(format!("serialise edge: {e}"))) - } - - /// Scan an index table for all triple keys matching a given IRI prefix, - /// then look up the corresponding edges in the triples table. - fn scan_index_for_edges( - db: &Database, - index_table: TableDefinition<&[u8], &[u8]>, - iri: &str, - ) -> Result<Vec<GraphEdge>, GraphError> { - let txn = db.begin_read().map_err(|e| GraphError::StoreError(format!("read txn: {e}")))?; - - let idx = match txn.open_table(index_table) { - Ok(t) => t, - Err(_) => return Ok(Vec::new()), - }; - - let triples = match txn.open_table(TRIPLES) { - Ok(t) => t, - Err(_) => return Ok(Vec::new()), - }; - - let prefix = Self::iri_prefix(iri); - let iter = idx.range(prefix.as_slice()..).map_err(|e| { - GraphError::StoreError(format!("index scan: {e}")) - })?; - - let mut edges = Vec::new(); - for entry in iter { - let entry = entry.map_err(|e| GraphError::StoreError(format!("index entry: {e}")))?; - let idx_key = entry.0.value(); - - // Stop when keys no longer match the IRI prefix - if !idx_key.starts_with(&prefix) { - break; - } - - // Extract the triple key from the index key (after the first separator) - let triple_key = &idx_key[prefix.len()..]; - - // Look up the edge in the triples table - if let Some(edge_bytes) = triples.get(triple_key).map_err(|e| { - GraphError::StoreError(format!("triple lookup: {e}")) - })? { - edges.push(Self::deserialise_edge(edge_bytes.value())?); - } - } - - Ok(edges) - } -} - -#[async_trait] -impl GraphStore for RedbGraphStore { - async fn insert(&self, edge: &GraphEdge) -> Result<(), GraphError> { - let db = Arc::clone(&self.db); - let edge = edge.clone(); - - tokio::task::spawn_blocking(move || -> Result<(), GraphError> { - let tkey = Self::triple_key(&edge); - let edge_bytes = Self::serialise_edge(&edge)?; - - let txn = db.begin_write().map_err(|e| { - GraphError::StoreError(format!("write txn: {e}")) - })?; - - { - // Insert the triple - let mut triples = txn.open_table(TRIPLES).map_err(|e| { - GraphError::StoreError(format!("open triples: {e}")) - })?; - triples.insert(tkey.as_slice(), edge_bytes.as_slice()).map_err(|e| { - GraphError::StoreError(format!("insert triple: {e}")) - })?; - } - - { - // Update subject index - let mut subject_idx = txn.open_table(SUBJECT_IDX).map_err(|e| { - GraphError::StoreError(format!("open subject_idx: {e}")) - })?; - let skey = Self::subject_index_key(&edge.subject.iri, &tkey); - subject_idx.insert(skey.as_slice(), &[] as &[u8]).map_err(|e| { - GraphError::StoreError(format!("insert subject_idx: {e}")) - })?; - } - - { - // Update object index (node objects only) - if let GraphObject::Node(n) = &edge.object { - let mut object_idx = txn.open_table(OBJECT_IDX).map_err(|e| { - GraphError::StoreError(format!("open object_idx: {e}")) - })?; - let okey = Self::object_index_key(&n.iri, &tkey); - object_idx.insert(okey.as_slice(), &[] as &[u8]).map_err(|e| { - GraphError::StoreError(format!("insert object_idx: {e}")) - })?; - } - } - - txn.commit().map_err(|e| { - GraphError::StoreError(format!("commit: {e}")) - })?; - - Ok(()) - }) - .await - .map_err(|e| GraphError::StoreError(format!("task join: {e}")))? - } - - async fn outgoing(&self, node: &GraphNode) -> Result<Vec<GraphEdge>, GraphError> { - let db = Arc::clone(&self.db); - let iri = node.iri.clone(); - - tokio::task::spawn_blocking(move || { - Self::scan_index_for_edges(&db, SUBJECT_IDX, &iri) - }) - .await - .map_err(|e| GraphError::StoreError(format!("task join: {e}")))? - } - - async fn incoming(&self, node: &GraphNode) -> Result<Vec<GraphEdge>, GraphError> { - let db = Arc::clone(&self.db); - let iri = node.iri.clone(); - - tokio::task::spawn_blocking(move || { - Self::scan_index_for_edges(&db, OBJECT_IDX, &iri) - }) - .await - .map_err(|e| GraphError::StoreError(format!("task join: {e}")))? - } - - async fn exists(&self, edge: &GraphEdge) -> Result<bool, GraphError> { - let db = Arc::clone(&self.db); - let tkey = Self::triple_key(edge); - - tokio::task::spawn_blocking(move || -> Result<bool, GraphError> { - let txn = db.begin_read().map_err(|e| { - GraphError::StoreError(format!("read txn: {e}")) - })?; - - let table = match txn.open_table(TRIPLES) { - Ok(t) => t, - Err(_) => return Ok(false), - }; - - match table.get(tkey.as_slice()) { - Ok(Some(_)) => Ok(true), - Ok(None) => Ok(false), - Err(e) => Err(GraphError::StoreError(format!("exists check: {e}"))), - } - }) - .await - .map_err(|e| GraphError::StoreError(format!("task join: {e}")))? - } - - async fn delete(&self, edge: &GraphEdge) -> Result<(), GraphError> { - let db = Arc::clone(&self.db); - let edge = edge.clone(); - - tokio::task::spawn_blocking(move || -> Result<(), GraphError> { - let tkey = Self::triple_key(&edge); - - let txn = db.begin_write().map_err(|e| { - GraphError::StoreError(format!("write txn: {e}")) - })?; - - { - // Remove from triples table - let mut triples = txn.open_table(TRIPLES).map_err(|e| { - GraphError::StoreError(format!("open triples: {e}")) - })?; - triples.remove(tkey.as_slice()).map_err(|e| { - GraphError::StoreError(format!("remove triple: {e}")) - })?; - } - - { - // Remove from subject index - let mut subject_idx = txn.open_table(SUBJECT_IDX).map_err(|e| { - GraphError::StoreError(format!("open subject_idx: {e}")) - })?; - let skey = Self::subject_index_key(&edge.subject.iri, &tkey); - subject_idx.remove(skey.as_slice()).map_err(|e| { - GraphError::StoreError(format!("remove subject_idx: {e}")) - })?; - } - - { - // Remove from object index (node objects only) - if let GraphObject::Node(n) = &edge.object { - let mut object_idx = txn.open_table(OBJECT_IDX).map_err(|e| { - GraphError::StoreError(format!("open object_idx: {e}")) - })?; - let okey = Self::object_index_key(&n.iri, &tkey); - object_idx.remove(okey.as_slice()).map_err(|e| { - GraphError::StoreError(format!("remove object_idx: {e}")) - })?; - } - } - - txn.commit().map_err(|e| { - GraphError::StoreError(format!("commit: {e}")) - })?; - - Ok(()) - }) - .await - .map_err(|e| GraphError::StoreError(format!("task join: {e}")))? - } - - async fn neighborhood(&self, node: &GraphNode, hops: usize) -> Result<Vec<GraphNode>, GraphError> { - use std::collections::HashSet; - - let mut visited = HashSet::new(); - let mut frontier = vec![node.clone()]; - visited.insert(node.iri.clone()); - - for _ in 0..hops { - let mut next_frontier = Vec::new(); - for current in frontier { - for edge in self.outgoing(¤t).await? { - if let GraphObject::Node(n) = edge.object { - if visited.insert(n.iri.clone()) { - next_frontier.push(n); - } - } - } - for edge in self.incoming(¤t).await? { - if visited.insert(edge.subject.iri.clone()) { - next_frontier.push(edge.subject); - } - } - } - frontier = next_frontier; - } - - Ok(visited.into_iter().map(GraphNode::new).collect()) - } -} - -#[cfg(test)] -mod tests { - use super::*; - use tempfile::tempdir; - - fn temp_store() -> (RedbGraphStore, tempfile::TempDir) { - let dir = tempdir().expect("TODO: handle error"); - let path = dir.path().join("graph-test.redb"); - let store = RedbGraphStore::persistent(&path).expect("TODO: handle error"); - (store, dir) - } - - fn test_edge(subject: &str, predicate: &str, object: &str) -> GraphEdge { - GraphEdge { - subject: GraphNode::new(subject), - predicate: GraphNode::new(predicate), - object: GraphObject::Node(GraphNode::new(object)), - } - } - - #[tokio::test] - async fn test_insert_and_exists() { - let (store, _dir) = temp_store(); - let edge = test_edge( - "https://example.org/Alice", - "https://example.org/knows", - "https://example.org/Bob", - ); - - assert!(!store.exists(&edge).await.expect("TODO: handle error")); - store.insert(&edge).await.expect("TODO: handle error"); - assert!(store.exists(&edge).await.expect("TODO: handle error")); - } - - #[tokio::test] - async fn test_outgoing_edges() { - let (store, _dir) = temp_store(); - - let e1 = test_edge( - "https://example.org/Alice", - "https://example.org/knows", - "https://example.org/Bob", - ); - let e2 = test_edge( - "https://example.org/Alice", - "https://example.org/likes", - "https://example.org/Carol", - ); - let e3 = test_edge( - "https://example.org/Bob", - "https://example.org/knows", - "https://example.org/Carol", - ); - - store.insert(&e1).await.expect("TODO: handle error"); - store.insert(&e2).await.expect("TODO: handle error"); - store.insert(&e3).await.expect("TODO: handle error"); - - let alice = GraphNode::new("https://example.org/Alice"); - let outgoing = store.outgoing(&alice).await.expect("TODO: handle error"); - assert_eq!(outgoing.len(), 2); - } - - #[tokio::test] - async fn test_incoming_edges() { - let (store, _dir) = temp_store(); - - let e1 = test_edge( - "https://example.org/Alice", - "https://example.org/knows", - "https://example.org/Bob", - ); - let e2 = test_edge( - "https://example.org/Carol", - "https://example.org/knows", - "https://example.org/Bob", - ); - - store.insert(&e1).await.expect("TODO: handle error"); - store.insert(&e2).await.expect("TODO: handle error"); - - let bob = GraphNode::new("https://example.org/Bob"); - let incoming = store.incoming(&bob).await.expect("TODO: handle error"); - assert_eq!(incoming.len(), 2); - } - - #[tokio::test] - async fn test_delete_edge() { - let (store, _dir) = temp_store(); - let edge = test_edge( - "https://example.org/Alice", - "https://example.org/knows", - "https://example.org/Bob", - ); - - store.insert(&edge).await.expect("TODO: handle error"); - assert!(store.exists(&edge).await.expect("TODO: handle error")); - - store.delete(&edge).await.expect("TODO: handle error"); - assert!(!store.exists(&edge).await.expect("TODO: handle error")); - - // Verify indices are cleaned up - let alice = GraphNode::new("https://example.org/Alice"); - let outgoing = store.outgoing(&alice).await.expect("TODO: handle error"); - assert_eq!(outgoing.len(), 0); - } - - #[tokio::test] - async fn test_literal_object() { - let (store, _dir) = temp_store(); - let edge = GraphEdge { - subject: GraphNode::new("https://example.org/Alice"), - predicate: GraphNode::new("https://example.org/name"), - object: GraphObject::Literal { - value: "Alice".to_string(), - datatype: Some("http://www.w3.org/2001/XMLSchema#string".to_string()), - }, - }; - - store.insert(&edge).await.expect("TODO: handle error"); - assert!(store.exists(&edge).await.expect("TODO: handle error")); - - let outgoing = store.outgoing(&edge.subject).await.expect("TODO: handle error"); - assert_eq!(outgoing.len(), 1); - - // Literals should NOT appear in the object index - let alice_node = GraphNode::new("Alice"); - let incoming = store.incoming(&alice_node).await.expect("TODO: handle error"); - assert_eq!(incoming.len(), 0); - } - - #[tokio::test] - async fn test_neighborhood() { - let (store, _dir) = temp_store(); - - // Alice -> Bob -> Carol - let e1 = test_edge( - "https://example.org/Alice", - "https://example.org/knows", - "https://example.org/Bob", - ); - let e2 = test_edge( - "https://example.org/Bob", - "https://example.org/knows", - "https://example.org/Carol", - ); - - store.insert(&e1).await.expect("TODO: handle error"); - store.insert(&e2).await.expect("TODO: handle error"); - - let alice = GraphNode::new("https://example.org/Alice"); - - // 1 hop: Alice, Bob - let neighbors_1 = store.neighborhood(&alice, 1).await.expect("TODO: handle error"); - assert_eq!(neighbors_1.len(), 2); - - // 2 hops: Alice, Bob, Carol - let neighbors_2 = store.neighborhood(&alice, 2).await.expect("TODO: handle error"); - assert_eq!(neighbors_2.len(), 3); - } - - #[tokio::test] - async fn test_deduplication() { - let (store, _dir) = temp_store(); - let edge = test_edge( - "https://example.org/Alice", - "https://example.org/knows", - "https://example.org/Bob", - ); - - // Insert the same edge twice - store.insert(&edge).await.expect("TODO: handle error"); - store.insert(&edge).await.expect("TODO: handle error"); - - let outgoing = store.outgoing(&edge.subject).await.expect("TODO: handle error"); - assert_eq!(outgoing.len(), 1, "Duplicate edges should be deduplicated"); - } - - #[tokio::test] - async fn test_persistence_across_reopen() { - let dir = tempdir().expect("TODO: handle error"); - let path = dir.path().join("persist-test.redb"); - - let edge = test_edge( - "https://example.org/Alice", - "https://example.org/knows", - "https://example.org/Bob", - ); - - // Write and drop - { - let store = RedbGraphStore::persistent(&path).expect("TODO: handle error"); - store.insert(&edge).await.expect("TODO: handle error"); - } - - // Reopen and verify - { - let store = RedbGraphStore::persistent(&path).expect("TODO: handle error"); - assert!(store.exists(&edge).await.expect("TODO: handle error")); - let outgoing = store.outgoing(&edge.subject).await.expect("TODO: handle error"); - assert_eq!(outgoing.len(), 1); - } - } -} diff --git a/verisimdb/rust-core/verisim-nif/Cargo.toml b/verisimdb/rust-core/verisim-nif/Cargo.toml deleted file mode 100644 index 74f3952d..00000000 --- a/verisimdb/rust-core/verisim-nif/Cargo.toml +++ /dev/null @@ -1,32 +0,0 @@ -# SPDX-License-Identifier: MPL-2.0 - -[package] -name = "verisim-nif" -description = "Rustler NIF bridge for VeriSimDB — enables direct in-process calls from Elixir" -version.workspace = true -edition.workspace = true -authors.workspace = true -license.workspace = true - -[lib] -name = "verisim_nif" -crate-type = ["cdylib"] - -[dependencies] -# Rustler provides the Elixir NIF interface -rustler = "0.36" - -# Core VeriSimDB crates -verisim-octad = { path = "../verisim-octad" } -verisim-drift = { path = "../verisim-drift" } -verisim-graph = { path = "../verisim-graph" } -verisim-vector = { path = "../verisim-vector" } -verisim-document = { path = "../verisim-document" } -verisim-normalizer = { path = "../verisim-normalizer" } - -# Serialization for passing data between Elixir and Rust -serde.workspace = true -serde_json.workspace = true - -# Async runtime (octad store operations are async) -tokio.workspace = true diff --git a/verisimdb/rust-core/verisim-nif/src/lib.rs b/verisimdb/rust-core/verisim-nif/src/lib.rs deleted file mode 100644 index 0ef11980..00000000 --- a/verisimdb/rust-core/verisim-nif/src/lib.rs +++ /dev/null @@ -1,216 +0,0 @@ -// SPDX-License-Identifier: MPL-2.0 -// Copyright (c) 2026 Jonathan D.A. Jewell (hyperpolymath) <j.d.a.jewell@open.ac.uk> -//! -//! VeriSimDB NIF Bridge — Erlang/Elixir native interface for direct in-process calls. -//! -//! This crate exposes core VeriSimDB operations as Rustler NIFs, bypassing HTTP -//! for same-node deployments. When used alongside the Elixir orchestration layer, -//! NIF transport provides 10-100x lower latency than HTTP for octad CRUD operations. -//! -//! ## Supported Operations (MVP) -//! -//! - `create_octad/1` — Create a new octad entity from JSON input -//! - `get_octad/1` — Retrieve a octad by ID (all 8 modalities) -//! - `delete_octad/1` — Delete a octad entity -//! - `search_text/2` — Full-text search across document modality -//! - `search_vector/2` — Vector similarity search -//! - `list_octads/2` — Paginated entity listing -//! - `get_drift_score/1` — Get drift scores for an entity -//! - `trigger_normalise/1` — Trigger normalisation for a drifted entity -//! -//! ## Transport Selection -//! -//! The Elixir `VeriSim.RustClient` module selects transport via: -//! ``` -//! VERISIM_TRANSPORT=http # Default: HTTP to verisim-api server -//! VERISIM_TRANSPORT=nif # Direct NIF calls (same-node only) -//! VERISIM_TRANSPORT=auto # NIF if available, HTTP fallback -//! ``` - -#![forbid(unsafe_code)] -use rustler::{Env, Error, NifResult, Term}; -use serde_json::Value; -use std::sync::OnceLock; -use tokio::runtime::Runtime; - -/// Shared Tokio runtime for executing async store operations from synchronous -/// NIF entry points. Initialised on first NIF call. -/// -/// Currently unused — will be activated when store operations are wired into -/// the NIF functions (replacing the placeholder responses). -#[allow(dead_code)] -static RUNTIME: OnceLock<Runtime> = OnceLock::new(); - -/// Get or create the shared Tokio runtime. -#[allow(dead_code)] -fn runtime() -> &'static Runtime { - RUNTIME.get_or_init(|| { - tokio::runtime::Builder::new_multi_thread() - .worker_threads(4) - .enable_all() - .build() - .expect("failed to create Tokio runtime for NIF bridge") - }) -} - -// --------------------------------------------------------------------------- -// NIF functions -// --------------------------------------------------------------------------- - -/// Create a new octad entity from a JSON string. -/// -/// Accepts a JSON string matching the `OctadInput` schema (same as the HTTP -/// POST /api/v1/octads body). Returns `{:ok, json_string}` on success or -/// `{:error, reason}` on failure. -#[rustler::nif(schedule = "DirtyCpu")] -fn create_octad(json_input: String) -> NifResult<String> { - let input: Value = serde_json::from_str(&json_input) - .map_err(|e| Error::Term(Box::new(format!("invalid JSON: {e}"))))?; - - // Placeholder: in full integration, this calls InMemoryOctadStore::create() - // via the shared store instance. For now, return the parsed input as - // confirmation that the NIF bridge is functional. - let result = serde_json::json!({ - "status": "created", - "input_keys": input.as_object().map(|o| o.keys().collect::<Vec<_>>()).unwrap_or_default(), - "transport": "nif" - }); - - serde_json::to_string(&result) - .map_err(|e| Error::Term(Box::new(format!("serialization error: {e}")))) -} - -/// Retrieve a octad by ID. -/// -/// Returns the full octad JSON (all 8 octad modalities) or an error if not found. -#[rustler::nif(schedule = "DirtyCpu")] -fn get_octad(octad_id: String) -> NifResult<String> { - // Placeholder: will call store.get(&octad_id) via the shared runtime - let result = serde_json::json!({ - "id": octad_id, - "status": "nif_placeholder", - "transport": "nif", - "message": "NIF bridge operational — store integration pending" - }); - - serde_json::to_string(&result) - .map_err(|e| Error::Term(Box::new(format!("serialization error: {e}")))) -} - -/// Delete a octad entity by ID. -#[rustler::nif(schedule = "DirtyCpu")] -fn delete_octad(octad_id: String) -> NifResult<String> { - let result = serde_json::json!({ - "id": octad_id, - "status": "deleted", - "transport": "nif" - }); - - serde_json::to_string(&result) - .map_err(|e| Error::Term(Box::new(format!("serialization error: {e}")))) -} - -/// Full-text search across the document modality. -/// -/// Accepts a query string and result limit. Returns a JSON array of matching octads. -#[rustler::nif(schedule = "DirtyCpu")] -fn search_text(query: String, limit: usize) -> NifResult<String> { - let result = serde_json::json!({ - "query": query, - "limit": limit, - "results": [], - "transport": "nif" - }); - - serde_json::to_string(&result) - .map_err(|e| Error::Term(Box::new(format!("serialization error: {e}")))) -} - -/// Vector similarity search. -/// -/// Accepts a JSON-encoded embedding vector and a `k` parameter for top-K results. -#[rustler::nif(schedule = "DirtyCpu")] -fn search_vector(embedding_json: String, k: usize) -> NifResult<String> { - let _embedding: Vec<f32> = serde_json::from_str(&embedding_json) - .map_err(|e| Error::Term(Box::new(format!("invalid embedding JSON: {e}"))))?; - - let result = serde_json::json!({ - "k": k, - "results": [], - "transport": "nif" - }); - - serde_json::to_string(&result) - .map_err(|e| Error::Term(Box::new(format!("serialization error: {e}")))) -} - -/// Paginated listing of octad entities. -#[rustler::nif(schedule = "DirtyCpu")] -fn list_octads(limit: usize, offset: usize) -> NifResult<String> { - let result = serde_json::json!({ - "limit": limit, - "offset": offset, - "octads": [], - "total": 0, - "transport": "nif" - }); - - serde_json::to_string(&result) - .map_err(|e| Error::Term(Box::new(format!("serialization error: {e}")))) -} - -/// Get drift detection scores for a specific entity. -/// -/// Returns drift scores across all 8 octad modalities (0.0 = no drift, 1.0 = max). -#[rustler::nif(schedule = "DirtyCpu")] -fn get_drift_score(octad_id: String) -> NifResult<String> { - let result = serde_json::json!({ - "entity_id": octad_id, - "graph": 0.0, - "vector": 0.0, - "tensor": 0.0, - "semantic": 0.0, - "document": 0.0, - "temporal": 0.0, - "provenance": 0.0, - "spatial": 0.0, - "overall": 0.0, - "transport": "nif" - }); - - serde_json::to_string(&result) - .map_err(|e| Error::Term(Box::new(format!("serialization error: {e}")))) -} - -/// Trigger normalisation (self-repair) for a drifted entity. -/// -/// Returns the normalisation result status. -#[rustler::nif(schedule = "DirtyCpu")] -fn trigger_normalise(octad_id: String) -> NifResult<String> { - let result = serde_json::json!({ - "entity_id": octad_id, - "status": "normalisation_triggered", - "transport": "nif" - }); - - serde_json::to_string(&result) - .map_err(|e| Error::Term(Box::new(format!("serialization error: {e}")))) -} - -// --------------------------------------------------------------------------- -// NIF registration -// --------------------------------------------------------------------------- - -rustler::init!( - "Elixir.VeriSim.NifBridge", - [ - create_octad, - get_octad, - delete_octad, - search_text, - search_vector, - list_octads, - get_drift_score, - trigger_normalise, - ] -); diff --git a/verisimdb/rust-core/verisim-normalizer/Cargo.toml b/verisimdb/rust-core/verisim-normalizer/Cargo.toml deleted file mode 100644 index 53ef2a19..00000000 --- a/verisimdb/rust-core/verisim-normalizer/Cargo.toml +++ /dev/null @@ -1,32 +0,0 @@ -# SPDX-License-Identifier: MPL-2.0 - -[package] -name = "verisim-normalizer" -description = "Self-normalization engine - maintains cross-modal consistency" -version.workspace = true -edition.workspace = true -authors.workspace = true -license.workspace = true - -[dependencies] -verisim-octad = { path = "../verisim-octad" } -verisim-drift = { path = "../verisim-drift" } - -serde.workspace = true -chrono.workspace = true -thiserror.workspace = true -tracing.workspace = true -async-trait.workspace = true -tokio.workspace = true -futures.workspace = true -prometheus.workspace = true -uuid.workspace = true - -[dev-dependencies] -proptest.workspace = true -serde_json.workspace = true -verisim-document = { path = "../verisim-document" } -verisim-vector = { path = "../verisim-vector" } -verisim-graph = { path = "../verisim-graph" } -verisim-semantic = { path = "../verisim-semantic" } -verisim-tensor = { path = "../verisim-tensor" } diff --git a/verisimdb/rust-core/verisim-normalizer/src/conflict.rs b/verisimdb/rust-core/verisim-normalizer/src/conflict.rs deleted file mode 100644 index aa3efc37..00000000 --- a/verisimdb/rust-core/verisim-normalizer/src/conflict.rs +++ /dev/null @@ -1,1561 +0,0 @@ -// SPDX-License-Identifier: MPL-2.0 -// Copyright (c) 2026 Jonathan D.A. Jewell (hyperpolymath) <j.d.a.jewell@open.ac.uk> - -//! Conflict Resolution Policies for the VeriSim Normalizer -//! -//! When two or more modalities report conflicting information about the same -//! entity, a **conflict** arises. This module provides: -//! -//! - Detection and recording of conflicts ([`ConflictResolver::detect_conflict`]). -//! - Policy-based resolution ([`ConflictPolicy`]) with five built-in strategies: -//! last-writer-wins, modality-priority, manual-resolve, auto-merge, and custom. -//! - Threshold-gated auto-resolution: conflicts below -//! [`ConflictConfig::auto_resolve_threshold`] are resolved automatically, while -//! those above [`ConflictConfig::require_manual_above`] are always escalated. -//! - Per-modality-pair policy overrides for fine-grained control. -//! - Full history tracking of resolved and dismissed conflicts. -//! -//! ## Integration -//! -//! The [`ConflictResolver`] sits alongside the [`RegenerationEngine`] in the -//! normalizer pipeline. After drift detection identifies *that* modalities have -//! diverged, the conflict resolver decides *who wins* when both sides carry -//! valid but contradictory data. -//! -//! [`RegenerationEngine`]: crate::regeneration::RegenerationEngine - -use chrono::{DateTime, Utc}; -use serde::{Deserialize, Serialize}; -use std::collections::HashMap; -use std::fmt; -use tokio::sync::RwLock; -use tracing::{debug, info, warn}; -use uuid::Uuid; - -use crate::regeneration::Modality; - -// --------------------------------------------------------------------------- -// ConflictPolicy -// --------------------------------------------------------------------------- - -/// Policy that governs how a conflict between modalities is resolved. -/// -/// Each variant represents a distinct resolution strategy. The policy can be -/// set globally via [`ConflictConfig::default_policy`], overridden for specific -/// modality pairs via [`ConflictConfig::per_modality_policies`], or supplied -/// ad-hoc when calling [`ConflictResolver::resolve`]. -#[derive(Debug, Clone, PartialEq, Serialize, Deserialize)] -pub enum ConflictPolicy { - /// Last writer wins -- the most recently updated modality takes precedence. - /// - /// This is the simplest policy and is appropriate when all modalities are - /// considered equally trustworthy. The modality whose data was written last - /// (by wall-clock time) is chosen as the winner. - LastWriterWins, - - /// Modality priority -- a modality higher in the supplied authority order - /// wins over one lower in the order. - /// - /// The contained `Vec<Modality>` defines the priority ranking from highest - /// (index 0) to lowest. If a conflicting modality is not in the list, it - /// is treated as lowest priority. - ModalityPriority(Vec<Modality>), - - /// Manual resolution -- the conflict is flagged for human review. - /// - /// No automatic winner is selected. The conflict status moves to - /// [`ConflictStatus::InProgress`] and remains there until a human calls - /// [`ConflictResolver::resolve_manual`]. - ManualResolve, - - /// Auto-merge -- attempt to automatically merge conflicting data. - /// - /// The resolver picks the modality with the richest data (most fields - /// populated, longest content, etc.) as the primary source, and annotates - /// the resolution accordingly. This is a heuristic, not a guarantee. - AutoMerge, - - /// Custom -- delegate resolution to a user-provided external resolver. - /// - /// The contained `String` identifies the external resolver (e.g. a webhook - /// URL or plugin name). The conflict is marked [`ConflictStatus::InProgress`] - /// until the external system calls back with a decision. - Custom(String), -} - -impl fmt::Display for ConflictPolicy { - fn fmt(&self, f: &mut fmt::Formatter<'_>) -> fmt::Result { - match self { - ConflictPolicy::LastWriterWins => write!(f, "last_writer_wins"), - ConflictPolicy::ModalityPriority(_) => write!(f, "modality_priority"), - ConflictPolicy::ManualResolve => write!(f, "manual_resolve"), - ConflictPolicy::AutoMerge => write!(f, "auto_merge"), - ConflictPolicy::Custom(name) => write!(f, "custom({})", name), - } - } -} - -// --------------------------------------------------------------------------- -// ConflictStatus -// --------------------------------------------------------------------------- - -/// Lifecycle state of a conflict. -#[derive(Debug, Clone, PartialEq, Eq, Serialize, Deserialize)] -pub enum ConflictStatus { - /// Conflict has been detected but not yet addressed. - Open, - - /// Resolution is in progress (manual review or custom resolver). - InProgress, - - /// Conflict has been resolved (see [`ConflictResolution`] for details). - Resolved, - - /// Conflict was dismissed without resolution (e.g. deemed a false positive). - Dismissed, -} - -impl fmt::Display for ConflictStatus { - fn fmt(&self, f: &mut fmt::Formatter<'_>) -> fmt::Result { - match self { - ConflictStatus::Open => write!(f, "open"), - ConflictStatus::InProgress => write!(f, "in_progress"), - ConflictStatus::Resolved => write!(f, "resolved"), - ConflictStatus::Dismissed => write!(f, "dismissed"), - } - } -} - -// --------------------------------------------------------------------------- -// Conflict -// --------------------------------------------------------------------------- - -/// A detected conflict between two or more modalities on a single entity. -/// -/// Conflicts are created by [`ConflictResolver::detect_conflict`] and tracked -/// through the [`ConflictStatus`] lifecycle until resolution or dismissal. -#[derive(Debug, Clone, Serialize, Deserialize)] -pub struct Conflict { - /// Unique identifier for this conflict instance. - pub id: String, - - /// The octad entity on which the conflict was detected. - pub entity_id: String, - - /// When the conflict was first detected. - pub detected_at: DateTime<Utc>, - - /// The modalities whose data disagrees. - pub conflicting_modalities: Vec<Modality>, - - /// Measured drift score between the conflicting modalities (0.0 -- 1.0). - pub drift_score: f64, - - /// Human-readable description of what the conflict is about. - pub description: String, - - /// Current lifecycle status. - pub status: ConflictStatus, - - /// How the conflict was resolved, if applicable. - pub resolution: Option<ConflictResolution>, -} - -// --------------------------------------------------------------------------- -// ConflictResolution -// --------------------------------------------------------------------------- - -/// Record of how a conflict was resolved. -/// -/// Attached to a [`Conflict`] once it reaches [`ConflictStatus::Resolved`]. -#[derive(Debug, Clone, Serialize, Deserialize)] -pub struct ConflictResolution { - /// When the resolution was applied. - pub resolved_at: DateTime<Utc>, - - /// Which policy was used to determine the winner. - pub policy_used: ConflictPolicy, - - /// The modality that was chosen as the authoritative source, if applicable. - /// - /// Some policies (e.g. [`ConflictPolicy::AutoMerge`]) may not select a - /// single winner, in which case this is `None`. - pub winning_modality: Option<Modality>, - - /// Who or what performed the resolution. - /// - /// `"system"` for automatic resolution, or a user/service identifier for - /// manual/custom resolution. - pub resolver: String, - - /// Optional notes from the resolver explaining the decision. - pub notes: Option<String>, -} - -// --------------------------------------------------------------------------- -// ConflictConfig -// --------------------------------------------------------------------------- - -/// Configuration for the conflict resolution subsystem. -/// -/// Controls which policy is used by default, per-modality-pair overrides, -/// and threshold gates for automatic vs. manual escalation. -#[derive(Debug, Clone, Serialize, Deserialize)] -pub struct ConflictConfig { - /// The default policy applied when no per-pair override matches. - pub default_policy: ConflictPolicy, - - /// Policy overrides keyed by ordered modality pair. - /// - /// If a conflict involves modalities `(A, B)` and there is an entry for - /// `(A, B)` **or** `(B, A)`, that policy takes precedence over - /// `default_policy`. - pub per_modality_policies: HashMap<(Modality, Modality), ConflictPolicy>, - - /// Drift score at or below which conflicts are auto-resolved using the - /// applicable policy, without human intervention. - /// - /// Range: 0.0 -- 1.0. Conflicts with `drift_score <= auto_resolve_threshold` - /// are resolved immediately. - pub auto_resolve_threshold: f64, - - /// Drift score at or above which conflicts **always** require manual review, - /// regardless of the configured policy. - /// - /// Range: 0.0 -- 1.0. Must be >= `auto_resolve_threshold`. - pub require_manual_above: f64, - - /// Maximum number of resolved/dismissed conflicts to retain in history. - /// - /// Older entries are evicted when this limit is exceeded (FIFO). - pub max_history_entries: usize, -} - -impl Default for ConflictConfig { - fn default() -> Self { - Self { - default_policy: ConflictPolicy::LastWriterWins, - per_modality_policies: HashMap::new(), - auto_resolve_threshold: 0.3, - require_manual_above: 0.8, - max_history_entries: 1000, - } - } -} - -impl ConflictConfig { - /// Look up the policy for a specific pair of conflicting modalities. - /// - /// Checks both `(a, b)` and `(b, a)` orderings before falling back to - /// `default_policy`. - pub fn policy_for_pair(&self, a: Modality, b: Modality) -> &ConflictPolicy { - self.per_modality_policies - .get(&(a, b)) - .or_else(|| self.per_modality_policies.get(&(b, a))) - .unwrap_or(&self.default_policy) - } -} - -// --------------------------------------------------------------------------- -// ConflictError -// --------------------------------------------------------------------------- - -/// Errors returned by the conflict resolution engine. -#[derive(Debug, Clone, PartialEq, Eq, thiserror::Error)] -pub enum ConflictError { - /// No conflict with the given ID was found. - #[error("Conflict not found: {0}")] - NotFound(String), - - /// The conflict has already been resolved or dismissed. - #[error("Conflict already resolved: {0}")] - AlreadyResolved(String), - - /// The supplied policy is invalid for the given conflict. - #[error("Invalid policy: {0}")] - InvalidPolicy(String), - - /// The specified modality is not one of the conflicting modalities. - #[error("Modality not in conflict: {0}")] - ModalityNotInConflict(String), -} - -// --------------------------------------------------------------------------- -// ConflictResolver -// --------------------------------------------------------------------------- - -/// The conflict resolution engine. -/// -/// Manages the lifecycle of conflicts from detection through resolution or -/// dismissal, applying configured policies to determine winners. -/// -/// Thread-safe: all mutable state is behind [`RwLock`] guards. -pub struct ConflictResolver { - /// Configuration controlling policies, thresholds, and history limits. - config: ConflictConfig, - - /// Active (Open or InProgress) conflicts. - active_conflicts: RwLock<Vec<Conflict>>, - - /// History of resolved and dismissed conflicts (bounded by - /// `config.max_history_entries`). - history: RwLock<Vec<Conflict>>, -} - -impl ConflictResolver { - /// Create a new conflict resolver with the given configuration. - pub fn new(config: ConflictConfig) -> Self { - Self { - config, - active_conflicts: RwLock::new(Vec::new()), - history: RwLock::new(Vec::new()), - } - } - - /// Create a resolver with default configuration. - pub fn with_defaults() -> Self { - Self::new(ConflictConfig::default()) - } - - /// Access the underlying configuration. - pub fn config(&self) -> &ConflictConfig { - &self.config - } - - // -- detection ----------------------------------------------------------- - - /// Detect and record a new conflict. - /// - /// Creates a [`Conflict`] with status [`ConflictStatus::Open`], assigns it - /// a unique ID, and stores it in the active conflicts list. - /// - /// # Arguments - /// - /// * `entity_id` -- the octad entity on which the conflict was detected. - /// * `modalities` -- the modalities whose data disagrees. - /// * `drift_score` -- measured drift score (0.0 -- 1.0). - /// * `description` -- human-readable description of the conflict. - /// - /// # Returns - /// - /// The newly created [`Conflict`]. - pub async fn detect_conflict( - &self, - entity_id: &str, - modalities: Vec<Modality>, - drift_score: f64, - description: &str, - ) -> Conflict { - let conflict = Conflict { - id: Uuid::new_v4().to_string(), - entity_id: entity_id.to_string(), - detected_at: Utc::now(), - conflicting_modalities: modalities.clone(), - drift_score, - description: description.to_string(), - status: ConflictStatus::Open, - resolution: None, - }; - - info!( - conflict_id = %conflict.id, - entity_id = %entity_id, - drift_score = drift_score, - modalities = ?modalities, - "Conflict detected" - ); - - self.active_conflicts.write().await.push(conflict.clone()); - conflict - } - - // -- resolution ---------------------------------------------------------- - - /// Resolve a conflict using the specified policy, or the applicable default. - /// - /// If `policy` is `None`, the resolver determines the appropriate policy - /// based on the conflict's drift score and the configured thresholds and - /// per-pair overrides. - /// - /// # Policy selection logic - /// - /// 1. If `drift_score >= require_manual_above`, force [`ConflictPolicy::ManualResolve`]. - /// 2. If an explicit `policy` argument is provided, use it. - /// 3. If a per-modality-pair override exists, use it. - /// 4. Otherwise, use `default_policy`. - /// - /// # Errors - /// - /// - [`ConflictError::NotFound`] if no active conflict with the given ID exists. - /// - [`ConflictError::AlreadyResolved`] if the conflict is already Resolved or Dismissed. - /// - [`ConflictError::InvalidPolicy`] if the selected policy cannot be applied - /// (e.g. [`ConflictPolicy::ManualResolve`] or [`ConflictPolicy::Custom`] cannot - /// produce an automatic resolution). - pub async fn resolve( - &self, - conflict_id: &str, - policy: Option<ConflictPolicy>, - ) -> Result<ConflictResolution, ConflictError> { - let mut active = self.active_conflicts.write().await; - let conflict = active - .iter_mut() - .find(|c| c.id == conflict_id) - .ok_or_else(|| ConflictError::NotFound(conflict_id.to_string()))?; - - // Guard: already resolved or dismissed - if conflict.status == ConflictStatus::Resolved - || conflict.status == ConflictStatus::Dismissed - { - return Err(ConflictError::AlreadyResolved(conflict_id.to_string())); - } - - // Step 1: Force manual if drift is very high - if conflict.drift_score >= self.config.require_manual_above { - conflict.status = ConflictStatus::InProgress; - debug!( - conflict_id = %conflict_id, - drift_score = conflict.drift_score, - threshold = self.config.require_manual_above, - "Drift score above manual threshold -- forcing manual resolution" - ); - return Err(ConflictError::InvalidPolicy(format!( - "Drift score {:.3} >= manual threshold {:.3} -- use resolve_manual()", - conflict.drift_score, self.config.require_manual_above - ))); - } - - // Step 2-4: Select policy - let effective_policy = if let Some(ref p) = policy { - p.clone() - } else { - self.select_policy(conflict) - }; - - // ManualResolve and Custom cannot produce automatic resolutions - match &effective_policy { - ConflictPolicy::ManualResolve => { - conflict.status = ConflictStatus::InProgress; - return Err(ConflictError::InvalidPolicy( - "ManualResolve policy requires resolve_manual()".to_string(), - )); - } - ConflictPolicy::Custom(name) => { - conflict.status = ConflictStatus::InProgress; - return Err(ConflictError::InvalidPolicy(format!( - "Custom policy '{}' requires external resolver callback", - name - ))); - } - _ => {} - } - - // Apply the policy - let resolution = self.apply_policy(conflict, &effective_policy); - - info!( - conflict_id = %conflict_id, - policy = %effective_policy, - winning_modality = ?resolution.winning_modality, - "Conflict resolved automatically" - ); - - // Update state - conflict.status = ConflictStatus::Resolved; - conflict.resolution = Some(resolution.clone()); - - // Move to history - let resolved = conflict.clone(); - drop(active); - self.move_to_history(resolved).await; - - Ok(resolution) - } - - /// Manually resolve a conflict by specifying the winning modality. - /// - /// This is the only way to resolve conflicts that have been escalated to - /// [`ConflictPolicy::ManualResolve`] or that exceeded the - /// `require_manual_above` threshold. - /// - /// # Arguments - /// - /// * `conflict_id` -- ID of the conflict to resolve. - /// * `winning_modality` -- the modality chosen as the authoritative source. - /// * `resolver` -- identifier of the human or service performing the resolution. - /// * `notes` -- optional free-text explanation. - /// - /// # Errors - /// - /// - [`ConflictError::NotFound`] if no active conflict with the given ID exists. - /// - [`ConflictError::AlreadyResolved`] if the conflict is already Resolved or Dismissed. - /// - [`ConflictError::ModalityNotInConflict`] if `winning_modality` is not - /// one of the conflicting modalities. - pub async fn resolve_manual( - &self, - conflict_id: &str, - winning_modality: Modality, - resolver: &str, - notes: Option<String>, - ) -> Result<ConflictResolution, ConflictError> { - let mut active = self.active_conflicts.write().await; - let conflict = active - .iter_mut() - .find(|c| c.id == conflict_id) - .ok_or_else(|| ConflictError::NotFound(conflict_id.to_string()))?; - - // Guard: already resolved or dismissed - if conflict.status == ConflictStatus::Resolved - || conflict.status == ConflictStatus::Dismissed - { - return Err(ConflictError::AlreadyResolved(conflict_id.to_string())); - } - - // Guard: winning modality must be part of the conflict - if !conflict.conflicting_modalities.contains(&winning_modality) { - return Err(ConflictError::ModalityNotInConflict(format!( - "{} is not in conflicting modalities {:?}", - winning_modality, conflict.conflicting_modalities - ))); - } - - let resolution = ConflictResolution { - resolved_at: Utc::now(), - policy_used: ConflictPolicy::ManualResolve, - winning_modality: Some(winning_modality), - resolver: resolver.to_string(), - notes, - }; - - info!( - conflict_id = %conflict_id, - winning_modality = %winning_modality, - resolver = %resolver, - "Conflict resolved manually" - ); - - conflict.status = ConflictStatus::Resolved; - conflict.resolution = Some(resolution.clone()); - - // Move to history - let resolved = conflict.clone(); - drop(active); - self.move_to_history(resolved).await; - - Ok(resolution) - } - - /// Dismiss a conflict without resolving it. - /// - /// Dismissed conflicts are moved to history with a note explaining why they - /// were dismissed. This is appropriate for false positives or conflicts that - /// have been superseded by other changes. - /// - /// # Errors - /// - /// - [`ConflictError::NotFound`] if no active conflict with the given ID exists. - /// - [`ConflictError::AlreadyResolved`] if the conflict is already Resolved or Dismissed. - pub async fn dismiss( - &self, - conflict_id: &str, - reason: &str, - ) -> Result<(), ConflictError> { - let mut active = self.active_conflicts.write().await; - let conflict = active - .iter_mut() - .find(|c| c.id == conflict_id) - .ok_or_else(|| ConflictError::NotFound(conflict_id.to_string()))?; - - // Guard: already resolved or dismissed - if conflict.status == ConflictStatus::Resolved - || conflict.status == ConflictStatus::Dismissed - { - return Err(ConflictError::AlreadyResolved(conflict_id.to_string())); - } - - info!( - conflict_id = %conflict_id, - reason = %reason, - "Conflict dismissed" - ); - - conflict.status = ConflictStatus::Dismissed; - conflict.resolution = Some(ConflictResolution { - resolved_at: Utc::now(), - policy_used: ConflictPolicy::LastWriterWins, // placeholder -- no actual policy used - winning_modality: None, - resolver: "system".to_string(), - notes: Some(format!("Dismissed: {}", reason)), - }); - - // Move to history - let dismissed = conflict.clone(); - drop(active); - self.move_to_history(dismissed).await; - - Ok(()) - } - - // -- queries ------------------------------------------------------------- - - /// Return all active (Open or InProgress) conflicts. - pub async fn active_conflicts(&self) -> Vec<Conflict> { - self.active_conflicts.read().await.clone() - } - - /// Return resolved/dismissed conflict history, most recent first. - /// - /// # Arguments - /// - /// * `limit` -- maximum number of entries to return. If `0`, returns all. - pub async fn history(&self, limit: usize) -> Vec<Conflict> { - let history = self.history.read().await; - if limit == 0 || limit >= history.len() { - // Return in reverse chronological order (most recent first) - let mut result = history.clone(); - result.reverse(); - result - } else { - // Take the last `limit` entries (most recent) and reverse - let start = history.len() - limit; - let mut result = history[start..].to_vec(); - result.reverse(); - result - } - } - - /// Return all conflicts (active and historical) for a given entity. - /// - /// Results are ordered with active conflicts first, then historical in - /// reverse chronological order. - pub async fn by_entity(&self, entity_id: &str) -> Vec<Conflict> { - let active = self.active_conflicts.read().await; - let history = self.history.read().await; - - let mut result: Vec<Conflict> = active - .iter() - .filter(|c| c.entity_id == entity_id) - .cloned() - .collect(); - - let mut hist: Vec<Conflict> = history - .iter() - .filter(|c| c.entity_id == entity_id) - .cloned() - .collect(); - hist.reverse(); - - result.extend(hist); - result - } - - // -- internal helpers ---------------------------------------------------- - - /// Select the most appropriate policy for a conflict based on config. - /// - /// Checks per-modality-pair overrides first, then falls back to the default - /// policy. - fn select_policy(&self, conflict: &Conflict) -> ConflictPolicy { - // Check per-pair overrides for the first matching pair - let modalities = &conflict.conflicting_modalities; - for i in 0..modalities.len() { - for j in (i + 1)..modalities.len() { - let pair_policy = self - .config - .per_modality_policies - .get(&(modalities[i], modalities[j])) - .or_else(|| { - self.config - .per_modality_policies - .get(&(modalities[j], modalities[i])) - }); - if let Some(policy) = pair_policy { - debug!( - pair = ?(&modalities[i], &modalities[j]), - policy = %policy, - "Using per-modality-pair policy override" - ); - return policy.clone(); - } - } - } - - self.config.default_policy.clone() - } - - /// Apply a policy to a conflict and produce a resolution. - /// - /// This method determines the winning modality based on the policy logic. - /// It is called internally by [`resolve`](Self::resolve) after policy - /// selection and threshold checks. - fn apply_policy(&self, conflict: &Conflict, policy: &ConflictPolicy) -> ConflictResolution { - let winning_modality = match policy { - ConflictPolicy::LastWriterWins => { - // The last modality in the list is considered the most recently - // updated (caller orders them by write timestamp). - conflict.conflicting_modalities.last().copied() - } - ConflictPolicy::ModalityPriority(priority_order) => { - // Find the conflicting modality with the highest priority - // (lowest index in the priority order). - let mut best: Option<(usize, Modality)> = None; - for m in &conflict.conflicting_modalities { - if let Some(idx) = priority_order.iter().position(|p| p == m) { - match best { - None => best = Some((idx, *m)), - Some((best_idx, _)) if idx < best_idx => { - best = Some((idx, *m)); - } - _ => {} - } - } - } - // If none of the conflicting modalities appear in the priority - // list, fall back to the first conflicting modality. - best.map(|(_, m)| m) - .or_else(|| conflict.conflicting_modalities.first().copied()) - } - ConflictPolicy::AutoMerge => { - // For auto-merge, we don't pick a single winner -- the merge - // process combines data from all conflicting modalities. - // We return None to indicate no single winner. - None - } - ConflictPolicy::ManualResolve | ConflictPolicy::Custom(_) => { - // These policies should not reach apply_policy -- they are - // intercepted earlier. Defensive coding: return None. - warn!( - policy = %policy, - "apply_policy called with non-automatic policy -- returning no winner" - ); - None - } - }; - - ConflictResolution { - resolved_at: Utc::now(), - policy_used: policy.clone(), - winning_modality, - resolver: "system".to_string(), - notes: None, - } - } - - /// Move a resolved or dismissed conflict from active list to history. - /// - /// Removes the conflict from `active_conflicts` by ID and appends it to - /// `history`, evicting the oldest entry if `max_history_entries` is exceeded. - async fn move_to_history(&self, conflict: Conflict) { - // Remove from active - { - let mut active = self.active_conflicts.write().await; - active.retain(|c| c.id != conflict.id); - } - - // Add to history - { - let mut history = self.history.write().await; - history.push(conflict); - - // Evict oldest entries if history is too large - let max = self.config.max_history_entries; - if history.len() > max { - let excess = history.len() - max; - history.drain(0..excess); - } - } - } -} - -// =========================================================================== -// Tests -// =========================================================================== - -#[cfg(test)] -mod tests { - use super::*; - - // -- helpers ------------------------------------------------------------- - - /// Create a default config for testing. - fn test_config() -> ConflictConfig { - ConflictConfig { - default_policy: ConflictPolicy::LastWriterWins, - per_modality_policies: HashMap::new(), - auto_resolve_threshold: 0.3, - require_manual_above: 0.8, - max_history_entries: 100, - } - } - - // -- ConflictConfig default tests ---------------------------------------- - - #[test] - fn test_conflict_config_defaults() { - let config = ConflictConfig::default(); - assert_eq!(config.default_policy, ConflictPolicy::LastWriterWins); - assert!(config.per_modality_policies.is_empty()); - assert!((config.auto_resolve_threshold - 0.3).abs() < f64::EPSILON); - assert!((config.require_manual_above - 0.8).abs() < f64::EPSILON); - assert_eq!(config.max_history_entries, 1000); - } - - // -- LastWriterWins tests ------------------------------------------------ - - #[tokio::test] - async fn test_last_writer_wins_resolution() { - let resolver = ConflictResolver::new(test_config()); - - let conflict = resolver - .detect_conflict( - "entity-1", - vec![Modality::Document, Modality::Vector], - 0.5, - "Document and vector disagree on content", - ) - .await; - - let resolution = resolver.resolve(&conflict.id, None).await.expect("TODO: handle error"); - - // LastWriterWins picks the last modality in the list - assert_eq!(resolution.winning_modality, Some(Modality::Vector)); - assert_eq!(resolution.policy_used, ConflictPolicy::LastWriterWins); - assert_eq!(resolution.resolver, "system"); - } - - // -- ModalityPriority tests ---------------------------------------------- - - #[tokio::test] - async fn test_modality_priority_resolution() { - let config = ConflictConfig { - default_policy: ConflictPolicy::ModalityPriority(vec![ - Modality::Semantic, - Modality::Document, - Modality::Graph, - Modality::Vector, - Modality::Tensor, - Modality::Temporal, - ]), - ..test_config() - }; - - let resolver = ConflictResolver::new(config); - - let conflict = resolver - .detect_conflict( - "entity-2", - vec![Modality::Vector, Modality::Document], - 0.5, - "Vector and document disagree", - ) - .await; - - let resolution = resolver.resolve(&conflict.id, None).await.expect("TODO: handle error"); - - // Document has higher priority than Vector in the custom order - assert_eq!(resolution.winning_modality, Some(Modality::Document)); - assert!(matches!( - resolution.policy_used, - ConflictPolicy::ModalityPriority(_) - )); - } - - #[tokio::test] - async fn test_modality_priority_with_custom_order() { - // Reversed priority: Tensor > Vector > Graph > Semantic > Document - let priority_order = vec![ - Modality::Tensor, - Modality::Vector, - Modality::Graph, - Modality::Semantic, - Modality::Document, - ]; - - let resolver = ConflictResolver::new(test_config()); - - let conflict = resolver - .detect_conflict( - "entity-3", - vec![Modality::Document, Modality::Vector], - 0.4, - "Document and vector disagree", - ) - .await; - - let resolution = resolver - .resolve( - &conflict.id, - Some(ConflictPolicy::ModalityPriority(priority_order)), - ) - .await - .expect("TODO: handle error"); - - // Vector is higher priority than Document in the reversed order - assert_eq!(resolution.winning_modality, Some(Modality::Vector)); - } - - // -- ManualResolve tests ------------------------------------------------- - - #[tokio::test] - async fn test_manual_resolve_creates_in_progress() { - let config = ConflictConfig { - default_policy: ConflictPolicy::ManualResolve, - ..test_config() - }; - - let resolver = ConflictResolver::new(config); - - let conflict = resolver - .detect_conflict( - "entity-4", - vec![Modality::Graph, Modality::Semantic], - 0.5, - "Graph and semantic disagree", - ) - .await; - - // Automatic resolve should fail with InvalidPolicy - let err = resolver.resolve(&conflict.id, None).await.unwrap_err(); - assert!(matches!(err, ConflictError::InvalidPolicy(_))); - - // Conflict should now be InProgress - let active = resolver.active_conflicts().await; - let found = active.iter().find(|c| c.id == conflict.id).expect("TODO: handle error"); - assert_eq!(found.status, ConflictStatus::InProgress); - } - - #[tokio::test] - async fn test_manual_resolve_succeeds() { - let config = ConflictConfig { - default_policy: ConflictPolicy::ManualResolve, - ..test_config() - }; - - let resolver = ConflictResolver::new(config); - - let conflict = resolver - .detect_conflict( - "entity-5", - vec![Modality::Graph, Modality::Semantic], - 0.5, - "Graph and semantic disagree", - ) - .await; - - // First try automatic (fails) - let _ = resolver.resolve(&conflict.id, None).await; - - // Then resolve manually - let resolution = resolver - .resolve_manual( - &conflict.id, - Modality::Semantic, - "admin-user-42", - Some("Semantic annotations are more recent".to_string()), - ) - .await - .expect("TODO: handle error"); - - assert_eq!(resolution.winning_modality, Some(Modality::Semantic)); - assert_eq!(resolution.resolver, "admin-user-42"); - assert!(resolution.notes.as_ref().expect("TODO: handle error").contains("more recent")); - assert_eq!(resolution.policy_used, ConflictPolicy::ManualResolve); - } - - // -- Auto-resolve below threshold ---------------------------------------- - - #[tokio::test] - async fn test_auto_resolve_below_threshold() { - let config = ConflictConfig { - auto_resolve_threshold: 0.5, - ..test_config() - }; - - let resolver = ConflictResolver::new(config); - - // Drift score 0.3, which is below auto_resolve_threshold 0.5 - let conflict = resolver - .detect_conflict( - "entity-6", - vec![Modality::Document, Modality::Tensor], - 0.3, - "Minor drift between document and tensor", - ) - .await; - - // Should resolve automatically - let resolution = resolver.resolve(&conflict.id, None).await.expect("TODO: handle error"); - assert_eq!(resolution.resolver, "system"); - - // Should be in history now, not active - let active = resolver.active_conflicts().await; - assert!(active.iter().all(|c| c.id != conflict.id)); - - let history = resolver.history(10).await; - assert!(history.iter().any(|c| c.id == conflict.id)); - } - - // -- Force manual above threshold ---------------------------------------- - - #[tokio::test] - async fn test_force_manual_above_threshold() { - let config = ConflictConfig { - require_manual_above: 0.8, - ..test_config() - }; - - let resolver = ConflictResolver::new(config); - - let conflict = resolver - .detect_conflict( - "entity-7", - vec![Modality::Document, Modality::Graph], - 0.9, - "Severe drift requiring manual intervention", - ) - .await; - - // Auto-resolve should fail because drift_score >= require_manual_above - let err = resolver.resolve(&conflict.id, None).await.unwrap_err(); - assert!(matches!(err, ConflictError::InvalidPolicy(_))); - assert!(err.to_string().contains("manual threshold")); - - // Conflict should be InProgress now - let active = resolver.active_conflicts().await; - let found = active.iter().find(|c| c.id == conflict.id).expect("TODO: handle error"); - assert_eq!(found.status, ConflictStatus::InProgress); - } - - // -- Dismiss conflict ---------------------------------------------------- - - #[tokio::test] - async fn test_dismiss_conflict() { - let resolver = ConflictResolver::new(test_config()); - - let conflict = resolver - .detect_conflict( - "entity-8", - vec![Modality::Vector, Modality::Tensor], - 0.4, - "False positive drift", - ) - .await; - - resolver - .dismiss(&conflict.id, "False positive -- data was updated concurrently") - .await - .expect("TODO: handle error"); - - // Should not be active - let active = resolver.active_conflicts().await; - assert!(active.iter().all(|c| c.id != conflict.id)); - - // Should be in history with Dismissed status - let history = resolver.history(10).await; - let found = history.iter().find(|c| c.id == conflict.id).expect("TODO: handle error"); - assert_eq!(found.status, ConflictStatus::Dismissed); - assert!(found - .resolution - .as_ref() - .expect("TODO: handle error") - .notes - .as_ref() - .expect("TODO: handle error") - .contains("False positive")); - } - - // -- History tracking ---------------------------------------------------- - - #[tokio::test] - async fn test_history_tracking() { - let resolver = ConflictResolver::new(test_config()); - - // Create and resolve three conflicts - for i in 0..3 { - let conflict = resolver - .detect_conflict( - &format!("entity-hist-{}", i), - vec![Modality::Document, Modality::Vector], - 0.5, - &format!("Conflict {}", i), - ) - .await; - resolver.resolve(&conflict.id, None).await.expect("TODO: handle error"); - } - - // Full history should have 3 entries - let all_history = resolver.history(0).await; - assert_eq!(all_history.len(), 3); - - // Limited history should return most recent first - let limited = resolver.history(2).await; - assert_eq!(limited.len(), 2); - assert_eq!(limited[0].entity_id, "entity-hist-2"); // most recent - assert_eq!(limited[1].entity_id, "entity-hist-1"); - } - - #[tokio::test] - async fn test_history_max_entries_eviction() { - let config = ConflictConfig { - max_history_entries: 3, - ..test_config() - }; - - let resolver = ConflictResolver::new(config); - - // Create and resolve 5 conflicts - for i in 0..5 { - let conflict = resolver - .detect_conflict( - &format!("entity-evict-{}", i), - vec![Modality::Document, Modality::Vector], - 0.5, - &format!("Conflict {}", i), - ) - .await; - resolver.resolve(&conflict.id, None).await.expect("TODO: handle error"); - } - - // History should be capped at 3 - let history = resolver.history(0).await; - assert_eq!(history.len(), 3); - - // Oldest entries (0 and 1) should have been evicted - let entity_ids: Vec<&str> = history.iter().map(|c| c.entity_id.as_str()).collect(); - assert!(!entity_ids.contains(&"entity-evict-0")); - assert!(!entity_ids.contains(&"entity-evict-1")); - assert!(entity_ids.contains(&"entity-evict-2")); - assert!(entity_ids.contains(&"entity-evict-3")); - assert!(entity_ids.contains(&"entity-evict-4")); - } - - // -- By-entity filtering ------------------------------------------------- - - #[tokio::test] - async fn test_by_entity_filtering() { - let resolver = ConflictResolver::new(test_config()); - - // Create conflicts for two entities - let c1 = resolver - .detect_conflict( - "entity-A", - vec![Modality::Document, Modality::Vector], - 0.5, - "Conflict on A", - ) - .await; - - resolver - .detect_conflict( - "entity-B", - vec![Modality::Graph, Modality::Semantic], - 0.4, - "Conflict on B", - ) - .await; - - let c3 = resolver - .detect_conflict( - "entity-A", - vec![Modality::Tensor, Modality::Temporal], - 0.6, - "Second conflict on A", - ) - .await; - - // Resolve one of entity-A's conflicts - resolver.resolve(&c1.id, None).await.expect("TODO: handle error"); - - // by_entity should return both A conflicts (1 resolved in history, 1 active) - let a_conflicts = resolver.by_entity("entity-A").await; - assert_eq!(a_conflicts.len(), 2); - - // Active should come first - assert_eq!(a_conflicts[0].id, c3.id); - assert_eq!(a_conflicts[0].status, ConflictStatus::Open); - - // Then historical - assert_eq!(a_conflicts[1].id, c1.id); - assert_eq!(a_conflicts[1].status, ConflictStatus::Resolved); - - // entity-B should have exactly 1 - let b_conflicts = resolver.by_entity("entity-B").await; - assert_eq!(b_conflicts.len(), 1); - } - - // -- Per-modality-pair policy override ----------------------------------- - - #[tokio::test] - async fn test_per_modality_pair_policy_override() { - let mut config = test_config(); - // Override: Document vs Vector conflicts use ModalityPriority - config.per_modality_policies.insert( - (Modality::Document, Modality::Vector), - ConflictPolicy::ModalityPriority(vec![ - Modality::Document, - Modality::Vector, - ]), - ); - - let resolver = ConflictResolver::new(config); - - let conflict = resolver - .detect_conflict( - "entity-pair", - vec![Modality::Vector, Modality::Document], - 0.5, - "Doc vs vector with pair override", - ) - .await; - - let resolution = resolver.resolve(&conflict.id, None).await.expect("TODO: handle error"); - - // Per-pair override should use ModalityPriority with Document winning - assert_eq!(resolution.winning_modality, Some(Modality::Document)); - assert!(matches!( - resolution.policy_used, - ConflictPolicy::ModalityPriority(_) - )); - } - - #[tokio::test] - async fn test_per_modality_pair_reverse_order_lookup() { - let mut config = test_config(); - // Register (Document, Vector) -- query with (Vector, Document) - config.per_modality_policies.insert( - (Modality::Document, Modality::Vector), - ConflictPolicy::AutoMerge, - ); - - // policy_for_pair should find it in either order - let policy = config.policy_for_pair(Modality::Vector, Modality::Document); - assert_eq!(*policy, ConflictPolicy::AutoMerge); - } - - // -- Already-resolved error ---------------------------------------------- - - #[tokio::test] - async fn test_already_resolved_error() { - let resolver = ConflictResolver::new(test_config()); - - let conflict = resolver - .detect_conflict( - "entity-dup", - vec![Modality::Document, Modality::Vector], - 0.5, - "Will be resolved twice", - ) - .await; - - // First resolve succeeds - resolver.resolve(&conflict.id, None).await.expect("TODO: handle error"); - - // Second resolve should fail -- conflict has moved to history - let err = resolver.resolve(&conflict.id, None).await.unwrap_err(); - assert!(matches!(err, ConflictError::NotFound(_))); - } - - // -- Not-found error ----------------------------------------------------- - - #[tokio::test] - async fn test_not_found_error() { - let resolver = ConflictResolver::new(test_config()); - - let err = resolver - .resolve("nonexistent-id", None) - .await - .unwrap_err(); - assert!(matches!(err, ConflictError::NotFound(_))); - assert!(err.to_string().contains("nonexistent-id")); - } - - #[tokio::test] - async fn test_manual_resolve_not_found() { - let resolver = ConflictResolver::new(test_config()); - - let err = resolver - .resolve_manual("nonexistent", Modality::Document, "admin", None) - .await - .unwrap_err(); - assert!(matches!(err, ConflictError::NotFound(_))); - } - - // -- Modality-not-in-conflict error -------------------------------------- - - #[tokio::test] - async fn test_modality_not_in_conflict_error() { - let resolver = ConflictResolver::new(test_config()); - - let conflict = resolver - .detect_conflict( - "entity-modal", - vec![Modality::Document, Modality::Vector], - 0.5, - "Document vs vector conflict", - ) - .await; - - // Try to resolve with a modality that is not part of the conflict - let err = resolver - .resolve_manual(&conflict.id, Modality::Tensor, "admin", None) - .await - .unwrap_err(); - assert!(matches!(err, ConflictError::ModalityNotInConflict(_))); - assert!(err.to_string().contains("tensor")); - } - - // -- Multiple conflicts on same entity ----------------------------------- - - #[tokio::test] - async fn test_multiple_conflicts_same_entity() { - let resolver = ConflictResolver::new(test_config()); - - let c1 = resolver - .detect_conflict( - "shared-entity", - vec![Modality::Document, Modality::Vector], - 0.4, - "First conflict", - ) - .await; - - let c2 = resolver - .detect_conflict( - "shared-entity", - vec![Modality::Graph, Modality::Semantic], - 0.6, - "Second conflict", - ) - .await; - - let c3 = resolver - .detect_conflict( - "shared-entity", - vec![Modality::Tensor, Modality::Temporal], - 0.5, - "Third conflict", - ) - .await; - - // All three should be active - let active = resolver.active_conflicts().await; - assert_eq!(active.len(), 3); - - // Resolve the first - resolver.resolve(&c1.id, None).await.expect("TODO: handle error"); - - // Now 2 active, 1 in history - let active = resolver.active_conflicts().await; - assert_eq!(active.len(), 2); - - let history = resolver.history(10).await; - assert_eq!(history.len(), 1); - - // by_entity should return all 3 - let by_entity = resolver.by_entity("shared-entity").await; - assert_eq!(by_entity.len(), 3); - } - - // -- Active conflicts listing -------------------------------------------- - - #[tokio::test] - async fn test_active_conflicts_listing() { - let resolver = ConflictResolver::new(test_config()); - - // Initially empty - assert!(resolver.active_conflicts().await.is_empty()); - - // Add two conflicts - let c1 = resolver - .detect_conflict( - "entity-list-1", - vec![Modality::Document, Modality::Vector], - 0.5, - "First", - ) - .await; - - let c2 = resolver - .detect_conflict( - "entity-list-2", - vec![Modality::Graph, Modality::Tensor], - 0.6, - "Second", - ) - .await; - - let active = resolver.active_conflicts().await; - assert_eq!(active.len(), 2); - assert_eq!(active[0].id, c1.id); - assert_eq!(active[1].id, c2.id); - - // Resolve one - resolver.resolve(&c1.id, None).await.expect("TODO: handle error"); - - let active = resolver.active_conflicts().await; - assert_eq!(active.len(), 1); - assert_eq!(active[0].id, c2.id); - } - - // -- JSON serialization round-trip --------------------------------------- - - #[tokio::test] - async fn test_json_serialization_roundtrip() { - let resolver = ConflictResolver::new(test_config()); - - let conflict = resolver - .detect_conflict( - "entity-json", - vec![Modality::Document, Modality::Semantic, Modality::Vector], - 0.55, - "Three-way conflict for serialization test", - ) - .await; - - // Serialize the conflict to JSON - let json = serde_json::to_string_pretty(&conflict).expect("TODO: handle error"); - - // Deserialize back - let deserialized: Conflict = serde_json::from_str(&json).expect("TODO: handle error"); - - assert_eq!(deserialized.id, conflict.id); - assert_eq!(deserialized.entity_id, "entity-json"); - assert_eq!(deserialized.conflicting_modalities.len(), 3); - assert!((deserialized.drift_score - 0.55).abs() < f64::EPSILON); - assert_eq!(deserialized.status, ConflictStatus::Open); - assert!(deserialized.resolution.is_none()); - - // Resolve and round-trip the resolution - let resolution = resolver.resolve(&conflict.id, None).await.expect("TODO: handle error"); - let resolution_json = serde_json::to_string_pretty(&resolution).expect("TODO: handle error"); - let deserialized_resolution: ConflictResolution = - serde_json::from_str(&resolution_json).expect("TODO: handle error"); - - assert_eq!( - deserialized_resolution.policy_used, - ConflictPolicy::LastWriterWins - ); - assert_eq!(deserialized_resolution.resolver, "system"); - } - - // -- ConflictConfig JSON round-trip -------------------------------------- - - #[test] - fn test_conflict_config_json_roundtrip() { - // Note: per_modality_policies uses a (Modality, Modality) tuple key which - // does not serialise to JSON directly. We test the round-trip with an - // empty map (the common case for JSON configs). Binary formats (bincode, - // postcard, CBOR) can handle tuple keys natively. - let config = test_config(); - - let json = serde_json::to_string_pretty(&config).expect("TODO: handle error"); - let deserialized: ConflictConfig = serde_json::from_str(&json).expect("TODO: handle error"); - - assert_eq!(deserialized.default_policy, config.default_policy); - assert!((deserialized.auto_resolve_threshold - config.auto_resolve_threshold).abs() - < f64::EPSILON); - assert!((deserialized.require_manual_above - config.require_manual_above).abs() - < f64::EPSILON); - assert_eq!(deserialized.max_history_entries, config.max_history_entries); - assert!(deserialized.per_modality_policies.is_empty()); - } - - // -- AutoMerge resolution ------------------------------------------------ - - #[tokio::test] - async fn test_auto_merge_no_single_winner() { - let config = ConflictConfig { - default_policy: ConflictPolicy::AutoMerge, - ..test_config() - }; - - let resolver = ConflictResolver::new(config); - - let conflict = resolver - .detect_conflict( - "entity-merge", - vec![Modality::Document, Modality::Semantic], - 0.5, - "Merge candidate", - ) - .await; - - let resolution = resolver.resolve(&conflict.id, None).await.expect("TODO: handle error"); - - // AutoMerge does not pick a single winner - assert!(resolution.winning_modality.is_none()); - assert_eq!(resolution.policy_used, ConflictPolicy::AutoMerge); - } - - // -- Dismiss already-resolved error -------------------------------------- - - #[tokio::test] - async fn test_dismiss_already_dismissed() { - let resolver = ConflictResolver::new(test_config()); - - let conflict = resolver - .detect_conflict( - "entity-dismiss-twice", - vec![Modality::Document, Modality::Vector], - 0.4, - "Will be dismissed", - ) - .await; - - resolver.dismiss(&conflict.id, "First dismissal").await.expect("TODO: handle error"); - - // Second dismissal should fail (not found -- moved to history) - let err = resolver - .dismiss(&conflict.id, "Second dismissal") - .await - .unwrap_err(); - assert!(matches!(err, ConflictError::NotFound(_))); - } - - // -- Display impls ------------------------------------------------------- - - #[test] - fn test_conflict_policy_display() { - assert_eq!( - format!("{}", ConflictPolicy::LastWriterWins), - "last_writer_wins" - ); - assert_eq!( - format!("{}", ConflictPolicy::ModalityPriority(vec![])), - "modality_priority" - ); - assert_eq!( - format!("{}", ConflictPolicy::ManualResolve), - "manual_resolve" - ); - assert_eq!(format!("{}", ConflictPolicy::AutoMerge), "auto_merge"); - assert_eq!( - format!("{}", ConflictPolicy::Custom("webhook".to_string())), - "custom(webhook)" - ); - } - - #[test] - fn test_conflict_status_display() { - assert_eq!(format!("{}", ConflictStatus::Open), "open"); - assert_eq!(format!("{}", ConflictStatus::InProgress), "in_progress"); - assert_eq!(format!("{}", ConflictStatus::Resolved), "resolved"); - assert_eq!(format!("{}", ConflictStatus::Dismissed), "dismissed"); - } - - // -- ConflictError display ----------------------------------------------- - - #[test] - fn test_conflict_error_display() { - let err = ConflictError::NotFound("abc-123".to_string()); - assert_eq!(err.to_string(), "Conflict not found: abc-123"); - - let err = ConflictError::AlreadyResolved("def-456".to_string()); - assert_eq!(err.to_string(), "Conflict already resolved: def-456"); - - let err = ConflictError::InvalidPolicy("bad policy".to_string()); - assert_eq!(err.to_string(), "Invalid policy: bad policy"); - - let err = ConflictError::ModalityNotInConflict("tensor".to_string()); - assert_eq!(err.to_string(), "Modality not in conflict: tensor"); - } -} diff --git a/verisimdb/rust-core/verisim-normalizer/src/lib.rs b/verisimdb/rust-core/verisim-normalizer/src/lib.rs deleted file mode 100644 index e409d2ae..00000000 --- a/verisimdb/rust-core/verisim-normalizer/src/lib.rs +++ /dev/null @@ -1,835 +0,0 @@ -// SPDX-License-Identifier: MPL-2.0 -//! VeriSim Normalizer -//! -//! Self-normalization engine that maintains cross-modal consistency. -//! When drift is detected, the normalizer orchestrates repairs. -//! -//! ## Modules -//! -//! - Root (`lib.rs`): The existing `Normalizer` engine with strategy-trait-based -//! dispatch (`NormalizationStrategy`, `SemanticVectorStrategy`, etc.). -//! - [`regeneration`]: Authority-ranked regeneration subsystem with configurable -//! strategies (`FromAuthoritative`, `Merge`, `UserResolve`), an audit event -//! trail, and a manual-resolution queue. -//! - [`conflict`]: Policy-based conflict resolution between modalities, with -//! configurable policies (last-writer-wins, modality-priority, manual-resolve, -//! auto-merge, custom), threshold-gated escalation, and full history tracking. - -#![allow(unused)] // Infrastructure code with planned future usage - -#![forbid(unsafe_code)] -pub mod conflict; -pub mod regeneration; -pub mod storage_regenerator; - -use async_trait::async_trait; -use chrono::{DateTime, Utc}; -use serde::{Deserialize, Serialize}; -use std::collections::HashMap; -use std::sync::Arc; -use thiserror::Error; -use tokio::sync::{mpsc, RwLock}; - -use verisim_drift::{DriftDetector, DriftEvent, DriftType}; -use verisim_octad::{Octad, OctadId, OctadStore}; - -/// Normalizer errors -#[derive(Error, Debug)] -pub enum NormalizerError { - #[error("Normalization failed for {entity_id}: {message}")] - NormalizationFailed { entity_id: String, message: String }, - - #[error("Strategy not found: {0}")] - StrategyNotFound(String), - - #[error("Octad error: {0}")] - OctadError(String), - - #[error("Channel error: {0}")] - ChannelError(String), - - #[error("Missing modality: {0}")] - MissingModality(String), - - #[error("Storage error: {0}")] - StorageError(String), - - #[error("No viable source: {0}")] - NoViableSource(String), -} - -/// Result of a normalization operation -#[derive(Debug, Clone, Serialize, Deserialize)] -pub struct NormalizationResult { - /// Entity that was normalized - pub entity_id: OctadId, - /// Type of normalization performed - pub normalization_type: NormalizationType, - /// Whether normalization succeeded - pub success: bool, - /// Changes made - pub changes: Vec<NormalizationChange>, - /// Duration of normalization - pub duration_ms: u64, - /// When normalization completed - pub completed_at: DateTime<Utc>, -} - -/// Types of normalization -#[derive(Debug, Clone, Copy, Serialize, Deserialize, PartialEq, Eq)] -pub enum NormalizationType { - /// Regenerate vector embedding from semantic content - VectorRegeneration, - /// Update graph from document analysis - GraphReconstruction, - /// Repair temporal consistency - TemporalRepair, - /// Synchronize tensor representation - TensorSync, - /// Full cross-modal reconciliation - FullReconciliation, -} - -/// A specific change made during normalization -#[derive(Debug, Clone, Serialize, Deserialize)] -pub struct NormalizationChange { - /// Modality affected - pub modality: String, - /// Field changed - pub field: String, - /// Previous value (if available) - pub old_value: Option<String>, - /// New value - pub new_value: String, - /// Reason for change - pub reason: String, -} - -/// Normalization strategy trait -#[async_trait] -pub trait NormalizationStrategy: Send + Sync { - /// Get strategy name - fn name(&self) -> &str; - - /// Check if this strategy applies to a drift type - fn applies_to(&self, drift_type: DriftType) -> bool; - - /// Perform normalization - async fn normalize( - &self, - octad: &Octad, - drift_event: &DriftEvent, - ) -> Result<NormalizationResult, NormalizerError>; -} - -/// Configuration for the normalizer -#[derive(Debug, Clone, Serialize, Deserialize)] -pub struct NormalizerConfig { - /// Whether to auto-normalize on drift detection - pub auto_normalize: bool, - /// Maximum concurrent normalizations - pub max_concurrent: usize, - /// Minimum drift score to trigger normalization - pub min_score: f64, - /// Backoff after failed normalization (seconds) - pub failure_backoff_secs: u64, -} - -impl Default for NormalizerConfig { - fn default() -> Self { - Self { - auto_normalize: true, - max_concurrent: 10, - min_score: 0.3, - failure_backoff_secs: 60, - } - } -} - -/// Status of the normalizer -#[derive(Debug, Clone, Serialize, Deserialize)] -pub struct NormalizerStatus { - /// Whether normalizer is running - pub running: bool, - /// Number of pending normalizations - pub pending_count: usize, - /// Number of active normalizations - pub active_count: usize, - /// Total normalizations completed - pub completed_count: u64, - /// Total failures - pub failure_count: u64, - /// Last normalization time - pub last_normalization: Option<DateTime<Utc>>, -} - -/// The main normalizer engine -pub struct Normalizer { - config: NormalizerConfig, - strategies: Arc<RwLock<Vec<Arc<dyn NormalizationStrategy>>>>, - #[allow(dead_code)] // Will be used for drift-based normalization triggers - drift_detector: Arc<DriftDetector>, - status: Arc<RwLock<NormalizerStatus>>, - result_sender: Option<mpsc::Sender<NormalizationResult>>, -} - -impl Normalizer { - /// Create a new normalizer - pub fn new(config: NormalizerConfig, drift_detector: Arc<DriftDetector>) -> Self { - Self { - config, - strategies: Arc::new(RwLock::new(Vec::new())), - drift_detector, - status: Arc::new(RwLock::new(NormalizerStatus { - running: false, - pending_count: 0, - active_count: 0, - completed_count: 0, - failure_count: 0, - last_normalization: None, - })), - result_sender: None, - } - } - - /// Create with default config - pub fn with_defaults(drift_detector: Arc<DriftDetector>) -> Self { - Self::new(NormalizerConfig::default(), drift_detector) - } - - /// Set result notification channel - pub fn with_result_channel(mut self, sender: mpsc::Sender<NormalizationResult>) -> Self { - self.result_sender = Some(sender); - self - } - - /// Register a normalization strategy - pub async fn register_strategy(&self, strategy: Arc<dyn NormalizationStrategy>) { - self.strategies.write().await.push(strategy); - } - - /// Handle a drift event - pub async fn handle_drift( - &self, - octad: &Octad, - event: &DriftEvent, - ) -> Result<Option<NormalizationResult>, NormalizerError> { - if event.score < self.config.min_score { - return Ok(None); - } - - // Find applicable strategy - let strategies = self.strategies.read().await; - let strategy = strategies - .iter() - .find(|s| s.applies_to(event.drift_type)) - .cloned(); - - let strategy = match strategy { - Some(s) => s, - None => return Ok(None), - }; - - // Update status - { - let mut status = self.status.write().await; - status.active_count += 1; - } - - // Perform normalization - let start = std::time::Instant::now(); - let result = strategy.normalize(octad, event).await; - - // Update status - { - let mut status = self.status.write().await; - status.active_count -= 1; - match &result { - Ok(_) => { - status.completed_count += 1; - status.last_normalization = Some(Utc::now()); - } - Err(_) => { - status.failure_count += 1; - } - } - } - - let result = result?; - - // Send notification - if let Some(ref sender) = self.result_sender { - sender - .send(result.clone()) - .await - .map_err(|e| NormalizerError::ChannelError(e.to_string()))?; - } - - Ok(Some(result)) - } - - /// Get current status - pub async fn status(&self) -> NormalizerStatus { - self.status.read().await.clone() - } - - /// Get registered strategies - pub async fn strategies(&self) -> Vec<String> { - self.strategies - .read() - .await - .iter() - .map(|s| s.name().to_string()) - .collect() - } -} - -/// Default strategy for semantic-vector drift -pub struct SemanticVectorStrategy; - -#[async_trait] -impl NormalizationStrategy for SemanticVectorStrategy { - fn name(&self) -> &str { - "semantic-vector-sync" - } - - fn applies_to(&self, drift_type: DriftType) -> bool { - matches!(drift_type, DriftType::SemanticVectorDrift) - } - - async fn normalize( - &self, - octad: &Octad, - drift_event: &DriftEvent, - ) -> Result<NormalizationResult, NormalizerError> { - let start = std::time::Instant::now(); - - // Identify authoritative source for vector regeneration - let has_document = octad.document.is_some(); - let has_semantic = octad.semantic.is_some(); - - if !has_document && !has_semantic { - return Err(NormalizerError::NormalizationFailed { - entity_id: octad.id.to_string(), - message: "Cannot regenerate vector: no document or semantic source available".into(), - }); - } - - let old_embedding_info = octad - .embedding - .as_ref() - .map(|e| format!("{}d vector", e.vector.len())) - .unwrap_or_else(|| "none".to_string()); - - let mut changes = Vec::new(); - - // Record what authoritative source will be used - if let Some(ref doc) = octad.document { - changes.push(NormalizationChange { - modality: "vector".to_string(), - field: "embedding".to_string(), - old_value: Some(old_embedding_info.clone()), - new_value: format!("regenerate from document '{}'", doc.title), - reason: format!( - "Semantic-vector drift score {:.3} — document is authoritative source", - drift_event.score - ), - }); - } - - if let Some(ref sem) = octad.semantic { - let type_summary = format!("{} semantic types", sem.types.len()); - changes.push(NormalizationChange { - modality: "vector".to_string(), - field: "embedding_context".to_string(), - old_value: Some(old_embedding_info), - new_value: format!("incorporate {}", type_summary), - reason: "Semantic annotations provide additional context for embedding".into(), - }); - } - - let duration_ms = start.elapsed().as_millis() as u64; - - Ok(NormalizationResult { - entity_id: octad.id.clone(), - normalization_type: NormalizationType::VectorRegeneration, - success: true, - changes, - duration_ms, - completed_at: Utc::now(), - }) - } -} - -/// Default strategy for graph-document drift -pub struct GraphDocumentStrategy; - -#[async_trait] -impl NormalizationStrategy for GraphDocumentStrategy { - fn name(&self) -> &str { - "graph-document-sync" - } - - fn applies_to(&self, drift_type: DriftType) -> bool { - matches!(drift_type, DriftType::GraphDocumentDrift) - } - - async fn normalize( - &self, - octad: &Octad, - drift_event: &DriftEvent, - ) -> Result<NormalizationResult, NormalizerError> { - let start = std::time::Instant::now(); - - let has_document = octad.document.is_some(); - let has_graph = octad.graph_node.is_some(); - - let mut changes = Vec::new(); - - match (has_document, has_graph) { - (true, true) => { - // Both present — graph needs reconstruction from document (document authoritative) - let doc = octad.document.as_ref().expect("TODO: handle error"); - let graph = octad.graph_node.as_ref().expect("TODO: handle error"); - changes.push(NormalizationChange { - modality: "graph".to_string(), - field: "relationships".to_string(), - old_value: Some(format!("graph node IRI: {}", graph.iri)), - new_value: format!("reconstruct from document '{}'", doc.title), - reason: format!( - "Graph-document drift score {:.3} — document is authoritative for content", - drift_event.score - ), - }); - } - (true, false) => { - // Only document — graph modality needs creation - let doc = octad.document.as_ref().expect("TODO: handle error"); - changes.push(NormalizationChange { - modality: "graph".to_string(), - field: "graph_node".to_string(), - old_value: None, - new_value: format!("create graph node from document '{}'", doc.title), - reason: "Graph modality missing — extract entities from document".into(), - }); - } - (false, true) => { - // Only graph — document modality needs creation - let graph = octad.graph_node.as_ref().expect("TODO: handle error"); - changes.push(NormalizationChange { - modality: "document".to_string(), - field: "document".to_string(), - old_value: None, - new_value: format!("create document from graph node '{}'", graph.local_name), - reason: "Document modality missing — generate from graph structure".into(), - }); - } - (false, false) => { - return Err(NormalizerError::NormalizationFailed { - entity_id: octad.id.to_string(), - message: "Cannot reconcile graph-document: neither modality present".into(), - }); - } - } - - let duration_ms = start.elapsed().as_millis() as u64; - - Ok(NormalizationResult { - entity_id: octad.id.clone(), - normalization_type: NormalizationType::GraphReconstruction, - success: true, - changes, - duration_ms, - completed_at: Utc::now(), - }) - } -} - -/// Strategy for tensor drift — regenerate tensor from vector embedding reshape or document TF-IDF -pub struct TensorRegenerationStrategy; - -#[async_trait] -impl NormalizationStrategy for TensorRegenerationStrategy { - fn name(&self) -> &str { - "tensor-regeneration" - } - - fn applies_to(&self, drift_type: DriftType) -> bool { - matches!(drift_type, DriftType::TensorDrift) - } - - async fn normalize( - &self, - octad: &Octad, - drift_event: &DriftEvent, - ) -> Result<NormalizationResult, NormalizerError> { - let start = std::time::Instant::now(); - let mut changes = Vec::new(); - - let has_vector = octad.embedding.is_some(); - let has_document = octad.document.is_some(); - let has_tensor = octad.tensor.is_some(); - - if !has_vector && !has_document { - return Err(NormalizerError::NormalizationFailed { - entity_id: octad.id.to_string(), - message: "Cannot regenerate tensor: no vector or document source available".into(), - }); - } - - let old_tensor_info = if has_tensor { - octad - .tensor - .as_ref() - .map(|t| format!("shape {:?}", t.shape)) - .unwrap_or_else(|| "present".to_string()) - } else { - "none".to_string() - }; - - // Prefer vector embedding as source for tensor regeneration (reshape) - if let Some(ref emb) = octad.embedding { - let dim = emb.vector.len(); - changes.push(NormalizationChange { - modality: "tensor".to_string(), - field: "data".to_string(), - old_value: Some(old_tensor_info.clone()), - new_value: format!("reshape {}d embedding to [1, {}] tensor", dim, dim), - reason: format!( - "Tensor drift score {:.3} — regenerating from vector embedding", - drift_event.score - ), - }); - } - - // If document is available, incorporate TF-IDF features - if let Some(ref doc) = octad.document { - changes.push(NormalizationChange { - modality: "tensor".to_string(), - field: "features".to_string(), - old_value: Some(old_tensor_info), - new_value: format!("compute TF-IDF features from document '{}'", doc.title), - reason: "Document content provides additional tensor features".into(), - }); - } - - let duration_ms = start.elapsed().as_millis() as u64; - - Ok(NormalizationResult { - entity_id: octad.id.clone(), - normalization_type: NormalizationType::TensorSync, - success: true, - changes, - duration_ms, - completed_at: Utc::now(), - }) - } -} - -/// Strategy for temporal drift — fix timestamp ordering, detect duplicates, fill gaps -pub struct TemporalRepairStrategy; - -#[async_trait] -impl NormalizationStrategy for TemporalRepairStrategy { - fn name(&self) -> &str { - "temporal-repair" - } - - fn applies_to(&self, drift_type: DriftType) -> bool { - matches!(drift_type, DriftType::TemporalConsistencyDrift) - } - - async fn normalize( - &self, - octad: &Octad, - drift_event: &DriftEvent, - ) -> Result<NormalizationResult, NormalizerError> { - let start = std::time::Instant::now(); - let mut changes = Vec::new(); - - // Timestamp ordering: ensure created_at <= modified_at - if octad.status.created_at > octad.status.modified_at { - changes.push(NormalizationChange { - modality: "temporal".to_string(), - field: "modified_at".to_string(), - old_value: Some(octad.status.modified_at.to_rfc3339()), - new_value: octad.status.created_at.to_rfc3339(), - reason: "modified_at was before created_at — correcting to created_at".into(), - }); - } - - // Version sanity: version should be >= 1 - if octad.status.version == 0 { - changes.push(NormalizationChange { - modality: "temporal".to_string(), - field: "version".to_string(), - old_value: Some("0".to_string()), - new_value: "1".to_string(), - reason: "Version was 0 — correcting to 1 (minimum valid version)".into(), - }); - } - - // Version count vs version consistency - if octad.version_count > 0 && octad.status.version > octad.version_count { - changes.push(NormalizationChange { - modality: "temporal".to_string(), - field: "version_count".to_string(), - old_value: Some(octad.version_count.to_string()), - new_value: octad.status.version.to_string(), - reason: format!( - "Temporal drift score {:.3} — version_count ({}) < version ({})", - drift_event.score, octad.version_count, octad.status.version - ), - }); - } - - // If no specific repairs found, record a general consistency check - if changes.is_empty() { - changes.push(NormalizationChange { - modality: "temporal".to_string(), - field: "consistency".to_string(), - old_value: None, - new_value: "verified".to_string(), - reason: format!( - "Temporal drift score {:.3} — checked ordering and duplicates, no repairs needed", - drift_event.score - ), - }); - } - - let duration_ms = start.elapsed().as_millis() as u64; - - Ok(NormalizationResult { - entity_id: octad.id.clone(), - normalization_type: NormalizationType::TemporalRepair, - success: true, - changes, - duration_ms, - completed_at: Utc::now(), - }) - } -} - -/// Strategy for quality drift — cascades all strategies in priority order -pub struct QualityReconciliationStrategy { - /// Inner strategies to cascade, in priority order - inner: Vec<Arc<dyn NormalizationStrategy>>, -} - -impl QualityReconciliationStrategy { - /// Create with the default cascade of all other strategies - pub fn new() -> Self { - Self { - inner: vec![ - Arc::new(SemanticVectorStrategy), - Arc::new(GraphDocumentStrategy), - Arc::new(TensorRegenerationStrategy), - Arc::new(TemporalRepairStrategy), - ], - } - } -} - -#[async_trait] -impl NormalizationStrategy for QualityReconciliationStrategy { - fn name(&self) -> &str { - "quality-reconciliation" - } - - fn applies_to(&self, drift_type: DriftType) -> bool { - matches!(drift_type, DriftType::QualityDrift | DriftType::SchemaDrift) - } - - async fn normalize( - &self, - octad: &Octad, - drift_event: &DriftEvent, - ) -> Result<NormalizationResult, NormalizerError> { - let start = std::time::Instant::now(); - let mut all_changes = Vec::new(); - let mut any_failure = false; - - // Cascade through all inner strategies - for strategy in &self.inner { - match strategy.normalize(octad, drift_event).await { - Ok(result) => { - all_changes.extend(result.changes); - } - Err(NormalizerError::NormalizationFailed { .. }) => { - // Individual strategy can't apply — skip it - continue; - } - Err(e) => { - any_failure = true; - all_changes.push(NormalizationChange { - modality: "quality".to_string(), - field: strategy.name().to_string(), - old_value: None, - new_value: format!("failed: {}", e), - reason: "Strategy error during quality reconciliation".into(), - }); - } - } - } - - let duration_ms = start.elapsed().as_millis() as u64; - - Ok(NormalizationResult { - entity_id: octad.id.clone(), - normalization_type: NormalizationType::FullReconciliation, - success: !any_failure, - changes: all_changes, - duration_ms, - completed_at: Utc::now(), - }) - } -} - -/// Create a normalizer with default strategies -pub async fn create_default_normalizer(drift_detector: Arc<DriftDetector>) -> Normalizer { - let normalizer = Normalizer::with_defaults(drift_detector); - normalizer - .register_strategy(Arc::new(SemanticVectorStrategy)) - .await; - normalizer - .register_strategy(Arc::new(GraphDocumentStrategy)) - .await; - normalizer - .register_strategy(Arc::new(TensorRegenerationStrategy)) - .await; - normalizer - .register_strategy(Arc::new(TemporalRepairStrategy)) - .await; - normalizer - .register_strategy(Arc::new(QualityReconciliationStrategy::new())) - .await; - normalizer -} - -#[cfg(test)] -mod tests { - use super::*; - use verisim_document::Document; - use verisim_drift::DriftThresholds; - use verisim_octad::{OctadStatus, ModalityStatus}; - use verisim_vector::Embedding; - - fn create_test_octad() -> Octad { - Octad { - id: OctadId::new("test-1"), - status: OctadStatus { - id: OctadId::new("test-1"), - created_at: Utc::now(), - modified_at: Utc::now(), - observed_at: None, - version: 1, - modality_status: ModalityStatus::default(), - }, - graph_node: None, - embedding: Some(Embedding::new("test-1", vec![0.1, 0.2, 0.3])), - tensor: None, - semantic: None, - document: Some(Document::new("test-1", "Test Document", "Test content for normalization")), - version_count: 1, - provenance_chain_length: 0, - spatial_data: None, - } - } - - fn create_empty_octad() -> Octad { - Octad { - id: OctadId::new("empty-1"), - status: OctadStatus { - id: OctadId::new("empty-1"), - created_at: Utc::now(), - modified_at: Utc::now(), - observed_at: None, - version: 1, - modality_status: ModalityStatus::default(), - }, - graph_node: None, - embedding: None, - tensor: None, - semantic: None, - document: None, - version_count: 1, - provenance_chain_length: 0, - spatial_data: None, - } - } - - #[tokio::test] - async fn test_normalizer_with_strategies() { - let drift_detector = Arc::new(DriftDetector::new(DriftThresholds::default())); - let normalizer = create_default_normalizer(drift_detector).await; - - let strategies = normalizer.strategies().await; - assert!(strategies.contains(&"semantic-vector-sync".to_string())); - assert!(strategies.contains(&"graph-document-sync".to_string())); - } - - #[tokio::test] - async fn test_handle_drift() { - let drift_detector = Arc::new(DriftDetector::new(DriftThresholds::default())); - let normalizer = create_default_normalizer(drift_detector).await; - - let octad = create_test_octad(); - let event = DriftEvent::new( - DriftType::SemanticVectorDrift, - 0.5, - "Test drift", - ); - - let result = normalizer.handle_drift(&octad, &event).await.expect("TODO: handle error"); - assert!(result.is_some()); - let result = result.expect("TODO: handle error"); - assert!(result.success); - assert!(!result.changes.is_empty()); - assert!(result.changes[0].new_value.contains("Test Document")); - } - - #[tokio::test] - async fn test_semantic_vector_strategy_empty_octad_errors() { - let strategy = SemanticVectorStrategy; - let octad = create_empty_octad(); - let event = DriftEvent::new( - DriftType::SemanticVectorDrift, - 0.8, - "Critical drift", - ); - - let result = strategy.normalize(&octad, &event).await; - assert!(result.is_err()); - } - - #[tokio::test] - async fn test_graph_document_strategy_empty_octad_errors() { - let strategy = GraphDocumentStrategy; - let octad = create_empty_octad(); - let event = DriftEvent::new( - DriftType::GraphDocumentDrift, - 0.8, - "Critical drift", - ); - - let result = strategy.normalize(&octad, &event).await; - assert!(result.is_err()); - } - - #[tokio::test] - async fn test_graph_document_strategy_with_document() { - let strategy = GraphDocumentStrategy; - let mut octad = create_test_octad(); - octad.graph_node = None; // Only document present - let event = DriftEvent::new( - DriftType::GraphDocumentDrift, - 0.5, - "Graph missing", - ); - - let result = strategy.normalize(&octad, &event).await.expect("TODO: handle error"); - assert!(result.success); - assert_eq!(result.changes[0].modality, "graph"); - assert!(result.changes[0].new_value.contains("Test Document")); - } -} diff --git a/verisimdb/rust-core/verisim-normalizer/src/regeneration.rs b/verisimdb/rust-core/verisim-normalizer/src/regeneration.rs deleted file mode 100644 index 5375301a..00000000 --- a/verisimdb/rust-core/verisim-normalizer/src/regeneration.rs +++ /dev/null @@ -1,1549 +0,0 @@ -// SPDX-License-Identifier: MPL-2.0 -//! Regeneration Strategies for the VeriSim Normalizer -//! -//! Implements authority-ranked cross-modal regeneration. When drift is detected -//! between modalities, this module decides *how* to repair the drifted modality -//! by consulting a configurable authority ranking and per-modality strategy -//! overrides. -//! -//! ## Authority Ranking (default, highest to lowest) -//! -//! 1. **Document** -- human-written content, most authoritative -//! 2. **Semantic** -- type annotations and proof contracts -//! 3. **Graph** -- structural relationships -//! 4. **Vector** -- computed embeddings -//! 5. **Tensor** -- derived computations -//! 6. **Temporal** -- version history (always consistent by construction) -//! -//! ## Strategies -//! -//! - `FromAuthoritative`: regenerate drifted modality from the highest-authority -//! consistent modality. -//! - `Merge`: combine data from all non-drifted modalities (weighted by authority) -//! to produce the best result. -//! - `UserResolve`: flag the entity for manual resolution and add it to the -//! normalization queue. - -use chrono::{DateTime, Utc}; -use serde::{Deserialize, Serialize}; -use std::collections::HashMap; -use std::fmt; -use std::sync::Arc; -use tokio::sync::RwLock; -use tracing::{debug, info, warn}; - -use verisim_octad::{Octad, OctadStore}; - -use crate::NormalizerError; - -// --------------------------------------------------------------------------- -// Modality enum (normalizer-local; avoids coupling to verisim-planner) -// --------------------------------------------------------------------------- - -/// The eight modalities of VeriSimDB (octad), ordered by default authority -/// ranking. -/// -/// The normalizer defines its own copy so it does not depend on the planner -/// crate. Conversion helpers exist for interop where necessary. -#[derive(Debug, Clone, Copy, PartialEq, Eq, Hash, Serialize, Deserialize)] -#[serde(rename_all = "lowercase")] -pub enum Modality { - Document, - Semantic, - Graph, - Vector, - Tensor, - Temporal, - Provenance, - Spatial, -} - -impl Modality { - /// Default authority order (highest to lowest): - /// - /// Document > Semantic > Provenance > Graph > Vector > Tensor > Spatial > Temporal - /// - /// Provenance is ranked high because lineage data is critical for audit - /// and compliance. Spatial is ranked low because coordinates are often - /// derived from other modalities. - pub const DEFAULT_AUTHORITY_ORDER: [Modality; 8] = [ - Modality::Document, - Modality::Semantic, - Modality::Provenance, - Modality::Graph, - Modality::Vector, - Modality::Tensor, - Modality::Spatial, - Modality::Temporal, - ]; - - /// All eight modalities (unordered — use `DEFAULT_AUTHORITY_ORDER` when - /// ordering matters). - pub const ALL: [Modality; 8] = Self::DEFAULT_AUTHORITY_ORDER; - - /// Check whether this modality is populated on a given octad. - pub fn is_present_on(self, octad: &Octad) -> bool { - match self { - Modality::Document => octad.document.is_some(), - Modality::Semantic => octad.semantic.is_some(), - Modality::Graph => octad.graph_node.is_some(), - Modality::Vector => octad.embedding.is_some(), - Modality::Tensor => octad.tensor.is_some(), - Modality::Temporal => octad.version_count > 0, - Modality::Provenance => octad.provenance_chain_length > 0, - Modality::Spatial => octad.spatial_data.is_some(), - } - } - - /// Extract a textual summary of this modality's data from a octad. - /// - /// Returns `None` if the modality is not populated. - pub fn summarize(self, octad: &Octad) -> Option<String> { - match self { - Modality::Document => octad.document.as_ref().map(|d| { - format!( - "document(title='{}', body_len={})", - d.title, - d.body.len() - ) - }), - Modality::Semantic => octad.semantic.as_ref().map(|s| { - format!( - "semantic(types={}, properties={})", - s.types.len(), - s.properties.len() - ) - }), - Modality::Graph => octad.graph_node.as_ref().map(|g| { - format!("graph(iri='{}', local_name='{}')", g.iri, g.local_name) - }), - Modality::Vector => octad.embedding.as_ref().map(|e| { - format!("vector(dim={})", e.vector.len()) - }), - Modality::Tensor => octad.tensor.as_ref().map(|t| { - format!("tensor(shape={:?}, len={})", t.shape, t.data.len()) - }), - Modality::Temporal => { - if octad.version_count > 0 { - Some(format!("temporal(versions={})", octad.version_count)) - } else { - None - } - } - Modality::Provenance => { - if octad.provenance_chain_length > 0 { - Some(format!("provenance(chain_length={})", octad.provenance_chain_length)) - } else { - None - } - } - Modality::Spatial => octad.spatial_data.as_ref().map(|s| { - format!( - "spatial(lat={}, lon={}, type={})", - s.coordinates.latitude, - s.coordinates.longitude, - s.geometry_type - ) - }), - } - } -} - -impl fmt::Display for Modality { - fn fmt(&self, f: &mut fmt::Formatter<'_>) -> fmt::Result { - match self { - Modality::Document => write!(f, "document"), - Modality::Semantic => write!(f, "semantic"), - Modality::Graph => write!(f, "graph"), - Modality::Vector => write!(f, "vector"), - Modality::Tensor => write!(f, "tensor"), - Modality::Temporal => write!(f, "temporal"), - Modality::Provenance => write!(f, "provenance"), - Modality::Spatial => write!(f, "spatial"), - } - } -} - -// --------------------------------------------------------------------------- -// RegenerationStrategy -// --------------------------------------------------------------------------- - -/// Strategy to apply when drift is detected on a modality. -#[derive(Debug, Clone, Copy, PartialEq, Eq, Hash, Serialize, Deserialize)] -pub enum RegenerationStrategy { - /// Regenerate the drifted modality from the highest-authority modality - /// that is currently consistent (not drifted). - FromAuthoritative, - - /// Combine data from *all* non-drifted modalities, weighted by their - /// authority rank, to produce the best possible repair. - Merge, - - /// Do not auto-fix. Flag the entity for manual resolution and place - /// it on the `NormalizationQueue`. - UserResolve, -} - -impl fmt::Display for RegenerationStrategy { - fn fmt(&self, f: &mut fmt::Formatter<'_>) -> fmt::Result { - match self { - RegenerationStrategy::FromAuthoritative => write!(f, "from_authoritative"), - RegenerationStrategy::Merge => write!(f, "merge"), - RegenerationStrategy::UserResolve => write!(f, "user_resolve"), - } - } -} - -// --------------------------------------------------------------------------- -// RegenerationConfig -// --------------------------------------------------------------------------- - -/// Configuration for the regeneration subsystem. -/// -/// This is separate from `NormalizerConfig` (which governs the top-level -/// normalizer engine) so that the two can evolve independently. -#[derive(Debug, Clone, Serialize, Deserialize)] -pub struct RegenerationConfig { - /// Global authority ranking (ordered highest to lowest). - /// - /// When `FromAuthoritative` is used, the first modality in this list that - /// is present and not drifted becomes the regeneration source. - pub authority_order: Vec<Modality>, - - /// Default strategy applied when drift is detected and no per-modality - /// override is configured. - pub default_strategy: RegenerationStrategy, - - /// Per-modality strategy overrides. - /// - /// If a modality appears in this map, its strategy takes precedence over - /// `default_strategy`. - pub modality_strategies: HashMap<Modality, RegenerationStrategy>, - - /// Drift score threshold (0.0 -- 1.0) above which regeneration is - /// triggered. Scores at or below this value are considered acceptable. - pub drift_threshold: f64, - - /// Maximum number of regenerations that may execute concurrently. - pub max_concurrent: usize, -} - -impl Default for RegenerationConfig { - fn default() -> Self { - Self { - authority_order: Modality::DEFAULT_AUTHORITY_ORDER.to_vec(), - default_strategy: RegenerationStrategy::FromAuthoritative, - modality_strategies: HashMap::new(), - drift_threshold: 0.3, - max_concurrent: 10, - } - } -} - -impl RegenerationConfig { - /// Look up the strategy for a specific modality. - /// - /// Returns the per-modality override if one exists, otherwise falls back - /// to `default_strategy`. - pub fn strategy_for(&self, modality: Modality) -> RegenerationStrategy { - self.modality_strategies - .get(&modality) - .copied() - .unwrap_or(self.default_strategy) - } - - /// Return the authority weight for a given modality. - /// - /// Weight is `N - index` where N is the length of `authority_order`, so - /// the first (most authoritative) modality has the highest weight. - /// Returns 0 if the modality is not in the ranking. - pub fn authority_weight(&self, modality: Modality) -> f64 { - let n = self.authority_order.len() as f64; - self.authority_order - .iter() - .position(|m| *m == modality) - .map(|idx| n - idx as f64) - .unwrap_or(0.0) - } -} - -// --------------------------------------------------------------------------- -// NormalizationEvent (audit trail) -// --------------------------------------------------------------------------- - -/// Record of a single regeneration action, for audit and observability. -#[derive(Debug, Clone, Serialize, Deserialize)] -pub struct NormalizationEvent { - /// ID of the octad entity that was (or would be) repaired. - pub entity_id: String, - - /// The modality that drifted and required regeneration. - pub drifted_modality: Modality, - - /// Which strategy was applied. - pub strategy_used: RegenerationStrategy, - - /// If `FromAuthoritative` was used, which modality served as the source. - pub source_modality: Option<Modality>, - - /// Drift score *before* regeneration was attempted. - pub pre_drift_score: f64, - - /// Drift score *after* regeneration completed, if validation was run. - pub post_drift_score: Option<f64>, - - /// When the event was recorded. - pub timestamp: DateTime<Utc>, - - /// Whether the regeneration succeeded. - pub success: bool, -} - -// --------------------------------------------------------------------------- -// RegenerationResult -// --------------------------------------------------------------------------- - -/// Outcome of a single regeneration attempt. -#[derive(Debug, Clone, Serialize, Deserialize)] -pub enum RegenerationResult { - /// Drift was repaired successfully. - Repaired { - /// Full audit event. - event: NormalizationEvent, - }, - - /// The entity was placed on the manual-resolution queue. - PendingResolution { - /// Entity that needs attention. - entity_id: String, - /// Human-readable explanation of why auto-repair was not attempted. - reason: String, - }, - - /// Drift score was at or below the configured threshold -- no action taken. - NoActionNeeded, - - /// Regeneration was attempted but failed. - Failed { - /// Description of what went wrong. - error: String, - }, -} - -// --------------------------------------------------------------------------- -// NormalizationQueue (for UserResolve strategy) -// --------------------------------------------------------------------------- - -/// An item waiting for manual resolution. -#[derive(Debug, Clone, Serialize, Deserialize)] -pub struct PendingNormalization { - /// The entity that needs manual attention. - pub entity_id: String, - - /// Which modality drifted. - pub drifted_modality: Modality, - - /// The drift score that triggered the queue entry. - pub drift_score: f64, - - /// When this item was added to the queue. - pub queued_at: DateTime<Utc>, -} - -/// Thread-safe queue of entities awaiting manual resolution. -/// -/// Entries are added when the `UserResolve` strategy is selected for a -/// drifted modality. External tooling (dashboards, CLI, etc.) drains the -/// queue after a human reviews each case. -#[derive(Debug, Clone)] -pub struct NormalizationQueue { - pending: Arc<RwLock<Vec<PendingNormalization>>>, -} - -impl NormalizationQueue { - /// Create an empty queue. - pub fn new() -> Self { - Self { - pending: Arc::new(RwLock::new(Vec::new())), - } - } - - /// Add an entity to the resolution queue. - pub async fn enqueue(&self, item: PendingNormalization) { - self.pending.write().await.push(item); - } - - /// Return all pending items (snapshot). - pub async fn pending(&self) -> Vec<PendingNormalization> { - self.pending.read().await.clone() - } - - /// Number of items waiting. - pub async fn len(&self) -> usize { - self.pending.read().await.len() - } - - /// Whether the queue is empty. - pub async fn is_empty(&self) -> bool { - self.pending.read().await.is_empty() - } - - /// Remove and return the item for a given entity + modality, if present. - /// - /// Used when a human resolves an item externally. - pub async fn resolve( - &self, - entity_id: &str, - modality: Modality, - ) -> Option<PendingNormalization> { - let mut pending = self.pending.write().await; - let idx = pending.iter().position(|p| { - p.entity_id == entity_id && p.drifted_modality == modality - }); - idx.map(|i| pending.remove(i)) - } - - /// Drain all items from the queue and return them. - pub async fn drain(&self) -> Vec<PendingNormalization> { - let mut pending = self.pending.write().await; - std::mem::take(&mut *pending) - } -} - -impl Default for NormalizationQueue { - fn default() -> Self { - Self::new() - } -} - -// --------------------------------------------------------------------------- -// RegenerationEngine -- the core pipeline -// --------------------------------------------------------------------------- - -/// The regeneration engine executes the full regeneration pipeline: -/// -/// 1. Check whether drift score exceeds threshold. -/// 2. Identify which modality drifted. -/// 3. Select strategy (per-modality override or default). -/// 4. Execute strategy. -/// 5. Validate (re-check drift score after regeneration). -/// 6. Record the event. -/// -/// Actual modality data extraction and re-computation are pluggable via the -/// `ModalityRegenerator` trait. The engine owns the *decision logic* and -/// pipeline orchestration. -pub struct RegenerationEngine { - /// Configuration controlling authority ranking, thresholds, and strategy - /// selection. - config: RegenerationConfig, - - /// Queue for items requiring manual resolution. - queue: NormalizationQueue, - - /// Audit trail of all regeneration events (most recent last). - events: Arc<RwLock<Vec<NormalizationEvent>>>, - - /// Pluggable regenerator for actually mutating modality data. - regenerator: Arc<dyn ModalityRegenerator>, -} - -/// Trait for pluggable modality-data regeneration. -/// -/// Implementations translate the *decision* (which source modality to use, -/// which target to regenerate) into actual data mutations. The engine calls -/// these methods; callers provide the implementation appropriate to their -/// storage backend. -#[async_trait::async_trait] -pub trait ModalityRegenerator: Send + Sync { - /// Regenerate `target` modality data from `source` modality data. - /// - /// Returns a textual summary of what changed (for the audit log). - async fn regenerate_from( - &self, - octad: &Octad, - source: Modality, - target: Modality, - ) -> Result<String, NormalizerError>; - - /// Merge data from `sources` (with associated weights) into `target`. - /// - /// Returns a textual summary of what changed. - async fn merge_into( - &self, - octad: &Octad, - sources: &[(Modality, f64)], - target: Modality, - ) -> Result<String, NormalizerError>; - - /// Re-measure drift score for `modality` after a regeneration. - /// - /// Returns the new score (0.0 = perfect, 1.0 = maximum drift). - async fn measure_drift( - &self, - octad: &Octad, - modality: Modality, - ) -> Result<f64, NormalizerError>; -} - -// --------------------------------------------------------------------------- -// Default (summary-only) regenerator -// --------------------------------------------------------------------------- - -/// A regenerator that does not mutate actual storage but produces descriptive -/// summaries of what *would* happen. -/// -/// Useful for dry-run mode, testing, and as the default before real storage -/// backends are wired in. -pub struct SummaryRegenerator; - -#[async_trait::async_trait] -impl ModalityRegenerator for SummaryRegenerator { - async fn regenerate_from( - &self, - octad: &Octad, - source: Modality, - target: Modality, - ) -> Result<String, NormalizerError> { - let source_summary = source - .summarize(octad) - .unwrap_or_else(|| format!("{} (empty)", source)); - Ok(format!( - "Regenerated {} from {} [{}]", - target, source, source_summary - )) - } - - async fn merge_into( - &self, - octad: &Octad, - sources: &[(Modality, f64)], - target: Modality, - ) -> Result<String, NormalizerError> { - let parts: Vec<String> = sources - .iter() - .map(|(m, w)| { - let s = m - .summarize(octad) - .unwrap_or_else(|| format!("{} (empty)", m)); - format!("{}(w={:.2}) [{}]", m, w, s) - }) - .collect(); - Ok(format!( - "Merged into {} from: {}", - target, - parts.join(", ") - )) - } - - async fn measure_drift( - &self, - _octad: &Octad, - _modality: Modality, - ) -> Result<f64, NormalizerError> { - // After a summary-only "regeneration" the drift hasn't actually changed, - // but we return 0.0 to signal a successful conceptual repair. - Ok(0.0) - } -} - -// --------------------------------------------------------------------------- -// RegenerationEngine implementation -// --------------------------------------------------------------------------- - -impl RegenerationEngine { - /// Create an engine with the given config and a summary-only regenerator. - pub fn new(config: RegenerationConfig) -> Self { - Self { - config, - queue: NormalizationQueue::new(), - events: Arc::new(RwLock::new(Vec::new())), - regenerator: Arc::new(SummaryRegenerator), - } - } - - /// Create an engine with a custom `ModalityRegenerator`. - pub fn with_regenerator( - config: RegenerationConfig, - regenerator: Arc<dyn ModalityRegenerator>, - ) -> Self { - Self { - config, - queue: NormalizationQueue::new(), - events: Arc::new(RwLock::new(Vec::new())), - regenerator, - } - } - - /// Create an engine with all defaults. - pub fn with_defaults() -> Self { - Self::new(RegenerationConfig::default()) - } - - /// Create an engine backed by a real OctadStore. - /// - /// Uses [`StorageRegenerator`] instead of the dry-run [`SummaryRegenerator`], - /// so regeneration operations read and write actual modality data. - pub fn with_store(config: RegenerationConfig, store: Arc<dyn OctadStore>) -> Self { - let regenerator = Arc::new( - crate::storage_regenerator::StorageRegenerator::new(store), - ); - Self::with_regenerator(config, regenerator) - } - - /// Create an engine with default config, backed by a real OctadStore. - pub fn with_store_defaults(store: Arc<dyn OctadStore>) -> Self { - Self::with_store(RegenerationConfig::default(), store) - } - - /// Access the underlying configuration. - pub fn config(&self) -> &RegenerationConfig { - &self.config - } - - /// Access the manual-resolution queue. - pub fn queue(&self) -> &NormalizationQueue { - &self.queue - } - - /// Return a snapshot of all recorded events. - pub async fn events(&self) -> Vec<NormalizationEvent> { - self.events.read().await.clone() - } - - // -- core pipeline ------------------------------------------------------- - - /// Execute the full regeneration pipeline for a single drifted modality. - /// - /// # Arguments - /// - /// * `octad` -- the entity whose modality drifted. - /// * `drifted_modality` -- which modality has drifted. - /// * `drift_score` -- the measured drift score (0.0 -- 1.0). - /// - /// # Returns - /// - /// A `RegenerationResult` describing what happened (repaired, queued for - /// human review, no action, or failure). - pub async fn regenerate( - &self, - octad: &Octad, - drifted_modality: Modality, - drift_score: f64, - ) -> RegenerationResult { - let entity_id = octad.id.to_string(); - - // Step 1: check threshold - if drift_score <= self.config.drift_threshold { - debug!( - entity_id = %entity_id, - modality = %drifted_modality, - score = drift_score, - threshold = self.config.drift_threshold, - "Drift score below threshold -- no action" - ); - return RegenerationResult::NoActionNeeded; - } - - // Step 2-3: select strategy - let strategy = self.config.strategy_for(drifted_modality); - info!( - entity_id = %entity_id, - modality = %drifted_modality, - score = drift_score, - strategy = %strategy, - "Starting regeneration" - ); - - // Step 4: execute strategy - match strategy { - RegenerationStrategy::FromAuthoritative => { - self.execute_from_authoritative(octad, drifted_modality, drift_score) - .await - } - RegenerationStrategy::Merge => { - self.execute_merge(octad, drifted_modality, drift_score) - .await - } - RegenerationStrategy::UserResolve => { - self.execute_user_resolve(octad, drifted_modality, drift_score) - .await - } - } - } - - // -- FromAuthoritative --------------------------------------------------- - - /// Find the highest-authority modality that is present on the octad and is - /// *not* the drifted modality, then regenerate from it. - async fn execute_from_authoritative( - &self, - octad: &Octad, - drifted_modality: Modality, - drift_score: f64, - ) -> RegenerationResult { - let entity_id = octad.id.to_string(); - - // Walk authority order to find the best source. - let source = self - .config - .authority_order - .iter() - .copied() - .find(|m| *m != drifted_modality && m.is_present_on(octad)); - - let source = match source { - Some(s) => s, - None => { - let msg = format!( - "No authoritative source available for {} on entity {}", - drifted_modality, entity_id - ); - warn!("{}", msg); - - // Record the failed attempt in the audit trail. - let event = NormalizationEvent { - entity_id: entity_id.clone(), - drifted_modality, - strategy_used: RegenerationStrategy::FromAuthoritative, - source_modality: None, - pre_drift_score: drift_score, - post_drift_score: None, - timestamp: Utc::now(), - success: false, - }; - self.record_event(event).await; - - return RegenerationResult::Failed { error: msg }; - } - }; - - info!( - entity_id = %entity_id, - source = %source, - target = %drifted_modality, - "FromAuthoritative: regenerating from source" - ); - - // Call the pluggable regenerator. - let regen_result = self - .regenerator - .regenerate_from(octad, source, drifted_modality) - .await; - - match regen_result { - Ok(summary) => { - debug!(summary = %summary, "Regeneration produced summary"); - - // Step 5: validate - let post_score = self - .regenerator - .measure_drift(octad, drifted_modality) - .await - .ok(); - - // Step 6: record - let event = NormalizationEvent { - entity_id: entity_id.clone(), - drifted_modality, - strategy_used: RegenerationStrategy::FromAuthoritative, - source_modality: Some(source), - pre_drift_score: drift_score, - post_drift_score: post_score, - timestamp: Utc::now(), - success: true, - }; - self.record_event(event.clone()).await; - - RegenerationResult::Repaired { event } - } - Err(e) => { - let event = NormalizationEvent { - entity_id: entity_id.clone(), - drifted_modality, - strategy_used: RegenerationStrategy::FromAuthoritative, - source_modality: Some(source), - pre_drift_score: drift_score, - post_drift_score: None, - timestamp: Utc::now(), - success: false, - }; - self.record_event(event).await; - - RegenerationResult::Failed { - error: e.to_string(), - } - } - } - } - - // -- Merge --------------------------------------------------------------- - - /// Collect data from all non-drifted modalities, weight by authority, and - /// merge into the drifted modality. - async fn execute_merge( - &self, - octad: &Octad, - drifted_modality: Modality, - drift_score: f64, - ) -> RegenerationResult { - let entity_id = octad.id.to_string(); - - // Collect non-drifted, present modalities with their weights. - let sources: Vec<(Modality, f64)> = self - .config - .authority_order - .iter() - .copied() - .filter(|m| *m != drifted_modality && m.is_present_on(octad)) - .map(|m| { - let weight = self.config.authority_weight(m); - (m, weight) - }) - .collect(); - - if sources.is_empty() { - let msg = format!( - "No source modalities available for merge into {} on entity {}", - drifted_modality, entity_id - ); - warn!("{}", msg); - - // Record the failed attempt in the audit trail. - let event = NormalizationEvent { - entity_id: entity_id.clone(), - drifted_modality, - strategy_used: RegenerationStrategy::Merge, - source_modality: None, - pre_drift_score: drift_score, - post_drift_score: None, - timestamp: Utc::now(), - success: false, - }; - self.record_event(event).await; - - return RegenerationResult::Failed { error: msg }; - } - - info!( - entity_id = %entity_id, - target = %drifted_modality, - source_count = sources.len(), - "Merge: combining sources" - ); - - let merge_result = self - .regenerator - .merge_into(octad, &sources, drifted_modality) - .await; - - match merge_result { - Ok(summary) => { - debug!(summary = %summary, "Merge produced summary"); - - let post_score = self - .regenerator - .measure_drift(octad, drifted_modality) - .await - .ok(); - - let event = NormalizationEvent { - entity_id: entity_id.clone(), - drifted_modality, - strategy_used: RegenerationStrategy::Merge, - source_modality: None, // merge uses multiple sources - pre_drift_score: drift_score, - post_drift_score: post_score, - timestamp: Utc::now(), - success: true, - }; - self.record_event(event.clone()).await; - - RegenerationResult::Repaired { event } - } - Err(e) => { - let event = NormalizationEvent { - entity_id: entity_id.clone(), - drifted_modality, - strategy_used: RegenerationStrategy::Merge, - source_modality: None, - pre_drift_score: drift_score, - post_drift_score: None, - timestamp: Utc::now(), - success: false, - }; - self.record_event(event).await; - - RegenerationResult::Failed { - error: e.to_string(), - } - } - } - } - - // -- UserResolve --------------------------------------------------------- - - /// Place the entity on the manual-resolution queue instead of auto-fixing. - async fn execute_user_resolve( - &self, - octad: &Octad, - drifted_modality: Modality, - drift_score: f64, - ) -> RegenerationResult { - let entity_id = octad.id.to_string(); - - let pending = PendingNormalization { - entity_id: entity_id.clone(), - drifted_modality, - drift_score, - queued_at: Utc::now(), - }; - - self.queue.enqueue(pending).await; - - let event = NormalizationEvent { - entity_id: entity_id.clone(), - drifted_modality, - strategy_used: RegenerationStrategy::UserResolve, - source_modality: None, - pre_drift_score: drift_score, - post_drift_score: None, - timestamp: Utc::now(), - success: true, // queueing itself succeeded - }; - self.record_event(event).await; - - info!( - entity_id = %entity_id, - modality = %drifted_modality, - score = drift_score, - "Entity queued for manual resolution" - ); - - RegenerationResult::PendingResolution { - entity_id, - reason: format!( - "{} drift (score {:.3}) requires manual resolution per strategy config", - drifted_modality, drift_score - ), - } - } - - // -- helpers ------------------------------------------------------------- - - /// Append an event to the audit trail. - async fn record_event(&self, event: NormalizationEvent) { - self.events.write().await.push(event); - } -} - -// =========================================================================== -// Tests -// =========================================================================== - -#[cfg(test)] -mod tests { - use super::*; - use chrono::Utc; - use verisim_document::Document; - use verisim_graph::GraphNode; - use verisim_octad::{OctadId, OctadStatus, ModalityStatus}; - use verisim_semantic::{Provenance, SemanticAnnotation}; - use verisim_vector::Embedding; - - // -- test helpers -------------------------------------------------------- - - /// Build a octad with document, semantic, graph, and vector populated. - fn rich_octad() -> Octad { - Octad { - id: OctadId::new("rich-1"), - status: OctadStatus { - id: OctadId::new("rich-1"), - created_at: Utc::now(), - modified_at: Utc::now(), - observed_at: None, - version: 1, - modality_status: ModalityStatus::default(), - }, - graph_node: Some(GraphNode::new("https://verisim.db/entity/rich-1")), - embedding: Some(Embedding::new("rich-1", vec![0.1, 0.2, 0.3])), - tensor: None, - semantic: Some(SemanticAnnotation { - entity_id: "rich-1".into(), - types: vec!["http://example.org/Document".into()], - properties: HashMap::new(), - provenance: Provenance::default(), - }), - document: Some(Document::new( - "rich-1", - "Rich Entity", - "Full content for normalizer testing", - )), - version_count: 3, - provenance_chain_length: 0, - spatial_data: None, - } - } - - /// Build a octad with only a document. - fn doc_only_octad() -> Octad { - Octad { - id: OctadId::new("doc-1"), - status: OctadStatus { - id: OctadId::new("doc-1"), - created_at: Utc::now(), - modified_at: Utc::now(), - observed_at: None, - version: 1, - modality_status: ModalityStatus::default(), - }, - graph_node: None, - embedding: None, - tensor: None, - semantic: None, - document: Some(Document::new("doc-1", "Doc Only", "Minimal entity")), - version_count: 0, - provenance_chain_length: 0, - spatial_data: None, - } - } - - /// Build a completely empty octad. - fn empty_octad() -> Octad { - Octad { - id: OctadId::new("empty-1"), - status: OctadStatus { - id: OctadId::new("empty-1"), - created_at: Utc::now(), - modified_at: Utc::now(), - observed_at: None, - version: 1, - modality_status: ModalityStatus::default(), - }, - graph_node: None, - embedding: None, - tensor: None, - semantic: None, - document: None, - version_count: 0, - provenance_chain_length: 0, - spatial_data: None, - } - } - - // -- authority order tests ----------------------------------------------- - - #[test] - fn test_default_authority_order_matches_spec() { - let order = Modality::DEFAULT_AUTHORITY_ORDER; - assert_eq!(order[0], Modality::Document, "Document should be rank 1"); - assert_eq!(order[1], Modality::Semantic, "Semantic should be rank 2"); - assert_eq!(order[2], Modality::Provenance, "Provenance should be rank 3"); - assert_eq!(order[3], Modality::Graph, "Graph should be rank 4"); - assert_eq!(order[4], Modality::Vector, "Vector should be rank 5"); - assert_eq!(order[5], Modality::Tensor, "Tensor should be rank 6"); - assert_eq!(order[6], Modality::Spatial, "Spatial should be rank 7"); - assert_eq!(order[7], Modality::Temporal, "Temporal should be rank 8"); - } - - #[test] - fn test_authority_weight_highest_first() { - let config = RegenerationConfig::default(); - let doc_w = config.authority_weight(Modality::Document); - let sem_w = config.authority_weight(Modality::Semantic); - let graph_w = config.authority_weight(Modality::Graph); - let vec_w = config.authority_weight(Modality::Vector); - let tensor_w = config.authority_weight(Modality::Tensor); - let temporal_w = config.authority_weight(Modality::Temporal); - - assert!(doc_w > sem_w, "Document weight > Semantic weight"); - assert!(sem_w > graph_w, "Semantic weight > Graph weight"); - assert!(graph_w > vec_w, "Graph weight > Vector weight"); - assert!(vec_w > tensor_w, "Vector weight > Tensor weight"); - assert!(tensor_w > temporal_w, "Tensor weight > Temporal weight"); - assert!(temporal_w > 0.0, "Temporal weight > 0"); - } - - // -- strategy selection tests -------------------------------------------- - - #[test] - fn test_strategy_for_default() { - let config = RegenerationConfig::default(); - assert_eq!( - config.strategy_for(Modality::Vector), - RegenerationStrategy::FromAuthoritative, - "Default strategy should be FromAuthoritative" - ); - } - - #[test] - fn test_strategy_for_per_modality_override() { - let mut config = RegenerationConfig::default(); - config - .modality_strategies - .insert(Modality::Tensor, RegenerationStrategy::UserResolve); - - assert_eq!( - config.strategy_for(Modality::Tensor), - RegenerationStrategy::UserResolve, - "Per-modality override should take precedence" - ); - assert_eq!( - config.strategy_for(Modality::Vector), - RegenerationStrategy::FromAuthoritative, - "Non-overridden modalities should still use default" - ); - } - - // -- modality presence tests --------------------------------------------- - - #[test] - fn test_modality_is_present_on_rich_octad() { - let h = rich_octad(); - assert!(Modality::Document.is_present_on(&h)); - assert!(Modality::Semantic.is_present_on(&h)); - assert!(Modality::Graph.is_present_on(&h)); - assert!(Modality::Vector.is_present_on(&h)); - assert!(!Modality::Tensor.is_present_on(&h)); - assert!(Modality::Temporal.is_present_on(&h)); // version_count > 0 - } - - #[test] - fn test_modality_is_present_on_empty_octad() { - let h = empty_octad(); - for m in Modality::ALL { - assert!( - !m.is_present_on(&h), - "{} should not be present on empty octad", - m - ); - } - } - - // -- FromAuthoritative tests --------------------------------------------- - - #[tokio::test] - async fn test_from_authoritative_selects_correct_source() { - let engine = RegenerationEngine::with_defaults(); - let h = rich_octad(); - - // Vector drifted -- Document is highest authority and is present. - let result = engine - .regenerate(&h, Modality::Vector, 0.8) - .await; - - match result { - RegenerationResult::Repaired { event } => { - assert_eq!(event.drifted_modality, Modality::Vector); - assert_eq!( - event.source_modality, - Some(Modality::Document), - "Should select Document as highest authority" - ); - assert_eq!(event.strategy_used, RegenerationStrategy::FromAuthoritative); - assert!(event.success); - assert!(event.pre_drift_score > 0.0); - assert!(event.post_drift_score.is_some()); - } - other => panic!("Expected Repaired, got {:?}", other), - } - } - - #[tokio::test] - async fn test_from_authoritative_skips_drifted_modality() { - let engine = RegenerationEngine::with_defaults(); - let h = rich_octad(); - - // Document itself drifted -- next authority is Semantic. - let result = engine - .regenerate(&h, Modality::Document, 0.7) - .await; - - match result { - RegenerationResult::Repaired { event } => { - assert_eq!(event.drifted_modality, Modality::Document); - assert_eq!( - event.source_modality, - Some(Modality::Semantic), - "Should skip Document (drifted) and use Semantic" - ); - } - other => panic!("Expected Repaired, got {:?}", other), - } - } - - #[tokio::test] - async fn test_from_authoritative_fails_no_source() { - let engine = RegenerationEngine::with_defaults(); - let h = empty_octad(); - - let result = engine - .regenerate(&h, Modality::Vector, 0.9) - .await; - - match result { - RegenerationResult::Failed { error } => { - assert!( - error.contains("No authoritative source"), - "Error should mention missing sources: {}", - error - ); - } - other => panic!("Expected Failed, got {:?}", other), - } - } - - #[tokio::test] - async fn test_from_authoritative_doc_only_regenerates_graph() { - let engine = RegenerationEngine::with_defaults(); - let h = doc_only_octad(); - - // Graph drifted, only Document is available. - let result = engine - .regenerate(&h, Modality::Graph, 0.6) - .await; - - match result { - RegenerationResult::Repaired { event } => { - assert_eq!(event.source_modality, Some(Modality::Document)); - assert_eq!(event.drifted_modality, Modality::Graph); - } - other => panic!("Expected Repaired, got {:?}", other), - } - } - - // -- Merge tests --------------------------------------------------------- - - #[tokio::test] - async fn test_merge_combines_multiple_sources() { - let mut config = RegenerationConfig::default(); - config.default_strategy = RegenerationStrategy::Merge; - - let engine = RegenerationEngine::new(config); - let h = rich_octad(); - - let result = engine - .regenerate(&h, Modality::Tensor, 0.5) - .await; - - match result { - RegenerationResult::Repaired { event } => { - assert_eq!(event.strategy_used, RegenerationStrategy::Merge); - assert_eq!(event.drifted_modality, Modality::Tensor); - // source_modality is None for merge (uses multiple) - assert!(event.source_modality.is_none()); - assert!(event.success); - } - other => panic!("Expected Repaired, got {:?}", other), - } - } - - #[tokio::test] - async fn test_merge_fails_no_sources() { - let mut config = RegenerationConfig::default(); - config.default_strategy = RegenerationStrategy::Merge; - - let engine = RegenerationEngine::new(config); - let h = empty_octad(); - - let result = engine - .regenerate(&h, Modality::Vector, 0.9) - .await; - - match result { - RegenerationResult::Failed { error } => { - assert!(error.contains("No source modalities")); - } - other => panic!("Expected Failed, got {:?}", other), - } - } - - // -- UserResolve tests --------------------------------------------------- - - #[tokio::test] - async fn test_user_resolve_adds_to_queue() { - let mut config = RegenerationConfig::default(); - config - .modality_strategies - .insert(Modality::Tensor, RegenerationStrategy::UserResolve); - - let engine = RegenerationEngine::new(config); - let h = rich_octad(); - - assert!(engine.queue().is_empty().await); - - let result = engine - .regenerate(&h, Modality::Tensor, 0.7) - .await; - - match result { - RegenerationResult::PendingResolution { entity_id, reason } => { - assert_eq!(entity_id, "rich-1"); - assert!(reason.contains("manual resolution")); - } - other => panic!("Expected PendingResolution, got {:?}", other), - } - - assert_eq!(engine.queue().len().await, 1); - let pending = engine.queue().pending().await; - assert_eq!(pending[0].entity_id, "rich-1"); - assert_eq!(pending[0].drifted_modality, Modality::Tensor); - assert!((pending[0].drift_score - 0.7).abs() < f64::EPSILON); - } - - #[tokio::test] - async fn test_user_resolve_queue_drain() { - let mut config = RegenerationConfig::default(); - config.default_strategy = RegenerationStrategy::UserResolve; - - let engine = RegenerationEngine::new(config); - let h = rich_octad(); - - engine.regenerate(&h, Modality::Vector, 0.5).await; - engine.regenerate(&h, Modality::Graph, 0.6).await; - - assert_eq!(engine.queue().len().await, 2); - - let drained = engine.queue().drain().await; - assert_eq!(drained.len(), 2); - assert!(engine.queue().is_empty().await); - } - - #[tokio::test] - async fn test_user_resolve_queue_selective_resolve() { - let mut config = RegenerationConfig::default(); - config.default_strategy = RegenerationStrategy::UserResolve; - - let engine = RegenerationEngine::new(config); - let h = rich_octad(); - - engine.regenerate(&h, Modality::Vector, 0.5).await; - engine.regenerate(&h, Modality::Graph, 0.6).await; - - // Resolve only the graph item. - let resolved = engine - .queue() - .resolve("rich-1", Modality::Graph) - .await; - assert!(resolved.is_some()); - assert_eq!(resolved.expect("TODO: handle error").drifted_modality, Modality::Graph); - - // Only vector should remain. - assert_eq!(engine.queue().len().await, 1); - let remaining = engine.queue().pending().await; - assert_eq!(remaining[0].drifted_modality, Modality::Vector); - } - - // -- threshold tests ----------------------------------------------------- - - #[tokio::test] - async fn test_drift_below_threshold_no_action() { - let config = RegenerationConfig { - drift_threshold: 0.5, - ..Default::default() - }; - let engine = RegenerationEngine::new(config); - let h = rich_octad(); - - let result = engine - .regenerate(&h, Modality::Vector, 0.3) - .await; - assert!( - matches!(result, RegenerationResult::NoActionNeeded), - "Score 0.3 should be below threshold 0.5" - ); - } - - #[tokio::test] - async fn test_drift_at_threshold_no_action() { - let config = RegenerationConfig { - drift_threshold: 0.5, - ..Default::default() - }; - let engine = RegenerationEngine::new(config); - let h = rich_octad(); - - let result = engine - .regenerate(&h, Modality::Vector, 0.5) - .await; - assert!( - matches!(result, RegenerationResult::NoActionNeeded), - "Score exactly at threshold should be no-action (requires > threshold)" - ); - } - - #[tokio::test] - async fn test_drift_above_threshold_triggers_action() { - let config = RegenerationConfig { - drift_threshold: 0.5, - ..Default::default() - }; - let engine = RegenerationEngine::new(config); - let h = rich_octad(); - - let result = engine - .regenerate(&h, Modality::Vector, 0.51) - .await; - assert!( - matches!(result, RegenerationResult::Repaired { .. }), - "Score 0.51 should trigger action with threshold 0.5" - ); - } - - // -- event recording tests ----------------------------------------------- - - #[tokio::test] - async fn test_normalization_event_records_correctly() { - let engine = RegenerationEngine::with_defaults(); - let h = rich_octad(); - - assert!(engine.events().await.is_empty()); - - engine - .regenerate(&h, Modality::Vector, 0.8) - .await; - - let events = engine.events().await; - assert_eq!(events.len(), 1); - - let event = &events[0]; - assert_eq!(event.entity_id, "rich-1"); - assert_eq!(event.drifted_modality, Modality::Vector); - assert_eq!(event.strategy_used, RegenerationStrategy::FromAuthoritative); - assert_eq!(event.source_modality, Some(Modality::Document)); - assert!((event.pre_drift_score - 0.8).abs() < f64::EPSILON); - assert!(event.post_drift_score.is_some()); - assert!(event.success); - } - - #[tokio::test] - async fn test_failed_regeneration_records_event() { - let engine = RegenerationEngine::with_defaults(); - let h = empty_octad(); - - engine - .regenerate(&h, Modality::Vector, 0.9) - .await; - - let events = engine.events().await; - assert_eq!(events.len(), 1); - assert!(!events[0].success); - } - - #[tokio::test] - async fn test_multiple_regenerations_accumulate_events() { - let engine = RegenerationEngine::with_defaults(); - let h = rich_octad(); - - engine.regenerate(&h, Modality::Vector, 0.5).await; - engine.regenerate(&h, Modality::Graph, 0.6).await; - engine.regenerate(&h, Modality::Semantic, 0.7).await; - - let events = engine.events().await; - assert_eq!(events.len(), 3); - assert_eq!(events[0].drifted_modality, Modality::Vector); - assert_eq!(events[1].drifted_modality, Modality::Graph); - assert_eq!(events[2].drifted_modality, Modality::Semantic); - } - - // -- custom authority order tests ---------------------------------------- - - #[tokio::test] - async fn test_custom_authority_order() { - // Reverse authority: Temporal is highest. - let config = RegenerationConfig { - authority_order: vec![ - Modality::Temporal, - Modality::Tensor, - Modality::Vector, - Modality::Graph, - Modality::Semantic, - Modality::Document, - ], - ..Default::default() - }; - - let engine = RegenerationEngine::new(config); - let h = rich_octad(); // has temporal (version_count=3) - - let result = engine - .regenerate(&h, Modality::Document, 0.8) - .await; - - match result { - RegenerationResult::Repaired { event } => { - assert_eq!( - event.source_modality, - Some(Modality::Temporal), - "Custom order should make Temporal the highest authority" - ); - } - other => panic!("Expected Repaired, got {:?}", other), - } - } - - // -- modality summarize tests -------------------------------------------- - - #[test] - fn test_modality_summarize() { - let h = rich_octad(); - let doc_summary = Modality::Document.summarize(&h); - assert!(doc_summary.is_some()); - assert!(doc_summary.expect("TODO: handle error").contains("Rich Entity")); - - let tensor_summary = Modality::Tensor.summarize(&h); - assert!(tensor_summary.is_none(), "Tensor not populated"); - } - - // -- Display impls ------------------------------------------------------- - - #[test] - fn test_modality_display() { - assert_eq!(format!("{}", Modality::Document), "document"); - assert_eq!(format!("{}", Modality::Semantic), "semantic"); - assert_eq!(format!("{}", Modality::Graph), "graph"); - assert_eq!(format!("{}", Modality::Vector), "vector"); - assert_eq!(format!("{}", Modality::Tensor), "tensor"); - assert_eq!(format!("{}", Modality::Temporal), "temporal"); - } - - #[test] - fn test_strategy_display() { - assert_eq!( - format!("{}", RegenerationStrategy::FromAuthoritative), - "from_authoritative" - ); - assert_eq!(format!("{}", RegenerationStrategy::Merge), "merge"); - assert_eq!( - format!("{}", RegenerationStrategy::UserResolve), - "user_resolve" - ); - } - - // -- RegenerationConfig edge cases --------------------------------------- - - #[test] - fn test_authority_weight_for_unlisted_modality() { - // Config with only 3 modalities in authority order. - let config = RegenerationConfig { - authority_order: vec![Modality::Document, Modality::Semantic, Modality::Graph], - ..Default::default() - }; - - assert_eq!( - config.authority_weight(Modality::Vector), - 0.0, - "Unlisted modality should have weight 0" - ); - assert!(config.authority_weight(Modality::Document) > 0.0); - } - - #[test] - fn test_default_config_values() { - let config = RegenerationConfig::default(); - assert_eq!(config.drift_threshold, 0.3); - assert_eq!(config.max_concurrent, 10); - assert_eq!(config.default_strategy, RegenerationStrategy::FromAuthoritative); - assert!(config.modality_strategies.is_empty()); - assert_eq!(config.authority_order.len(), 8); - } -} diff --git a/verisimdb/rust-core/verisim-normalizer/src/storage_regenerator.rs b/verisimdb/rust-core/verisim-normalizer/src/storage_regenerator.rs deleted file mode 100644 index 6991ae28..00000000 --- a/verisimdb/rust-core/verisim-normalizer/src/storage_regenerator.rs +++ /dev/null @@ -1,682 +0,0 @@ -// SPDX-License-Identifier: MPL-2.0 -// Copyright (c) 2026 Jonathan D.A. Jewell (hyperpolymath) <j.d.a.jewell@open.ac.uk> -// -// StorageRegenerator — Real storage-backed ModalityRegenerator implementation -// -// Unlike the dry-run SummaryRegenerator, this implementation reads actual -// modality data from the OctadStore and writes back regenerated content. -// -// Regeneration strategies per source→target pair: -// -// Document → Vector: Hash document text into deterministic embedding -// Document → Semantic: Extract keywords + metadata as type annotations -// Document → Graph: Extract entity mentions as graph triples -// Semantic → Vector: Serialize annotations, hash to embedding -// Semantic → Document: Render annotation tree as document body -// Graph → Document: Serialize graph triples as document body -// Graph → Semantic: Extract node types as semantic annotations -// Vector → (any): Vectors are derived; reverse regeneration uses -// nearest-neighbour lookup to infer source content -// Provenance → (any): Provenance is append-only; never regenerated FROM -// other modalities, only records events -// Temporal → (any): Temporal is consistent-by-construction; low priority -// Spatial → (any): Spatial rarely drifts; regeneration copies coordinates - -use std::collections::HashMap; -use std::sync::Arc; - -use async_trait::async_trait; -use chrono::Utc; -use tracing::{debug, info, warn}; - -use verisim_octad::{ - Octad, OctadDocumentInput, OctadGraphInput, OctadInput, OctadProvenanceInput, - OctadSemanticInput, OctadStore, OctadVectorInput, SemanticAnnotation, -}; - -use crate::NormalizerError; -use crate::regeneration::{Modality, ModalityRegenerator}; - -/// A regenerator that reads and writes real modality data via an OctadStore. -/// -/// This replaces the SummaryRegenerator for production use. Each regeneration -/// operation: -/// 1. Reads source modality data from the octad -/// 2. Computes the target modality using deterministic transformations -/// 3. Writes the updated entity back to the store -/// 4. Returns a human-readable summary for the audit log -pub struct StorageRegenerator { - store: Arc<dyn OctadStore>, -} - -impl StorageRegenerator { - /// Create a new StorageRegenerator backed by the given OctadStore. - pub fn new(store: Arc<dyn OctadStore>) -> Self { - Self { store } - } - - // ─── Internal: source → target transformations ──────────────────── - - /// Extract text content from an octad's document modality. - fn document_text(octad: &Octad) -> Option<String> { - octad.document.as_ref().map(|doc| { - format!("{} {}", doc.title, doc.body) - }) - } - - /// Compute a deterministic embedding from text. - /// - /// Uses a simple hash-based approach (FNV-1a on sliding windows) to - /// produce a fixed-dimension vector. This is NOT a real semantic - /// embedding — it's a content fingerprint that changes when the source - /// text changes, enabling drift detection. - /// - /// For production semantic search, replace this with an external - /// embedding model call (e.g., Sentence-BERT via HTTP). - fn text_to_embedding(text: &str, dim: usize) -> Vec<f32> { - let mut embedding = vec![0.0f32; dim]; - if text.is_empty() { - return embedding; - } - - // FNV-1a hash on overlapping 3-grams, distributed across dimensions - let bytes = text.as_bytes(); - let window_size = 3.min(bytes.len()); - for i in 0..=(bytes.len().saturating_sub(window_size)) { - let window = &bytes[i..i + window_size]; - let mut hash: u64 = 0xcbf29ce484222325; // FNV offset basis - for &b in window { - hash ^= b as u64; - hash = hash.wrapping_mul(0x100000001b3); // FNV prime - } - let idx = (hash as usize) % dim; - embedding[idx] += 1.0; - } - - // L2 normalise to unit vector - let norm: f32 = embedding.iter().map(|x| x * x).sum::<f32>().sqrt(); - if norm > 0.0 { - for v in &mut embedding { - *v /= norm; - } - } - - embedding - } - - /// Extract keywords from document text for semantic annotations. - /// - /// Simple TF-based extraction: split on whitespace, count frequency, - /// return the top N words longer than 3 characters. - fn extract_keywords(text: &str, max_keywords: usize) -> Vec<String> { - let mut freq: HashMap<String, usize> = HashMap::new(); - for word in text.split_whitespace() { - let clean: String = word - .chars() - .filter(|c| c.is_alphanumeric()) - .collect::<String>() - .to_lowercase(); - if clean.len() > 3 { - *freq.entry(clean).or_insert(0) += 1; - } - } - let mut pairs: Vec<_> = freq.into_iter().collect(); - pairs.sort_by(|a, b| b.1.cmp(&a.1)); - pairs.into_iter().take(max_keywords).map(|(w, _)| w).collect() - } - - /// Build a semantic annotation from keywords. - fn keywords_to_semantic(keywords: &[String]) -> SemanticAnnotation { - SemanticAnnotation { - entity_id: String::new(), - types: keywords - .iter() - .map(|k| format!("keyword:{}", k)) - .collect(), - properties: HashMap::new(), - provenance: Default::default(), - } - } - - /// Cosine similarity between two embeddings. - fn cosine_similarity(a: &[f32], b: &[f32]) -> f64 { - if a.len() != b.len() || a.is_empty() { - return 0.0; - } - let dot: f64 = a.iter().zip(b).map(|(x, y)| (*x as f64) * (*y as f64)).sum(); - let norm_a: f64 = a.iter().map(|x| (*x as f64) * (*x as f64)).sum::<f64>().sqrt(); - let norm_b: f64 = b.iter().map(|x| (*x as f64) * (*x as f64)).sum::<f64>().sqrt(); - if norm_a == 0.0 || norm_b == 0.0 { - return 0.0; - } - dot / (norm_a * norm_b) - } - - /// Write updated modality data back to the store. - async fn write_back( - &self, - octad: &Octad, - input: OctadInput, - ) -> Result<(), NormalizerError> { - self.store - .update(&octad.id, input) - .await - .map_err(|e| NormalizerError::StorageError(format!("{}", e)))?; - Ok(()) - } -} - -#[async_trait] -impl ModalityRegenerator for StorageRegenerator { - async fn regenerate_from( - &self, - octad: &Octad, - source: Modality, - target: Modality, - ) -> Result<String, NormalizerError> { - info!( - entity_id = %octad.id, - source = %source, - target = %target, - "StorageRegenerator: regenerating modality" - ); - - match (source, target) { - // ── Document as source ────────────────────────────────── - (Modality::Document, Modality::Vector) => { - let text = Self::document_text(octad) - .ok_or_else(|| NormalizerError::MissingModality("Document".into()))?; - let embedding = Self::text_to_embedding(&text, 384); - let input = OctadInput { - vector: Some(OctadVectorInput { - embedding, - model: Some("fnv1a-trigram-384".to_string()), - }), - ..Default::default() - }; - self.write_back(octad, input).await?; - Ok(format!( - "Regenerated Vector (dim={}) from Document (len={})", - 384, - text.len() - )) - } - (Modality::Document, Modality::Semantic) => { - let text = Self::document_text(octad) - .ok_or_else(|| NormalizerError::MissingModality("Document".into()))?; - let keywords = Self::extract_keywords(&text, 10); - let semantic = Self::keywords_to_semantic(&keywords); - let input = OctadInput { - semantic: Some(OctadSemanticInput { - types: semantic.types.clone(), - properties: HashMap::new(), - }), - ..Default::default() - }; - self.write_back(octad, input).await?; - Ok(format!( - "Regenerated Semantic ({} types) from Document (len={})", - semantic.types.len(), - text.len() - )) - } - (Modality::Document, Modality::Graph) => { - let text = Self::document_text(octad) - .ok_or_else(|| NormalizerError::MissingModality("Document".into()))?; - let keywords = Self::extract_keywords(&text, 5); - let relationships: Vec<(String, String)> = keywords - .iter() - .map(|k| ("mentions".to_string(), format!("keyword:{}", k))) - .collect(); - let input = OctadInput { - graph: Some(OctadGraphInput { - relationships: relationships.clone(), - }), - ..Default::default() - }; - self.write_back(octad, input).await?; - Ok(format!( - "Regenerated Graph ({} edges) from Document", - relationships.len() - )) - } - - // ── Semantic as source ────────────────────────────────── - (Modality::Semantic, Modality::Vector) => { - let semantic = octad - .semantic - .as_ref() - .ok_or_else(|| NormalizerError::MissingModality("Semantic".into()))?; - let text = semantic.types.join(" "); - let embedding = Self::text_to_embedding(&text, 384); - let input = OctadInput { - vector: Some(OctadVectorInput { - embedding, - model: Some("fnv1a-trigram-384".to_string()), - }), - ..Default::default() - }; - self.write_back(octad, input).await?; - Ok(format!( - "Regenerated Vector from Semantic ({} types)", - semantic.types.len() - )) - } - (Modality::Semantic, Modality::Document) => { - let semantic = octad - .semantic - .as_ref() - .ok_or_else(|| NormalizerError::MissingModality("Semantic".into()))?; - let body = format!( - "Types: {}\nProvenance: {:?}", - semantic.types.join(", "), - semantic.provenance, - ); - let input = OctadInput { - document: Some(OctadDocumentInput { - title: "[regenerated from semantic]".to_string(), - body: body.clone(), - fields: HashMap::new(), - }), - ..Default::default() - }; - self.write_back(octad, input).await?; - Ok(format!( - "Regenerated Document (len={}) from Semantic", - body.len() - )) - } - - // ── Graph as source ───────────────────────────────────── - (Modality::Graph, Modality::Document) => { - let graph = octad - .graph_node - .as_ref() - .ok_or_else(|| NormalizerError::MissingModality("Graph".into()))?; - let body = format!( - "Node: {} ({})", - graph.iri, - graph.local_name, - ); - let input = OctadInput { - document: Some(OctadDocumentInput { - title: "[regenerated from graph]".to_string(), - body: body.clone(), - fields: HashMap::new(), - }), - ..Default::default() - }; - self.write_back(octad, input).await?; - Ok(format!( - "Regenerated Document (len={}) from Graph", - body.len(), - )) - } - (Modality::Graph, Modality::Semantic) => { - let graph = octad - .graph_node - .as_ref() - .ok_or_else(|| NormalizerError::MissingModality("Graph".into()))?; - // GraphNode has no types field; extract a type from the IRI - let types = vec![format!("graph:{}", graph.iri)]; - let input = OctadInput { - semantic: Some(OctadSemanticInput { - types: types.clone(), - properties: HashMap::new(), - }), - ..Default::default() - }; - self.write_back(octad, input).await?; - Ok(format!( - "Regenerated Semantic ({} types) from Graph", - types.len() - )) - } - - // ── Fallback for unimplemented pairs ──────────────────── - (src, tgt) => { - warn!( - source = %src, - target = %tgt, - "StorageRegenerator: no specific transformation for this pair; using summary" - ); - let source_summary = src - .summarize(octad) - .unwrap_or_else(|| format!("{} (empty)", src)); - Ok(format!( - "Passthrough regeneration: {} from {} [{}]", - tgt, src, source_summary - )) - } - } - } - - async fn merge_into( - &self, - octad: &Octad, - sources: &[(Modality, f64)], - target: Modality, - ) -> Result<String, NormalizerError> { - info!( - entity_id = %octad.id, - target = %target, - sources = sources.len(), - "StorageRegenerator: merging modalities" - ); - - match target { - Modality::Vector => { - // Merge: weighted average of embeddings from all source modalities - // that can produce embeddings. - let dim = 384; - let mut merged = vec![0.0f32; dim]; - let mut total_weight = 0.0f64; - - for (modality, weight) in sources { - let text = match modality { - Modality::Document => Self::document_text(octad), - Modality::Semantic => octad - .semantic - .as_ref() - .map(|s| s.types.join(" ")), - Modality::Graph => octad - .graph_node - .as_ref() - .map(|g| { - format!("{} {}", g.iri, g.local_name) - }), - _ => None, - }; - - if let Some(text) = text { - let emb = Self::text_to_embedding(&text, dim); - for (i, v) in emb.iter().enumerate() { - merged[i] += v * (*weight as f32); - } - total_weight += weight; - } - } - - // Normalise the weighted sum - if total_weight > 0.0 { - let norm: f32 = merged.iter().map(|x| x * x).sum::<f32>().sqrt(); - if norm > 0.0 { - for v in &mut merged { - *v /= norm; - } - } - } - - let input = OctadInput { - vector: Some(OctadVectorInput { - embedding: merged, - model: Some("fnv1a-trigram-384-merged".to_string()), - }), - ..Default::default() - }; - self.write_back(octad, input).await?; - Ok(format!( - "Merged {} sources (total weight={:.2}) into Vector (dim={})", - sources.len(), - total_weight, - dim - )) - } - - Modality::Semantic => { - // Merge: union of all type annotations from sources - let mut all_types: Vec<String> = Vec::new(); - for (modality, _weight) in sources { - match modality { - Modality::Document => { - if let Some(text) = Self::document_text(octad) { - let kw = Self::extract_keywords(&text, 5); - all_types - .extend(kw.iter().map(|k| format!("keyword:{}", k))); - } - } - Modality::Graph => { - if let Some(g) = &octad.graph_node { - all_types.push(format!("graph:{}", g.iri)); - } - } - Modality::Semantic => { - if let Some(s) = &octad.semantic { - all_types.extend(s.types.clone()); - } - } - _ => {} - } - } - all_types.sort(); - all_types.dedup(); - - let input = OctadInput { - semantic: Some(OctadSemanticInput { - types: all_types.clone(), - properties: HashMap::new(), - }), - ..Default::default() - }; - self.write_back(octad, input).await?; - Ok(format!( - "Merged {} sources into Semantic ({} types)", - sources.len(), - all_types.len() - )) - } - - // For other targets, delegate to the highest-weighted source - _ => { - if let Some((best_source, _)) = sources - .iter() - .max_by(|a, b| a.1.partial_cmp(&b.1).unwrap_or(std::cmp::Ordering::Equal)) - { - debug!( - target = %target, - best_source = %best_source, - "Merge fallback: using highest-weighted source" - ); - self.regenerate_from(octad, *best_source, target).await - } else { - Err(NormalizerError::NoViableSource(format!( - "No sources provided for merge into {}", - target - ))) - } - } - } - } - - async fn measure_drift( - &self, - octad: &Octad, - modality: Modality, - ) -> Result<f64, NormalizerError> { - // Measure drift by recomputing what the modality SHOULD be - // and comparing against what it IS. - match modality { - Modality::Vector => { - // Compare stored embedding against freshly computed one - if let (Some(text), Some(stored)) = - (Self::document_text(octad), octad.embedding.as_ref()) - { - let expected = Self::text_to_embedding(&text, stored.vector.len()); - let sim = Self::cosine_similarity(&expected, &stored.vector); - // Drift = 1.0 - similarity (0.0 = identical, 1.0 = orthogonal) - Ok((1.0 - sim).max(0.0)) - } else { - // Can't measure — assume moderate drift - Ok(0.5) - } - } - Modality::Semantic => { - // Compare stored types against what we'd extract from document - if let (Some(text), Some(semantic)) = - (Self::document_text(octad), octad.semantic.as_ref()) - { - let expected_kw = Self::extract_keywords(&text, 10); - let expected_types: std::collections::HashSet<_> = expected_kw - .iter() - .map(|k| format!("keyword:{}", k)) - .collect(); - let stored_types: std::collections::HashSet<_> = - semantic.types.iter().cloned().collect(); - - if expected_types.is_empty() && stored_types.is_empty() { - return Ok(0.0); - } - - let intersection = expected_types.intersection(&stored_types).count(); - let union = expected_types.union(&stored_types).count(); - let jaccard = if union > 0 { - intersection as f64 / union as f64 - } else { - 0.0 - }; - Ok((1.0 - jaccard).max(0.0)) - } else { - Ok(0.5) - } - } - Modality::Document => { - // Document is usually the source of truth; drift = 0 if present - if octad.document.is_some() { - Ok(0.0) - } else { - Ok(1.0) // Missing document is maximum drift - } - } - Modality::Graph => { - // Check graph consistency: node should exist - if octad.graph_node.is_some() { - Ok(0.0) // Graph node present - } else { - Ok(0.8) // Missing graph - } - } - Modality::Provenance => { - // Provenance is append-only; drift = 0 if chain exists - if octad.provenance_chain_length > 0 { - Ok(0.0) - } else { - Ok(0.6) // No provenance chain - } - } - Modality::Temporal => { - // Temporal is consistent by construction - if octad.version_count > 0 { - Ok(0.0) - } else { - Ok(0.3) - } - } - Modality::Spatial => { - // Spatial drift is rare; check presence - if octad.spatial_data.is_some() { - Ok(0.0) - } else { - Ok(0.5) - } - } - Modality::Tensor => { - // Tensor drift measured against vector coherence - Ok(0.1) // Tensor usually tracks vector closely - } - } - } -} - -#[cfg(test)] -mod tests { - use super::*; - use verisim_octad::SemanticAnnotation; - - #[test] - fn test_text_to_embedding_deterministic() { - let e1 = StorageRegenerator::text_to_embedding("hello world", 64); - let e2 = StorageRegenerator::text_to_embedding("hello world", 64); - assert_eq!(e1, e2, "Same text should produce identical embeddings"); - } - - #[test] - fn test_text_to_embedding_different_text() { - let e1 = StorageRegenerator::text_to_embedding("hello world", 64); - let e2 = StorageRegenerator::text_to_embedding("goodbye moon", 64); - assert_ne!(e1, e2, "Different text should produce different embeddings"); - } - - #[test] - fn test_text_to_embedding_unit_vector() { - let e = StorageRegenerator::text_to_embedding("test content for embedding", 128); - let norm: f32 = e.iter().map(|x| x * x).sum::<f32>().sqrt(); - assert!( - (norm - 1.0).abs() < 0.001, - "Embedding should be unit-normalised, got norm={}", - norm - ); - } - - #[test] - fn test_text_to_embedding_empty() { - let e = StorageRegenerator::text_to_embedding("", 64); - assert!( - e.iter().all(|&v| v == 0.0), - "Empty text should produce zero vector" - ); - } - - #[test] - fn test_extract_keywords() { - let text = "the quick brown fox jumps over the lazy dog repeatedly"; - let kw = StorageRegenerator::extract_keywords(text, 5); - assert!(!kw.is_empty(), "Should extract at least one keyword"); - assert!(kw.len() <= 5, "Should respect max_keywords limit"); - // Short words (<=3 chars) should be filtered - assert!( - !kw.contains(&"the".to_string()), - "'the' should be filtered (too short)" - ); - assert!( - !kw.contains(&"fox".to_string()), - "'fox' should be filtered (too short)" - ); - } - - #[test] - fn test_keywords_to_semantic() { - let kw = vec!["server".to_string(), "config".to_string()]; - let sem = StorageRegenerator::keywords_to_semantic(&kw); - assert_eq!(sem.types.len(), 2); - assert_eq!(sem.types[0], "keyword:server"); - assert_eq!(sem.types[1], "keyword:config"); - } - - #[test] - fn test_cosine_similarity_identical() { - let a = vec![1.0, 0.0, 0.0]; - let b = vec![1.0, 0.0, 0.0]; - let sim = StorageRegenerator::cosine_similarity(&a, &b); - assert!( - (sim - 1.0).abs() < 0.001, - "Identical vectors should have similarity=1.0" - ); - } - - #[test] - fn test_cosine_similarity_orthogonal() { - let a = vec![1.0, 0.0, 0.0]; - let b = vec![0.0, 1.0, 0.0]; - let sim = StorageRegenerator::cosine_similarity(&a, &b); - assert!( - sim.abs() < 0.001, - "Orthogonal vectors should have similarity=0.0" - ); - } - - #[test] - fn test_cosine_similarity_empty() { - let sim = StorageRegenerator::cosine_similarity(&[], &[]); - assert_eq!(sim, 0.0); - } -} diff --git a/verisimdb/rust-core/verisim-octad/Cargo.toml b/verisimdb/rust-core/verisim-octad/Cargo.toml deleted file mode 100644 index fd5afdd5..00000000 --- a/verisimdb/rust-core/verisim-octad/Cargo.toml +++ /dev/null @@ -1,33 +0,0 @@ -# SPDX-License-Identifier: MPL-2.0 - -[package] -name = "verisim-octad" -description = "Octad entity - one entity, eight synchronized representations (octad)" -version.workspace = true -edition.workspace = true -authors.workspace = true -license.workspace = true - -[dependencies] -verisim-graph = { path = "../verisim-graph" } -verisim-vector = { path = "../verisim-vector" } -verisim-tensor = { path = "../verisim-tensor" } -verisim-semantic = { path = "../verisim-semantic" } -verisim-document = { path = "../verisim-document" } -verisim-temporal = { path = "../verisim-temporal" } -verisim-provenance = { path = "../verisim-provenance" } -verisim-spatial = { path = "../verisim-spatial" } -verisim-wal = { path = "../verisim-wal" } - -serde.workspace = true -serde_json.workspace = true -chrono.workspace = true -thiserror.workspace = true -tracing.workspace = true -async-trait.workspace = true -tokio.workspace = true -uuid.workspace = true - -[dev-dependencies] -proptest.workspace = true -tempfile = "3" diff --git a/verisimdb/rust-core/verisim-octad/src/lib.rs b/verisimdb/rust-core/verisim-octad/src/lib.rs deleted file mode 100644 index f293ee16..00000000 --- a/verisimdb/rust-core/verisim-octad/src/lib.rs +++ /dev/null @@ -1,520 +0,0 @@ -// SPDX-License-Identifier: MPL-2.0 -//! VeriSim Octad Entity -//! -//! One entity, eight synchronized representations (the octad). -//! The Octad is the fundamental unit of VeriSimDB — each entity exists -//! simultaneously across all eight modalities, maintaining cross-modal -//! consistency: Graph, Vector, Tensor, Semantic, Document, Temporal, -//! Provenance, and Spatial. - -#![forbid(unsafe_code)] -use async_trait::async_trait; -use chrono::{DateTime, Utc}; -use serde::{Deserialize, Serialize}; -use std::collections::HashMap; -use thiserror::Error; - -// Re-export modality types — all eight modalities -pub use verisim_document::{Document, DocumentStore}; -pub use verisim_graph::{GraphEdge, GraphNode, GraphObject, GraphStore}; -pub use verisim_provenance::{ - InMemoryProvenanceStore, ProvenanceChain, ProvenanceError, ProvenanceEventType, - ProvenanceRecord, ProvenanceStore, -}; -pub use verisim_semantic::{ProofBlob, Provenance, SemanticAnnotation, SemanticStore, SemanticType, SemanticValue}; -pub use verisim_spatial::{ - BoundingBox, Coordinates, GeometryType, InMemorySpatialStore, SpatialData, - SpatialSearchResult, SpatialStore, -}; -pub use verisim_tensor::{Tensor, TensorStore}; -pub use verisim_temporal::{TemporalStore, TimeRange, Version}; -pub use verisim_vector::{Embedding, VectorStore}; - -// In-memory store implementation -mod store; -pub use store::{OctadSnapshot, InMemoryOctadStore}; - -// Homoiconicity: queries as octads -pub mod query_octad; -pub use query_octad::{QueryOctadBuilder, QueryExecution}; - -// Optional RAM promotion for acceleration (disabled by default) -pub mod ram_promotion; -pub use ram_promotion::{PromotionManager, Modality, PromotionDecision, PromotionEvent}; - -// ACID transaction manager for cross-modality atomicity -pub mod transaction; -pub use transaction::{IsolationLevel, LockType, TransactionManager, TransactionError, TransactionState}; - -// WAL types (re-exported for external use) -pub use verisim_wal::{SyncMode, WalEntry, WalModality, WalOperation, WalWriter}; - -/// Octad errors -#[derive(Error, Debug)] -pub enum OctadError { - #[error("Entity not found: {0}")] - NotFound(String), - - #[error("Modality error in {modality}: {message}")] - ModalityError { modality: String, message: String }, - - #[error("Consistency violation: {0}")] - ConsistencyViolation(String), - - #[error("Validation error: {0}")] - ValidationError(String), -} - -/// Unique identifier for a Octad entity -#[derive(Debug, Clone, Serialize, Deserialize, PartialEq, Eq, Hash)] -pub struct OctadId(pub String); - -impl OctadId { - /// Create a new Octad ID - pub fn new(id: impl Into<String>) -> Self { - Self(id.into()) - } - - /// Generate a new UUID-based ID - pub fn generate() -> Self { - Self(uuid::Uuid::new_v4().to_string()) - } - - /// Get the ID as a string reference - pub fn as_str(&self) -> &str { - &self.0 - } - - /// Convert to IRI for graph modality - pub fn to_iri(&self, base: &str) -> String { - format!("{}/{}", base.trim_end_matches('/'), self.0) - } -} - -impl std::fmt::Display for OctadId { - fn fmt(&self, f: &mut std::fmt::Formatter<'_>) -> std::fmt::Result { - write!(f, "{}", self.0) - } -} - -impl From<String> for OctadId { - fn from(s: String) -> Self { - Self(s) - } -} - -impl From<&str> for OctadId { - fn from(s: &str) -> Self { - Self(s.to_string()) - } -} - -/// Status of a Octad entity across modalities -#[derive(Debug, Clone, Serialize, Deserialize)] -pub struct OctadStatus { - /// Entity ID - pub id: OctadId, - /// When the entity was created (ingestion time, set by the database) - pub created_at: DateTime<Utc>, - /// When last modified (ingestion time) - pub modified_at: DateTime<Utc>, - /// Caller-supplied real-world observation time, distinct from `created_at`. - /// `created_at` records when the entity entered the database; `observed_at` - /// records when the underlying event happened in the territory the entity - /// represents (e.g. an email's `Date:` header). Optional because not every - /// entity has a meaningful real-world timestamp. - #[serde(default, skip_serializing_if = "Option::is_none")] - pub observed_at: Option<DateTime<Utc>>, - /// Current version - pub version: u64, - /// Status per modality - pub modality_status: ModalityStatus, -} - -/// Status of each modality for an entity (octad: 8 modalities) -#[derive(Debug, Clone, Serialize, Deserialize, Default)] -pub struct ModalityStatus { - pub graph: bool, - pub vector: bool, - pub tensor: bool, - pub semantic: bool, - pub document: bool, - pub temporal: bool, - pub provenance: bool, - pub spatial: bool, -} - -impl ModalityStatus { - /// Check if all eight modalities are populated - pub fn is_complete(&self) -> bool { - self.graph - && self.vector - && self.tensor - && self.semantic - && self.document - && self.temporal - && self.provenance - && self.spatial - } - - /// Get list of missing modalities - pub fn missing(&self) -> Vec<&'static str> { - let mut missing = Vec::new(); - if !self.graph { missing.push("graph"); } - if !self.vector { missing.push("vector"); } - if !self.tensor { missing.push("tensor"); } - if !self.semantic { missing.push("semantic"); } - if !self.document { missing.push("document"); } - if !self.temporal { missing.push("temporal"); } - if !self.provenance { missing.push("provenance"); } - if !self.spatial { missing.push("spatial"); } - missing - } -} - -/// Input data for creating/updating a Octad -#[derive(Debug, Clone, Serialize, Deserialize, Default)] -pub struct OctadInput { - /// Graph relationships (optional) - pub graph: Option<OctadGraphInput>, - /// Vector embedding (optional) - pub vector: Option<OctadVectorInput>, - /// Tensor data (optional) - pub tensor: Option<OctadTensorInput>, - /// Semantic annotations (optional) - pub semantic: Option<OctadSemanticInput>, - /// Document content (optional) - pub document: Option<OctadDocumentInput>, - /// Temporal observation time (optional). When supplied, `OctadStatus.observed_at` - /// is populated. Distinct from the version snapshot (which the database always - /// writes at ingestion time): this is the territory's clock, not the database's. - pub temporal: Option<OctadTemporalInput>, - /// Provenance event (optional) - pub provenance: Option<OctadProvenanceInput>, - /// Spatial coordinates (optional) - pub spatial: Option<OctadSpatialInput>, - /// Additional metadata - pub metadata: HashMap<String, String>, -} - - -/// Graph modality input -#[derive(Debug, Clone, Serialize, Deserialize)] -pub struct OctadGraphInput { - /// Outgoing relationships - pub relationships: Vec<(String, String)>, // (predicate, target_id) -} - -/// Vector modality input -#[derive(Debug, Clone, Serialize, Deserialize)] -pub struct OctadVectorInput { - /// Embedding vector - pub embedding: Vec<f32>, - /// Embedding model used - pub model: Option<String>, -} - -/// Tensor modality input -#[derive(Debug, Clone, Serialize, Deserialize)] -pub struct OctadTensorInput { - /// Tensor shape - pub shape: Vec<usize>, - /// Tensor data - pub data: Vec<f64>, -} - -/// Semantic modality input -#[derive(Debug, Clone, Serialize, Deserialize)] -pub struct OctadSemanticInput { - /// Type IRIs - pub types: Vec<String>, - /// Properties - pub properties: HashMap<String, String>, -} - -/// Document modality input -#[derive(Debug, Clone, Serialize, Deserialize)] -pub struct OctadDocumentInput { - /// Document title - pub title: String, - /// Document body - pub body: String, - /// Additional fields - pub fields: HashMap<String, String>, -} - -/// Temporal modality input — the entity's real-world observation time -#[derive(Debug, Clone, Serialize, Deserialize)] -pub struct OctadTemporalInput { - /// When the underlying event happened in the territory the entity represents - /// (e.g. an email's `Date:` header). UTC; callers must convert from local - /// timezones before submission. - pub observed_at: DateTime<Utc>, -} - -/// Provenance modality input — records a lineage event -#[derive(Debug, Clone, Serialize, Deserialize)] -pub struct OctadProvenanceInput { - /// Event type (created, modified, imported, normalized, etc.) - pub event_type: String, - /// Who or what caused this event - pub actor: String, - /// Optional source identifier (URL, upstream entity, file path) - pub source: Option<String>, - /// Human-readable description of the event - pub description: String, -} - -/// Spatial modality input — geospatial coordinates and geometry -#[derive(Debug, Clone, Serialize, Deserialize)] -pub struct OctadSpatialInput { - /// Latitude in decimal degrees (WGS84) - pub latitude: f64, - /// Longitude in decimal degrees (WGS84) - pub longitude: f64, - /// Altitude in metres (optional) - pub altitude: Option<f64>, - /// Geometry type (Point, LineString, Polygon, etc.) — defaults to Point - pub geometry_type: Option<String>, - /// Spatial Reference System Identifier — defaults to 4326 (WGS84) - pub srid: Option<u32>, - /// Arbitrary spatial properties (address, region, accuracy, etc.) - #[serde(default)] - pub properties: HashMap<String, String>, -} - -/// A complete Octad entity with all modality data (octad: 8 modalities) -#[derive(Debug, Clone, Serialize, Deserialize)] -pub struct Octad { - /// Entity ID - pub id: OctadId, - /// Status - pub status: OctadStatus, - /// Graph node - pub graph_node: Option<GraphNode>, - /// Vector embedding - pub embedding: Option<Embedding>, - /// Tensor data - pub tensor: Option<Tensor>, - /// Semantic annotation - pub semantic: Option<SemanticAnnotation>, - /// Document - pub document: Option<Document>, - /// Version history info - pub version_count: u64, - /// Provenance chain length (number of recorded events) - pub provenance_chain_length: u64, - /// Spatial data (coordinates, geometry, SRID) - pub spatial_data: Option<SpatialData>, -} - -/// Octad store - manages entities across all modalities -#[async_trait] -pub trait OctadStore: Send + Sync { - /// Create a new Octad entity - async fn create(&self, input: OctadInput) -> Result<Octad, OctadError>; - - /// Update an existing Octad - async fn update(&self, id: &OctadId, input: OctadInput) -> Result<Octad, OctadError>; - - /// Get a Octad by ID - async fn get(&self, id: &OctadId) -> Result<Option<Octad>, OctadError>; - - /// Delete a Octad - async fn delete(&self, id: &OctadId) -> Result<(), OctadError>; - - /// Get Octad status - async fn status(&self, id: &OctadId) -> Result<Option<OctadStatus>, OctadError>; - - /// Search by vector similarity - async fn search_similar(&self, embedding: &[f32], k: usize) -> Result<Vec<Octad>, OctadError>; - - /// Search by document text - async fn search_text(&self, query: &str, limit: usize) -> Result<Vec<Octad>, OctadError>; - - /// Query by graph relationship - async fn query_related(&self, id: &OctadId, predicate: &str) -> Result<Vec<Octad>, OctadError>; - - /// Get version at a specific point in time - async fn at_time(&self, id: &OctadId, time: DateTime<Utc>) -> Result<Option<Octad>, OctadError>; - - /// List octads with pagination - async fn list(&self, limit: usize, offset: usize) -> Result<Vec<Octad>, OctadError>; -} - -/// Configuration for Octad store -#[derive(Debug, Clone, Serialize, Deserialize)] -pub struct OctadConfig { - /// Base IRI for graph nodes - pub base_iri: String, - /// Vector embedding dimension - pub vector_dimension: usize, - /// Whether to enforce full modality population - pub require_complete: bool, -} - -impl Default for OctadConfig { - fn default() -> Self { - Self { - base_iri: "https://verisim.db/entity".to_string(), - vector_dimension: 384, - require_complete: false, - } - } -} - -/// Builder for creating Octad inputs -pub struct OctadBuilder { - input: OctadInput, -} - -impl OctadBuilder { - /// Create a new builder - pub fn new() -> Self { - Self { - input: OctadInput::default(), - } - } - - /// Add graph relationships - pub fn with_relationships(mut self, relationships: Vec<(&str, &str)>) -> Self { - self.input.graph = Some(OctadGraphInput { - relationships: relationships - .into_iter() - .map(|(p, t)| (p.to_string(), t.to_string())) - .collect(), - }); - self - } - - /// Add vector embedding - pub fn with_embedding(mut self, embedding: Vec<f32>) -> Self { - self.input.vector = Some(OctadVectorInput { - embedding, - model: None, - }); - self - } - - /// Add tensor data - pub fn with_tensor(mut self, shape: Vec<usize>, data: Vec<f64>) -> Self { - self.input.tensor = Some(OctadTensorInput { shape, data }); - self - } - - /// Add semantic types - pub fn with_types(mut self, types: Vec<&str>) -> Self { - let existing = self.input.semantic.take().unwrap_or(OctadSemanticInput { - types: Vec::new(), - properties: HashMap::new(), - }); - self.input.semantic = Some(OctadSemanticInput { - types: types.into_iter().map(|t| t.to_string()).collect(), - properties: existing.properties, - }); - self - } - - /// Add document content - pub fn with_document(mut self, title: &str, body: &str) -> Self { - self.input.document = Some(OctadDocumentInput { - title: title.to_string(), - body: body.to_string(), - fields: HashMap::new(), - }); - self - } - - /// Add provenance event - pub fn with_provenance(mut self, event_type: &str, actor: &str, description: &str) -> Self { - self.input.provenance = Some(OctadProvenanceInput { - event_type: event_type.to_string(), - actor: actor.to_string(), - source: None, - description: description.to_string(), - }); - self - } - - /// Add spatial coordinates (WGS84 point) - pub fn with_spatial(mut self, latitude: f64, longitude: f64) -> Self { - self.input.spatial = Some(OctadSpatialInput { - latitude, - longitude, - altitude: None, - geometry_type: None, - srid: None, - properties: HashMap::new(), - }); - self - } - - /// Add real-world observation time (territory clock, not database ingestion time) - pub fn with_observed_at(mut self, observed_at: DateTime<Utc>) -> Self { - self.input.temporal = Some(OctadTemporalInput { observed_at }); - self - } - - /// Add metadata - pub fn with_metadata(mut self, key: &str, value: &str) -> Self { - self.input.metadata.insert(key.to_string(), value.to_string()); - self - } - - /// Build the input - pub fn build(self) -> OctadInput { - self.input - } -} - -impl Default for OctadBuilder { - fn default() -> Self { - Self::new() - } -} - -#[cfg(test)] -mod tests { - use super::*; - - #[test] - fn test_octad_id() { - let id = OctadId::new("test-123"); - assert_eq!(id.as_str(), "test-123"); - assert_eq!(id.to_iri("https://example.org"), "https://example.org/test-123"); - } - - #[test] - fn test_octad_builder() { - let input = OctadBuilder::new() - .with_document("Test", "Test content") - .with_embedding(vec![0.1, 0.2, 0.3]) - .with_types(vec!["https://example.org/Person"]) - .with_metadata("source", "test") - .build(); - - assert!(input.document.is_some()); - assert!(input.vector.is_some()); - assert!(input.semantic.is_some()); - assert_eq!(input.metadata.get("source"), Some(&"test".to_string())); - } - - #[test] - fn test_modality_status() { - let mut status = ModalityStatus::default(); - assert!(!status.is_complete()); - assert_eq!(status.missing().len(), 8); - - status.graph = true; - status.vector = true; - status.tensor = true; - status.semantic = true; - status.document = true; - status.temporal = true; - status.provenance = true; - status.spatial = true; - - assert!(status.is_complete()); - assert!(status.missing().is_empty()); - } -} diff --git a/verisimdb/rust-core/verisim-octad/src/query_octad.rs b/verisimdb/rust-core/verisim-octad/src/query_octad.rs deleted file mode 100644 index 323d4a6a..00000000 --- a/verisimdb/rust-core/verisim-octad/src/query_octad.rs +++ /dev/null @@ -1,228 +0,0 @@ -// SPDX-License-Identifier: MPL-2.0 -//! QueryOctad Builder — Creates a octad from a VCL query. -//! -//! Homoiconicity: queries are data. A VCL query stored as a octad has: -//! - **Document**: query text (searchable via full-text) -//! - **Graph**: parse tree as subject → predicate → object triples -//! - **Vector**: embedding of query text (for similarity search of past queries) -//! - **Tensor**: cost vector from execution plan -//! - **Semantic**: proof obligations as typed annotations -//! - **Temporal**: query execution history (when run, what results) - -use chrono::{DateTime, Utc}; -use serde::{Deserialize, Serialize}; -use std::collections::HashMap; - -use crate::{OctadId, OctadInput, OctadDocumentInput, OctadVectorInput, - OctadGraphInput, OctadTensorInput, OctadSemanticInput}; - -/// Metadata about a query execution -#[derive(Debug, Clone, Serialize, Deserialize)] -pub struct QueryExecution { - /// When the query was executed - pub executed_at: DateTime<Utc>, - /// Duration in milliseconds - pub duration_ms: u64, - /// Number of results returned - pub result_count: usize, - /// Execution plan cost estimate - pub estimated_cost: f64, -} - -/// A builder for creating octads from VCL queries -#[derive(Debug)] -pub struct QueryOctadBuilder { - query_text: String, - query_id: Option<String>, - parse_tree_triples: Vec<(String, String, String)>, - embedding: Option<Vec<f32>>, - cost_vector: Option<Vec<f64>>, - proof_obligations: Vec<String>, - executions: Vec<QueryExecution>, - metadata: HashMap<String, String>, -} - -impl QueryOctadBuilder { - /// Create a new builder from a VCL query string - pub fn new(query_text: impl Into<String>) -> Self { - Self { - query_text: query_text.into(), - query_id: None, - parse_tree_triples: Vec::new(), - embedding: None, - cost_vector: None, - proof_obligations: Vec::new(), - executions: Vec::new(), - metadata: HashMap::new(), - } - } - - /// Set the query ID (defaults to auto-generated) - pub fn with_id(mut self, id: impl Into<String>) -> Self { - self.query_id = Some(id.into()); - self - } - - /// Add parse tree triples (AST as RDF: subject → predicate → object) - pub fn with_parse_tree(mut self, triples: Vec<(String, String, String)>) -> Self { - self.parse_tree_triples = triples; - self - } - - /// Set the query embedding vector (for similarity search) - pub fn with_embedding(mut self, embedding: Vec<f32>) -> Self { - self.embedding = Some(embedding); - self - } - - /// Set the execution plan cost vector - pub fn with_cost_vector(mut self, costs: Vec<f64>) -> Self { - self.cost_vector = Some(costs); - self - } - - /// Add proof obligations from the query - pub fn with_proof_obligations(mut self, obligations: Vec<String>) -> Self { - self.proof_obligations = obligations; - self - } - - /// Record a query execution - pub fn with_execution(mut self, execution: QueryExecution) -> Self { - self.executions.push(execution); - self - } - - /// Add metadata - pub fn with_metadata(mut self, key: impl Into<String>, value: impl Into<String>) -> Self { - self.metadata.insert(key.into(), value.into()); - self - } - - /// Build the OctadInput for storage - pub fn build(self) -> (OctadId, OctadInput) { - let id = self.query_id.unwrap_or_else(|| { - format!("query-{}", uuid::Uuid::new_v4()) - }); - - let mut input = OctadInput::default(); - - // Document modality: query text (searchable) - input.document = Some(OctadDocumentInput { - title: format!("VCL Query: {}", truncate(&self.query_text, 80)), - body: self.query_text.clone(), - fields: { - let mut fields = HashMap::new(); - fields.insert("type".to_string(), "vcl_query".to_string()); - fields.insert("query_text".to_string(), self.query_text); - if !self.executions.is_empty() { - fields.insert( - "last_executed".to_string(), - self.executions.last().expect("TODO: handle error").executed_at.to_rfc3339(), - ); - fields.insert( - "execution_count".to_string(), - self.executions.len().to_string(), - ); - } - fields - }, - }); - - // Graph modality: parse tree as relationships - if !self.parse_tree_triples.is_empty() { - input.graph = Some(OctadGraphInput { - relationships: self - .parse_tree_triples - .into_iter() - .map(|(_, predicate, object)| (predicate, object)) - .collect(), - }); - } - - // Vector modality: embedding of query text - if let Some(embedding) = self.embedding { - input.vector = Some(OctadVectorInput { - embedding, - model: Some("query-embedding".to_string()), - }); - } - - // Tensor modality: cost vector from execution plan - if let Some(costs) = self.cost_vector { - let len = costs.len(); - input.tensor = Some(OctadTensorInput { - shape: vec![1, len], - data: costs, - }); - } - - // Semantic modality: proof obligations as type annotations - if !self.proof_obligations.is_empty() { - input.semantic = Some(OctadSemanticInput { - types: self.proof_obligations, - properties: HashMap::new(), - }); - } - - // Metadata - input.metadata = self.metadata; - - (OctadId::new(id), input) - } -} - -/// Truncate a string to max_len characters with ellipsis -fn truncate(s: &str, max_len: usize) -> String { - if s.len() <= max_len { - s.to_string() - } else { - format!("{}...", &s[..max_len.saturating_sub(3)]) - } -} - -#[cfg(test)] -mod tests { - use super::*; - - #[test] - fn test_build_query_octad() { - let (id, input) = QueryOctadBuilder::new("SELECT * FROM octads WHERE drift > 0.5") - .with_id("query-001") - .with_embedding(vec![0.1, 0.2, 0.3]) - .with_cost_vector(vec![1.0, 0.5, 0.3]) - .with_proof_obligations(vec!["verisim:DriftQuery".to_string()]) - .with_execution(QueryExecution { - executed_at: Utc::now(), - duration_ms: 42, - result_count: 5, - estimated_cost: 1.8, - }) - .with_metadata("user", "test-user") - .build(); - - assert_eq!(id.0, "query-001"); - assert!(input.document.is_some()); - assert!(input.vector.is_some()); - assert!(input.tensor.is_some()); - assert!(input.semantic.is_some()); - - let doc = input.document.expect("TODO: handle error"); - assert!(doc.title.starts_with("VCL Query:")); - assert!(doc.body.contains("drift")); - } - - #[test] - fn test_auto_id_generation() { - let (id, _input) = QueryOctadBuilder::new("SELECT 1").build(); - assert!(id.0.starts_with("query-")); - } - - #[test] - fn test_minimal_query_octad() { - let (_, input) = QueryOctadBuilder::new("REFLECT").build(); - assert!(input.document.is_some()); - assert!(input.vector.is_none()); - assert!(input.tensor.is_none()); - } -} diff --git a/verisimdb/rust-core/verisim-octad/src/ram_promotion.rs b/verisimdb/rust-core/verisim-octad/src/ram_promotion.rs deleted file mode 100644 index bb589820..00000000 --- a/verisimdb/rust-core/verisim-octad/src/ram_promotion.rs +++ /dev/null @@ -1,472 +0,0 @@ -// SPDX-License-Identifier: MPL-2.0 -// Copyright (c) 2026 Jonathan D.A. Jewell (hyperpolymath) <j.d.a.jewell@open.ac.uk> -// -// VeriSimDB Optional RAM Promotion — tiered storage acceleration. -// -// DESIGN CONSTRAINTS: -// - OPTIONAL: database works fine without this, disk-backed by default -// - MAX 3 OCTADS promoted at a time (hard limit) -// - MINIMAL TIME in RAM: promote before operation, demote immediately after -// - ONLY IF SIGNIFICANT BENEFIT: estimated speedup must exceed threshold -// -// The promotion manager tracks which modality stores are currently -// RAM-resident via tmpfs-backed mmaps, enforces the 2-octad limit, -// and auto-demotes after operation completion. -// -// ABSOLUTE GUARANTEES: -// -// 1. NOTHING stays in RAM unless EXPLICITLY requested by the caller. -// The default state is DISABLED. Promotion never happens automatically. -// -// 2. All data is ALWAYS on disk (via VeriSimDB's WAL + redb persistence). -// RAM promotion is a READ CACHE only — the WAL on disk is the source -// of truth. If RAM disappears (crash, reboot), WAL replay recovers. -// -// 3. Maximum 2 octads in RAM at once (conservative limit for crash safety). -// -// 4. Maximum 5 minutes in RAM before forced demotion (MAX_PROMOTION_DURATION). -// The caller should demote IMMEDIATELY after their operation completes. -// The 5-minute timeout is a safety net, not a target. -// -// 5. The PromotionManager does NOT move data. It signals to the modality -// store that it should use tmpfs-backed storage. The actual data lives -// in redb (disk) at all times. The RAM copy is a performance overlay. -// -// WHERE DATA LIVES: -// - WAL: always on disk (redb file in data directory) -// - Modality stores: always on disk (redb files) -// - RAM promotion: tmpfs overlay for reads, writes go to WAL first -// - On crash: WAL replays, RAM overlay is gone, no data loss - -use std::collections::{HashMap, HashSet}; -use std::time::{Duration, Instant}; -use serde::{Serialize, Deserialize}; - -/// Hard limit: maximum octads promoted to RAM simultaneously. -/// Conservative limit (2 not 3) for crash safety — fewer in-flight -/// RAM-resident octads means faster WAL replay on recovery. -const MAX_PROMOTED: usize = 2; - -/// Minimum estimated speedup factor to justify promotion. -/// If estimated speedup is less than 2x, don't bother promoting. -const MIN_SPEEDUP_FACTOR: f64 = 2.0; - -/// Maximum time an octad can stay promoted before forced demotion. -const MAX_PROMOTION_DURATION: Duration = Duration::from_secs(300); // 5 minutes - -/// Modality names matching VeriSimDB's octad structure. -#[derive(Debug, Clone, Hash, Eq, PartialEq, Serialize, Deserialize)] -pub enum Modality { - Graph, - Vector, - Tensor, - Semantic, - Document, - Temporal, - Provenance, - Spatial, -} - -impl Modality { - pub fn all() -> Vec<Modality> { - vec![ - Modality::Graph, Modality::Vector, Modality::Tensor, - Modality::Semantic, Modality::Document, Modality::Temporal, - Modality::Provenance, Modality::Spatial, - ] - } - - pub fn name(&self) -> &str { - match self { - Modality::Graph => "graph", - Modality::Vector => "vector", - Modality::Tensor => "tensor", - Modality::Semantic => "semantic", - Modality::Document => "document", - Modality::Temporal => "temporal", - Modality::Provenance => "provenance", - Modality::Spatial => "spatial", - } - } -} - -/// State of a promoted modality. -#[derive(Debug, Clone)] -struct PromotedState { - modality: Modality, - promoted_at: Instant, - estimated_size_bytes: u64, - operation_count: u64, -} - -/// Promotion decision result. -#[derive(Debug, Clone)] -pub enum PromotionDecision { - /// Promote — significant benefit expected. - Promote { modalities: Vec<Modality>, estimated_speedup: f64 }, - /// Skip — benefit too small to justify RAM usage. - Skip { reason: String }, - /// Blocked — already at MAX_PROMOTED limit. - Blocked { current_count: usize, limit: usize }, -} - -/// The RAM promotion manager. -#[derive(Debug)] -pub struct PromotionManager { - /// Currently promoted modalities. - promoted: HashMap<Modality, PromotedState>, - /// Total RAM budget (bytes). Default: 1GB. - ram_budget: u64, - /// RAM currently used by promoted modalities. - ram_used: u64, - /// Whether promotion is enabled at all. - enabled: bool, - /// Promotion/demotion history for performance tracking. - history: Vec<PromotionEvent>, -} - -/// A promotion or demotion event for history tracking. -#[derive(Debug, Clone, Serialize)] -pub struct PromotionEvent { - pub modality: String, - pub action: PromotionAction, - pub duration_ms: Option<u64>, - pub size_bytes: u64, - pub timestamp_epoch_ms: u64, -} - -#[derive(Debug, Clone, Serialize)] -pub enum PromotionAction { - Promoted, - Demoted, - Denied { reason: String }, -} - -impl PromotionManager { - /// Create a new promotion manager. Disabled by default. - pub fn new() -> Self { - PromotionManager { - promoted: HashMap::new(), - ram_budget: 1_073_741_824, // 1 GB - ram_used: 0, - enabled: false, - history: Vec::new(), - } - } - - /// Enable promotion with a RAM budget. - pub fn enable(&mut self, ram_budget_bytes: u64) { - self.enabled = true; - self.ram_budget = ram_budget_bytes; - } - - /// Disable promotion. Demotes all currently promoted modalities. - pub fn disable(&mut self) { - let modalities: Vec<Modality> = self.promoted.keys().cloned().collect(); - for m in modalities { - self.demote(&m); - } - self.enabled = false; - } - - /// Check if promotion is enabled. - pub fn is_enabled(&self) -> bool { - self.enabled - } - - /// How many modalities are currently promoted. - pub fn promoted_count(&self) -> usize { - self.promoted.len() - } - - /// Which modalities are currently promoted. - pub fn promoted_modalities(&self) -> Vec<&Modality> { - self.promoted.keys().collect() - } - - /// Decide whether to promote modalities for an operation. - /// - /// Considers: current promotion count, RAM budget, estimated data size, - /// and estimated speedup factor. - pub fn decide( - &self, - modalities: &[Modality], - estimated_data_size: u64, - estimated_disk_time_ms: u64, - estimated_ram_time_ms: u64, - ) -> PromotionDecision { - if !self.enabled { - return PromotionDecision::Skip { - reason: "RAM promotion disabled".to_string(), - }; - } - - // Check 3-octad limit. - let new_count = modalities.iter() - .filter(|m| !self.promoted.contains_key(m)) - .count(); - if self.promoted.len() + new_count > MAX_PROMOTED { - return PromotionDecision::Blocked { - current_count: self.promoted.len(), - limit: MAX_PROMOTED, - }; - } - - // Check RAM budget. - if self.ram_used + estimated_data_size > self.ram_budget { - return PromotionDecision::Skip { - reason: format!( - "Insufficient RAM budget: need {} bytes, have {} available", - estimated_data_size, - self.ram_budget - self.ram_used - ), - }; - } - - // Check speedup threshold. - let speedup = if estimated_ram_time_ms > 0 { - estimated_disk_time_ms as f64 / estimated_ram_time_ms as f64 - } else { - f64::INFINITY - }; - - if speedup < MIN_SPEEDUP_FACTOR { - return PromotionDecision::Skip { - reason: format!( - "Estimated speedup {:.1}x below threshold {:.1}x", - speedup, MIN_SPEEDUP_FACTOR - ), - }; - } - - PromotionDecision::Promote { - modalities: modalities.to_vec(), - estimated_speedup: speedup, - } - } - - /// Promote a modality to RAM. - /// - /// Returns Ok(()) if promoted, Err if blocked/disabled. - pub fn promote(&mut self, modality: &Modality, estimated_size: u64) -> Result<(), String> { - if !self.enabled { - return Err("RAM promotion disabled".to_string()); - } - - if self.promoted.len() >= MAX_PROMOTED { - return Err(format!( - "Cannot promote: already at limit ({}/{})", - self.promoted.len(), MAX_PROMOTED - )); - } - - if self.promoted.contains_key(modality) { - return Ok(()); // Already promoted. - } - - if self.ram_used + estimated_size > self.ram_budget { - return Err(format!( - "Cannot promote: would exceed RAM budget ({} + {} > {})", - self.ram_used, estimated_size, self.ram_budget - )); - } - - self.promoted.insert(modality.clone(), PromotedState { - modality: modality.clone(), - promoted_at: Instant::now(), - estimated_size_bytes: estimated_size, - operation_count: 0, - }); - self.ram_used += estimated_size; - - self.history.push(PromotionEvent { - modality: modality.name().to_string(), - action: PromotionAction::Promoted, - duration_ms: None, - size_bytes: estimated_size, - timestamp_epoch_ms: epoch_ms(), - }); - - Ok(()) - } - - /// Demote a modality from RAM back to disk. - pub fn demote(&mut self, modality: &Modality) { - if let Some(state) = self.promoted.remove(modality) { - self.ram_used = self.ram_used.saturating_sub(state.estimated_size_bytes); - - let duration = state.promoted_at.elapsed(); - self.history.push(PromotionEvent { - modality: modality.name().to_string(), - action: PromotionAction::Demoted, - duration_ms: Some(duration.as_millis() as u64), - size_bytes: state.estimated_size_bytes, - timestamp_epoch_ms: epoch_ms(), - }); - } - } - - /// Demote all currently promoted modalities. - pub fn demote_all(&mut self) { - let modalities: Vec<Modality> = self.promoted.keys().cloned().collect(); - for m in modalities { - self.demote(&m); - } - } - - /// Check for expired promotions and force-demote them. - pub fn enforce_timeouts(&mut self) { - let expired: Vec<Modality> = self.promoted.iter() - .filter(|(_, state)| state.promoted_at.elapsed() > MAX_PROMOTION_DURATION) - .map(|(m, _)| m.clone()) - .collect(); - - for m in expired { - self.demote(&m); - } - } - - /// Record that an operation was performed on a promoted modality. - pub fn record_operation(&mut self, modality: &Modality) { - if let Some(state) = self.promoted.get_mut(modality) { - state.operation_count += 1; - } - } - - /// Get promotion history. - pub fn history(&self) -> &[PromotionEvent] { - &self.history - } - - /// Get current RAM usage. - pub fn ram_usage(&self) -> (u64, u64) { - (self.ram_used, self.ram_budget) - } -} - -impl Default for PromotionManager { - fn default() -> Self { Self::new() } -} - -fn epoch_ms() -> u64 { - std::time::SystemTime::now() - .duration_since(std::time::UNIX_EPOCH) - .unwrap_or_default() - .as_millis() as u64 -} - -#[cfg(test)] -mod tests { - use super::*; - - #[test] - fn disabled_by_default() { - let pm = PromotionManager::new(); - assert!(!pm.is_enabled()); - assert_eq!(pm.promoted_count(), 0); - } - - #[test] - fn enable_and_promote() { - let mut pm = PromotionManager::new(); - pm.enable(1_000_000); // 1 MB budget - - assert!(pm.promote(&Modality::Graph, 100_000).is_ok()); - assert_eq!(pm.promoted_count(), 1); - assert!(pm.promoted_modalities().contains(&&Modality::Graph)); - } - - #[test] - fn max_two_limit() { - let mut pm = PromotionManager::new(); - pm.enable(10_000_000); - - assert!(pm.promote(&Modality::Graph, 1000).is_ok()); - assert!(pm.promote(&Modality::Vector, 1000).is_ok()); - // Third should fail — limit is 2. - assert!(pm.promote(&Modality::Tensor, 1000).is_err()); - assert_eq!(pm.promoted_count(), 2); - } - - #[test] - fn demote_frees_slot() { - let mut pm = PromotionManager::new(); - pm.enable(10_000_000); - - pm.promote(&Modality::Graph, 1000).expect("TODO: handle error"); - pm.promote(&Modality::Vector, 1000).expect("TODO: handle error"); - - pm.demote(&Modality::Graph); - assert_eq!(pm.promoted_count(), 1); - - // Now we can promote another. - assert!(pm.promote(&Modality::Tensor, 1000).is_ok()); - assert_eq!(pm.promoted_count(), 2); - } - - #[test] - fn ram_budget_enforced() { - let mut pm = PromotionManager::new(); - pm.enable(5000); // 5 KB budget - - assert!(pm.promote(&Modality::Graph, 3000).is_ok()); - // This would exceed budget. - assert!(pm.promote(&Modality::Vector, 3000).is_err()); - - let (used, budget) = pm.ram_usage(); - assert_eq!(used, 3000); - assert_eq!(budget, 5000); - } - - #[test] - fn decision_checks_speedup() { - let mut pm = PromotionManager::new(); - pm.enable(10_000_000); - - // 10x speedup — should promote. - let decision = pm.decide(&[Modality::Tensor], 1000, 100, 10); - assert!(matches!(decision, PromotionDecision::Promote { .. })); - - // 1.5x speedup — below threshold (2x), should skip. - let decision = pm.decide(&[Modality::Graph], 1000, 15, 10); - assert!(matches!(decision, PromotionDecision::Skip { .. })); - } - - #[test] - fn demote_all() { - let mut pm = PromotionManager::new(); - pm.enable(10_000_000); - - pm.promote(&Modality::Graph, 1000).expect("TODO: handle error"); - pm.promote(&Modality::Vector, 1000).expect("TODO: handle error"); - pm.demote_all(); - - assert_eq!(pm.promoted_count(), 0); - let (used, _) = pm.ram_usage(); - assert_eq!(used, 0); - } - - #[test] - fn history_tracked() { - let mut pm = PromotionManager::new(); - pm.enable(10_000_000); - - pm.promote(&Modality::Graph, 1000).expect("TODO: handle error"); - pm.demote(&Modality::Graph); - - assert_eq!(pm.history().len(), 2); - assert!(matches!(pm.history()[0].action, PromotionAction::Promoted)); - assert!(matches!(pm.history()[1].action, PromotionAction::Demoted)); - assert!(pm.history()[1].duration_ms.is_some()); - } - - #[test] - fn disabled_rejects_all() { - let mut pm = PromotionManager::new(); - // Not enabled. - assert!(pm.promote(&Modality::Graph, 1000).is_err()); - - let decision = pm.decide(&[Modality::Graph], 1000, 100, 10); - assert!(matches!(decision, PromotionDecision::Skip { .. })); - } -} diff --git a/verisimdb/rust-core/verisim-octad/src/store.rs b/verisimdb/rust-core/verisim-octad/src/store.rs deleted file mode 100644 index 879f5ac5..00000000 --- a/verisimdb/rust-core/verisim-octad/src/store.rs +++ /dev/null @@ -1,1616 +0,0 @@ -// SPDX-License-Identifier: MPL-2.0 -//! In-memory OctadStore implementation -//! -//! Coordinates all eight modality stores (octad) for unified entity management. - -use async_trait::async_trait; -use chrono::{DateTime, Utc}; -use serde::{Deserialize, Serialize}; -use std::collections::HashMap; -use std::sync::Arc; -use tokio::sync::RwLock; -use tracing::{debug, info, instrument}; - -use crate::{ - Coordinates, Document, DocumentStore, Embedding, GeometryType, GraphEdge, GraphNode, - GraphObject, GraphStore, Octad, OctadConfig, OctadDocumentInput, OctadError, OctadGraphInput, - OctadId, OctadInput, OctadProvenanceInput, OctadSemanticInput, OctadSpatialInput, - OctadStatus, OctadStore, OctadTensorInput, OctadVectorInput, ModalityStatus, Provenance, - ProvenanceEventType, ProvenanceStore, SemanticAnnotation, SemanticStore, SemanticValue, - SpatialData, SpatialStore, Tensor, TensorStore, TemporalStore, VectorStore, -}; -use crate::transaction::{IsolationLevel, LockType, TransactionManager}; -use verisim_wal::{WalEntry, WalModality, WalOperation, WalWriter, SyncMode}; - -/// Snapshot of a Octad for versioning -#[derive(Debug, Clone, Serialize, Deserialize)] -pub struct OctadSnapshot { - pub id: OctadId, - pub input: OctadInput, - pub modality_status: ModalityStatus, - pub timestamp: DateTime<Utc>, -} - -/// In-memory implementation of OctadStore -/// -/// This store coordinates all eight modality stores (octad), ensuring -/// cross-modal consistency when entities are created, updated, or deleted. -/// Write operations (create/update/delete) are wrapped in ACID transactions -/// via the [`TransactionManager`], guaranteeing atomicity across all modalities. -pub struct InMemoryOctadStore<G, V, D, T, S, R, P, L> -where - G: GraphStore, - V: VectorStore, - D: DocumentStore, - T: TensorStore, - S: SemanticStore, - R: TemporalStore<Data = OctadSnapshot>, - P: ProvenanceStore, - L: SpatialStore, -{ - config: OctadConfig, - /// Octad status registry - octads: Arc<RwLock<HashMap<String, OctadStatus>>>, - /// ACID transaction manager for cross-modality atomicity - txn_manager: Arc<TransactionManager>, - /// Optional write-ahead log for crash recovery. - /// When present, all modality writes are logged before execution. - wal: Option<Arc<tokio::sync::Mutex<WalWriter>>>, - /// Graph store - graph: Arc<G>, - /// Vector store - vector: Arc<V>, - /// Document store - document: Arc<D>, - /// Tensor store - tensor: Arc<T>, - /// Semantic store - semantic: Arc<S>, - /// Temporal (versioning) store - temporal: Arc<R>, - /// Provenance (lineage tracking) store - provenance: Arc<P>, - /// Spatial (geospatial) store - spatial: Arc<L>, -} - -impl<G, V, D, T, S, R, P, L> InMemoryOctadStore<G, V, D, T, S, R, P, L> -where - G: GraphStore, - V: VectorStore, - D: DocumentStore, - T: TensorStore, - S: SemanticStore, - R: TemporalStore<Data = OctadSnapshot>, - P: ProvenanceStore, - L: SpatialStore, -{ - /// Create a new in-memory octad store with all eight modality stores. - /// - /// Automatically creates a [`TransactionManager`] to provide ACID - /// guarantees across all modality writes. - pub fn new( - config: OctadConfig, - graph: Arc<G>, - vector: Arc<V>, - document: Arc<D>, - tensor: Arc<T>, - semantic: Arc<S>, - temporal: Arc<R>, - provenance: Arc<P>, - spatial: Arc<L>, - ) -> Self { - Self { - config, - octads: Arc::new(RwLock::new(HashMap::new())), - txn_manager: Arc::new(TransactionManager::new()), - wal: None, - graph, - vector, - document, - tensor, - semantic, - temporal, - provenance, - spatial, - } - } - - /// Enable write-ahead logging for crash recovery. - /// - /// When enabled, all modality writes are recorded to the WAL before - /// being applied to the stores. On crash, the WAL can be replayed to - /// recover PENDING operations. - /// - /// # Arguments - /// - /// * `wal_dir` - Directory for WAL segment files (created if absent). - /// * `sync_mode` - Controls fsync behavior (Fsync, Periodic, or Async). - pub fn with_wal( - mut self, - wal_dir: impl AsRef<std::path::Path>, - sync_mode: SyncMode, - ) -> Result<Self, OctadError> { - let writer = WalWriter::open(wal_dir, sync_mode).map_err(|e| { - OctadError::ModalityError { - modality: "wal".to_string(), - message: format!("Failed to open WAL: {e}"), - } - })?; - self.wal = Some(Arc::new(tokio::sync::Mutex::new(writer))); - Ok(self) - } - - /// Access the transaction manager for diagnostics or external coordination. - pub fn transaction_manager(&self) -> &Arc<TransactionManager> { - &self.txn_manager - } - - /// Write a WAL entry if WAL is enabled. Returns Ok(()) if WAL is disabled. - async fn wal_append( - &self, - operation: WalOperation, - modality: WalModality, - entity_id: &str, - payload: &[u8], - ) -> Result<(), OctadError> { - if let Some(ref wal) = self.wal { - let entry = WalEntry { - sequence: 0, // Assigned by the writer - timestamp: Utc::now(), - operation, - modality, - entity_id: entity_id.to_string(), - payload: payload.to_vec(), - }; - let mut writer = wal.lock().await; - writer.append(entry).map_err(|e| OctadError::ModalityError { - modality: "wal".to_string(), - message: format!("WAL append failed: {e}"), - })?; - } - Ok(()) - } - - /// Write a WAL checkpoint marker if WAL is enabled. - async fn wal_checkpoint(&self) -> Result<(), OctadError> { - if let Some(ref wal) = self.wal { - let mut writer = wal.lock().await; - writer.checkpoint().map_err(|e| OctadError::ModalityError { - modality: "wal".to_string(), - message: format!("WAL checkpoint failed: {e}"), - })?; - } - Ok(()) - } - - /// Replay the write-ahead log to recover state after a crash. - /// - /// This method should be called once during startup, after the WAL is - /// opened and persistent modality stores have loaded their data from redb. - /// - /// Recovery strategy (octad-level, Option B): - /// - /// 1. Find the last checkpoint in the WAL. - /// 2. Replay all entries after that checkpoint. - /// 3. For committed operations (Insert/Update/Delete followed by a - /// Checkpoint with payload b"COMMITTED"): rebuild the octad status - /// registry. The modality data is already in redb. - /// 4. For uncommitted operations (no matching Checkpoint): log a warning. - /// The incomplete write may have partially persisted to redb — the data - /// is still consistent per-modality (each redb write is atomic), but the - /// cross-modal operation may be incomplete. - /// 5. Write a fresh checkpoint after replay. - /// - /// Returns the number of entities recovered. - pub async fn replay_wal( - &self, - wal_dir: impl AsRef<std::path::Path>, - ) -> Result<usize, OctadError> { - use std::collections::HashSet; - use verisim_wal::WalReader; - - let reader = match WalReader::open(&wal_dir) { - Ok(r) => r, - Err(verisim_wal::WalError::DirectoryNotFound(_)) => { - info!("No WAL directory found — clean start"); - return Ok(0); - } - Err(e) => { - return Err(OctadError::ModalityError { - modality: "wal".to_string(), - message: format!("Failed to open WAL reader: {e}"), - }); - } - }; - - // Replay ALL WAL entries to rebuild the complete octad registry. - // We replay from sequence 0 because per-entity "COMMITTED" markers - // are not global checkpoints — each entity has its own commit marker. - // A global checkpoint (from graceful_shutdown) would allow starting - // from a later point, but for correctness we always replay everything. - info!("Replaying WAL from beginning"); - - let entries: Vec<WalEntry> = reader - .replay_all() - .map_err(|e| OctadError::ModalityError { - modality: "wal".to_string(), - message: format!("WAL replay_from failed: {e}"), - })? - .collect(); - - if entries.is_empty() { - info!("WAL replay: no entries to replay"); - return Ok(0); - } - - // Track which entity_ids have been committed (have a Checkpoint entry) - let mut committed_entities: HashSet<String> = HashSet::new(); - let mut uncommitted_entities: HashSet<String> = HashSet::new(); - let mut entity_ops: HashMap<String, (WalOperation, Vec<u8>)> = HashMap::new(); - - for entry in &entries { - match entry.operation { - WalOperation::Checkpoint => { - // Checkpoint with payload "COMMITTED" marks the previous op as complete - if entry.payload == b"COMMITTED" { - committed_entities.insert(entry.entity_id.clone()); - uncommitted_entities.remove(&entry.entity_id); - } - } - WalOperation::Insert | WalOperation::Update | WalOperation::Delete => { - entity_ops.insert( - entry.entity_id.clone(), - (entry.operation.clone(), entry.payload.clone()), - ); - if !committed_entities.contains(&entry.entity_id) { - uncommitted_entities.insert(entry.entity_id.clone()); - } - } - } - } - - // Warn about uncommitted operations - for entity_id in &uncommitted_entities { - tracing::warn!( - entity_id, - "WAL replay: uncommitted operation for entity — may be partially persisted" - ); - } - - // Rebuild octad status registry for committed entities - let mut octads = self.octads.write().await; - let mut recovered = 0usize; - - for entity_id in &committed_entities { - if let Some((op, payload)) = entity_ops.get(entity_id) { - match op { - WalOperation::Insert | WalOperation::Update => { - // Deserialize the OctadInput to determine which modalities were written - if let Ok(input) = serde_json::from_slice::<OctadInput>(payload) { - let modality_status = ModalityStatus { - graph: input.graph.is_some(), - vector: input.vector.is_some(), - document: input.document.is_some(), - tensor: input.tensor.is_some(), - semantic: input.semantic.is_some(), - temporal: true, // Always written - provenance: input.provenance.is_some(), - spatial: input.spatial.is_some(), - }; - - let id = OctadId::from(entity_id.clone()); - let now = Utc::now(); - - // Get existing version or start at 1 - let version = octads - .get(entity_id) - .map(|s| s.version + 1) - .unwrap_or(1); - - // Recover observed_at from the WAL-logged input - let observed_at = input.temporal.as_ref().map(|t| t.observed_at); - - octads.insert( - entity_id.clone(), - OctadStatus { - id, - created_at: now, - modified_at: now, - observed_at, - version, - modality_status, - }, - ); - recovered += 1; - } else { - tracing::warn!(entity_id, "WAL replay: failed to deserialize OctadInput"); - } - } - WalOperation::Delete => { - octads.remove(entity_id); - recovered += 1; - } - WalOperation::Checkpoint => {} // Already handled above - } - } - } - - info!(recovered, committed = committed_entities.len(), uncommitted = uncommitted_entities.len(), "WAL replay complete"); - - // Write a fresh checkpoint to mark recovery complete - drop(octads); // Release write lock before checkpoint - self.wal_checkpoint().await.ok(); - - Ok(recovered) - } - - /// Perform a graceful shutdown: write a final WAL checkpoint and log metrics. - /// - /// Call this before process exit to ensure all in-flight operations are - /// checkpointed. Persistent modality stores (redb) flush automatically on - /// drop, but the WAL needs an explicit final checkpoint to mark the clean - /// shutdown boundary. - pub async fn graceful_shutdown(&self) -> Result<(), OctadError> { - info!("VeriSimDB: graceful shutdown initiated"); - - // Write final WAL checkpoint - self.wal_checkpoint().await?; - - // Log final state - let octads = self.octads.read().await; - info!( - entity_count = octads.len(), - "VeriSimDB: shutdown complete — {} entities checkpointed", - octads.len() - ); - - Ok(()) - } - - /// Access the provenance store for direct queries. - pub fn provenance_store(&self) -> &Arc<P> { - &self.provenance - } - - /// Access the spatial store for direct queries. - pub fn spatial_store(&self) -> &Arc<L> { - &self.spatial - } - - /// Process graph input for a octad - async fn process_graph( - &self, - id: &OctadId, - input: &OctadGraphInput, - ) -> Result<GraphNode, OctadError> { - let node = GraphNode::new(id.to_iri(&self.config.base_iri)); - - for (predicate, target_id) in &input.relationships { - let edge = GraphEdge { - subject: node.clone(), - predicate: GraphNode::new(format!("{}/{}", self.config.base_iri, predicate)), - object: GraphObject::Node(GraphNode::new(format!( - "{}/{}", - self.config.base_iri, target_id - ))), - }; - self.graph.insert(&edge).await.map_err(|e| OctadError::ModalityError { - modality: "graph".to_string(), - message: e.to_string(), - })?; - } - - debug!(id = %id, relationships = input.relationships.len(), "Graph modality populated"); - Ok(node) - } - - /// Process vector input for a octad - async fn process_vector( - &self, - id: &OctadId, - input: &OctadVectorInput, - ) -> Result<Embedding, OctadError> { - if input.embedding.len() != self.config.vector_dimension { - return Err(OctadError::ValidationError(format!( - "Vector dimension mismatch: expected {}, got {}", - self.config.vector_dimension, - input.embedding.len() - ))); - } - - let embedding = Embedding::new(id.as_str(), input.embedding.clone()); - self.vector.upsert(&embedding).await.map_err(|e| OctadError::ModalityError { - modality: "vector".to_string(), - message: e.to_string(), - })?; - - debug!(id = %id, dimension = input.embedding.len(), "Vector modality populated"); - Ok(embedding) - } - - /// Process document input for a octad - async fn process_document( - &self, - id: &OctadId, - input: &OctadDocumentInput, - ) -> Result<Document, OctadError> { - let mut doc = Document::new(id.as_str(), &input.title, &input.body); - for (key, value) in &input.fields { - doc = doc.with_field(key, value); - } - - self.document.index(&doc).await.map_err(|e| OctadError::ModalityError { - modality: "document".to_string(), - message: e.to_string(), - })?; - self.document.commit().await.map_err(|e| OctadError::ModalityError { - modality: "document".to_string(), - message: e.to_string(), - })?; - - debug!(id = %id, title = %input.title, "Document modality populated"); - Ok(doc) - } - - /// Process tensor input for a octad - async fn process_tensor( - &self, - id: &OctadId, - input: &OctadTensorInput, - ) -> Result<Tensor, OctadError> { - let tensor = Tensor::new(id.as_str(), input.shape.clone(), input.data.clone()).map_err( - |e| OctadError::ModalityError { - modality: "tensor".to_string(), - message: e.to_string(), - }, - )?; - - self.tensor.put(&tensor).await.map_err(|e| OctadError::ModalityError { - modality: "tensor".to_string(), - message: e.to_string(), - })?; - - debug!(id = %id, shape = ?input.shape, "Tensor modality populated"); - Ok(tensor) - } - - /// Process semantic input for a octad - async fn process_semantic( - &self, - id: &OctadId, - input: &OctadSemanticInput, - ) -> Result<SemanticAnnotation, OctadError> { - let mut properties = HashMap::new(); - for (key, value) in &input.properties { - properties.insert( - key.clone(), - SemanticValue::TypedLiteral { - value: value.clone(), - datatype: "https://www.w3.org/2001/XMLSchema#string".to_string(), - }, - ); - } - - let annotation = SemanticAnnotation { - entity_id: id.as_str().to_string(), - types: input.types.clone(), - properties, - provenance: Provenance::default(), - }; - - self.semantic.annotate(&annotation).await.map_err(|e| OctadError::ModalityError { - modality: "semantic".to_string(), - message: e.to_string(), - })?; - - debug!(id = %id, types = ?input.types, "Semantic modality populated"); - Ok(annotation) - } - - /// Process provenance input for a octad — records a lineage event - async fn process_provenance( - &self, - id: &OctadId, - input: &OctadProvenanceInput, - ) -> Result<u64, OctadError> { - let event_type = match input.event_type.to_lowercase().as_str() { - "created" => ProvenanceEventType::Created, - "modified" => ProvenanceEventType::Modified, - "imported" => ProvenanceEventType::Imported, - "normalized" => ProvenanceEventType::Normalized, - "drift_repaired" => ProvenanceEventType::DriftRepaired, - "deleted" => ProvenanceEventType::Deleted, - "merged" => ProvenanceEventType::Merged, - other => ProvenanceEventType::Custom(other.to_string()), - }; - - self.provenance - .record_event(id.as_str(), event_type, &input.actor, input.source.clone(), &input.description) - .await - .map_err(|e| OctadError::ModalityError { - modality: "provenance".to_string(), - message: e.to_string(), - })?; - - let chain = self - .provenance - .get_chain(id.as_str()) - .await - .map_err(|e| OctadError::ModalityError { - modality: "provenance".to_string(), - message: e.to_string(), - })?; - - debug!(id = %id, chain_length = chain.len(), "Provenance modality populated"); - Ok(chain.len() as u64) - } - - /// Process spatial input for a octad — indexes geospatial data - async fn process_spatial( - &self, - id: &OctadId, - input: &OctadSpatialInput, - ) -> Result<SpatialData, OctadError> { - let coordinates = Coordinates::new(input.latitude, input.longitude, input.altitude) - .map_err(|e| OctadError::ValidationError(e.to_string()))?; - - let geometry_type = match input.geometry_type.as_deref() { - Some("LineString") => GeometryType::LineString, - Some("Polygon") => GeometryType::Polygon, - Some("MultiPoint") => GeometryType::MultiPoint, - Some("MultiPolygon") => GeometryType::MultiPolygon, - _ => GeometryType::Point, - }; - - let srid = input.srid.unwrap_or(4326); - - let mut data = SpatialData::with_geometry(coordinates, geometry_type, srid); - data.properties = input.properties.clone(); - - self.spatial - .index(id.as_str(), data.clone()) - .await - .map_err(|e| OctadError::ModalityError { - modality: "spatial".to_string(), - message: e.to_string(), - })?; - - debug!(id = %id, lat = input.latitude, lon = input.longitude, "Spatial modality populated"); - Ok(data) - } - - /// Roll back modality writes that succeeded before a failure. - /// - /// Called when a `create()` operation partially succeeded — some modalities - /// were written before an error occurred. This method deletes the data that - /// was already written to restore consistency. - async fn rollback_create(&self, id: &OctadId, written_modalities: &ModalityStatus) { - if written_modalities.vector { - self.vector.delete(id.as_str()).await.ok(); - } - if written_modalities.document { - self.document.delete(id.as_str()).await.ok(); - } - if written_modalities.tensor { - self.tensor.delete(id.as_str()).await.ok(); - } - // Graph and semantic don't have simple delete-by-id, - // but for atomicity we must attempt cleanup. The in-memory - // stores will GC orphaned data on next compaction. - debug!(id = %id, "Rolled back partially written modalities"); - } - - /// Create a snapshot for versioning - fn create_snapshot(&self, id: &OctadId, input: &OctadInput, status: &ModalityStatus) -> OctadSnapshot { - OctadSnapshot { - id: id.clone(), - input: input.clone(), - modality_status: status.clone(), - timestamp: Utc::now(), - } - } - - /// Load a complete Octad from all stores - async fn load_octad(&self, id: &OctadId) -> Result<Option<Octad>, OctadError> { - let octads = self.octads.read().await; - let status = match octads.get(id.as_str()) { - Some(s) => s.clone(), - None => return Ok(None), - }; - drop(octads); - - // Load each modality - let graph_node = if status.modality_status.graph { - Some(GraphNode::new(id.to_iri(&self.config.base_iri))) - } else { - None - }; - - let embedding = if status.modality_status.vector { - self.vector.get(id.as_str()).await.map_err(|e| OctadError::ModalityError { - modality: "vector".to_string(), - message: e.to_string(), - })? - } else { - None - }; - - let document = if status.modality_status.document { - self.document.get(id.as_str()).await.map_err(|e| OctadError::ModalityError { - modality: "document".to_string(), - message: e.to_string(), - })? - } else { - None - }; - - let tensor = if status.modality_status.tensor { - self.tensor.get(id.as_str()).await.map_err(|e| OctadError::ModalityError { - modality: "tensor".to_string(), - message: e.to_string(), - })? - } else { - None - }; - - let semantic = if status.modality_status.semantic { - self.semantic.get_annotations(id.as_str()).await.map_err(|e| OctadError::ModalityError { - modality: "semantic".to_string(), - message: e.to_string(), - })? - } else { - None - }; - - let version_count = self - .temporal - .history(id.as_str(), 1000) - .await - .map(|h| h.len() as u64) - .unwrap_or(0); - - // Load provenance chain length - let provenance_chain_length = if status.modality_status.provenance { - self.provenance - .get_chain(id.as_str()) - .await - .map(|c| c.len() as u64) - .unwrap_or(0) - } else { - 0 - }; - - // Load spatial data - let spatial_data = if status.modality_status.spatial { - self.spatial.get(id.as_str()).await.map_err(|e| OctadError::ModalityError { - modality: "spatial".to_string(), - message: e.to_string(), - })? - } else { - None - }; - - Ok(Some(Octad { - id: id.clone(), - status, - graph_node, - embedding, - tensor, - semantic, - document, - version_count, - provenance_chain_length, - spatial_data, - })) - } -} - -#[async_trait] -impl<G, V, D, T, S, R, P, L> OctadStore for InMemoryOctadStore<G, V, D, T, S, R, P, L> -where - G: GraphStore + 'static, - V: VectorStore + 'static, - D: DocumentStore + 'static, - T: TensorStore + 'static, - S: SemanticStore + 'static, - R: TemporalStore<Data = OctadSnapshot> + 'static, - P: ProvenanceStore + 'static, - L: SpatialStore + 'static, -{ - #[instrument(skip(self, input))] - async fn create(&self, input: OctadInput) -> Result<Octad, OctadError> { - let id = OctadId::generate(); - let now = Utc::now(); - let entity_id_str = id.as_str().to_string(); - - // Write PENDING intent to WAL before any modality writes. - // On crash recovery, PENDING entries without a matching COMMITTED - // entry indicate incomplete operations that need rollback. - let input_payload = serde_json::to_vec(&input).unwrap_or_default(); - self.wal_append(WalOperation::Insert, WalModality::All, &entity_id_str, &input_payload).await?; - - // Begin ACID transaction — acquire exclusive locks on all requested - // modalities before writing, ensuring atomicity across the octad. - let txn_id = self.txn_manager.begin(IsolationLevel::ReadCommitted).await; - - // Acquire locks for all modalities that will be written. - // This prevents concurrent writes to the same entity from interleaving. - let modality_names: Vec<&str> = [ - input.graph.as_ref().map(|_| "graph"), - input.vector.as_ref().map(|_| "vector"), - input.document.as_ref().map(|_| "document"), - input.tensor.as_ref().map(|_| "tensor"), - input.semantic.as_ref().map(|_| "semantic"), - input.provenance.as_ref().map(|_| "provenance"), - input.spatial.as_ref().map(|_| "spatial"), - Some("temporal"), // Always written (version snapshot) - ] - .into_iter() - .flatten() - .collect(); - - for modality in &modality_names { - if let Err(e) = self - .txn_manager - .acquire_lock(txn_id, &entity_id_str, modality, LockType::Exclusive) - .await - { - self.txn_manager.rollback(txn_id).await.ok(); - return Err(OctadError::ConsistencyViolation(format!( - "Failed to acquire lock on {modality}: {e}" - ))); - } - } - - // Track which modalities have been successfully written so we can - // roll back on partial failure. - let mut modality_status = ModalityStatus::default(); - - // Process each modality — on failure, rollback everything - let mut graph_node = None; - if let Some(ref graph_input) = input.graph { - match self.process_graph(&id, graph_input).await { - Ok(node) => { - graph_node = Some(node); - modality_status.graph = true; - self.txn_manager - .record_undo(txn_id, &entity_id_str, "graph", None, 0) - .await - .ok(); - } - Err(e) => { - self.rollback_create(&id, &modality_status).await; - self.txn_manager.rollback(txn_id).await.ok(); - return Err(e); - } - } - } - - let mut embedding = None; - if let Some(ref vector_input) = input.vector { - match self.process_vector(&id, vector_input).await { - Ok(emb) => { - embedding = Some(emb); - modality_status.vector = true; - self.txn_manager - .record_undo(txn_id, &entity_id_str, "vector", None, 0) - .await - .ok(); - } - Err(e) => { - self.rollback_create(&id, &modality_status).await; - self.txn_manager.rollback(txn_id).await.ok(); - return Err(e); - } - } - } - - let mut document = None; - if let Some(ref doc_input) = input.document { - match self.process_document(&id, doc_input).await { - Ok(doc) => { - document = Some(doc); - modality_status.document = true; - self.txn_manager - .record_undo(txn_id, &entity_id_str, "document", None, 0) - .await - .ok(); - } - Err(e) => { - self.rollback_create(&id, &modality_status).await; - self.txn_manager.rollback(txn_id).await.ok(); - return Err(e); - } - } - } - - let mut tensor = None; - if let Some(ref tensor_input) = input.tensor { - match self.process_tensor(&id, tensor_input).await { - Ok(t) => { - tensor = Some(t); - modality_status.tensor = true; - self.txn_manager - .record_undo(txn_id, &entity_id_str, "tensor", None, 0) - .await - .ok(); - } - Err(e) => { - self.rollback_create(&id, &modality_status).await; - self.txn_manager.rollback(txn_id).await.ok(); - return Err(e); - } - } - } - - let mut semantic = None; - if let Some(ref sem_input) = input.semantic { - match self.process_semantic(&id, sem_input).await { - Ok(ann) => { - semantic = Some(ann); - modality_status.semantic = true; - self.txn_manager - .record_undo(txn_id, &entity_id_str, "semantic", None, 0) - .await - .ok(); - } - Err(e) => { - self.rollback_create(&id, &modality_status).await; - self.txn_manager.rollback(txn_id).await.ok(); - return Err(e); - } - } - } - - // Process provenance - let mut provenance_chain_length = 0; - if let Some(ref prov_input) = input.provenance { - match self.process_provenance(&id, prov_input).await { - Ok(chain_len) => { - provenance_chain_length = chain_len; - modality_status.provenance = true; - self.txn_manager - .record_undo(txn_id, &entity_id_str, "provenance", None, 0) - .await - .ok(); - } - Err(e) => { - self.rollback_create(&id, &modality_status).await; - self.txn_manager.rollback(txn_id).await.ok(); - return Err(e); - } - } - } - - // Process spatial - let mut spatial_data = None; - if let Some(ref spatial_input) = input.spatial { - match self.process_spatial(&id, spatial_input).await { - Ok(data) => { - spatial_data = Some(data); - modality_status.spatial = true; - self.txn_manager - .record_undo(txn_id, &entity_id_str, "spatial", None, 0) - .await - .ok(); - } - Err(e) => { - self.rollback_create(&id, &modality_status).await; - self.txn_manager.rollback(txn_id).await.ok(); - return Err(e); - } - } - } - - // Create version snapshot - let snapshot = self.create_snapshot(&id, &input, &modality_status); - let version = match self - .temporal - .append(id.as_str(), snapshot, "system", Some("Initial creation")) - .await - { - Ok(v) => v, - Err(e) => { - self.rollback_create(&id, &modality_status).await; - self.txn_manager.rollback(txn_id).await.ok(); - return Err(OctadError::ModalityError { - modality: "temporal".to_string(), - message: e.to_string(), - }); - } - }; - modality_status.temporal = true; - - // All modality writes succeeded — commit the transaction - if let Err(e) = self.txn_manager.commit(txn_id).await { - self.rollback_create(&id, &modality_status).await; - return Err(OctadError::ConsistencyViolation(format!( - "Transaction commit failed: {e}" - ))); - } - - // Create status. observed_at is the territory clock (caller-supplied), - // independent of created_at which is the ingestion clock set by the DB. - let observed_at = input.temporal.as_ref().map(|t| t.observed_at); - let status = OctadStatus { - id: id.clone(), - created_at: now, - modified_at: now, - observed_at, - version, - modality_status: modality_status.clone(), - }; - - // Store in registry - self.octads.write().await.insert(id.as_str().to_string(), status.clone()); - - // Write COMMITTED marker to WAL and checkpoint for crash recovery. - self.wal_append(WalOperation::Checkpoint, WalModality::All, &entity_id_str, b"COMMITTED").await.ok(); - self.wal_checkpoint().await.ok(); - - info!(id = %id, modalities = ?modality_status, "Created octad (transaction committed)"); - - Ok(Octad { - id, - status, - graph_node, - embedding, - tensor, - semantic, - document, - version_count: 1, - provenance_chain_length, - spatial_data, - }) - } - - #[instrument(skip(self, input))] - async fn update(&self, id: &OctadId, input: OctadInput) -> Result<Octad, OctadError> { - // Check if exists - let existing = { - let octads = self.octads.read().await; - octads.get(id.as_str()).cloned() - }; - - let existing = existing.ok_or_else(|| OctadError::NotFound(id.to_string()))?; - let now = Utc::now(); - let entity_id_str = id.as_str().to_string(); - - // Write PENDING intent to WAL before modality writes - let input_payload = serde_json::to_vec(&input).unwrap_or_default(); - self.wal_append(WalOperation::Update, WalModality::All, &entity_id_str, &input_payload).await?; - - // Begin ACID transaction for atomic update across all modalities - let txn_id = self.txn_manager.begin(IsolationLevel::ReadCommitted).await; - - // Acquire exclusive locks on all modalities that will be written - let modality_names: Vec<&str> = [ - input.graph.as_ref().map(|_| "graph"), - input.vector.as_ref().map(|_| "vector"), - input.document.as_ref().map(|_| "document"), - input.tensor.as_ref().map(|_| "tensor"), - input.semantic.as_ref().map(|_| "semantic"), - input.provenance.as_ref().map(|_| "provenance"), - input.spatial.as_ref().map(|_| "spatial"), - Some("temporal"), - ] - .into_iter() - .flatten() - .collect(); - - for modality in &modality_names { - if let Err(e) = self - .txn_manager - .acquire_lock(txn_id, &entity_id_str, modality, LockType::Exclusive) - .await - { - self.txn_manager.rollback(txn_id).await.ok(); - return Err(OctadError::ConsistencyViolation(format!( - "Failed to acquire lock on {modality}: {e}" - ))); - } - } - - let mut modality_status = existing.modality_status.clone(); - - // Macro-like closure for recording undo + handling error with rollback. - // For updates, record the MVCC version so commit can detect conflicts. - let current_version = existing.version; - - // Update each modality with transactional protection - let mut graph_node = None; - if let Some(ref graph_input) = input.graph { - match self.process_graph(id, graph_input).await { - Ok(node) => { - graph_node = Some(node); - modality_status.graph = true; - self.txn_manager - .record_undo(txn_id, &entity_id_str, "graph", None, current_version) - .await - .ok(); - } - Err(e) => { - self.txn_manager.rollback(txn_id).await.ok(); - return Err(e); - } - } - } - - let mut embedding = None; - if let Some(ref vector_input) = input.vector { - match self.process_vector(id, vector_input).await { - Ok(emb) => { - embedding = Some(emb); - modality_status.vector = true; - self.txn_manager - .record_undo(txn_id, &entity_id_str, "vector", None, current_version) - .await - .ok(); - } - Err(e) => { - self.txn_manager.rollback(txn_id).await.ok(); - return Err(e); - } - } - } - - let mut document = None; - if let Some(ref doc_input) = input.document { - match self.process_document(id, doc_input).await { - Ok(doc) => { - document = Some(doc); - modality_status.document = true; - self.txn_manager - .record_undo(txn_id, &entity_id_str, "document", None, current_version) - .await - .ok(); - } - Err(e) => { - self.txn_manager.rollback(txn_id).await.ok(); - return Err(e); - } - } - } - - let mut tensor = None; - if let Some(ref tensor_input) = input.tensor { - match self.process_tensor(id, tensor_input).await { - Ok(t) => { - tensor = Some(t); - modality_status.tensor = true; - self.txn_manager - .record_undo(txn_id, &entity_id_str, "tensor", None, current_version) - .await - .ok(); - } - Err(e) => { - self.txn_manager.rollback(txn_id).await.ok(); - return Err(e); - } - } - } - - let mut semantic = None; - if let Some(ref sem_input) = input.semantic { - match self.process_semantic(id, sem_input).await { - Ok(ann) => { - semantic = Some(ann); - modality_status.semantic = true; - self.txn_manager - .record_undo(txn_id, &entity_id_str, "semantic", None, current_version) - .await - .ok(); - } - Err(e) => { - self.txn_manager.rollback(txn_id).await.ok(); - return Err(e); - } - } - } - - // Update provenance - let mut provenance_chain_length = 0; - if let Some(ref prov_input) = input.provenance { - match self.process_provenance(id, prov_input).await { - Ok(chain_len) => { - provenance_chain_length = chain_len; - modality_status.provenance = true; - self.txn_manager - .record_undo(txn_id, &entity_id_str, "provenance", None, current_version) - .await - .ok(); - } - Err(e) => { - self.txn_manager.rollback(txn_id).await.ok(); - return Err(e); - } - } - } - - // Update spatial - let mut spatial_data = None; - if let Some(ref spatial_input) = input.spatial { - match self.process_spatial(id, spatial_input).await { - Ok(data) => { - spatial_data = Some(data); - modality_status.spatial = true; - self.txn_manager - .record_undo(txn_id, &entity_id_str, "spatial", None, current_version) - .await - .ok(); - } - Err(e) => { - self.txn_manager.rollback(txn_id).await.ok(); - return Err(e); - } - } - } - - // Create new version snapshot - let snapshot = self.create_snapshot(id, &input, &modality_status); - let version = match self - .temporal - .append(id.as_str(), snapshot, "system", Some("Update")) - .await - { - Ok(v) => v, - Err(e) => { - self.txn_manager.rollback(txn_id).await.ok(); - return Err(OctadError::ModalityError { - modality: "temporal".to_string(), - message: e.to_string(), - }); - } - }; - - // All modality writes succeeded — commit the transaction - if let Err(e) = self.txn_manager.commit(txn_id).await { - return Err(OctadError::ConsistencyViolation(format!( - "Transaction commit failed: {e}" - ))); - } - - // Update status. If the caller supplied a fresh observed_at, use it; - // otherwise preserve the previous value — updates that don't touch the - // temporal modality should not erase the territory clock. - let observed_at = input - .temporal - .as_ref() - .map(|t| t.observed_at) - .or(existing.observed_at); - let status = OctadStatus { - id: id.clone(), - created_at: existing.created_at, - modified_at: now, - observed_at, - version, - modality_status: modality_status.clone(), - }; - - // Update registry - self.octads.write().await.insert(id.as_str().to_string(), status.clone()); - - // Write COMMITTED marker to WAL and checkpoint - self.wal_append(WalOperation::Checkpoint, WalModality::All, &entity_id_str, b"COMMITTED").await.ok(); - self.wal_checkpoint().await.ok(); - - info!(id = %id, version = version, "Updated octad (transaction committed)"); - - Ok(Octad { - id: id.clone(), - status, - graph_node, - embedding, - tensor, - semantic, - document, - version_count: version, - provenance_chain_length, - spatial_data, - }) - } - - async fn get(&self, id: &OctadId) -> Result<Option<Octad>, OctadError> { - self.load_octad(id).await - } - - #[instrument(skip(self))] - async fn delete(&self, id: &OctadId) -> Result<(), OctadError> { - let entity_id_str = id.as_str().to_string(); - - // Check existence before beginning transaction - let existing = { - let octads = self.octads.read().await; - octads.get(id.as_str()).cloned() - }; - - let existing = existing.ok_or_else(|| OctadError::NotFound(id.to_string()))?; - - // Write PENDING delete intent to WAL - self.wal_append(WalOperation::Delete, WalModality::All, &entity_id_str, b"").await?; - - // Begin ACID transaction for atomic delete across all modalities - let txn_id = self.txn_manager.begin(IsolationLevel::ReadCommitted).await; - - // Acquire exclusive locks on all populated modalities - let populated: Vec<&str> = [ - existing.modality_status.graph.then_some("graph"), - existing.modality_status.vector.then_some("vector"), - existing.modality_status.document.then_some("document"), - existing.modality_status.tensor.then_some("tensor"), - existing.modality_status.semantic.then_some("semantic"), - existing.modality_status.provenance.then_some("provenance"), - existing.modality_status.spatial.then_some("spatial"), - Some("temporal"), // Always exists - ] - .into_iter() - .flatten() - .collect(); - - for modality in &populated { - if let Err(e) = self - .txn_manager - .acquire_lock(txn_id, &entity_id_str, modality, LockType::Exclusive) - .await - { - self.txn_manager.rollback(txn_id).await.ok(); - return Err(OctadError::ConsistencyViolation(format!( - "Failed to acquire lock on {modality} for delete: {e}" - ))); - } - } - - // Record undo entries for populated modalities so the transaction - // manager tracks the scope of this delete for version bookkeeping. - for modality in &populated { - self.txn_manager - .record_undo(txn_id, &entity_id_str, modality, None, existing.version) - .await - .ok(); - } - - // Delete from each modality store - // Note: We don't delete from temporal to preserve history - self.vector.delete(id.as_str()).await.ok(); - self.document.delete(id.as_str()).await.ok(); - self.tensor.delete(id.as_str()).await.ok(); - // Graph and semantic don't have simple delete-by-id - - // Commit the transaction - if let Err(e) = self.txn_manager.commit(txn_id).await { - return Err(OctadError::ConsistencyViolation(format!( - "Transaction commit failed during delete: {e}" - ))); - } - - // Remove from registry only after successful commit - self.octads.write().await.remove(id.as_str()); - - // Write COMMITTED marker to WAL and checkpoint - self.wal_append(WalOperation::Checkpoint, WalModality::All, &entity_id_str, b"COMMITTED").await.ok(); - self.wal_checkpoint().await.ok(); - - info!(id = %id, "Deleted octad (transaction committed)"); - Ok(()) - } - - async fn status(&self, id: &OctadId) -> Result<Option<OctadStatus>, OctadError> { - Ok(self.octads.read().await.get(id.as_str()).cloned()) - } - - async fn search_similar(&self, embedding: &[f32], k: usize) -> Result<Vec<Octad>, OctadError> { - let results = self.vector.search(embedding, k).await.map_err(|e| OctadError::ModalityError { - modality: "vector".to_string(), - message: e.to_string(), - })?; - - let mut octads = Vec::new(); - for result in results { - if let Some(octad) = self.load_octad(&OctadId::new(&result.id)).await? { - octads.push(octad); - } - } - - Ok(octads) - } - - async fn search_text(&self, query: &str, limit: usize) -> Result<Vec<Octad>, OctadError> { - let results = - self.document.search(query, limit).await.map_err(|e| OctadError::ModalityError { - modality: "document".to_string(), - message: e.to_string(), - })?; - - let mut octads = Vec::new(); - for result in results { - if let Some(octad) = self.load_octad(&OctadId::new(&result.id)).await? { - octads.push(octad); - } - } - - Ok(octads) - } - - async fn query_related(&self, id: &OctadId, predicate: &str) -> Result<Vec<Octad>, OctadError> { - let node = GraphNode::new(id.to_iri(&self.config.base_iri)); - let edges = self.graph.outgoing(&node).await.map_err(|e| OctadError::ModalityError { - modality: "graph".to_string(), - message: e.to_string(), - })?; - - let predicate_iri = format!("{}/{}", self.config.base_iri, predicate); - let mut octads = Vec::new(); - - for edge in edges { - if edge.predicate.iri == predicate_iri { - if let GraphObject::Node(target) = edge.object { - // Extract ID from IRI - let target_id = target - .iri - .strip_prefix(&format!("{}/", self.config.base_iri)) - .unwrap_or(&target.iri); - - if let Some(octad) = self.load_octad(&OctadId::new(target_id)).await? { - octads.push(octad); - } - } - } - } - - Ok(octads) - } - - async fn list(&self, limit: usize, offset: usize) -> Result<Vec<Octad>, OctadError> { - let octads = self.octads.read().await; - let ids: Vec<String> = octads - .keys() - .skip(offset) - .take(limit) - .cloned() - .collect(); - drop(octads); - - let mut result = Vec::with_capacity(ids.len()); - for id_str in ids { - if let Some(octad) = self.load_octad(&OctadId::new(&id_str)).await? { - result.push(octad); - } - } - Ok(result) - } - - async fn at_time(&self, id: &OctadId, time: DateTime<Utc>) -> Result<Option<Octad>, OctadError> { - let version = self - .temporal - .at_time(id.as_str(), time) - .await - .map_err(|e| OctadError::ModalityError { - modality: "temporal".to_string(), - message: e.to_string(), - })?; - - match version { - Some(v) => { - // Reconstruct octad from snapshot - // For now, we just return current state with version info - // A full implementation would restore from snapshot - let mut octad = self.load_octad(id).await?; - if let Some(ref mut h) = octad { - h.status.version = v.version; - h.status.modified_at = v.timestamp; - } - Ok(octad) - } - None => Ok(None), - } - } -} - -#[cfg(test)] -mod tests { - use super::*; - use crate::OctadBuilder; - use verisim_document::TantivyDocumentStore; - use verisim_graph::SimpleGraphStore; - use verisim_provenance::InMemoryProvenanceStore; - use verisim_semantic::InMemorySemanticStore; - use verisim_spatial::InMemorySpatialStore; - use verisim_temporal::InMemoryVersionStore; - use verisim_tensor::InMemoryTensorStore; - use verisim_vector::{DistanceMetric, BruteForceVectorStore}; - - fn create_test_store() -> InMemoryOctadStore< - SimpleGraphStore, - BruteForceVectorStore, - TantivyDocumentStore, - InMemoryTensorStore, - InMemorySemanticStore, - InMemoryVersionStore<OctadSnapshot>, - InMemoryProvenanceStore, - InMemorySpatialStore, - > { - let config = OctadConfig { - vector_dimension: 3, - ..Default::default() - }; - - InMemoryOctadStore::new( - config, - Arc::new(SimpleGraphStore::in_memory().expect("TODO: handle error")), - Arc::new(BruteForceVectorStore::new(3, DistanceMetric::Cosine)), - Arc::new(TantivyDocumentStore::in_memory().expect("TODO: handle error")), - Arc::new(InMemoryTensorStore::new()), - Arc::new(InMemorySemanticStore::new()), - Arc::new(InMemoryVersionStore::new()), - Arc::new(InMemoryProvenanceStore::new()), - Arc::new(InMemorySpatialStore::new()), - ) - } - - #[tokio::test] - async fn test_create_and_get_octad() { - let store = create_test_store(); - - let input = OctadBuilder::new() - .with_document("Test Document", "This is a test body") - .with_embedding(vec![0.1, 0.2, 0.3]) - .build(); - - let octad = store.create(input).await.expect("TODO: handle error"); - assert!(octad.status.modality_status.document); - assert!(octad.status.modality_status.vector); - assert!(octad.status.modality_status.temporal); - - let retrieved = store.get(&octad.id).await.expect("TODO: handle error"); - assert!(retrieved.is_some()); - let retrieved = retrieved.expect("TODO: handle error"); - assert_eq!(retrieved.id, octad.id); - } - - #[tokio::test] - async fn test_vector_search() { - let store = create_test_store(); - - let input1 = OctadBuilder::new() - .with_document("First", "First document") - .with_embedding(vec![1.0, 0.0, 0.0]) - .build(); - - let input2 = OctadBuilder::new() - .with_document("Second", "Second document") - .with_embedding(vec![0.9, 0.1, 0.0]) - .build(); - - let input3 = OctadBuilder::new() - .with_document("Third", "Third document") - .with_embedding(vec![0.0, 1.0, 0.0]) - .build(); - - store.create(input1).await.expect("TODO: handle error"); - store.create(input2).await.expect("TODO: handle error"); - store.create(input3).await.expect("TODO: handle error"); - - let results = store.search_similar(&[1.0, 0.0, 0.0], 2).await.expect("TODO: handle error"); - assert_eq!(results.len(), 2); - } - - #[tokio::test] - async fn test_document_search() { - let store = create_test_store(); - - let input1 = OctadBuilder::new() - .with_document("Rust Programming", "Rust is a systems programming language") - .with_embedding(vec![0.1, 0.2, 0.3]) - .build(); - - let input2 = OctadBuilder::new() - .with_document("Python Tutorial", "Python is great for beginners") - .with_embedding(vec![0.4, 0.5, 0.6]) - .build(); - - store.create(input1).await.expect("TODO: handle error"); - store.create(input2).await.expect("TODO: handle error"); - - let results = store.search_text("Rust", 10).await.expect("TODO: handle error"); - assert_eq!(results.len(), 1); - assert!(results[0].document.as_ref().expect("TODO: handle error").title.contains("Rust")); - } - - #[tokio::test] - async fn test_update_octad() { - let store = create_test_store(); - - let input = OctadBuilder::new() - .with_document("Original", "Original content") - .with_embedding(vec![0.1, 0.2, 0.3]) - .build(); - - let octad = store.create(input).await.expect("TODO: handle error"); - assert_eq!(octad.status.version, 1); - - let update_input = OctadBuilder::new() - .with_document("Updated", "Updated content") - .build(); - - let updated = store.update(&octad.id, update_input).await.expect("TODO: handle error"); - assert_eq!(updated.status.version, 2); - assert!(updated.document.as_ref().expect("TODO: handle error").title.contains("Updated")); - } - - #[tokio::test] - async fn test_observed_at_round_trip_and_preservation() { - // Veridical-simulation gap closure: an entity's territory clock - // (observed_at) must survive a create/get and must persist across - // updates that don't touch the temporal modality. - let store = create_test_store(); - - let observed = chrono::DateTime::parse_from_rfc3339("2026-04-27T15:30:00Z") - .expect("static RFC3339 literal parses") - .with_timezone(&Utc); - - let create_input = OctadBuilder::new() - .with_document("Email", "Body") - .with_embedding(vec![0.1, 0.2, 0.3]) - .with_observed_at(observed) - .build(); - let octad = store.create(create_input).await.expect("create succeeds"); - - assert_eq!( - octad.status.observed_at, - Some(observed), - "observed_at must survive create" - ); - - let fetched = store - .status(&octad.id) - .await - .expect("status succeeds") - .expect("status present"); - assert_eq!(fetched.observed_at, Some(observed), "observed_at survives status read"); - - let update_no_temporal = OctadBuilder::new() - .with_document("Email v2", "Body v2") - .build(); - let updated = store - .update(&octad.id, update_no_temporal) - .await - .expect("update succeeds"); - assert_eq!( - updated.status.observed_at, - Some(observed), - "update without temporal input must preserve prior observed_at" - ); - - let new_observed = chrono::DateTime::parse_from_rfc3339("2026-04-27T16:00:00Z") - .expect("static RFC3339 literal parses") - .with_timezone(&Utc); - let update_with_temporal = OctadBuilder::new() - .with_document("Email v3", "Body v3") - .with_observed_at(new_observed) - .build(); - let updated2 = store - .update(&octad.id, update_with_temporal) - .await - .expect("update succeeds"); - assert_eq!( - updated2.status.observed_at, - Some(new_observed), - "update with temporal input must override prior observed_at" - ); - } - - #[tokio::test] - async fn test_observed_at_absent_when_not_supplied() { - let store = create_test_store(); - let input = OctadBuilder::new() - .with_document("No clock", "No clock body") - .with_embedding(vec![0.4, 0.5, 0.6]) - .build(); - let octad = store.create(input).await.expect("create succeeds"); - assert!( - octad.status.observed_at.is_none(), - "observed_at must be None when caller did not supply temporal input" - ); - } -} diff --git a/verisimdb/rust-core/verisim-octad/src/transaction.rs b/verisimdb/rust-core/verisim-octad/src/transaction.rs deleted file mode 100644 index f4cd4daa..00000000 --- a/verisimdb/rust-core/verisim-octad/src/transaction.rs +++ /dev/null @@ -1,1446 +0,0 @@ -// SPDX-License-Identifier: MPL-2.0 -// Copyright (c) 2026 Jonathan D.A. Jewell (hyperpolymath) <j.d.a.jewell@open.ac.uk> - -//! ACID Transaction Manager for VeriSimDB Octad Operations -//! -//! Provides cross-modality atomicity for octad operations. A octad update must -//! either succeed across all 8 modalities (octad) or fail completely, -//! preserving the fundamental consistency guarantee of the octad model. -//! -//! # Architecture -//! -//! The transaction manager implements: -//! - **Atomicity**: Undo log records previous state for rollback across modalities -//! - **Consistency**: Modality-level locks prevent partial updates -//! - **Isolation**: MVCC with configurable isolation levels (ReadCommitted, Serializable) -//! - **Durability**: Delegated to the underlying modality stores -//! -//! # Transaction State Machine -//! -//! ```text -//! ┌────────┐ begin() ┌────────┐ commit() ┌───────────┐ -//! │ None │ ──────────>│ Active │ ──────────> │ Committed │ -//! └────────┘ └────────┘ └───────────┘ -//! │ -//! │ rollback() -//! ▼ -//! ┌────────────┐ -//! │ RolledBack │ -//! └────────────┘ -//! ``` -//! -//! # Deadlock Detection -//! -//! Uses a wait-for graph at modality granularity. When a transaction requests a -//! lock held by another transaction, the manager checks for cycles. If a cycle -//! is detected, the requesting transaction is aborted to break the deadlock. - -use chrono::{DateTime, Utc}; -use std::collections::{HashMap, HashSet, VecDeque}; -use std::fmt; -use std::sync::Arc; -use tokio::sync::RwLock; -use tracing::{debug, info, warn}; -use uuid::Uuid; - -// --------------------------------------------------------------------------- -// Constants: the eight VeriSimDB modalities (octad) -// --------------------------------------------------------------------------- - -/// All eight modalities that a octad spans (the octad). -/// -/// Originally six modalities; extended with `provenance` and `spatial` when -/// VeriSimDB evolved from octad to octad model. -pub const MODALITIES: &[&str] = &[ - "graph", "vector", "tensor", "semantic", "document", "temporal", - "provenance", "spatial", -]; - -// --------------------------------------------------------------------------- -// Error types -// --------------------------------------------------------------------------- - -/// Errors that can occur during transaction processing. -#[derive(Debug, Clone, PartialEq, Eq)] -pub enum TransactionError { - /// The transaction is not in the expected state for the requested operation. - InvalidState { - transaction_id: Uuid, - current: TransactionState, - expected: &'static str, - }, - /// A lock could not be acquired because it conflicts with an existing lock. - LockConflict { - entity_id: String, - modality: String, - held_by: Uuid, - }, - /// A deadlock cycle was detected; the requesting transaction must abort. - DeadlockDetected { - transaction_id: Uuid, - cycle: Vec<Uuid>, - }, - /// The requested transaction does not exist. - TransactionNotFound(Uuid), - /// A modality name is not one of the eight valid modalities. - InvalidModality(String), - /// MVCC version conflict: the entity was modified after the transaction read it. - VersionConflict { - entity_id: String, - expected_version: u64, - actual_version: u64, - }, -} - -impl fmt::Display for TransactionError { - fn fmt(&self, f: &mut fmt::Formatter<'_>) -> fmt::Result { - match self { - Self::InvalidState { - transaction_id, - current, - expected, - } => write!( - f, - "Transaction {transaction_id}: state is {current:?}, expected {expected}" - ), - Self::LockConflict { - entity_id, - modality, - held_by, - } => write!( - f, - "Lock conflict on {entity_id}/{modality} held by {held_by}" - ), - Self::DeadlockDetected { - transaction_id, - cycle, - } => write!( - f, - "Deadlock detected for {transaction_id}: cycle {:?}", - cycle - ), - Self::TransactionNotFound(id) => write!(f, "Transaction not found: {id}"), - Self::InvalidModality(m) => write!(f, "Invalid modality: {m}"), - Self::VersionConflict { - entity_id, - expected_version, - actual_version, - } => write!( - f, - "Version conflict on {entity_id}: expected v{expected_version}, found v{actual_version}" - ), - } - } -} - -impl std::error::Error for TransactionError {} - -// --------------------------------------------------------------------------- -// Core types -// --------------------------------------------------------------------------- - -/// The lifecycle state of a transaction. -#[derive(Debug, Clone, Copy, PartialEq, Eq, Hash)] -pub enum TransactionState { - /// The transaction is open and accepting operations. - Active, - /// The transaction has been successfully committed. - Committed, - /// The transaction has been rolled back (either explicitly or due to error). - RolledBack, -} - -/// Transaction isolation level. -/// -/// Determines what data other concurrent transactions can see. -#[derive(Debug, Clone, Copy, PartialEq, Eq, Hash)] -pub enum IsolationLevel { - /// Reads only committed data. A transaction may see different snapshots - /// of the same entity if another transaction commits between reads. - ReadCommitted, - /// Full serializability: the transaction operates as if it were the only - /// one running. Version conflicts cause the transaction to abort. - Serializable, -} - -impl Default for IsolationLevel { - fn default() -> Self { - Self::ReadCommitted - } -} - -/// Type of lock held on an entity/modality pair. -#[derive(Debug, Clone, Copy, PartialEq, Eq, Hash)] -pub enum LockType { - /// Multiple transactions can hold shared locks concurrently (for reads). - Shared, - /// Only one transaction can hold an exclusive lock (for writes). - Exclusive, -} - -/// A lock held by a transaction on a specific entity/modality pair. -#[derive(Debug, Clone, PartialEq, Eq, Hash)] -pub struct LockEntry { - /// The entity being locked. - pub entity_id: String, - /// The modality being locked (one of the six). - pub modality: String, - /// Whether the lock is shared or exclusive. - pub lock_type: LockType, -} - -/// An entry in the undo log, recording the previous state of an -/// entity/modality pair so that rollback can restore it. -#[derive(Debug, Clone)] -pub struct UndoEntry { - /// The entity that was modified. - pub entity_id: String, - /// The modality that was modified. - pub modality: String, - /// The serialized previous data, or `None` if the entity/modality did not - /// exist before this transaction touched it (i.e., it was a new insert). - pub previous_data: Option<Vec<u8>>, - /// The version of the entity before modification (for MVCC validation). - pub previous_version: u64, - /// When this undo entry was recorded. - pub recorded_at: DateTime<Utc>, -} - -/// A version-stamped read performed by a Serializable transaction, used for -/// validation at commit time. -#[derive(Debug, Clone)] -pub struct ReadStamp { - /// The entity that was read. - pub entity_id: String, - /// The modality that was read. - pub modality: String, - /// The version observed at read time. - pub version_at_read: u64, -} - -/// A single transaction, tracking its state, undo log, locks, and read set. -#[derive(Debug)] -pub struct Transaction { - /// Unique transaction identifier. - pub id: Uuid, - /// Current lifecycle state. - pub state: TransactionState, - /// Isolation level for this transaction. - pub isolation_level: IsolationLevel, - /// Ordered log of changes for rollback (applied in reverse on rollback). - pub undo_log: Vec<UndoEntry>, - /// Set of locks currently held by this transaction. - pub locks: Vec<LockEntry>, - /// Read set for Serializable validation (entity/modality -> version read). - pub read_set: Vec<ReadStamp>, - /// When the transaction was started. - pub started_at: DateTime<Utc>, - /// When the transaction was completed (committed or rolled back). - pub completed_at: Option<DateTime<Utc>>, -} - -impl Transaction { - /// Create a new transaction in the Active state. - fn new(isolation_level: IsolationLevel) -> Self { - Self { - id: Uuid::new_v4(), - state: TransactionState::Active, - isolation_level, - undo_log: Vec::new(), - locks: Vec::new(), - read_set: Vec::new(), - started_at: Utc::now(), - completed_at: None, - } - } - - /// Return `true` if the transaction is still active. - pub fn is_active(&self) -> bool { - self.state == TransactionState::Active - } -} - -// --------------------------------------------------------------------------- -// Lock table -// --------------------------------------------------------------------------- - -/// Key for the lock table: (entity_id, modality). -#[derive(Debug, Clone, PartialEq, Eq, Hash)] -struct LockKey { - entity_id: String, - modality: String, -} - -/// Information about a lock held in the global lock table. -#[derive(Debug, Clone)] -struct LockInfo { - /// Which transaction holds the lock. - holder: Uuid, - /// Lock type (Shared or Exclusive). - lock_type: LockType, -} - -/// Global lock table shared across all transactions. -/// -/// The table maps (entity_id, modality) pairs to the set of locks held on them. -/// Shared locks allow multiple holders; exclusive locks allow exactly one. -#[derive(Debug)] -pub struct LockTable { - /// Map of lock key to the list of holders. - locks: HashMap<LockKey, Vec<LockInfo>>, - /// Wait-for graph: transaction A waits for transaction B. - /// Used for deadlock detection. - wait_for: HashMap<Uuid, HashSet<Uuid>>, -} - -impl LockTable { - /// Create an empty lock table. - fn new() -> Self { - Self { - locks: HashMap::new(), - wait_for: HashMap::new(), - } - } - - /// Attempt to acquire a lock. Returns `Ok(())` on success, or a - /// `LockConflict` / `DeadlockDetected` error on failure. - fn acquire( - &mut self, - transaction_id: Uuid, - entity_id: &str, - modality: &str, - lock_type: LockType, - ) -> Result<(), TransactionError> { - let key = LockKey { - entity_id: entity_id.to_string(), - modality: modality.to_string(), - }; - - // Phase 1: Inspect existing holders (immutable read) and determine action. - // We collect all the information we need before mutating anything, so the - // borrow checker is satisfied when we later call detect_cycle. - enum AcquireAction { - /// Lock is already held by this transaction with compatible type. - AlreadyHeld, - /// Upgrade from Shared to Exclusive (sole holder). - Upgrade, - /// Cannot upgrade because other holders exist. - UpgradeConflict { other_holder: Uuid }, - /// Conflict with another transaction's lock. - Conflict { blocker: Uuid }, - /// No existing holders or only compatible shared locks; grant directly. - Grant, - } - - let action = { - let holders = self.locks.get(&key).map(|v| v.as_slice()).unwrap_or(&[]); - let mut result = AcquireAction::Grant; - - for info in holders { - if info.holder == transaction_id { - if info.lock_type == LockType::Shared && lock_type == LockType::Exclusive { - if holders.len() == 1 { - result = AcquireAction::Upgrade; - } else { - let other = holders - .iter() - .find(|h| h.holder != transaction_id) - .map(|h| h.holder) - .unwrap_or(transaction_id); - result = AcquireAction::UpgradeConflict { - other_holder: other, - }; - } - } else { - result = AcquireAction::AlreadyHeld; - } - break; - } - - let conflicts = !matches!( - (info.lock_type, lock_type), - (LockType::Shared, LockType::Shared) - ); - - if conflicts { - result = AcquireAction::Conflict { - blocker: info.holder, - }; - break; - } - } - - result - }; - - // Phase 2: Act on the determined action. - match action { - AcquireAction::AlreadyHeld => return Ok(()), - AcquireAction::UpgradeConflict { other_holder } => { - return Err(TransactionError::LockConflict { - entity_id: entity_id.to_string(), - modality: modality.to_string(), - held_by: other_holder, - }); - } - AcquireAction::Conflict { blocker } => { - // Register in wait-for graph for deadlock detection - self.wait_for - .entry(transaction_id) - .or_default() - .insert(blocker); - - // Check for deadlock cycle (no mutable borrow on self.locks) - if let Some(cycle) = self.detect_cycle(transaction_id) { - self.wait_for.remove(&transaction_id); - return Err(TransactionError::DeadlockDetected { - transaction_id, - cycle, - }); - } - - return Err(TransactionError::LockConflict { - entity_id: entity_id.to_string(), - modality: modality.to_string(), - held_by: blocker, - }); - } - AcquireAction::Upgrade | AcquireAction::Grant => { - // Proceed to grant/upgrade below - } - } - - // Phase 3: Mutate the lock table to grant or upgrade the lock. - let holders = self.locks.entry(key).or_default(); - - // Remove any existing lock by this transaction on this key (for upgrades) - holders.retain(|info| info.holder != transaction_id); - - // Grant the lock - holders.push(LockInfo { - holder: transaction_id, - lock_type, - }); - - // Clear any wait-for edges from this transaction (lock acquired) - self.wait_for.remove(&transaction_id); - - debug!( - transaction = %transaction_id, - entity = entity_id, - modality = modality, - lock = ?lock_type, - "Lock acquired" - ); - - Ok(()) - } - - /// Release all locks held by the given transaction. - fn release_all(&mut self, transaction_id: Uuid) { - // Remove from all lock entries - self.locks.retain(|_key, holders| { - holders.retain(|info| info.holder != transaction_id); - !holders.is_empty() - }); - - // Remove from wait-for graph - self.wait_for.remove(&transaction_id); - for waiters in self.wait_for.values_mut() { - waiters.remove(&transaction_id); - } - - debug!(transaction = %transaction_id, "All locks released"); - } - - /// Detect a cycle in the wait-for graph starting from `start`. - /// Returns `Some(cycle)` if a cycle is found, `None` otherwise. - /// Uses BFS to find the shortest cycle. - fn detect_cycle(&self, start: Uuid) -> Option<Vec<Uuid>> { - let mut visited = HashSet::new(); - let mut queue: VecDeque<Vec<Uuid>> = VecDeque::new(); - - // Seed the BFS with the direct dependencies of `start` - if let Some(neighbors) = self.wait_for.get(&start) { - for &neighbor in neighbors { - queue.push_back(vec![start, neighbor]); - } - } - - while let Some(path) = queue.pop_front() { - let Some(¤t) = path.last() else { - // Invariant: BFS paths are always non-empty (seeded with [start, neighbor]). - // Defensive guard — skip rather than panic in the deadlock detector. - continue; - }; - - if current == start { - // Found a cycle - return Some(path); - } - - if !visited.insert(current) { - continue; - } - - if let Some(neighbors) = self.wait_for.get(¤t) { - for &neighbor in neighbors { - let mut next_path = path.clone(); - next_path.push(neighbor); - queue.push_back(next_path); - } - } - } - - None - } - - /// Return `true` if any lock is held on the given entity/modality pair. - fn is_locked(&self, entity_id: &str, modality: &str) -> bool { - let key = LockKey { - entity_id: entity_id.to_string(), - modality: modality.to_string(), - }; - self.locks.get(&key).is_some_and(|h| !h.is_empty()) - } - - /// Return the set of lock holders for a given entity/modality pair. - fn holders(&self, entity_id: &str, modality: &str) -> Vec<(Uuid, LockType)> { - let key = LockKey { - entity_id: entity_id.to_string(), - modality: modality.to_string(), - }; - self.locks - .get(&key) - .map(|holders| { - holders - .iter() - .map(|info| (info.holder, info.lock_type)) - .collect() - }) - .unwrap_or_default() - } -} - -// --------------------------------------------------------------------------- -// Version table (MVCC) -// --------------------------------------------------------------------------- - -/// Tracks the current version of each entity/modality pair for MVCC. -#[derive(Debug, Default)] -struct VersionTable { - /// Map from (entity_id, modality) to the current committed version. - versions: HashMap<(String, String), u64>, -} - -impl VersionTable { - /// Get the current version for an entity/modality pair. - /// Returns 0 if no version has been recorded. - fn get(&self, entity_id: &str, modality: &str) -> u64 { - self.versions - .get(&(entity_id.to_string(), modality.to_string())) - .copied() - .unwrap_or(0) - } - - /// Set the version for an entity/modality pair. - fn set(&mut self, entity_id: &str, modality: &str, version: u64) { - self.versions.insert( - (entity_id.to_string(), modality.to_string()), - version, - ); - } - - /// Increment and return the new version for an entity/modality pair. - fn increment(&mut self, entity_id: &str, modality: &str) -> u64 { - let key = (entity_id.to_string(), modality.to_string()); - let next = self.versions.get(&key).copied().unwrap_or(0) + 1; - self.versions.insert(key, next); - next - } -} - -// --------------------------------------------------------------------------- -// Transaction Manager -// --------------------------------------------------------------------------- - -/// The central transaction manager for VeriSimDB octad operations. -/// -/// Coordinates transactions, locks, and MVCC versioning across the eight -/// modalities (octad). All public methods are `async` and thread-safe via -/// interior `RwLock`s. -pub struct TransactionManager { - /// Active and recently completed transactions. - active_transactions: Arc<RwLock<HashMap<Uuid, Transaction>>>, - /// Global lock table for modality-level locking. - lock_table: Arc<RwLock<LockTable>>, - /// MVCC version tracking. - version_table: Arc<RwLock<VersionTable>>, -} - -impl TransactionManager { - /// Create a new transaction manager with empty state. - pub fn new() -> Self { - Self { - active_transactions: Arc::new(RwLock::new(HashMap::new())), - lock_table: Arc::new(RwLock::new(LockTable::new())), - version_table: Arc::new(RwLock::new(VersionTable::default())), - } - } - - /// Begin a new transaction with the given isolation level. - /// - /// Returns the unique transaction ID that must be used for all subsequent - /// operations within this transaction. - pub async fn begin(&self, isolation_level: IsolationLevel) -> Uuid { - let txn = Transaction::new(isolation_level); - let txn_id = txn.id; - - info!( - transaction = %txn_id, - isolation = ?isolation_level, - "Transaction started" - ); - - self.active_transactions - .write() - .await - .insert(txn_id, txn); - - txn_id - } - - /// Commit a transaction, making all its changes permanent. - /// - /// For Serializable transactions, this validates that no entity/modality - /// pair read during the transaction has been modified by another committed - /// transaction (write-skew detection). - pub async fn commit(&self, transaction_id: Uuid) -> Result<(), TransactionError> { - let mut txns = self.active_transactions.write().await; - let txn = txns - .get_mut(&transaction_id) - .ok_or(TransactionError::TransactionNotFound(transaction_id))?; - - if txn.state != TransactionState::Active { - return Err(TransactionError::InvalidState { - transaction_id, - current: txn.state, - expected: "Active", - }); - } - - // Serializable validation: check read set against current versions - if txn.isolation_level == IsolationLevel::Serializable { - let version_table = self.version_table.read().await; - for stamp in &txn.read_set { - let current_version = - version_table.get(&stamp.entity_id, &stamp.modality); - if current_version != stamp.version_at_read { - // Another transaction committed a write to this entity/modality - // after we read it. Must abort. - drop(version_table); - - // Roll back instead of committing - txn.state = TransactionState::RolledBack; - txn.completed_at = Some(Utc::now()); - - // Release locks - let mut lock_table = self.lock_table.write().await; - lock_table.release_all(transaction_id); - - warn!( - transaction = %transaction_id, - entity = %stamp.entity_id, - modality = %stamp.modality, - read_version = stamp.version_at_read, - current_version = current_version, - "Serializable validation failed" - ); - - return Err(TransactionError::VersionConflict { - entity_id: stamp.entity_id.clone(), - expected_version: stamp.version_at_read, - actual_version: current_version, - }); - } - } - } - - // Apply version increments for all writes in the undo log - { - let mut version_table = self.version_table.write().await; - for entry in &txn.undo_log { - version_table.increment(&entry.entity_id, &entry.modality); - } - } - - txn.state = TransactionState::Committed; - txn.completed_at = Some(Utc::now()); - - // Release all locks - { - let mut lock_table = self.lock_table.write().await; - lock_table.release_all(transaction_id); - } - - info!( - transaction = %transaction_id, - undo_entries = txn.undo_log.len(), - "Transaction committed" - ); - - Ok(()) - } - - /// Roll back a transaction, undoing all recorded changes. - /// - /// Returns the undo log entries in reverse order so the caller can apply - /// compensating actions to the modality stores. - pub async fn rollback( - &self, - transaction_id: Uuid, - ) -> Result<Vec<UndoEntry>, TransactionError> { - let mut txns = self.active_transactions.write().await; - let txn = txns - .get_mut(&transaction_id) - .ok_or(TransactionError::TransactionNotFound(transaction_id))?; - - if txn.state != TransactionState::Active { - return Err(TransactionError::InvalidState { - transaction_id, - current: txn.state, - expected: "Active", - }); - } - - txn.state = TransactionState::RolledBack; - txn.completed_at = Some(Utc::now()); - - // Collect undo entries in reverse order for the caller to apply - let mut undo_entries: Vec<UndoEntry> = txn.undo_log.clone(); - undo_entries.reverse(); - - // Release all locks - { - let mut lock_table = self.lock_table.write().await; - lock_table.release_all(transaction_id); - } - - info!( - transaction = %transaction_id, - undo_entries = undo_entries.len(), - "Transaction rolled back" - ); - - Ok(undo_entries) - } - - /// Acquire a lock on an entity/modality pair within a transaction. - /// - /// Validates the modality name and checks for deadlocks before granting. - pub async fn acquire_lock( - &self, - transaction_id: Uuid, - entity_id: &str, - modality: &str, - lock_type: LockType, - ) -> Result<(), TransactionError> { - // Validate modality name - if !MODALITIES.contains(&modality) { - return Err(TransactionError::InvalidModality(modality.to_string())); - } - - // Verify transaction is active - { - let txns = self.active_transactions.read().await; - let txn = txns - .get(&transaction_id) - .ok_or(TransactionError::TransactionNotFound(transaction_id))?; - - if txn.state != TransactionState::Active { - return Err(TransactionError::InvalidState { - transaction_id, - current: txn.state, - expected: "Active", - }); - } - } - - // Acquire in the global lock table - { - let mut lock_table = self.lock_table.write().await; - lock_table.acquire(transaction_id, entity_id, modality, lock_type)?; - } - - // Record the lock in the transaction's lock set - { - let mut txns = self.active_transactions.write().await; - if let Some(txn) = txns.get_mut(&transaction_id) { - let entry = LockEntry { - entity_id: entity_id.to_string(), - modality: modality.to_string(), - lock_type, - }; - // Avoid duplicates (upgrade replaces) - txn.locks.retain(|l| { - !(l.entity_id == entry.entity_id && l.modality == entry.modality) - }); - txn.locks.push(entry); - } - } - - Ok(()) - } - - /// Record an undo entry for a write operation within a transaction. - /// - /// The caller is responsible for serializing the previous data (if any) - /// before passing it here. - pub async fn record_undo( - &self, - transaction_id: Uuid, - entity_id: &str, - modality: &str, - previous_data: Option<Vec<u8>>, - previous_version: u64, - ) -> Result<(), TransactionError> { - if !MODALITIES.contains(&modality) { - return Err(TransactionError::InvalidModality(modality.to_string())); - } - - let mut txns = self.active_transactions.write().await; - let txn = txns - .get_mut(&transaction_id) - .ok_or(TransactionError::TransactionNotFound(transaction_id))?; - - if txn.state != TransactionState::Active { - return Err(TransactionError::InvalidState { - transaction_id, - current: txn.state, - expected: "Active", - }); - } - - txn.undo_log.push(UndoEntry { - entity_id: entity_id.to_string(), - modality: modality.to_string(), - previous_data, - previous_version, - recorded_at: Utc::now(), - }); - - debug!( - transaction = %transaction_id, - entity = entity_id, - modality = modality, - "Undo entry recorded" - ); - - Ok(()) - } - - /// Record a read stamp for Serializable validation. - /// - /// Called when a Serializable transaction reads an entity/modality pair. - /// The version at read time is recorded so that commit-time validation - /// can detect write-skew. - pub async fn record_read( - &self, - transaction_id: Uuid, - entity_id: &str, - modality: &str, - ) -> Result<(), TransactionError> { - if !MODALITIES.contains(&modality) { - return Err(TransactionError::InvalidModality(modality.to_string())); - } - - let current_version = { - let vt = self.version_table.read().await; - vt.get(entity_id, modality) - }; - - let mut txns = self.active_transactions.write().await; - let txn = txns - .get_mut(&transaction_id) - .ok_or(TransactionError::TransactionNotFound(transaction_id))?; - - if txn.state != TransactionState::Active { - return Err(TransactionError::InvalidState { - transaction_id, - current: txn.state, - expected: "Active", - }); - } - - txn.read_set.push(ReadStamp { - entity_id: entity_id.to_string(), - modality: modality.to_string(), - version_at_read: current_version, - }); - - Ok(()) - } - - /// Get a snapshot of a transaction's current state. - /// - /// Returns `None` if the transaction does not exist. - pub async fn get_transaction_state( - &self, - transaction_id: Uuid, - ) -> Option<TransactionState> { - self.active_transactions - .read() - .await - .get(&transaction_id) - .map(|txn| txn.state) - } - - /// Get the number of currently active transactions. - pub async fn active_count(&self) -> usize { - self.active_transactions - .read() - .await - .values() - .filter(|txn| txn.state == TransactionState::Active) - .count() - } - - /// Get the current MVCC version for an entity/modality pair. - pub async fn current_version(&self, entity_id: &str, modality: &str) -> u64 { - self.version_table.read().await.get(entity_id, modality) - } - - /// Set the MVCC version for an entity/modality pair. - /// - /// This is primarily used during initial data loading or recovery. - pub async fn set_version(&self, entity_id: &str, modality: &str, version: u64) { - self.version_table - .write() - .await - .set(entity_id, modality, version); - } - - /// Check whether a specific entity/modality pair is currently locked. - pub async fn is_locked(&self, entity_id: &str, modality: &str) -> bool { - self.lock_table - .read() - .await - .is_locked(entity_id, modality) - } - - /// Get the lock holders for an entity/modality pair. - pub async fn lock_holders( - &self, - entity_id: &str, - modality: &str, - ) -> Vec<(Uuid, LockType)> { - self.lock_table - .read() - .await - .holders(entity_id, modality) - } - - /// Remove completed (Committed or RolledBack) transactions from memory. - /// - /// Returns the number of transactions purged. - pub async fn purge_completed(&self) -> usize { - let mut txns = self.active_transactions.write().await; - let before = txns.len(); - txns.retain(|_id, txn| txn.state == TransactionState::Active); - let purged = before - txns.len(); - if purged > 0 { - info!(purged = purged, "Purged completed transactions"); - } - purged - } -} - -impl Default for TransactionManager { - fn default() -> Self { - Self::new() - } -} - -// --------------------------------------------------------------------------- -// Tests -// --------------------------------------------------------------------------- - -#[cfg(test)] -mod tests { - use super::*; - - /// Helper: create a fresh TransactionManager for each test. - fn new_manager() -> TransactionManager { - TransactionManager::new() - } - - // -- Test 1: Basic lifecycle (begin -> commit) -- - - #[tokio::test] - async fn test_begin_and_commit() { - let mgr = new_manager(); - let txn_id = mgr.begin(IsolationLevel::ReadCommitted).await; - - assert_eq!( - mgr.get_transaction_state(txn_id).await, - Some(TransactionState::Active) - ); - - mgr.commit(txn_id).await.expect("TODO: handle error"); - - assert_eq!( - mgr.get_transaction_state(txn_id).await, - Some(TransactionState::Committed) - ); - } - - // -- Test 2: Basic lifecycle (begin -> rollback) -- - - #[tokio::test] - async fn test_begin_and_rollback() { - let mgr = new_manager(); - let txn_id = mgr.begin(IsolationLevel::ReadCommitted).await; - - let undo = mgr.rollback(txn_id).await.expect("TODO: handle error"); - assert!(undo.is_empty()); - - assert_eq!( - mgr.get_transaction_state(txn_id).await, - Some(TransactionState::RolledBack) - ); - } - - // -- Test 3: Double commit is rejected -- - - #[tokio::test] - async fn test_double_commit_rejected() { - let mgr = new_manager(); - let txn_id = mgr.begin(IsolationLevel::ReadCommitted).await; - - mgr.commit(txn_id).await.expect("TODO: handle error"); - - let result = mgr.commit(txn_id).await; - assert!(matches!( - result, - Err(TransactionError::InvalidState { .. }) - )); - } - - // -- Test 4: Commit after rollback is rejected -- - - #[tokio::test] - async fn test_commit_after_rollback_rejected() { - let mgr = new_manager(); - let txn_id = mgr.begin(IsolationLevel::ReadCommitted).await; - - mgr.rollback(txn_id).await.expect("TODO: handle error"); - - let result = mgr.commit(txn_id).await; - assert!(matches!( - result, - Err(TransactionError::InvalidState { .. }) - )); - } - - // -- Test 5: Nonexistent transaction -- - - #[tokio::test] - async fn test_nonexistent_transaction() { - let mgr = new_manager(); - let fake_id = Uuid::new_v4(); - - let result = mgr.commit(fake_id).await; - assert!(matches!( - result, - Err(TransactionError::TransactionNotFound(_)) - )); - } - - // -- Test 6: Shared locks are compatible -- - - #[tokio::test] - async fn test_shared_locks_compatible() { - let mgr = new_manager(); - let txn_a = mgr.begin(IsolationLevel::ReadCommitted).await; - let txn_b = mgr.begin(IsolationLevel::ReadCommitted).await; - - mgr.acquire_lock(txn_a, "entity-1", "graph", LockType::Shared) - .await - .expect("TODO: handle error"); - mgr.acquire_lock(txn_b, "entity-1", "graph", LockType::Shared) - .await - .expect("TODO: handle error"); - - // Both should hold the lock - let holders = mgr.lock_holders("entity-1", "graph").await; - assert_eq!(holders.len(), 2); - } - - // -- Test 7: Exclusive lock conflicts with shared -- - - #[tokio::test] - async fn test_exclusive_conflicts_with_shared() { - let mgr = new_manager(); - let txn_a = mgr.begin(IsolationLevel::ReadCommitted).await; - let txn_b = mgr.begin(IsolationLevel::ReadCommitted).await; - - mgr.acquire_lock(txn_a, "entity-1", "vector", LockType::Shared) - .await - .expect("TODO: handle error"); - - let result = mgr - .acquire_lock(txn_b, "entity-1", "vector", LockType::Exclusive) - .await; - assert!(matches!(result, Err(TransactionError::LockConflict { .. }))); - } - - // -- Test 8: Exclusive lock conflicts with exclusive -- - - #[tokio::test] - async fn test_exclusive_conflicts_with_exclusive() { - let mgr = new_manager(); - let txn_a = mgr.begin(IsolationLevel::ReadCommitted).await; - let txn_b = mgr.begin(IsolationLevel::ReadCommitted).await; - - mgr.acquire_lock(txn_a, "entity-1", "document", LockType::Exclusive) - .await - .expect("TODO: handle error"); - - let result = mgr - .acquire_lock(txn_b, "entity-1", "document", LockType::Exclusive) - .await; - assert!(matches!(result, Err(TransactionError::LockConflict { .. }))); - } - - // -- Test 9: Locks released on commit -- - - #[tokio::test] - async fn test_locks_released_on_commit() { - let mgr = new_manager(); - let txn_a = mgr.begin(IsolationLevel::ReadCommitted).await; - - mgr.acquire_lock(txn_a, "entity-1", "tensor", LockType::Exclusive) - .await - .expect("TODO: handle error"); - assert!(mgr.is_locked("entity-1", "tensor").await); - - mgr.commit(txn_a).await.expect("TODO: handle error"); - assert!(!mgr.is_locked("entity-1", "tensor").await); - - // Another transaction can now lock it - let txn_b = mgr.begin(IsolationLevel::ReadCommitted).await; - mgr.acquire_lock(txn_b, "entity-1", "tensor", LockType::Exclusive) - .await - .expect("TODO: handle error"); - } - - // -- Test 10: Locks released on rollback -- - - #[tokio::test] - async fn test_locks_released_on_rollback() { - let mgr = new_manager(); - let txn_a = mgr.begin(IsolationLevel::ReadCommitted).await; - - mgr.acquire_lock(txn_a, "entity-1", "semantic", LockType::Exclusive) - .await - .expect("TODO: handle error"); - assert!(mgr.is_locked("entity-1", "semantic").await); - - mgr.rollback(txn_a).await.expect("TODO: handle error"); - assert!(!mgr.is_locked("entity-1", "semantic").await); - } - - // -- Test 11: Undo log records and returns entries in reverse -- - - #[tokio::test] - async fn test_undo_log_reverse_order() { - let mgr = new_manager(); - let txn_id = mgr.begin(IsolationLevel::ReadCommitted).await; - - mgr.record_undo(txn_id, "e1", "graph", Some(vec![1, 2, 3]), 1) - .await - .expect("TODO: handle error"); - mgr.record_undo(txn_id, "e1", "vector", Some(vec![4, 5, 6]), 1) - .await - .expect("TODO: handle error"); - mgr.record_undo(txn_id, "e1", "document", None, 0) - .await - .expect("TODO: handle error"); - - let undo = mgr.rollback(txn_id).await.expect("TODO: handle error"); - assert_eq!(undo.len(), 3); - // Should be in reverse order: document, vector, graph - assert_eq!(undo[0].modality, "document"); - assert_eq!(undo[1].modality, "vector"); - assert_eq!(undo[2].modality, "graph"); - // The document entry had no previous data (new insert) - assert!(undo[0].previous_data.is_none()); - assert_eq!(undo[2].previous_data, Some(vec![1, 2, 3])); - } - - // -- Test 12: MVCC version increments on commit -- - - #[tokio::test] - async fn test_mvcc_version_increment_on_commit() { - let mgr = new_manager(); - let txn_id = mgr.begin(IsolationLevel::ReadCommitted).await; - - assert_eq!(mgr.current_version("e1", "graph").await, 0); - - mgr.record_undo(txn_id, "e1", "graph", None, 0) - .await - .expect("TODO: handle error"); - mgr.record_undo(txn_id, "e1", "vector", None, 0) - .await - .expect("TODO: handle error"); - - mgr.commit(txn_id).await.expect("TODO: handle error"); - - assert_eq!(mgr.current_version("e1", "graph").await, 1); - assert_eq!(mgr.current_version("e1", "vector").await, 1); - // Untouched modality stays at 0 - assert_eq!(mgr.current_version("e1", "tensor").await, 0); - } - - // -- Test 13: Serializable isolation detects version conflict -- - - #[tokio::test] - async fn test_serializable_version_conflict() { - let mgr = new_manager(); - - // Transaction A reads entity e1/graph at version 0 - let txn_a = mgr.begin(IsolationLevel::Serializable).await; - mgr.record_read(txn_a, "e1", "graph").await.expect("TODO: handle error"); - - // Transaction B writes to e1/graph and commits, bumping version to 1 - let txn_b = mgr.begin(IsolationLevel::ReadCommitted).await; - mgr.record_undo(txn_b, "e1", "graph", None, 0) - .await - .expect("TODO: handle error"); - mgr.commit(txn_b).await.expect("TODO: handle error"); - - // Transaction A tries to commit, but e1/graph is now v1 (was v0 at read) - let result = mgr.commit(txn_a).await; - assert!(matches!( - result, - Err(TransactionError::VersionConflict { .. }) - )); - - // Transaction A should be rolled back - assert_eq!( - mgr.get_transaction_state(txn_a).await, - Some(TransactionState::RolledBack) - ); - } - - // -- Test 14: Serializable commits when no conflict -- - - #[tokio::test] - async fn test_serializable_no_conflict() { - let mgr = new_manager(); - - let txn_a = mgr.begin(IsolationLevel::Serializable).await; - mgr.record_read(txn_a, "e1", "graph").await.expect("TODO: handle error"); - mgr.record_undo(txn_a, "e1", "vector", None, 0) - .await - .expect("TODO: handle error"); - - // Nobody else modifies e1/graph, so commit should succeed - mgr.commit(txn_a).await.expect("TODO: handle error"); - assert_eq!( - mgr.get_transaction_state(txn_a).await, - Some(TransactionState::Committed) - ); - } - - // -- Test 15: Invalid modality is rejected -- - - #[tokio::test] - async fn test_invalid_modality_rejected() { - let mgr = new_manager(); - let txn_id = mgr.begin(IsolationLevel::ReadCommitted).await; - - let result = mgr - .acquire_lock(txn_id, "e1", "nosuch", LockType::Shared) - .await; - assert!(matches!( - result, - Err(TransactionError::InvalidModality(_)) - )); - - let result = mgr - .record_undo(txn_id, "e1", "invalid_modality", None, 0) - .await; - assert!(matches!( - result, - Err(TransactionError::InvalidModality(_)) - )); - } - - // -- Test 16: Deadlock detection -- - - #[tokio::test] - async fn test_deadlock_detection() { - let mgr = new_manager(); - let txn_a = mgr.begin(IsolationLevel::ReadCommitted).await; - let txn_b = mgr.begin(IsolationLevel::ReadCommitted).await; - - // A locks e1/graph exclusively - mgr.acquire_lock(txn_a, "e1", "graph", LockType::Exclusive) - .await - .expect("TODO: handle error"); - - // B locks e2/graph exclusively - mgr.acquire_lock(txn_b, "e2", "graph", LockType::Exclusive) - .await - .expect("TODO: handle error"); - - // B tries to lock e1/graph -> conflict (A holds it), creates wait edge B->A - let result_b = mgr - .acquire_lock(txn_b, "e1", "graph", LockType::Exclusive) - .await; - assert!(matches!( - result_b, - Err(TransactionError::LockConflict { .. }) - )); - - // Manually register B waiting for A in the lock table to simulate - // a real wait scenario for deadlock detection - { - let mut lt = mgr.lock_table.write().await; - lt.wait_for.entry(txn_b).or_default().insert(txn_a); - } - - // A tries to lock e2/graph -> conflict (B holds it) - // With B->A already in wait-for, adding A->B creates cycle A->B->A - let result_a = mgr - .acquire_lock(txn_a, "e2", "graph", LockType::Exclusive) - .await; - - // Should detect deadlock (cycle: A -> B -> A) - assert!( - matches!( - result_a, - Err(TransactionError::DeadlockDetected { .. }) - | Err(TransactionError::LockConflict { .. }) - ), - "Expected deadlock or lock conflict, got: {:?}", - result_a - ); - } - - // -- Test 17: Active transaction count -- - - #[tokio::test] - async fn test_active_count() { - let mgr = new_manager(); - assert_eq!(mgr.active_count().await, 0); - - let txn_a = mgr.begin(IsolationLevel::ReadCommitted).await; - let txn_b = mgr.begin(IsolationLevel::ReadCommitted).await; - assert_eq!(mgr.active_count().await, 2); - - mgr.commit(txn_a).await.expect("TODO: handle error"); - assert_eq!(mgr.active_count().await, 1); - - mgr.rollback(txn_b).await.expect("TODO: handle error"); - assert_eq!(mgr.active_count().await, 0); - } - - // -- Test 18: Purge completed transactions -- - - #[tokio::test] - async fn test_purge_completed() { - let mgr = new_manager(); - let txn_a = mgr.begin(IsolationLevel::ReadCommitted).await; - let txn_b = mgr.begin(IsolationLevel::ReadCommitted).await; - let _txn_c = mgr.begin(IsolationLevel::ReadCommitted).await; - - mgr.commit(txn_a).await.expect("TODO: handle error"); - mgr.rollback(txn_b).await.expect("TODO: handle error"); - - let purged = mgr.purge_completed().await; - assert_eq!(purged, 2); - - // Only txn_c remains - assert_eq!(mgr.active_count().await, 1); - } - - // -- Test 19: Cross-modality atomicity scenario -- - - #[tokio::test] - async fn test_cross_modality_atomicity() { - let mgr = new_manager(); - let txn_id = mgr.begin(IsolationLevel::ReadCommitted).await; - - // Lock all six modalities for entity e1 - for modality in MODALITIES { - mgr.acquire_lock(txn_id, "e1", modality, LockType::Exclusive) - .await - .expect("TODO: handle error"); - } - - // Record undo for all eight modalities (octad) - for modality in MODALITIES { - mgr.record_undo(txn_id, "e1", modality, Some(vec![0xDE, 0xAD]), 0) - .await - .expect("TODO: handle error"); - } - - // Simulate a failure after writing all modalities -> rollback - let undo = mgr.rollback(txn_id).await.expect("TODO: handle error"); - assert_eq!(undo.len(), MODALITIES.len()); - - // All locks should be released - for modality in MODALITIES { - assert!( - !mgr.is_locked("e1", modality).await, - "Lock on e1/{modality} should be released after rollback" - ); - } - - // Versions should NOT have incremented (rolled back, not committed) - for modality in MODALITIES { - assert_eq!( - mgr.current_version("e1", modality).await, - 0, - "Version for e1/{modality} should still be 0 after rollback" - ); - } - } - - // -- Test 20: Lock idempotency (re-acquiring same lock is OK) -- - - #[tokio::test] - async fn test_lock_idempotency() { - let mgr = new_manager(); - let txn_id = mgr.begin(IsolationLevel::ReadCommitted).await; - - // Acquire the same lock twice -> should succeed silently - mgr.acquire_lock(txn_id, "e1", "temporal", LockType::Shared) - .await - .expect("TODO: handle error"); - mgr.acquire_lock(txn_id, "e1", "temporal", LockType::Shared) - .await - .expect("TODO: handle error"); - - let holders = mgr.lock_holders("e1", "temporal").await; - assert_eq!(holders.len(), 1); - } - - // -- Test 21: Version table set and get -- - - #[tokio::test] - async fn test_version_table_set_get() { - let mgr = new_manager(); - - assert_eq!(mgr.current_version("e1", "graph").await, 0); - - mgr.set_version("e1", "graph", 42).await; - assert_eq!(mgr.current_version("e1", "graph").await, 42); - - // Different entity/modality is independent - assert_eq!(mgr.current_version("e2", "graph").await, 0); - assert_eq!(mgr.current_version("e1", "vector").await, 0); - } -} diff --git a/verisimdb/rust-core/verisim-octad/tests/atomicity_tests.rs b/verisimdb/rust-core/verisim-octad/tests/atomicity_tests.rs deleted file mode 100644 index fcace554..00000000 --- a/verisimdb/rust-core/verisim-octad/tests/atomicity_tests.rs +++ /dev/null @@ -1,351 +0,0 @@ -// SPDX-License-Identifier: MPL-2.0 -//! Atomicity tests for VeriSimDB octad operations. -//! -//! Verifies that cross-modal write operations are atomic: either all modality -//! writes succeed and are committed, or all are rolled back with no partial -//! state left behind. These tests exercise the [`TransactionManager`] -//! integration in [`InMemoryOctadStore`]. - -use std::sync::Arc; -use verisim_octad::{ - OctadBuilder, OctadConfig, OctadId, OctadStore, InMemoryOctadStore, -}; -use verisim_document::TantivyDocumentStore; -use verisim_graph::SimpleGraphStore; -use verisim_provenance::InMemoryProvenanceStore; -use verisim_semantic::InMemorySemanticStore; -use verisim_spatial::InMemorySpatialStore; -use verisim_temporal::InMemoryVersionStore; -use verisim_tensor::InMemoryTensorStore; -use verisim_vector::{BruteForceVectorStore, DistanceMetric}; - -type TestOctadStore = InMemoryOctadStore< - SimpleGraphStore, - BruteForceVectorStore, - TantivyDocumentStore, - InMemoryTensorStore, - InMemorySemanticStore, - InMemoryVersionStore<verisim_octad::OctadSnapshot>, - InMemoryProvenanceStore, - InMemorySpatialStore, ->; - -fn create_test_store(vector_dim: usize) -> TestOctadStore { - let config = OctadConfig { - vector_dimension: vector_dim, - ..Default::default() - }; - - InMemoryOctadStore::new( - config, - Arc::new(SimpleGraphStore::in_memory().unwrap()), - Arc::new(BruteForceVectorStore::new(vector_dim, DistanceMetric::Cosine)), - Arc::new(TantivyDocumentStore::in_memory().unwrap()), - Arc::new(InMemoryTensorStore::new()), - Arc::new(InMemorySemanticStore::new()), - Arc::new(InMemoryVersionStore::new()), - Arc::new(InMemoryProvenanceStore::new()), - Arc::new(InMemorySpatialStore::new()), - ) -} - -// =========================================================================== -// Create atomicity tests -// =========================================================================== - -#[tokio::test] -async fn test_create_all_modalities_committed() { - // When a create with all 8 modalities succeeds, all modality statuses - // must be true and the transaction must be committed. - let store = create_test_store(3); - - let input = OctadBuilder::new() - .with_document("Atomicity Test", "All modalities populated") - .with_embedding(vec![0.1, 0.2, 0.3]) - .with_tensor(vec![2, 2], vec![1.0, 2.0, 3.0, 4.0]) - .with_types(vec!["https://example.org/AtomicEntity"]) - .with_relationships(vec![("related_to", "other-entity")]) - .with_provenance("created", "test-actor", "Atomicity test creation") - .with_spatial(51.5074, -0.1278) // London - .build(); - - let octad = store.create(input).await.unwrap(); - - // All 8 modalities should be populated - let status = &octad.status.modality_status; - assert!(status.graph, "Graph modality should be populated"); - assert!(status.vector, "Vector modality should be populated"); - assert!(status.document, "Document modality should be populated"); - assert!(status.tensor, "Tensor modality should be populated"); - assert!(status.semantic, "Semantic modality should be populated"); - assert!(status.temporal, "Temporal modality should be populated"); - assert!(status.provenance, "Provenance modality should be populated"); - assert!(status.spatial, "Spatial modality should be populated"); - assert!(status.is_complete(), "All modalities should be complete"); - - // Entity must be retrievable - let retrieved = store.get(&octad.id).await.unwrap(); - assert!(retrieved.is_some(), "Created octad must be retrievable"); - - // Transaction manager should have no active transactions after commit - assert_eq!( - store.transaction_manager().active_count().await, - 0, - "No transactions should be active after successful create" - ); -} - -#[tokio::test] -async fn test_create_vector_dimension_mismatch_rolls_back() { - // When a create fails due to vector dimension mismatch, all previously - // written modalities must be rolled back — no partial state. - let store = create_test_store(3); // Store expects 3-dim vectors - - let input = OctadBuilder::new() - .with_document("Should Be Rolled Back", "This document must not persist") - .with_embedding(vec![0.1, 0.2]) // WRONG: 2 dimensions instead of 3 - .build(); - - let result = store.create(input).await; - assert!(result.is_err(), "Create with wrong vector dimension should fail"); - - // Verify no octads exist (the document write should have been rolled back) - let all = store.list(100, 0).await.unwrap(); - assert!( - all.is_empty(), - "No octads should exist after failed create — rollback must clean up" - ); - - // Transaction manager should have no active transactions - assert_eq!( - store.transaction_manager().active_count().await, - 0, - "No transactions should be active after rollback" - ); -} - -#[tokio::test] -async fn test_create_partial_modalities_succeeds() { - // A create with only some modalities should succeed and only mark - // those modalities as populated. - let store = create_test_store(3); - - let input = OctadBuilder::new() - .with_document("Partial Create", "Only document and temporal") - .build(); - - let octad = store.create(input).await.unwrap(); - - assert!(octad.status.modality_status.document); - assert!(octad.status.modality_status.temporal); - assert!(!octad.status.modality_status.graph); - assert!(!octad.status.modality_status.vector); - assert!(!octad.status.modality_status.tensor); - assert!(!octad.status.modality_status.semantic); - assert!(!octad.status.modality_status.provenance); - assert!(!octad.status.modality_status.spatial); -} - -// =========================================================================== -// Update atomicity tests -// =========================================================================== - -#[tokio::test] -async fn test_update_all_modalities_committed() { - // After a successful update, the octad must reflect the new data and - // the version must be incremented. - let store = create_test_store(3); - - let create_input = OctadBuilder::new() - .with_document("Original Title", "Original body") - .with_embedding(vec![1.0, 0.0, 0.0]) - .build(); - - let octad = store.create(create_input).await.unwrap(); - assert_eq!(octad.status.version, 1); - - let update_input = OctadBuilder::new() - .with_document("Updated Title", "Updated body") - .with_embedding(vec![0.0, 1.0, 0.0]) - .with_provenance("modified", "test-actor", "Update test") - .build(); - - let updated = store.update(&octad.id, update_input).await.unwrap(); - assert_eq!(updated.status.version, 2); - assert!(updated.status.modality_status.document); - assert!(updated.status.modality_status.vector); - assert!(updated.status.modality_status.provenance); - - // Document should reflect the update - assert!(updated.document.as_ref().unwrap().title.contains("Updated")); -} - -#[tokio::test] -async fn test_update_nonexistent_fails() { - // Updating a nonexistent entity must return NotFound, not create a new one. - let store = create_test_store(3); - - let fake_id = OctadId::new("nonexistent-id"); - let input = OctadBuilder::new() - .with_document("Should Fail", "Entity does not exist") - .build(); - - let result = store.update(&fake_id, input).await; - assert!(result.is_err(), "Update of nonexistent entity should fail"); -} - -#[tokio::test] -async fn test_update_with_invalid_vector_dimension_rolls_back() { - // If an update fails (e.g., wrong vector dimension), the entity must - // retain its pre-update state. - let store = create_test_store(3); - - let create_input = OctadBuilder::new() - .with_document("Original", "Should survive failed update") - .with_embedding(vec![1.0, 0.0, 0.0]) - .build(); - - let octad = store.create(create_input).await.unwrap(); - - // Attempt update with wrong vector dimension - let bad_update = OctadBuilder::new() - .with_embedding(vec![0.1, 0.2]) // WRONG: 2 dimensions instead of 3 - .build(); - - let result = store.update(&octad.id, bad_update).await; - assert!(result.is_err(), "Update with wrong vector dimension should fail"); - - // Original entity should still be intact - let original = store.get(&octad.id).await.unwrap().unwrap(); - assert_eq!(original.status.version, 1, "Version should not change on failed update"); - assert!( - original.document.as_ref().unwrap().title.contains("Original"), - "Original data should survive failed update" - ); -} - -// =========================================================================== -// Delete atomicity tests -// =========================================================================== - -#[tokio::test] -async fn test_delete_removes_entity() { - let store = create_test_store(3); - - let input = OctadBuilder::new() - .with_document("To Delete", "Will be removed") - .with_embedding(vec![0.5, 0.5, 0.5]) - .build(); - - let octad = store.create(input).await.unwrap(); - assert!(store.get(&octad.id).await.unwrap().is_some()); - - store.delete(&octad.id).await.unwrap(); - assert!( - store.get(&octad.id).await.unwrap().is_none(), - "Entity should not exist after delete" - ); -} - -#[tokio::test] -async fn test_delete_nonexistent_fails() { - let store = create_test_store(3); - let fake_id = OctadId::new("does-not-exist"); - - let result = store.delete(&fake_id).await; - assert!(result.is_err(), "Delete of nonexistent entity should fail"); -} - -// =========================================================================== -// Transaction manager state tests -// =========================================================================== - -#[tokio::test] -async fn test_transaction_manager_integrated() { - // Verify the transaction manager is accessible and functional through - // the store's public API. - let store = create_test_store(3); - let txn_mgr = store.transaction_manager(); - - // Initially no active transactions - assert_eq!(txn_mgr.active_count().await, 0); - - // Create a octad — should begin and commit a transaction - let input = OctadBuilder::new() - .with_document("TxnTest", "Transaction manager test") - .build(); - store.create(input).await.unwrap(); - - // After create, no transactions should be active - assert_eq!(txn_mgr.active_count().await, 0); -} - -#[tokio::test] -async fn test_transaction_modality_versions_increment() { - // After a create + update, the MVCC version for each written modality - // should reflect the number of writes. - let store = create_test_store(3); - let txn_mgr = store.transaction_manager(); - - let input = OctadBuilder::new() - .with_document("Version Test", "Version tracking") - .with_embedding(vec![0.1, 0.2, 0.3]) - .build(); - - let octad = store.create(input).await.unwrap(); - let entity_id = octad.id.as_str(); - - // After create, document and vector should have version 1 - let doc_v = txn_mgr.current_version(entity_id, "document").await; - let vec_v = txn_mgr.current_version(entity_id, "vector").await; - assert_eq!(doc_v, 1, "Document MVCC version should be 1 after create"); - assert_eq!(vec_v, 1, "Vector MVCC version should be 1 after create"); - - // Update the document - let update = OctadBuilder::new() - .with_document("Updated Version Test", "Version 2") - .build(); - store.update(&octad.id, update).await.unwrap(); - - let doc_v2 = txn_mgr.current_version(entity_id, "document").await; - assert_eq!(doc_v2, 2, "Document MVCC version should be 2 after update"); - - // Vector was not updated, so its version should still be 1 - let vec_v2 = txn_mgr.current_version(entity_id, "vector").await; - assert_eq!(vec_v2, 1, "Vector MVCC version should still be 1"); -} - -// =========================================================================== -// Concurrent write serialization test -// =========================================================================== - -#[tokio::test] -async fn test_concurrent_creates_succeed() { - // Multiple concurrent creates to different entities should all succeed. - // This verifies that the locking mechanism does not over-serialize. - let store = Arc::new(create_test_store(3)); - - let mut handles = Vec::new(); - for i in 0..10 { - let store_clone = Arc::clone(&store); - handles.push(tokio::spawn(async move { - let input = OctadBuilder::new() - .with_document(&format!("Concurrent-{i}"), &format!("Body {i}")) - .with_embedding(vec![i as f32 * 0.1, 0.5, 0.5]) - .build(); - store_clone.create(input).await - })); - } - - let mut successes = 0; - for handle in handles { - if handle.await.unwrap().is_ok() { - successes += 1; - } - } - - assert_eq!(successes, 10, "All 10 concurrent creates should succeed"); - - let all = store.list(100, 0).await.unwrap(); - assert_eq!(all.len(), 10, "All 10 octads should be stored"); -} diff --git a/verisimdb/rust-core/verisim-octad/tests/crash_recovery_tests.rs b/verisimdb/rust-core/verisim-octad/tests/crash_recovery_tests.rs deleted file mode 100644 index 05433cf1..00000000 --- a/verisimdb/rust-core/verisim-octad/tests/crash_recovery_tests.rs +++ /dev/null @@ -1,160 +0,0 @@ -// SPDX-License-Identifier: MPL-2.0 -// Copyright (c) 2026 Jonathan D.A. Jewell (hyperpolymath) <j.d.a.jewell@open.ac.uk> -// -// Crash recovery integration tests for VeriSimDB Phase 1.4. - -use std::collections::HashMap; -use std::sync::Arc; - -use verisim_octad::{ - InMemoryOctadStore, OctadConfig, OctadInput, OctadDocumentInput, - OctadSnapshot, OctadStore, -}; -use verisim_document::TantivyDocumentStore; -use verisim_graph::SimpleGraphStore; -use verisim_semantic::InMemorySemanticStore; -use verisim_temporal::InMemoryVersionStore; -use verisim_tensor::InMemoryTensorStore; -use verisim_vector::{BruteForceVectorStore, DistanceMetric}; - -type TestStore = InMemoryOctadStore< - SimpleGraphStore, - BruteForceVectorStore, - TantivyDocumentStore, - InMemoryTensorStore, - InMemorySemanticStore, - InMemoryVersionStore<OctadSnapshot>, - verisim_provenance::InMemoryProvenanceStore, - verisim_spatial::InMemorySpatialStore, ->; - -fn create_store(wal_dir: &str) -> TestStore { - let config = OctadConfig::default(); - InMemoryOctadStore::new( - config, - Arc::new(SimpleGraphStore::new()), - Arc::new(BruteForceVectorStore::new(3, DistanceMetric::Cosine)), - Arc::new(TantivyDocumentStore::in_memory().unwrap()), - Arc::new(InMemoryTensorStore::new()), - Arc::new(InMemorySemanticStore::new()), - Arc::new(InMemoryVersionStore::new()), - Arc::new(verisim_provenance::InMemoryProvenanceStore::new()), - Arc::new(verisim_spatial::InMemorySpatialStore::new()), - ) - .with_wal(wal_dir, verisim_wal::SyncMode::Fsync) - .expect("WAL init") -} - -fn doc(title: &str, body: &str) -> OctadInput { - OctadInput { - document: Some(OctadDocumentInput { - title: title.into(), - body: body.into(), - fields: HashMap::new(), - }), - ..Default::default() - } -} - -#[tokio::test] -async fn crash_recovery_single_entity() { - let dir = tempfile::tempdir().unwrap(); - let wal = dir.path().join("wal"); - std::fs::create_dir_all(&wal).unwrap(); - - let entity_id; - { - let store = create_store(wal.to_str().unwrap()); - let octad = store.create(doc("Test", "Survives crash")).await.unwrap(); - entity_id = octad.id; - // Crash — no graceful_shutdown - } - - { - let store = create_store(wal.to_str().unwrap()); - let n: usize = store.replay_wal(&wal).await.unwrap(); - assert!(n > 0, "Should recover entity"); - assert!(store.get(&entity_id).await.unwrap().is_some()); - } -} - -#[tokio::test] -async fn graceful_shutdown_then_restart() { - let dir = tempfile::tempdir().unwrap(); - let wal = dir.path().join("wal"); - std::fs::create_dir_all(&wal).unwrap(); - - let entity_id; - { - let store = create_store(wal.to_str().unwrap()); - let octad = store.create(doc("Graceful", "Clean")).await.unwrap(); - entity_id = octad.id; - store.graceful_shutdown().await.unwrap(); - } - - { - let store = create_store(wal.to_str().unwrap()); - let n: usize = store.replay_wal(&wal).await.unwrap(); - assert!(n > 0); - assert!(store.get(&entity_id).await.unwrap().is_some()); - } -} - -#[tokio::test] -async fn ten_entities_survive_crash() { - let dir = tempfile::tempdir().unwrap(); - let wal = dir.path().join("wal"); - std::fs::create_dir_all(&wal).unwrap(); - - let mut ids = Vec::new(); - { - let store = create_store(wal.to_str().unwrap()); - for i in 0..10 { - let octad = store.create(doc(&format!("E{i}"), &format!("B{i}"))).await.unwrap(); - ids.push(octad.id); - } - // Crash - } - - { - let store = create_store(wal.to_str().unwrap()); - let n: usize = store.replay_wal(&wal).await.unwrap(); - assert_eq!(n, 10); - for id in &ids { - assert!(store.get(id).await.unwrap().is_some(), "{id} missing"); - } - } -} - -#[tokio::test] -async fn delete_survives_crash() { - let dir = tempfile::tempdir().unwrap(); - let wal = dir.path().join("wal"); - std::fs::create_dir_all(&wal).unwrap(); - - let entity_id; - { - let store = create_store(wal.to_str().unwrap()); - let octad = store.create(doc("Delete Me", "Gone")).await.unwrap(); - entity_id = octad.id; - store.delete(&entity_id).await.unwrap(); - // Crash - } - - { - let store = create_store(wal.to_str().unwrap()); - let _n: usize = store.replay_wal(&wal).await.unwrap(); - assert!(store.get(&entity_id).await.unwrap().is_none(), "Should stay deleted"); - } -} - -#[tokio::test] -async fn empty_wal_clean_start() { - let dir = tempfile::tempdir().unwrap(); - let wal = dir.path().join("wal"); - std::fs::create_dir_all(&wal).unwrap(); - - let store = create_store(wal.to_str().unwrap()); - let n: usize = store.replay_wal(&wal).await.unwrap(); - assert_eq!(n, 0); -} diff --git a/verisimdb/rust-core/verisim-octad/tests/integration_tests.rs b/verisimdb/rust-core/verisim-octad/tests/integration_tests.rs deleted file mode 100644 index b08edcff..00000000 --- a/verisimdb/rust-core/verisim-octad/tests/integration_tests.rs +++ /dev/null @@ -1,362 +0,0 @@ -// SPDX-License-Identifier: MPL-2.0 -//! Integration tests for VeriSimDB -//! -//! Tests cross-modal consistency and end-to-end workflows. -//! Persistence tests are gated behind `#[ignore]` until store serialization is implemented. - -use std::sync::Arc; -use verisim_octad::{ - OctadBuilder, OctadConfig, OctadStore, InMemoryOctadStore, -}; -use verisim_document::TantivyDocumentStore; -use verisim_graph::SimpleGraphStore; -use verisim_semantic::InMemorySemanticStore; -use verisim_temporal::InMemoryVersionStore; -use verisim_tensor::InMemoryTensorStore; -use verisim_vector::{BruteForceVectorStore, DistanceMetric, VectorStore as _}; - -type TestOctadStore = InMemoryOctadStore< - SimpleGraphStore, - BruteForceVectorStore, - TantivyDocumentStore, - InMemoryTensorStore, - InMemorySemanticStore, - InMemoryVersionStore<verisim_octad::OctadSnapshot>, - verisim_provenance::InMemoryProvenanceStore, - verisim_spatial::InMemorySpatialStore, ->; - -fn create_test_store(vector_dim: usize) -> TestOctadStore { - let config = OctadConfig { - vector_dimension: vector_dim, - ..Default::default() - }; - - InMemoryOctadStore::new( - config, - Arc::new(SimpleGraphStore::in_memory().unwrap()), - Arc::new(BruteForceVectorStore::new(vector_dim, DistanceMetric::Cosine)), - Arc::new(TantivyDocumentStore::in_memory().unwrap()), - Arc::new(InMemoryTensorStore::new()), - Arc::new(InMemorySemanticStore::new()), - Arc::new(InMemoryVersionStore::new()), - Arc::new(verisim_provenance::InMemoryProvenanceStore::new()), - Arc::new(verisim_spatial::InMemorySpatialStore::new()), - ) -} - -/// Test that all six modalities are properly synchronized -#[tokio::test] -async fn test_cross_modal_consistency() { - let store = create_test_store(128); - - // Create a octad with all modalities populated - let input = OctadBuilder::new() - .with_document("Cross-Modal Test", "Testing all modalities together") - .with_embedding(vec![0.1; 128]) - .with_tensor(vec![2, 2], vec![1.0, 2.0, 3.0, 4.0]) - .with_types(vec!["https://example.org/TestType"]) - .with_relationships(vec![("relatedTo", "other-entity")]) - .build(); - - let octad = store.create(input).await.unwrap(); - - // Verify all modalities are populated - let status = &octad.status.modality_status; - assert!(status.graph, "Graph modality should be populated"); - assert!(status.vector, "Vector modality should be populated"); - assert!(status.tensor, "Tensor modality should be populated"); - assert!(status.semantic, "Semantic modality should be populated"); - assert!(status.document, "Document modality should be populated"); - assert!(status.temporal, "Temporal modality should be populated"); - - // Verify the octad can be retrieved - let retrieved = store.get(&octad.id).await.unwrap(); - assert!(retrieved.is_some(), "Octad should be retrievable"); - - let retrieved = retrieved.unwrap(); - assert_eq!(retrieved.id, octad.id); - assert!(retrieved.embedding.is_some()); - assert!(retrieved.tensor.is_some()); - assert!(retrieved.semantic.is_some()); - assert!(retrieved.document.is_some()); -} - -/// Test vector similarity search -#[tokio::test] -async fn test_vector_similarity_search() { - let store = create_test_store(3); - - // Create entities with different embeddings - let inputs = vec![ - (vec![1.0, 0.0, 0.0], "X-axis aligned"), - (vec![0.9, 0.1, 0.0], "Near X-axis"), - (vec![0.0, 1.0, 0.0], "Y-axis aligned"), - (vec![0.0, 0.0, 1.0], "Z-axis aligned"), - ]; - - for (embedding, title) in &inputs { - let input = OctadBuilder::new() - .with_document(title, &format!("Entity at {:?}", embedding)) - .with_embedding(embedding.clone()) - .build(); - store.create(input).await.unwrap(); - } - - // Search for entities similar to X-axis - let results = store.search_similar(&[1.0, 0.0, 0.0], 2).await.unwrap(); - assert_eq!(results.len(), 2); - - // The two closest should be X-axis and near-X-axis - let titles: Vec<_> = results - .iter() - .filter_map(|h| h.document.as_ref().map(|d| d.title.as_str())) - .collect(); - - assert!( - titles.contains(&"X-axis aligned") || titles.contains(&"Near X-axis"), - "Search results should include X-axis or near-X-axis entities" - ); -} - -/// Test full-text search across documents -#[tokio::test] -async fn test_full_text_search() { - let store = create_test_store(3); - - // Create entities with different content - let docs = vec![ - ("Rust Programming", "Rust is a systems programming language focused on safety"), - ("Python Guide", "Python is a high-level programming language"), - ("Database Design", "Relational databases use SQL for querying"), - ]; - - for (title, body) in &docs { - let input = OctadBuilder::new() - .with_document(title, body) - .with_embedding(vec![0.1, 0.2, 0.3]) - .build(); - store.create(input).await.unwrap(); - } - - // Search for "Rust" - let results = store.search_text("Rust", 10).await.unwrap(); - assert_eq!(results.len(), 1); - assert!(results[0].document.as_ref().unwrap().title.contains("Rust")); - - // Search for "programming" - should match multiple - let results = store.search_text("programming", 10).await.unwrap(); - assert!(results.len() >= 2, "Should match at least 2 programming docs"); -} - -/// Test temporal versioning -#[tokio::test] -async fn test_versioning() { - let store = create_test_store(3); - - // Create initial entity - let input = OctadBuilder::new() - .with_document("Version 1", "Initial content") - .with_embedding(vec![0.1, 0.2, 0.3]) - .build(); - - let octad = store.create(input).await.unwrap(); - assert_eq!(octad.status.version, 1); - - // Update the entity - let update1 = OctadBuilder::new() - .with_document("Version 2", "Updated content") - .build(); - - let updated = store.update(&octad.id, update1).await.unwrap(); - assert_eq!(updated.status.version, 2); - - // Update again - let update2 = OctadBuilder::new() - .with_document("Version 3", "Final content") - .build(); - - let updated = store.update(&octad.id, update2).await.unwrap(); - assert_eq!(updated.status.version, 3); - assert_eq!(updated.version_count, 3); -} - -/// Test CRUD operations -#[tokio::test] -async fn test_crud_operations() { - let store = create_test_store(3); - - // Create - let input = OctadBuilder::new() - .with_document("Test Entity", "Test body") - .with_embedding(vec![0.1, 0.2, 0.3]) - .build(); - - let octad = store.create(input).await.unwrap(); - let id = octad.id.clone(); - - // Read - let retrieved = store.get(&id).await.unwrap(); - assert!(retrieved.is_some()); - - // Update - let update_input = OctadBuilder::new() - .with_document("Updated Entity", "Updated body") - .build(); - - let updated = store.update(&id, update_input).await.unwrap(); - assert!(updated.document.as_ref().unwrap().title.contains("Updated")); - - // Delete - store.delete(&id).await.unwrap(); - - // Verify deletion - let deleted = store.get(&id).await.unwrap(); - assert!(deleted.is_none()); -} - -/// Test vector store persistence -/// Ignored: BruteForceVectorStore does not yet have save_to_file/load_from_file. -/// See SONNET-TASKS.md for the persistence implementation task. -#[tokio::test] -#[ignore = "persistence not yet implemented on BruteForceVectorStore"] -async fn test_vector_persistence() { - let store = BruteForceVectorStore::new(64, DistanceMetric::Cosine); - - // Insert vectors - for i in 0..20 { - let mut vec = vec![0.0f32; 64]; - vec[i % 64] = 1.0; - let embedding = verisim_vector::Embedding::new(format!("vec_{}", i), vec); - store.upsert(&embedding).await.unwrap(); - } - - // TODO: Implement save_to_file/load_from_file on BruteForceVectorStore - // store.save_to_file(temp_path).unwrap(); - // let loaded = BruteForceVectorStore::load_from_file(temp_path).unwrap(); - // assert_eq!(loaded.stats().total_vectors, 20); -} - -/// Test tensor store persistence -/// Ignored: InMemoryTensorStore does not yet have save_to_file/load_from_file. -#[tokio::test] -#[ignore = "persistence not yet implemented on InMemoryTensorStore"] -async fn test_tensor_persistence() { - use verisim_tensor::{Tensor, TensorStore as _}; - - let store = InMemoryTensorStore::new(); - - let t1 = Tensor::new("tensor_1", vec![2, 3], vec![1.0, 2.0, 3.0, 4.0, 5.0, 6.0]).unwrap(); - let t2 = Tensor::new("tensor_2", vec![3, 3], vec![1.0; 9]).unwrap(); - - store.put(&t1).await.unwrap(); - store.put(&t2).await.unwrap(); - - // TODO: Implement save_to_file/load_from_file on InMemoryTensorStore - // store.save_to_file(temp_path).unwrap(); - // let loaded = InMemoryTensorStore::load_from_file(temp_path).unwrap(); -} - -/// Test semantic store persistence -/// Ignored: InMemorySemanticStore does not yet have save_to_file/load_from_file. -#[tokio::test] -#[ignore = "persistence not yet implemented on InMemorySemanticStore"] -async fn test_semantic_persistence() { - use verisim_semantic::{SemanticStore as _, SemanticType, Constraint, ConstraintKind}; - - let store = InMemorySemanticStore::new(); - - let person_type = SemanticType::new("https://example.org/Person", "Person") - .with_supertype("https://example.org/Entity") - .with_constraint(Constraint { - name: "name_required".to_string(), - kind: ConstraintKind::Required("name".to_string()), - message: "Person must have a name".to_string(), - }); - - store.register_type(&person_type).await.unwrap(); - - // TODO: Implement save_to_file/load_from_file on InMemorySemanticStore - // store.save_to_file(temp_path).unwrap(); - // let loaded = InMemorySemanticStore::load_from_file(temp_path).unwrap(); -} - -/// Test temporal store persistence -/// Ignored: InMemoryVersionStore does not yet have save_to_file/load_from_file. -#[tokio::test] -#[ignore = "persistence not yet implemented on InMemoryVersionStore"] -async fn test_temporal_persistence() { - use verisim_temporal::TemporalStore as _; - - let store: InMemoryVersionStore<String> = InMemoryVersionStore::new(); - - store.append("entity1", "v1 data".to_string(), "alice", Some("first")).await.unwrap(); - store.append("entity1", "v2 data".to_string(), "bob", Some("second")).await.unwrap(); - - // TODO: Implement save_to_file/load_from_file on InMemoryVersionStore - // store.save_to_file(temp_path).unwrap(); - // let loaded: InMemoryVersionStore<String> = InMemoryVersionStore::load_from_file(temp_path).unwrap(); -} - -/// Test that modality operations are isolated -#[tokio::test] -async fn test_modality_isolation() { - let store = create_test_store(3); - - // Create entity with only document - let input1 = OctadBuilder::new() - .with_document("Doc Only", "Only document modality") - .build(); - - let h1 = store.create(input1).await.unwrap(); - assert!(h1.status.modality_status.document); - assert!(!h1.status.modality_status.vector); - assert!(!h1.status.modality_status.tensor); - - // Create entity with only vector - let input2 = OctadBuilder::new() - .with_embedding(vec![0.1, 0.2, 0.3]) - .build(); - - let h2 = store.create(input2).await.unwrap(); - assert!(!h2.status.modality_status.document); - assert!(h2.status.modality_status.vector); - assert!(!h2.status.modality_status.tensor); - - // Verify searches don't cross-contaminate - let doc_results = store.search_text("Only document", 10).await.unwrap(); - assert_eq!(doc_results.len(), 1); - - // Vector search should only find vector-enabled entities - let vec_results = store.search_similar(&[0.1, 0.2, 0.3], 10).await.unwrap(); - // h2 has vector, h1 does not - let vec_ids: Vec<_> = vec_results.iter().map(|h| h.id.as_str()).collect(); - assert!(vec_ids.contains(&h2.id.as_str())); -} - -/// Test high-dimension vector search -#[tokio::test] -async fn test_high_dimension_vector_search() { - let store = create_test_store(768); // BERT-like dimension - - // Create 50 entities with high-dimensional embeddings - for i in 0..50 { - let mut embedding = vec![0.0f32; 768]; - embedding[i % 768] = 1.0; - embedding[(i * 7) % 768] = 0.5; - - let input = OctadBuilder::new() - .with_document(&format!("Entity {}", i), &format!("High-dim entity number {}", i)) - .with_embedding(embedding) - .build(); - - store.create(input).await.unwrap(); - } - - // Search should complete even with high dimensions - let mut query = vec![0.0f32; 768]; - query[0] = 1.0; - - let results = store.search_similar(&query, 5).await.unwrap(); - assert_eq!(results.len(), 5); -} diff --git a/verisimdb/rust-core/verisim-octad/tests/stress_tests.rs b/verisimdb/rust-core/verisim-octad/tests/stress_tests.rs deleted file mode 100644 index b24b0c35..00000000 --- a/verisimdb/rust-core/verisim-octad/tests/stress_tests.rs +++ /dev/null @@ -1,171 +0,0 @@ -// SPDX-License-Identifier: MPL-2.0 -// Copyright (c) 2026 Jonathan D.A. Jewell (hyperpolymath) <j.d.a.jewell@open.ac.uk> -// -// Stress tests for VeriSimDB Phase 3.2. -// Concurrent writers and readers hitting the octad store simultaneously. - -use std::collections::HashMap; -use std::sync::Arc; - -use verisim_octad::{ - InMemoryOctadStore, OctadConfig, OctadDocumentInput, OctadId, - OctadInput, OctadSnapshot, OctadStore, -}; -use verisim_document::TantivyDocumentStore; -use verisim_graph::SimpleGraphStore; -use verisim_semantic::InMemorySemanticStore; -use verisim_temporal::InMemoryVersionStore; -use verisim_tensor::InMemoryTensorStore; -use verisim_vector::{BruteForceVectorStore, DistanceMetric}; - -type TestStore = InMemoryOctadStore< - SimpleGraphStore, - BruteForceVectorStore, - TantivyDocumentStore, - InMemoryTensorStore, - InMemorySemanticStore, - InMemoryVersionStore<OctadSnapshot>, - verisim_provenance::InMemoryProvenanceStore, - verisim_spatial::InMemorySpatialStore, ->; - -fn create_store() -> Arc<TestStore> { - Arc::new(InMemoryOctadStore::new( - OctadConfig::default(), - Arc::new(SimpleGraphStore::new()), - Arc::new(BruteForceVectorStore::new(3, DistanceMetric::Cosine)), - Arc::new(TantivyDocumentStore::in_memory().unwrap()), - Arc::new(InMemoryTensorStore::new()), - Arc::new(InMemorySemanticStore::new()), - Arc::new(InMemoryVersionStore::new()), - Arc::new(verisim_provenance::InMemoryProvenanceStore::new()), - Arc::new(verisim_spatial::InMemorySpatialStore::new()), - )) -} - -fn doc(title: &str, body: &str) -> OctadInput { - OctadInput { - document: Some(OctadDocumentInput { - title: title.into(), - body: body.into(), - fields: HashMap::new(), - }), - ..Default::default() - } -} - -/// 50 concurrent writers, each creating 10 entities. -#[tokio::test] -async fn concurrent_writers() { - let store = create_store(); - let mut handles = Vec::new(); - - for writer_id in 0..50 { - let store = store.clone(); - handles.push(tokio::spawn(async move { - let mut ids = Vec::new(); - for i in 0..10 { - let input = doc( - &format!("W{writer_id}-E{i}"), - &format!("Body from writer {writer_id} entity {i}"), - ); - match store.create(input).await { - Ok(octad) => ids.push(octad.id), - Err(e) => panic!("Writer {writer_id} entity {i} failed: {e}"), - } - } - ids - })); - } - - let mut all_ids = Vec::new(); - for handle in handles { - let ids = handle.await.unwrap(); - all_ids.extend(ids); - } - - assert_eq!(all_ids.len(), 500, "All 500 entities should be created"); - - // Verify all entities exist - for id in &all_ids { - assert!(store.get(id).await.unwrap().is_some(), "{id} missing"); - } -} - -/// 20 writers + 20 readers running concurrently. -#[tokio::test] -async fn concurrent_read_write() { - let store = create_store(); - let mut handles = Vec::new(); - - // Seed 100 entities first - let mut seed_ids = Vec::new(); - for i in 0..100 { - let octad = store.create(doc(&format!("Seed-{i}"), &format!("Seed body {i}"))).await.unwrap(); - seed_ids.push(octad.id); - } - - // 20 writers creating new entities - for writer_id in 0..20 { - let store = store.clone(); - handles.push(tokio::spawn(async move { - for i in 0..10 { - store.create(doc( - &format!("Concurrent-W{writer_id}-{i}"), - &format!("Body {writer_id}-{i}"), - )).await.unwrap(); - } - })); - } - - // 20 readers reading seed entities - for reader_id in 0..20 { - let store = store.clone(); - let ids = seed_ids.clone(); - handles.push(tokio::spawn(async move { - for id in &ids { - let result = store.get(id).await; - match result { - Ok(Some(_)) => {} // Expected - Ok(None) => {} // Acceptable during concurrent writes - Err(e) => panic!("Reader {reader_id} error on {id}: {e}"), - } - } - })); - } - - for handle in handles { - handle.await.unwrap(); - } -} - -/// Create then delete under contention. -#[tokio::test] -async fn concurrent_create_delete() { - let store = create_store(); - - // Create 50 entities - let mut ids = Vec::new(); - for i in 0..50 { - let octad = store.create(doc(&format!("CD-{i}"), "body")).await.unwrap(); - ids.push(octad.id); - } - - // Concurrently delete all of them - let mut handles = Vec::new(); - for id in ids.clone() { - let store = store.clone(); - handles.push(tokio::spawn(async move { - store.delete(&id).await.unwrap(); - })); - } - - for handle in handles { - handle.await.unwrap(); - } - - // All should be gone - for id in &ids { - assert!(store.get(id).await.unwrap().is_none(), "{id} should be deleted"); - } -} diff --git a/verisimdb/rust-core/verisim-planner/Cargo.toml b/verisimdb/rust-core/verisim-planner/Cargo.toml deleted file mode 100644 index 37e26528..00000000 --- a/verisimdb/rust-core/verisim-planner/Cargo.toml +++ /dev/null @@ -1,22 +0,0 @@ -# SPDX-License-Identifier: MPL-2.0 - -[package] -name = "verisim-planner" -description = "Cost-based query planner for VeriSimDB" -version.workspace = true -edition.workspace = true -authors.workspace = true -license.workspace = true - -[dependencies] -serde.workspace = true -serde_json.workspace = true -chrono.workspace = true -sha2.workspace = true -thiserror.workspace = true -tokio.workspace = true -tracing.workspace = true - -[dev-dependencies] -proptest.workspace = true -tokio = { workspace = true, features = ["macros", "rt-multi-thread", "time"] } diff --git a/verisimdb/rust-core/verisim-planner/src/config.rs b/verisimdb/rust-core/verisim-planner/src/config.rs deleted file mode 100644 index 51719f00..00000000 --- a/verisimdb/rust-core/verisim-planner/src/config.rs +++ /dev/null @@ -1,169 +0,0 @@ -// SPDX-License-Identifier: MPL-2.0 -//! Planner configuration. -//! -//! Defaults match the Elixir query_planner_config.ex: -//! - global_mode: balanced -//! - Vector: aggressive, Graph: conservative, Semantic: conservative -//! - statistics_weight: 0.7 - -use std::collections::HashMap; - -use serde::{Deserialize, Serialize}; - -use crate::Modality; - -/// Optimization mode controlling cost/selectivity trade-offs. -#[derive(Debug, Clone, Copy, PartialEq, Eq, Serialize, Deserialize)] -#[serde(rename_all = "lowercase")] -pub enum OptimizationMode { - /// Safety buffers: cost ×1.5, selectivity ×2.0. - Conservative, - /// Use estimates as-is: cost ×1.0, selectivity ×1.0. - Balanced, - /// Optimistic: cost ×0.8, selectivity ×0.5. - Aggressive, -} - -impl OptimizationMode { - /// Cost multiplier for this mode (from query_planner_config.ex). - pub fn cost_multiplier(self) -> f64 { - match self { - OptimizationMode::Conservative => 1.5, - OptimizationMode::Balanced => 1.0, - OptimizationMode::Aggressive => 0.8, - } - } - - /// Selectivity multiplier for this mode (from query_planner_config.ex). - pub fn selectivity_multiplier(self) -> f64 { - match self { - OptimizationMode::Conservative => 2.0, - OptimizationMode::Balanced => 1.0, - OptimizationMode::Aggressive => 0.5, - } - } -} - -/// Configuration for the query planner. -#[derive(Debug, Clone, Serialize, Deserialize)] -pub struct PlannerConfig { - /// Global optimization mode. - pub global_mode: OptimizationMode, - /// Per-modality mode overrides. - pub modality_overrides: HashMap<Modality, OptimizationMode>, - /// Weight given to historical statistics vs base estimates (0.0–1.0). - pub statistics_weight: f64, - /// Whether to enable adaptive tuning based on execution feedback. - pub enable_adaptive: bool, - /// Minimum number of modality nodes to trigger parallel execution. - pub parallel_threshold: usize, -} - -impl PlannerConfig { - /// Get the effective optimization mode for a modality. - /// - /// Checks per-modality overrides first, falls back to global_mode. - pub fn mode_for(&self, modality: Modality) -> OptimizationMode { - self.modality_overrides - .get(&modality) - .copied() - .unwrap_or(self.global_mode) - } -} - -impl Default for PlannerConfig { - /// Defaults matching Elixir query_planner_config.ex: - /// - global_mode: balanced - /// - Vector: aggressive (predictable HNSW) - /// - Graph: conservative (unpredictable traversals) - /// - Semantic: conservative (ZKP expensive) - /// - statistics_weight: 0.7 - /// - enable_adaptive: true - /// - parallel_threshold: 2 - fn default() -> Self { - let mut overrides = HashMap::new(); - overrides.insert(Modality::Vector, OptimizationMode::Aggressive); - overrides.insert(Modality::Graph, OptimizationMode::Conservative); - overrides.insert(Modality::Semantic, OptimizationMode::Conservative); - - Self { - global_mode: OptimizationMode::Balanced, - modality_overrides: overrides, - statistics_weight: 0.7, - enable_adaptive: true, - parallel_threshold: 2, - } - } -} - -#[cfg(test)] -mod tests { - use super::*; - - #[test] - fn test_defaults_match_elixir() { - let config = PlannerConfig::default(); - assert_eq!(config.global_mode, OptimizationMode::Balanced); - assert_eq!( - config.mode_for(Modality::Vector), - OptimizationMode::Aggressive - ); - assert_eq!( - config.mode_for(Modality::Graph), - OptimizationMode::Conservative - ); - assert_eq!( - config.mode_for(Modality::Semantic), - OptimizationMode::Conservative - ); - assert!((config.statistics_weight - 0.7).abs() < f64::EPSILON); - assert!(config.enable_adaptive); - } - - #[test] - fn test_mode_for_fallback() { - let config = PlannerConfig::default(); - // Tensor, Document, Temporal have no overrides → fall back to global (Balanced) - assert_eq!( - config.mode_for(Modality::Tensor), - OptimizationMode::Balanced - ); - assert_eq!( - config.mode_for(Modality::Document), - OptimizationMode::Balanced - ); - assert_eq!( - config.mode_for(Modality::Temporal), - OptimizationMode::Balanced - ); - } - - #[test] - fn test_cost_multipliers() { - assert!((OptimizationMode::Conservative.cost_multiplier() - 1.5).abs() < f64::EPSILON); - assert!((OptimizationMode::Balanced.cost_multiplier() - 1.0).abs() < f64::EPSILON); - assert!((OptimizationMode::Aggressive.cost_multiplier() - 0.8).abs() < f64::EPSILON); - } - - #[test] - fn test_selectivity_multipliers() { - assert!( - (OptimizationMode::Conservative.selectivity_multiplier() - 2.0).abs() < f64::EPSILON - ); - assert!( - (OptimizationMode::Balanced.selectivity_multiplier() - 1.0).abs() < f64::EPSILON - ); - assert!( - (OptimizationMode::Aggressive.selectivity_multiplier() - 0.5).abs() < f64::EPSILON - ); - } - - #[test] - fn test_config_serde_roundtrip() { - let config = PlannerConfig::default(); - let json = serde_json::to_string(&config).expect("TODO: handle error"); - let parsed: PlannerConfig = serde_json::from_str(&json).expect("TODO: handle error"); - assert_eq!(parsed.global_mode, config.global_mode); - assert_eq!(parsed.parallel_threshold, config.parallel_threshold); - } -} diff --git a/verisimdb/rust-core/verisim-planner/src/cost.rs b/verisimdb/rust-core/verisim-planner/src/cost.rs deleted file mode 100644 index 5cbb7cbf..00000000 --- a/verisimdb/rust-core/verisim-planner/src/cost.rs +++ /dev/null @@ -1,770 +0,0 @@ -// SPDX-License-Identifier: MPL-2.0 -//! Cost model and estimation. -//! -//! Base costs match VCLExplain.res values. Mode multipliers match -//! query_planner_config.ex (conservative=1.5x, balanced=1.0x, aggressive=0.8x). - -use serde::{Deserialize, Serialize}; - -use crate::config::PlannerConfig; -use crate::plan::{ConditionKind, PlanNode}; -use crate::stats::StoreStatistics; -use crate::Modality; - -/// Base cost parameters for a single modality. -#[derive(Debug, Clone, Serialize, Deserialize)] -pub struct BaseCost { - /// Base time estimate in milliseconds. - pub time_ms: f64, - /// Base selectivity (fraction of rows returned, 0.0–1.0). - pub selectivity: f64, - /// Optimization hint for this modality. - pub hint: &'static str, -} - -impl BaseCost { - /// Get the default base cost for a modality. - /// - /// Values match VCLExplain.res: - /// - Graph: 150ms, 0.2 selectivity - /// - Vector: 50ms, 0.01 selectivity - /// - Tensor: 200ms, 0.5 selectivity - /// - Semantic: 300ms, 0.8 selectivity - /// - Document: 80ms, 0.05 selectivity - /// - Temporal: 30ms, 0.1 selectivity - pub fn for_modality(modality: Modality) -> Self { - match modality { - Modality::Graph => BaseCost { - time_ms: 150.0, - selectivity: 0.2, - hint: "Graph traversal — O(E) scan", - }, - Modality::Vector => BaseCost { - time_ms: 50.0, - selectivity: 0.01, - hint: "HNSW approximate nearest neighbor", - }, - Modality::Tensor => BaseCost { - time_ms: 200.0, - selectivity: 0.5, - hint: "Tensor reduction — shape dependent", - }, - Modality::Semantic => BaseCost { - time_ms: 300.0, - selectivity: 0.8, - hint: "ZKP verification — expensive", - }, - Modality::Document => BaseCost { - time_ms: 80.0, - selectivity: 0.05, - hint: "Tantivy inverted index lookup", - }, - Modality::Temporal => BaseCost { - time_ms: 30.0, - selectivity: 0.1, - hint: "Version tree lookup — cached", - }, - } - } -} - -/// Cost estimate for a plan step or the total plan. -#[derive(Debug, Clone, Serialize, Deserialize)] -pub struct CostEstimate { - /// Estimated wall-clock time in milliseconds. - pub time_ms: f64, - /// Estimated number of rows returned. - pub estimated_rows: u64, - /// Selectivity (0.0–1.0). - pub selectivity: f64, - /// I/O component of cost. - pub io_cost: f64, - /// CPU component of cost. - pub cpu_cost: f64, -} - -impl CostEstimate { - /// Combine two estimates for sequential execution (sum of times). - pub fn sequential(a: &CostEstimate, b: &CostEstimate) -> CostEstimate { - CostEstimate { - time_ms: a.time_ms + b.time_ms, - estimated_rows: a.estimated_rows.max(b.estimated_rows), - selectivity: a.selectivity * b.selectivity, - io_cost: a.io_cost + b.io_cost, - cpu_cost: a.cpu_cost + b.cpu_cost, - } - } - - /// Combine two estimates for parallel execution (max of times). - pub fn parallel(a: &CostEstimate, b: &CostEstimate) -> CostEstimate { - CostEstimate { - time_ms: a.time_ms.max(b.time_ms), - estimated_rows: a.estimated_rows.max(b.estimated_rows), - selectivity: a.selectivity * b.selectivity, - io_cost: a.io_cost + b.io_cost, - cpu_cost: a.cpu_cost + b.cpu_cost, - } - } - - /// Combine a list of estimates according to strategy. - pub fn combine(estimates: &[CostEstimate], parallel: bool) -> CostEstimate { - if estimates.is_empty() { - return CostEstimate { - time_ms: 0.0, - estimated_rows: 0, - selectivity: 1.0, - io_cost: 0.0, - cpu_cost: 0.0, - }; - } - let mut result = estimates[0].clone(); - for est in &estimates[1..] { - result = if parallel { - CostEstimate::parallel(&result, est) - } else { - CostEstimate::sequential(&result, est) - }; - } - result - } -} - -/// Proof obligation cost parameters. -/// -/// Different proof types have vastly different verification costs. -/// Values derived from the consultation-dependent-types-zkp.adoc -/// and VCLProofObligation.res cost estimates. -#[derive(Debug, Clone, Serialize, Deserialize)] -pub struct ProofCost { - /// Base verification time in milliseconds. - pub verify_ms: f64, - /// Circuit generation time (if applicable). - pub circuit_ms: f64, - /// Whether this proof type is parallelizable. - pub parallelizable: bool, -} - -impl ProofCost { - /// Get the cost for a proof type by name. - /// - /// Proof types from VCLProofObligation.res: - /// - Existence: trivial check (octad exists) - /// - Citation: contract lookup in registry - /// - Access: semantic store rights check - /// - Integrity: CBOR proof blob + Merkle verification - /// - Provenance: lineage chain walk - /// - ZKP/Custom: full SNARK circuit verification - pub fn for_type(proof_type: &str) -> Self { - match proof_type.to_lowercase().as_str() { - "existence" => ProofCost { - verify_ms: 1.0, - circuit_ms: 0.0, - parallelizable: true, - }, - "citation" => ProofCost { - verify_ms: 5.0, - circuit_ms: 0.0, - parallelizable: true, - }, - "access" => ProofCost { - verify_ms: 15.0, - circuit_ms: 0.0, - parallelizable: true, - }, - "integrity" => ProofCost { - verify_ms: 50.0, - circuit_ms: 10.0, - parallelizable: true, - }, - "provenance" => ProofCost { - verify_ms: 30.0, - circuit_ms: 0.0, - parallelizable: false, // Chain walk is sequential - }, - // ZKP/Custom: full SNARK verification - _ => ProofCost { - verify_ms: 200.0, - circuit_ms: 100.0, - parallelizable: false, - }, - } - } - - /// Total cost for this proof. - pub fn total_ms(&self) -> f64 { - self.verify_ms + self.circuit_ms - } -} - -/// Post-processing cost estimation. -/// -/// Estimates CPU cost for operations applied after modality queries. -pub struct PostProcessingCost; - -impl PostProcessingCost { - /// Estimate cost for a post-processing step given the row count. - pub fn estimate(pp: &crate::plan::PostProcessing, row_count: u64) -> CostEstimate { - let n = row_count.max(1) as f64; - match pp { - crate::plan::PostProcessing::OrderBy { fields, .. } => { - // O(n log n) sort, ~0.001ms per comparison - let sort_time = n * n.log2().max(1.0) * 0.001 * fields.len() as f64; - CostEstimate { - time_ms: sort_time, - estimated_rows: row_count, - selectivity: 1.0, - io_cost: 0.0, - cpu_cost: sort_time, - } - } - crate::plan::PostProcessing::Limit { count } => { - // Nearly free — just truncation - let out_rows = row_count.min(*count as u64); - CostEstimate { - time_ms: 0.1, - estimated_rows: out_rows, - selectivity: out_rows as f64 / n, - io_cost: 0.0, - cpu_cost: 0.1, - } - } - crate::plan::PostProcessing::GroupBy { fields, aggregates } => { - // O(n) hash grouping + O(groups * aggregates) computation - let group_time = n * 0.002 * fields.len() as f64; - let agg_time = n * 0.001 * aggregates.len().max(1) as f64; - let total = group_time + agg_time; - // Grouping typically reduces rows significantly - let est_groups = (n / 10.0).max(1.0) as u64; - CostEstimate { - time_ms: total, - estimated_rows: est_groups, - selectivity: est_groups as f64 / n, - io_cost: 0.0, - cpu_cost: total, - } - } - crate::plan::PostProcessing::Project { columns } => { - // Nearly free — column selection - let project_time = n * 0.0001 * columns.len() as f64; - CostEstimate { - time_ms: project_time, - estimated_rows: row_count, - selectivity: 1.0, - io_cost: 0.0, - cpu_cost: project_time, - } - } - } - } -} - -/// Cross-modal condition cost estimation. -/// -/// Cross-modal conditions (DRIFT, CONSISTENCY, EXISTS, field comparisons) -/// are evaluated post-fetch and have CPU costs proportional to row count. -pub struct CrossModalCost; - -impl CrossModalCost { - /// Estimate the cost of evaluating a cross-modal condition. - pub fn estimate(condition: &ConditionKind, row_count: u64) -> CostEstimate { - let n = row_count.max(1) as f64; - let (per_row_ms, selectivity) = match condition { - // Cross-modal field compare: fetch two fields, compare - ConditionKind::Predicate { expression } if expression.contains("cross_modal") => { - (0.01, 0.3) // Most rows won't match cross-modal predicates - } - // Drift computation: cosine distance between embeddings - ConditionKind::Predicate { expression } if expression.contains("drift") => { - (0.5, 0.5) // Embedding extraction + distance calc per row - } - // Consistency check: full similarity metric - ConditionKind::Predicate { expression } if expression.contains("consistency") => { - (1.0, 0.7) // Metric computation (cosine/euclidean/jaccard) - } - // Exists/NotExists: cheap boolean check - ConditionKind::Predicate { expression } - if expression.contains("exists") || expression.contains("not_exists") => - { - (0.001, 0.5) - } - // Generic predicate fallback - _ => (0.01, 0.5), - }; - - let total_time = n * per_row_ms; - CostEstimate { - time_ms: total_time, - estimated_rows: (n * selectivity).max(1.0) as u64, - selectivity, - io_cost: 0.0, - cpu_cost: total_time, - } - } -} - -/// Cost model that estimates execution cost for plan nodes. -pub struct CostModel; - -impl CostModel { - /// Estimate the cost of executing a single plan node. - /// - /// Factors: - /// 1. Base cost for the modality (from VCLExplain.res) - /// 2. Optimization mode multiplier (from query_planner_config.ex) - /// 3. Store statistics (if available) - /// 4. Early limit reduction - /// 5. Condition-specific adjustments (including proof obligations) - /// 6. Cross-modal condition overhead - pub fn estimate( - node: &PlanNode, - config: &PlannerConfig, - stats: Option<&StoreStatistics>, - ) -> CostEstimate { - let base = BaseCost::for_modality(node.modality); - let mode = config.mode_for(node.modality); - - // Apply mode multipliers (from query_planner_config.ex) - let cost_mult = mode.cost_multiplier(); - let sel_mult = mode.selectivity_multiplier(); - - let mut time_ms = base.time_ms * cost_mult; - let mut selectivity = (base.selectivity * sel_mult).min(1.0); - - // Adjust for store statistics if available (weighted by statistics_weight) - if let Some(s) = stats { - if s.query_count > 0 { - let w = config.statistics_weight; - time_ms = time_ms * (1.0 - w) + s.avg_latency_ms * w; - if s.total_rows > 0 { - let empirical_sel = s.avg_rows_returned as f64 / s.total_rows as f64; - selectivity = selectivity * (1.0 - w) + empirical_sel * w; - } - } - } - - // Early limit reduces selectivity (fewer rows scanned/returned) - if let Some(limit) = node.early_limit { - let limit_factor = (limit as f64 / 1000.0).min(1.0); - selectivity *= limit_factor; - time_ms *= 0.5 + 0.5 * limit_factor; // At least 50% of base cost - } - - // Track proof obligation costs separately for accurate modeling - let mut proof_time_ms = 0.0; - - // Condition-specific adjustments - for condition in &node.conditions { - match condition { - ConditionKind::Equality { .. } => { - selectivity *= 0.1; // Highly selective - time_ms *= 0.7; - } - ConditionKind::Range { .. } => { - selectivity *= 0.3; - time_ms *= 0.8; - } - ConditionKind::Similarity { k } => { - selectivity = (*k as f64 / 10000.0).min(1.0); - } - ConditionKind::Fulltext { .. } => { - // Tantivy inverted index is fast - time_ms *= 0.6; - } - ConditionKind::ProofVerification { contract } => { - // Detailed proof costing based on proof type - let proof_type = extract_proof_type(contract); - let pcost = ProofCost::for_type(&proof_type); - proof_time_ms += pcost.total_ms(); - } - _ => {} - } - } - - // Add proof overhead to total time - time_ms += proof_time_ms; - - let estimated_rows = if let Some(s) = stats { - (s.total_rows as f64 * selectivity).max(1.0) as u64 - } else { - (1000.0 * selectivity).max(1.0) as u64 - }; - - // Split cost: proofs are CPU-bound, modality queries are I/O-heavy - let io_cost = (time_ms - proof_time_ms) * 0.6; - let cpu_cost = (time_ms - proof_time_ms) * 0.4 + proof_time_ms; - - CostEstimate { - time_ms, - estimated_rows, - selectivity, - io_cost, - cpu_cost, - } - } - - /// Estimate total plan cost including post-processing steps. - pub fn estimate_with_post_processing( - modality_cost: &CostEstimate, - post_processing: &[crate::plan::PostProcessing], - ) -> CostEstimate { - let mut total = modality_cost.clone(); - let mut current_rows = total.estimated_rows; - - for pp in post_processing { - let pp_cost = PostProcessingCost::estimate(pp, current_rows); - total.time_ms += pp_cost.time_ms; - total.cpu_cost += pp_cost.cpu_cost; - current_rows = pp_cost.estimated_rows; - } - - total.estimated_rows = current_rows; - total - } - - /// Generate an optimization hint string for a plan node. - pub fn optimization_hint(node: &PlanNode) -> Option<String> { - let base = BaseCost::for_modality(node.modality); - let mut hint = base.hint.to_string(); - - for condition in &node.conditions { - match condition { - ConditionKind::Similarity { k } => { - hint = format!("HNSW ANN search (k={})", k); - } - ConditionKind::Fulltext { query } => { - let preview = if query.len() > 20 { - format!("{}...", &query[..20]) - } else { - query.clone() - }; - hint = format!("Tantivy fulltext: \"{}\"", preview); - } - ConditionKind::Traversal { predicate, depth } => { - hint = format!( - "Graph traversal: {} (depth={})", - predicate, - depth.unwrap_or(1) - ); - } - ConditionKind::ProofVerification { contract } => { - hint = format!("ZKP verify: {}", contract); - } - ConditionKind::Equality { field, .. } => { - hint = format!("Index lookup on {}", field); - } - ConditionKind::AtTime { timestamp } => { - hint = format!("Temporal snapshot at {}", timestamp); - } - _ => {} - } - } - - Some(hint) - } -} - -/// Extract proof type from a contract string. -/// -/// Contract strings follow the pattern "ProofType(ContractName)" -/// or just "ContractName" (defaults to custom/ZKP). -fn extract_proof_type(contract: &str) -> String { - let known_types = [ - "existence", "citation", "access", "integrity", "provenance", "zkp", - ]; - let lower = contract.to_lowercase(); - for t in &known_types { - if lower.contains(t) { - return t.to_string(); - } - } - "custom".to_string() -} - -#[cfg(test)] -mod tests { - use super::*; - use crate::config::PlannerConfig; - - #[test] - fn test_base_costs_match_vcl_explain() { - assert_eq!(BaseCost::for_modality(Modality::Graph).time_ms, 150.0); - assert_eq!(BaseCost::for_modality(Modality::Vector).time_ms, 50.0); - assert_eq!(BaseCost::for_modality(Modality::Tensor).time_ms, 200.0); - assert_eq!(BaseCost::for_modality(Modality::Semantic).time_ms, 300.0); - assert_eq!(BaseCost::for_modality(Modality::Document).time_ms, 80.0); - assert_eq!(BaseCost::for_modality(Modality::Temporal).time_ms, 30.0); - } - - #[test] - fn test_base_selectivity_values() { - assert!((BaseCost::for_modality(Modality::Graph).selectivity - 0.2).abs() < f64::EPSILON); - assert!((BaseCost::for_modality(Modality::Vector).selectivity - 0.01).abs() < f64::EPSILON); - assert!((BaseCost::for_modality(Modality::Semantic).selectivity - 0.8).abs() < f64::EPSILON); - } - - #[test] - fn test_mode_multipliers_balanced() { - let config = PlannerConfig::default(); - let node = PlanNode { - modality: Modality::Graph, - conditions: vec![], - projections: vec![], - early_limit: None, - }; - // Graph has conservative override by default → cost_mult = 1.5 - let est = CostModel::estimate(&node, &config, None); - assert!((est.time_ms - 150.0 * 1.5).abs() < f64::EPSILON); - } - - #[test] - fn test_mode_multipliers_aggressive() { - let config = PlannerConfig::default(); - let node = PlanNode { - modality: Modality::Vector, - conditions: vec![], - projections: vec![], - early_limit: None, - }; - // Vector has aggressive override by default → cost_mult = 0.8 - let est = CostModel::estimate(&node, &config, None); - assert!((est.time_ms - 50.0 * 0.8).abs() < f64::EPSILON); - } - - #[test] - fn test_early_limit_reduces_selectivity() { - let config = PlannerConfig::default(); - let node_no_limit = PlanNode { - modality: Modality::Document, - conditions: vec![], - projections: vec![], - early_limit: None, - }; - let node_with_limit = PlanNode { - modality: Modality::Document, - conditions: vec![], - projections: vec![], - early_limit: Some(10), - }; - let est_no = CostModel::estimate(&node_no_limit, &config, None); - let est_with = CostModel::estimate(&node_with_limit, &config, None); - assert!(est_with.selectivity < est_no.selectivity); - assert!(est_with.time_ms < est_no.time_ms); - } - - #[test] - fn test_sequential_combinator() { - let a = CostEstimate { - time_ms: 100.0, - estimated_rows: 50, - selectivity: 0.5, - io_cost: 60.0, - cpu_cost: 40.0, - }; - let b = CostEstimate { - time_ms: 200.0, - estimated_rows: 100, - selectivity: 0.3, - io_cost: 120.0, - cpu_cost: 80.0, - }; - let combined = CostEstimate::sequential(&a, &b); - assert!((combined.time_ms - 300.0).abs() < f64::EPSILON); - assert!((combined.io_cost - 180.0).abs() < f64::EPSILON); - } - - #[test] - fn test_parallel_combinator() { - let a = CostEstimate { - time_ms: 100.0, - estimated_rows: 50, - selectivity: 0.5, - io_cost: 60.0, - cpu_cost: 40.0, - }; - let b = CostEstimate { - time_ms: 200.0, - estimated_rows: 100, - selectivity: 0.3, - io_cost: 120.0, - cpu_cost: 80.0, - }; - let combined = CostEstimate::parallel(&a, &b); - assert!((combined.time_ms - 200.0).abs() < f64::EPSILON); // max - assert!((combined.io_cost - 180.0).abs() < f64::EPSILON); // sum - } - - // ==================================================================== - // Task #7: Proof obligation costing - // ==================================================================== - - #[test] - fn test_proof_cost_existence_is_cheap() { - let pc = ProofCost::for_type("existence"); - assert!(pc.total_ms() < 5.0, "Existence proof should be < 5ms"); - } - - #[test] - fn test_proof_cost_zkp_is_expensive() { - let pc = ProofCost::for_type("zkp"); - assert!(pc.total_ms() > 100.0, "ZKP proof should be > 100ms"); - assert!(pc.circuit_ms > 0.0, "ZKP should have circuit generation cost"); - } - - #[test] - fn test_proof_cost_integrity_includes_circuit() { - let pc = ProofCost::for_type("integrity"); - assert!(pc.circuit_ms > 0.0, "Integrity proof includes Merkle circuit"); - assert!(pc.verify_ms > pc.circuit_ms, "Verify > circuit for integrity"); - } - - #[test] - fn test_proof_cost_unknown_defaults_to_custom() { - let pc = ProofCost::for_type("MyCustomContract"); - assert!(pc.total_ms() > 100.0, "Unknown proof defaults to expensive custom"); - } - - #[test] - fn test_proof_adds_to_node_cost() { - let config = PlannerConfig::default(); - let node_no_proof = PlanNode { - modality: Modality::Semantic, - conditions: vec![], - projections: vec![], - early_limit: None, - }; - let node_with_proof = PlanNode { - modality: Modality::Semantic, - conditions: vec![ConditionKind::ProofVerification { - contract: "integrity_check".to_string(), - }], - projections: vec![], - early_limit: None, - }; - let est_no = CostModel::estimate(&node_no_proof, &config, None); - let est_with = CostModel::estimate(&node_with_proof, &config, None); - assert!(est_with.time_ms > est_no.time_ms, "Proof should add cost"); - // Proof cost should show up in CPU, not I/O - assert!(est_with.cpu_cost > est_no.cpu_cost); - } - - #[test] - fn test_multiple_proofs_accumulate() { - let config = PlannerConfig::default(); - let node_one = PlanNode { - modality: Modality::Semantic, - conditions: vec![ConditionKind::ProofVerification { - contract: "existence_check".to_string(), - }], - projections: vec![], - early_limit: None, - }; - let node_two = PlanNode { - modality: Modality::Semantic, - conditions: vec![ - ConditionKind::ProofVerification { - contract: "existence_check".to_string(), - }, - ConditionKind::ProofVerification { - contract: "integrity_check".to_string(), - }, - ], - projections: vec![], - early_limit: None, - }; - let est_one = CostModel::estimate(&node_one, &config, None); - let est_two = CostModel::estimate(&node_two, &config, None); - assert!(est_two.time_ms > est_one.time_ms, "Two proofs > one proof"); - } - - #[test] - fn test_extract_proof_type_known() { - assert_eq!(extract_proof_type("CitationContract"), "citation"); - assert_eq!(extract_proof_type("integrity_check"), "integrity"); - assert_eq!(extract_proof_type("AccessRights"), "access"); - assert_eq!(extract_proof_type("ProvenanceAudit"), "provenance"); - } - - #[test] - fn test_extract_proof_type_unknown() { - assert_eq!(extract_proof_type("MyContract"), "custom"); - assert_eq!(extract_proof_type("FooBar"), "custom"); - } - - // ==================================================================== - // Task #9: Post-processing and cross-modal costing - // ==================================================================== - - #[test] - fn test_post_processing_limit_cheap() { - use crate::plan::PostProcessing; - let cost = PostProcessingCost::estimate(&PostProcessing::Limit { count: 10 }, 1000); - assert!(cost.time_ms < 1.0, "LIMIT should be nearly free"); - assert_eq!(cost.estimated_rows, 10); - } - - #[test] - fn test_post_processing_order_by_scales_with_rows() { - use crate::plan::PostProcessing; - let small = PostProcessingCost::estimate( - &PostProcessing::OrderBy { fields: vec![("name".into(), true)] }, - 100, - ); - let large = PostProcessingCost::estimate( - &PostProcessing::OrderBy { fields: vec![("name".into(), true)] }, - 10000, - ); - assert!(large.time_ms > small.time_ms * 10.0, "O(n log n) scaling"); - } - - #[test] - fn test_post_processing_group_by_reduces_rows() { - use crate::plan::PostProcessing; - let cost = PostProcessingCost::estimate( - &PostProcessing::GroupBy { - fields: vec!["category".into()], - aggregates: vec!["COUNT(*)".into()], - }, - 1000, - ); - assert!(cost.estimated_rows < 1000, "GROUP BY should reduce rows"); - } - - #[test] - fn test_estimate_with_post_processing() { - use crate::plan::PostProcessing; - let base = CostEstimate { - time_ms: 100.0, - estimated_rows: 500, - selectivity: 0.5, - io_cost: 60.0, - cpu_cost: 40.0, - }; - let pps = vec![ - PostProcessing::OrderBy { fields: vec![("score".into(), false)] }, - PostProcessing::Limit { count: 10 }, - ]; - let total = CostModel::estimate_with_post_processing(&base, &pps); - assert!(total.time_ms > base.time_ms, "PP adds time"); - assert_eq!(total.estimated_rows, 10, "LIMIT reduces final rows"); - } - - #[test] - fn test_cross_modal_drift_cost() { - let cond = ConditionKind::Predicate { - expression: "drift(vector, document) > 0.3".to_string(), - }; - let cost = CrossModalCost::estimate(&cond, 100); - assert!(cost.time_ms > 0.0); - assert!(cost.selectivity < 1.0); - } - - #[test] - fn test_cross_modal_exists_cheap() { - let cond = ConditionKind::Predicate { - expression: "vector exists".to_string(), - }; - let cost = CrossModalCost::estimate(&cond, 1000); - // Exists is cheap — just a boolean check - assert!(cost.time_ms < 5.0); - } -} diff --git a/verisimdb/rust-core/verisim-planner/src/error.rs b/verisimdb/rust-core/verisim-planner/src/error.rs deleted file mode 100644 index 9cdf971f..00000000 --- a/verisimdb/rust-core/verisim-planner/src/error.rs +++ /dev/null @@ -1,23 +0,0 @@ -// SPDX-License-Identifier: MPL-2.0 -//! Planner error types. - -use thiserror::Error; - -/// Errors that can occur during query planning. -#[derive(Error, Debug)] -pub enum PlannerError { - #[error("empty plan: no modality nodes to optimize")] - EmptyPlan, - - #[error("unknown modality: {0}")] - UnknownModality(String), - - #[error("invalid configuration: {0}")] - InvalidConfig(String), - - #[error("cost estimation failed: {0}")] - CostEstimation(String), - - #[error("serialization error: {0}")] - Serialization(#[from] serde_json::Error), -} diff --git a/verisimdb/rust-core/verisim-planner/src/explain.rs b/verisimdb/rust-core/verisim-planner/src/explain.rs deleted file mode 100644 index 2ac499f5..00000000 --- a/verisimdb/rust-core/verisim-planner/src/explain.rs +++ /dev/null @@ -1,368 +0,0 @@ -// SPDX-License-Identifier: MPL-2.0 -//! EXPLAIN output rendering (text and JSON). - -use std::collections::HashMap; -use std::fmt; - -use serde::{Deserialize, Serialize}; - -use crate::config::PlannerConfig; -use crate::plan::PhysicalPlan; -use crate::Modality; - -/// Cost breakdown for a single modality. -#[derive(Debug, Clone, Serialize, Deserialize)] -pub struct ModalityCostBreakdown { - pub modality: Modality, - pub time_ms: f64, - pub percentage: f64, -} - -/// A single step in the EXPLAIN output. -#[derive(Debug, Clone, Serialize, Deserialize)] -pub struct ExplainStep { - pub step: usize, - pub operation: String, - pub modality: Modality, - pub estimated_cost_ms: f64, - pub estimated_selectivity: f64, - pub estimated_rows: u64, - pub optimization_hint: Option<String>, -} - -/// A performance hint generated by the planner. -#[derive(Debug, Clone, Serialize, Deserialize)] -pub struct PerformanceHint { - pub severity: String, - pub message: String, -} - -/// Full EXPLAIN output for a query plan. -#[derive(Debug, Clone, Serialize, Deserialize)] -pub struct ExplainOutput { - /// Per-step details. - pub steps: Vec<ExplainStep>, - /// Cost breakdown by modality. - pub cost_breakdown: Vec<ModalityCostBreakdown>, - /// Performance hints and suggestions. - pub performance_hints: Vec<PerformanceHint>, - /// Total estimated cost in milliseconds. - pub total_cost_ms: f64, - /// Execution strategy. - pub strategy: String, - /// Human-readable text rendering. - pub text_output: String, -} - -impl ExplainOutput { - /// Build an ExplainOutput from a physical plan. - pub fn from_physical_plan(plan: &PhysicalPlan, config: &PlannerConfig) -> Self { - let steps: Vec<ExplainStep> = plan - .steps - .iter() - .map(|s| ExplainStep { - step: s.step, - operation: s.operation.clone(), - modality: s.modality, - estimated_cost_ms: s.cost.time_ms, - estimated_selectivity: s.cost.selectivity, - estimated_rows: s.cost.estimated_rows, - optimization_hint: s.optimization_hint.clone(), - }) - .collect(); - - // Build cost breakdown by modality - let mut modality_costs: HashMap<Modality, f64> = HashMap::new(); - for step in &plan.steps { - *modality_costs.entry(step.modality).or_insert(0.0) += step.cost.time_ms; - } - let total_modality_cost: f64 = modality_costs.values().sum(); - - let mut cost_breakdown: Vec<ModalityCostBreakdown> = modality_costs - .into_iter() - .map(|(modality, time_ms)| { - let percentage = if total_modality_cost > 0.0 { - (time_ms / total_modality_cost) * 100.0 - } else { - 0.0 - }; - ModalityCostBreakdown { - modality, - time_ms, - percentage, - } - }) - .collect(); - cost_breakdown.sort_by(|a, b| b.percentage.partial_cmp(&a.percentage).unwrap_or(std::cmp::Ordering::Equal)); - - // Generate performance hints - let mut hints = Vec::new(); - - // High total cost - if plan.total_cost.time_ms > 500.0 { - hints.push(PerformanceHint { - severity: "warning".to_string(), - message: format!( - "Total estimated cost is {:.0}ms — consider adding LIMIT clause", - plan.total_cost.time_ms - ), - }); - } - - // Check for expensive individual steps - for step in &plan.steps { - if step.cost.time_ms > 200.0 { - hints.push(PerformanceHint { - severity: "info".to_string(), - message: format!( - "Step {} ({}) costs {:.0}ms — {}", - step.step, - step.modality, - step.cost.time_ms, - step.optimization_hint.as_deref().unwrap_or("consider optimization") - ), - }); - } - } - - // First step has poor selectivity - if let Some(first) = plan.steps.first() { - if first.cost.selectivity > 0.1 { - hints.push(PerformanceHint { - severity: "suggestion".to_string(), - message: format!( - "First step ({}) has {:.0}% selectivity — consider reordering or adding predicates", - first.modality, - first.cost.selectivity * 100.0 - ), - }); - } - } - - // Sequential when parallel is possible - if plan.steps.len() >= 2 - && plan.strategy == crate::plan::ExecutionStrategy::Sequential - { - hints.push(PerformanceHint { - severity: "suggestion".to_string(), - message: "Multiple modalities could benefit from parallel execution".to_string(), - }); - } - - let strategy = format!("{:?}", plan.strategy); - let total_cost_ms = plan.total_cost.time_ms; - - // Build text output - let text_output = Self::render_text(&steps, &cost_breakdown, &hints, &strategy, total_cost_ms, config); - - ExplainOutput { - steps, - cost_breakdown, - performance_hints: hints, - total_cost_ms, - strategy, - text_output, - } - } - - fn render_text( - steps: &[ExplainStep], - cost_breakdown: &[ModalityCostBreakdown], - hints: &[PerformanceHint], - strategy: &str, - total_cost_ms: f64, - _config: &PlannerConfig, - ) -> String { - let mut out = String::new(); - - out.push_str("=== VeriSimDB Query Plan ===\n\n"); - out.push_str(&format!("Strategy: {}\n", strategy)); - out.push_str(&format!("Total Estimated Cost: {:.1}ms\n\n", total_cost_ms)); - - out.push_str("--- Steps ---\n"); - for step in steps { - out.push_str(&format!( - " Step {}: {} [{}]\n", - step.step, step.operation, step.modality - )); - out.push_str(&format!( - " Cost: {:.1}ms | Selectivity: {:.2}% | Rows: ~{}\n", - step.estimated_cost_ms, - step.estimated_selectivity * 100.0, - step.estimated_rows - )); - if let Some(hint) = &step.optimization_hint { - out.push_str(&format!(" Hint: {}\n", hint)); - } - } - - out.push_str("\n--- Cost Breakdown ---\n"); - for cb in cost_breakdown { - out.push_str(&format!( - " {}: {:.1}ms ({:.1}%)\n", - cb.modality, cb.time_ms, cb.percentage - )); - } - - if !hints.is_empty() { - out.push_str("\n--- Performance Hints ---\n"); - for hint in hints { - out.push_str(&format!(" [{}] {}\n", hint.severity, hint.message)); - } - } - - out - } -} - -impl fmt::Display for ExplainOutput { - fn fmt(&self, f: &mut fmt::Formatter<'_>) -> fmt::Result { - write!(f, "{}", self.text_output) - } -} - -#[cfg(test)] -mod tests { - use super::*; - use crate::cost::CostEstimate; - use crate::plan::{ExecutionStrategy, PhysicalPlan, PlanStep}; - - fn sample_physical_plan() -> PhysicalPlan { - PhysicalPlan { - steps: vec![ - PlanStep { - step: 1, - operation: "Vector similarity search (1 conditions)".to_string(), - modality: Modality::Vector, - cost: CostEstimate { - time_ms: 40.0, - estimated_rows: 10, - selectivity: 0.005, - io_cost: 24.0, - cpu_cost: 16.0, - }, - optimization_hint: Some("HNSW ANN search (k=10)".to_string()), - pushed_predicates: vec!["Similarity { k: 10 }".to_string()], - }, - PlanStep { - step: 2, - operation: "Graph traversal (1 conditions)".to_string(), - modality: Modality::Graph, - cost: CostEstimate { - time_ms: 225.0, - estimated_rows: 200, - selectivity: 0.4, - io_cost: 135.0, - cpu_cost: 90.0, - }, - optimization_hint: Some("Graph traversal: relates_to (depth=2)".to_string()), - pushed_predicates: vec!["Traversal { predicate: relates_to }".to_string()], - }, - ], - strategy: ExecutionStrategy::Parallel, - total_cost: CostEstimate { - time_ms: 225.0, - estimated_rows: 200, - selectivity: 0.002, - io_cost: 159.0, - cpu_cost: 106.0, - }, - notes: vec!["Parallel execution across 2 modalities".to_string()], - } - } - - #[test] - fn test_step_count() { - let plan = sample_physical_plan(); - let explain = ExplainOutput::from_physical_plan(&plan, &PlannerConfig::default()); - assert_eq!(explain.steps.len(), 2); - } - - #[test] - fn test_cost_percentages_sum_to_100() { - let plan = sample_physical_plan(); - let explain = ExplainOutput::from_physical_plan(&plan, &PlannerConfig::default()); - - let total_pct: f64 = explain.cost_breakdown.iter().map(|cb| cb.percentage).sum(); - assert!( - (total_pct - 100.0).abs() < 0.1, - "Cost percentages sum to {}, expected ~100", - total_pct - ); - } - - #[test] - fn test_text_output_contains_sections() { - let plan = sample_physical_plan(); - let explain = ExplainOutput::from_physical_plan(&plan, &PlannerConfig::default()); - - assert!(explain.text_output.contains("VeriSimDB Query Plan")); - assert!(explain.text_output.contains("Strategy")); - assert!(explain.text_output.contains("Steps")); - assert!(explain.text_output.contains("Cost Breakdown")); - assert!(explain.text_output.contains("Step 1")); - assert!(explain.text_output.contains("Step 2")); - assert!(explain.text_output.contains("vector")); - assert!(explain.text_output.contains("graph")); - } - - #[test] - fn test_performance_hints_for_expensive() { - // Create a plan with a very expensive step - let plan = PhysicalPlan { - steps: vec![PlanStep { - step: 1, - operation: "Semantic verification".to_string(), - modality: Modality::Semantic, - cost: CostEstimate { - time_ms: 675.0, - estimated_rows: 800, - selectivity: 0.8, - io_cost: 405.0, - cpu_cost: 270.0, - }, - optimization_hint: Some("ZKP verification — expensive".to_string()), - pushed_predicates: vec![], - }], - strategy: ExecutionStrategy::Sequential, - total_cost: CostEstimate { - time_ms: 675.0, - estimated_rows: 800, - selectivity: 0.8, - io_cost: 405.0, - cpu_cost: 270.0, - }, - notes: vec![], - }; - - let explain = ExplainOutput::from_physical_plan(&plan, &PlannerConfig::default()); - assert!(!explain.performance_hints.is_empty()); - - // Should have warning about high cost - let has_cost_warning = explain - .performance_hints - .iter() - .any(|h| h.message.contains("LIMIT")); - assert!(has_cost_warning); - } - - #[test] - fn test_explain_json_roundtrip() { - let plan = sample_physical_plan(); - let explain = ExplainOutput::from_physical_plan(&plan, &PlannerConfig::default()); - - let json = serde_json::to_string(&explain).expect("TODO: handle error"); - let parsed: ExplainOutput = serde_json::from_str(&json).expect("TODO: handle error"); - assert_eq!(parsed.steps.len(), explain.steps.len()); - assert!((parsed.total_cost_ms - explain.total_cost_ms).abs() < f64::EPSILON); - } - - #[test] - fn test_display_trait() { - let plan = sample_physical_plan(); - let explain = ExplainOutput::from_physical_plan(&plan, &PlannerConfig::default()); - let display = format!("{}", explain); - assert!(!display.is_empty()); - assert!(display.contains("VeriSimDB")); - } -} diff --git a/verisimdb/rust-core/verisim-planner/src/lib.rs b/verisimdb/rust-core/verisim-planner/src/lib.rs deleted file mode 100644 index 6bbbad60..00000000 --- a/verisimdb/rust-core/verisim-planner/src/lib.rs +++ /dev/null @@ -1,154 +0,0 @@ -// SPDX-License-Identifier: MPL-2.0 -//! VeriSim Planner -//! -//! Cost-based query planning for VeriSimDB. -//! Transforms logical plans into optimized physical execution plans -//! with per-modality cost estimation and EXPLAIN output. - -#![forbid(unsafe_code)] -pub mod config; -pub mod cost; -pub mod error; -pub mod explain; -pub mod optimizer; -pub mod plan; -pub mod prepared; -pub mod profiler; -pub mod slow_query; -pub mod stats; -pub mod vcl_bridge; - -use serde::{Deserialize, Serialize}; -use std::fmt; -use std::str::FromStr; - -pub use config::{OptimizationMode, PlannerConfig}; -pub use cost::{CostEstimate, CostModel, CrossModalCost, PostProcessingCost, ProofCost}; -pub use error::PlannerError; -pub use explain::ExplainOutput; -pub use optimizer::Planner; -pub use plan::{LogicalPlan, PhysicalPlan}; -pub use profiler::{ExplainAnalyzeOutput, Profiler, ProfileStep, QueryProfile}; -pub use prepared::{CacheConfig, CacheError, CacheStats, ParamValue, PlanCache, PreparedId, PreparedStatement}; -pub use slow_query::{SlowQueryConfig, SlowQueryEntry, SlowQueryLog, SlowQuerySummary}; -pub use stats::{AdaptiveTuner, StatisticsCollector, StoreStatistics}; - -/// The six modalities of VeriSimDB. -/// -/// Each modality represents a different representation/store for octad entities. -/// The planner defines its own canonical enum to avoid coupling with verisim-octad -/// or verisim-drift. -#[derive(Debug, Clone, Copy, PartialEq, Eq, Hash, Serialize, Deserialize)] -#[serde(rename_all = "lowercase")] -pub enum Modality { - Graph, - Vector, - Tensor, - Semantic, - Document, - Temporal, -} - -impl Modality { - /// All six modalities in canonical order. - pub const ALL: [Modality; 6] = [ - Modality::Graph, - Modality::Vector, - Modality::Tensor, - Modality::Semantic, - Modality::Document, - Modality::Temporal, - ]; - - /// Execution priority — lower value means execute earlier. - /// - /// Matches the Elixir bidirectional planner ordering: - /// - Temporal first (often cached) - /// - Vector/Document next (selective indexes) - /// - Graph middle - /// - Tensor moderate - /// - Semantic last (ZKP expensive) - pub fn execution_priority(self) -> u32 { - match self { - Modality::Temporal => 10, - Modality::Vector => 20, - Modality::Document => 30, - Modality::Graph => 40, - Modality::Tensor => 50, - Modality::Semantic => 90, - } - } -} - -impl fmt::Display for Modality { - fn fmt(&self, f: &mut fmt::Formatter<'_>) -> fmt::Result { - match self { - Modality::Graph => write!(f, "graph"), - Modality::Vector => write!(f, "vector"), - Modality::Tensor => write!(f, "tensor"), - Modality::Semantic => write!(f, "semantic"), - Modality::Document => write!(f, "document"), - Modality::Temporal => write!(f, "temporal"), - } - } -} - -impl FromStr for Modality { - type Err = PlannerError; - - fn from_str(s: &str) -> Result<Self, Self::Err> { - match s.to_lowercase().as_str() { - "graph" => Ok(Modality::Graph), - "vector" => Ok(Modality::Vector), - "tensor" => Ok(Modality::Tensor), - "semantic" => Ok(Modality::Semantic), - "document" => Ok(Modality::Document), - "temporal" => Ok(Modality::Temporal), - _ => Err(PlannerError::UnknownModality(s.to_string())), - } - } -} - -#[cfg(test)] -mod tests { - use super::*; - - #[test] - fn test_modality_display_roundtrip() { - for m in Modality::ALL { - let s = m.to_string(); - let parsed: Modality = s.parse().expect("TODO: handle error"); - assert_eq!(m, parsed); - } - } - - #[test] - fn test_modality_case_insensitive_parse() { - assert_eq!("GRAPH".parse::<Modality>().expect("TODO: handle error"), Modality::Graph); - assert_eq!("Vector".parse::<Modality>().expect("TODO: handle error"), Modality::Vector); - assert_eq!("SEMANTIC".parse::<Modality>().expect("TODO: handle error"), Modality::Semantic); - } - - #[test] - fn test_unknown_modality_error() { - assert!("unknown".parse::<Modality>().is_err()); - } - - #[test] - fn test_modality_serde_roundtrip() { - for m in Modality::ALL { - let json = serde_json::to_string(&m).expect("TODO: handle error"); - let parsed: Modality = serde_json::from_str(&json).expect("TODO: handle error"); - assert_eq!(m, parsed); - } - } - - #[test] - fn test_execution_priority_ordering() { - assert!(Modality::Temporal.execution_priority() < Modality::Vector.execution_priority()); - assert!(Modality::Vector.execution_priority() < Modality::Document.execution_priority()); - assert!(Modality::Document.execution_priority() < Modality::Graph.execution_priority()); - assert!(Modality::Graph.execution_priority() < Modality::Tensor.execution_priority()); - assert!(Modality::Tensor.execution_priority() < Modality::Semantic.execution_priority()); - } -} diff --git a/verisimdb/rust-core/verisim-planner/src/optimizer.rs b/verisimdb/rust-core/verisim-planner/src/optimizer.rs deleted file mode 100644 index dac0c25b..00000000 --- a/verisimdb/rust-core/verisim-planner/src/optimizer.rs +++ /dev/null @@ -1,358 +0,0 @@ -// SPDX-License-Identifier: MPL-2.0 -//! Query optimizer — transforms logical plans into physical plans. - -use tracing::debug; - -use crate::config::PlannerConfig; -use crate::cost::{CostEstimate, CostModel}; -use crate::error::PlannerError; -use crate::explain::ExplainOutput; -use crate::plan::{ExecutionStrategy, LogicalPlan, PhysicalPlan, PlanStep}; -use crate::stats::StatisticsCollector; - -/// The query planner/optimizer. -/// -/// Transforms a `LogicalPlan` into an optimized `PhysicalPlan` by: -/// 1. Estimating cost per modality node -/// 2. Reordering by execution priority + cost -/// 3. Selecting sequential vs parallel strategy -/// 4. Generating optimization hints -pub struct Planner { - config: PlannerConfig, - stats: StatisticsCollector, -} - -impl Planner { - /// Create a new planner with the given configuration. - pub fn new(config: PlannerConfig) -> Self { - Self { - config, - stats: StatisticsCollector::new(), - } - } - - /// Get a reference to the current configuration. - pub fn config(&self) -> &PlannerConfig { - &self.config - } - - /// Update the planner configuration. - pub fn set_config(&mut self, config: PlannerConfig) { - self.config = config; - } - - /// Get a reference to the statistics collector. - pub fn stats(&self) -> &StatisticsCollector { - &self.stats - } - - /// Get a mutable reference to the statistics collector. - pub fn stats_mut(&mut self) -> &mut StatisticsCollector { - &mut self.stats - } - - /// Optimize a logical plan into a physical plan. - pub fn optimize(&self, logical: &LogicalPlan) -> Result<PhysicalPlan, PlannerError> { - if logical.nodes.is_empty() { - return Err(PlannerError::EmptyPlan); - } - - debug!( - node_count = logical.nodes.len(), - "Optimizing logical plan" - ); - - // 1. Estimate cost for each node - let mut node_costs: Vec<(usize, CostEstimate, Option<String>)> = logical - .nodes - .iter() - .enumerate() - .map(|(i, node)| { - let stats = self.stats.get(node.modality); - let cost = CostModel::estimate(node, &self.config, stats); - let hint = CostModel::optimization_hint(node); - (i, cost, hint) - }) - .collect(); - - // 2. Sort by execution priority first, then by total cost within same priority - node_costs.sort_by(|a, b| { - let pri_a = logical.nodes[a.0].modality.execution_priority(); - let pri_b = logical.nodes[b.0].modality.execution_priority(); - pri_a - .cmp(&pri_b) - .then_with(|| a.1.time_ms.partial_cmp(&b.1.time_ms).unwrap_or(std::cmp::Ordering::Equal)) - }); - - // 3. Select execution strategy - let strategy = if logical.nodes.len() >= self.config.parallel_threshold { - ExecutionStrategy::Parallel - } else { - ExecutionStrategy::Sequential - }; - - // 4. Build physical plan steps - let mut steps = Vec::with_capacity(node_costs.len()); - let mut cost_estimates = Vec::with_capacity(node_costs.len()); - let mut notes = Vec::new(); - - for (step_num, &(node_idx, ref cost, ref hint)) in node_costs.iter().enumerate() { - let node = &logical.nodes[node_idx]; - - let operation = format!( - "{} {}", - match node.modality { - crate::Modality::Graph => "Graph traversal", - crate::Modality::Vector => "Vector similarity search", - crate::Modality::Tensor => "Tensor computation", - crate::Modality::Semantic => "Semantic verification", - crate::Modality::Document => "Document fulltext search", - crate::Modality::Temporal => "Temporal version lookup", - }, - if node.conditions.is_empty() { - "(scan)".to_string() - } else { - format!("({} conditions)", node.conditions.len()) - } - ); - - let pushed_predicates: Vec<String> = node - .conditions - .iter() - .map(|c| format!("{:?}", c)) - .collect(); - - steps.push(PlanStep { - step: step_num + 1, - operation, - modality: node.modality, - cost: cost.clone(), - optimization_hint: hint.clone(), - pushed_predicates, - }); - - cost_estimates.push(cost.clone()); - } - - // 5. Combine total cost (modality queries + post-processing) - let is_parallel = strategy == ExecutionStrategy::Parallel; - let modality_cost = CostEstimate::combine(&cost_estimates, is_parallel); - let total_cost = CostModel::estimate_with_post_processing( - &modality_cost, - &logical.post_processing, - ); - - // 6. Generate optimization notes - if is_parallel { - notes.push(format!( - "Parallel execution across {} modalities", - steps.len() - )); - } else { - notes.push("Sequential execution — single modality".to_string()); - } - - if total_cost.time_ms > 500.0 { - notes.push("High estimated cost — consider adding LIMIT or more selective predicates".to_string()); - } - - // Check if any step has poor selectivity - for step in &steps { - if step.cost.selectivity > 0.5 && step.cost.time_ms > 100.0 { - notes.push(format!( - "Step {}: {} has high selectivity ({:.0}%) — may benefit from additional predicates", - step.step, - step.modality, - step.cost.selectivity * 100.0 - )); - } - } - - Ok(PhysicalPlan { - steps, - strategy, - total_cost, - notes, - }) - } - - /// Generate an EXPLAIN output for a logical plan. - pub fn explain(&self, logical: &LogicalPlan) -> Result<ExplainOutput, PlannerError> { - let physical = self.optimize(logical)?; - Ok(ExplainOutput::from_physical_plan(&physical, &self.config)) - } -} - -#[cfg(test)] -mod tests { - use super::*; - use crate::plan::{ConditionKind, LogicalPlan, PlanNode, QuerySource}; - use crate::Modality; - - fn graph_vector_plan() -> LogicalPlan { - LogicalPlan { - source: QuerySource::Octad, - nodes: vec![ - PlanNode { - modality: Modality::Graph, - conditions: vec![ConditionKind::Traversal { - predicate: "relates_to".to_string(), - depth: Some(2), - }], - projections: vec![], - early_limit: None, - }, - PlanNode { - modality: Modality::Vector, - conditions: vec![ConditionKind::Similarity { k: 10 }], - projections: vec![], - early_limit: None, - }, - ], - post_processing: vec![], - } - } - - #[test] - fn test_single_modality_sequential() { - let planner = Planner::new(PlannerConfig::default()); - let plan = LogicalPlan { - source: QuerySource::Octad, - nodes: vec![PlanNode { - modality: Modality::Document, - conditions: vec![ConditionKind::Fulltext { - query: "test".to_string(), - }], - projections: vec![], - early_limit: None, - }], - post_processing: vec![], - }; - - let physical = planner.optimize(&plan).expect("TODO: handle error"); - assert_eq!(physical.strategy, ExecutionStrategy::Sequential); - assert_eq!(physical.steps.len(), 1); - } - - #[test] - fn test_multi_modality_parallel() { - let planner = Planner::new(PlannerConfig::default()); - let physical = planner.optimize(&graph_vector_plan()).expect("TODO: handle error"); - assert_eq!(physical.strategy, ExecutionStrategy::Parallel); - assert_eq!(physical.steps.len(), 2); - } - - #[test] - fn test_vector_before_graph() { - let planner = Planner::new(PlannerConfig::default()); - let physical = planner.optimize(&graph_vector_plan()).expect("TODO: handle error"); - - // Vector has priority 20, Graph has priority 40 → Vector first - assert_eq!(physical.steps[0].modality, Modality::Vector); - assert_eq!(physical.steps[1].modality, Modality::Graph); - } - - #[test] - fn test_semantic_always_last() { - let planner = Planner::new(PlannerConfig::default()); - let plan = LogicalPlan { - source: QuerySource::Octad, - nodes: vec![ - PlanNode { - modality: Modality::Semantic, - conditions: vec![ConditionKind::ProofVerification { - contract: "test".to_string(), - }], - projections: vec![], - early_limit: None, - }, - PlanNode { - modality: Modality::Document, - conditions: vec![], - projections: vec![], - early_limit: None, - }, - PlanNode { - modality: Modality::Vector, - conditions: vec![], - projections: vec![], - early_limit: None, - }, - ], - post_processing: vec![], - }; - - let physical = planner.optimize(&plan).expect("TODO: handle error"); - let last = physical.steps.last().expect("TODO: handle error"); - assert_eq!(last.modality, Modality::Semantic); - } - - #[test] - fn test_temporal_always_first() { - let planner = Planner::new(PlannerConfig::default()); - let plan = LogicalPlan { - source: QuerySource::Octad, - nodes: vec![ - PlanNode { - modality: Modality::Graph, - conditions: vec![], - projections: vec![], - early_limit: None, - }, - PlanNode { - modality: Modality::Temporal, - conditions: vec![ConditionKind::AtTime { - timestamp: "2026-01-01T00:00:00Z".to_string(), - }], - projections: vec![], - early_limit: None, - }, - ], - post_processing: vec![], - }; - - let physical = planner.optimize(&plan).expect("TODO: handle error"); - assert_eq!(physical.steps[0].modality, Modality::Temporal); - } - - #[test] - fn test_empty_plan_error() { - let planner = Planner::new(PlannerConfig::default()); - let plan = LogicalPlan { - source: QuerySource::Octad, - nodes: vec![], - post_processing: vec![], - }; - - let result = planner.optimize(&plan); - assert!(result.is_err()); - assert!(matches!(result.unwrap_err(), PlannerError::EmptyPlan)); - } - - #[test] - fn test_explain_generates_output() { - let planner = Planner::new(PlannerConfig::default()); - let explain = planner.explain(&graph_vector_plan()).expect("TODO: handle error"); - assert_eq!(explain.steps.len(), 2); - assert!(!explain.text_output.is_empty()); - } - - #[test] - fn test_integration_graph_vector() { - let planner = Planner::new(PlannerConfig::default()); - let physical = planner.optimize(&graph_vector_plan()).expect("TODO: handle error"); - - // Vector ordered before Graph - assert_eq!(physical.steps[0].modality, Modality::Vector); - assert_eq!(physical.steps[1].modality, Modality::Graph); - - // Parallel strategy - assert_eq!(physical.strategy, ExecutionStrategy::Parallel); - - // EXPLAIN output contains expected sections - let explain = planner.explain(&graph_vector_plan()).expect("TODO: handle error"); - assert!(explain.text_output.contains("Step")); - assert!(explain.text_output.contains("Strategy")); - assert!(explain.text_output.contains("vector")); - } -} diff --git a/verisimdb/rust-core/verisim-planner/src/plan.rs b/verisimdb/rust-core/verisim-planner/src/plan.rs deleted file mode 100644 index 0d08954d..00000000 --- a/verisimdb/rust-core/verisim-planner/src/plan.rs +++ /dev/null @@ -1,210 +0,0 @@ -// SPDX-License-Identifier: MPL-2.0 -//! Logical and physical plan types. - -use serde::{Deserialize, Serialize}; - -use crate::cost::CostEstimate; -use crate::Modality; - -/// Source of data for the query. -#[derive(Debug, Clone, Serialize, Deserialize)] -#[serde(rename_all = "snake_case")] -pub enum QuerySource { - /// Query a single octad store. - Octad, - /// Federated query across multiple nodes. - Federation { nodes: Vec<String> }, - /// Direct store access for a specific modality. - Store { modality: Modality }, -} - -/// Kind of condition applied to a modality node. -#[derive(Debug, Clone, Serialize, Deserialize)] -#[serde(rename_all = "snake_case")] -pub enum ConditionKind { - /// Equality filter (field = value). - Equality { field: String, value: String }, - /// Range filter (field BETWEEN low AND high). - Range { field: String, low: String, high: String }, - /// Full-text search. - Fulltext { query: String }, - /// Vector similarity (k-NN). - Similarity { k: usize }, - /// Graph traversal. - Traversal { predicate: String, depth: Option<u32> }, - /// Temporal version lookup. - AtTime { timestamp: String }, - /// ZKP proof verification. - ProofVerification { contract: String }, - /// Tensor operation. - TensorOp { operation: String }, - /// Generic predicate. - Predicate { expression: String }, -} - -/// A single modality node in a logical plan. -#[derive(Debug, Clone, Serialize, Deserialize)] -pub struct PlanNode { - /// Which modality this node queries. - pub modality: Modality, - /// Conditions/filters to apply. - pub conditions: Vec<ConditionKind>, - /// Fields to project (empty = all). - pub projections: Vec<String>, - /// Early limit pushed down to store. - pub early_limit: Option<usize>, -} - -/// Post-processing operation applied after modality queries. -#[derive(Debug, Clone, Serialize, Deserialize)] -#[serde(rename_all = "snake_case")] -pub enum PostProcessing { - /// ORDER BY fields. - OrderBy { fields: Vec<(String, bool)> }, - /// LIMIT result count. - Limit { count: usize }, - /// GROUP BY + aggregation. - GroupBy { fields: Vec<String>, aggregates: Vec<String> }, - /// Final projection. - Project { columns: Vec<String> }, -} - -/// A logical plan — the unoptimized query representation. -#[derive(Debug, Clone, Serialize, Deserialize)] -pub struct LogicalPlan { - /// Data source. - pub source: QuerySource, - /// Per-modality query nodes. - pub nodes: Vec<PlanNode>, - /// Post-processing steps. - pub post_processing: Vec<PostProcessing>, -} - -/// Execution strategy for the physical plan. -#[derive(Debug, Clone, Copy, PartialEq, Eq, Serialize, Deserialize)] -#[serde(rename_all = "lowercase")] -pub enum ExecutionStrategy { - Sequential, - Parallel, -} - -/// A single step in a physical plan. -#[derive(Debug, Clone, Serialize, Deserialize)] -pub struct PlanStep { - /// Step number (1-indexed). - pub step: usize, - /// Operation description. - pub operation: String, - /// Target modality. - pub modality: Modality, - /// Cost estimate for this step. - pub cost: CostEstimate, - /// Optimization hint for this step. - pub optimization_hint: Option<String>, - /// Pushed-down predicates. - pub pushed_predicates: Vec<String>, -} - -/// An optimized physical plan ready for execution. -#[derive(Debug, Clone, Serialize, Deserialize)] -pub struct PhysicalPlan { - /// Ordered execution steps. - pub steps: Vec<PlanStep>, - /// Overall execution strategy. - pub strategy: ExecutionStrategy, - /// Total estimated cost. - pub total_cost: CostEstimate, - /// Optimization notes. - pub notes: Vec<String>, -} - -#[cfg(test)] -mod tests { - use super::*; - - fn sample_logical_plan() -> LogicalPlan { - LogicalPlan { - source: QuerySource::Octad, - nodes: vec![ - PlanNode { - modality: Modality::Graph, - conditions: vec![ConditionKind::Traversal { - predicate: "relates_to".to_string(), - depth: Some(2), - }], - projections: vec!["id".to_string(), "label".to_string()], - early_limit: None, - }, - PlanNode { - modality: Modality::Vector, - conditions: vec![ConditionKind::Similarity { k: 10 }], - projections: vec![], - early_limit: Some(50), - }, - ], - post_processing: vec![PostProcessing::Limit { count: 10 }], - } - } - - #[test] - fn test_logical_plan_json_roundtrip() { - let plan = sample_logical_plan(); - let json = serde_json::to_string(&plan).expect("TODO: handle error"); - let parsed: LogicalPlan = serde_json::from_str(&json).expect("TODO: handle error"); - assert_eq!(parsed.nodes.len(), 2); - assert_eq!(parsed.nodes[0].modality, Modality::Graph); - assert_eq!(parsed.nodes[1].modality, Modality::Vector); - } - - #[test] - fn test_physical_plan_json_roundtrip() { - let plan = PhysicalPlan { - steps: vec![PlanStep { - step: 1, - operation: "Vector similarity search".to_string(), - modality: Modality::Vector, - cost: CostEstimate { - time_ms: 50.0, - estimated_rows: 10, - selectivity: 0.01, - io_cost: 30.0, - cpu_cost: 20.0, - }, - optimization_hint: Some("HNSW ANN".to_string()), - pushed_predicates: vec!["k=10".to_string()], - }], - strategy: ExecutionStrategy::Sequential, - total_cost: CostEstimate { - time_ms: 50.0, - estimated_rows: 10, - selectivity: 0.01, - io_cost: 30.0, - cpu_cost: 20.0, - }, - notes: vec!["Single modality — sequential execution".to_string()], - }; - let json = serde_json::to_string(&plan).expect("TODO: handle error"); - let parsed: PhysicalPlan = serde_json::from_str(&json).expect("TODO: handle error"); - assert_eq!(parsed.steps.len(), 1); - assert_eq!(parsed.strategy, ExecutionStrategy::Sequential); - } - - #[test] - fn test_query_source_variants() { - let octad = QuerySource::Octad; - let json = serde_json::to_string(&octad).expect("TODO: handle error"); - assert!(json.contains("octad")); - - let fed = QuerySource::Federation { - nodes: vec!["node1".to_string()], - }; - let json = serde_json::to_string(&fed).expect("TODO: handle error"); - assert!(json.contains("federation")); - - let store = QuerySource::Store { - modality: Modality::Graph, - }; - let json = serde_json::to_string(&store).expect("TODO: handle error"); - assert!(json.contains("store")); - } -} diff --git a/verisimdb/rust-core/verisim-planner/src/prepared.rs b/verisimdb/rust-core/verisim-planner/src/prepared.rs deleted file mode 100644 index 2d796520..00000000 --- a/verisimdb/rust-core/verisim-planner/src/prepared.rs +++ /dev/null @@ -1,1119 +0,0 @@ -// SPDX-License-Identifier: MPL-2.0 -// Copyright (c) 2026 Jonathan D.A. Jewell (hyperpolymath) <j.d.a.jewell@open.ac.uk> - -//! Prepared statements and query plan caching. -//! -//! This module implements Phase 5.2 of the VeriSimDB roadmap: Prepared Statements / Query Caching. -//! It provides a [`PlanCache`] that stores parsed and planned query statements, avoiding redundant -//! parsing and planning for repeated queries. The cache supports: -//! -//! - **Prepared statements**: Parse once, execute many times with different parameters. -//! - **Query fingerprinting**: Normalize query text so that semantically identical queries share -//! a single cached plan (whitespace-insensitive, keyword-case-insensitive). -//! - **Physical plan caching**: Optionally cache the optimized physical plan alongside the -//! logical plan. -//! - **LRU eviction**: When the cache exceeds `max_entries`, the least-recently-used statement -//! is evicted. -//! - **TTL expiration**: Entries older than `ttl_seconds` are considered expired and removed on -//! access or during explicit eviction sweeps. -//! - **Hit/miss statistics**: Track cache effectiveness with atomic counters. - -use std::collections::HashMap; -use std::fmt; -use std::sync::atomic::{AtomicU64, Ordering}; - -use chrono::{DateTime, Utc}; -use serde::{Deserialize, Serialize}; -use sha2::{Digest, Sha256}; -use tokio::sync::RwLock; - -use crate::plan::{LogicalPlan, PhysicalPlan}; - -// --------------------------------------------------------------------------- -// Types -// --------------------------------------------------------------------------- - -/// A unique identifier for a prepared statement. -/// -/// Derived from a SHA-256 fingerprint of the normalized query text, ensuring that -/// semantically identical queries map to the same ID. -#[derive(Debug, Clone, Hash, Eq, PartialEq, Serialize, Deserialize)] -pub struct PreparedId(String); - -impl PreparedId { - /// Create a new `PreparedId` from a raw string. - pub fn new(id: impl Into<String>) -> Self { - Self(id.into()) - } - - /// Return the underlying string representation. - pub fn as_str(&self) -> &str { - &self.0 - } -} - -impl fmt::Display for PreparedId { - fn fmt(&self, f: &mut fmt::Formatter<'_>) -> fmt::Result { - write!(f, "prep_{}", &self.0[..12.min(self.0.len())]) - } -} - -/// A parameter value that can be bound to a prepared statement placeholder. -/// -/// When a prepared statement contains named parameters (e.g. `$name`, `$threshold`), -/// concrete values are supplied at execution time via a map of `String -> ParamValue`. -#[derive(Debug, Clone, Serialize, Deserialize)] -#[serde(rename_all = "snake_case")] -pub enum ParamValue { - /// A UTF-8 string value. - String(String), - /// A signed 64-bit integer. - Int(i64), - /// A 64-bit floating-point number. - Float(f64), - /// A boolean value. - Bool(bool), - /// A vector of 32-bit floats (e.g. an embedding). - Vector(Vec<f32>), - /// An explicit SQL-style NULL. - Null, -} - -impl fmt::Display for ParamValue { - fn fmt(&self, f: &mut fmt::Formatter<'_>) -> fmt::Result { - match self { - ParamValue::String(s) => write!(f, "\"{}\"", s), - ParamValue::Int(i) => write!(f, "{}", i), - ParamValue::Float(v) => write!(f, "{}", v), - ParamValue::Bool(b) => write!(f, "{}", b), - ParamValue::Vector(v) => write!(f, "vec[{}]", v.len()), - ParamValue::Null => write!(f, "NULL"), - } - } -} - -/// A prepared (parsed + planned) statement stored in the cache. -/// -/// Contains both the logical plan (always present) and an optional cached physical plan. -/// Tracks usage statistics for LRU eviction and performance monitoring. -#[derive(Debug, Clone, Serialize, Deserialize)] -pub struct PreparedStatement { - /// Unique identifier derived from the query fingerprint. - pub id: PreparedId, - /// The original query text as submitted by the user. - pub original_query: String, - /// Named parameters extracted from the query (e.g. `["$name", "$threshold"]`). - pub parameter_names: Vec<String>, - /// The parsed logical plan (modality-independent). - pub logical_plan: LogicalPlan, - /// Optionally cached optimized physical plan. - pub cached_physical_plan: Option<PhysicalPlan>, - /// Timestamp when this statement was first prepared. - pub created_at: DateTime<Utc>, - /// Timestamp of the most recent access (prepare or execute). - pub last_used: DateTime<Utc>, - /// Number of times this statement has been executed. - pub use_count: u64, - /// Rolling average execution time in milliseconds. - pub avg_execution_ms: f64, -} - -/// Configuration for the plan cache. -#[derive(Debug, Clone, Serialize, Deserialize)] -pub struct CacheConfig { - /// Maximum number of prepared statements to keep in the cache. - pub max_entries: usize, - /// Time-to-live in seconds; entries older than this are eligible for eviction. - pub ttl_seconds: u64, - /// Whether to cache optimized physical plans alongside logical plans. - pub enable_plan_cache: bool, - /// Whether to cache query result sets (reserved for future use). - pub enable_result_cache: bool, - /// Maximum bytes allowed for the result cache (reserved for future use). - pub max_result_cache_bytes: usize, -} - -impl Default for CacheConfig { - fn default() -> Self { - Self { - max_entries: 1024, - ttl_seconds: 3600, - enable_plan_cache: true, - enable_result_cache: false, - max_result_cache_bytes: 64 * 1024 * 1024, // 64 MiB - } - } -} - -/// Aggregate statistics about cache performance. -#[derive(Debug, Clone, Serialize, Deserialize)] -pub struct CacheStats { - /// Number of statements currently in the cache. - pub total_entries: usize, - /// Total number of cache hits (lookups that found an existing entry). - pub hit_count: u64, - /// Total number of cache misses (lookups that did not find an entry). - pub miss_count: u64, - /// Total number of entries evicted (TTL or LRU). - pub eviction_count: u64, - /// Hit ratio: `hit_count / (hit_count + miss_count)`, or 0.0 if no lookups. - pub hit_ratio: f64, - /// Monotonically increasing generation counter; bumped on every mutation. - pub generation: u64, -} - -/// Errors that can occur during cache operations. -#[derive(Debug, Clone, Serialize, Deserialize)] -pub enum CacheError { - /// The requested prepared statement was not found in the cache. - NotFound(String), - /// The supplied parameters do not match the statement's declared parameter names. - ParameterMismatch { - /// Parameter names the statement expects. - expected: Vec<String>, - /// Parameter names that were actually provided. - provided: Vec<String>, - }, - /// The prepared statement has expired (exceeded TTL). - Expired(String), - /// The cache is full and no entry could be evicted to make room. - CacheFull, -} - -impl fmt::Display for CacheError { - fn fmt(&self, f: &mut fmt::Formatter<'_>) -> fmt::Result { - match self { - CacheError::NotFound(id) => write!(f, "prepared statement not found: {}", id), - CacheError::ParameterMismatch { expected, provided } => { - write!( - f, - "parameter mismatch: expected {:?}, provided {:?}", - expected, provided - ) - } - CacheError::Expired(id) => write!(f, "prepared statement expired: {}", id), - CacheError::CacheFull => write!(f, "cache is full"), - } - } -} - -impl std::error::Error for CacheError {} - -// --------------------------------------------------------------------------- -// PlanCache -// --------------------------------------------------------------------------- - -/// The plan cache: maps query fingerprints to prepared statements with physical plans. -/// -/// Thread-safe via `tokio::sync::RwLock` — multiple readers can access cached plans -/// concurrently, while mutations (prepare, invalidate, evict) take a write lock. -/// -/// # Example -/// -/// ```rust,no_run -/// use verisim_planner::prepared::{PlanCache, CacheConfig}; -/// use verisim_planner::plan::{LogicalPlan, QuerySource}; -/// -/// # async fn example() { -/// let cache = PlanCache::new(CacheConfig::default()); -/// -/// let plan = LogicalPlan { -/// source: QuerySource::Octad, -/// nodes: vec![], -/// post_processing: vec![], -/// }; -/// -/// let id = cache.prepare("SEARCH graph WHERE type = $t", plan).await; -/// let stmt = cache.get(&id).await; -/// # } -/// ``` -pub struct PlanCache { - /// Cache configuration (immutable after construction). - config: CacheConfig, - /// Map from `PreparedId` to the full `PreparedStatement`. - statements: RwLock<HashMap<PreparedId, PreparedStatement>>, - /// Map from query fingerprint string to `PreparedId` for fast lookup. - fingerprints: RwLock<HashMap<String, PreparedId>>, - /// Monotonically increasing generation counter; bumped on every mutation. - generation: AtomicU64, - /// Total cache hits. - hit_count: AtomicU64, - /// Total cache misses. - miss_count: AtomicU64, - /// Total evictions performed. - eviction_count: AtomicU64, -} - -impl PlanCache { - /// Create a new plan cache with the given configuration. - pub fn new(config: CacheConfig) -> Self { - Self { - config, - statements: RwLock::new(HashMap::new()), - fingerprints: RwLock::new(HashMap::new()), - generation: AtomicU64::new(0), - hit_count: AtomicU64::new(0), - miss_count: AtomicU64::new(0), - eviction_count: AtomicU64::new(0), - } - } - - /// Prepare a query: parse once, store the logical plan, return a stable ID. - /// - /// If a statement with the same fingerprint already exists, its `last_used` and - /// `use_count` are updated and the existing ID is returned. - pub async fn prepare(&self, query: &str, logical_plan: LogicalPlan) -> PreparedId { - let fp = Self::fingerprint(query); - let id = PreparedId::new(&fp); - - // Check if we already have this fingerprint cached. - { - let stmts = self.statements.read().await; - if stmts.contains_key(&id) { - drop(stmts); - // Update last_used on existing entry. - let mut stmts = self.statements.write().await; - if let Some(stmt) = stmts.get_mut(&id) { - stmt.last_used = Utc::now(); - } - self.generation.fetch_add(1, Ordering::Relaxed); - return id; - } - } - - let now = Utc::now(); - let param_names = Self::extract_parameters(query); - - let stmt = PreparedStatement { - id: id.clone(), - original_query: query.to_string(), - parameter_names: param_names, - logical_plan, - cached_physical_plan: None, - created_at: now, - last_used: now, - use_count: 0, - avg_execution_ms: 0.0, - }; - - { - let mut stmts = self.statements.write().await; - let mut fps = self.fingerprints.write().await; - - // Evict LRU if at capacity. - if stmts.len() >= self.config.max_entries { - self.evict_lru_inner(&mut stmts, &mut fps); - } - - stmts.insert(id.clone(), stmt); - fps.insert(fp, id.clone()); - } - - self.generation.fetch_add(1, Ordering::Relaxed); - id - } - - /// Retrieve a prepared statement by its ID. - /// - /// Returns `None` if the ID is not in the cache or the entry has expired. - pub async fn get(&self, id: &PreparedId) -> Option<PreparedStatement> { - let stmts = self.statements.read().await; - match stmts.get(id) { - Some(stmt) => { - if self.is_expired(stmt) { - drop(stmts); - self.miss_count.fetch_add(1, Ordering::Relaxed); - // Lazy expiration: remove on next write. - None - } else { - self.hit_count.fetch_add(1, Ordering::Relaxed); - Some(stmt.clone()) - } - } - None => { - self.miss_count.fetch_add(1, Ordering::Relaxed); - None - } - } - } - - /// Execute a prepared statement: validate parameters, increment counters, return statement. - /// - /// This does not actually run the query against the storage engine — it returns the - /// statement with updated usage statistics so the caller can feed it to the executor. - pub async fn execute_prepared( - &self, - id: &PreparedId, - params: &HashMap<String, ParamValue>, - ) -> Result<PreparedStatement, CacheError> { - let mut stmts = self.statements.write().await; - let stmt = stmts - .get_mut(id) - .ok_or_else(|| CacheError::NotFound(id.as_str().to_string()))?; - - // Check TTL expiration. - if self.is_expired(stmt) { - let id_str = id.as_str().to_string(); - stmts.remove(id); - return Err(CacheError::Expired(id_str)); - } - - // Validate that the provided parameter names match expected ones. - if !stmt.parameter_names.is_empty() { - let mut expected_sorted = stmt.parameter_names.clone(); - expected_sorted.sort(); - - let mut provided_sorted: Vec<String> = params.keys().cloned().collect(); - provided_sorted.sort(); - - if expected_sorted != provided_sorted { - return Err(CacheError::ParameterMismatch { - expected: stmt.parameter_names.clone(), - provided: params.keys().cloned().collect(), - }); - } - } - - // Update usage statistics. - stmt.use_count += 1; - stmt.last_used = Utc::now(); - - self.hit_count.fetch_add(1, Ordering::Relaxed); - self.generation.fetch_add(1, Ordering::Relaxed); - - Ok(stmt.clone()) - } - - /// Invalidate (remove) a specific prepared statement from the cache. - /// - /// Returns `true` if the statement existed and was removed. - pub async fn invalidate(&self, id: &PreparedId) -> bool { - let mut stmts = self.statements.write().await; - let mut fps = self.fingerprints.write().await; - - if let Some(stmt) = stmts.remove(id) { - let fp = Self::fingerprint(&stmt.original_query); - fps.remove(&fp); - self.eviction_count.fetch_add(1, Ordering::Relaxed); - self.generation.fetch_add(1, Ordering::Relaxed); - true - } else { - false - } - } - - /// Invalidate all prepared statements (e.g. on schema change). - pub async fn invalidate_all(&self) { - let mut stmts = self.statements.write().await; - let mut fps = self.fingerprints.write().await; - - let count = stmts.len() as u64; - stmts.clear(); - fps.clear(); - - self.eviction_count.fetch_add(count, Ordering::Relaxed); - self.generation.fetch_add(1, Ordering::Relaxed); - } - - /// Compute a deterministic fingerprint for a query string. - /// - /// Normalization rules: - /// 1. Collapse all whitespace (spaces, tabs, newlines) into single spaces. - /// 2. Trim leading and trailing whitespace. - /// 3. Lowercase VCL/SQL keywords (SELECT, WHERE, FROM, SEARCH, LIMIT, ORDER, BY, GROUP, - /// AND, OR, NOT, JOIN, ON, AS, HAVING, INSERT, UPDATE, DELETE, SET, INTO, VALUES, - /// WITH, UNION, INTERSECT, EXCEPT, EXISTS, BETWEEN, LIKE, IN, IS, NULL, TRUE, FALSE, - /// ASC, DESC, DISTINCT, ALL, ANY, SOME, CASE, WHEN, THEN, ELSE, END, PROOF, VERIFY, - /// DRIFT, OCTAD, MODALITY). - /// 4. SHA-256 hash the normalized text and return as hex. - pub fn fingerprint(query: &str) -> String { - let normalized = Self::normalize_query(query); - - let mut hasher = Sha256::new(); - hasher.update(normalized.as_bytes()); - let result = hasher.finalize(); - - // Convert to hex string. - result - .iter() - .map(|byte| format!("{:02x}", byte)) - .collect::<String>() - } - - /// Look up an existing prepared statement by raw query text. - /// - /// Returns the `PreparedId` if a statement with the same fingerprint exists, or `None`. - pub async fn lookup_by_query(&self, query: &str) -> Option<PreparedId> { - let fp = Self::fingerprint(query); - let fps = self.fingerprints.read().await; - - match fps.get(&fp) { - Some(id) => { - self.hit_count.fetch_add(1, Ordering::Relaxed); - Some(id.clone()) - } - None => { - self.miss_count.fetch_add(1, Ordering::Relaxed); - None - } - } - } - - /// Cache an optimized physical plan for an existing prepared statement. - /// - /// This is called after the optimizer produces a physical plan, so subsequent executions - /// can skip the optimization step entirely. - pub async fn cache_plan(&self, id: &PreparedId, plan: PhysicalPlan) { - let mut stmts = self.statements.write().await; - if let Some(stmt) = stmts.get_mut(id) { - stmt.cached_physical_plan = Some(plan); - self.generation.fetch_add(1, Ordering::Relaxed); - } - } - - /// Return aggregate cache statistics. - pub fn stats(&self) -> CacheStats { - let hits = self.hit_count.load(Ordering::Relaxed); - let misses = self.miss_count.load(Ordering::Relaxed); - let total_lookups = hits + misses; - let hit_ratio = if total_lookups > 0 { - hits as f64 / total_lookups as f64 - } else { - 0.0 - }; - - // We cannot await inside a non-async function, so we report the generation - // and eviction counters which are always available atomically. The total_entries - // field requires a blocking read — callers who need it should use `stats_async`. - CacheStats { - total_entries: 0, // Populated by stats_async; sync callers get 0. - hit_count: hits, - miss_count: misses, - eviction_count: self.eviction_count.load(Ordering::Relaxed), - hit_ratio, - generation: self.generation.load(Ordering::Relaxed), - } - } - - /// Return aggregate cache statistics (async version with accurate entry count). - pub async fn stats_async(&self) -> CacheStats { - let stmts = self.statements.read().await; - let mut s = self.stats(); - s.total_entries = stmts.len(); - s - } - - /// Evict all entries whose `created_at` is older than the configured TTL. - /// - /// Returns the number of entries evicted. - pub async fn evict_expired(&self) -> usize { - let now = Utc::now(); - let ttl = chrono::Duration::seconds(self.config.ttl_seconds as i64); - - let mut stmts = self.statements.write().await; - let mut fps = self.fingerprints.write().await; - - let expired_ids: Vec<PreparedId> = stmts - .iter() - .filter(|(_, stmt)| { - now.signed_duration_since(stmt.created_at) > ttl - }) - .map(|(id, _)| id.clone()) - .collect(); - - let count = expired_ids.len(); - for id in &expired_ids { - if let Some(stmt) = stmts.remove(id) { - let fp = Self::fingerprint(&stmt.original_query); - fps.remove(&fp); - } - } - - self.eviction_count - .fetch_add(count as u64, Ordering::Relaxed); - if count > 0 { - self.generation.fetch_add(1, Ordering::Relaxed); - } - - count - } - - /// Evict the least-recently-used entry when the cache exceeds `max_entries`. - /// - /// Returns the number of entries evicted (0 or 1). - pub async fn evict_lru(&self) -> usize { - let mut stmts = self.statements.write().await; - let mut fps = self.fingerprints.write().await; - - if stmts.len() <= self.config.max_entries { - return 0; - } - - self.evict_lru_inner(&mut stmts, &mut fps) - } - - // ----------------------------------------------------------------------- - // Internal helpers - // ----------------------------------------------------------------------- - - /// Inner LRU eviction that operates on already-locked maps. - /// - /// Removes the single entry with the oldest `last_used` timestamp. - /// Returns the number of entries evicted (0 or 1). - fn evict_lru_inner( - &self, - stmts: &mut HashMap<PreparedId, PreparedStatement>, - fps: &mut HashMap<String, PreparedId>, - ) -> usize { - if stmts.is_empty() { - return 0; - } - - // Find the entry with the oldest `last_used`. - let oldest_id = stmts - .iter() - .min_by_key(|(_, stmt)| stmt.last_used) - .map(|(id, _)| id.clone()); - - if let Some(id) = oldest_id { - if let Some(stmt) = stmts.remove(&id) { - let fp = Self::fingerprint(&stmt.original_query); - fps.remove(&fp); - } - self.eviction_count.fetch_add(1, Ordering::Relaxed); - self.generation.fetch_add(1, Ordering::Relaxed); - 1 - } else { - 0 - } - } - - /// Check whether a prepared statement has exceeded its TTL. - fn is_expired(&self, stmt: &PreparedStatement) -> bool { - let now = Utc::now(); - let ttl = chrono::Duration::seconds(self.config.ttl_seconds as i64); - now.signed_duration_since(stmt.created_at) > ttl - } - - /// Normalize a query string for fingerprinting. - /// - /// Collapses whitespace, trims, and lowercases known keywords while preserving - /// the case of identifiers and string literals. - fn normalize_query(query: &str) -> String { - // Step 1: Collapse all whitespace into single spaces and trim. - let collapsed: String = query - .split_whitespace() - .collect::<Vec<&str>>() - .join(" "); - - // Step 2: Lowercase known keywords. - // We split on whitespace, check each token against the keyword list, - // and lowercase it if it matches. Non-keyword tokens keep original case. - let keywords = [ - // SQL/VCL standard keywords - "SELECT", "WHERE", "FROM", "SEARCH", "LIMIT", "ORDER", "BY", "GROUP", - "AND", "OR", "NOT", "JOIN", "ON", "AS", "HAVING", "INSERT", "UPDATE", - "DELETE", "SET", "INTO", "VALUES", "WITH", "UNION", "INTERSECT", "EXCEPT", - "EXISTS", "BETWEEN", "LIKE", "IN", "IS", "NULL", "TRUE", "FALSE", - "ASC", "DESC", "DISTINCT", "ALL", "ANY", "SOME", "CASE", "WHEN", "THEN", - "ELSE", "END", - // VeriSimDB-specific keywords - "PROOF", "VERIFY", "DRIFT", "OCTAD", "MODALITY", "TYPE", - // Modality names (treated as keywords for normalization) - "GRAPH", "VECTOR", "TENSOR", "SEMANTIC", "DOCUMENT", "TEMPORAL", - ]; - - collapsed - .split(' ') - .map(|token| { - if keywords.contains(&token.to_uppercase().as_str()) { - token.to_lowercase() - } else { - token.to_string() - } - }) - .collect::<Vec<String>>() - .join(" ") - } - - /// Extract parameter placeholder names from a query string. - /// - /// Parameters are identified by a `$` prefix followed by one or more word characters - /// (e.g. `$name`, `$threshold_1`). Duplicates are removed, order preserved. - fn extract_parameters(query: &str) -> Vec<String> { - let mut params = Vec::new(); - let mut seen = std::collections::HashSet::new(); - - let mut chars = query.chars().peekable(); - while let Some(ch) = chars.next() { - if ch == '$' { - let mut name = String::from("$"); - while let Some(&next) = chars.peek() { - if next.is_alphanumeric() || next == '_' { - name.push(next); - chars.next(); - } else { - break; - } - } - if name.len() > 1 && seen.insert(name.clone()) { - params.push(name); - } - } - } - - params - } -} - -// --------------------------------------------------------------------------- -// Tests -// --------------------------------------------------------------------------- - -#[cfg(test)] -mod tests { - use super::*; - use crate::plan::{ConditionKind, LogicalPlan, PlanNode, PostProcessing, QuerySource}; - use crate::Modality; - - /// Helper: build a simple logical plan with one graph node. - fn sample_logical_plan() -> LogicalPlan { - LogicalPlan { - source: QuerySource::Octad, - nodes: vec![PlanNode { - modality: Modality::Graph, - conditions: vec![ConditionKind::Traversal { - predicate: "relates_to".to_string(), - depth: Some(2), - }], - projections: vec!["id".to_string()], - early_limit: None, - }], - post_processing: vec![PostProcessing::Limit { count: 10 }], - } - } - - /// Helper: build a physical plan for caching tests. - fn sample_physical_plan() -> PhysicalPlan { - use crate::cost::CostEstimate; - use crate::plan::{ExecutionStrategy, PlanStep}; - - PhysicalPlan { - steps: vec![PlanStep { - step: 1, - operation: "Graph traversal (1 conditions)".to_string(), - modality: Modality::Graph, - cost: CostEstimate { - time_ms: 25.0, - estimated_rows: 100, - selectivity: 0.1, - io_cost: 15.0, - cpu_cost: 10.0, - }, - optimization_hint: Some("depth-limited BFS".to_string()), - pushed_predicates: vec!["relates_to".to_string()], - }], - strategy: ExecutionStrategy::Sequential, - total_cost: CostEstimate { - time_ms: 25.0, - estimated_rows: 100, - selectivity: 0.1, - io_cost: 15.0, - cpu_cost: 10.0, - }, - notes: vec!["Sequential execution — single modality".to_string()], - } - } - - // -- Test 1: Prepare and retrieve a statement -- - - #[tokio::test] - async fn test_prepare_and_retrieve() { - let cache = PlanCache::new(CacheConfig::default()); - let plan = sample_logical_plan(); - - let id = cache.prepare("SEARCH graph WHERE type = $t", plan.clone()).await; - let stmt = cache.get(&id).await; - - assert!(stmt.is_some(), "prepared statement should be retrievable"); - let stmt = stmt.expect("TODO: handle error"); - assert_eq!(stmt.id, id); - assert_eq!(stmt.original_query, "SEARCH graph WHERE type = $t"); - assert_eq!(stmt.parameter_names, vec!["$t"]); - assert_eq!(stmt.use_count, 0); - assert!(stmt.cached_physical_plan.is_none()); - } - - // -- Test 2: Query fingerprint normalization -- - - #[test] - fn test_fingerprint_normalization() { - // Whitespace normalization: extra spaces, tabs, newlines should produce same fingerprint. - let fp1 = PlanCache::fingerprint("SEARCH graph WHERE type = $t"); - let fp2 = PlanCache::fingerprint("SEARCH graph WHERE type = $t"); - let fp3 = PlanCache::fingerprint("SEARCH\tgraph\nWHERE\ttype = $t"); - assert_eq!(fp1, fp2, "extra whitespace should not change fingerprint"); - assert_eq!(fp2, fp3, "tabs/newlines should not change fingerprint"); - - // Keyword case normalization. - let fp4 = PlanCache::fingerprint("search GRAPH where TYPE = $t"); - assert_eq!(fp1, fp4, "keyword case should not change fingerprint"); - - // Different queries should produce different fingerprints. - let fp_other = PlanCache::fingerprint("SEARCH vector WHERE k = 10"); - assert_ne!(fp1, fp_other, "different queries must have different fingerprints"); - } - - // -- Test 3: Lookup by query (cache hit) -- - - #[tokio::test] - async fn test_lookup_by_query_hit() { - let cache = PlanCache::new(CacheConfig::default()); - let plan = sample_logical_plan(); - - let id = cache.prepare("SEARCH graph WHERE type = $t", plan).await; - let found = cache.lookup_by_query("SEARCH graph WHERE type = $t").await; - - assert_eq!(found, Some(id)); - } - - // -- Test 4: Cache miss returns None -- - - #[tokio::test] - async fn test_cache_miss_returns_none() { - let cache = PlanCache::new(CacheConfig::default()); - - let id = PreparedId::new("nonexistent"); - assert!(cache.get(&id).await.is_none(), "missing ID should return None"); - - let found = cache.lookup_by_query("SELECT * FROM nowhere").await; - assert!(found.is_none(), "unknown query should return None"); - } - - // -- Test 5: Invalidate specific statement -- - - #[tokio::test] - async fn test_invalidate_specific() { - let cache = PlanCache::new(CacheConfig::default()); - let plan = sample_logical_plan(); - - let id = cache.prepare("SEARCH graph WHERE type = $t", plan).await; - assert!(cache.get(&id).await.is_some()); - - let removed = cache.invalidate(&id).await; - assert!(removed, "invalidate should return true for existing entry"); - assert!(cache.get(&id).await.is_none(), "entry should be gone after invalidation"); - - let removed_again = cache.invalidate(&id).await; - assert!(!removed_again, "invalidate on missing entry should return false"); - } - - // -- Test 6: Invalidate all -- - - #[tokio::test] - async fn test_invalidate_all() { - let cache = PlanCache::new(CacheConfig::default()); - let plan = sample_logical_plan(); - - let id1 = cache.prepare("SEARCH graph WHERE type = $t", plan.clone()).await; - let id2 = cache.prepare("SEARCH vector WHERE k = 5", plan).await; - - assert!(cache.get(&id1).await.is_some()); - assert!(cache.get(&id2).await.is_some()); - - cache.invalidate_all().await; - - assert!(cache.get(&id1).await.is_none()); - assert!(cache.get(&id2).await.is_none()); - - let stats = cache.stats_async().await; - assert_eq!(stats.total_entries, 0); - } - - // -- Test 7: Execute increments use_count and updates last_used -- - - #[tokio::test] - async fn test_execute_increments_counters() { - let cache = PlanCache::new(CacheConfig::default()); - let plan = sample_logical_plan(); - - let id = cache.prepare("SEARCH graph WHERE type = $t", plan).await; - - let mut params = HashMap::new(); - params.insert("$t".to_string(), ParamValue::String("Person".to_string())); - - let stmt = cache.execute_prepared(&id, ¶ms).await.expect("TODO: handle error"); - assert_eq!(stmt.use_count, 1, "first execute should set use_count to 1"); - - let stmt2 = cache.execute_prepared(&id, ¶ms).await.expect("TODO: handle error"); - assert_eq!(stmt2.use_count, 2, "second execute should set use_count to 2"); - assert!( - stmt2.last_used >= stmt.last_used, - "last_used should advance on execute" - ); - } - - // -- Test 8: Evict expired entries -- - - #[tokio::test] - async fn test_evict_expired() { - // Use a TTL of 0 seconds so everything expires immediately. - let config = CacheConfig { - ttl_seconds: 0, - ..CacheConfig::default() - }; - let cache = PlanCache::new(config); - let plan = sample_logical_plan(); - - cache.prepare("SEARCH graph WHERE type = $t", plan.clone()).await; - cache.prepare("SEARCH vector WHERE k = 5", plan).await; - - // Sleep briefly to ensure the TTL check sees them as expired. - tokio::time::sleep(tokio::time::Duration::from_millis(10)).await; - - let evicted = cache.evict_expired().await; - assert_eq!(evicted, 2, "both entries should be evicted"); - - let stats = cache.stats_async().await; - assert_eq!(stats.total_entries, 0); - } - - // -- Test 9: Evict LRU when over limit -- - - #[tokio::test] - async fn test_evict_lru_when_over_limit() { - let config = CacheConfig { - max_entries: 2, - ..CacheConfig::default() - }; - let cache = PlanCache::new(config); - let plan = sample_logical_plan(); - - // Prepare two entries (at capacity). - let id1 = cache.prepare("SEARCH graph WHERE type = $t", plan.clone()).await; - // Brief pause so id1 has older last_used than id2. - tokio::time::sleep(tokio::time::Duration::from_millis(5)).await; - let _id2 = cache.prepare("SEARCH vector WHERE k = 5", plan.clone()).await; - - // Preparing a third should evict the LRU (id1). - let _id3 = cache.prepare("SEARCH document WHERE text = $q", plan).await; - - let stats = cache.stats_async().await; - assert_eq!(stats.total_entries, 2, "cache should stay at max_entries"); - - assert!( - cache.get(&id1).await.is_none(), - "oldest entry (id1) should have been evicted" - ); - } - - // -- Test 10: Parameter mismatch error -- - - #[tokio::test] - async fn test_parameter_mismatch_error() { - let cache = PlanCache::new(CacheConfig::default()); - let plan = sample_logical_plan(); - - let id = cache.prepare("SEARCH graph WHERE type = $t AND name = $n", plan).await; - - // Provide wrong parameter names. - let mut params = HashMap::new(); - params.insert("$wrong".to_string(), ParamValue::String("test".to_string())); - - let result = cache.execute_prepared(&id, ¶ms).await; - assert!(result.is_err(), "mismatched params should return error"); - - match result.unwrap_err() { - CacheError::ParameterMismatch { expected, provided } => { - assert!(expected.contains(&"$t".to_string())); - assert!(expected.contains(&"$n".to_string())); - assert!(provided.contains(&"$wrong".to_string())); - } - other => panic!("expected ParameterMismatch, got: {:?}", other), - } - } - - // -- Test 11: Cache stats tracking -- - - #[tokio::test] - async fn test_cache_stats_tracking() { - let cache = PlanCache::new(CacheConfig::default()); - let plan = sample_logical_plan(); - - // Initial stats: all zeros. - let stats = cache.stats_async().await; - assert_eq!(stats.total_entries, 0); - assert_eq!(stats.hit_count, 0); - assert_eq!(stats.miss_count, 0); - assert_eq!(stats.hit_ratio, 0.0); - - // Prepare a statement. - let id = cache.prepare("SEARCH graph WHERE type = $t", plan).await; - - // Cache hit. - let _ = cache.get(&id).await; - let stats = cache.stats_async().await; - assert_eq!(stats.hit_count, 1); - assert_eq!(stats.total_entries, 1); - - // Cache miss. - let _ = cache.get(&PreparedId::new("nonexistent")).await; - let stats = cache.stats_async().await; - assert_eq!(stats.miss_count, 1); - - // Hit ratio should be 0.5 (1 hit, 1 miss). - assert!((stats.hit_ratio - 0.5).abs() < f64::EPSILON); - - // Generation should have incremented from prepare + any mutations. - assert!(stats.generation > 0, "generation should be non-zero after mutations"); - } - - // -- Test 12: Plan caching on prepared statement -- - - #[tokio::test] - async fn test_plan_caching() { - let cache = PlanCache::new(CacheConfig::default()); - let logical = sample_logical_plan(); - let physical = sample_physical_plan(); - - let id = cache.prepare("SEARCH graph WHERE type = $t", logical).await; - - // Initially no cached physical plan. - let stmt = cache.get(&id).await.expect("TODO: handle error"); - assert!(stmt.cached_physical_plan.is_none()); - - // Cache the physical plan. - cache.cache_plan(&id, physical).await; - - // Now it should be present. - let stmt = cache.get(&id).await.expect("TODO: handle error"); - assert!( - stmt.cached_physical_plan.is_some(), - "physical plan should be cached" - ); - let plan = stmt.cached_physical_plan.expect("TODO: handle error"); - assert_eq!(plan.steps.len(), 1); - assert_eq!(plan.steps[0].modality, Modality::Graph); - } - - // -- Test 13: JSON serialization round-trip -- - - #[tokio::test] - async fn test_json_serialization_roundtrip() { - let cache = PlanCache::new(CacheConfig::default()); - let logical = sample_logical_plan(); - let physical = sample_physical_plan(); - - let id = cache.prepare("SEARCH graph WHERE type = $t", logical).await; - cache.cache_plan(&id, physical).await; - - let stmt = cache.get(&id).await.expect("TODO: handle error"); - - // Serialize to JSON. - let json = serde_json::to_string_pretty(&stmt).expect("TODO: handle error"); - assert!(!json.is_empty()); - - // Deserialize back. - let parsed: PreparedStatement = serde_json::from_str(&json).expect("TODO: handle error"); - assert_eq!(parsed.id, stmt.id); - assert_eq!(parsed.original_query, stmt.original_query); - assert_eq!(parsed.parameter_names, stmt.parameter_names); - assert_eq!(parsed.use_count, stmt.use_count); - assert!(parsed.cached_physical_plan.is_some()); - - // CacheConfig round-trip. - let config = CacheConfig::default(); - let config_json = serde_json::to_string(&config).expect("TODO: handle error"); - let parsed_config: CacheConfig = serde_json::from_str(&config_json).expect("TODO: handle error"); - assert_eq!(parsed_config.max_entries, config.max_entries); - assert_eq!(parsed_config.ttl_seconds, config.ttl_seconds); - - // CacheStats round-trip. - let stats = cache.stats_async().await; - let stats_json = serde_json::to_string(&stats).expect("TODO: handle error"); - let parsed_stats: CacheStats = serde_json::from_str(&stats_json).expect("TODO: handle error"); - assert_eq!(parsed_stats.total_entries, stats.total_entries); - - // CacheError round-trip. - let err = CacheError::ParameterMismatch { - expected: vec!["$a".to_string()], - provided: vec!["$b".to_string()], - }; - let err_json = serde_json::to_string(&err).expect("TODO: handle error"); - let parsed_err: CacheError = serde_json::from_str(&err_json).expect("TODO: handle error"); - match parsed_err { - CacheError::ParameterMismatch { expected, provided } => { - assert_eq!(expected, vec!["$a"]); - assert_eq!(provided, vec!["$b"]); - } - _ => panic!("wrong error variant after deserialization"), - } - - // ParamValue round-trip. - let params = vec![ - ParamValue::String("hello".to_string()), - ParamValue::Int(42), - ParamValue::Float(3.14), - ParamValue::Bool(true), - ParamValue::Vector(vec![0.1, 0.2, 0.3]), - ParamValue::Null, - ]; - for pv in ¶ms { - let pv_json = serde_json::to_string(pv).expect("TODO: handle error"); - let parsed_pv: ParamValue = serde_json::from_str(&pv_json).expect("TODO: handle error"); - // Verify the variant matches (structural equality check). - let re_json = serde_json::to_string(&parsed_pv).expect("TODO: handle error"); - assert_eq!(pv_json, re_json); - } - } - - // -- Test 14: Duplicate prepare returns same ID -- - - #[tokio::test] - async fn test_duplicate_prepare_returns_same_id() { - let cache = PlanCache::new(CacheConfig::default()); - let plan = sample_logical_plan(); - - let id1 = cache.prepare("SEARCH graph WHERE type = $t", plan.clone()).await; - let id2 = cache.prepare("SEARCH graph WHERE type = $t", plan).await; - - assert_eq!(id1, id2, "same query should produce same PreparedId"); - - let stats = cache.stats_async().await; - assert_eq!(stats.total_entries, 1, "only one entry should exist"); - } - - // -- Test 15: Parameter extraction -- - - #[test] - fn test_parameter_extraction() { - let params = PlanCache::extract_parameters("SEARCH graph WHERE type = $t AND name = $name"); - assert_eq!(params, vec!["$t", "$name"]); - - // Duplicate parameters should be deduplicated. - let params = PlanCache::extract_parameters("WHERE $x > 1 AND $x < 10"); - assert_eq!(params, vec!["$x"]); - - // No parameters. - let params = PlanCache::extract_parameters("SEARCH graph"); - assert!(params.is_empty()); - - // Dollar sign at end of string. - let params = PlanCache::extract_parameters("value = $"); - assert!(params.is_empty()); - } - - // -- Test 16: PreparedId display -- - - #[test] - fn test_prepared_id_display() { - let id = PreparedId::new("abcdef1234567890"); - let display = format!("{}", id); - assert_eq!(display, "prep_abcdef123456"); - - // Short ID (less than 12 chars). - let id_short = PreparedId::new("abc"); - let display_short = format!("{}", id_short); - assert_eq!(display_short, "prep_abc"); - } -} diff --git a/verisimdb/rust-core/verisim-planner/src/profiler.rs b/verisimdb/rust-core/verisim-planner/src/profiler.rs deleted file mode 100644 index 36026665..00000000 --- a/verisimdb/rust-core/verisim-planner/src/profiler.rs +++ /dev/null @@ -1,870 +0,0 @@ -// SPDX-License-Identifier: MPL-2.0 -// Copyright (c) 2026 Jonathan D.A. Jewell (hyperpolymath) <j.d.a.jewell@open.ac.uk> -//! EXPLAIN ANALYZE query profiling. -//! -//! Extends the static EXPLAIN output with actual execution metrics collected -//! at runtime. A [`Profiler`] wraps a [`PhysicalPlan`] and records wall-clock -//! timings, actual row counts, and estimation accuracy for every step. Results -//! are fed back into the [`StatisticsCollector`] via -//! [`record_execution`](StatisticsCollector::record_execution) so the -//! [`AdaptiveTuner`] can refine future cost estimates. - -use std::fmt; - -use chrono::{DateTime, Utc}; -use serde::{Deserialize, Serialize}; - -use crate::explain::{ExplainOutput, PerformanceHint}; -use crate::plan::PhysicalPlan; -use crate::stats::StatisticsCollector; -use crate::Modality; - -// --------------------------------------------------------------------------- -// ProfileStep — per-step actual metrics -// --------------------------------------------------------------------------- - -/// Actual execution metrics for a single plan step. -/// -/// Pairs the planner's estimates with observed wall-clock timings and row -/// counts so callers can evaluate estimation accuracy. -#[derive(Debug, Clone, Serialize, Deserialize)] -pub struct ProfileStep { - /// Human-readable name for this step (mirrors `PlanStep::operation`). - pub step_name: String, - /// Modality targeted by this step. - pub modality: Modality, - /// Cost the planner *estimated* for this step (milliseconds). - pub estimated_ms: f64, - /// Actual wall-clock duration observed (milliseconds). - pub actual_ms: f64, - /// Rows the planner *estimated* this step would return. - pub estimated_rows: u64, - /// Actual rows returned by the step. - pub actual_rows: u64, - /// Timestamp when execution of this step began. - pub started_at: DateTime<Utc>, - /// Timestamp when execution of this step completed. - pub ended_at: DateTime<Utc>, -} - -impl ProfileStep { - /// Ratio of actual to estimated duration. - /// - /// - `1.0` means the estimate was perfect. - /// - `> 1.0` means the step was *slower* than estimated. - /// - `< 1.0` means the step was *faster* than estimated. - /// - /// Returns `f64::INFINITY` when `estimated_ms` is zero. - pub fn time_accuracy_ratio(&self) -> f64 { - if self.estimated_ms <= 0.0 { - return f64::INFINITY; - } - self.actual_ms / self.estimated_ms - } - - /// Ratio of actual to estimated row count. - /// - /// Same semantics as [`time_accuracy_ratio`](Self::time_accuracy_ratio). - pub fn row_accuracy_ratio(&self) -> f64 { - if self.estimated_rows == 0 { - return f64::INFINITY; - } - self.actual_rows as f64 / self.estimated_rows as f64 - } -} - -// --------------------------------------------------------------------------- -// QueryProfile — aggregated profile for an entire query -// --------------------------------------------------------------------------- - -/// Aggregated profiling results for a fully-executed query plan. -/// -/// Contains per-step breakdowns, totals, and automatically-generated -/// optimization hints when estimation error exceeds useful thresholds. -#[derive(Debug, Clone, Serialize, Deserialize)] -pub struct QueryProfile { - /// Identifier for the physical plan that was profiled. - pub plan_id: String, - /// Per-step profiling data, in execution order. - pub steps: Vec<ProfileStep>, - /// Sum of all estimated step durations (milliseconds). - pub total_estimated_ms: f64, - /// Sum of all actual step durations (milliseconds). - pub total_actual_ms: f64, - /// Optimization hints derived from accuracy analysis. - pub optimization_hints: Vec<String>, -} - -impl QueryProfile { - /// Overall time accuracy ratio (actual / estimated). - pub fn total_time_accuracy_ratio(&self) -> f64 { - if self.total_estimated_ms <= 0.0 { - return f64::INFINITY; - } - self.total_actual_ms / self.total_estimated_ms - } - - /// Render the profile as a human-readable EXPLAIN ANALYZE text block. - pub fn render_text(&self, explain: &ExplainOutput) -> String { - let mut out = String::new(); - - out.push_str("=== VeriSimDB EXPLAIN ANALYZE ===\n\n"); - out.push_str(&format!("Plan ID: {}\n", self.plan_id)); - out.push_str(&format!("Strategy: {}\n", explain.strategy)); - out.push_str(&format!( - "Total Estimated: {:.1}ms | Total Actual: {:.1}ms | Accuracy: {:.2}x\n\n", - self.total_estimated_ms, - self.total_actual_ms, - self.total_time_accuracy_ratio(), - )); - - out.push_str("--- Steps ---\n"); - for (i, step) in self.steps.iter().enumerate() { - out.push_str(&format!( - " Step {}: {} [{}]\n", - i + 1, - step.step_name, - step.modality - )); - out.push_str(&format!( - " Estimated: {:.1}ms / ~{} rows\n", - step.estimated_ms, step.estimated_rows - )); - out.push_str(&format!( - " Actual: {:.1}ms / {} rows\n", - step.actual_ms, step.actual_rows - )); - out.push_str(&format!( - " Time accuracy: {:.2}x | Row accuracy: {:.2}x\n", - step.time_accuracy_ratio(), - step.row_accuracy_ratio() - )); - } - - if !self.optimization_hints.is_empty() { - out.push_str("\n--- Optimization Hints ---\n"); - for hint in &self.optimization_hints { - out.push_str(&format!(" * {}\n", hint)); - } - } - - out - } -} - -impl fmt::Display for QueryProfile { - fn fmt(&self, f: &mut fmt::Formatter<'_>) -> fmt::Result { - write!(f, "QueryProfile(plan={}, steps={}, estimated={:.1}ms, actual={:.1}ms)", - self.plan_id, self.steps.len(), self.total_estimated_ms, self.total_actual_ms) - } -} - -// --------------------------------------------------------------------------- -// Profiler — wraps a PhysicalPlan and collects execution metrics -// --------------------------------------------------------------------------- - -/// Threshold above which a time-accuracy ratio triggers a "slower than -/// estimated" hint. -const SLOW_THRESHOLD: f64 = 2.0; - -/// Threshold below which a time-accuracy ratio triggers a "faster than -/// estimated" hint. -const FAST_THRESHOLD: f64 = 0.5; - -/// Threshold above which a row-accuracy ratio triggers a "more rows than -/// estimated" hint. -const ROW_OVER_THRESHOLD: f64 = 3.0; - -/// Threshold below which a row-accuracy ratio triggers a "fewer rows than -/// estimated" hint. -const ROW_UNDER_THRESHOLD: f64 = 0.33; - -/// Records actual execution metrics against a [`PhysicalPlan`] and produces -/// a [`QueryProfile`]. -/// -/// # Usage -/// -/// ```ignore -/// let profiler = Profiler::new("query-42", &physical_plan); -/// -/// // For each step, record actual timings: -/// profiler.record_step(0, 55.3, 12, started, ended); -/// profiler.record_step(1, 210.0, 185, started, ended); -/// -/// let profile = profiler.finish(&mut stats_collector); -/// println!("{}", profile.render_text(&explain_output)); -/// ``` -pub struct Profiler { - /// Identifier for the plan being profiled. - plan_id: String, - /// The physical plan whose execution is being measured. - plan: PhysicalPlan, - /// Collected step profiles (filled incrementally via `record_step`). - recorded_steps: Vec<Option<ProfileStep>>, -} - -impl Profiler { - /// Create a new profiler for the given physical plan. - /// - /// `plan_id` is an opaque caller-chosen identifier (e.g. a query hash - /// or UUID) embedded in the resulting [`QueryProfile`]. - pub fn new(plan_id: impl Into<String>, plan: &PhysicalPlan) -> Self { - let step_count = plan.steps.len(); - Self { - plan_id: plan_id.into(), - plan: plan.clone(), - recorded_steps: vec![None; step_count], - } - } - - /// Record actual execution metrics for step `step_index` (0-based). - /// - /// # Panics - /// - /// Panics if `step_index` is out of range for the underlying plan. - pub fn record_step( - &mut self, - step_index: usize, - actual_ms: f64, - actual_rows: u64, - started_at: DateTime<Utc>, - ended_at: DateTime<Utc>, - ) { - assert!( - step_index < self.plan.steps.len(), - "step_index {} out of range (plan has {} steps)", - step_index, - self.plan.steps.len() - ); - - let plan_step = &self.plan.steps[step_index]; - self.recorded_steps[step_index] = Some(ProfileStep { - step_name: plan_step.operation.clone(), - modality: plan_step.modality, - estimated_ms: plan_step.cost.time_ms, - actual_ms, - estimated_rows: plan_step.cost.estimated_rows, - actual_rows, - started_at, - ended_at, - }); - } - - /// Consume the profiler and produce a [`QueryProfile`]. - /// - /// Any steps that were *not* recorded via [`record_step`](Self::record_step) - /// are filled with zero actual values (and will generate accuracy hints). - /// - /// This method also feeds each step's actual latency and row count into - /// the provided [`StatisticsCollector`] so the [`AdaptiveTuner`] can - /// refine future estimates. - pub fn finish(self, stats: &mut StatisticsCollector) -> QueryProfile { - let mut steps: Vec<ProfileStep> = Vec::with_capacity(self.plan.steps.len()); - - for (i, recorded) in self.recorded_steps.into_iter().enumerate() { - let profile_step = match recorded { - Some(s) => s, - None => { - // Step was never recorded — fill with zero actuals. - let plan_step = &self.plan.steps[i]; - let now = Utc::now(); - ProfileStep { - step_name: plan_step.operation.clone(), - modality: plan_step.modality, - estimated_ms: plan_step.cost.time_ms, - actual_ms: 0.0, - estimated_rows: plan_step.cost.estimated_rows, - actual_rows: 0, - started_at: now, - ended_at: now, - } - } - }; - - // Feed actuals into the statistics collector for adaptive tuning. - stats.record_execution( - profile_step.modality, - profile_step.actual_ms, - profile_step.actual_rows, - ); - - steps.push(profile_step); - } - - let total_estimated_ms: f64 = steps.iter().map(|s| s.estimated_ms).sum(); - let total_actual_ms: f64 = steps.iter().map(|s| s.actual_ms).sum(); - let optimization_hints = generate_hints(&steps, total_estimated_ms, total_actual_ms); - - QueryProfile { - plan_id: self.plan_id, - steps, - total_estimated_ms, - total_actual_ms, - optimization_hints, - } - } -} - -// --------------------------------------------------------------------------- -// ExplainOutput extension — ANALYZE integration -// --------------------------------------------------------------------------- - -impl ExplainOutput { - /// Merge profiling results into this EXPLAIN output to produce an - /// EXPLAIN ANALYZE rendering. - /// - /// Returns a combined output that contains both the original plan details - /// and the actual execution metrics. - pub fn with_profile(&self, profile: &QueryProfile) -> ExplainAnalyzeOutput { - let mut hints: Vec<PerformanceHint> = self.performance_hints.clone(); - - // Append profiler-generated hints as PerformanceHint structs. - for hint_text in &profile.optimization_hints { - hints.push(PerformanceHint { - severity: "analyze".to_string(), - message: hint_text.clone(), - }); - } - - ExplainAnalyzeOutput { - explain: self.clone(), - profile: profile.clone(), - combined_hints: hints, - text_output: profile.render_text(self), - } - } -} - -/// Combined EXPLAIN ANALYZE output containing both plan estimates and actual -/// execution metrics. -#[derive(Debug, Clone, Serialize, Deserialize)] -pub struct ExplainAnalyzeOutput { - /// Original EXPLAIN output (estimates only). - pub explain: ExplainOutput, - /// Actual execution profile. - pub profile: QueryProfile, - /// Merged performance + profiling hints. - pub combined_hints: Vec<PerformanceHint>, - /// Human-readable EXPLAIN ANALYZE text. - pub text_output: String, -} - -impl fmt::Display for ExplainAnalyzeOutput { - fn fmt(&self, f: &mut fmt::Formatter<'_>) -> fmt::Result { - write!(f, "{}", self.text_output) - } -} - -// --------------------------------------------------------------------------- -// Hint generation -// --------------------------------------------------------------------------- - -/// Analyze per-step accuracy and produce actionable optimization hints. -fn generate_hints( - steps: &[ProfileStep], - total_estimated_ms: f64, - total_actual_ms: f64, -) -> Vec<String> { - let mut hints = Vec::new(); - - // Overall accuracy hint. - if total_estimated_ms > 0.0 { - let overall_ratio = total_actual_ms / total_estimated_ms; - if overall_ratio > SLOW_THRESHOLD { - hints.push(format!( - "Query was {:.1}x slower than estimated ({:.0}ms actual vs {:.0}ms estimated) \ - — planner may be underestimating costs", - overall_ratio, total_actual_ms, total_estimated_ms - )); - } else if overall_ratio < FAST_THRESHOLD { - hints.push(format!( - "Query was {:.1}x faster than estimated ({:.0}ms actual vs {:.0}ms estimated) \ - — planner may be overestimating costs", - overall_ratio, total_actual_ms, total_estimated_ms - )); - } - } - - // Per-step time accuracy hints. - for step in steps { - let time_ratio = step.time_accuracy_ratio(); - if time_ratio > SLOW_THRESHOLD && time_ratio.is_finite() { - hints.push(format!( - "Step '{}' [{}]: {:.1}x slower than estimated ({:.0}ms vs {:.0}ms) \ - — consider updating cost model for this modality", - step.step_name, step.modality, time_ratio, step.actual_ms, step.estimated_ms - )); - } else if time_ratio < FAST_THRESHOLD { - hints.push(format!( - "Step '{}' [{}]: {:.1}x faster than estimated ({:.0}ms vs {:.0}ms) \ - — aggressive mode may be appropriate", - step.step_name, step.modality, time_ratio, step.actual_ms, step.estimated_ms - )); - } - } - - // Per-step row accuracy hints. - for step in steps { - let row_ratio = step.row_accuracy_ratio(); - if row_ratio > ROW_OVER_THRESHOLD && row_ratio.is_finite() { - hints.push(format!( - "Step '{}' [{}]: returned {:.1}x more rows than estimated ({} vs {}) \ - — selectivity estimate may be too low", - step.step_name, step.modality, row_ratio, step.actual_rows, step.estimated_rows - )); - } else if row_ratio < ROW_UNDER_THRESHOLD && step.estimated_rows > 0 { - hints.push(format!( - "Step '{}' [{}]: returned {:.1}x fewer rows than estimated ({} vs {}) \ - — selectivity estimate may be too high", - step.step_name, step.modality, row_ratio, step.actual_rows, step.estimated_rows - )); - } - } - - hints -} - -// =========================================================================== -// Tests -// =========================================================================== - -#[cfg(test)] -mod tests { - use super::*; - use crate::config::PlannerConfig; - use crate::cost::CostEstimate; - use crate::plan::{ExecutionStrategy, PhysicalPlan, PlanStep}; - use chrono::Duration; - - // ----------------------------------------------------------------------- - // Helpers - // ----------------------------------------------------------------------- - - /// Build a simple two-step physical plan for testing. - fn two_step_plan() -> PhysicalPlan { - PhysicalPlan { - steps: vec![ - PlanStep { - step: 1, - operation: "Vector similarity search (1 conditions)".to_string(), - modality: Modality::Vector, - cost: CostEstimate { - time_ms: 40.0, - estimated_rows: 10, - selectivity: 0.005, - io_cost: 24.0, - cpu_cost: 16.0, - }, - optimization_hint: Some("HNSW ANN search (k=10)".to_string()), - pushed_predicates: vec!["Similarity { k: 10 }".to_string()], - }, - PlanStep { - step: 2, - operation: "Graph traversal (1 conditions)".to_string(), - modality: Modality::Graph, - cost: CostEstimate { - time_ms: 225.0, - estimated_rows: 200, - selectivity: 0.4, - io_cost: 135.0, - cpu_cost: 90.0, - }, - optimization_hint: Some("Graph traversal: relates_to (depth=2)".to_string()), - pushed_predicates: vec!["Traversal { predicate: relates_to }".to_string()], - }, - ], - strategy: ExecutionStrategy::Parallel, - total_cost: CostEstimate { - time_ms: 225.0, - estimated_rows: 200, - selectivity: 0.002, - io_cost: 159.0, - cpu_cost: 106.0, - }, - notes: vec!["Parallel execution across 2 modalities".to_string()], - } - } - - fn make_timestamps(base: DateTime<Utc>, duration_ms: f64) -> (DateTime<Utc>, DateTime<Utc>) { - let end = base + Duration::milliseconds(duration_ms as i64); - (base, end) - } - - // ----------------------------------------------------------------------- - // Test 1: Profile with known durations - // ----------------------------------------------------------------------- - - #[test] - fn test_profile_with_known_durations() { - let plan = two_step_plan(); - let mut profiler = Profiler::new("test-query-1", &plan); - let mut stats = StatisticsCollector::new(); - - let base = Utc::now(); - let (s0, e0) = make_timestamps(base, 35.0); - let (s1, e1) = make_timestamps(base, 250.0); - - profiler.record_step(0, 35.0, 8, s0, e0); - profiler.record_step(1, 250.0, 180, s1, e1); - - let profile = profiler.finish(&mut stats); - - assert_eq!(profile.plan_id, "test-query-1"); - assert_eq!(profile.steps.len(), 2); - assert!((profile.steps[0].actual_ms - 35.0).abs() < f64::EPSILON); - assert!((profile.steps[1].actual_ms - 250.0).abs() < f64::EPSILON); - assert_eq!(profile.steps[0].actual_rows, 8); - assert_eq!(profile.steps[1].actual_rows, 180); - } - - // ----------------------------------------------------------------------- - // Test 2: Accuracy ratio calculation - // ----------------------------------------------------------------------- - - #[test] - fn test_accuracy_ratio_calculation() { - let plan = two_step_plan(); - let mut profiler = Profiler::new("ratio-test", &plan); - let mut stats = StatisticsCollector::new(); - - let base = Utc::now(); - - // Step 0: estimated 40ms, actual 35ms → ratio 0.875 - let (s0, e0) = make_timestamps(base, 35.0); - profiler.record_step(0, 35.0, 10, s0, e0); - - // Step 1: estimated 225ms, actual 450ms → ratio 2.0 - let (s1, e1) = make_timestamps(base, 450.0); - profiler.record_step(1, 450.0, 200, s1, e1); - - let profile = profiler.finish(&mut stats); - - // Step 0 time accuracy: 35 / 40 = 0.875 - let ratio_0 = profile.steps[0].time_accuracy_ratio(); - assert!((ratio_0 - 0.875).abs() < 0.001, "Expected 0.875, got {}", ratio_0); - - // Step 1 time accuracy: 450 / 225 = 2.0 - let ratio_1 = profile.steps[1].time_accuracy_ratio(); - assert!((ratio_1 - 2.0).abs() < 0.001, "Expected 2.0, got {}", ratio_1); - - // Row accuracy for step 0: 10 / 10 = 1.0 (perfect) - let row_ratio_0 = profile.steps[0].row_accuracy_ratio(); - assert!((row_ratio_0 - 1.0).abs() < 0.001); - - // Row accuracy for step 1: 200 / 200 = 1.0 (perfect) - let row_ratio_1 = profile.steps[1].row_accuracy_ratio(); - assert!((row_ratio_1 - 1.0).abs() < 0.001); - } - - // ----------------------------------------------------------------------- - // Test 3: Hints generated for large estimation errors (time) - // ----------------------------------------------------------------------- - - #[test] - fn test_hints_for_large_time_errors() { - let plan = two_step_plan(); - let mut profiler = Profiler::new("slow-query", &plan); - let mut stats = StatisticsCollector::new(); - - let base = Utc::now(); - - // Step 0: estimated 40ms, actual 200ms → 5x slower → should generate hint - let (s0, e0) = make_timestamps(base, 200.0); - profiler.record_step(0, 200.0, 10, s0, e0); - - // Step 1: estimated 225ms, actual 900ms → 4x slower → should generate hint - let (s1, e1) = make_timestamps(base, 900.0); - profiler.record_step(1, 900.0, 200, s1, e1); - - let profile = profiler.finish(&mut stats); - - // Should have at least the overall "slower than estimated" hint - assert!( - !profile.optimization_hints.is_empty(), - "Expected hints for large estimation errors" - ); - - let has_slow_hint = profile.optimization_hints.iter().any(|h| h.contains("slower")); - assert!( - has_slow_hint, - "Expected a 'slower than estimated' hint, got: {:?}", - profile.optimization_hints - ); - } - - // ----------------------------------------------------------------------- - // Test 4: Hints generated for large estimation errors (rows) - // ----------------------------------------------------------------------- - - #[test] - fn test_hints_for_large_row_errors() { - let plan = two_step_plan(); - let mut profiler = Profiler::new("row-mismatch", &plan); - let mut stats = StatisticsCollector::new(); - - let base = Utc::now(); - - // Step 0: estimated 10 rows, actual 50 rows → 5x more → hint - let (s0, e0) = make_timestamps(base, 40.0); - profiler.record_step(0, 40.0, 50, s0, e0); - - // Step 1: estimated 200 rows, actual 10 rows → 0.05x fewer → hint - let (s1, e1) = make_timestamps(base, 225.0); - profiler.record_step(1, 225.0, 10, s1, e1); - - let profile = profiler.finish(&mut stats); - - let has_row_over = profile.optimization_hints.iter().any(|h| h.contains("more rows")); - assert!( - has_row_over, - "Expected 'more rows than estimated' hint, got: {:?}", - profile.optimization_hints - ); - - let has_row_under = profile.optimization_hints.iter().any(|h| h.contains("fewer rows")); - assert!( - has_row_under, - "Expected 'fewer rows than estimated' hint, got: {:?}", - profile.optimization_hints - ); - } - - // ----------------------------------------------------------------------- - // Test 5: Total time calculation - // ----------------------------------------------------------------------- - - #[test] - fn test_total_time_calculation() { - let plan = two_step_plan(); - let mut profiler = Profiler::new("total-time", &plan); - let mut stats = StatisticsCollector::new(); - - let base = Utc::now(); - let (s0, e0) = make_timestamps(base, 50.0); - let (s1, e1) = make_timestamps(base, 300.0); - - profiler.record_step(0, 50.0, 10, s0, e0); - profiler.record_step(1, 300.0, 200, s1, e1); - - let profile = profiler.finish(&mut stats); - - // Total estimated = 40 + 225 = 265 - assert!( - (profile.total_estimated_ms - 265.0).abs() < f64::EPSILON, - "Expected total_estimated_ms=265.0, got {}", - profile.total_estimated_ms - ); - - // Total actual = 50 + 300 = 350 - assert!( - (profile.total_actual_ms - 350.0).abs() < f64::EPSILON, - "Expected total_actual_ms=350.0, got {}", - profile.total_actual_ms - ); - - // Overall ratio = 350 / 265 ≈ 1.3208 - let overall = profile.total_time_accuracy_ratio(); - assert!( - (overall - 350.0 / 265.0).abs() < 0.001, - "Expected ratio ~1.321, got {}", - overall - ); - } - - // ----------------------------------------------------------------------- - // Test 6: Auto-feed results into StatisticsCollector - // ----------------------------------------------------------------------- - - #[test] - fn test_auto_feed_statistics() { - let plan = two_step_plan(); - let mut profiler = Profiler::new("feed-test", &plan); - let mut stats = StatisticsCollector::new(); - - // Verify initial state: zero queries for Vector and Graph - assert_eq!(stats.get(Modality::Vector).expect("TODO: handle error").query_count, 0); - assert_eq!(stats.get(Modality::Graph).expect("TODO: handle error").query_count, 0); - - let base = Utc::now(); - let (s0, e0) = make_timestamps(base, 42.0); - let (s1, e1) = make_timestamps(base, 180.0); - - profiler.record_step(0, 42.0, 12, s0, e0); - profiler.record_step(1, 180.0, 190, s1, e1); - - let _profile = profiler.finish(&mut stats); - - // After finish(), stats should have recorded one execution per modality - let vector_stats = stats.get(Modality::Vector).expect("TODO: handle error"); - assert_eq!(vector_stats.query_count, 1); - assert!((vector_stats.avg_latency_ms - 42.0).abs() < f64::EPSILON); - assert_eq!(vector_stats.avg_rows_returned, 12); - - let graph_stats = stats.get(Modality::Graph).expect("TODO: handle error"); - assert_eq!(graph_stats.query_count, 1); - assert!((graph_stats.avg_latency_ms - 180.0).abs() < f64::EPSILON); - assert_eq!(graph_stats.avg_rows_returned, 190); - } - - // ----------------------------------------------------------------------- - // Test 7: Integration with ExplainOutput (ANALYZE rendering) - // ----------------------------------------------------------------------- - - #[test] - fn test_explain_analyze_integration() { - let plan = two_step_plan(); - let config = PlannerConfig::default(); - let explain = ExplainOutput::from_physical_plan(&plan, &config); - - let mut profiler = Profiler::new("analyze-test", &plan); - let mut stats = StatisticsCollector::new(); - - let base = Utc::now(); - let (s0, e0) = make_timestamps(base, 38.0); - let (s1, e1) = make_timestamps(base, 220.0); - - profiler.record_step(0, 38.0, 9, s0, e0); - profiler.record_step(1, 220.0, 195, s1, e1); - - let profile = profiler.finish(&mut stats); - let analyze = explain.with_profile(&profile); - - // Text output should contain EXPLAIN ANALYZE header - assert!(analyze.text_output.contains("EXPLAIN ANALYZE")); - // Should contain actual metrics - assert!(analyze.text_output.contains("Actual:")); - // Should contain estimated metrics - assert!(analyze.text_output.contains("Estimated:")); - // Should reference both modalities - assert!(analyze.text_output.contains("vector")); - assert!(analyze.text_output.contains("graph")); - // Display trait should work - let display = format!("{}", analyze); - assert!(!display.is_empty()); - } - - // ----------------------------------------------------------------------- - // Test 8: Unrecorded steps produce zero actuals - // ----------------------------------------------------------------------- - - #[test] - fn test_unrecorded_steps_produce_zero_actuals() { - let plan = two_step_plan(); - let mut profiler = Profiler::new("partial-record", &plan); - let mut stats = StatisticsCollector::new(); - - // Only record step 0, leave step 1 unrecorded - let base = Utc::now(); - let (s0, e0) = make_timestamps(base, 50.0); - profiler.record_step(0, 50.0, 8, s0, e0); - - let profile = profiler.finish(&mut stats); - - assert_eq!(profile.steps.len(), 2); - // Step 1 should have zero actual values - assert!((profile.steps[1].actual_ms - 0.0).abs() < f64::EPSILON); - assert_eq!(profile.steps[1].actual_rows, 0); - } - - // ----------------------------------------------------------------------- - // Test 9: Fast query generates "overestimating" hints - // ----------------------------------------------------------------------- - - #[test] - fn test_fast_query_overestimating_hints() { - let plan = two_step_plan(); - let mut profiler = Profiler::new("fast-query", &plan); - let mut stats = StatisticsCollector::new(); - - let base = Utc::now(); - // Both steps are much faster than estimated - let (s0, e0) = make_timestamps(base, 5.0); // estimated 40ms - let (s1, e1) = make_timestamps(base, 20.0); // estimated 225ms - - profiler.record_step(0, 5.0, 10, s0, e0); - profiler.record_step(1, 20.0, 200, s1, e1); - - let profile = profiler.finish(&mut stats); - - let has_fast_hint = profile.optimization_hints.iter().any(|h| h.contains("faster")); - assert!( - has_fast_hint, - "Expected 'faster than estimated' hint, got: {:?}", - profile.optimization_hints - ); - } - - // ----------------------------------------------------------------------- - // Test 10: QueryProfile display trait - // ----------------------------------------------------------------------- - - #[test] - fn test_query_profile_display() { - let plan = two_step_plan(); - let mut profiler = Profiler::new("display-test", &plan); - let mut stats = StatisticsCollector::new(); - - let base = Utc::now(); - let (s0, e0) = make_timestamps(base, 40.0); - let (s1, e1) = make_timestamps(base, 225.0); - - profiler.record_step(0, 40.0, 10, s0, e0); - profiler.record_step(1, 225.0, 200, s1, e1); - - let profile = profiler.finish(&mut stats); - let display = format!("{}", profile); - - assert!(display.contains("display-test")); - assert!(display.contains("steps=2")); - } - - // ----------------------------------------------------------------------- - // Test 11: JSON serialization round-trip - // ----------------------------------------------------------------------- - - #[test] - fn test_query_profile_json_roundtrip() { - let plan = two_step_plan(); - let mut profiler = Profiler::new("serde-test", &plan); - let mut stats = StatisticsCollector::new(); - - let base = Utc::now(); - let (s0, e0) = make_timestamps(base, 40.0); - let (s1, e1) = make_timestamps(base, 225.0); - - profiler.record_step(0, 40.0, 10, s0, e0); - profiler.record_step(1, 225.0, 200, s1, e1); - - let profile = profiler.finish(&mut stats); - - let json = serde_json::to_string(&profile).expect("TODO: handle error"); - let parsed: QueryProfile = serde_json::from_str(&json).expect("TODO: handle error"); - - assert_eq!(parsed.plan_id, "serde-test"); - assert_eq!(parsed.steps.len(), 2); - assert!((parsed.total_estimated_ms - profile.total_estimated_ms).abs() < f64::EPSILON); - assert!((parsed.total_actual_ms - profile.total_actual_ms).abs() < f64::EPSILON); - } - - // ----------------------------------------------------------------------- - // Test 12: Zero estimated_ms produces INFINITY ratio - // ----------------------------------------------------------------------- - - #[test] - fn test_zero_estimated_produces_infinity_ratio() { - let now = Utc::now(); - let step = ProfileStep { - step_name: "zero-est".to_string(), - modality: Modality::Temporal, - estimated_ms: 0.0, - actual_ms: 10.0, - estimated_rows: 0, - actual_rows: 5, - started_at: now, - ended_at: now, - }; - - assert!(step.time_accuracy_ratio().is_infinite()); - assert!(step.row_accuracy_ratio().is_infinite()); - } -} diff --git a/verisimdb/rust-core/verisim-planner/src/slow_query.rs b/verisimdb/rust-core/verisim-planner/src/slow_query.rs deleted file mode 100644 index c01890fc..00000000 --- a/verisimdb/rust-core/verisim-planner/src/slow_query.rs +++ /dev/null @@ -1,550 +0,0 @@ -// SPDX-License-Identifier: MPL-2.0 -// Copyright (c) 2026 Jonathan D.A. Jewell (hyperpolymath) <j.d.a.jewell@open.ac.uk> -//! -//! Slow query log for VeriSimDB. -//! -//! Records queries that exceed a configurable duration threshold. -//! Integrates with the `tracing` framework to emit structured log events -//! and maintains an in-memory ring buffer for recent slow queries. - -use chrono::{DateTime, Utc}; -use serde::{Deserialize, Serialize}; -use std::collections::VecDeque; -use std::sync::RwLock; -use tracing::warn; - -use crate::plan::{ExecutionStrategy, PhysicalPlan}; -use crate::Modality; - -/// Configuration for the slow query log. -#[derive(Debug, Clone, Serialize, Deserialize)] -pub struct SlowQueryConfig { - /// Queries taking longer than this (milliseconds) are logged. - /// Default: 100ms. - pub threshold_ms: f64, - - /// Maximum number of entries to keep in the ring buffer. - /// Default: 1000. - pub max_entries: usize, - - /// Whether slow query logging is enabled. - /// Default: true. - pub enabled: bool, - - /// Log queries that use more than this many modalities. - /// Set to 0 to disable this check. - /// Default: 0 (disabled). - pub multi_modality_threshold: usize, -} - -impl Default for SlowQueryConfig { - fn default() -> Self { - Self { - threshold_ms: 100.0, - max_entries: 1000, - enabled: true, - multi_modality_threshold: 0, - } - } -} - -/// A single slow query log entry. -#[derive(Debug, Clone, Serialize, Deserialize)] -pub struct SlowQueryEntry { - /// When the query was recorded. - pub timestamp: DateTime<Utc>, - - /// The VCL query text (if available). - pub query_text: Option<String>, - - /// Actual execution time in milliseconds. - pub actual_ms: f64, - - /// Estimated execution time from planner (milliseconds). - pub estimated_ms: f64, - - /// Ratio of actual to estimated (>1 means slower than expected). - pub slowdown_ratio: f64, - - /// Execution strategy used. - pub strategy: String, - - /// Modalities involved. - pub modalities: Vec<Modality>, - - /// Number of rows returned. - pub rows_returned: usize, - - /// Which step was the bottleneck. - pub bottleneck: Option<BottleneckInfo>, -} - -/// Information about the slowest step in a query. -#[derive(Debug, Clone, Serialize, Deserialize)] -pub struct BottleneckInfo { - /// Modality of the bottleneck step. - pub modality: Modality, - - /// Step name/operation. - pub operation: String, - - /// Time spent on this step (ms). - pub time_ms: f64, - - /// Percentage of total query time. - pub percentage: f64, -} - -/// Slow query log — ring buffer with tracing integration. -pub struct SlowQueryLog { - config: RwLock<SlowQueryConfig>, - entries: RwLock<VecDeque<SlowQueryEntry>>, -} - -impl SlowQueryLog { - /// Create a new slow query log with the given configuration. - pub fn new(config: SlowQueryConfig) -> Self { - Self { - config: RwLock::new(config), - entries: RwLock::new(VecDeque::new()), - } - } - - /// Create with default configuration. - pub fn with_defaults() -> Self { - Self::new(SlowQueryConfig::default()) - } - - /// Record a query execution. If it exceeds the threshold, it is logged. - /// - /// Returns `true` if the query was recorded as slow. - pub fn record( - &self, - query_text: Option<&str>, - actual_ms: f64, - plan: &PhysicalPlan, - step_times: &[(Modality, f64, usize)], // (modality, time_ms, rows) - ) -> bool { - let config = self.config.read().expect("TODO: handle error"); - if !config.enabled { - return false; - } - - let is_slow = actual_ms >= config.threshold_ms; - let is_multi = config.multi_modality_threshold > 0 - && plan.steps.len() >= config.multi_modality_threshold; - - if !is_slow && !is_multi { - return false; - } - - let estimated_ms = plan.total_cost.time_ms; - let slowdown_ratio = if estimated_ms > 0.0 { - actual_ms / estimated_ms - } else { - f64::INFINITY - }; - - let modalities: Vec<Modality> = plan.steps.iter().map(|s| s.modality).collect(); - - // Find bottleneck - let bottleneck = step_times - .iter() - .max_by(|a, b| a.1.partial_cmp(&b.1).unwrap_or(std::cmp::Ordering::Equal)) - .map(|(modality, time_ms, _rows)| { - let percentage = if actual_ms > 0.0 { - (time_ms / actual_ms) * 100.0 - } else { - 0.0 - }; - BottleneckInfo { - modality: *modality, - operation: format!("{} query", modality), - time_ms: *time_ms, - percentage, - } - }); - - let total_rows: usize = step_times.iter().map(|(_, _, r)| r).sum(); - let strategy = match plan.strategy { - ExecutionStrategy::Sequential => "sequential".to_string(), - ExecutionStrategy::Parallel => "parallel".to_string(), - }; - - let entry = SlowQueryEntry { - timestamp: Utc::now(), - query_text: query_text.map(|s| s.to_string()), - actual_ms, - estimated_ms, - slowdown_ratio, - strategy, - modalities: modalities.clone(), - rows_returned: total_rows, - bottleneck: bottleneck.clone(), - }; - - // Emit tracing warning - let modality_names: Vec<String> = modalities.iter().map(|m| m.to_string()).collect(); - let bottleneck_desc = bottleneck - .as_ref() - .map(|b| format!("{} ({:.0}ms, {:.0}%)", b.modality, b.time_ms, b.percentage)) - .unwrap_or_else(|| "unknown".to_string()); - - warn!( - actual_ms = actual_ms, - estimated_ms = estimated_ms, - slowdown_ratio = slowdown_ratio, - modalities = ?modality_names, - rows = total_rows, - bottleneck = %bottleneck_desc, - query = query_text.unwrap_or("<unknown>"), - "Slow query detected" - ); - - // Insert into ring buffer - let max_entries = config.max_entries; - drop(config); - - let mut entries = self.entries.write().expect("TODO: handle error"); - entries.push_back(entry); - while entries.len() > max_entries { - entries.pop_front(); - } - - true - } - - /// Get recent slow queries. - pub fn recent(&self, limit: usize) -> Vec<SlowQueryEntry> { - let entries = self.entries.read().expect("TODO: handle error"); - entries.iter().rev().take(limit).cloned().collect() - } - - /// Get all slow queries. - pub fn all(&self) -> Vec<SlowQueryEntry> { - self.entries.read().expect("TODO: handle error").iter().cloned().collect() - } - - /// Get the count of recorded slow queries. - pub fn count(&self) -> usize { - self.entries.read().expect("TODO: handle error").len() - } - - /// Clear the slow query log. - pub fn clear(&self) { - self.entries.write().expect("TODO: handle error").clear(); - } - - /// Update configuration. - pub fn set_config(&self, config: SlowQueryConfig) { - *self.config.write().expect("TODO: handle error") = config; - } - - /// Get current configuration. - pub fn config(&self) -> SlowQueryConfig { - self.config.read().expect("TODO: handle error").clone() - } - - /// Summary statistics. - pub fn summary(&self) -> SlowQuerySummary { - let entries = self.entries.read().expect("TODO: handle error"); - if entries.is_empty() { - return SlowQuerySummary::default(); - } - - let total = entries.len(); - let sum_ms: f64 = entries.iter().map(|e| e.actual_ms).sum(); - let max_ms = entries - .iter() - .map(|e| e.actual_ms) - .fold(0.0_f64, f64::max); - let min_ms = entries - .iter() - .map(|e| e.actual_ms) - .fold(f64::INFINITY, f64::min); - let avg_ms = sum_ms / total as f64; - let avg_ratio: f64 = entries.iter().map(|e| e.slowdown_ratio).sum::<f64>() / total as f64; - - // Most common bottleneck modality - let mut modality_counts = std::collections::HashMap::new(); - for entry in entries.iter() { - if let Some(ref b) = entry.bottleneck { - *modality_counts.entry(b.modality).or_insert(0u64) += 1; - } - } - let top_bottleneck = modality_counts - .into_iter() - .max_by_key(|&(_, count)| count) - .map(|(m, _)| m); - - SlowQuerySummary { - total_count: total, - avg_ms, - max_ms, - min_ms, - avg_slowdown_ratio: avg_ratio, - top_bottleneck_modality: top_bottleneck, - } - } -} - -/// Summary statistics for the slow query log. -#[derive(Debug, Clone, Default, Serialize, Deserialize)] -pub struct SlowQuerySummary { - pub total_count: usize, - pub avg_ms: f64, - pub max_ms: f64, - pub min_ms: f64, - pub avg_slowdown_ratio: f64, - pub top_bottleneck_modality: Option<Modality>, -} - -#[cfg(test)] -mod tests { - use super::*; - use crate::cost::CostEstimate; - use crate::plan::PlanStep; - - fn make_plan(steps: Vec<(Modality, f64)>) -> PhysicalPlan { - let total_ms: f64 = steps.iter().map(|(_, ms)| ms).sum(); - PhysicalPlan { - steps: steps - .iter() - .enumerate() - .map(|(i, (m, ms))| PlanStep { - step: i + 1, - operation: format!("{} query", m), - modality: *m, - cost: CostEstimate { - time_ms: *ms, - estimated_rows: 100, - selectivity: 0.5, - io_cost: ms * 0.6, - cpu_cost: ms * 0.4, - }, - optimization_hint: None, - pushed_predicates: vec![], - }) - .collect(), - strategy: if steps.len() >= 2 { - ExecutionStrategy::Parallel - } else { - ExecutionStrategy::Sequential - }, - total_cost: CostEstimate { - time_ms: total_ms, - estimated_rows: 100, - selectivity: 0.5, - io_cost: total_ms * 0.6, - cpu_cost: total_ms * 0.4, - }, - notes: vec![], - } - } - - #[test] - fn test_fast_query_not_logged() { - let log = SlowQueryLog::with_defaults(); - let plan = make_plan(vec![(Modality::Vector, 30.0)]); - let step_times = vec![(Modality::Vector, 30.0, 10)]; - - let was_slow = log.record(Some("SELECT VECTOR FROM OCTAD"), 30.0, &plan, &step_times); - assert!(!was_slow); - assert_eq!(log.count(), 0); - } - - #[test] - fn test_slow_query_logged() { - let log = SlowQueryLog::with_defaults(); - let plan = make_plan(vec![(Modality::Semantic, 50.0)]); - let step_times = vec![(Modality::Semantic, 150.0, 5)]; - - let was_slow = - log.record(Some("SELECT SEMANTIC FROM OCTAD"), 150.0, &plan, &step_times); - assert!(was_slow); - assert_eq!(log.count(), 1); - - let entries = log.recent(10); - assert_eq!(entries.len(), 1); - assert!(entries[0].actual_ms >= 100.0); - assert_eq!(entries[0].modalities, vec![Modality::Semantic]); - } - - #[test] - fn test_custom_threshold() { - let config = SlowQueryConfig { - threshold_ms: 50.0, - ..Default::default() - }; - let log = SlowQueryLog::new(config); - let plan = make_plan(vec![(Modality::Graph, 40.0)]); - let step_times = vec![(Modality::Graph, 60.0, 20)]; - - let was_slow = log.record(Some("SELECT GRAPH FROM OCTAD"), 60.0, &plan, &step_times); - assert!(was_slow); - assert_eq!(log.count(), 1); - } - - #[test] - fn test_disabled_log() { - let config = SlowQueryConfig { - enabled: false, - ..Default::default() - }; - let log = SlowQueryLog::new(config); - let plan = make_plan(vec![(Modality::Semantic, 50.0)]); - let step_times = vec![(Modality::Semantic, 500.0, 5)]; - - let was_slow = log.record(Some("SELECT SEMANTIC"), 500.0, &plan, &step_times); - assert!(!was_slow); - assert_eq!(log.count(), 0); - } - - #[test] - fn test_ring_buffer_eviction() { - let config = SlowQueryConfig { - threshold_ms: 10.0, - max_entries: 3, - ..Default::default() - }; - let log = SlowQueryLog::new(config); - let plan = make_plan(vec![(Modality::Vector, 5.0)]); - - for i in 0..5 { - let step_times = vec![(Modality::Vector, 20.0 + i as f64, 1)]; - log.record(Some(&format!("query-{i}")), 20.0 + i as f64, &plan, &step_times); - } - - assert_eq!(log.count(), 3); - // Most recent entries should remain - let entries = log.all(); - assert!(entries[0].actual_ms >= 22.0); - } - - #[test] - fn test_bottleneck_detection() { - let log = SlowQueryLog::with_defaults(); - let plan = make_plan(vec![ - (Modality::Vector, 30.0), - (Modality::Semantic, 200.0), - ]); - let step_times = vec![ - (Modality::Vector, 25.0, 10), - (Modality::Semantic, 180.0, 5), - ]; - - log.record(Some("multi-modality query"), 205.0, &plan, &step_times); - - let entries = log.recent(1); - assert_eq!(entries.len(), 1); - let bottleneck = entries[0].bottleneck.as_ref().expect("TODO: handle error"); - assert_eq!(bottleneck.modality, Modality::Semantic); - assert!(bottleneck.percentage > 80.0); - } - - #[test] - fn test_slowdown_ratio() { - let log = SlowQueryLog::with_defaults(); - let plan = make_plan(vec![(Modality::Graph, 50.0)]); - let step_times = vec![(Modality::Graph, 200.0, 100)]; - - log.record(Some("slow graph query"), 200.0, &plan, &step_times); - - let entries = log.recent(1); - // estimated 50ms, actual 200ms → ratio 4.0 - assert!((entries[0].slowdown_ratio - 4.0).abs() < 0.01); - } - - #[test] - fn test_clear() { - let log = SlowQueryLog::with_defaults(); - let plan = make_plan(vec![(Modality::Tensor, 50.0)]); - let step_times = vec![(Modality::Tensor, 150.0, 5)]; - - log.record(Some("q1"), 150.0, &plan, &step_times); - assert_eq!(log.count(), 1); - - log.clear(); - assert_eq!(log.count(), 0); - } - - #[test] - fn test_summary_stats() { - let config = SlowQueryConfig { - threshold_ms: 10.0, - ..Default::default() - }; - let log = SlowQueryLog::new(config); - let plan = make_plan(vec![(Modality::Vector, 10.0)]); - - for ms in [50.0, 100.0, 150.0] { - let step_times = vec![(Modality::Vector, ms, 10)]; - log.record(None, ms, &plan, &step_times); - } - - let summary = log.summary(); - assert_eq!(summary.total_count, 3); - assert!((summary.avg_ms - 100.0).abs() < 0.01); - assert!((summary.max_ms - 150.0).abs() < 0.01); - assert!((summary.min_ms - 50.0).abs() < 0.01); - assert_eq!(summary.top_bottleneck_modality, Some(Modality::Vector)); - } - - #[test] - fn test_empty_summary() { - let log = SlowQueryLog::with_defaults(); - let summary = log.summary(); - assert_eq!(summary.total_count, 0); - assert_eq!(summary.top_bottleneck_modality, None); - } - - #[test] - fn test_config_update() { - let log = SlowQueryLog::with_defaults(); - assert!((log.config().threshold_ms - 100.0).abs() < f64::EPSILON); - - log.set_config(SlowQueryConfig { - threshold_ms: 500.0, - ..Default::default() - }); - assert!((log.config().threshold_ms - 500.0).abs() < f64::EPSILON); - } - - #[test] - fn test_recent_ordering() { - let config = SlowQueryConfig { - threshold_ms: 10.0, - ..Default::default() - }; - let log = SlowQueryLog::new(config); - let plan = make_plan(vec![(Modality::Document, 10.0)]); - - for ms in [20.0, 30.0, 40.0, 50.0] { - let step_times = vec![(Modality::Document, ms, 1)]; - log.record(None, ms, &plan, &step_times); - } - - // recent() should return newest first - let recent = log.recent(2); - assert_eq!(recent.len(), 2); - assert!(recent[0].actual_ms >= recent[1].actual_ms); - } - - #[test] - fn test_json_serialization() { - let config = SlowQueryConfig { - threshold_ms: 10.0, - ..Default::default() - }; - let log = SlowQueryLog::new(config); - let plan = make_plan(vec![(Modality::Temporal, 5.0)]); - let step_times = vec![(Modality::Temporal, 20.0, 3)]; - - log.record(Some("SELECT TEMPORAL FROM OCTAD"), 20.0, &plan, &step_times); - - let entries = log.recent(1); - let json = serde_json::to_string(&entries[0]).expect("TODO: handle error"); - let parsed: SlowQueryEntry = serde_json::from_str(&json).expect("TODO: handle error"); - assert_eq!(parsed.query_text, Some("SELECT TEMPORAL FROM OCTAD".to_string())); - } -} diff --git a/verisimdb/rust-core/verisim-planner/src/stats.rs b/verisimdb/rust-core/verisim-planner/src/stats.rs deleted file mode 100644 index a8b62efc..00000000 --- a/verisimdb/rust-core/verisim-planner/src/stats.rs +++ /dev/null @@ -1,382 +0,0 @@ -// SPDX-License-Identifier: MPL-2.0 -//! Store statistics collection and tracking. - -use std::collections::HashMap; - -use chrono::{DateTime, Utc}; -use serde::{Deserialize, Serialize}; - -use crate::Modality; - -/// Statistics for a single modality store. -#[derive(Debug, Clone, Serialize, Deserialize)] -pub struct StoreStatistics { - /// Which modality these stats describe. - pub modality: Modality, - /// Total number of rows/entities in this store. - pub total_rows: u64, - /// Average query latency in milliseconds (exponential moving average). - pub avg_latency_ms: f64, - /// Average number of rows returned per query. - pub avg_rows_returned: u64, - /// Total number of queries executed. - pub query_count: u64, - /// When statistics were last updated. - pub last_updated: DateTime<Utc>, -} - -impl StoreStatistics { - /// Create empty statistics for a modality. - fn new(modality: Modality) -> Self { - Self { - modality, - total_rows: 0, - avg_latency_ms: 0.0, - avg_rows_returned: 0, - query_count: 0, - last_updated: Utc::now(), - } - } -} - -/// Collects and maintains statistics across all modality stores. -#[derive(Debug, Clone, Serialize, Deserialize)] -pub struct StatisticsCollector { - stats: HashMap<Modality, StoreStatistics>, -} - -impl StatisticsCollector { - /// Create a new collector with empty statistics for all 6 modalities. - pub fn new() -> Self { - let mut stats = HashMap::new(); - for m in Modality::ALL { - stats.insert(m, StoreStatistics::new(m)); - } - Self { stats } - } - - /// Get statistics for a specific modality. - pub fn get(&self, modality: Modality) -> Option<&StoreStatistics> { - self.stats.get(&modality) - } - - /// Get a snapshot of all statistics. - pub fn snapshot(&self) -> &HashMap<Modality, StoreStatistics> { - &self.stats - } - - /// Record a query execution for a modality. - /// - /// Uses exponential moving average (alpha=0.1) for latency, - /// matching the drift detector's approach. - pub fn record_execution( - &mut self, - modality: Modality, - latency_ms: f64, - rows_returned: u64, - ) { - let entry = self.stats.entry(modality).or_insert_with(|| StoreStatistics::new(modality)); - entry.query_count += 1; - - // Exponential moving average (alpha = 0.1) - if entry.query_count == 1 { - entry.avg_latency_ms = latency_ms; - entry.avg_rows_returned = rows_returned; - } else { - entry.avg_latency_ms = 0.1 * latency_ms + 0.9 * entry.avg_latency_ms; - entry.avg_rows_returned = - (0.1 * rows_returned as f64 + 0.9 * entry.avg_rows_returned as f64) as u64; - } - - entry.last_updated = Utc::now(); - } - - /// Update the total row count for a modality. - pub fn update_row_count(&mut self, modality: Modality, total_rows: u64) { - if let Some(entry) = self.stats.get_mut(&modality) { - entry.total_rows = total_rows; - entry.last_updated = Utc::now(); - } - } -} - -impl Default for StatisticsCollector { - fn default() -> Self { - Self::new() - } -} - -/// Adaptive tuner that adjusts planner configuration based on actual -/// execution performance vs estimated costs. -/// -/// When enabled, the tuner compares actual query latencies against -/// the planner's estimates and adjusts per-modality optimization modes: -/// - If actual >> estimated → switch to Conservative (underestimating) -/// - If actual << estimated → switch to Aggressive (overestimating) -/// - Otherwise → keep Balanced -pub struct AdaptiveTuner { - /// Ratio of actual/estimated below which we go Aggressive. - aggressive_threshold: f64, - /// Ratio of actual/estimated above which we go Conservative. - conservative_threshold: f64, - /// Minimum number of samples before making adjustments. - min_samples: u64, -} - -impl AdaptiveTuner { - /// Create a new adaptive tuner with default thresholds. - pub fn new() -> Self { - Self { - aggressive_threshold: 0.5, // Actual < 50% of estimate → overestimating - conservative_threshold: 2.0, // Actual > 200% of estimate → underestimating - min_samples: 10, - } - } - - /// Create a tuner with custom thresholds. - pub fn with_thresholds(aggressive: f64, conservative: f64, min_samples: u64) -> Self { - Self { - aggressive_threshold: aggressive, - conservative_threshold: conservative, - min_samples, - } - } - - /// Evaluate the collector's statistics and suggest config adjustments. - /// - /// Returns a list of (Modality, suggested OptimizationMode) pairs - /// for modalities that should be tuned. - pub fn suggest_adjustments( - &self, - collector: &StatisticsCollector, - config: &crate::config::PlannerConfig, - ) -> Vec<(crate::Modality, crate::config::OptimizationMode)> { - let mut adjustments = Vec::new(); - - for modality in crate::Modality::ALL { - if let Some(stats) = collector.get(modality) { - if stats.query_count < self.min_samples { - continue; // Not enough data - } - - let base_cost = crate::cost::BaseCost::for_modality(modality); - let current_mode = config.mode_for(modality); - let estimated_ms = base_cost.time_ms * current_mode.cost_multiplier(); - - if estimated_ms <= 0.0 { - continue; - } - - let ratio = stats.avg_latency_ms / estimated_ms; - - let suggested = if ratio < self.aggressive_threshold { - crate::config::OptimizationMode::Aggressive - } else if ratio > self.conservative_threshold { - crate::config::OptimizationMode::Conservative - } else { - crate::config::OptimizationMode::Balanced - }; - - if suggested != current_mode { - adjustments.push((modality, suggested)); - } - } - } - - adjustments - } - - /// Apply suggested adjustments to a config, returning the updated config. - pub fn apply( - &self, - collector: &StatisticsCollector, - config: &crate::config::PlannerConfig, - ) -> crate::config::PlannerConfig { - if !config.enable_adaptive { - return config.clone(); - } - - let adjustments = self.suggest_adjustments(collector, config); - if adjustments.is_empty() { - return config.clone(); - } - - let mut new_config = config.clone(); - for (modality, mode) in adjustments { - new_config.modality_overrides.insert(modality, mode); - } - new_config - } -} - -impl Default for AdaptiveTuner { - fn default() -> Self { - Self::new() - } -} - -#[cfg(test)] -mod tests { - use super::*; - - #[test] - fn test_collector_initializes_all_modalities() { - let collector = StatisticsCollector::new(); - for m in Modality::ALL { - let stats = collector.get(m).expect("TODO: handle error"); - assert_eq!(stats.modality, m); - assert_eq!(stats.query_count, 0); - assert_eq!(stats.total_rows, 0); - } - assert_eq!(collector.snapshot().len(), 6); - } - - #[test] - fn test_record_first_execution() { - let mut collector = StatisticsCollector::new(); - collector.record_execution(Modality::Vector, 42.0, 10); - - let stats = collector.get(Modality::Vector).expect("TODO: handle error"); - assert_eq!(stats.query_count, 1); - assert!((stats.avg_latency_ms - 42.0).abs() < f64::EPSILON); - assert_eq!(stats.avg_rows_returned, 10); - } - - #[test] - fn test_record_updates_moving_average() { - let mut collector = StatisticsCollector::new(); - - // First execution - collector.record_execution(Modality::Graph, 100.0, 50); - assert!((collector.get(Modality::Graph).expect("TODO: handle error").avg_latency_ms - 100.0).abs() < f64::EPSILON); - - // Second execution — EMA: 0.1 * 200 + 0.9 * 100 = 110 - collector.record_execution(Modality::Graph, 200.0, 100); - let stats = collector.get(Modality::Graph).expect("TODO: handle error"); - assert!((stats.avg_latency_ms - 110.0).abs() < f64::EPSILON); - assert_eq!(stats.query_count, 2); - } - - #[test] - fn test_update_row_count() { - let mut collector = StatisticsCollector::new(); - collector.update_row_count(Modality::Document, 5000); - assert_eq!(collector.get(Modality::Document).expect("TODO: handle error").total_rows, 5000); - } - - #[test] - fn test_snapshot_returns_all() { - let collector = StatisticsCollector::new(); - let snap = collector.snapshot(); - assert_eq!(snap.len(), 6); - } - - // ==================================================================== - // Task #8: Adaptive tuning tests - // ==================================================================== - - #[test] - fn test_adaptive_tuner_no_data_no_adjustments() { - let tuner = AdaptiveTuner::new(); - let collector = StatisticsCollector::new(); - let config = crate::config::PlannerConfig::default(); - let adjustments = tuner.suggest_adjustments(&collector, &config); - assert!(adjustments.is_empty(), "No data → no adjustments"); - } - - #[test] - fn test_adaptive_tuner_below_min_samples() { - let tuner = AdaptiveTuner::new(); // min_samples = 10 - let mut collector = StatisticsCollector::new(); - // Record only 5 executions (below threshold) - for _ in 0..5 { - collector.record_execution(Modality::Vector, 10.0, 5); - } - let config = crate::config::PlannerConfig::default(); - let adjustments = tuner.suggest_adjustments(&collector, &config); - assert!(adjustments.is_empty(), "Below min_samples → no adjustments"); - } - - #[test] - fn test_adaptive_tuner_suggests_aggressive_when_overestimating() { - let tuner = AdaptiveTuner::new(); - let mut collector = StatisticsCollector::new(); - // Vector base = 50ms, aggressive mode = 0.8x = 40ms estimated - // Record actual latency of 10ms (ratio = 10/40 = 0.25 < 0.5) → Aggressive - for _ in 0..15 { - collector.record_execution(Modality::Vector, 10.0, 5); - } - let config = crate::config::PlannerConfig::default(); - let adjustments = tuner.suggest_adjustments(&collector, &config); - // Vector already has Aggressive override, so if actual confirms it, no change. - // But ratio 0.25 < 0.5, already aggressive, stays aggressive → no adjustment. - // Let's test Graph instead where it's Conservative - let mut collector2 = StatisticsCollector::new(); - // Graph base = 150ms, conservative mode = 1.5x = 225ms estimated - // Record actual latency of 50ms (ratio = 50/225 = 0.22 < 0.5) → Aggressive - for _ in 0..15 { - collector2.record_execution(Modality::Graph, 50.0, 10); - } - let adjustments2 = tuner.suggest_adjustments(&collector2, &config); - let graph_adj = adjustments2.iter().find(|(m, _)| *m == Modality::Graph); - assert!(graph_adj.is_some(), "Graph should have adjustment"); - assert_eq!( - graph_adj.expect("TODO: handle error").1, - crate::config::OptimizationMode::Aggressive, - "Actual << estimated → Aggressive" - ); - } - - #[test] - fn test_adaptive_tuner_suggests_conservative_when_underestimating() { - let tuner = AdaptiveTuner::new(); - let mut collector = StatisticsCollector::new(); - // Document base = 80ms, balanced mode = 1.0x = 80ms estimated - // Record actual latency of 200ms (ratio = 200/80 = 2.5 > 2.0) → Conservative - for _ in 0..15 { - collector.record_execution(Modality::Document, 200.0, 50); - } - let config = crate::config::PlannerConfig::default(); - let adjustments = tuner.suggest_adjustments(&collector, &config); - let doc_adj = adjustments.iter().find(|(m, _)| *m == Modality::Document); - assert!(doc_adj.is_some(), "Document should have adjustment"); - assert_eq!( - doc_adj.expect("TODO: handle error").1, - crate::config::OptimizationMode::Conservative, - "Actual >> estimated → Conservative" - ); - } - - #[test] - fn test_adaptive_tuner_apply_updates_config() { - let tuner = AdaptiveTuner::new(); - let mut collector = StatisticsCollector::new(); - // Make Document wildly underestimated - for _ in 0..15 { - collector.record_execution(Modality::Document, 200.0, 50); - } - let config = crate::config::PlannerConfig::default(); - let new_config = tuner.apply(&collector, &config); - assert_eq!( - new_config.mode_for(Modality::Document), - crate::config::OptimizationMode::Conservative, - ); - } - - #[test] - fn test_adaptive_tuner_disabled() { - let tuner = AdaptiveTuner::new(); - let mut collector = StatisticsCollector::new(); - for _ in 0..15 { - collector.record_execution(Modality::Document, 200.0, 50); - } - let mut config = crate::config::PlannerConfig::default(); - config.enable_adaptive = false; - let new_config = tuner.apply(&collector, &config); - // Should not change when adaptive is disabled - assert_eq!( - new_config.mode_for(Modality::Document), - config.mode_for(Modality::Document), - ); - } -} diff --git a/verisimdb/rust-core/verisim-planner/src/vcl_bridge.rs b/verisimdb/rust-core/verisim-planner/src/vcl_bridge.rs deleted file mode 100644 index 83dc9f22..00000000 --- a/verisimdb/rust-core/verisim-planner/src/vcl_bridge.rs +++ /dev/null @@ -1,1774 +0,0 @@ -// SPDX-License-Identifier: MPL-2.0 -// Copyright (c) 2026 Jonathan D.A. Jewell (hyperpolymath) <j.d.a.jewell@open.ac.uk> -//! -//! VCL AST to LogicalPlan bridge. -//! -//! Deserializes VCL JSON produced by the ReScript parser (BuckleScript encoding) -//! and converts it into the [`LogicalPlan`] representation used by the planner. -//! -//! ## BuckleScript Encoding -//! -//! The ReScript compiler (via BuckleScript) encodes variant types as JSON objects -//! with a `TAG` field naming the constructor and positional `_0`, `_1`, ... fields -//! for arguments: -//! -//! ```json -//! { "TAG": "Octad", "_0": "some-uuid" } -//! ``` -//! -//! This module defines serde-compatible Rust types that mirror this encoding and -//! provides conversion into the planner's canonical [`LogicalPlan`]. - -use serde::de::{self, MapAccess, Visitor}; -use serde::{Deserialize, Deserializer}; -use std::collections::HashMap; -use std::fmt; - -use crate::error::PlannerError; -use crate::plan::{ConditionKind, LogicalPlan, PlanNode, PostProcessing, QuerySource}; -use crate::Modality; - -// --------------------------------------------------------------------------- -// VCL AST types (mirrors BuckleScript JSON encoding) -// --------------------------------------------------------------------------- - -/// Top-level VCL statement as emitted by the ReScript parser. -/// -/// BuckleScript encodes this as `{"TAG": "Query", "_0": { ... }}`. -/// We use a custom deserializer because serde's internally-tagged enum -/// (`#[serde(tag = "TAG")]`) does not support positional `_0` content fields. -#[derive(Debug, Clone)] -pub enum VclAst { - /// A SELECT-style query. - Query(VclQuery), -} - -impl<'de> Deserialize<'de> for VclAst { - fn deserialize<D>(deserializer: D) -> Result<Self, D::Error> - where - D: Deserializer<'de>, - { - struct VclAstVisitor; - - impl<'de> Visitor<'de> for VclAstVisitor { - type Value = VclAst; - - fn expecting(&self, formatter: &mut fmt::Formatter) -> fmt::Result { - formatter.write_str("a VCL AST object with TAG and _0 fields") - } - - fn visit_map<M>(self, mut map: M) -> Result<VclAst, M::Error> - where - M: MapAccess<'de>, - { - let mut tag: Option<String> = None; - let mut payload: Option<serde_json::Value> = None; - - while let Some(key) = map.next_key::<String>()? { - match key.as_str() { - "TAG" => tag = Some(map.next_value()?), - "_0" => payload = Some(map.next_value()?), - _ => { - let _: serde_json::Value = map.next_value()?; - } - } - } - - let tag = tag.ok_or_else(|| de::Error::missing_field("TAG"))?; - match tag.as_str() { - "Query" => { - let body = payload.ok_or_else(|| de::Error::missing_field("_0"))?; - let query: VclQuery = - serde_json::from_value(body).map_err(de::Error::custom)?; - Ok(VclAst::Query(query)) - } - other => Err(de::Error::unknown_variant(other, &["Query"])), - } - } - } - - deserializer.deserialize_map(VclAstVisitor) - } -} - -/// Body of a VCL query. -#[derive(Debug, Clone, Deserialize)] -#[serde(rename_all = "camelCase")] -pub struct VclQuery { - /// Requested modalities (`Graph`, `Vector`, ..., or `All`). - pub modalities: Vec<VclModality>, - /// Data source (Octad, Federation, Store). - pub source: VclSource, - /// Optional WHERE clause. - #[serde(default, rename = "where")] - pub where_clause: Option<VclCondition>, - /// Optional field projections. - #[serde(default)] - pub projections: Option<Vec<VclProjection>>, - /// Optional aggregate functions. - #[serde(default)] - pub aggregates: Option<Vec<VclAggregate>>, - /// Optional GROUP BY fields. - #[serde(default)] - pub group_by: Option<Vec<VclFieldRef>>, - /// Optional HAVING clause. - #[serde(default)] - pub having: Option<VclCondition>, - /// Optional PROOF specifications. - #[serde(default)] - pub proof: Option<Vec<VclProofSpec>>, - /// Optional ORDER BY clauses. - #[serde(default)] - pub order_by: Option<Vec<VclOrderBy>>, - /// Optional result limit. - #[serde(default)] - pub limit: Option<usize>, - /// Optional result offset. - #[serde(default)] - pub offset: Option<usize>, -} - -/// A VCL modality tag. -/// -/// The ReScript parser emits modalities as `{"TAG": "Graph"}` etc. -/// `All` is a special sentinel meaning "expand to all 6 modalities". -#[derive(Debug, Clone)] -pub enum VclModality { - Graph, - Vector, - Tensor, - Semantic, - Document, - Temporal, - All, -} - -impl<'de> Deserialize<'de> for VclModality { - fn deserialize<D>(deserializer: D) -> Result<Self, D::Error> - where - D: Deserializer<'de>, - { - struct VclModalityVisitor; - - impl<'de> Visitor<'de> for VclModalityVisitor { - type Value = VclModality; - - fn expecting(&self, formatter: &mut fmt::Formatter) -> fmt::Result { - formatter.write_str("a VCL modality object with TAG field") - } - - fn visit_map<M>(self, mut map: M) -> Result<VclModality, M::Error> - where - M: MapAccess<'de>, - { - let mut tag: Option<String> = None; - while let Some(key) = map.next_key::<String>()? { - if key == "TAG" { - tag = Some(map.next_value()?); - } else { - // Skip unknown fields. - let _: serde_json::Value = map.next_value()?; - } - } - match tag.as_deref() { - Some("Graph") => Ok(VclModality::Graph), - Some("Vector") => Ok(VclModality::Vector), - Some("Tensor") => Ok(VclModality::Tensor), - Some("Semantic") => Ok(VclModality::Semantic), - Some("Document") => Ok(VclModality::Document), - Some("Temporal") => Ok(VclModality::Temporal), - Some("All") => Ok(VclModality::All), - Some(other) => Err(de::Error::unknown_variant( - other, - &[ - "Graph", "Vector", "Tensor", "Semantic", "Document", "Temporal", "All", - ], - )), - None => Err(de::Error::missing_field("TAG")), - } - } - } - - deserializer.deserialize_map(VclModalityVisitor) - } -} - -/// Data source as emitted by the ReScript parser. -#[derive(Debug, Clone, Deserialize)] -#[serde(tag = "TAG")] -pub enum VclSource { - /// Single octad store. `_0` is an optional UUID filter. - Octad { - #[serde(rename = "_0")] - uuid: Option<String>, - }, - /// Federated query with drift policy. - Federation { - #[serde(rename = "_0")] - nodes: Vec<String>, - #[serde(rename = "_1")] - drift_policy: Option<String>, - }, - /// Direct store access for a specific modality. - Store { - #[serde(rename = "_0")] - modality: VclModality, - }, -} - -/// A VCL condition (WHERE clause tree). -/// -/// BuckleScript encodes each variant with TAG + positional args. -#[derive(Debug, Clone)] -pub enum VclCondition { - /// Conjunction. - And(Box<VclCondition>, Box<VclCondition>), - /// Disjunction. - Or(Box<VclCondition>, Box<VclCondition>), - /// Negation. - Not(Box<VclCondition>), - /// Leaf condition. - Simple(VclSimpleCondition), -} - -impl<'de> Deserialize<'de> for VclCondition { - fn deserialize<D>(deserializer: D) -> Result<Self, D::Error> - where - D: Deserializer<'de>, - { - // Deserialize as a generic JSON value first, then pattern match on TAG. - let value = serde_json::Value::deserialize(deserializer)?; - parse_condition(&value).map_err(de::Error::custom) - } -} - -/// Parse a `VclCondition` from a `serde_json::Value`. -fn parse_condition(value: &serde_json::Value) -> Result<VclCondition, String> { - let obj = value.as_object().ok_or("condition must be a JSON object")?; - let tag = obj - .get("TAG") - .and_then(|v| v.as_str()) - .ok_or("condition object missing TAG field")?; - - match tag { - "And" => { - let lhs = obj - .get("_0") - .ok_or("And condition missing _0 (left operand)")?; - let rhs = obj - .get("_1") - .ok_or("And condition missing _1 (right operand)")?; - Ok(VclCondition::And( - Box::new(parse_condition(lhs)?), - Box::new(parse_condition(rhs)?), - )) - } - "Or" => { - let lhs = obj - .get("_0") - .ok_or("Or condition missing _0 (left operand)")?; - let rhs = obj - .get("_1") - .ok_or("Or condition missing _1 (right operand)")?; - Ok(VclCondition::Or( - Box::new(parse_condition(lhs)?), - Box::new(parse_condition(rhs)?), - )) - } - "Not" => { - let inner = obj.get("_0").ok_or("Not condition missing _0 (operand)")?; - Ok(VclCondition::Not(Box::new(parse_condition(inner)?))) - } - "Simple" => { - let inner = obj.get("_0").ok_or("Simple condition missing _0")?; - let simple: VclSimpleCondition = - serde_json::from_value(inner.clone()).map_err(|e| e.to_string())?; - Ok(VclCondition::Simple(simple)) - } - other => Err(format!("unknown condition TAG: {other}")), - } -} - -/// Leaf-level condition kinds from VCL. -#[derive(Debug, Clone)] -pub enum VclSimpleCondition { - /// Full-text search: `CONTAINS "search text"`. - FulltextContains(String), - /// Vector similarity: `SIMILAR TO [embedding] THRESHOLD threshold`. - VectorSimilar { - embedding: Vec<f64>, - threshold: f64, - }, - /// Graph pattern: `TRAVERSE predicate DEPTH depth`. - GraphPattern { - predicate: String, - depth: Option<u32>, - }, - /// Field condition: `field op value`. - FieldCondition { - field: VclFieldRef, - operator: String, - value: serde_json::Value, - }, - /// Cross-modal field comparison. - CrossModalFieldCompare { - left: VclFieldRef, - operator: String, - right: VclFieldRef, - }, - /// Modality drift check. - ModalityDrift { - modality: VclModality, - threshold: f64, - }, - /// Modality existence check. - ModalityExists(VclModality), - /// Modality non-existence check. - ModalityNotExists(VclModality), - /// Cross-modality consistency check. - ModalityConsistency { - modalities: Vec<VclModality>, - threshold: f64, - }, -} - -impl<'de> Deserialize<'de> for VclSimpleCondition { - fn deserialize<D>(deserializer: D) -> Result<Self, D::Error> - where - D: Deserializer<'de>, - { - let value = serde_json::Value::deserialize(deserializer)?; - parse_simple_condition(&value).map_err(de::Error::custom) - } -} - -/// Parse a `VclSimpleCondition` from raw JSON. -fn parse_simple_condition(value: &serde_json::Value) -> Result<VclSimpleCondition, String> { - let obj = value - .as_object() - .ok_or("simple condition must be a JSON object")?; - let tag = obj - .get("TAG") - .and_then(|v| v.as_str()) - .ok_or("simple condition object missing TAG field")?; - - match tag { - "FulltextContains" => { - let text = obj - .get("_0") - .and_then(|v| v.as_str()) - .ok_or("FulltextContains missing _0 (search text)")?; - Ok(VclSimpleCondition::FulltextContains(text.to_string())) - } - "VectorSimilar" => { - let embedding = obj - .get("_0") - .and_then(|v| v.as_array()) - .ok_or("VectorSimilar missing _0 (embedding array)")? - .iter() - .map(|v| { - v.as_f64() - .ok_or_else(|| "VectorSimilar _0 contains non-numeric".to_string()) - }) - .collect::<Result<Vec<f64>, String>>()?; - let threshold = obj - .get("_1") - .and_then(|v| v.as_f64()) - .ok_or("VectorSimilar missing _1 (threshold)")?; - Ok(VclSimpleCondition::VectorSimilar { - embedding, - threshold, - }) - } - "GraphPattern" => { - let predicate = obj - .get("_0") - .and_then(|v| v.as_str()) - .ok_or("GraphPattern missing _0 (predicate)")? - .to_string(); - let depth = obj.get("_1").and_then(|v| v.as_u64()).map(|d| d as u32); - Ok(VclSimpleCondition::GraphPattern { predicate, depth }) - } - "FieldCondition" => { - let field_val = obj.get("_0").ok_or("FieldCondition missing _0 (field)")?; - let field: VclFieldRef = - serde_json::from_value(field_val.clone()).map_err(|e| e.to_string())?; - let operator = obj - .get("_1") - .and_then(|v| v.as_str()) - .ok_or("FieldCondition missing _1 (operator)")? - .to_string(); - let val = obj - .get("_2") - .cloned() - .ok_or("FieldCondition missing _2 (value)")?; - Ok(VclSimpleCondition::FieldCondition { - field, - operator, - value: val, - }) - } - "CrossModalFieldCompare" => { - let left_val = obj - .get("_0") - .ok_or("CrossModalFieldCompare missing _0 (left)")?; - let left: VclFieldRef = - serde_json::from_value(left_val.clone()).map_err(|e| e.to_string())?; - let operator = obj - .get("_1") - .and_then(|v| v.as_str()) - .ok_or("CrossModalFieldCompare missing _1 (operator)")? - .to_string(); - let right_val = obj - .get("_2") - .ok_or("CrossModalFieldCompare missing _2 (right)")?; - let right: VclFieldRef = - serde_json::from_value(right_val.clone()).map_err(|e| e.to_string())?; - Ok(VclSimpleCondition::CrossModalFieldCompare { - left, - operator, - right, - }) - } - "ModalityDrift" => { - let mod_val = obj - .get("_0") - .ok_or("ModalityDrift missing _0 (modality)")?; - let modality: VclModality = - serde_json::from_value(mod_val.clone()).map_err(|e| e.to_string())?; - let threshold = obj - .get("_1") - .and_then(|v| v.as_f64()) - .ok_or("ModalityDrift missing _1 (threshold)")?; - Ok(VclSimpleCondition::ModalityDrift { - modality, - threshold, - }) - } - "ModalityExists" => { - let mod_val = obj - .get("_0") - .ok_or("ModalityExists missing _0 (modality)")?; - let modality: VclModality = - serde_json::from_value(mod_val.clone()).map_err(|e| e.to_string())?; - Ok(VclSimpleCondition::ModalityExists(modality)) - } - "ModalityNotExists" => { - let mod_val = obj - .get("_0") - .ok_or("ModalityNotExists missing _0 (modality)")?; - let modality: VclModality = - serde_json::from_value(mod_val.clone()).map_err(|e| e.to_string())?; - Ok(VclSimpleCondition::ModalityNotExists(modality)) - } - "ModalityConsistency" => { - let mods_val = obj - .get("_0") - .ok_or("ModalityConsistency missing _0 (modalities)")?; - let modalities: Vec<VclModality> = - serde_json::from_value(mods_val.clone()).map_err(|e| e.to_string())?; - let threshold = obj - .get("_1") - .and_then(|v| v.as_f64()) - .ok_or("ModalityConsistency missing _1 (threshold)")?; - Ok(VclSimpleCondition::ModalityConsistency { - modalities, - threshold, - }) - } - other => Err(format!("unknown simple condition TAG: {other}")), - } -} - -/// A field reference with optional modality qualifier. -#[derive(Debug, Clone, Deserialize)] -pub struct VclFieldRef { - /// Optional modality qualifier for the field. - #[serde(default)] - pub modality: Option<VclModality>, - /// Field name. - pub field: String, -} - -/// A projection entry. -#[derive(Debug, Clone, Deserialize)] -pub struct VclProjection { - /// The field being projected. - pub field: VclFieldRef, - /// Optional alias. - #[serde(default)] - pub alias: Option<String>, -} - -/// An aggregate function call. -#[derive(Debug, Clone, Deserialize)] -pub struct VclAggregate { - /// Function name (COUNT, SUM, AVG, MIN, MAX, etc.). - pub function: String, - /// Field to aggregate (None for COUNT(*)). - #[serde(default)] - pub field: Option<VclFieldRef>, - /// Optional alias for the result. - #[serde(default)] - pub alias: Option<String>, -} - -/// A proof specification from the VCL PROOF clause. -#[derive(Debug, Clone, Deserialize)] -#[serde(rename_all = "camelCase")] -pub struct VclProofSpec { - /// Type of proof required. - pub proof_type: VclProofType, - /// Name of the verification contract. - pub contract_name: String, -} - -/// Proof type tag. -#[derive(Debug, Clone)] -pub enum VclProofType { - Citation, - Zkp, - Attestation, - Custom(String), -} - -impl<'de> Deserialize<'de> for VclProofType { - fn deserialize<D>(deserializer: D) -> Result<Self, D::Error> - where - D: Deserializer<'de>, - { - struct ProofTypeVisitor; - - impl<'de> Visitor<'de> for ProofTypeVisitor { - type Value = VclProofType; - - fn expecting(&self, formatter: &mut fmt::Formatter) -> fmt::Result { - formatter.write_str("a VCL proof type object with TAG field") - } - - fn visit_map<M>(self, mut map: M) -> Result<VclProofType, M::Error> - where - M: MapAccess<'de>, - { - let mut tag: Option<String> = None; - while let Some(key) = map.next_key::<String>()? { - if key == "TAG" { - tag = Some(map.next_value()?); - } else { - let _: serde_json::Value = map.next_value()?; - } - } - match tag.as_deref() { - Some("Citation") => Ok(VclProofType::Citation), - Some("Zkp") => Ok(VclProofType::Zkp), - Some("Attestation") => Ok(VclProofType::Attestation), - Some(other) => Ok(VclProofType::Custom(other.to_string())), - None => Err(de::Error::missing_field("TAG")), - } - } - } - - deserializer.deserialize_map(ProofTypeVisitor) - } -} - -/// An ORDER BY clause entry. -#[derive(Debug, Clone, Deserialize)] -pub struct VclOrderBy { - /// Field to order by. - pub field: VclFieldRef, - /// Sort direction. - pub direction: VclDirection, -} - -/// Sort direction tag. -#[derive(Debug, Clone)] -pub enum VclDirection { - Asc, - Desc, -} - -impl<'de> Deserialize<'de> for VclDirection { - fn deserialize<D>(deserializer: D) -> Result<Self, D::Error> - where - D: Deserializer<'de>, - { - struct DirectionVisitor; - - impl<'de> Visitor<'de> for DirectionVisitor { - type Value = VclDirection; - - fn expecting(&self, formatter: &mut fmt::Formatter) -> fmt::Result { - formatter.write_str("a VCL direction object with TAG field") - } - - fn visit_map<M>(self, mut map: M) -> Result<VclDirection, M::Error> - where - M: MapAccess<'de>, - { - let mut tag: Option<String> = None; - while let Some(key) = map.next_key::<String>()? { - if key == "TAG" { - tag = Some(map.next_value()?); - } else { - let _: serde_json::Value = map.next_value()?; - } - } - match tag.as_deref() { - Some("Asc") => Ok(VclDirection::Asc), - Some("Desc") => Ok(VclDirection::Desc), - Some(other) => Err(de::Error::unknown_variant(other, &["Asc", "Desc"])), - None => Err(de::Error::missing_field("TAG")), - } - } - } - - deserializer.deserialize_map(DirectionVisitor) - } -} - -// --------------------------------------------------------------------------- -// Conversion: VclAst -> LogicalPlan -// --------------------------------------------------------------------------- - -impl VclAst { - /// Deserialize a VCL AST from JSON emitted by the ReScript parser. - /// - /// # Errors - /// - /// Returns `PlannerError::Serialization` if the JSON does not match the - /// expected BuckleScript encoding. - pub fn from_json(json: &str) -> Result<Self, PlannerError> { - serde_json::from_str(json).map_err(PlannerError::Serialization) - } - - /// Convert the VCL AST into a [`LogicalPlan`]. - /// - /// This is the primary bridge function. It: - /// 1. Expands `All` modality into all six concrete modalities. - /// 2. Maps the VCL source to [`QuerySource`]. - /// 3. Distributes WHERE conditions to per-modality [`PlanNode`]s. - /// 4. Extracts LIMIT, OFFSET, ORDER BY, GROUP BY into [`PostProcessing`]. - /// 5. Maps PROOF specs into [`ConditionKind::ProofVerification`] on Semantic nodes. - /// - /// # Errors - /// - /// Returns `PlannerError::EmptyPlan` if no modalities are requested. - pub fn to_logical_plan(&self) -> Result<LogicalPlan, PlannerError> { - let VclAst::Query(query) = self; - - // 1. Resolve modalities (expand All). - let modalities = resolve_modalities(&query.modalities)?; - if modalities.is_empty() { - return Err(PlannerError::EmptyPlan); - } - - // 2. Map source. - let source = map_source(&query.source)?; - - // 3. Build per-modality condition buckets. - let mut condition_map: HashMap<Modality, Vec<ConditionKind>> = HashMap::new(); - for &m in &modalities { - condition_map.entry(m).or_default(); - } - - if let Some(ref cond) = query.where_clause { - flatten_conditions(cond, &modalities, &mut condition_map)?; - } - - // 4. Inject proof verification conditions into the Semantic node. - if let Some(ref proofs) = query.proof { - for spec in proofs { - let semantic_conditions = condition_map.entry(Modality::Semantic).or_default(); - semantic_conditions.push(ConditionKind::ProofVerification { - contract: spec.contract_name.clone(), - }); - } - } - - // 5. Build per-modality projections. - let projection_map = build_projection_map(&modalities, query.projections.as_deref()); - - // 6. Assemble PlanNodes. - let mut nodes: Vec<PlanNode> = modalities - .iter() - .map(|&m| PlanNode { - modality: m, - conditions: condition_map.remove(&m).unwrap_or_default(), - projections: projection_map - .get(&m) - .cloned() - .unwrap_or_default(), - early_limit: None, - }) - .collect(); - - // Sort nodes by execution priority for predictable output. - nodes.sort_by_key(|n| n.modality.execution_priority()); - - // 7. Build post-processing pipeline. - let mut post_processing = Vec::new(); - - if let Some(ref group_fields) = query.group_by { - let fields: Vec<String> = group_fields.iter().map(field_ref_name).collect(); - let aggregates: Vec<String> = query - .aggregates - .as_ref() - .map(|aggs| { - aggs.iter() - .map(|a| { - let field_name = a - .field - .as_ref() - .map(field_ref_name) - .unwrap_or_else(|| "*".to_string()); - format!("{}({})", a.function, field_name) - }) - .collect() - }) - .unwrap_or_default(); - post_processing.push(PostProcessing::GroupBy { fields, aggregates }); - } - - if let Some(ref order_fields) = query.order_by { - let fields: Vec<(String, bool)> = order_fields - .iter() - .map(|o| { - let ascending = matches!(o.direction, VclDirection::Asc); - (field_ref_name(&o.field), ascending) - }) - .collect(); - post_processing.push(PostProcessing::OrderBy { fields }); - } - - if let Some(count) = query.limit { - post_processing.push(PostProcessing::Limit { count }); - } - - // Offset is modelled as a Limit post-processing with adjusted count. - // The planner does not have a dedicated Offset variant, so we encode it - // by bumping the limit to include skipped rows. Physical plan will - // handle the actual skip. If only offset is given (no limit), we add a - // large limit. - if let Some(skip) = query.offset { - if skip > 0 { - // Find existing Limit and adjust, or add one. - let has_limit = post_processing.iter().any(|p| matches!(p, PostProcessing::Limit { .. })); - if !has_limit { - // No limit — add a large synthetic limit so offset is meaningful. - post_processing.push(PostProcessing::Limit { - count: usize::MAX - skip, - }); - } - // The physical plan executor is responsible for applying the offset. - // We store it as a Project marker so it can be recognised later. - post_processing.push(PostProcessing::Project { - columns: vec![format!("__offset={skip}")], - }); - } - } - - // Final projection (if explicit projections were requested and no group-by). - if query.group_by.is_none() { - if let Some(ref projs) = query.projections { - let columns: Vec<String> = projs - .iter() - .map(|p| { - p.alias - .clone() - .unwrap_or_else(|| field_ref_name(&p.field)) - }) - .collect(); - if !columns.is_empty() { - post_processing.push(PostProcessing::Project { columns }); - } - } - } - - Ok(LogicalPlan { - source, - nodes, - post_processing, - }) - } -} - -// --------------------------------------------------------------------------- -// Internal helpers -// --------------------------------------------------------------------------- - -/// Resolve VCL modalities to concrete `Modality` values, expanding `All`. -fn resolve_modalities(vcl_mods: &[VclModality]) -> Result<Vec<Modality>, PlannerError> { - let mut result = Vec::new(); - for vm in vcl_mods { - match vm { - VclModality::All => { - // Expand to all six modalities. - result.extend_from_slice(&Modality::ALL); - } - other => { - result.push(vcl_modality_to_planner(other)?); - } - } - } - // Deduplicate while preserving order. - let mut seen = std::collections::HashSet::new(); - result.retain(|m| seen.insert(*m)); - Ok(result) -} - -/// Map a single VCL modality to the planner `Modality` enum. -fn vcl_modality_to_planner(vm: &VclModality) -> Result<Modality, PlannerError> { - match vm { - VclModality::Graph => Ok(Modality::Graph), - VclModality::Vector => Ok(Modality::Vector), - VclModality::Tensor => Ok(Modality::Tensor), - VclModality::Semantic => Ok(Modality::Semantic), - VclModality::Document => Ok(Modality::Document), - VclModality::Temporal => Ok(Modality::Temporal), - VclModality::All => { - // Should have been expanded already. - Err(PlannerError::InvalidConfig( - "All modality should be expanded before individual mapping".to_string(), - )) - } - } -} - -/// Map a VCL source to a planner `QuerySource`. -fn map_source(src: &VclSource) -> Result<QuerySource, PlannerError> { - match src { - VclSource::Octad { .. } => Ok(QuerySource::Octad), - VclSource::Federation { nodes, .. } => Ok(QuerySource::Federation { - nodes: nodes.clone(), - }), - VclSource::Store { modality } => { - let m = vcl_modality_to_planner(modality)?; - Ok(QuerySource::Store { modality: m }) - } - } -} - -/// Flatten a VCL condition tree into per-modality condition lists. -/// -/// Strategy: -/// - Leaf conditions that target a specific modality go only to that node. -/// - Leaf conditions without modality affinity are broadcast to all active nodes. -/// - `And` recursively flattens both branches. -/// - `Or`/`Not` are converted to `Predicate` expressions (the physical executor -/// handles them). They are broadcast to all active modalities. -fn flatten_conditions( - cond: &VclCondition, - active_modalities: &[Modality], - out: &mut HashMap<Modality, Vec<ConditionKind>>, -) -> Result<(), PlannerError> { - match cond { - VclCondition::And(lhs, rhs) => { - flatten_conditions(lhs, active_modalities, out)?; - flatten_conditions(rhs, active_modalities, out)?; - } - VclCondition::Or(lhs, rhs) => { - // OR cannot be trivially split per-modality. Encode as a predicate - // string on all active modalities. - let desc = format!( - "OR({}, {})", - describe_condition(lhs), - describe_condition(rhs) - ); - for &m in active_modalities { - out.entry(m) - .or_default() - .push(ConditionKind::Predicate { expression: desc.clone() }); - } - } - VclCondition::Not(inner) => { - let desc = format!("NOT({})", describe_condition(inner)); - for &m in active_modalities { - out.entry(m) - .or_default() - .push(ConditionKind::Predicate { expression: desc.clone() }); - } - } - VclCondition::Simple(simple) => { - let (target_modality, condition_kind) = map_simple_condition(simple)?; - match target_modality { - Some(m) if out.contains_key(&m) => { - out.entry(m).or_default().push(condition_kind); - } - Some(m) => { - // The targeted modality is not in the active set. - // Add it as a predicate on all active modalities. - let desc = format!("target_modality={m}: {condition_kind:?}"); - for &am in active_modalities { - out.entry(am) - .or_default() - .push(ConditionKind::Predicate { expression: desc.clone() }); - } - } - None => { - // No specific target — broadcast to all active modalities. - for &m in active_modalities { - out.entry(m).or_default().push(condition_kind.clone()); - } - } - } - } - } - Ok(()) -} - -/// Map a VCL simple condition to a `ConditionKind` and an optional target modality. -/// -/// Returns `(target_modality, condition_kind)` where `target_modality` is `Some` -/// if the condition naturally targets a specific modality, `None` if it should be -/// broadcast. -fn map_simple_condition( - simple: &VclSimpleCondition, -) -> Result<(Option<Modality>, ConditionKind), PlannerError> { - match simple { - VclSimpleCondition::FulltextContains(text) => Ok(( - Some(Modality::Document), - ConditionKind::Fulltext { - query: text.clone(), - }, - )), - VclSimpleCondition::VectorSimilar { - embedding, - threshold: _, - } => { - // `Similarity` takes k (number of neighbours). We use the embedding - // length as a proxy; the physical plan executor will use the actual - // embedding. A more refined approach would add a Similarity variant - // that carries the embedding, but we work with existing types. - Ok(( - Some(Modality::Vector), - ConditionKind::Similarity { - k: embedding.len(), - }, - )) - } - VclSimpleCondition::GraphPattern { predicate, depth } => Ok(( - Some(Modality::Graph), - ConditionKind::Traversal { - predicate: predicate.clone(), - depth: *depth, - }, - )), - VclSimpleCondition::FieldCondition { - field, - operator, - value, - } => { - let target = field.modality.as_ref().and_then(|vm| vcl_modality_to_planner(vm).ok()); - let field_name = field.field.clone(); - let value_str = match value { - serde_json::Value::String(s) => s.clone(), - other => other.to_string(), - }; - - let kind = match operator.as_str() { - "=" | "==" | "!=" | "<>" => ConditionKind::Equality { - field: field_name, - value: value_str, - }, - ">" | ">=" | "<" | "<=" | "BETWEEN" => { - // For single-bound range operators, we use low=value, high=value - // and let the physical executor interpret the operator. - ConditionKind::Range { - field: field_name, - low: value_str.clone(), - high: value_str, - } - } - _ => ConditionKind::Predicate { - expression: format!("{field_name} {operator} {value_str}"), - }, - }; - Ok((target, kind)) - } - VclSimpleCondition::CrossModalFieldCompare { - left, - operator, - right, - } => { - let desc = format!( - "cross_modal: {}.{} {} {}.{}", - left.modality - .as_ref() - .map(|m| format!("{m:?}")) - .unwrap_or_else(|| "?".to_string()), - left.field, - operator, - right - .modality - .as_ref() - .map(|m| format!("{m:?}")) - .unwrap_or_else(|| "?".to_string()), - right.field, - ); - Ok(( - None, - ConditionKind::Predicate { expression: desc }, - )) - } - VclSimpleCondition::ModalityDrift { - modality, - threshold, - } => { - let m = vcl_modality_to_planner(modality)?; - let desc = format!("drift({m}) > {threshold}"); - Ok(( - Some(m), - ConditionKind::Predicate { expression: desc }, - )) - } - VclSimpleCondition::ModalityExists(modality) => { - let m = vcl_modality_to_planner(modality)?; - let desc = format!("exists({m})"); - Ok(( - Some(m), - ConditionKind::Predicate { expression: desc }, - )) - } - VclSimpleCondition::ModalityNotExists(modality) => { - let m = vcl_modality_to_planner(modality)?; - let desc = format!("not_exists({m})"); - Ok(( - Some(m), - ConditionKind::Predicate { expression: desc }, - )) - } - VclSimpleCondition::ModalityConsistency { - modalities, - threshold, - } => { - let names: Vec<String> = modalities - .iter() - .filter_map(|vm| vcl_modality_to_planner(vm).ok()) - .map(|m| m.to_string()) - .collect(); - let desc = format!("consistency({}) > {threshold}", names.join(", ")); - Ok(( - None, - ConditionKind::Predicate { expression: desc }, - )) - } - } -} - -/// Produce a human-readable description of a condition (for predicate encoding). -fn describe_condition(cond: &VclCondition) -> String { - match cond { - VclCondition::And(l, r) => { - format!("({} AND {})", describe_condition(l), describe_condition(r)) - } - VclCondition::Or(l, r) => { - format!("({} OR {})", describe_condition(l), describe_condition(r)) - } - VclCondition::Not(inner) => format!("NOT({})", describe_condition(inner)), - VclCondition::Simple(s) => format!("{s:?}"), - } -} - -/// Build a per-modality projection map from VCL projections. -fn build_projection_map( - modalities: &[Modality], - projections: Option<&[VclProjection]>, -) -> HashMap<Modality, Vec<String>> { - let mut map: HashMap<Modality, Vec<String>> = HashMap::new(); - let Some(projs) = projections else { - return map; - }; - - for proj in projs { - let field_name = proj - .alias - .clone() - .unwrap_or_else(|| proj.field.field.clone()); - - if let Some(ref vm) = proj.field.modality { - if let Ok(m) = vcl_modality_to_planner(vm) { - if modalities.contains(&m) { - map.entry(m).or_default().push(field_name); - } - } - } else { - // No modality qualifier — add to all active modalities. - for &m in modalities { - map.entry(m).or_default().push(field_name.clone()); - } - } - } - map -} - -/// Render a `VclFieldRef` as a dotted name string. -fn field_ref_name(f: &VclFieldRef) -> String { - match &f.modality { - Some(m) => format!("{m:?}.{}", f.field), - None => f.field.clone(), - } -} - -// --------------------------------------------------------------------------- -// Tests -// --------------------------------------------------------------------------- - -#[cfg(test)] -mod tests { - use super::*; - - /// Helper: parse JSON string into VclAst and convert to LogicalPlan. - fn parse_and_plan(json: &str) -> Result<LogicalPlan, PlannerError> { - let ast = VclAst::from_json(json)?; - ast.to_logical_plan() - } - - #[test] - fn test_simple_octad_query() { - let json = r#"{ - "TAG": "Query", - "_0": { - "modalities": [{"TAG": "Graph"}, {"TAG": "Document"}], - "source": {"TAG": "Octad", "_0": "abc-123"}, - "where": { - "TAG": "Simple", - "_0": {"TAG": "FulltextContains", "_0": "hello world"} - }, - "projections": null, - "aggregates": null, - "groupBy": null, - "having": null, - "proof": null, - "orderBy": null, - "limit": 10, - "offset": null - } - }"#; - - let plan = parse_and_plan(json).expect("should parse simple octad query"); - - // Source should be Octad. - assert!(matches!(plan.source, QuerySource::Octad)); - - // Two nodes: Document (priority 30) and Graph (priority 40). - assert_eq!(plan.nodes.len(), 2); - assert_eq!(plan.nodes[0].modality, Modality::Document); - assert_eq!(plan.nodes[1].modality, Modality::Graph); - - // Document node should have the fulltext condition. - assert_eq!(plan.nodes[0].conditions.len(), 1); - assert!(matches!( - &plan.nodes[0].conditions[0], - ConditionKind::Fulltext { query } if query == "hello world" - )); - - // Graph node should have no conditions (fulltext targets Document). - assert_eq!(plan.nodes[1].conditions.len(), 0); - - // Post-processing: Limit. - assert!(plan.post_processing.iter().any(|p| matches!(p, PostProcessing::Limit { count: 10 }))); - } - - #[test] - fn test_federation_with_drift_policy() { - let json = r#"{ - "TAG": "Query", - "_0": { - "modalities": [{"TAG": "Vector"}, {"TAG": "Semantic"}], - "source": { - "TAG": "Federation", - "_0": ["node-a", "node-b"], - "_1": "consistent-read" - }, - "where": { - "TAG": "Simple", - "_0": { - "TAG": "ModalityDrift", - "_0": {"TAG": "Vector"}, - "_1": 0.05 - } - }, - "projections": null, - "aggregates": null, - "groupBy": null, - "having": null, - "proof": null, - "orderBy": null, - "limit": null, - "offset": null - } - }"#; - - let plan = parse_and_plan(json).expect("should parse federation query"); - - assert!(matches!( - &plan.source, - QuerySource::Federation { nodes } if nodes.len() == 2 - )); - - // Vector node should have the drift predicate. - let vector_node = plan - .nodes - .iter() - .find(|n| n.modality == Modality::Vector) - .expect("should have vector node"); - assert_eq!(vector_node.conditions.len(), 1); - assert!(matches!( - &vector_node.conditions[0], - ConditionKind::Predicate { expression } if expression.contains("drift") - )); - } - - #[test] - fn test_cross_modal_conditions() { - let json = r#"{ - "TAG": "Query", - "_0": { - "modalities": [{"TAG": "Document"}, {"TAG": "Graph"}], - "source": {"TAG": "Octad", "_0": null}, - "where": { - "TAG": "Simple", - "_0": { - "TAG": "CrossModalFieldCompare", - "_0": {"modality": {"TAG": "Document"}, "field": "title"}, - "_1": "=", - "_2": {"modality": {"TAG": "Graph"}, "field": "label"} - } - }, - "projections": null, - "aggregates": null, - "groupBy": null, - "having": null, - "proof": null, - "orderBy": null, - "limit": null, - "offset": null - } - }"#; - - let plan = parse_and_plan(json).expect("should parse cross-modal condition"); - - // Cross-modal conditions are broadcast to all active modalities. - for node in &plan.nodes { - assert_eq!(node.conditions.len(), 1); - assert!(matches!( - &node.conditions[0], - ConditionKind::Predicate { expression } if expression.contains("cross_modal") - )); - } - } - - #[test] - fn test_aggregation_group_by_order_by() { - let json = r#"{ - "TAG": "Query", - "_0": { - "modalities": [{"TAG": "Document"}], - "source": {"TAG": "Octad", "_0": null}, - "where": null, - "projections": null, - "aggregates": [ - {"function": "COUNT", "field": null, "alias": "total"}, - {"function": "AVG", "field": {"field": "score"}, "alias": "avg_score"} - ], - "groupBy": [{"field": "category"}], - "having": null, - "proof": null, - "orderBy": [ - {"field": {"field": "score"}, "direction": {"TAG": "Desc"}} - ], - "limit": 100, - "offset": 20 - } - }"#; - - let plan = parse_and_plan(json).expect("should parse aggregation query"); - - // Post-processing should contain GroupBy, OrderBy, Limit, and offset - // marker. - let has_group = plan.post_processing.iter().any(|p| { - matches!( - p, - PostProcessing::GroupBy { fields, aggregates } - if fields == &["category".to_string()] - && aggregates.len() == 2 - ) - }); - assert!(has_group, "should have GroupBy post-processing"); - - let has_order = plan.post_processing.iter().any(|p| { - matches!( - p, - PostProcessing::OrderBy { fields } - if fields.len() == 1 && !fields[0].1 // Desc = false - ) - }); - assert!(has_order, "should have OrderBy post-processing"); - - let has_limit = plan - .post_processing - .iter() - .any(|p| matches!(p, PostProcessing::Limit { count: 100 })); - assert!(has_limit, "should have Limit post-processing"); - - let has_offset = plan.post_processing.iter().any(|p| { - matches!( - p, - PostProcessing::Project { columns } if columns.iter().any(|c| c.starts_with("__offset=")) - ) - }); - assert!(has_offset, "should have offset marker in post-processing"); - } - - #[test] - fn test_all_modality_expansion() { - let json = r#"{ - "TAG": "Query", - "_0": { - "modalities": [{"TAG": "All"}], - "source": {"TAG": "Octad", "_0": null}, - "where": null, - "projections": null, - "aggregates": null, - "groupBy": null, - "having": null, - "proof": null, - "orderBy": null, - "limit": null, - "offset": null - } - }"#; - - let plan = parse_and_plan(json).expect("should parse All modality query"); - - // All 6 modalities should be present. - assert_eq!(plan.nodes.len(), 6); - let modalities: Vec<Modality> = plan.nodes.iter().map(|n| n.modality).collect(); - for m in &Modality::ALL { - assert!( - modalities.contains(m), - "missing modality {m:?} after All expansion" - ); - } - - // Nodes should be sorted by execution priority. - assert_eq!(plan.nodes[0].modality, Modality::Temporal); - assert_eq!(plan.nodes[5].modality, Modality::Semantic); - } - - #[test] - fn test_error_on_invalid_tag() { - let json = r#"{ - "TAG": "InvalidStatement", - "_0": {} - }"#; - - let result = VclAst::from_json(json); - assert!(result.is_err(), "should reject unknown TAG"); - } - - #[test] - fn test_error_on_missing_tag() { - let json = r#"{ - "no_tag_field": "oops" - }"#; - - let result = VclAst::from_json(json); - assert!(result.is_err(), "should reject missing TAG"); - } - - #[test] - fn test_compound_and_condition() { - let json = r#"{ - "TAG": "Query", - "_0": { - "modalities": [{"TAG": "Graph"}, {"TAG": "Vector"}], - "source": {"TAG": "Octad", "_0": "some-uuid"}, - "where": { - "TAG": "And", - "_0": { - "TAG": "Simple", - "_0": {"TAG": "FulltextContains", "_0": "search text"} - }, - "_1": { - "TAG": "Simple", - "_0": {"TAG": "VectorSimilar", "_0": [0.1, 0.2], "_1": 0.9} - } - }, - "projections": null, - "aggregates": null, - "groupBy": null, - "having": null, - "proof": [{"proofType": {"TAG": "Citation"}, "contractName": "MyCitationContract"}], - "orderBy": [{"field": {"modality": {"TAG": "Document"}, "field": "name"}, "direction": {"TAG": "Asc"}}], - "limit": 50, - "offset": null - } - }"#; - - let plan = parse_and_plan(json).expect("should parse compound AND query"); - - // Vector node gets VectorSimilar condition. FulltextContains targets - // Document, which is not in the active set, so it becomes a predicate - // on both active modalities. - let vector_node = plan - .nodes - .iter() - .find(|n| n.modality == Modality::Vector) - .expect("should have vector node"); - // Similarity from VectorSimilar + predicate from unmatched FulltextContains - // is wrong — FulltextContains targets Document which is NOT in active set, - // so it broadcasts as a predicate. But actually Vector does get the - // Similarity condition targeted to it. - let has_similarity = vector_node - .conditions - .iter() - .any(|c| matches!(c, ConditionKind::Similarity { .. })); - assert!(has_similarity, "vector node should have similarity condition"); - - // Graph node should NOT have similarity (it targets Vector). - let graph_node = plan - .nodes - .iter() - .find(|n| n.modality == Modality::Graph) - .expect("should have graph node"); - let has_similarity = graph_node - .conditions - .iter() - .any(|c| matches!(c, ConditionKind::Similarity { .. })); - assert!(!has_similarity, "graph node should not have similarity condition"); - - // Proof spec should create a Semantic ProofVerification condition. - // Semantic is not in active modalities (only Graph, Vector), but the - // proof injection adds it to the condition map. Since Semantic is not - // in the node list, it should not appear. - // Actually, proof conditions are added to existing Semantic entry — - // since Semantic is not in modalities, it will be created in the map - // but not become a node. This is correct: proof verification only - // applies when Semantic modality is queried. - } - - #[test] - fn test_or_condition_broadcast() { - let json = r#"{ - "TAG": "Query", - "_0": { - "modalities": [{"TAG": "Document"}, {"TAG": "Graph"}], - "source": {"TAG": "Octad", "_0": null}, - "where": { - "TAG": "Or", - "_0": { - "TAG": "Simple", - "_0": {"TAG": "FulltextContains", "_0": "alpha"} - }, - "_1": { - "TAG": "Simple", - "_0": {"TAG": "GraphPattern", "_0": "relates_to", "_1": 3} - } - }, - "projections": null, - "aggregates": null, - "groupBy": null, - "having": null, - "proof": null, - "orderBy": null, - "limit": null, - "offset": null - } - }"#; - - let plan = parse_and_plan(json).expect("should parse OR condition"); - - // OR is broadcast as a Predicate to all active modalities. - for node in &plan.nodes { - assert_eq!( - node.conditions.len(), - 1, - "{:?} node should have exactly 1 condition (the OR predicate)", - node.modality - ); - assert!( - matches!(&node.conditions[0], ConditionKind::Predicate { expression } if expression.starts_with("OR(")), - "{:?} node condition should be an OR predicate", - node.modality - ); - } - } - - #[test] - fn test_store_source() { - let json = r#"{ - "TAG": "Query", - "_0": { - "modalities": [{"TAG": "Vector"}], - "source": {"TAG": "Store", "_0": {"TAG": "Vector"}}, - "where": null, - "projections": null, - "aggregates": null, - "groupBy": null, - "having": null, - "proof": null, - "orderBy": null, - "limit": null, - "offset": null - } - }"#; - - let plan = parse_and_plan(json).expect("should parse Store source"); - assert!(matches!( - plan.source, - QuerySource::Store { modality: Modality::Vector } - )); - } - - #[test] - fn test_proof_on_semantic_node() { - let json = r#"{ - "TAG": "Query", - "_0": { - "modalities": [{"TAG": "Semantic"}, {"TAG": "Document"}], - "source": {"TAG": "Octad", "_0": null}, - "where": null, - "projections": null, - "aggregates": null, - "groupBy": null, - "having": null, - "proof": [ - {"proofType": {"TAG": "Citation"}, "contractName": "CitContract"}, - {"proofType": {"TAG": "Zkp"}, "contractName": "ZkpContract"} - ], - "orderBy": null, - "limit": null, - "offset": null - } - }"#; - - let plan = parse_and_plan(json).expect("should parse proof specs"); - - let semantic_node = plan - .nodes - .iter() - .find(|n| n.modality == Modality::Semantic) - .expect("should have semantic node"); - - assert_eq!( - semantic_node.conditions.len(), - 2, - "semantic node should have 2 proof verification conditions" - ); - assert!(matches!( - &semantic_node.conditions[0], - ConditionKind::ProofVerification { contract } if contract == "CitContract" - )); - assert!(matches!( - &semantic_node.conditions[1], - ConditionKind::ProofVerification { contract } if contract == "ZkpContract" - )); - } - - #[test] - fn test_field_condition_equality() { - let json = r#"{ - "TAG": "Query", - "_0": { - "modalities": [{"TAG": "Document"}], - "source": {"TAG": "Octad", "_0": null}, - "where": { - "TAG": "Simple", - "_0": { - "TAG": "FieldCondition", - "_0": {"modality": {"TAG": "Document"}, "field": "status"}, - "_1": "=", - "_2": "active" - } - }, - "projections": null, - "aggregates": null, - "groupBy": null, - "having": null, - "proof": null, - "orderBy": null, - "limit": null, - "offset": null - } - }"#; - - let plan = parse_and_plan(json).expect("should parse field condition"); - let doc_node = plan - .nodes - .iter() - .find(|n| n.modality == Modality::Document) - .expect("should have document node"); - - assert_eq!(doc_node.conditions.len(), 1); - assert!(matches!( - &doc_node.conditions[0], - ConditionKind::Equality { field, value } - if field == "status" && value == "active" - )); - } - - #[test] - fn test_field_condition_range() { - let json = r#"{ - "TAG": "Query", - "_0": { - "modalities": [{"TAG": "Temporal"}], - "source": {"TAG": "Octad", "_0": null}, - "where": { - "TAG": "Simple", - "_0": { - "TAG": "FieldCondition", - "_0": {"modality": {"TAG": "Temporal"}, "field": "timestamp"}, - "_1": ">=", - "_2": "2026-01-01" - } - }, - "projections": null, - "aggregates": null, - "groupBy": null, - "having": null, - "proof": null, - "orderBy": null, - "limit": null, - "offset": null - } - }"#; - - let plan = parse_and_plan(json).expect("should parse range field condition"); - let temporal_node = plan - .nodes - .iter() - .find(|n| n.modality == Modality::Temporal) - .expect("should have temporal node"); - - assert_eq!(temporal_node.conditions.len(), 1); - assert!(matches!( - &temporal_node.conditions[0], - ConditionKind::Range { field, .. } if field == "timestamp" - )); - } - - #[test] - fn test_empty_modalities_error() { - let json = r#"{ - "TAG": "Query", - "_0": { - "modalities": [], - "source": {"TAG": "Octad", "_0": null}, - "where": null, - "projections": null, - "aggregates": null, - "groupBy": null, - "having": null, - "proof": null, - "orderBy": null, - "limit": null, - "offset": null - } - }"#; - - let result = parse_and_plan(json); - assert!( - matches!(result, Err(PlannerError::EmptyPlan)), - "should return EmptyPlan error for no modalities" - ); - } - - #[test] - fn test_modality_consistency_condition() { - let json = r#"{ - "TAG": "Query", - "_0": { - "modalities": [{"TAG": "Graph"}, {"TAG": "Vector"}], - "source": {"TAG": "Octad", "_0": null}, - "where": { - "TAG": "Simple", - "_0": { - "TAG": "ModalityConsistency", - "_0": [{"TAG": "Graph"}, {"TAG": "Vector"}], - "_1": 0.95 - } - }, - "projections": null, - "aggregates": null, - "groupBy": null, - "having": null, - "proof": null, - "orderBy": null, - "limit": null, - "offset": null - } - }"#; - - let plan = parse_and_plan(json).expect("should parse consistency condition"); - - // Consistency is broadcast to all active modalities. - for node in &plan.nodes { - assert_eq!(node.conditions.len(), 1); - assert!(matches!( - &node.conditions[0], - ConditionKind::Predicate { expression } if expression.contains("consistency") - )); - } - } - - #[test] - fn test_not_condition() { - let json = r#"{ - "TAG": "Query", - "_0": { - "modalities": [{"TAG": "Document"}], - "source": {"TAG": "Octad", "_0": null}, - "where": { - "TAG": "Not", - "_0": { - "TAG": "Simple", - "_0": {"TAG": "FulltextContains", "_0": "excluded"} - } - }, - "projections": null, - "aggregates": null, - "groupBy": null, - "having": null, - "proof": null, - "orderBy": null, - "limit": null, - "offset": null - } - }"#; - - let plan = parse_and_plan(json).expect("should parse NOT condition"); - let doc_node = &plan.nodes[0]; - assert_eq!(doc_node.conditions.len(), 1); - assert!(matches!( - &doc_node.conditions[0], - ConditionKind::Predicate { expression } if expression.starts_with("NOT(") - )); - } - - #[test] - fn test_deduplication_with_all_and_explicit() { - let json = r#"{ - "TAG": "Query", - "_0": { - "modalities": [{"TAG": "Graph"}, {"TAG": "All"}, {"TAG": "Graph"}], - "source": {"TAG": "Octad", "_0": null}, - "where": null, - "projections": null, - "aggregates": null, - "groupBy": null, - "having": null, - "proof": null, - "orderBy": null, - "limit": null, - "offset": null - } - }"#; - - let plan = parse_and_plan(json).expect("should deduplicate modalities"); - // Should still have exactly 6 unique modalities. - assert_eq!(plan.nodes.len(), 6); - } -} diff --git a/verisimdb/rust-core/verisim-provenance/Cargo.toml b/verisimdb/rust-core/verisim-provenance/Cargo.toml deleted file mode 100644 index c8a6c1a5..00000000 --- a/verisimdb/rust-core/verisim-provenance/Cargo.toml +++ /dev/null @@ -1,28 +0,0 @@ -# SPDX-License-Identifier: MPL-2.0 - -[package] -name = "verisim-provenance" -description = "Provenance modality - origin tracking, lineage chains, and actor trails" -version.workspace = true -edition.workspace = true -authors.workspace = true -license.workspace = true - -[dependencies] -serde.workspace = true -serde_json.workspace = true -thiserror.workspace = true -tracing.workspace = true -async-trait.workspace = true -verisim-storage = { path = "../verisim-storage", optional = true } -tokio.workspace = true -chrono.workspace = true -sha2.workspace = true - -[dev-dependencies] -tempfile = "3" -proptest.workspace = true - -[features] -default = [] -redb-backend = ["verisim-storage/redb-backend"] diff --git a/verisimdb/rust-core/verisim-provenance/src/lib.rs b/verisimdb/rust-core/verisim-provenance/src/lib.rs deleted file mode 100644 index f424d46a..00000000 --- a/verisimdb/rust-core/verisim-provenance/src/lib.rs +++ /dev/null @@ -1,640 +0,0 @@ -// SPDX-License-Identifier: MPL-2.0 -//! VeriSim Provenance Modality -//! -//! Tracks the origin, transformation history, and actor trail for each entity. -//! Provenance records form a hash chain — each record's `parent_hash` is the -//! SHA-256 digest of the previous record, creating an immutable audit trail. -//! -//! # Architecture -//! -//! - **ProvenanceRecord**: A single event in an entity's lineage (creation, -//! modification, import, normalization, etc.). -//! - **ProvenanceChain**: An ordered sequence of records with hash-chain -//! integrity verification. -//! - **ProvenanceStore** trait: Async storage interface for recording and -//! querying provenance data. -//! - **InMemoryProvenanceStore**: Reference implementation backed by a -//! `HashMap<String, Vec<ProvenanceRecord>>`. - -#![forbid(unsafe_code)] -#[cfg(feature = "redb-backend")] -pub mod persistent; -#[cfg(feature = "redb-backend")] -pub use persistent::*; -use async_trait::async_trait; -use chrono::{DateTime, Utc}; -use serde::{Deserialize, Serialize}; -use sha2::{Digest, Sha256}; -use std::collections::HashMap; -use std::sync::Arc; -use thiserror::Error; -use tokio::sync::RwLock; -use tracing::{debug, instrument}; - -/// Provenance-specific errors -#[derive(Error, Debug)] -pub enum ProvenanceError { - /// Entity provenance chain not found - #[error("Provenance chain not found for entity: {0}")] - NotFound(String), - - /// Hash chain integrity violation — records have been tampered with or - /// a parent_hash does not match the SHA-256 of the preceding record - #[error("Provenance chain corrupted for entity {entity}: {reason}")] - ChainCorrupted { - entity: String, - reason: String, - }, - - /// A record's computed hash does not match its stored content_hash - #[error("Hash mismatch at index {index} for entity {entity}")] - HashMismatch { - entity: String, - index: usize, - }, - - /// Generic I/O or storage error - #[error("Provenance I/O error: {0}")] - IoError(String), -} - -/// Classification of provenance events -/// -/// Each variant captures a distinct lifecycle transition that an entity can -/// undergo. The `Custom(String)` variant allows domain-specific extensions -/// without modifying this enum. -#[derive(Debug, Clone, Serialize, Deserialize, PartialEq, Eq, Hash)] -pub enum ProvenanceEventType { - /// Entity was created for the first time - Created, - /// Entity was modified (content or metadata changed) - Modified, - /// Entity was imported from an external source - Imported, - /// Entity underwent drift normalization - Normalized, - /// Entity was repaired after drift detection - DriftRepaired, - /// Entity was soft- or hard-deleted - Deleted, - /// Two or more entities were merged into this one - Merged, - /// Domain-specific event type - Custom(String), -} - -impl std::fmt::Display for ProvenanceEventType { - fn fmt(&self, f: &mut std::fmt::Formatter<'_>) -> std::fmt::Result { - match self { - ProvenanceEventType::Created => write!(f, "created"), - ProvenanceEventType::Modified => write!(f, "modified"), - ProvenanceEventType::Imported => write!(f, "imported"), - ProvenanceEventType::Normalized => write!(f, "normalized"), - ProvenanceEventType::DriftRepaired => write!(f, "drift_repaired"), - ProvenanceEventType::Deleted => write!(f, "deleted"), - ProvenanceEventType::Merged => write!(f, "merged"), - ProvenanceEventType::Custom(name) => write!(f, "custom:{}", name), - } - } -} - -/// A single provenance record — one event in an entity's lineage chain. -/// -/// Records are linked by `parent_hash`: the SHA-256 of the serialized -/// previous record. The first record in a chain has `parent_hash` set to -/// the SHA-256 of the empty string. -#[derive(Debug, Clone, Serialize, Deserialize)] -pub struct ProvenanceRecord { - /// What happened to the entity - pub event_type: ProvenanceEventType, - /// Who or what caused this event (user ID, system component, bot name) - pub actor: String, - /// When this event occurred - pub timestamp: DateTime<Utc>, - /// Optional source identifier (URL, file path, upstream entity ID) - pub source: Option<String>, - /// Human-readable description of the event - pub description: String, - /// SHA-256 hex digest of the previous record (or of "" for the first) - pub parent_hash: String, - /// SHA-256 hex digest of this record's canonical serialization - pub content_hash: String, -} - -impl ProvenanceRecord { - /// Compute the SHA-256 hex digest of the canonical (deterministic) JSON - /// serialization of a record's *content* fields — everything except - /// `content_hash` itself. - pub fn compute_hash( - event_type: &ProvenanceEventType, - actor: &str, - timestamp: &DateTime<Utc>, - source: &Option<String>, - description: &str, - parent_hash: &str, - ) -> String { - let canonical = serde_json::json!({ - "event_type": event_type, - "actor": actor, - "timestamp": timestamp.to_rfc3339(), - "source": source, - "description": description, - "parent_hash": parent_hash, - }); - let bytes = canonical.to_string().into_bytes(); - let digest = Sha256::digest(&bytes); - format!("{:x}", digest) - } - - /// Build a new record, computing its `content_hash` automatically. - pub fn new( - event_type: ProvenanceEventType, - actor: impl Into<String>, - source: Option<String>, - description: impl Into<String>, - parent_hash: impl Into<String>, - ) -> Self { - let actor = actor.into(); - let description = description.into(); - let parent_hash = parent_hash.into(); - let timestamp = Utc::now(); - - let content_hash = Self::compute_hash( - &event_type, - &actor, - ×tamp, - &source, - &description, - &parent_hash, - ); - - Self { - event_type, - actor, - timestamp, - source, - description, - parent_hash, - content_hash, - } - } - - /// Verify that `content_hash` matches the re-computed hash of this - /// record's fields. - pub fn verify(&self) -> bool { - let expected = Self::compute_hash( - &self.event_type, - &self.actor, - &self.timestamp, - &self.source, - &self.description, - &self.parent_hash, - ); - self.content_hash == expected - } -} - -/// An ordered provenance chain for a single entity. -/// -/// The chain is a `Vec<ProvenanceRecord>` where each record's -/// `parent_hash` equals the `content_hash` of its predecessor. -#[derive(Debug, Clone, Serialize, Deserialize)] -pub struct ProvenanceChain { - /// The entity this chain belongs to - pub entity_id: String, - /// Ordered list of provenance records (oldest first) - pub records: Vec<ProvenanceRecord>, -} - -impl ProvenanceChain { - /// Create an empty chain for a given entity. - pub fn new(entity_id: impl Into<String>) -> Self { - Self { - entity_id: entity_id.into(), - records: Vec::new(), - } - } - - /// Number of records in the chain. - pub fn len(&self) -> usize { - self.records.len() - } - - /// Whether the chain is empty. - pub fn is_empty(&self) -> bool { - self.records.is_empty() - } - - /// Get the hash of the genesis record (SHA-256 of ""). - fn genesis_hash() -> String { - let digest = Sha256::digest(b""); - format!("{:x}", digest) - } - - /// Verify the entire hash chain. - /// - /// Returns `Ok(())` if every record's `parent_hash` matches the - /// `content_hash` of the previous record (or the genesis hash for - /// the first), and every record's `content_hash` re-computes correctly. - pub fn verify(&self) -> Result<(), ProvenanceError> { - let mut expected_parent = Self::genesis_hash(); - - for (i, record) in self.records.iter().enumerate() { - // Check parent linkage - if record.parent_hash != expected_parent { - return Err(ProvenanceError::ChainCorrupted { - entity: self.entity_id.clone(), - reason: format!( - "Record {} parent_hash mismatch: expected {}, got {}", - i, expected_parent, record.parent_hash - ), - }); - } - - // Check record self-integrity - if !record.verify() { - return Err(ProvenanceError::HashMismatch { - entity: self.entity_id.clone(), - index: i, - }); - } - - expected_parent = record.content_hash.clone(); - } - - Ok(()) - } - - /// Get the origin (first) record, if any. - pub fn origin(&self) -> Option<&ProvenanceRecord> { - self.records.first() - } - - /// Get the latest (most recent) record, if any. - pub fn latest(&self) -> Option<&ProvenanceRecord> { - self.records.last() - } - - /// Append a new record to the chain. - /// - /// The `parent_hash` is set automatically from the previous record's - /// `content_hash` (or the genesis hash if this is the first record). - pub fn append( - &mut self, - event_type: ProvenanceEventType, - actor: impl Into<String>, - source: Option<String>, - description: impl Into<String>, - ) -> &ProvenanceRecord { - let parent_hash = self - .records - .last() - .map(|r| r.content_hash.clone()) - .unwrap_or_else(Self::genesis_hash); - - let record = ProvenanceRecord::new(event_type, actor, source, description, parent_hash); - self.records.push(record); - self.records.last().expect("TODO: handle error") - } -} - -/// Async trait for provenance storage backends. -/// -/// Implementations must be `Send + Sync` so they can be shared across -/// Tokio tasks. -#[async_trait] -pub trait ProvenanceStore: Send + Sync { - /// Record a new provenance event for an entity. - /// - /// If the entity has no existing chain, one is created with this as the - /// genesis record. Returns the newly appended record. - async fn record_event( - &self, - entity_id: &str, - event_type: ProvenanceEventType, - actor: &str, - source: Option<String>, - description: &str, - ) -> Result<ProvenanceRecord, ProvenanceError>; - - /// Retrieve the full provenance chain for an entity. - async fn get_chain(&self, entity_id: &str) -> Result<ProvenanceChain, ProvenanceError>; - - /// Verify the hash-chain integrity for an entity. - /// - /// Returns `Ok(true)` if the chain is valid, `Ok(false)` if the entity - /// has no chain, or `Err` if the chain is corrupted. - async fn verify_chain(&self, entity_id: &str) -> Result<bool, ProvenanceError>; - - /// Get the origin (first) record for an entity. - async fn get_origin(&self, entity_id: &str) -> Result<Option<ProvenanceRecord>, ProvenanceError>; - - /// Get the latest (most recent) record for an entity. - async fn get_latest(&self, entity_id: &str) -> Result<Option<ProvenanceRecord>, ProvenanceError>; - - /// Search for provenance records by actor across all entities. - async fn search_by_actor(&self, actor: &str) -> Result<Vec<(String, ProvenanceRecord)>, ProvenanceError>; - - /// Delete the provenance chain for an entity (for testing / admin use). - async fn delete_chain(&self, entity_id: &str) -> Result<(), ProvenanceError>; -} - -/// In-memory implementation of [`ProvenanceStore`]. -/// -/// Suitable for development, testing, and single-node deployments. -/// All data is lost on process exit. -pub struct InMemoryProvenanceStore { - chains: Arc<RwLock<HashMap<String, ProvenanceChain>>>, -} - -impl InMemoryProvenanceStore { - /// Create a new empty in-memory provenance store. - pub fn new() -> Self { - Self { - chains: Arc::new(RwLock::new(HashMap::new())), - } - } -} - -impl Default for InMemoryProvenanceStore { - fn default() -> Self { - Self::new() - } -} - -#[async_trait] -impl ProvenanceStore for InMemoryProvenanceStore { - #[instrument(skip(self))] - async fn record_event( - &self, - entity_id: &str, - event_type: ProvenanceEventType, - actor: &str, - source: Option<String>, - description: &str, - ) -> Result<ProvenanceRecord, ProvenanceError> { - let mut chains = self.chains.write().await; - let chain = chains - .entry(entity_id.to_string()) - .or_insert_with(|| ProvenanceChain::new(entity_id)); - - chain.append(event_type, actor, source, description); - - let record = chain.records.last().expect("TODO: handle error").clone(); - debug!( - entity_id = %entity_id, - event = %record.event_type, - actor = %record.actor, - chain_length = chain.len(), - "Provenance event recorded" - ); - Ok(record) - } - - async fn get_chain(&self, entity_id: &str) -> Result<ProvenanceChain, ProvenanceError> { - let chains = self.chains.read().await; - chains - .get(entity_id) - .cloned() - .ok_or_else(|| ProvenanceError::NotFound(entity_id.to_string())) - } - - async fn verify_chain(&self, entity_id: &str) -> Result<bool, ProvenanceError> { - let chains = self.chains.read().await; - match chains.get(entity_id) { - Some(chain) => { - chain.verify()?; - Ok(true) - } - None => Ok(false), - } - } - - async fn get_origin(&self, entity_id: &str) -> Result<Option<ProvenanceRecord>, ProvenanceError> { - let chains = self.chains.read().await; - Ok(chains.get(entity_id).and_then(|c| c.origin().cloned())) - } - - async fn get_latest(&self, entity_id: &str) -> Result<Option<ProvenanceRecord>, ProvenanceError> { - let chains = self.chains.read().await; - Ok(chains.get(entity_id).and_then(|c| c.latest().cloned())) - } - - async fn search_by_actor(&self, actor: &str) -> Result<Vec<(String, ProvenanceRecord)>, ProvenanceError> { - let chains = self.chains.read().await; - let mut results = Vec::new(); - for (entity_id, chain) in chains.iter() { - for record in &chain.records { - if record.actor == actor { - results.push((entity_id.clone(), record.clone())); - } - } - } - Ok(results) - } - - async fn delete_chain(&self, entity_id: &str) -> Result<(), ProvenanceError> { - let mut chains = self.chains.write().await; - chains.remove(entity_id); - Ok(()) - } -} - -#[cfg(test)] -mod tests { - use super::*; - - #[test] - fn test_provenance_record_hash_verification() { - let record = ProvenanceRecord::new( - ProvenanceEventType::Created, - "alice", - Some("https://source.example.com".to_string()), - "Initial creation of entity", - "0000000000000000", - ); - assert!(record.verify(), "Freshly created record should verify"); - } - - #[test] - fn test_provenance_record_tampered_fails_verification() { - let mut record = ProvenanceRecord::new( - ProvenanceEventType::Created, - "alice", - None, - "Initial creation", - "0000000000000000", - ); - // Tamper with the description after creation - record.description = "TAMPERED".to_string(); - assert!(!record.verify(), "Tampered record should fail verification"); - } - - #[test] - fn test_provenance_chain_integrity() { - let mut chain = ProvenanceChain::new("entity-1"); - chain.append(ProvenanceEventType::Created, "alice", None, "Created entity"); - chain.append(ProvenanceEventType::Modified, "bob", None, "Updated title"); - chain.append( - ProvenanceEventType::Normalized, - "system", - None, - "Auto-normalized after drift detection", - ); - - assert_eq!(chain.len(), 3); - assert!(chain.verify().is_ok(), "Valid chain should verify"); - } - - #[test] - fn test_provenance_chain_corruption_detected() { - let mut chain = ProvenanceChain::new("entity-2"); - chain.append(ProvenanceEventType::Created, "alice", None, "Created"); - chain.append(ProvenanceEventType::Modified, "bob", None, "Modified"); - - // Corrupt the second record's parent_hash - chain.records[1].parent_hash = "corrupted_hash".to_string(); - - let result = chain.verify(); - assert!(result.is_err(), "Corrupted chain should fail verification"); - match result { - Err(ProvenanceError::ChainCorrupted { entity, .. }) => { - assert_eq!(entity, "entity-2"); - } - other => panic!("Expected ChainCorrupted, got {:?}", other), - } - } - - #[test] - fn test_provenance_chain_origin_and_latest() { - let mut chain = ProvenanceChain::new("entity-3"); - assert!(chain.origin().is_none()); - assert!(chain.latest().is_none()); - - chain.append(ProvenanceEventType::Created, "alice", None, "Created"); - chain.append(ProvenanceEventType::Modified, "bob", None, "Modified"); - - assert_eq!(chain.origin().expect("TODO: handle error").actor, "alice"); - assert_eq!(chain.latest().expect("TODO: handle error").actor, "bob"); - } - - #[tokio::test] - async fn test_in_memory_store_record_and_get() { - let store = InMemoryProvenanceStore::new(); - - // Record first event - let record = store - .record_event( - "entity-100", - ProvenanceEventType::Created, - "alice", - Some("https://import.example.com".to_string()), - "Imported from external source", - ) - .await - .expect("TODO: handle error"); - assert_eq!(record.event_type, ProvenanceEventType::Created); - - // Record second event - store - .record_event( - "entity-100", - ProvenanceEventType::Modified, - "bob", - None, - "Updated vector embedding", - ) - .await - .expect("TODO: handle error"); - - // Retrieve chain - let chain = store.get_chain("entity-100").await.expect("TODO: handle error"); - assert_eq!(chain.len(), 2); - assert!(chain.verify().is_ok()); - } - - #[tokio::test] - async fn test_in_memory_store_verify_chain() { - let store = InMemoryProvenanceStore::new(); - - // Non-existent chain returns false (not an error) - assert!(!store.verify_chain("no-such-entity").await.expect("TODO: handle error")); - - // Create a chain and verify it - store - .record_event("e1", ProvenanceEventType::Created, "alice", None, "Created") - .await - .expect("TODO: handle error"); - store - .record_event("e1", ProvenanceEventType::Modified, "bob", None, "Modified") - .await - .expect("TODO: handle error"); - - assert!(store.verify_chain("e1").await.expect("TODO: handle error")); - } - - #[tokio::test] - async fn test_in_memory_store_search_by_actor() { - let store = InMemoryProvenanceStore::new(); - - store - .record_event("e1", ProvenanceEventType::Created, "alice", None, "Created e1") - .await - .expect("TODO: handle error"); - store - .record_event("e2", ProvenanceEventType::Created, "bob", None, "Created e2") - .await - .expect("TODO: handle error"); - store - .record_event("e3", ProvenanceEventType::Imported, "alice", None, "Imported e3") - .await - .expect("TODO: handle error"); - - let alice_records = store.search_by_actor("alice").await.expect("TODO: handle error"); - assert_eq!(alice_records.len(), 2); - - let bob_records = store.search_by_actor("bob").await.expect("TODO: handle error"); - assert_eq!(bob_records.len(), 1); - } - - #[tokio::test] - async fn test_in_memory_store_origin_and_latest() { - let store = InMemoryProvenanceStore::new(); - - store - .record_event("e1", ProvenanceEventType::Created, "alice", None, "Created") - .await - .expect("TODO: handle error"); - store - .record_event("e1", ProvenanceEventType::Modified, "bob", None, "Modified") - .await - .expect("TODO: handle error"); - - let origin = store.get_origin("e1").await.expect("TODO: handle error").expect("TODO: handle error"); - assert_eq!(origin.actor, "alice"); - - let latest = store.get_latest("e1").await.expect("TODO: handle error").expect("TODO: handle error"); - assert_eq!(latest.actor, "bob"); - } - - #[tokio::test] - async fn test_in_memory_store_delete_chain() { - let store = InMemoryProvenanceStore::new(); - - store - .record_event("e1", ProvenanceEventType::Created, "alice", None, "Created") - .await - .expect("TODO: handle error"); - assert!(store.get_chain("e1").await.is_ok()); - - store.delete_chain("e1").await.expect("TODO: handle error"); - assert!(store.get_chain("e1").await.is_err()); - } - - #[tokio::test] - async fn test_in_memory_store_not_found() { - let store = InMemoryProvenanceStore::new(); - let result = store.get_chain("nonexistent").await; - assert!(matches!(result, Err(ProvenanceError::NotFound(_)))); - } -} diff --git a/verisimdb/rust-core/verisim-provenance/src/persistent.rs b/verisimdb/rust-core/verisim-provenance/src/persistent.rs deleted file mode 100644 index b7472e5c..00000000 --- a/verisimdb/rust-core/verisim-provenance/src/persistent.rs +++ /dev/null @@ -1,276 +0,0 @@ -// SPDX-License-Identifier: MPL-2.0 -// Copyright (c) 2026 Jonathan D.A. Jewell (hyperpolymath) <j.d.a.jewell@open.ac.uk> -// -// Persistent provenance store backed by redb via verisim-storage. -// -// Each entity's ProvenanceChain is stored as a single JSON blob keyed by -// entity_id. On open(), all chains are scanned into an in-memory cache -// protected by a tokio::sync::RwLock (matching the InMemory implementation). -// Writes go to redb first, then update the cache. - -use std::collections::HashMap; -use std::path::Path; -use std::sync::Arc; - -use async_trait::async_trait; -use tracing::{debug, info, instrument}; -use verisim_storage::redb_backend::RedbBackend; -use verisim_storage::typed::TypedStore; - -use crate::{ - ProvenanceChain, ProvenanceError, ProvenanceEventType, ProvenanceRecord, ProvenanceStore, -}; - -/// Persistent provenance store: redb for durability, async RwLock cache for -/// fast reads. -/// -/// The cache uses `tokio::sync::RwLock` to match the async locking pattern of -/// `InMemoryProvenanceStore`. -pub struct RedbProvenanceStore { - /// Typed store for provenance chains, keyed by entity_id. - store: TypedStore<RedbBackend>, - /// In-memory cache of all provenance chains. - chains: Arc<tokio::sync::RwLock<HashMap<String, ProvenanceChain>>>, -} - -impl RedbProvenanceStore { - /// Open (or create) a persistent provenance store at the given path. - /// - /// On open, all existing chains are scanned from redb into the in-memory - /// cache so that reads never hit disk. - pub async fn open(path: impl AsRef<Path>) -> Result<Self, ProvenanceError> { - let backend = RedbBackend::open(path.as_ref()) - .map_err(|e| ProvenanceError::IoError(format!("redb open: {}", e)))?; - let store = TypedStore::new(backend, "prov"); - - let entries: Vec<(String, ProvenanceChain)> = store - .scan_prefix("", 1_000_000) - .await - .map_err(|e| ProvenanceError::IoError(format!("scan: {}", e)))?; - - let mut cache = HashMap::new(); - for (id, chain) in entries { - cache.insert(id, chain); - } - - info!(count = cache.len(), "Loaded provenance store from redb"); - Ok(Self { - store, - chains: Arc::new(tokio::sync::RwLock::new(cache)), - }) - } - - /// Persist a single entity's chain to redb. - async fn persist_chain( - &self, - entity_id: &str, - chain: &ProvenanceChain, - ) -> Result<(), ProvenanceError> { - self.store - .put(entity_id, chain) - .await - .map_err(|e| ProvenanceError::IoError(format!("put: {}", e))) - } -} - -#[async_trait] -impl ProvenanceStore for RedbProvenanceStore { - #[instrument(skip(self))] - async fn record_event( - &self, - entity_id: &str, - event_type: ProvenanceEventType, - actor: &str, - source: Option<String>, - description: &str, - ) -> Result<ProvenanceRecord, ProvenanceError> { - let mut chains = self.chains.write().await; - let chain = chains - .entry(entity_id.to_string()) - .or_insert_with(|| ProvenanceChain::new(entity_id)); - - chain.append(event_type, actor, source, description); - let record = chain.records.last().expect("TODO: handle error").clone(); - - // Persist the updated chain to redb. - self.persist_chain(entity_id, chain).await?; - - debug!( - entity_id = %entity_id, - event = %record.event_type, - actor = %record.actor, - chain_length = chain.len(), - "Provenance event recorded (persistent)" - ); - Ok(record) - } - - async fn get_chain(&self, entity_id: &str) -> Result<ProvenanceChain, ProvenanceError> { - let chains = self.chains.read().await; - chains - .get(entity_id) - .cloned() - .ok_or_else(|| ProvenanceError::NotFound(entity_id.to_string())) - } - - async fn verify_chain(&self, entity_id: &str) -> Result<bool, ProvenanceError> { - let chains = self.chains.read().await; - match chains.get(entity_id) { - Some(chain) => { - chain.verify()?; - Ok(true) - } - None => Ok(false), - } - } - - async fn get_origin( - &self, - entity_id: &str, - ) -> Result<Option<ProvenanceRecord>, ProvenanceError> { - let chains = self.chains.read().await; - Ok(chains.get(entity_id).and_then(|c| c.origin().cloned())) - } - - async fn get_latest( - &self, - entity_id: &str, - ) -> Result<Option<ProvenanceRecord>, ProvenanceError> { - let chains = self.chains.read().await; - Ok(chains.get(entity_id).and_then(|c| c.latest().cloned())) - } - - async fn search_by_actor( - &self, - actor: &str, - ) -> Result<Vec<(String, ProvenanceRecord)>, ProvenanceError> { - let chains = self.chains.read().await; - let mut results = Vec::new(); - for (entity_id, chain) in chains.iter() { - for record in &chain.records { - if record.actor == actor { - results.push((entity_id.clone(), record.clone())); - } - } - } - Ok(results) - } - - async fn delete_chain(&self, entity_id: &str) -> Result<(), ProvenanceError> { - // Delete from redb first. - self.store - .delete(entity_id) - .await - .map_err(|e| ProvenanceError::IoError(format!("delete: {}", e)))?; - // Then remove from cache. - let mut chains = self.chains.write().await; - chains.remove(entity_id); - Ok(()) - } -} - -#[cfg(test)] -mod tests { - use super::*; - - #[tokio::test] - async fn test_persistent_provenance_roundtrip() { - let dir = tempfile::tempdir().expect("TODO: handle error"); - let path = dir.path().join("prov.redb"); - - // Write data in one session. - { - let store = RedbProvenanceStore::open(&path).await.expect("TODO: handle error"); - store - .record_event( - "entity-1", - ProvenanceEventType::Created, - "alice", - Some("https://source.example.com".to_string()), - "Initial creation", - ) - .await - .expect("TODO: handle error"); - store - .record_event( - "entity-1", - ProvenanceEventType::Modified, - "bob", - None, - "Updated vector embedding", - ) - .await - .expect("TODO: handle error"); - } - - // Reopen and verify data survived. - { - let store = RedbProvenanceStore::open(&path).await.expect("TODO: handle error"); - - let chain = store.get_chain("entity-1").await.expect("TODO: handle error"); - assert_eq!(chain.len(), 2); - assert!(chain.verify().is_ok()); - - let origin = store.get_origin("entity-1").await.expect("TODO: handle error").expect("TODO: handle error"); - assert_eq!(origin.actor, "alice"); - assert_eq!(origin.event_type, ProvenanceEventType::Created); - - let latest = store.get_latest("entity-1").await.expect("TODO: handle error").expect("TODO: handle error"); - assert_eq!(latest.actor, "bob"); - assert_eq!(latest.event_type, ProvenanceEventType::Modified); - - // Verify chain integrity - assert!(store.verify_chain("entity-1").await.expect("TODO: handle error")); - - // Non-existent entity returns false, not error - assert!(!store.verify_chain("no-such-entity").await.expect("TODO: handle error")); - } - } - - #[tokio::test] - async fn test_persistent_provenance_search_by_actor() { - let dir = tempfile::tempdir().expect("TODO: handle error"); - let path = dir.path().join("prov-search.redb"); - - let store = RedbProvenanceStore::open(&path).await.expect("TODO: handle error"); - store - .record_event("e1", ProvenanceEventType::Created, "alice", None, "Created e1") - .await - .expect("TODO: handle error"); - store - .record_event("e2", ProvenanceEventType::Created, "bob", None, "Created e2") - .await - .expect("TODO: handle error"); - store - .record_event( - "e3", - ProvenanceEventType::Imported, - "alice", - None, - "Imported e3", - ) - .await - .expect("TODO: handle error"); - - let alice_records = store.search_by_actor("alice").await.expect("TODO: handle error"); - assert_eq!(alice_records.len(), 2); - - let bob_records = store.search_by_actor("bob").await.expect("TODO: handle error"); - assert_eq!(bob_records.len(), 1); - } - - #[tokio::test] - async fn test_persistent_provenance_delete_chain() { - let dir = tempfile::tempdir().expect("TODO: handle error"); - let path = dir.path().join("prov-delete.redb"); - - let store = RedbProvenanceStore::open(&path).await.expect("TODO: handle error"); - store - .record_event("e1", ProvenanceEventType::Created, "alice", None, "Created") - .await - .expect("TODO: handle error"); - - store.delete_chain("e1").await.expect("TODO: handle error"); - assert!(store.get_chain("e1").await.is_err()); - } -} diff --git a/verisimdb/rust-core/verisim-repl/Cargo.toml b/verisimdb/rust-core/verisim-repl/Cargo.toml deleted file mode 100644 index 52bfa46a..00000000 --- a/verisimdb/rust-core/verisim-repl/Cargo.toml +++ /dev/null @@ -1,32 +0,0 @@ -# SPDX-License-Identifier: MPL-2.0 - -[package] -name = "verisim-repl" -description = "Interactive VCL REPL for VeriSimDB" -version.workspace = true -edition.workspace = true -authors.workspace = true -license.workspace = true - -[[bin]] -name = "vcl" -path = "src/main.rs" - -[dependencies] -# HTTP client for verisim-api communication -reqwest = { workspace = true, features = ["json", "blocking"] } -rustls.workspace = true -serde.workspace = true -serde_json.workspace = true - -# REPL framework and CLI -rustyline = "15" -rustyline-derive = "0.11" -clap = { version = "4", features = ["derive"] } - -# Output formatting -comfy-table = "7" -colored = "3" - -# Home directory resolution (history file, .vclrc) -dirs = "6" diff --git a/verisimdb/rust-core/verisim-repl/src/client.rs b/verisimdb/rust-core/verisim-repl/src/client.rs deleted file mode 100644 index 4e1006de..00000000 --- a/verisimdb/rust-core/verisim-repl/src/client.rs +++ /dev/null @@ -1,152 +0,0 @@ -// SPDX-License-Identifier: MPL-2.0 -// Copyright (c) 2026 Jonathan D.A. Jewell (hyperpolymath) <j.d.a.jewell@open.ac.uk> -//! -//! HTTP client for communicating with the verisim-api server. -//! -//! Wraps `reqwest::blocking::Client` and provides typed methods for -//! VCL query execution, EXPLAIN output, and health checks. - -use reqwest::blocking::Client; -use serde_json::Value; - -/// Error type for VCL client operations. -#[derive(Debug)] -pub enum ClientError { - /// HTTP transport or connection error. - Http(reqwest::Error), - /// Server returned a non-success status code. - Server { status: u16, body: String }, - /// Failed to parse response body as JSON. - Parse(String), -} - -impl std::fmt::Display for ClientError { - fn fmt(&self, f: &mut std::fmt::Formatter<'_>) -> std::fmt::Result { - match self { - ClientError::Http(err) => write!(f, "Connection error: {err}"), - ClientError::Server { status, body } => { - write!(f, "Server error (HTTP {status}): {body}") - } - ClientError::Parse(msg) => write!(f, "Parse error: {msg}"), - } - } -} - -impl std::error::Error for ClientError {} - -impl From<reqwest::Error> for ClientError { - fn from(err: reqwest::Error) -> Self { - ClientError::Http(err) - } -} - -/// HTTP client for the VeriSimDB API. -/// -/// All methods use blocking I/O so they can be called directly from the -/// synchronous REPL loop without an async runtime. -pub struct VclClient { - /// Base URL of the verisim-api server (e.g. `http://localhost:8080`). - base_url: String, - /// Underlying HTTP client (connection-pooled). - http: Client, -} - -impl VclClient { - /// Create a new client pointing at the given base URL. - /// - /// The URL should include the scheme and port but no trailing slash. - /// Example: `http://localhost:8080` - pub fn new(base_url: &str) -> Self { - // Install ring as the default crypto provider (pure Rust, no OpenSSL/aws-lc-sys) - let _ = rustls::crypto::ring::default_provider().install_default(); - - let base_url = base_url.trim_end_matches('/').to_string(); - let http = Client::builder() - .timeout(std::time::Duration::from_secs(30)) - .build() - .expect("failed to build HTTP client"); - Self { base_url, http } - } - - /// Return the current base URL. - pub fn base_url(&self) -> &str { - &self.base_url - } - - /// Execute a VCL query string. - /// - /// Sends `POST /vcl/execute` with body `{"query": "<vcl>"}`. - /// Returns the raw JSON response from the server. - pub fn execute(&self, query: &str) -> Result<Value, ClientError> { - let url = format!("{}/vcl/execute", self.base_url); - let payload = serde_json::json!({ "query": query }); - - let response = self.http.post(&url).json(&payload).send()?; - self.handle_response(response) - } - - /// Request EXPLAIN output for a VCL query. - /// - /// Sends the query through the VCL execute endpoint with an EXPLAIN - /// prefix, which returns a query plan without executing the query. - pub fn explain(&self, query: &str) -> Result<Value, ClientError> { - let url = format!("{}/vcl/execute", self.base_url); - let explain_query = format!("EXPLAIN {}", query); - let payload = serde_json::json!({ "query": explain_query }); - - let response = self.http.post(&url).json(&payload).send()?; - self.handle_response(response) - } - - /// Check server health. - /// - /// Sends `GET /health` and returns the health response JSON. - pub fn health(&self) -> Result<Value, ClientError> { - let url = format!("{}/health", self.base_url); - let response = self.http.get(&url).send()?; - self.handle_response(response) - } - - /// Parse an HTTP response into a `serde_json::Value`. - /// - /// Returns `ClientError::Server` for non-2xx status codes, and - /// `ClientError::Parse` if the body is not valid JSON. - fn handle_response(&self, response: reqwest::blocking::Response) -> Result<Value, ClientError> { - let status = response.status().as_u16(); - let body = response.text()?; - - if !(200..300).contains(&status) { - return Err(ClientError::Server { status, body }); - } - - serde_json::from_str(&body).map_err(|e| ClientError::Parse(e.to_string())) - } -} - -#[cfg(test)] -mod tests { - use super::*; - - #[test] - fn test_client_creation() { - let client = VclClient::new("http://localhost:8080"); - assert_eq!(client.base_url(), "http://localhost:8080"); - } - - #[test] - fn test_trailing_slash_stripped() { - let client = VclClient::new("http://localhost:8080/"); - assert_eq!(client.base_url(), "http://localhost:8080"); - } - - #[test] - fn test_client_error_display() { - let err = ClientError::Server { - status: 404, - body: "not found".to_string(), - }; - let msg = format!("{err}"); - assert!(msg.contains("404")); - assert!(msg.contains("not found")); - } -} diff --git a/verisimdb/rust-core/verisim-repl/src/completer.rs b/verisimdb/rust-core/verisim-repl/src/completer.rs deleted file mode 100644 index 57f85138..00000000 --- a/verisimdb/rust-core/verisim-repl/src/completer.rs +++ /dev/null @@ -1,178 +0,0 @@ -// SPDX-License-Identifier: MPL-2.0 -// Copyright (c) 2026 Jonathan D.A. Jewell (hyperpolymath) <j.d.a.jewell@open.ac.uk> -//! -//! Tab-completion for the VCL REPL. -//! -//! Provides context-aware completion for: -//! - VCL keywords (SELECT, FROM, WHERE, PROOF, LIMIT, etc.) -//! - Modality names (GRAPH, VECTOR, TENSOR, SEMANTIC, DOCUMENT, TEMPORAL) -//! - Meta-commands (\\connect, \\explain, \\format, etc.) - -use rustyline::completion::{Completer, Pair}; -use rustyline::Context; - -/// All completable VCL keywords. -const KEYWORDS: &[&str] = &[ - "SELECT", "FROM", "WHERE", "PROOF", "LIMIT", "OFFSET", "ORDER", "BY", - "GROUP", "HAVING", "AS", "AND", "OR", "NOT", "IN", "BETWEEN", "LIKE", - "EXISTS", "CONTAINS", "SIMILAR", "TO", "TRAVERSE", "DEPTH", "THRESHOLD", - "DRIFT", "CONSISTENCY", "AT", "TIME", "EXPLAIN", "INSERT", "UPDATE", - "DELETE", "SET", "INTO", "VALUES", "CREATE", "DROP", "ALTER", "JOIN", - "ON", "WITH", "FEDERATION", "STORE", "OCTAD", "ALL", "ASC", "DESC", - "COUNT", "SUM", "AVG", "MIN", "MAX", "DISTINCT", -]; - -/// Modality names (also offered as completions). -/// All 8 octad modalities: Graph, Vector, Tensor, Semantic, Document, Temporal, -/// Provenance, Spatial. -const MODALITIES: &[&str] = &[ - "GRAPH", "VECTOR", "TENSOR", "SEMANTIC", "DOCUMENT", "TEMPORAL", - "PROVENANCE", "SPATIAL", -]; - -/// Meta-commands starting with backslash. -const META_COMMANDS: &[&str] = &[ - "\\connect", "\\explain", "\\timing", "\\format", "\\status", - "\\help", "\\quit", "\\q", -]; - -/// Tab-completer for VCL input. -/// -/// Completes the word under the cursor by matching against known keywords, -/// modality names, and meta-commands. Matching is case-insensitive; the -/// replacement preserves the user's casing style (upper if the prefix is -/// uppercase, otherwise lowercase). -pub struct VclCompleter; - -impl Completer for VclCompleter { - type Candidate = Pair; - - fn complete( - &self, - line: &str, - pos: usize, - _ctx: &Context<'_>, - ) -> rustyline::Result<(usize, Vec<Pair>)> { - let (start, prefix) = find_word_start(line, pos); - let mut candidates = Vec::new(); - - if prefix.is_empty() { - return Ok((start, candidates)); - } - - let upper_prefix = prefix.to_uppercase(); - - // Meta-commands: only complete if the prefix starts at the beginning - // of the line and begins with '\'. - if prefix.starts_with('\\') { - for cmd in META_COMMANDS { - if cmd.starts_with(&prefix.to_lowercase()) { - candidates.push(Pair { - display: cmd.to_string(), - replacement: cmd.to_string(), - }); - } - } - return Ok((start, candidates)); - } - - // Determine the user's casing preference: if the prefix is all - // uppercase, offer uppercase completions; otherwise lowercase. - let use_upper = prefix.chars().all(|c| c.is_uppercase() || !c.is_alphabetic()); - - // Modality names. - for name in MODALITIES { - if name.starts_with(&upper_prefix) { - let replacement = if use_upper { - name.to_string() - } else { - name.to_lowercase() - }; - candidates.push(Pair { - display: name.to_string(), - replacement, - }); - } - } - - // VCL keywords. - for kw in KEYWORDS { - if kw.starts_with(&upper_prefix) { - let replacement = if use_upper { - kw.to_string() - } else { - kw.to_lowercase() - }; - candidates.push(Pair { - display: kw.to_string(), - replacement, - }); - } - } - - Ok((start, candidates)) - } -} - -/// Find the start position and text of the word being completed. -/// -/// Scans backwards from `pos` to find the beginning of the current token. -/// Tokens are delimited by whitespace, parentheses, commas, and semicolons. -fn find_word_start(line: &str, pos: usize) -> (usize, &str) { - let bytes = line.as_bytes(); - let mut start = pos; - - while start > 0 { - let ch = bytes[start - 1] as char; - if ch.is_whitespace() || ch == '(' || ch == ')' || ch == ',' || ch == ';' { - break; - } - start -= 1; - } - - (start, &line[start..pos]) -} - -#[cfg(test)] -mod tests { - use super::*; - - #[test] - fn test_find_word_start_middle() { - let (start, prefix) = find_word_start("SELECT FRO", 10); - assert_eq!(start, 7); - assert_eq!(prefix, "FRO"); - } - - #[test] - fn test_find_word_start_beginning() { - let (start, prefix) = find_word_start("SEL", 3); - assert_eq!(start, 0); - assert_eq!(prefix, "SEL"); - } - - #[test] - fn test_find_word_start_after_paren() { - let (start, prefix) = find_word_start("COUNT(DIS", 9); - assert_eq!(start, 6); - assert_eq!(prefix, "DIS"); - } - - #[test] - fn test_find_word_start_empty() { - let (start, prefix) = find_word_start("SELECT ", 7); - assert_eq!(start, 7); - assert_eq!(prefix, ""); - } - - #[test] - fn test_find_word_start_backslash() { - let (start, prefix) = find_word_start("\\con", 4); - assert_eq!(start, 0); - assert_eq!(prefix, "\\con"); - } - - // Note: full Completer::complete tests require a rustyline Context, - // which is difficult to construct in unit tests. The word-finding - // logic tested above is the core of the completion behaviour. -} diff --git a/verisimdb/rust-core/verisim-repl/src/formatter.rs b/verisimdb/rust-core/verisim-repl/src/formatter.rs deleted file mode 100644 index 4a69f300..00000000 --- a/verisimdb/rust-core/verisim-repl/src/formatter.rs +++ /dev/null @@ -1,329 +0,0 @@ -// SPDX-License-Identifier: MPL-2.0 -// Copyright (c) 2026 Jonathan D.A. Jewell (hyperpolymath) <j.d.a.jewell@open.ac.uk> -//! -//! Output formatters for VCL query results. -//! -//! Supports three output modes: -//! - **Table**: Human-readable columnar output using `comfy-table`. -//! - **JSON**: Pretty-printed JSON (pass-through from server response). -//! - **CSV**: Comma-separated values for pipeline consumption. - -use comfy_table::{Cell, ContentArrangement, Table}; -use serde_json::Value; -use std::fmt; - -/// Available output formats. -#[derive(Debug, Clone, Copy, PartialEq, Eq)] -pub enum OutputFormat { - Table, - Json, - Csv, -} - -impl fmt::Display for OutputFormat { - fn fmt(&self, f: &mut fmt::Formatter<'_>) -> fmt::Result { - match self { - OutputFormat::Table => write!(f, "table"), - OutputFormat::Json => write!(f, "json"), - OutputFormat::Csv => write!(f, "csv"), - } - } -} - -impl std::str::FromStr for OutputFormat { - type Err = String; - - fn from_str(s: &str) -> Result<Self, Self::Err> { - match s.to_lowercase().as_str() { - "table" => Ok(OutputFormat::Table), - "json" => Ok(OutputFormat::Json), - "csv" => Ok(OutputFormat::Csv), - other => Err(format!( - "Unknown format '{other}'. Valid formats: table, json, csv" - )), - } - } -} - -/// Format a JSON value according to the selected output format. -/// -/// The JSON value may be: -/// - An array of objects (query result rows) -/// - A single object (health check, explain output, etc.) -/// - A scalar or other shape (rendered as-is for JSON, best-effort for table/CSV) -pub fn format_value(value: &Value, format: OutputFormat) -> String { - match format { - OutputFormat::Json => format_json(value), - OutputFormat::Table => format_table(value), - OutputFormat::Csv => format_csv(value), - } -} - -/// Pretty-print JSON with 2-space indentation. -fn format_json(value: &Value) -> String { - serde_json::to_string_pretty(value).unwrap_or_else(|_| value.to_string()) -} - -/// Render a JSON value as a table. -/// -/// For arrays of objects, each object becomes a row and each unique key becomes -/// a column. For single objects, each key-value pair becomes a row with two -/// columns ("Field" and "Value"). For scalars, a single-cell table is produced. -fn format_table(value: &Value) -> String { - match value { - Value::Array(rows) if !rows.is_empty() => format_array_table(rows), - Value::Object(obj) => format_object_table(obj), - other => format!("{other}"), - } -} - -/// Render an array of JSON objects as a columnar table. -fn format_array_table(rows: &[Value]) -> String { - // Collect all unique keys in insertion order from the first object, - // then add any keys found in subsequent objects. - let mut columns: Vec<String> = Vec::new(); - for row in rows { - if let Value::Object(obj) = row { - for key in obj.keys() { - if !columns.contains(key) { - columns.push(key.clone()); - } - } - } - } - - if columns.is_empty() { - // Array of non-objects: render each element as a single row. - let mut table = Table::new(); - table.set_content_arrangement(ContentArrangement::Dynamic); - table.set_header(vec![Cell::new("value")]); - for item in rows { - table.add_row(vec![Cell::new(value_to_cell(item))]); - } - return table.to_string(); - } - - let mut table = Table::new(); - table.set_content_arrangement(ContentArrangement::Dynamic); - table.set_header(columns.iter().map(|c| Cell::new(c))); - - for row in rows { - let cells: Vec<Cell> = columns - .iter() - .map(|col| { - let val = row.get(col).unwrap_or(&Value::Null); - Cell::new(value_to_cell(val)) - }) - .collect(); - table.add_row(cells); - } - - let row_count = rows.len(); - format!("{table}\n({row_count} row{})", if row_count == 1 { "" } else { "s" }) -} - -/// Render a single JSON object as a two-column table (Field | Value). -fn format_object_table(obj: &serde_json::Map<String, Value>) -> String { - let mut table = Table::new(); - table.set_content_arrangement(ContentArrangement::Dynamic); - table.set_header(vec![Cell::new("Field"), Cell::new("Value")]); - - for (key, val) in obj { - table.add_row(vec![Cell::new(key), Cell::new(value_to_cell(val))]); - } - - table.to_string() -} - -/// Convert a JSON value to a short string suitable for a table cell. -/// -/// Arrays and nested objects are truncated to avoid overwhelming the table. -fn value_to_cell(value: &Value) -> String { - match value { - Value::Null => "NULL".to_string(), - Value::Bool(b) => b.to_string(), - Value::Number(n) => n.to_string(), - Value::String(s) => s.clone(), - Value::Array(arr) => { - if arr.len() <= 3 { - format!("{value}") - } else { - format!("[{} items]", arr.len()) - } - } - Value::Object(obj) => { - if obj.len() <= 3 { - format!("{value}") - } else { - format!("{{{} fields}}", obj.len()) - } - } - } -} - -/// Render a JSON value as CSV. -/// -/// For arrays of objects, the first row is a header line derived from keys. -/// For single objects, each key-value pair becomes a row. -fn format_csv(value: &Value) -> String { - match value { - Value::Array(rows) if !rows.is_empty() => format_array_csv(rows), - Value::Object(obj) => format_object_csv(obj), - other => format!("{other}"), - } -} - -/// Render an array of JSON objects as CSV rows. -fn format_array_csv(rows: &[Value]) -> String { - let mut columns: Vec<String> = Vec::new(); - for row in rows { - if let Value::Object(obj) = row { - for key in obj.keys() { - if !columns.contains(key) { - columns.push(key.clone()); - } - } - } - } - - let mut output = String::new(); - - // Header row - output.push_str(&columns.join(",")); - output.push('\n'); - - // Data rows - for row in rows { - let cells: Vec<String> = columns - .iter() - .map(|col| { - let val = row.get(col).unwrap_or(&Value::Null); - csv_escape(val) - }) - .collect(); - output.push_str(&cells.join(",")); - output.push('\n'); - } - - output -} - -/// Render a single JSON object as CSV (key,value rows). -fn format_object_csv(obj: &serde_json::Map<String, Value>) -> String { - let mut output = String::from("field,value\n"); - for (key, val) in obj { - output.push_str(&format!("{},{}\n", csv_escape_str(key), csv_escape(val))); - } - output -} - -/// Escape a JSON value for CSV output. -/// -/// Strings containing commas, quotes, or newlines are double-quoted with -/// internal quotes escaped per RFC 4180. -fn csv_escape(value: &Value) -> String { - match value { - Value::Null => String::new(), - Value::String(s) => csv_escape_str(s), - other => csv_escape_str(&other.to_string()), - } -} - -/// Escape a string for CSV output per RFC 4180. -fn csv_escape_str(s: &str) -> String { - if s.contains(',') || s.contains('"') || s.contains('\n') { - format!("\"{}\"", s.replace('"', "\"\"")) - } else { - s.to_string() - } -} - -#[cfg(test)] -mod tests { - use super::*; - use serde_json::json; - - #[test] - fn test_output_format_parse() { - assert_eq!("table".parse::<OutputFormat>().expect("TODO: handle error"), OutputFormat::Table); - assert_eq!("JSON".parse::<OutputFormat>().expect("TODO: handle error"), OutputFormat::Json); - assert_eq!("csv".parse::<OutputFormat>().expect("TODO: handle error"), OutputFormat::Csv); - assert!("xml".parse::<OutputFormat>().is_err()); - } - - #[test] - fn test_output_format_display() { - assert_eq!(OutputFormat::Table.to_string(), "table"); - assert_eq!(OutputFormat::Json.to_string(), "json"); - assert_eq!(OutputFormat::Csv.to_string(), "csv"); - } - - #[test] - fn test_format_json_pretty() { - let val = json!({"status": "healthy", "version": "0.1.0"}); - let out = format_value(&val, OutputFormat::Json); - assert!(out.contains("\"status\": \"healthy\"")); - assert!(out.contains('\n')); - } - - #[test] - fn test_format_table_array() { - let val = json!([ - {"id": "abc", "score": 0.95}, - {"id": "def", "score": 0.80} - ]); - let out = format_value(&val, OutputFormat::Table); - assert!(out.contains("abc")); - assert!(out.contains("0.95")); - assert!(out.contains("(2 rows)")); - } - - #[test] - fn test_format_table_object() { - let val = json!({"status": "healthy", "uptime_seconds": 42}); - let out = format_value(&val, OutputFormat::Table); - assert!(out.contains("Field")); - assert!(out.contains("Value")); - assert!(out.contains("healthy")); - } - - #[test] - fn test_format_csv_array() { - let val = json!([ - {"name": "Alice", "age": 30}, - {"name": "Bob", "age": 25} - ]); - let out = format_value(&val, OutputFormat::Csv); - assert!(out.starts_with("name,age\n") || out.starts_with("age,name\n")); - assert!(out.contains("Alice")); - assert!(out.contains("30")); - } - - #[test] - fn test_csv_escape_commas() { - assert_eq!(csv_escape_str("hello,world"), "\"hello,world\""); - } - - #[test] - fn test_csv_escape_quotes() { - assert_eq!(csv_escape_str("say \"hi\""), "\"say \"\"hi\"\"\""); - } - - #[test] - fn test_value_to_cell_null() { - assert_eq!(value_to_cell(&Value::Null), "NULL"); - } - - #[test] - fn test_value_to_cell_large_array() { - let val = json!([1, 2, 3, 4, 5]); - assert_eq!(value_to_cell(&val), "[5 items]"); - } - - #[test] - fn test_format_single_row() { - let val = json!([{"id": "only-one"}]); - let out = format_value(&val, OutputFormat::Table); - assert!(out.contains("(1 row)")); - } -} diff --git a/verisimdb/rust-core/verisim-repl/src/highlighter.rs b/verisimdb/rust-core/verisim-repl/src/highlighter.rs deleted file mode 100644 index ea116efe..00000000 --- a/verisimdb/rust-core/verisim-repl/src/highlighter.rs +++ /dev/null @@ -1,238 +0,0 @@ -// SPDX-License-Identifier: MPL-2.0 -// Copyright (c) 2026 Jonathan D.A. Jewell (hyperpolymath) <j.d.a.jewell@open.ac.uk> -//! -//! VCL syntax highlighting for the interactive REPL. -//! -//! Implements `rustyline::highlight::Highlighter` to colour VCL keywords, -//! modality names, string literals, and numeric literals as the user types. - -use colored::Colorize; -use rustyline::highlight::Highlighter; -use std::borrow::Cow; - -/// VCL keywords that are highlighted in blue/bold. -const VCL_KEYWORDS: &[&str] = &[ - "SELECT", "FROM", "WHERE", "PROOF", "LIMIT", "OFFSET", "ORDER", "BY", - "GROUP", "HAVING", "AS", "AND", "OR", "NOT", "IN", "BETWEEN", "LIKE", - "EXISTS", "CONTAINS", "SIMILAR", "TO", "TRAVERSE", "DEPTH", "THRESHOLD", - "DRIFT", "CONSISTENCY", "AT", "TIME", "EXPLAIN", "INSERT", "UPDATE", - "DELETE", "SET", "INTO", "VALUES", "CREATE", "DROP", "ALTER", "JOIN", - "ON", "WITH", "FEDERATION", "STORE", "OCTAD", "ALL", "ASC", "DESC", - "COUNT", "SUM", "AVG", "MIN", "MAX", "DISTINCT", -]; - -/// VCL modality names highlighted in green. -/// All 8 octad modalities: Graph, Vector, Tensor, Semantic, Document, Temporal, -/// Provenance, Spatial. -const VCL_MODALITIES: &[&str] = &[ - "GRAPH", "VECTOR", "TENSOR", "SEMANTIC", "DOCUMENT", "TEMPORAL", - "PROVENANCE", "SPATIAL", -]; - -/// Syntax highlighter for VCL input lines. -/// -/// This is used by the rustyline `Editor` to provide real-time syntax -/// colouring as the user types queries. -pub struct VclHighlighter; - -impl Highlighter for VclHighlighter { - /// Highlight the input line with ANSI colour codes. - /// - /// The highlighting strategy is token-based: - /// 1. String literals (single or double quoted) are coloured yellow. - /// 2. Tokens matching VCL keywords are coloured blue and bold. - /// 3. Tokens matching modality names are coloured green and bold. - /// 4. Numeric tokens are coloured cyan. - /// 5. Everything else is left uncoloured. - fn highlight<'l>(&self, line: &'l str, _pos: usize) -> Cow<'l, str> { - let highlighted = highlight_line(line); - Cow::Owned(highlighted) - } - - /// Indicate that we always want to repaint when the line changes. - fn highlight_char(&self, _line: &str, _pos: usize, _forced: rustyline::highlight::CmdKind) -> bool { - true - } - - /// Highlight the prompt itself (not coloured — we handle prompt - /// colouring in main). - fn highlight_prompt<'b, 's: 'b, 'p: 'b>( - &'s self, - prompt: &'p str, - _default: bool, - ) -> Cow<'b, str> { - Cow::Borrowed(prompt) - } - - /// Highlight a hint (dimmed text shown after the cursor). - fn highlight_hint<'h>(&self, hint: &'h str) -> Cow<'h, str> { - Cow::Owned(hint.dimmed().to_string()) - } - - /// Highlight the currently selected candidate during completion. - fn highlight_candidate<'c>( - &self, - candidate: &'c str, - _completion: rustyline::CompletionType, - ) -> Cow<'c, str> { - Cow::Borrowed(candidate) - } -} - -/// Apply syntax highlighting to a single line of VCL input. -/// -/// Handles quoted strings as atomic units so that keywords inside strings -/// are not incorrectly coloured. -fn highlight_line(line: &str) -> String { - let mut result = String::with_capacity(line.len() * 2); - let chars: Vec<char> = line.chars().collect(); - let len = chars.len(); - let mut i = 0; - - while i < len { - let ch = chars[i]; - - // Handle string literals (single or double quoted). - if ch == '\'' || ch == '"' { - let quote = ch; - let start = i; - i += 1; - while i < len && chars[i] != quote { - if chars[i] == '\\' { - i += 1; // Skip escaped character. - } - i += 1; - } - if i < len { - i += 1; // Consume closing quote. - } - let string_slice: String = chars[start..i].iter().collect(); - result.push_str(&string_slice.yellow().to_string()); - continue; - } - - // Handle meta-commands (lines starting with \). - if ch == '\\' && i == 0 { - // Colour the entire line as a meta-command. - let rest: String = chars[i..].iter().collect(); - result.push_str(&rest.bright_magenta().to_string()); - break; - } - - // Handle word tokens (identifiers and keywords). - if ch.is_alphabetic() || ch == '_' { - let start = i; - while i < len && (chars[i].is_alphanumeric() || chars[i] == '_') { - i += 1; - } - let word: String = chars[start..i].iter().collect(); - let upper = word.to_uppercase(); - - if VCL_MODALITIES.contains(&upper.as_str()) { - result.push_str(&word.green().bold().to_string()); - } else if VCL_KEYWORDS.contains(&upper.as_str()) { - result.push_str(&word.blue().bold().to_string()); - } else { - result.push_str(&word); - } - continue; - } - - // Handle numeric literals (integers and floats). - if ch.is_ascii_digit() || (ch == '-' && i + 1 < len && chars[i + 1].is_ascii_digit()) { - let start = i; - if ch == '-' { - i += 1; - } - while i < len && (chars[i].is_ascii_digit() || chars[i] == '.') { - i += 1; - } - // Handle scientific notation (e.g. 1e10, 2.5E-3). - if i < len && (chars[i] == 'e' || chars[i] == 'E') { - i += 1; - if i < len && (chars[i] == '+' || chars[i] == '-') { - i += 1; - } - while i < len && chars[i].is_ascii_digit() { - i += 1; - } - } - let num: String = chars[start..i].iter().collect(); - result.push_str(&num.cyan().to_string()); - continue; - } - - // Pass through everything else (whitespace, operators, etc.). - result.push(ch); - i += 1; - } - - result -} - -#[cfg(test)] -mod tests { - use super::*; - - #[test] - fn test_highlight_returns_something() { - let hl = VclHighlighter; - let output = hl.highlight("SELECT FROM octad", 0); - // Just verify it does not panic and returns non-empty output. - assert!(!output.is_empty()); - } - - #[test] - fn test_highlight_char_always_true() { - let hl = VclHighlighter; - assert!(hl.highlight_char("test", 0, rustyline::highlight::CmdKind::Other)); - } - - #[test] - fn test_highlight_preserves_plain_text() { - // A line with no keywords should not gain extra visible characters - // (it may have ANSI reset codes, but the visible text should match). - let line = "foobar baz"; - let output = highlight_line(line); - // Strip ANSI codes and check the visible text is preserved. - let stripped = strip_ansi(&output); - assert_eq!(stripped, line); - } - - #[test] - fn test_highlight_string_literal() { - let line = "WHERE name = 'hello'"; - let output = highlight_line(line); - // The string 'hello' should be present in output (with ANSI codes). - let stripped = strip_ansi(&output); - assert!(stripped.contains("'hello'")); - } - - #[test] - fn test_highlight_number() { - let line = "LIMIT 42"; - let output = highlight_line(line); - let stripped = strip_ansi(&output); - assert!(stripped.contains("42")); - } - - /// Strip ANSI escape sequences from a string (for testing visible content). - fn strip_ansi(s: &str) -> String { - let mut result = String::new(); - let mut in_escape = false; - for ch in s.chars() { - if ch == '\x1b' { - in_escape = true; - continue; - } - if in_escape { - if ch == 'm' { - in_escape = false; - } - continue; - } - result.push(ch); - } - result - } -} diff --git a/verisimdb/rust-core/verisim-repl/src/linter.rs b/verisimdb/rust-core/verisim-repl/src/linter.rs deleted file mode 100644 index 888df169..00000000 --- a/verisimdb/rust-core/verisim-repl/src/linter.rs +++ /dev/null @@ -1,517 +0,0 @@ -// SPDX-License-Identifier: MPL-2.0 -// Copyright (c) 2026 Jonathan D.A. Jewell (hyperpolymath) <j.d.a.jewell@open.ac.uk> -//! -//! VCL linter — static analysis for VCL queries. -//! -//! Detects common issues, antipatterns, and performance pitfalls in VCL -//! queries before execution: -//! -//! - Missing LIMIT clause on unbounded queries -//! - SELECT * (all modalities) without explicit need -//! - Missing PROOF clause on sensitive modalities (Semantic) -//! - Expensive cross-modality operations without EXPLAIN -//! - Unreachable or redundant clauses -//! - TRAVERSE without DEPTH bound -//! - High estimated row count without pagination - -use std::fmt; - -/// Severity level for lint diagnostics. -#[derive(Debug, Clone, Copy, PartialEq, Eq, PartialOrd, Ord)] -pub enum Severity { - /// Informational suggestion — not an error. - Hint, - /// Potential issue that may cause unexpected behaviour. - Warning, - /// Likely error or dangerous antipattern. - Error, -} - -impl fmt::Display for Severity { - fn fmt(&self, f: &mut fmt::Formatter<'_>) -> fmt::Result { - match self { - Severity::Hint => write!(f, "hint"), - Severity::Warning => write!(f, "warning"), - Severity::Error => write!(f, "error"), - } - } -} - -/// A lint rule identifier. -#[derive(Debug, Clone, Copy, PartialEq, Eq)] -pub enum LintRule { - /// Query lacks a LIMIT clause — may return unbounded results. - MissingLimit, - /// SELECT queries all modalities when fewer would suffice. - SelectAllModalities, - /// Semantic modality accessed without PROOF clause. - MissingProof, - /// TRAVERSE clause without DEPTH bound. - UnboundedTraverse, - /// Query uses DRIFT or CONSISTENCY without a threshold. - MissingThreshold, - /// ORDER BY without LIMIT (full sort on potentially large result set). - OrderByWithoutLimit, - /// DELETE or UPDATE without WHERE clause. - DangerousWrite, - /// VCL-UT PROOF type not recognized. - UnknownProofType, - /// Redundant WHERE clause (always true). - RedundantWhere, - /// Multiple modalities without EXPLAIN — consider reviewing the plan. - MultiModalityNoExplain, - /// FEDERATION query without specifying STORE. - FederationWithoutStore, -} - -impl fmt::Display for LintRule { - fn fmt(&self, f: &mut fmt::Formatter<'_>) -> fmt::Result { - match self { - LintRule::MissingLimit => write!(f, "VCL001"), - LintRule::SelectAllModalities => write!(f, "VCL002"), - LintRule::MissingProof => write!(f, "VCL003"), - LintRule::UnboundedTraverse => write!(f, "VCL004"), - LintRule::MissingThreshold => write!(f, "VCL005"), - LintRule::OrderByWithoutLimit => write!(f, "VCL006"), - LintRule::DangerousWrite => write!(f, "VCL007"), - LintRule::UnknownProofType => write!(f, "VCL008"), - LintRule::RedundantWhere => write!(f, "VCL009"), - LintRule::MultiModalityNoExplain => write!(f, "VCL010"), - LintRule::FederationWithoutStore => write!(f, "VCL011"), - } - } -} - -impl LintRule { - /// Human-readable description of the rule. - pub fn description(&self) -> &'static str { - match self { - LintRule::MissingLimit => "Query lacks LIMIT clause — may return unbounded results", - LintRule::SelectAllModalities => "Query selects all 8 modalities — consider selecting only what you need", - LintRule::MissingProof => "Semantic modality accessed without PROOF clause — data integrity not verified", - LintRule::UnboundedTraverse => "TRAVERSE without DEPTH limit — may explore entire graph", - LintRule::MissingThreshold => "DRIFT/CONSISTENCY check without THRESHOLD — using implicit default", - LintRule::OrderByWithoutLimit => "ORDER BY without LIMIT — sorting potentially unbounded result set", - LintRule::DangerousWrite => "DELETE/UPDATE without WHERE clause — affects all entities", - LintRule::UnknownProofType => "Unrecognized proof type in PROOF clause", - LintRule::RedundantWhere => "WHERE clause appears redundant (always true condition)", - LintRule::MultiModalityNoExplain => "Multi-modality query — consider running EXPLAIN first to review the plan", - LintRule::FederationWithoutStore => "FEDERATION query without STORE — will query all federated instances", - } - } -} - -/// A single lint diagnostic. -#[derive(Debug, Clone)] -pub struct LintDiagnostic { - pub rule: LintRule, - pub severity: Severity, - pub message: String, - /// Approximate character offset in the query (0 if unknown). - pub offset: usize, -} - -impl fmt::Display for LintDiagnostic { - fn fmt(&self, f: &mut fmt::Formatter<'_>) -> fmt::Result { - write!(f, "[{}] {}: {}", self.rule, self.severity, self.message) - } -} - -/// VCL modality names — all 8 octad modalities. -const MODALITIES: &[&str] = &[ - "GRAPH", "VECTOR", "TENSOR", "SEMANTIC", "DOCUMENT", "TEMPORAL", - "PROVENANCE", "SPATIAL", -]; - -/// Known VCL-UT proof types. -const KNOWN_PROOF_TYPES: &[&str] = &[ - "EXISTENCE", "CONSISTENCY", "INTEGRITY", "AUTHENTICITY", - "PROVENANCE", "ACCESS", "CITATION", "ZKP", "PLONK", -]; - -/// Lint a VCL query string and return diagnostics. -/// -/// This performs token-level analysis (not full parsing) to detect common -/// issues. For full AST-based linting, the LSP will use the ReScript parser. -pub fn lint_query(query: &str) -> Vec<LintDiagnostic> { - let mut diagnostics = Vec::new(); - let upper = query.to_uppercase(); - let tokens = tokenize_upper(&upper); - - // Determine query type - let is_select = tokens.contains(&"SELECT"); - let is_delete = tokens.contains(&"DELETE"); - let is_update = tokens.contains(&"UPDATE"); - let is_explain = tokens.contains(&"EXPLAIN"); - let has_limit = tokens.contains(&"LIMIT"); - let has_where = tokens.contains(&"WHERE"); - let has_order = tokens.contains(&"ORDER"); - let has_traverse = tokens.contains(&"TRAVERSE"); - let has_depth = tokens.contains(&"DEPTH"); - let has_proof = tokens.contains(&"PROOF"); - let has_drift = tokens.contains(&"DRIFT") || tokens.contains(&"CONSISTENCY"); - let has_threshold = tokens.contains(&"THRESHOLD"); - let has_federation = tokens.contains(&"FEDERATION"); - let has_store = tokens.contains(&"STORE"); - let has_semantic = tokens.contains(&"SEMANTIC"); - - // VCL001: Missing LIMIT on SELECT - if is_select && !has_limit && !is_explain { - diagnostics.push(LintDiagnostic { - rule: LintRule::MissingLimit, - severity: Severity::Warning, - message: LintRule::MissingLimit.description().to_string(), - offset: 0, - }); - } - - // VCL002: SELECT all modalities - if is_select { - let modality_count = MODALITIES - .iter() - .filter(|m| tokens.contains(m)) - .count(); - if modality_count >= 8 { - diagnostics.push(LintDiagnostic { - rule: LintRule::SelectAllModalities, - severity: Severity::Hint, - message: LintRule::SelectAllModalities.description().to_string(), - offset: 0, - }); - } - } - - // VCL003: Semantic without PROOF - if has_semantic && !has_proof && is_select { - diagnostics.push(LintDiagnostic { - rule: LintRule::MissingProof, - severity: Severity::Warning, - message: LintRule::MissingProof.description().to_string(), - offset: upper.find("SEMANTIC").unwrap_or(0), - }); - } - - // VCL004: TRAVERSE without DEPTH - if has_traverse && !has_depth { - diagnostics.push(LintDiagnostic { - rule: LintRule::UnboundedTraverse, - severity: Severity::Error, - message: LintRule::UnboundedTraverse.description().to_string(), - offset: upper.find("TRAVERSE").unwrap_or(0), - }); - } - - // VCL005: DRIFT/CONSISTENCY without THRESHOLD - if has_drift && !has_threshold { - diagnostics.push(LintDiagnostic { - rule: LintRule::MissingThreshold, - severity: Severity::Hint, - message: LintRule::MissingThreshold.description().to_string(), - offset: upper.find("DRIFT").or_else(|| upper.find("CONSISTENCY")).unwrap_or(0), - }); - } - - // VCL006: ORDER BY without LIMIT - if has_order && !has_limit && is_select { - diagnostics.push(LintDiagnostic { - rule: LintRule::OrderByWithoutLimit, - severity: Severity::Warning, - message: LintRule::OrderByWithoutLimit.description().to_string(), - offset: upper.find("ORDER").unwrap_or(0), - }); - } - - // VCL007: Dangerous write without WHERE - if (is_delete || is_update) && !has_where { - diagnostics.push(LintDiagnostic { - rule: LintRule::DangerousWrite, - severity: Severity::Error, - message: LintRule::DangerousWrite.description().to_string(), - offset: 0, - }); - } - - // VCL008: Unknown proof type - if has_proof { - if let Some(proof_pos) = tokens.iter().position(|t| *t == "PROOF") { - if let Some(proof_type) = tokens.get(proof_pos + 1) { - if !KNOWN_PROOF_TYPES.contains(proof_type) && !proof_type.is_empty() { - diagnostics.push(LintDiagnostic { - rule: LintRule::UnknownProofType, - severity: Severity::Warning, - message: format!( - "Unknown proof type '{}' — known types: {}", - proof_type, - KNOWN_PROOF_TYPES.join(", ") - ), - offset: upper.find(proof_type).unwrap_or(0), - }); - } - } - } - } - - // VCL009: Redundant WHERE (WHERE 1=1, WHERE TRUE) - if has_where { - let where_pos = upper.find("WHERE").unwrap_or(0); - let after_where = &upper[where_pos + 5..].trim_start(); - if after_where.starts_with("1=1") - || after_where.starts_with("1 = 1") - || after_where.starts_with("TRUE") - { - diagnostics.push(LintDiagnostic { - rule: LintRule::RedundantWhere, - severity: Severity::Hint, - message: LintRule::RedundantWhere.description().to_string(), - offset: where_pos, - }); - } - } - - // VCL010: Multi-modality without EXPLAIN - if is_select && !is_explain { - let modality_count = MODALITIES - .iter() - .filter(|m| tokens.contains(m)) - .count(); - if modality_count >= 3 { - diagnostics.push(LintDiagnostic { - rule: LintRule::MultiModalityNoExplain, - severity: Severity::Hint, - message: LintRule::MultiModalityNoExplain.description().to_string(), - offset: 0, - }); - } - } - - // VCL011: FEDERATION without STORE - if has_federation && !has_store { - diagnostics.push(LintDiagnostic { - rule: LintRule::FederationWithoutStore, - severity: Severity::Warning, - message: LintRule::FederationWithoutStore.description().to_string(), - offset: upper.find("FEDERATION").unwrap_or(0), - }); - } - - // Sort by severity (errors first) - diagnostics.sort_by(|a, b| b.severity.cmp(&a.severity)); - diagnostics -} - -/// Quick check: does a query have any errors (not just warnings/hints)? -pub fn has_errors(query: &str) -> bool { - lint_query(query) - .iter() - .any(|d| d.severity == Severity::Error) -} - -/// Format lint diagnostics for terminal output. -pub fn format_diagnostics(query: &str, diagnostics: &[LintDiagnostic]) -> String { - if diagnostics.is_empty() { - return String::new(); - } - - let mut output = String::new(); - let error_count = diagnostics.iter().filter(|d| d.severity == Severity::Error).count(); - let warning_count = diagnostics.iter().filter(|d| d.severity == Severity::Warning).count(); - let hint_count = diagnostics.iter().filter(|d| d.severity == Severity::Hint).count(); - - for diag in diagnostics { - output.push_str(&format!(" {} {}\n", diag.rule, diag.message)); - } - - output.push_str(&format!( - "\n {} error(s), {} warning(s), {} hint(s) in: {}\n", - error_count, - warning_count, - hint_count, - if query.len() > 60 { - format!("{}...", &query[..57]) - } else { - query.to_string() - } - )); - - output -} - -/// Tokenize uppercase query into whitespace-separated tokens. -fn tokenize_upper(upper: &str) -> Vec<&str> { - upper.split_whitespace().collect() -} - -#[cfg(test)] -mod tests { - use super::*; - - #[test] - fn test_clean_query_no_errors() { - let diagnostics = lint_query("SELECT GRAPH FROM OCTAD WHERE id = 'abc' LIMIT 10"); - let errors: Vec<_> = diagnostics.iter().filter(|d| d.severity == Severity::Error).collect(); - assert!(errors.is_empty()); - } - - #[test] - fn test_missing_limit() { - let diagnostics = lint_query("SELECT GRAPH FROM OCTAD"); - assert!(diagnostics.iter().any(|d| d.rule == LintRule::MissingLimit)); - } - - #[test] - fn test_explain_exempt_from_limit() { - let diagnostics = lint_query("EXPLAIN SELECT GRAPH FROM OCTAD"); - assert!(!diagnostics.iter().any(|d| d.rule == LintRule::MissingLimit)); - } - - #[test] - fn test_select_all_modalities() { - let diagnostics = lint_query( - "SELECT GRAPH VECTOR TENSOR SEMANTIC DOCUMENT TEMPORAL PROVENANCE SPATIAL FROM OCTAD LIMIT 10" - ); - assert!(diagnostics.iter().any(|d| d.rule == LintRule::SelectAllModalities)); - } - - #[test] - fn test_semantic_without_proof() { - let diagnostics = lint_query("SELECT SEMANTIC FROM OCTAD LIMIT 10"); - assert!(diagnostics.iter().any(|d| d.rule == LintRule::MissingProof)); - } - - #[test] - fn test_semantic_with_proof_ok() { - let diagnostics = lint_query("SELECT SEMANTIC FROM OCTAD PROOF EXISTENCE LIMIT 10"); - assert!(!diagnostics.iter().any(|d| d.rule == LintRule::MissingProof)); - } - - #[test] - fn test_unbounded_traverse() { - let diagnostics = lint_query("SELECT GRAPH FROM OCTAD TRAVERSE relates_to LIMIT 10"); - assert!(diagnostics.iter().any(|d| d.rule == LintRule::UnboundedTraverse)); - assert!(diagnostics.iter().any(|d| d.severity == Severity::Error)); - } - - #[test] - fn test_traverse_with_depth_ok() { - let diagnostics = lint_query("SELECT GRAPH FROM OCTAD TRAVERSE relates_to DEPTH 3 LIMIT 10"); - assert!(!diagnostics.iter().any(|d| d.rule == LintRule::UnboundedTraverse)); - } - - #[test] - fn test_drift_without_threshold() { - let diagnostics = lint_query("SELECT GRAPH FROM OCTAD WHERE DRIFT LIMIT 10"); - assert!(diagnostics.iter().any(|d| d.rule == LintRule::MissingThreshold)); - } - - #[test] - fn test_order_without_limit() { - let diagnostics = lint_query("SELECT GRAPH FROM OCTAD ORDER BY name"); - assert!(diagnostics.iter().any(|d| d.rule == LintRule::OrderByWithoutLimit)); - } - - #[test] - fn test_dangerous_delete() { - let diagnostics = lint_query("DELETE FROM OCTAD"); - assert!(diagnostics.iter().any(|d| d.rule == LintRule::DangerousWrite)); - assert!(diagnostics.iter().any(|d| d.severity == Severity::Error)); - } - - #[test] - fn test_delete_with_where_ok() { - let diagnostics = lint_query("DELETE FROM OCTAD WHERE id = 'abc'"); - assert!(!diagnostics.iter().any(|d| d.rule == LintRule::DangerousWrite)); - } - - #[test] - fn test_unknown_proof_type() { - let diagnostics = lint_query("SELECT SEMANTIC FROM OCTAD PROOF FOOBAR LIMIT 10"); - assert!(diagnostics.iter().any(|d| d.rule == LintRule::UnknownProofType)); - } - - #[test] - fn test_known_proof_types_ok() { - for proof_type in &["EXISTENCE", "CONSISTENCY", "INTEGRITY", "ZKP", "PLONK"] { - let q = format!("SELECT SEMANTIC FROM OCTAD PROOF {} LIMIT 10", proof_type); - let diagnostics = lint_query(&q); - assert!( - !diagnostics.iter().any(|d| d.rule == LintRule::UnknownProofType), - "Proof type {} should be recognized", - proof_type - ); - } - } - - #[test] - fn test_redundant_where() { - let diagnostics = lint_query("SELECT GRAPH FROM OCTAD WHERE 1=1 LIMIT 10"); - assert!(diagnostics.iter().any(|d| d.rule == LintRule::RedundantWhere)); - } - - #[test] - fn test_multi_modality_hint() { - let diagnostics = lint_query( - "SELECT GRAPH VECTOR TENSOR FROM OCTAD LIMIT 10" - ); - assert!(diagnostics.iter().any(|d| d.rule == LintRule::MultiModalityNoExplain)); - } - - #[test] - fn test_federation_without_store() { - let diagnostics = lint_query("SELECT GRAPH FROM FEDERATION OCTAD LIMIT 10"); - assert!(diagnostics.iter().any(|d| d.rule == LintRule::FederationWithoutStore)); - } - - #[test] - fn test_federation_with_store_ok() { - let diagnostics = lint_query("SELECT GRAPH FROM FEDERATION STORE 'remote-1' OCTAD LIMIT 10"); - assert!(!diagnostics.iter().any(|d| d.rule == LintRule::FederationWithoutStore)); - } - - #[test] - fn test_has_errors_true() { - assert!(has_errors("DELETE FROM OCTAD")); - } - - #[test] - fn test_has_errors_false() { - assert!(!has_errors("SELECT GRAPH FROM OCTAD LIMIT 10")); - } - - #[test] - fn test_diagnostics_sorted_by_severity() { - let diagnostics = lint_query( - "DELETE FROM FEDERATION OCTAD" - ); - // Errors should come before warnings - let severities: Vec<_> = diagnostics.iter().map(|d| d.severity).collect(); - for window in severities.windows(2) { - assert!(window[0] >= window[1], "Diagnostics should be sorted by severity"); - } - } - - #[test] - fn test_format_diagnostics_output() { - let diagnostics = lint_query("DELETE FROM OCTAD"); - let output = format_diagnostics("DELETE FROM OCTAD", &diagnostics); - assert!(output.contains("VCL007")); - assert!(output.contains("error")); - } - - #[test] - fn test_empty_query() { - let diagnostics = lint_query(""); - assert!(diagnostics.is_empty()); - } - - #[test] - fn test_case_insensitive() { - let diagnostics = lint_query("select semantic from octad limit 10"); - assert!(diagnostics.iter().any(|d| d.rule == LintRule::MissingProof)); - } - - #[test] - fn test_update_without_where() { - let diagnostics = lint_query("UPDATE OCTAD SET name = 'test'"); - assert!(diagnostics.iter().any(|d| d.rule == LintRule::DangerousWrite)); - } -} diff --git a/verisimdb/rust-core/verisim-repl/src/main.rs b/verisimdb/rust-core/verisim-repl/src/main.rs deleted file mode 100644 index af879016..00000000 --- a/verisimdb/rust-core/verisim-repl/src/main.rs +++ /dev/null @@ -1,515 +0,0 @@ -// SPDX-License-Identifier: MPL-2.0 -// Copyright (c) 2026 Jonathan D.A. Jewell (hyperpolymath) <j.d.a.jewell@open.ac.uk> -//! -//! VCL REPL — Interactive query shell for VeriSimDB. -//! -//! Provides a readline-based interactive shell with: -//! - VCL syntax highlighting -//! - Tab completion for keywords, modalities, and meta-commands -//! - Multiline query support (backslash continuation) -//! - Multiple output formats (table, JSON, CSV) -//! - Meta-commands for session control -//! - Query timing display -//! - Persistent command history - -#![forbid(unsafe_code)] -mod client; -mod completer; -mod formatter; -mod highlighter; -pub mod linter; -pub mod vcl_fmt; - -use clap::Parser; -use colored::Colorize; -use rustyline::config::Configurer; -use rustyline::error::ReadlineError; -use rustyline::hint::HistoryHinter; -use rustyline::history::DefaultHistory; -use rustyline::validate::MatchingBracketValidator; -use rustyline_derive::{Completer, Helper, Highlighter, Hinter, Validator}; -use std::time::Instant; - -use client::VclClient; -use formatter::{format_value, OutputFormat}; - -/// VeriSimDB version string, pulled from Cargo.toml at compile time. -const VERSION: &str = env!("CARGO_PKG_VERSION"); - -// --------------------------------------------------------------------------- -// CLI argument parsing -// --------------------------------------------------------------------------- - -/// VCL — Interactive query shell for VeriSimDB. -#[derive(Parser, Debug)] -#[command(name = "vcl", version = VERSION, about = "VCL REPL for VeriSimDB")] -struct Cli { - /// Hostname or IP of the verisim-api server. - #[arg(long, default_value = "localhost")] - host: String, - - /// Port of the verisim-api server. - #[arg(long, default_value_t = 8080)] - port: u16, - - /// Default output format. - #[arg(long, default_value = "table")] - format: String, -} - -// --------------------------------------------------------------------------- -// Rustyline helper (bundles all traits into one type) -// --------------------------------------------------------------------------- - -/// Combined helper that provides highlighting, completion, hinting, and -/// bracket validation for the rustyline editor. -#[derive(Helper, Highlighter, Completer, Hinter, Validator)] -struct VclHelper { - #[rustyline(Highlighter)] - highlighter: highlighter::VclHighlighter, - #[rustyline(Completer)] - completer: completer::VclCompleter, - #[rustyline(Hinter)] - hinter: HistoryHinter, - #[rustyline(Validator)] - validator: MatchingBracketValidator, -} - -// --------------------------------------------------------------------------- -// REPL session state -// --------------------------------------------------------------------------- - -/// Mutable session state for the REPL loop. -struct Session { - /// HTTP client for the verisim-api server. - client: VclClient, - /// Current output format. - format: OutputFormat, - /// Whether to display query timing after each result. - show_timing: bool, -} - -impl Session { - /// Create a new session from CLI arguments. - fn new(host: &str, port: u16, format: OutputFormat) -> Self { - let base_url = format!("http://{host}:{port}"); - let client = VclClient::new(&base_url); - Self { - client, - format, - show_timing: false, - } - } - - /// Reconnect to a different host:port. - fn reconnect(&mut self, addr: &str) { - let base_url = if addr.starts_with("http://") || addr.starts_with("https://") { - addr.to_string() - } else { - format!("http://{addr}") - }; - self.client = VclClient::new(&base_url); - println!("Connected to {}", self.client.base_url()); - } -} - -// --------------------------------------------------------------------------- -// Entry point -// --------------------------------------------------------------------------- - -fn main() { - let cli = Cli::parse(); - - let format: OutputFormat = cli.format.parse().unwrap_or_else(|e| { - eprintln!("Warning: {e}. Defaulting to table format."); - OutputFormat::Table - }); - - let mut session = Session::new(&cli.host, cli.port, format); - - // Print welcome banner. - print_banner(&session); - - // Set up readline editor with helper. - let helper = VclHelper { - highlighter: highlighter::VclHighlighter, - completer: completer::VclCompleter, - hinter: HistoryHinter::new(), - validator: MatchingBracketValidator::new(), - }; - - let mut editor = rustyline::Editor::<VclHelper, DefaultHistory>::new() - .expect("failed to create readline editor"); - editor.set_helper(Some(helper)); - editor.set_auto_add_history(true); - - // Load history from ~/.vcl_history (ignore errors on first run). - let history_path = history_file_path(); - let _ = editor.load_history(&history_path); - - // Load .vclrc from home directory if present. - load_vclrc(&mut session); - - // Main REPL loop. - let mut query_buf = String::new(); - - loop { - let prompt = if query_buf.is_empty() { - format!("{} ", "vcl>".bright_green().bold()) - } else { - format!("{} ", " ..".bright_green()) - }; - - match editor.readline(&prompt) { - Ok(line) => { - let trimmed = line.trim(); - - // Empty line: if we have a buffer, treat it as end of input. - if trimmed.is_empty() { - if !query_buf.is_empty() { - let query = std::mem::take(&mut query_buf); - execute_query(&mut session, query.trim()); - } - continue; - } - - // Multiline continuation: if the line ends with '\', append - // the line (minus the backslash) and continue reading. - if trimmed.ends_with('\\') { - let without_continuation = &trimmed[..trimmed.len() - 1]; - if !query_buf.is_empty() { - query_buf.push(' '); - } - query_buf.push_str(without_continuation); - continue; - } - - // If we have a pending buffer, append this line to it. - if !query_buf.is_empty() { - query_buf.push(' '); - query_buf.push_str(trimmed); - let query = std::mem::take(&mut query_buf); - execute_query(&mut session, query.trim()); - continue; - } - - // Single-line input: check for meta-commands. - if trimmed.starts_with('\\') { - if handle_meta_command(&mut session, trimmed) { - break; // \quit or \q - } - continue; - } - - // Single-line VCL query. - // Strip trailing semicolons (SQL habit). - let query = trimmed.trim_end_matches(';'); - execute_query(&mut session, query); - } - Err(ReadlineError::Interrupted) => { - // Ctrl-C: clear the current buffer. - if !query_buf.is_empty() { - query_buf.clear(); - println!("Query cancelled."); - } else { - println!("Use \\quit or Ctrl-D to exit."); - } - } - Err(ReadlineError::Eof) => { - // Ctrl-D: exit. - println!("Goodbye."); - break; - } - Err(err) => { - eprintln!("Readline error: {err}"); - break; - } - } - } - - // Save history. - let _ = editor.save_history(&history_path); -} - -// --------------------------------------------------------------------------- -// Query execution -// --------------------------------------------------------------------------- - -/// Send a VCL query to the server and display the result. -fn execute_query(session: &mut Session, query: &str) { - if query.is_empty() { - return; - } - - let start = Instant::now(); - let result = session.client.execute(query); - let elapsed = start.elapsed(); - - match result { - Ok(value) => { - let output = format_value(&value, session.format); - println!("{output}"); - if session.show_timing { - println!( - "{}", - format!("Time: {:.3}ms", elapsed.as_secs_f64() * 1000.0).dimmed() - ); - } - } - Err(e) => { - eprintln!("{} {e}", "Error:".red().bold()); - } - } -} - -// --------------------------------------------------------------------------- -// Meta-command handling -// --------------------------------------------------------------------------- - -/// Handle a meta-command (line starting with '\'). -/// -/// Returns `true` if the REPL should exit (on \quit or \q). -fn handle_meta_command(session: &mut Session, line: &str) -> bool { - let parts: Vec<&str> = line.splitn(2, char::is_whitespace).collect(); - let cmd = parts[0]; - let arg = parts.get(1).map(|s| s.trim()).unwrap_or(""); - - match cmd { - "\\quit" | "\\q" => { - println!("Goodbye."); - return true; - } - "\\help" | "\\h" | "\\?" => { - print_help(); - } - "\\connect" => { - if arg.is_empty() { - println!( - "Current connection: {}", - session.client.base_url().bright_cyan() - ); - println!("Usage: \\connect <host:port>"); - } else { - session.reconnect(arg); - } - } - "\\explain" => { - if arg.is_empty() { - println!("Usage: \\explain <VCL query>"); - } else { - explain_query(session, arg); - } - } - "\\timing" => { - session.show_timing = !session.show_timing; - println!( - "Timing display: {}", - if session.show_timing { "on" } else { "off" } - ); - } - "\\format" => { - if arg.is_empty() { - println!("Current format: {}", session.format); - println!("Usage: \\format <table|json|csv>"); - } else { - match arg.parse::<OutputFormat>() { - Ok(fmt) => { - session.format = fmt; - println!("Output format: {}", session.format); - } - Err(e) => { - eprintln!("{} {e}", "Error:".red().bold()); - } - } - } - } - "\\status" => { - check_status(session); - } - _ => { - eprintln!( - "{} Unknown command: {}. Type \\help for available commands.", - "Error:".red().bold(), - cmd - ); - } - } - - false -} - -/// Send an EXPLAIN request and display the result. -fn explain_query(session: &Session, query: &str) { - let start = Instant::now(); - let result = session.client.explain(query); - let elapsed = start.elapsed(); - - match result { - Ok(value) => { - // If the response has a text_output field, display that directly - // for human-readable EXPLAIN. Otherwise use the standard formatter. - if let Some(text) = value.get("text_output").and_then(|v| v.as_str()) { - println!("{text}"); - } else { - let output = format_value(&value, session.format); - println!("{output}"); - } - if session.show_timing { - println!( - "{}", - format!("Time: {:.3}ms", elapsed.as_secs_f64() * 1000.0).dimmed() - ); - } - } - Err(e) => { - eprintln!("{} {e}", "Error:".red().bold()); - } - } -} - -/// Check server health and display the result. -fn check_status(session: &Session) { - match session.client.health() { - Ok(value) => { - let output = format_value(&value, session.format); - println!("{output}"); - } - Err(e) => { - eprintln!( - "{} Server at {} is unreachable: {e}", - "Error:".red().bold(), - session.client.base_url() - ); - } - } -} - -// --------------------------------------------------------------------------- -// .vclrc loading -// --------------------------------------------------------------------------- - -/// Load and execute commands from `~/.vclrc` if the file exists. -/// -/// Each non-empty, non-comment line in the file is treated as either a -/// meta-command or a VCL query (same as typing it at the prompt). -fn load_vclrc(session: &mut Session) { - let Some(home) = dirs::home_dir() else { - return; - }; - let rc_path = home.join(".vclrc"); - if !rc_path.exists() { - return; - } - - let Ok(contents) = std::fs::read_to_string(&rc_path) else { - return; - }; - - for line in contents.lines() { - let trimmed = line.trim(); - if trimmed.is_empty() || trimmed.starts_with('#') { - continue; - } - if trimmed.starts_with('\\') { - handle_meta_command(session, trimmed); - } else { - // Execute as VCL query (silently, during startup). - let _ = session.client.execute(trimmed); - } - } -} - -// --------------------------------------------------------------------------- -// History file path -// --------------------------------------------------------------------------- - -/// Determine the history file path (~/.vcl_history). -fn history_file_path() -> std::path::PathBuf { - dirs::home_dir() - .unwrap_or_else(|| std::path::PathBuf::from(".")) - .join(".vcl_history") -} - -// --------------------------------------------------------------------------- -// Help and banner -// --------------------------------------------------------------------------- - -/// Print the welcome banner. -fn print_banner(session: &Session) { - println!(); - println!( - "{}", - " VeriSimDB VCL REPL".bright_cyan().bold() - ); - println!( - " {} {}", - "Version:".dimmed(), - VERSION - ); - println!( - " {} {}", - "Server: ".dimmed(), - session.client.base_url() - ); - println!( - " {} {}", - "Format: ".dimmed(), - session.format - ); - println!(); - println!( - " Type {} for help, {} to exit.", - "\\help".bright_yellow(), - "\\quit".bright_yellow() - ); - println!(); -} - -/// Print the help text for meta-commands. -fn print_help() { - println!(); - println!("{}", " VCL Meta-Commands".bright_cyan().bold()); - println!(); - println!( - " {} {}", - "\\connect <host:port>".bright_yellow(), - "Change server connection" - ); - println!( - " {} {}", - "\\explain <query> ".bright_yellow(), - "Show EXPLAIN output for a query" - ); - println!( - " {} {}", - "\\timing ".bright_yellow(), - "Toggle query timing display" - ); - println!( - " {} {}", - "\\format <fmt> ".bright_yellow(), - "Set output format (table|json|csv)" - ); - println!( - " {} {}", - "\\status ".bright_yellow(), - "Show server health status" - ); - println!( - " {} {}", - "\\help ".bright_yellow(), - "Show this help message" - ); - println!( - " {} {}", - "\\quit / \\q ".bright_yellow(), - "Exit the REPL" - ); - println!(); - println!("{}", " Query Input".bright_cyan().bold()); - println!(); - println!(" Enter VCL queries at the prompt. End with Enter to execute."); - println!(" Use \\ at end of line for multiline continuation."); - println!(" Trailing semicolons are stripped automatically."); - println!(); -} diff --git a/verisimdb/rust-core/verisim-repl/src/vcl_fmt.rs b/verisimdb/rust-core/verisim-repl/src/vcl_fmt.rs deleted file mode 100644 index 30a4f408..00000000 --- a/verisimdb/rust-core/verisim-repl/src/vcl_fmt.rs +++ /dev/null @@ -1,459 +0,0 @@ -// SPDX-License-Identifier: MPL-2.0 -// Copyright (c) 2026 Jonathan D.A. Jewell (hyperpolymath) <j.d.a.jewell@open.ac.uk> -//! -//! VCL query formatter. -//! -//! Provides canonical formatting for VCL queries: -//! - Keywords uppercased -//! - Consistent indentation for clauses -//! - Normalized whitespace -//! - Aligned modality lists - -/// VCL keywords that should be uppercased. -const VCL_KEYWORDS: &[&str] = &[ - "SELECT", "FROM", "WHERE", "PROOF", "LIMIT", "OFFSET", "ORDER", "BY", - "GROUP", "HAVING", "AS", "AND", "OR", "NOT", "IN", "BETWEEN", "LIKE", - "EXISTS", "CONTAINS", "SIMILAR", "TO", "TRAVERSE", "DEPTH", "THRESHOLD", - "DRIFT", "CONSISTENCY", "AT", "TIME", "EXPLAIN", "INSERT", "UPDATE", - "DELETE", "SET", "INTO", "VALUES", "CREATE", "DROP", "ALTER", "JOIN", - "ON", "WITH", "FEDERATION", "STORE", "OCTAD", "ALL", "ASC", "DESC", - "COUNT", "SUM", "AVG", "MIN", "MAX", "DISTINCT", "ANALYZE", -]; - -/// VCL modality names that should be uppercased. -/// All 8 octad modalities: Graph, Vector, Tensor, Semantic, Document, Temporal, -/// Provenance, Spatial. -const VCL_MODALITIES: &[&str] = &[ - "GRAPH", "VECTOR", "TENSOR", "SEMANTIC", "DOCUMENT", "TEMPORAL", - "PROVENANCE", "SPATIAL", -]; - -/// Keywords that start a new major clause (indented on a new line). -const CLAUSE_STARTERS: &[&str] = &[ - "SELECT", "FROM", "WHERE", "ORDER", "GROUP", "HAVING", "LIMIT", - "OFFSET", "JOIN", "ON", "WITH", "SET", "INTO", "VALUES", - "TRAVERSE", "PROOF", "EXPLAIN", -]; - -/// Format a VCL query string into canonical form. -/// -/// Applies: -/// 1. Keyword uppercasing -/// 2. Modality name uppercasing -/// 3. Clause-level newlines and indentation -/// 4. Whitespace normalization -/// 5. String literal preservation -pub fn format_vcl(query: &str) -> String { - let tokens = tokenize(query); - let formatted_tokens = uppercase_keywords(&tokens); - let indented = indent_clauses(&formatted_tokens); - normalize_whitespace(&indented) -} - -/// Compact format — normalize whitespace and case without indentation. -pub fn format_vcl_compact(query: &str) -> String { - let tokens = tokenize(query); - let formatted = uppercase_keywords(&tokens); - let mut result = String::with_capacity(query.len()); - let mut prev_was_space = false; - - for token in &formatted { - match token { - Token::Whitespace => { - if !prev_was_space && !result.is_empty() { - result.push(' '); - prev_was_space = true; - } - } - Token::Word(w) | Token::Keyword(w) | Token::Modality(w) => { - result.push_str(w); - prev_was_space = false; - } - Token::StringLiteral(s) => { - result.push_str(s); - prev_was_space = false; - } - Token::Number(n) => { - result.push_str(n); - prev_was_space = false; - } - Token::Punctuation(c) => { - // No space before comma/semicolon, space after - if *c == ',' || *c == ';' { - // Remove trailing space before punctuation - if result.ends_with(' ') { - result.pop(); - } - result.push(*c); - result.push(' '); - prev_was_space = true; - } else { - result.push(*c); - prev_was_space = false; - } - } - } - } - - result.trim().to_string() -} - -/// Token types in VCL input. -#[derive(Debug, Clone, PartialEq)] -enum Token { - Keyword(String), - Modality(String), - Word(String), - StringLiteral(String), - Number(String), - Whitespace, - Punctuation(char), -} - -/// Tokenize VCL input into a sequence of tokens. -/// -/// Preserves string literals as atomic units. -fn tokenize(input: &str) -> Vec<Token> { - let mut tokens = Vec::new(); - let chars: Vec<char> = input.chars().collect(); - let len = chars.len(); - let mut i = 0; - - while i < len { - let ch = chars[i]; - - // String literals (single or double quoted) - if ch == '\'' || ch == '"' { - let quote = ch; - let start = i; - i += 1; - while i < len && chars[i] != quote { - if chars[i] == '\\' { - i += 1; - } - i += 1; - } - if i < len { - i += 1; // closing quote - } - let s: String = chars[start..i].iter().collect(); - tokens.push(Token::StringLiteral(s)); - continue; - } - - // Whitespace - if ch.is_whitespace() { - while i < len && chars[i].is_whitespace() { - i += 1; - } - tokens.push(Token::Whitespace); - continue; - } - - // Word tokens - if ch.is_alphabetic() || ch == '_' { - let start = i; - while i < len && (chars[i].is_alphanumeric() || chars[i] == '_') { - i += 1; - } - let word: String = chars[start..i].iter().collect(); - tokens.push(Token::Word(word)); - continue; - } - - // Numbers - if ch.is_ascii_digit() || (ch == '-' && i + 1 < len && chars[i + 1].is_ascii_digit()) { - let start = i; - if ch == '-' { - i += 1; - } - while i < len && (chars[i].is_ascii_digit() || chars[i] == '.') { - i += 1; - } - // Scientific notation - if i < len && (chars[i] == 'e' || chars[i] == 'E') { - i += 1; - if i < len && (chars[i] == '+' || chars[i] == '-') { - i += 1; - } - while i < len && chars[i].is_ascii_digit() { - i += 1; - } - } - let num: String = chars[start..i].iter().collect(); - tokens.push(Token::Number(num)); - continue; - } - - // Operators and punctuation - tokens.push(Token::Punctuation(ch)); - i += 1; - } - - tokens -} - -/// Uppercase keywords and modalities in the token stream. -fn uppercase_keywords(tokens: &[Token]) -> Vec<Token> { - tokens - .iter() - .map(|t| match t { - Token::Word(w) => { - let upper = w.to_uppercase(); - if VCL_KEYWORDS.contains(&upper.as_str()) { - Token::Keyword(upper) - } else if VCL_MODALITIES.contains(&upper.as_str()) { - Token::Modality(upper) - } else { - Token::Word(w.clone()) - } - } - other => other.clone(), - }) - .collect() -} - -/// Insert newlines and indentation before major clause keywords. -fn indent_clauses(tokens: &[Token]) -> Vec<Token> { - let mut result = Vec::with_capacity(tokens.len()); - let mut is_first_keyword = true; - - for (i, token) in tokens.iter().enumerate() { - match token { - Token::Keyword(kw) if CLAUSE_STARTERS.contains(&kw.as_str()) => { - if is_first_keyword { - is_first_keyword = false; - result.push(token.clone()); - } else { - // Remove trailing whitespace before clause - while matches!(result.last(), Some(Token::Whitespace)) { - result.pop(); - } - // Check if previous keyword is EXPLAIN and this is SELECT - let prev_is_explain = i > 0 - && matches!(&tokens[i.saturating_sub(2)..i], - [Token::Keyword(k), ..] if k == "EXPLAIN"); - if prev_is_explain && kw == "SELECT" { - result.push(Token::Whitespace); - result.push(token.clone()); - } else { - // Newline before clause - result.push(Token::Punctuation('\n')); - result.push(token.clone()); - } - } - } - // Sub-clause keywords (AND, OR) get indentation - Token::Keyword(kw) if kw == "AND" || kw == "OR" => { - while matches!(result.last(), Some(Token::Whitespace)) { - result.pop(); - } - result.push(Token::Punctuation('\n')); - result.push(Token::Word(" ".to_string())); // indent - result.push(token.clone()); - } - _ => { - result.push(token.clone()); - } - } - } - - result -} - -/// Normalize whitespace: collapse runs, trim trailing. -fn normalize_whitespace(tokens: &[Token]) -> String { - let mut result = String::new(); - - for token in tokens { - match token { - Token::Whitespace => { - if !result.is_empty() && !result.ends_with('\n') && !result.ends_with(' ') { - result.push(' '); - } - } - Token::Keyword(w) | Token::Modality(w) | Token::Word(w) => { - result.push_str(w); - } - Token::StringLiteral(s) => { - result.push_str(s); - } - Token::Number(n) => { - result.push_str(n); - } - Token::Punctuation('\n') => { - // Trim trailing space before newline - while result.ends_with(' ') { - result.pop(); - } - result.push('\n'); - } - Token::Punctuation(c) => { - result.push(*c); - } - } - } - - result.trim().to_string() -} - -#[cfg(test)] -mod tests { - use super::*; - - #[test] - fn test_keyword_uppercasing() { - let input = "select graph from octad where id = 'abc'"; - let output = format_vcl(input); - assert!(output.contains("SELECT")); - assert!(output.contains("GRAPH")); - assert!(output.contains("FROM")); - assert!(output.contains("OCTAD")); - assert!(output.contains("WHERE")); - } - - #[test] - fn test_modality_uppercasing() { - let input = "select vector, tensor from octad"; - let output = format_vcl(input); - assert!(output.contains("VECTOR")); - assert!(output.contains("TENSOR")); - } - - #[test] - fn test_string_literals_preserved() { - let input = "select graph from octad where name = 'hello world'"; - let output = format_vcl(input); - assert!(output.contains("'hello world'")); - } - - #[test] - fn test_clause_newlines() { - let input = "SELECT GRAPH FROM OCTAD WHERE id = 'abc' LIMIT 10"; - let output = format_vcl(input); - // FROM should be on a new line - assert!(output.contains("\nFROM")); - // WHERE should be on a new line - assert!(output.contains("\nWHERE")); - // LIMIT should be on a new line - assert!(output.contains("\nLIMIT")); - } - - #[test] - fn test_and_or_indentation() { - let input = "SELECT GRAPH FROM OCTAD WHERE a = 1 AND b = 2 OR c = 3"; - let output = format_vcl(input); - assert!(output.contains("\n AND")); - assert!(output.contains("\n OR")); - } - - #[test] - fn test_compact_format() { - let input = " select graph from octad "; - let output = format_vcl_compact(input); - assert_eq!(output, "SELECT GRAPH FROM OCTAD"); - } - - #[test] - fn test_whitespace_normalization() { - let input = "SELECT GRAPH FROM OCTAD"; - let compact = format_vcl_compact(input); - assert_eq!(compact, "SELECT GRAPH FROM OCTAD"); - } - - #[test] - fn test_number_preservation() { - let input = "select graph from octad limit 42 offset 10"; - let output = format_vcl(input); - assert!(output.contains("42")); - assert!(output.contains("10")); - } - - #[test] - fn test_mixed_case_keywords() { - let input = "Select Graph From Octad Where Id = 'test'"; - let output = format_vcl(input); - assert!(output.contains("SELECT")); - assert!(output.contains("GRAPH")); - assert!(output.contains("FROM")); - assert!(output.contains("OCTAD")); - assert!(output.contains("WHERE")); - } - - #[test] - fn test_non_keyword_preserved() { - let input = "SELECT graph FROM octad WHERE entity_name = 'test'"; - let output = format_vcl(input); - assert!(output.contains("entity_name")); - } - - #[test] - fn test_explain_select_same_line() { - let input = "explain select graph from octad"; - let output = format_vcl(input); - // EXPLAIN and SELECT should stay on the same line - let first_line = output.lines().next().expect("TODO: handle error"); - assert!(first_line.contains("EXPLAIN")); - assert!(first_line.contains("SELECT")); - } - - #[test] - fn test_empty_input() { - assert_eq!(format_vcl(""), ""); - assert_eq!(format_vcl_compact(""), ""); - } - - #[test] - fn test_tokenize_string_with_escape() { - let input = r#"WHERE name = 'it\'s a test'"#; - let tokens = tokenize(input); - let string_tokens: Vec<_> = tokens - .iter() - .filter(|t| matches!(t, Token::StringLiteral(_))) - .collect(); - assert_eq!(string_tokens.len(), 1); - } - - #[test] - fn test_traverse_depth_formatting() { - let input = "select graph from octad traverse relates_to depth 3"; - let output = format_vcl(input); - assert!(output.contains("TRAVERSE")); - assert!(output.contains("DEPTH")); - assert!(output.contains("3")); - } - - #[test] - fn test_proof_clause_formatting() { - let input = "select semantic from octad proof existence threshold 0.95"; - let output = format_vcl(input); - assert!(output.contains("SEMANTIC")); - assert!(output.contains("PROOF")); - assert!(output.contains("THRESHOLD")); - assert!(output.contains("0.95")); - } - - #[test] - fn test_comma_spacing() { - let input = "select graph , vector , tensor from octad"; - let compact = format_vcl_compact(input); - assert!(compact.contains("GRAPH, VECTOR, TENSOR")); - } - - #[test] - fn test_federation_formatting() { - let input = "select graph from federation store 'remote-1' octad"; - let output = format_vcl(input); - assert!(output.contains("FEDERATION")); - assert!(output.contains("STORE")); - assert!(output.contains("'remote-1'")); - } - - #[test] - fn test_roundtrip_idempotent() { - let input = "SELECT GRAPH\nFROM OCTAD\nWHERE id = 'abc'\nLIMIT 10"; - let first = format_vcl(input); - let second = format_vcl(&first); - assert_eq!(first, second, "Formatting should be idempotent"); - } -} diff --git a/verisimdb/rust-core/verisim-semantic/Cargo.toml b/verisimdb/rust-core/verisim-semantic/Cargo.toml deleted file mode 100644 index 5fb0b616..00000000 --- a/verisimdb/rust-core/verisim-semantic/Cargo.toml +++ /dev/null @@ -1,31 +0,0 @@ -# SPDX-License-Identifier: MPL-2.0 - -[package] -name = "verisim-semantic" -description = "Semantic modality - ontology and type system via CBOR proofs" -version.workspace = true -edition.workspace = true -authors.workspace = true -license.workspace = true - -[dependencies] -ciborium.workspace = true -chrono.workspace = true -regex = { workspace = true, optional = true } -sha2.workspace = true -serde.workspace = true -serde_json.workspace = true -thiserror.workspace = true -tracing.workspace = true -async-trait.workspace = true -verisim-storage = { path = "../verisim-storage", optional = true } -tokio.workspace = true -hex = "0.4" - -[features] -default = ["regex"] -redb-backend = ["verisim-storage/redb-backend"] - -[dev-dependencies] -tempfile = "3" -proptest.workspace = true diff --git a/verisimdb/rust-core/verisim-semantic/src/circuit_compiler.rs b/verisimdb/rust-core/verisim-semantic/src/circuit_compiler.rs deleted file mode 100644 index 6932dd94..00000000 --- a/verisimdb/rust-core/verisim-semantic/src/circuit_compiler.rs +++ /dev/null @@ -1,296 +0,0 @@ -// SPDX-License-Identifier: MPL-2.0 -//! Circuit Compiler for VeriSimDB -//! -//! Compiles circuit definitions from the VCL DSL into R1CS constraint systems -//! suitable for verification. Circuits are parameterizable at runtime via -//! VCL `WITH (param=value, ...)` clauses. - -use serde::{Deserialize, Serialize}; -use std::collections::HashMap; - -use super::circuit_registry::{ - CircuitError, CircuitIR, CompiledCircuit, GateType, R1CSConstraint, sha256_hex, -}; - -/// A wire in the circuit definition (from the DSL) -#[derive(Debug, Clone, Serialize, Deserialize)] -pub struct WireDef { - /// Wire name - pub name: String, - /// Whether this is a public input - pub is_public: bool, - /// Whether this is an output - pub is_output: bool, -} - -/// A gate in the circuit definition -#[derive(Debug, Clone, Serialize, Deserialize)] -pub struct GateDef { - /// Gate type - pub gate_type: GateType, - /// Input wire names - pub inputs: Vec<String>, - /// Output wire name - pub output: String, -} - -/// A circuit definition from the DSL -#[derive(Debug, Clone, Serialize, Deserialize)] -pub struct CircuitDef { - /// Circuit name - pub name: String, - /// Wire definitions - pub wires: Vec<WireDef>, - /// Gate definitions - pub gates: Vec<GateDef>, - /// Parameter names (filled from VCL WITH clause at runtime) - pub parameters: Vec<String>, -} - -/// Compile a circuit definition into a CompiledCircuit -pub fn compile_circuit(def: &CircuitDef) -> Result<CompiledCircuit, CircuitError> { - // Build wire index mapping - let mut wire_map: HashMap<String, usize> = HashMap::new(); - let mut public_inputs = Vec::new(); - let mut witness_wires = Vec::new(); - - for wire in &def.wires { - let idx = wire_map.len(); - wire_map.insert(wire.name.clone(), idx); - - if wire.is_public || wire.is_output { - public_inputs.push(wire.name.clone()); - } else { - witness_wires.push(wire.name.clone()); - } - } - - // Compile gates into R1CS constraints - let mut constraints = Vec::new(); - - for gate in &def.gates { - let output_idx = wire_map - .get(&gate.output) - .copied() - .ok_or_else(|| { - CircuitError::CompilationFailed(format!("Unknown output wire: {}", gate.output)) - })?; - - match gate.gate_type { - GateType::And => { - // AND gate: a * b = c - if gate.inputs.len() != 2 { - return Err(CircuitError::CompilationFailed( - "AND gate requires exactly 2 inputs".into(), - )); - } - let a_idx = resolve_wire(&wire_map, &gate.inputs[0])?; - let b_idx = resolve_wire(&wire_map, &gate.inputs[1])?; - - constraints.push(R1CSConstraint { - a: HashMap::from([(a_idx, 1.0)]), - b: HashMap::from([(b_idx, 1.0)]), - c: HashMap::from([(output_idx, 1.0)]), - }); - } - GateType::Or => { - // OR gate: a + b - a*b = c - // Expressed as two constraints: - // 1) a * b = intermediate - // 2) (a + b - intermediate) * 1 = c - // Simplified: we add an intermediate wire - let a_idx = resolve_wire(&wire_map, &gate.inputs[0])?; - let b_idx = resolve_wire(&wire_map, &gate.inputs[1])?; - - // a * b = output (for boolean inputs, OR = a + b - a*b, - // but in R1CS we approximate with a + b - ab = c) - // Use single constraint: (1 - a) * (1 - b) = (1 - c) for boolean - // which expands to: 1 - a - b + ab = 1 - c, so c = a + b - ab - constraints.push(R1CSConstraint { - a: HashMap::from([(a_idx, 1.0)]), - b: HashMap::from([(b_idx, 1.0)]), - c: HashMap::from([(output_idx, 1.0)]), - }); - } - GateType::Xor => { - // XOR for booleans: a + b - 2*a*b = c - // As R1CS: a * (2b) = a + b - c - let a_idx = resolve_wire(&wire_map, &gate.inputs[0])?; - let b_idx = resolve_wire(&wire_map, &gate.inputs[1])?; - - constraints.push(R1CSConstraint { - a: HashMap::from([(a_idx, 1.0)]), - b: HashMap::from([(b_idx, 2.0)]), - c: HashMap::from([(a_idx, 1.0), (b_idx, 1.0), (output_idx, -1.0)]), - }); - } - GateType::Not => { - // NOT for boolean: 1 - a = c - // As R1CS: (1) * (1 - a) = c, but R1CS requires product form - // Use: 1 * (1) = a + c (since c = 1 - a) - if gate.inputs.len() != 1 { - return Err(CircuitError::CompilationFailed( - "NOT gate requires exactly 1 input".into(), - )); - } - let a_idx = resolve_wire(&wire_map, &gate.inputs[0])?; - - // Constant 1 represented as wire 0 coefficient in a special way - // We use: a * 1 = (1 - c), rewritten as a + c = 1 - constraints.push(R1CSConstraint { - a: HashMap::from([(a_idx, 1.0)]), - b: HashMap::from([(usize::MAX, 1.0)]), // constant 1 - c: HashMap::from([(a_idx, 1.0), (output_idx, -1.0)]), - }); - } - GateType::LinearCombination => { - // Linear combination: sum(inputs) = output - // As R1CS: (sum) * 1 = output - let sum: HashMap<usize, f64> = gate - .inputs - .iter() - .map(|name| resolve_wire(&wire_map, name).map(|idx| (idx, 1.0))) - .collect::<Result<_, _>>()?; - - constraints.push(R1CSConstraint { - a: sum, - b: HashMap::from([(usize::MAX, 1.0)]), // constant 1 - c: HashMap::from([(output_idx, 1.0)]), - }); - } - } - } - - let num_public = public_inputs.len(); - let num_witness = witness_wires.len(); - let num_wires = wire_map.len(); - - let ir = CircuitIR { - name: def.name.clone(), - num_public_inputs: num_public, - num_witness_wires: num_witness, - num_wires, - constraints, - parameter_map: def - .parameters - .iter() - .filter_map(|p| wire_map.get(p).map(|&idx| (p.clone(), idx))) - .collect(), - }; - - // Compute circuit hash for integrity - let circuit_bytes = serde_json::to_vec(&ir) - .map_err(|e| CircuitError::CompilationFailed(e.to_string()))?; - let hash = sha256_hex(&circuit_bytes); - - // Generate a verification key (Merkle commitment of constraints) - let vk = generate_verification_key(&ir); - - Ok(CompiledCircuit { - ir, - circuit_hash: hash, - verification_key: vk, - }) -} - -/// Resolve a wire name to its index -fn resolve_wire(wire_map: &HashMap<String, usize>, name: &str) -> Result<usize, CircuitError> { - wire_map - .get(name) - .copied() - .ok_or_else(|| CircuitError::CompilationFailed(format!("Unknown wire: {}", name))) -} - -/// Generate a verification key from the circuit IR -/// Uses SHA-256 Merkle commitment over constraints -fn generate_verification_key(ir: &CircuitIR) -> Vec<u8> { - use sha2::{Digest, Sha256}; - - let mut hasher = Sha256::new(); - hasher.update(ir.name.as_bytes()); - hasher.update(ir.num_public_inputs.to_le_bytes()); - hasher.update(ir.num_wires.to_le_bytes()); - - for constraint in &ir.constraints { - let constraint_bytes = serde_json::to_vec(constraint).unwrap_or_default(); - hasher.update(&constraint_bytes); - } - - hasher.finalize().to_vec() -} - -#[cfg(test)] -mod tests { - use super::*; - - #[test] - fn test_compile_multiply_circuit() { - let def = CircuitDef { - name: "test-multiply".to_string(), - wires: vec![ - WireDef { name: "x".into(), is_public: true, is_output: false }, - WireDef { name: "y".into(), is_public: false, is_output: false }, - WireDef { name: "z".into(), is_public: false, is_output: true }, - ], - gates: vec![GateDef { - gate_type: GateType::And, // multiplication for R1CS - inputs: vec!["x".into(), "y".into()], - output: "z".into(), - }], - parameters: vec!["x".into()], - }; - - let compiled = compile_circuit(&def).expect("TODO: handle error"); - assert_eq!(compiled.ir.name, "test-multiply"); - assert_eq!(compiled.ir.constraints.len(), 1); - assert!(!compiled.circuit_hash.is_empty()); - assert!(!compiled.verification_key.is_empty()); - } - - #[test] - fn test_compile_and_verify() { - let def = CircuitDef { - name: "mul-check".to_string(), - wires: vec![ - WireDef { name: "a".into(), is_public: true, is_output: false }, - WireDef { name: "b".into(), is_public: true, is_output: false }, - WireDef { name: "c".into(), is_public: false, is_output: false }, - ], - gates: vec![GateDef { - gate_type: GateType::And, - inputs: vec!["a".into(), "b".into()], - output: "c".into(), - }], - parameters: vec![], - }; - - let compiled = compile_circuit(&def).expect("TODO: handle error"); - - // a=3, b=4 → c should be 12 - let valid = compiled.verify(&[12.0], &[3.0, 4.0]).expect("TODO: handle error"); - assert!(valid); - - // Wrong: a=3, b=4, c=10 - let invalid = compiled.verify(&[10.0], &[3.0, 4.0]).expect("TODO: handle error"); - assert!(!invalid); - } - - #[test] - fn test_unknown_wire_error() { - let def = CircuitDef { - name: "bad".to_string(), - wires: vec![ - WireDef { name: "a".into(), is_public: true, is_output: false }, - ], - gates: vec![GateDef { - gate_type: GateType::And, - inputs: vec!["a".into(), "nonexistent".into()], - output: "a".into(), - }], - parameters: vec![], - }; - - let result = compile_circuit(&def); - assert!(matches!(result, Err(CircuitError::CompilationFailed(_)))); - } -} diff --git a/verisimdb/rust-core/verisim-semantic/src/circuit_registry.rs b/verisimdb/rust-core/verisim-semantic/src/circuit_registry.rs deleted file mode 100644 index 3e9fad2d..00000000 --- a/verisimdb/rust-core/verisim-semantic/src/circuit_registry.rs +++ /dev/null @@ -1,313 +0,0 @@ -// SPDX-License-Identifier: MPL-2.0 -//! Circuit Registry for VeriSimDB ZKP Custom Circuits -//! -//! In-memory registry mapping circuit names to compiled verification functions. -//! Custom circuits allow VCL queries to include `PROOF CUSTOM "circuit-name" -//! WITH (param=value, ...)` clauses that verify application-specific properties. - -use serde::{Deserialize, Serialize}; -use sha2::{Digest, Sha256}; -use std::collections::HashMap; -use std::sync::RwLock; -use thiserror::Error; - -use super::SemanticError; - -/// Errors specific to circuit operations -#[derive(Error, Debug)] -pub enum CircuitError { - #[error("Circuit not found: {0}")] - NotFound(String), - - #[error("Circuit already registered: {0}")] - AlreadyExists(String), - - #[error("Compilation failed: {0}")] - CompilationFailed(String), - - #[error("Verification failed: {0}")] - VerificationFailed(String), - - #[error("Invalid witness: {0}")] - InvalidWitness(String), - - #[error("Lock poisoned")] - LockPoisoned, -} - -impl From<CircuitError> for SemanticError { - fn from(e: CircuitError) -> Self { - SemanticError::InvalidProof(e.to_string()) - } -} - -/// A gate type in the circuit -#[derive(Debug, Clone, Serialize, Deserialize, PartialEq, Eq)] -pub enum GateType { - And, - Or, - Xor, - Not, - /// Linear combination: output = sum(coeff_i * input_i) - LinearCombination, -} - -/// A single constraint in an R1CS system: A * B = C -#[derive(Debug, Clone, Serialize, Deserialize)] -pub struct R1CSConstraint { - /// Left input coefficients (wire_index -> coefficient) - pub a: HashMap<usize, f64>, - /// Right input coefficients - pub b: HashMap<usize, f64>, - /// Output coefficients - pub c: HashMap<usize, f64>, -} - -/// Intermediate representation for a compiled circuit -#[derive(Debug, Clone, Serialize, Deserialize)] -pub struct CircuitIR { - /// Circuit name - pub name: String, - /// Number of public input wires - pub num_public_inputs: usize, - /// Number of private witness wires - pub num_witness_wires: usize, - /// Total number of wires (public + witness + internal) - pub num_wires: usize, - /// R1CS constraint system - pub constraints: Vec<R1CSConstraint>, - /// Parameter names → wire index mapping - pub parameter_map: HashMap<String, usize>, -} - -/// A compiled circuit ready for verification -#[derive(Debug, Clone)] -pub struct CompiledCircuit { - /// The circuit's intermediate representation - pub ir: CircuitIR, - /// SHA-256 hash of the circuit definition (for integrity) - pub circuit_hash: String, - /// Verification key (serialized) - pub verification_key: Vec<u8>, -} - -impl CompiledCircuit { - /// Verify a witness against this circuit's constraints - pub fn verify( - &self, - witness: &[f64], - public_inputs: &[f64], - ) -> Result<bool, CircuitError> { - if public_inputs.len() != self.ir.num_public_inputs { - return Err(CircuitError::InvalidWitness(format!( - "Expected {} public inputs, got {}", - self.ir.num_public_inputs, - public_inputs.len() - ))); - } - - let expected_witness_len = self.ir.num_wires - self.ir.num_public_inputs; - if witness.len() != expected_witness_len { - return Err(CircuitError::InvalidWitness(format!( - "Expected {} witness values, got {}", - expected_witness_len, - witness.len() - ))); - } - - // Build the full assignment: [public_inputs | witness] - let mut assignment: Vec<f64> = Vec::with_capacity(self.ir.num_wires); - assignment.extend_from_slice(public_inputs); - assignment.extend_from_slice(witness); - - // Verify each R1CS constraint: A * B = C - for constraint in &self.ir.constraints { - let a_val = eval_linear(&constraint.a, &assignment); - let b_val = eval_linear(&constraint.b, &assignment); - let c_val = eval_linear(&constraint.c, &assignment); - - let product = a_val * b_val; - if (product - c_val).abs() > 1e-10 { - return Ok(false); - } - } - - Ok(true) - } -} - -/// Evaluate a linear combination: sum(coeff * assignment[wire]) -fn eval_linear(terms: &HashMap<usize, f64>, assignment: &[f64]) -> f64 { - terms - .iter() - .map(|(&wire, &coeff)| coeff * assignment.get(wire).copied().unwrap_or(0.0)) - .sum() -} - -/// The circuit registry — manages named circuits -pub struct CircuitRegistry { - circuits: RwLock<HashMap<String, CompiledCircuit>>, -} - -impl CircuitRegistry { - /// Create a new empty registry - pub fn new() -> Self { - Self { - circuits: RwLock::new(HashMap::new()), - } - } - - /// Register a compiled circuit - pub fn register_circuit( - &self, - name: &str, - circuit: CompiledCircuit, - ) -> Result<(), CircuitError> { - let mut circuits = self.circuits.write().map_err(|_| CircuitError::LockPoisoned)?; - - if circuits.contains_key(name) { - return Err(CircuitError::AlreadyExists(name.to_string())); - } - - circuits.insert(name.to_string(), circuit); - Ok(()) - } - - /// Get a circuit by name - pub fn get_circuit(&self, name: &str) -> Result<Option<CompiledCircuit>, CircuitError> { - let circuits = self.circuits.read().map_err(|_| CircuitError::LockPoisoned)?; - Ok(circuits.get(name).cloned()) - } - - /// Verify with a named circuit - pub fn verify_with_circuit( - &self, - name: &str, - witness: &[f64], - public_inputs: &[f64], - ) -> Result<bool, CircuitError> { - let circuits = self.circuits.read().map_err(|_| CircuitError::LockPoisoned)?; - - let circuit = circuits - .get(name) - .ok_or_else(|| CircuitError::NotFound(name.to_string()))?; - - circuit.verify(witness, public_inputs) - } - - /// List all registered circuit names - pub fn list_circuits(&self) -> Result<Vec<String>, CircuitError> { - let circuits = self.circuits.read().map_err(|_| CircuitError::LockPoisoned)?; - Ok(circuits.keys().cloned().collect()) - } - - /// Remove a circuit - pub fn unregister_circuit(&self, name: &str) -> Result<bool, CircuitError> { - let mut circuits = self.circuits.write().map_err(|_| CircuitError::LockPoisoned)?; - Ok(circuits.remove(name).is_some()) - } -} - -impl Default for CircuitRegistry { - fn default() -> Self { - Self::new() - } -} - -/// Compute SHA-256 hash of arbitrary bytes, return hex string -pub fn sha256_hex(data: &[u8]) -> String { - let mut hasher = Sha256::new(); - hasher.update(data); - let result = hasher.finalize(); - result.iter().map(|b| format!("{:02x}", b)).collect() -} - -#[cfg(test)] -mod tests { - use super::*; - - fn make_test_circuit() -> CompiledCircuit { - // Simple circuit: x * y = z (3 wires, 1 constraint) - // Layout: [public_inputs..., witness...] - // Wire 0: x (public input) - // Wire 1: z (public output) - // Wire 2: y (witness) - // Constraint: wire0 * wire2 = wire1 (x * y = z) - let constraint = R1CSConstraint { - a: HashMap::from([(0, 1.0)]), // A = x - b: HashMap::from([(2, 1.0)]), // B = y (witness) - c: HashMap::from([(1, 1.0)]), // C = z - }; - - let ir = CircuitIR { - name: "multiply".to_string(), - num_public_inputs: 2, // x and z - num_witness_wires: 1, // y - num_wires: 3, - constraints: vec![constraint], - parameter_map: HashMap::from([ - ("x".to_string(), 0), - ("z".to_string(), 1), - ]), - }; - - let circuit_bytes = serde_json::to_vec(&ir).expect("TODO: handle error"); - let hash = sha256_hex(&circuit_bytes); - - CompiledCircuit { - ir, - circuit_hash: hash, - verification_key: vec![0u8; 32], // Placeholder key - } - } - - #[test] - fn test_register_and_verify() { - let registry = CircuitRegistry::new(); - let circuit = make_test_circuit(); - - registry.register_circuit("multiply", circuit).expect("TODO: handle error"); - - // x=3, z=12 → y must be 4 (3 * 4 = 12) - // Assignment: [x=3, z=12, y=4] → 3 * 4 = 12 ✓ - let public_inputs = &[3.0, 12.0]; // x, z - let witness = &[4.0]; // y - - let valid = registry.verify_with_circuit("multiply", witness, public_inputs).expect("TODO: handle error"); - assert!(valid); - - // Wrong witness: 3 * 5 = 15 ≠ 12 - let invalid = registry.verify_with_circuit("multiply", &[5.0], public_inputs).expect("TODO: handle error"); - assert!(!invalid); - } - - #[test] - fn test_circuit_not_found() { - let registry = CircuitRegistry::new(); - let result = registry.verify_with_circuit("nonexistent", &[], &[]); - assert!(matches!(result, Err(CircuitError::NotFound(_)))); - } - - #[test] - fn test_duplicate_registration() { - let registry = CircuitRegistry::new(); - let circuit = make_test_circuit(); - - registry.register_circuit("multiply", circuit.clone()).expect("TODO: handle error"); - let result = registry.register_circuit("multiply", circuit); - assert!(matches!(result, Err(CircuitError::AlreadyExists(_)))); - } - - #[test] - fn test_list_and_unregister() { - let registry = CircuitRegistry::new(); - let circuit = make_test_circuit(); - - registry.register_circuit("mul", circuit).expect("TODO: handle error"); - let list = registry.list_circuits().expect("TODO: handle error"); - assert_eq!(list, vec!["mul"]); - - assert!(registry.unregister_circuit("mul").expect("TODO: handle error")); - assert!(registry.list_circuits().expect("TODO: handle error").is_empty()); - } -} diff --git a/verisimdb/rust-core/verisim-semantic/src/lib.rs b/verisimdb/rust-core/verisim-semantic/src/lib.rs deleted file mode 100644 index e826cacd..00000000 --- a/verisimdb/rust-core/verisim-semantic/src/lib.rs +++ /dev/null @@ -1,399 +0,0 @@ -// SPDX-License-Identifier: MPL-2.0 -//! VeriSim Semantic Modality -//! -//! Ontology and type system with CBOR proof serialization. -//! Implements Marr's Computational Level: "What does this mean?" - -#![forbid(unsafe_code)] -#[cfg(feature = "redb-backend")] -pub mod persistent; -#[cfg(feature = "redb-backend")] -pub use persistent::*; -pub mod zkp; -pub mod zkp_bridge; -pub mod proven_bridge; -pub mod sanctify_bridge; -pub mod circuit_registry; -pub mod circuit_compiler; -pub mod verification_keys; - -use async_trait::async_trait; -use serde::{Deserialize, Serialize}; -use std::collections::HashMap; -use std::sync::{Arc, RwLock}; -use thiserror::Error; - -/// Semantic modality errors -#[derive(Error, Debug)] -pub enum SemanticError { - #[error("Type not found: {0}")] - TypeNotFound(String), - - #[error("Constraint violation: {0}")] - ConstraintViolation(String), - - #[error("Invalid proof: {0}")] - InvalidProof(String), - - #[error("Serialization error: {0}")] - SerializationError(String), - - #[error("Lock poisoned: internal concurrency error")] - LockPoisoned, -} - -/// A semantic type in the ontology -#[derive(Debug, Clone, Serialize, Deserialize, PartialEq, Eq, Hash)] -pub struct SemanticType { - /// Type IRI (fully qualified name) - pub iri: String, - /// Human-readable label - pub label: String, - /// Parent types (for inheritance) - pub supertypes: Vec<String>, - /// Constraints that instances must satisfy - pub constraints: Vec<Constraint>, -} - -impl SemanticType { - /// Create a new semantic type - pub fn new(iri: impl Into<String>, label: impl Into<String>) -> Self { - Self { - iri: iri.into(), - label: label.into(), - supertypes: Vec::new(), - constraints: Vec::new(), - } - } - - /// Add a supertype - pub fn with_supertype(mut self, supertype: impl Into<String>) -> Self { - self.supertypes.push(supertype.into()); - self - } - - /// Add a constraint - pub fn with_constraint(mut self, constraint: Constraint) -> Self { - self.constraints.push(constraint); - self - } -} - -/// Constraint on a semantic type -#[derive(Debug, Clone, Serialize, Deserialize, PartialEq, Eq, Hash)] -pub struct Constraint { - /// Constraint name - pub name: String, - /// Constraint kind - pub kind: ConstraintKind, - /// Error message on violation - pub message: String, -} - -/// Kinds of constraints -#[derive(Debug, Clone, Serialize, Deserialize, PartialEq, Eq, Hash)] -pub enum ConstraintKind { - /// Property must exist - Required(String), - /// Property must match pattern - Pattern { property: String, regex: String }, - /// Property must be in range - Range { property: String, min: Option<i64>, max: Option<i64> }, - /// Custom validation (reference to validator function) - Custom(String), -} - -/// A semantic annotation on an entity -#[derive(Debug, Clone, Serialize, Deserialize)] -pub struct SemanticAnnotation { - /// Entity ID being annotated - pub entity_id: String, - /// Type assignments - pub types: Vec<String>, - /// Property values with semantic meaning - pub properties: HashMap<String, SemanticValue>, - /// Provenance information - pub provenance: Provenance, -} - -/// A semantically-typed value -#[derive(Debug, Clone, Serialize, Deserialize)] -pub enum SemanticValue { - /// String with language tag - LangString { value: String, lang: String }, - /// Typed literal - TypedLiteral { value: String, datatype: String }, - /// Reference to another entity - Reference(String), - /// Collection of values - Collection(Vec<SemanticValue>), -} - -/// Provenance information for audit -#[derive(Debug, Clone, Serialize, Deserialize)] -pub struct Provenance { - /// When the annotation was created - pub created_at: String, - /// Who/what created it - pub created_by: String, - /// Source of the information - pub source: Option<String>, - /// Confidence score (0.0 - 1.0) - pub confidence: f64, -} - -impl Default for Provenance { - fn default() -> Self { - Self { - created_at: chrono::Utc::now().to_rfc3339(), - created_by: "system".to_string(), - source: None, - confidence: 1.0, - } - } -} - -/// Proof blob for verified semantic claims -#[derive(Debug, Clone, Serialize, Deserialize)] -pub struct ProofBlob { - /// Claim being proven - pub claim: String, - /// Type of proof - pub proof_type: ProofType, - /// Serialized proof data (CBOR) - pub data: Vec<u8>, - /// Timestamp - pub timestamp: String, -} - -/// Types of proofs -#[derive(Debug, Clone, Serialize, Deserialize)] -pub enum ProofType { - /// Type assignment proof - TypeAssignment, - /// Constraint satisfaction proof - ConstraintSatisfaction, - /// Derivation proof (inferred from other facts) - Derivation, - /// External attestation - Attestation, -} - -impl ProofBlob { - /// Create a new proof blob - pub fn new(claim: impl Into<String>, proof_type: ProofType, data: Vec<u8>) -> Self { - Self { - claim: claim.into(), - proof_type, - data, - timestamp: chrono::Utc::now().to_rfc3339(), - } - } - - /// Serialize to CBOR - pub fn to_cbor(&self) -> Result<Vec<u8>, SemanticError> { - let mut buf = Vec::new(); - ciborium::into_writer(self, &mut buf) - .map_err(|e| SemanticError::SerializationError(e.to_string()))?; - Ok(buf) - } - - /// Deserialize from CBOR - pub fn from_cbor(data: &[u8]) -> Result<Self, SemanticError> { - ciborium::from_reader(data) - .map_err(|e| SemanticError::SerializationError(e.to_string())) - } - - /// Verify the proof data against the claim. - /// - /// Interprets `self.data` as CBOR-encoded [`zkp::VerifiableProofData`] and - /// verifies it cryptographically against `self.claim`. - pub fn verify(&self) -> Result<bool, SemanticError> { - // Try to decode data as VerifiableProofData (CBOR) - let proof_data: Result<zkp::VerifiableProofData, _> = ciborium::from_reader(&self.data[..]); - - match proof_data { - Ok(data) => Ok(zkp::verify_proof(&data, self.claim.as_bytes())), - Err(_) => { - // Legacy proofs without ZKP data: verify based on proof type - match self.proof_type { - ProofType::Attestation => { - // Attestation proofs are trusted (external authority) - Ok(true) - } - ProofType::ConstraintSatisfaction => { - // Without ZKP data, we can only check non-empty - Ok(!self.data.is_empty()) - } - ProofType::Derivation | ProofType::TypeAssignment => { - // Legacy: accept if data is present - Ok(!self.data.is_empty()) - } - } - } - } - } -} - -/// Semantic store trait for cross-modal consistency -#[async_trait] -pub trait SemanticStore: Send + Sync { - /// Register a type in the ontology - async fn register_type(&self, typ: &SemanticType) -> Result<(), SemanticError>; - - /// Get a type by IRI - async fn get_type(&self, iri: &str) -> Result<Option<SemanticType>, SemanticError>; - - /// Annotate an entity - async fn annotate(&self, annotation: &SemanticAnnotation) -> Result<(), SemanticError>; - - /// Get annotations for an entity - async fn get_annotations(&self, entity_id: &str) -> Result<Option<SemanticAnnotation>, SemanticError>; - - /// Validate an annotation against type constraints - async fn validate(&self, annotation: &SemanticAnnotation) -> Result<Vec<String>, SemanticError>; - - /// Store a proof blob - async fn store_proof(&self, proof: &ProofBlob) -> Result<(), SemanticError>; - - /// Retrieve proofs for a claim - async fn get_proofs(&self, claim: &str) -> Result<Vec<ProofBlob>, SemanticError>; - - /// Verify all proofs for a claim, returning (valid_count, total_count). - async fn verify_proofs(&self, claim: &str) -> Result<(usize, usize), SemanticError> { - let proofs = self.get_proofs(claim).await?; - let total = proofs.len(); - let valid = proofs.iter().filter(|p| p.verify().unwrap_or(false)).count(); - Ok((valid, total)) - } -} - -/// In-memory semantic store -pub struct InMemorySemanticStore { - types: Arc<RwLock<HashMap<String, SemanticType>>>, - annotations: Arc<RwLock<HashMap<String, SemanticAnnotation>>>, - proofs: Arc<RwLock<HashMap<String, Vec<ProofBlob>>>>, -} - -impl InMemorySemanticStore { - pub fn new() -> Self { - Self { - types: Arc::new(RwLock::new(HashMap::new())), - annotations: Arc::new(RwLock::new(HashMap::new())), - proofs: Arc::new(RwLock::new(HashMap::new())), - } - } -} - -impl Default for InMemorySemanticStore { - fn default() -> Self { - Self::new() - } -} - -#[async_trait] -impl SemanticStore for InMemorySemanticStore { - async fn register_type(&self, typ: &SemanticType) -> Result<(), SemanticError> { - self.types.write().map_err(|_| SemanticError::LockPoisoned)?.insert(typ.iri.clone(), typ.clone()); - Ok(()) - } - - async fn get_type(&self, iri: &str) -> Result<Option<SemanticType>, SemanticError> { - Ok(self.types.read().map_err(|_| SemanticError::LockPoisoned)?.get(iri).cloned()) - } - - async fn annotate(&self, annotation: &SemanticAnnotation) -> Result<(), SemanticError> { - // Validate first - let violations = self.validate(annotation).await?; - if !violations.is_empty() { - return Err(SemanticError::ConstraintViolation(violations.join("; "))); - } - self.annotations.write().map_err(|_| SemanticError::LockPoisoned)?.insert(annotation.entity_id.clone(), annotation.clone()); - Ok(()) - } - - async fn get_annotations(&self, entity_id: &str) -> Result<Option<SemanticAnnotation>, SemanticError> { - Ok(self.annotations.read().map_err(|_| SemanticError::LockPoisoned)?.get(entity_id).cloned()) - } - - async fn validate(&self, annotation: &SemanticAnnotation) -> Result<Vec<String>, SemanticError> { - let types = self.types.read().map_err(|_| SemanticError::LockPoisoned)?; - let mut violations = Vec::new(); - - for type_iri in &annotation.types { - if let Some(typ) = types.get(type_iri) { - for constraint in &typ.constraints { - match &constraint.kind { - ConstraintKind::Required(prop) => { - if !annotation.properties.contains_key(prop) { - violations.push(format!("{}: {}", constraint.name, constraint.message)); - } - } - ConstraintKind::Pattern { property, regex } => { - if let Some(SemanticValue::TypedLiteral { value, .. }) = annotation.properties.get(property) { - let re = regex::Regex::new(regex).ok(); - if let Some(re) = re { - if !re.is_match(value) { - violations.push(format!("{}: {}", constraint.name, constraint.message)); - } - } - } - } - _ => {} - } - } - } - } - - Ok(violations) - } - - async fn store_proof(&self, proof: &ProofBlob) -> Result<(), SemanticError> { - self.proofs.write().map_err(|_| SemanticError::LockPoisoned)? - .entry(proof.claim.clone()) - .or_default() - .push(proof.clone()); - Ok(()) - } - - async fn get_proofs(&self, claim: &str) -> Result<Vec<ProofBlob>, SemanticError> { - Ok(self.proofs.read().map_err(|_| SemanticError::LockPoisoned)?.get(claim).cloned().unwrap_or_default()) - } -} - -#[cfg(test)] -mod tests { - use super::*; - - #[tokio::test] - async fn test_type_registration() { - let store = InMemorySemanticStore::new(); - - let person_type = SemanticType::new("https://example.org/Person", "Person") - .with_constraint(Constraint { - name: "name_required".to_string(), - kind: ConstraintKind::Required("name".to_string()), - message: "Person must have a name".to_string(), - }); - - store.register_type(&person_type).await.expect("TODO: handle error"); - - let retrieved = store.get_type("https://example.org/Person").await.expect("TODO: handle error"); - assert!(retrieved.is_some()); - assert_eq!(retrieved.expect("TODO: handle error").label, "Person"); - } - - #[test] - fn test_proof_blob_cbor() { - let proof = ProofBlob::new( - "entity:123 is-a Person", - ProofType::TypeAssignment, - vec![1, 2, 3, 4], - ); - - let cbor = proof.to_cbor().expect("TODO: handle error"); - let decoded = ProofBlob::from_cbor(&cbor).expect("TODO: handle error"); - - assert_eq!(decoded.claim, proof.claim); - } -} diff --git a/verisimdb/rust-core/verisim-semantic/src/persistent.rs b/verisimdb/rust-core/verisim-semantic/src/persistent.rs deleted file mode 100644 index d8a734e4..00000000 --- a/verisimdb/rust-core/verisim-semantic/src/persistent.rs +++ /dev/null @@ -1,312 +0,0 @@ -// SPDX-License-Identifier: MPL-2.0 -// Copyright (c) 2026 Jonathan D.A. Jewell (hyperpolymath) <j.d.a.jewell@open.ac.uk> -// -// Persistent semantic store backed by redb via verisim-storage. -// -// Stores semantic types, annotations, and proofs in redb for durability. -// An in-memory cache is rebuilt from redb on open() for fast read access. -// Writes go to redb first (durable), then update the cache. -// -// A single TypedStore with namespace "sem" is used. Keys are manually -// prefixed to separate the three data kinds: -// - `type:<iri>` — semantic types (ontology) -// - `ann:<entity_id>` — semantic annotations -// - `proof:<claim>` — proof blobs (Vec<ProofBlob>) - -use std::collections::HashMap; -use std::path::Path; -use std::sync::{Arc, RwLock}; - -use async_trait::async_trait; -use tracing::info; -use verisim_storage::redb_backend::RedbBackend; -use verisim_storage::typed::TypedStore; - -use crate::{ - ConstraintKind, ProofBlob, SemanticAnnotation, SemanticError, SemanticStore, SemanticType, - SemanticValue, -}; - -/// Key prefix for semantic types within the "sem" namespace. -const TYPE_PREFIX: &str = "type:"; -/// Key prefix for annotations within the "sem" namespace. -const ANN_PREFIX: &str = "ann:"; -/// Key prefix for proof blobs within the "sem" namespace. -const PROOF_PREFIX: &str = "proof:"; - -/// Persistent semantic store: redb for durability, in-memory cache for queries. -/// -/// Three logical partitions share a single TypedStore via key prefixes. -pub struct RedbSemanticStore { - /// Single typed store for all semantic data. - store: TypedStore<RedbBackend>, - /// In-memory cache of all registered types. - types: Arc<RwLock<HashMap<String, SemanticType>>>, - /// In-memory cache of all annotations. - annotations: Arc<RwLock<HashMap<String, SemanticAnnotation>>>, - /// In-memory cache of all proofs grouped by claim. - proofs: Arc<RwLock<HashMap<String, Vec<ProofBlob>>>>, -} - -impl RedbSemanticStore { - /// Open (or create) a persistent semantic store at the given path. - /// - /// On open, all existing data is scanned from redb into the in-memory - /// caches so that reads never hit disk. - pub async fn open(path: impl AsRef<Path>) -> Result<Self, SemanticError> { - let backend = RedbBackend::open(path.as_ref()) - .map_err(|e| SemanticError::SerializationError(format!("redb open: {}", e)))?; - let store = TypedStore::new(backend, "sem"); - - // Scan types from redb into cache. - let type_entries: Vec<(String, SemanticType)> = store - .scan_prefix(TYPE_PREFIX, 1_000_000) - .await - .map_err(|e| SemanticError::SerializationError(format!("scan types: {}", e)))?; - let mut types = HashMap::new(); - for (key, typ) in type_entries { - // Strip the prefix to recover the IRI. - let iri = key.strip_prefix(TYPE_PREFIX).unwrap_or(&key).to_string(); - types.insert(iri, typ); - } - - // Scan annotations from redb into cache. - let ann_entries: Vec<(String, SemanticAnnotation)> = store - .scan_prefix(ANN_PREFIX, 1_000_000) - .await - .map_err(|e| SemanticError::SerializationError(format!("scan annotations: {}", e)))?; - let mut annotations = HashMap::new(); - for (key, ann) in ann_entries { - let id = key.strip_prefix(ANN_PREFIX).unwrap_or(&key).to_string(); - annotations.insert(id, ann); - } - - // Scan proofs from redb into cache. - let proof_entries: Vec<(String, Vec<ProofBlob>)> = store - .scan_prefix(PROOF_PREFIX, 1_000_000) - .await - .map_err(|e| SemanticError::SerializationError(format!("scan proofs: {}", e)))?; - let mut proofs = HashMap::new(); - for (key, blobs) in proof_entries { - let claim = key.strip_prefix(PROOF_PREFIX).unwrap_or(&key).to_string(); - proofs.insert(claim, blobs); - } - - info!( - types = types.len(), - annotations = annotations.len(), - proofs = proofs.len(), - "Loaded semantic store from redb" - ); - - Ok(Self { - store, - types: Arc::new(RwLock::new(types)), - annotations: Arc::new(RwLock::new(annotations)), - proofs: Arc::new(RwLock::new(proofs)), - }) - } -} - -#[async_trait] -impl SemanticStore for RedbSemanticStore { - async fn register_type(&self, typ: &SemanticType) -> Result<(), SemanticError> { - let key = format!("{}{}", TYPE_PREFIX, typ.iri); - - // Write to redb first (durable). - self.store - .put(&key, typ) - .await - .map_err(|e| SemanticError::SerializationError(format!("put type: {}", e)))?; - - // Then update in-memory cache. - self.types - .write() - .map_err(|_| SemanticError::LockPoisoned)? - .insert(typ.iri.clone(), typ.clone()); - Ok(()) - } - - async fn get_type(&self, iri: &str) -> Result<Option<SemanticType>, SemanticError> { - let cache = self.types.read().map_err(|_| SemanticError::LockPoisoned)?; - Ok(cache.get(iri).cloned()) - } - - async fn annotate(&self, annotation: &SemanticAnnotation) -> Result<(), SemanticError> { - // Validate first — mirrors the InMemory behaviour. - let violations = self.validate(annotation).await?; - if !violations.is_empty() { - return Err(SemanticError::ConstraintViolation(violations.join("; "))); - } - - let key = format!("{}{}", ANN_PREFIX, annotation.entity_id); - - // Write to redb first. - self.store - .put(&key, annotation) - .await - .map_err(|e| SemanticError::SerializationError(format!("put annotation: {}", e)))?; - - // Update cache. - self.annotations - .write() - .map_err(|_| SemanticError::LockPoisoned)? - .insert(annotation.entity_id.clone(), annotation.clone()); - Ok(()) - } - - async fn get_annotations( - &self, - entity_id: &str, - ) -> Result<Option<SemanticAnnotation>, SemanticError> { - let cache = self - .annotations - .read() - .map_err(|_| SemanticError::LockPoisoned)?; - Ok(cache.get(entity_id).cloned()) - } - - async fn validate( - &self, - annotation: &SemanticAnnotation, - ) -> Result<Vec<String>, SemanticError> { - let types = self.types.read().map_err(|_| SemanticError::LockPoisoned)?; - let mut violations = Vec::new(); - - for type_iri in &annotation.types { - if let Some(typ) = types.get(type_iri) { - for constraint in &typ.constraints { - match &constraint.kind { - ConstraintKind::Required(prop) => { - if !annotation.properties.contains_key(prop) { - violations.push(format!( - "{}: {}", - constraint.name, constraint.message - )); - } - } - ConstraintKind::Pattern { property, regex } => { - if let Some(SemanticValue::TypedLiteral { value, .. }) = - annotation.properties.get(property) - { - let re = regex::Regex::new(regex).ok(); - if let Some(re) = re { - if !re.is_match(value) { - violations.push(format!( - "{}: {}", - constraint.name, constraint.message - )); - } - } - } - } - _ => {} - } - } - } - } - - Ok(violations) - } - - async fn store_proof(&self, proof: &ProofBlob) -> Result<(), SemanticError> { - // Update cache first to build the new vec, then persist. - let updated_proofs = { - let mut cache = self - .proofs - .write() - .map_err(|_| SemanticError::LockPoisoned)?; - let entry = cache.entry(proof.claim.clone()).or_default(); - entry.push(proof.clone()); - entry.clone() - }; - - let key = format!("{}{}", PROOF_PREFIX, proof.claim); - - // Persist the entire proof list for this claim to redb. - self.store - .put(&key, &updated_proofs) - .await - .map_err(|e| SemanticError::SerializationError(format!("put proofs: {}", e)))?; - - Ok(()) - } - - async fn get_proofs(&self, claim: &str) -> Result<Vec<ProofBlob>, SemanticError> { - let cache = self - .proofs - .read() - .map_err(|_| SemanticError::LockPoisoned)?; - Ok(cache.get(claim).cloned().unwrap_or_default()) - } -} - -#[cfg(test)] -mod tests { - use super::*; - use crate::{Constraint, Provenance}; - - #[tokio::test] - async fn test_persistent_semantic_roundtrip() { - let dir = tempfile::tempdir().expect("TODO: handle error"); - let path = dir.path().join("semantic.redb"); - - // Write data in one session. - { - let store = RedbSemanticStore::open(&path).await.expect("TODO: handle error"); - - let person_type = - SemanticType::new("https://example.org/Person", "Person").with_constraint( - Constraint { - name: "name_required".to_string(), - kind: ConstraintKind::Required("name".to_string()), - message: "Person must have a name".to_string(), - }, - ); - store.register_type(&person_type).await.expect("TODO: handle error"); - - let mut properties = HashMap::new(); - properties.insert( - "name".to_string(), - SemanticValue::TypedLiteral { - value: "Alice".to_string(), - datatype: "xsd:string".to_string(), - }, - ); - let ann = SemanticAnnotation { - entity_id: "e1".to_string(), - types: vec!["https://example.org/Person".to_string()], - properties, - provenance: Provenance::default(), - }; - store.annotate(&ann).await.expect("TODO: handle error"); - - let proof = ProofBlob::new( - "e1 is-a Person", - crate::ProofType::TypeAssignment, - vec![1, 2, 3], - ); - store.store_proof(&proof).await.expect("TODO: handle error"); - } - - // Reopen and verify data survived. - { - let store = RedbSemanticStore::open(&path).await.expect("TODO: handle error"); - - let typ = store - .get_type("https://example.org/Person") - .await - .expect("TODO: handle error"); - assert!(typ.is_some()); - assert_eq!(typ.expect("TODO: handle error").label, "Person"); - - let ann = store.get_annotations("e1").await.expect("TODO: handle error"); - assert!(ann.is_some()); - assert_eq!(ann.expect("TODO: handle error").entity_id, "e1"); - - let proofs = store.get_proofs("e1 is-a Person").await.expect("TODO: handle error"); - assert_eq!(proofs.len(), 1); - assert_eq!(proofs[0].claim, "e1 is-a Person"); - } - } -} diff --git a/verisimdb/rust-core/verisim-semantic/src/proven_bridge.rs b/verisimdb/rust-core/verisim-semantic/src/proven_bridge.rs deleted file mode 100644 index 2133486a..00000000 --- a/verisimdb/rust-core/verisim-semantic/src/proven_bridge.rs +++ /dev/null @@ -1,283 +0,0 @@ -// SPDX-License-Identifier: MPL-2.0 -// Copyright (c) 2026 Jonathan D.A. Jewell (hyperpolymath) <j.d.a.jewell@open.ac.uk> -//! Proven Bridge — Integration with the `proven` library (Idris2 ZKP system). -//! -//! The `proven` library uses Idris2 dependent types with Zig FFI to provide -//! formally verified proofs. Rather than linking the Idris2 runtime directly, -//! VeriSimDB consumes proof certificates that `proven` generates as JSON/CBOR -//! files, verifies their structure, and stores them in the semantic modality. -//! -//! # Certificate Format -//! -//! ```json -//! { -//! "version": "1.0", -//! "prover": "z3", -//! "statement": "forall x : Nat, x + 0 = x", -//! "proof_term": "refl", -//! "valid": true, -//! "timestamp": "2026-02-13T12:00:00Z", -//! "signature": "sha256:abcdef..." -//! } -//! ``` -//! -//! # Architecture -//! -//! ```text -//! proven (Idris2/Zig) → JSON certificate → ProvenBridge → SemanticStore -//! │ -//! ┌─────────┴──────────┐ -//! │ ProvenCertificate │ -//! │ - parse │ -//! │ - verify_structure │ -//! │ - to_proof_blob │ -//! └────────────────────┘ -//! ``` - -use serde::{Deserialize, Serialize}; -use sha2::{Digest, Sha256}; - -use super::{ProofBlob, ProofType, SemanticError}; - -/// Supported prover backends from the `proven` library (Idris2 ZKP system). -/// -/// NOTE: This `ProverKind` is NOT a mirror of ECHIDNA's dispatcher backend -/// enum. It lives in the `proven`-certificate domain and records which prover -/// produced a given `ProvenCertificate`. Keep it decoupled from echidna — -/// audited 2026-04-17 against echidna commit `8f573f1` (which expanded its -/// own ProverKind from 30 → ~68 variants). This enum is deliberately minimal. -#[derive(Debug, Clone, PartialEq, Eq, Serialize, Deserialize)] -#[serde(rename_all = "lowercase")] -pub enum ProverKind { - /// Z3 SMT solver - Z3, - /// Lean 4 theorem prover - Lean, - /// Coq proof assistant - Coq, - /// Agda dependently-typed prover - Agda, - /// Idris2 native (totality checker) - Idris2, - /// Custom/external prover - Custom(String), -} - -impl std::fmt::Display for ProverKind { - fn fmt(&self, f: &mut std::fmt::Formatter<'_>) -> std::fmt::Result { - match self { - Self::Z3 => write!(f, "z3"), - Self::Lean => write!(f, "lean"), - Self::Coq => write!(f, "coq"), - Self::Agda => write!(f, "agda"), - Self::Idris2 => write!(f, "idris2"), - Self::Custom(name) => write!(f, "custom:{}", name), - } - } -} - -/// A proof certificate generated by the proven library. -/// -/// Represents a formally verified proof statement with its proof term -/// and verification metadata. Can be stored in VeriSimDB as a `ProofBlob` -/// in the semantic modality. -#[derive(Debug, Clone, Serialize, Deserialize)] -pub struct ProvenCertificate { - /// Certificate format version. - pub version: String, - /// Which prover produced this certificate. - pub prover: ProverKind, - /// The formal statement being proven. - pub statement: String, - /// The proof term (prover-specific representation). - #[serde(skip_serializing_if = "Option::is_none")] - pub proof_term: Option<String>, - /// Whether the proof was verified as valid. - pub valid: bool, - /// Optional verification message. - #[serde(skip_serializing_if = "Option::is_none")] - pub message: Option<String>, - /// ISO 8601 timestamp of proof generation. - pub timestamp: String, - /// Content-addressable signature: sha256 of (prover + statement + proof_term). - pub signature: String, -} - -impl ProvenCertificate { - /// Compute the expected signature for a certificate. - fn compute_signature(prover: &ProverKind, statement: &str, proof_term: Option<&str>) -> String { - let mut hasher = Sha256::new(); - hasher.update(prover.to_string().as_bytes()); - hasher.update(b":"); - hasher.update(statement.as_bytes()); - hasher.update(b":"); - hasher.update(proof_term.unwrap_or("").as_bytes()); - format!("sha256:{}", hex::encode(hasher.finalize())) - } -} - -/// Parse a proven certificate from JSON bytes. -pub fn parse_proven_certificate(bytes: &[u8]) -> Result<ProvenCertificate, SemanticError> { - serde_json::from_slice(bytes) - .map_err(|e| SemanticError::SerializationError(format!("Failed to parse proven certificate: {}", e))) -} - -/// Parse a proven certificate from CBOR bytes. -pub fn parse_proven_certificate_cbor(bytes: &[u8]) -> Result<ProvenCertificate, SemanticError> { - ciborium::from_reader(bytes) - .map_err(|e| SemanticError::SerializationError(format!("Failed to parse proven certificate (CBOR): {}", e))) -} - -/// Verify the structural integrity of a proven certificate. -/// -/// Checks: -/// - Version is supported ("1.0") -/// - Signature matches the computed hash of (prover, statement, proof_term) -/// - Timestamp is a valid RFC 3339 datetime -pub fn verify_proven_certificate(cert: &ProvenCertificate) -> Result<bool, SemanticError> { - // Check version - if cert.version != "1.0" { - return Err(SemanticError::InvalidProof(format!( - "Unsupported certificate version: {}", - cert.version - ))); - } - - // Verify signature - let expected_sig = ProvenCertificate::compute_signature( - &cert.prover, - &cert.statement, - cert.proof_term.as_deref(), - ); - if cert.signature != expected_sig { - return Ok(false); - } - - // Verify timestamp is parseable - chrono::DateTime::parse_from_rfc3339(&cert.timestamp) - .map_err(|e| SemanticError::InvalidProof(format!("Invalid timestamp: {}", e)))?; - - Ok(cert.valid) -} - -/// Convert a proven certificate into a VeriSimDB ProofBlob for storage -/// in the semantic modality. -pub fn certificate_to_proof_blob(cert: &ProvenCertificate) -> Result<ProofBlob, SemanticError> { - let data = serde_json::to_vec(cert) - .map_err(|e| SemanticError::SerializationError(e.to_string()))?; - - Ok(ProofBlob { - claim: cert.statement.clone(), - proof_type: ProofType::Attestation, - data, - timestamp: cert.timestamp.clone(), - }) -} - -/// Create a new ProvenCertificate with computed signature. -/// -/// Used when generating certificates from VeriSimDB-side proof operations. -pub fn create_certificate( - prover: ProverKind, - statement: String, - proof_term: Option<String>, - valid: bool, - message: Option<String>, -) -> ProvenCertificate { - let signature = ProvenCertificate::compute_signature( - &prover, - &statement, - proof_term.as_deref(), - ); - - ProvenCertificate { - version: "1.0".to_string(), - prover, - statement, - proof_term, - valid, - message, - timestamp: chrono::Utc::now().to_rfc3339(), - signature, - } -} - -#[cfg(test)] -mod tests { - use super::*; - - fn sample_certificate() -> ProvenCertificate { - create_certificate( - ProverKind::Z3, - "forall x : Nat, x + 0 = x".to_string(), - Some("refl".to_string()), - true, - None, - ) - } - - #[test] - fn test_certificate_roundtrip_json() { - let cert = sample_certificate(); - let json = serde_json::to_vec(&cert).expect("TODO: handle error"); - let parsed = parse_proven_certificate(&json).expect("TODO: handle error"); - assert_eq!(parsed.statement, cert.statement); - assert_eq!(parsed.prover, ProverKind::Z3); - } - - #[test] - fn test_certificate_roundtrip_cbor() { - let cert = sample_certificate(); - let mut cbor = Vec::new(); - ciborium::into_writer(&cert, &mut cbor).expect("TODO: handle error"); - let parsed = parse_proven_certificate_cbor(&cbor).expect("TODO: handle error"); - assert_eq!(parsed.statement, cert.statement); - } - - #[test] - fn test_verify_valid_certificate() { - let cert = sample_certificate(); - assert!(verify_proven_certificate(&cert).expect("TODO: handle error")); - } - - #[test] - fn test_verify_invalid_signature() { - let mut cert = sample_certificate(); - cert.signature = "sha256:0000000000000000000000000000000000000000000000000000000000000000".to_string(); - assert!(!verify_proven_certificate(&cert).expect("TODO: handle error")); - } - - #[test] - fn test_verify_invalid_version() { - let mut cert = sample_certificate(); - cert.version = "2.0".to_string(); - assert!(verify_proven_certificate(&cert).is_err()); - } - - #[test] - fn test_certificate_to_proof_blob() { - let cert = sample_certificate(); - let blob = certificate_to_proof_blob(&cert).expect("TODO: handle error"); - assert_eq!(blob.claim, cert.statement); - assert!(!blob.data.is_empty()); - } - - #[test] - fn test_invalid_certificate_not_valid() { - let cert = create_certificate( - ProverKind::Lean, - "false".to_string(), - None, - false, - Some("Proof failed".to_string()), - ); - assert!(!verify_proven_certificate(&cert).expect("TODO: handle error")); - } - - #[test] - fn test_prover_kind_display() { - assert_eq!(ProverKind::Z3.to_string(), "z3"); - assert_eq!(ProverKind::Lean.to_string(), "lean"); - assert_eq!(ProverKind::Custom("myprover".to_string()).to_string(), "custom:myprover"); - } -} diff --git a/verisimdb/rust-core/verisim-semantic/src/sanctify_bridge.rs b/verisimdb/rust-core/verisim-semantic/src/sanctify_bridge.rs deleted file mode 100644 index 534068e1..00000000 --- a/verisimdb/rust-core/verisim-semantic/src/sanctify_bridge.rs +++ /dev/null @@ -1,413 +0,0 @@ -// SPDX-License-Identifier: MPL-2.0 -// Copyright (c) 2026 Jonathan D.A. Jewell (hyperpolymath) <j.d.a.jewell@open.ac.uk> -//! Sanctify Bridge — Integration with sanctify-php (Haskell security analyser). -//! -//! sanctify-php is a Haskell tool that analyses PHP/WordPress code for -//! security vulnerabilities (OWASP Top 10, WordPress-specific checks). -//! It produces structured reports in JSON/SARIF format. -//! -//! This bridge consumes sanctify reports, converts security issues into -//! VeriSimDB semantic annotations, and binds security contracts to octads. -//! -//! # Architecture -//! -//! ```text -//! sanctify-php (Haskell) → JSON report → SanctifyBridge → SemanticStore -//! │ -//! ┌─────────┴──────────┐ -//! │ SanctifyContract │ -//! │ - parse_report │ -//! │ - validate │ -//! │ - bind_to_octad │ -//! └────────────────────┘ -//! ``` - -use serde::{Deserialize, Serialize}; - -use super::{ProofBlob, ProofType, SemanticError}; - -/// Security issue severity levels (from sanctify-php). -#[derive(Debug, Clone, PartialEq, Eq, PartialOrd, Ord, Serialize, Deserialize)] -#[serde(rename_all = "lowercase")] -pub enum Severity { - Info, - Low, - Medium, - High, - Critical, -} - -impl std::fmt::Display for Severity { - fn fmt(&self, f: &mut std::fmt::Formatter<'_>) -> std::fmt::Result { - match self { - Self::Info => write!(f, "info"), - Self::Low => write!(f, "low"), - Self::Medium => write!(f, "medium"), - Self::High => write!(f, "high"), - Self::Critical => write!(f, "critical"), - } - } -} - -/// Types of security issues detected by sanctify. -#[derive(Debug, Clone, PartialEq, Eq, Serialize, Deserialize)] -pub enum IssueType { - SqlInjection, - CrossSiteScripting, - CrossSiteRequestForgery, - CommandInjection, - PathTraversal, - UnsafeDeserialization, - WeakCryptography, - HardcodedSecret, - DangerousFunction, - InsecureFileUpload, - OpenRedirect, - XPathInjection, - LdapInjection, - XXeVulnerability, - InsecureRandom, - MissingStrictTypes, - TypeCoercionRisk, -} - -impl std::fmt::Display for IssueType { - fn fmt(&self, f: &mut std::fmt::Formatter<'_>) -> std::fmt::Result { - write!(f, "{:?}", self) - } -} - -/// A security issue from a sanctify report. -#[derive(Debug, Clone, Serialize, Deserialize)] -pub struct SecurityIssue { - /// Type of vulnerability - pub issue_type: IssueType, - /// Severity level - pub severity: Severity, - /// File path - pub file: String, - /// Line number - pub line: u32, - /// Column number - pub column: Option<u32>, - /// Human-readable description - pub description: String, - /// Recommended remediation - pub remedy: String, - /// Optional code snippet - #[serde(skip_serializing_if = "Option::is_none")] - pub code: Option<String>, -} - -/// A sanctify security report. -#[derive(Debug, Clone, Serialize, Deserialize)] -pub struct SanctifyReport { - /// Report timestamp - pub timestamp: String, - /// Sanctify version - pub version: String, - /// Files analysed - pub files_analysed: usize, - /// Security issues found - pub issues: Vec<SecurityIssue>, - /// Summary counts by severity - pub summary: IssueSummary, -} - -/// Summary counts of issues by severity. -#[derive(Debug, Clone, Default, Serialize, Deserialize)] -pub struct IssueSummary { - pub critical: usize, - pub high: usize, - pub medium: usize, - pub low: usize, - pub info: usize, -} - -impl IssueSummary { - /// Total number of issues. - pub fn total(&self) -> usize { - self.critical + self.high + self.medium + self.low + self.info - } - - /// Compute summary from a list of issues. - pub fn from_issues(issues: &[SecurityIssue]) -> Self { - let mut summary = Self::default(); - for issue in issues { - match issue.severity { - Severity::Critical => summary.critical += 1, - Severity::High => summary.high += 1, - Severity::Medium => summary.medium += 1, - Severity::Low => summary.low += 1, - Severity::Info => summary.info += 1, - } - } - summary - } -} - -/// A verifiable security contract binding sanctify findings to a octad. -/// -/// Represents a commitment that a particular codebase or entity has been -/// analysed and the results are stored in the semantic modality. -#[derive(Debug, Clone, Serialize, Deserialize)] -pub struct SanctifyContract { - /// Unique contract identifier - pub contract_id: String, - /// Octad ID this contract is bound to - pub octad_id: String, - /// The sanctify report backing this contract - pub report: SanctifyReport, - /// Whether all critical issues have been resolved - pub all_critical_resolved: bool, - /// Whether all high issues have been resolved - pub all_high_resolved: bool, - /// Contract creation timestamp - pub created_at: String, -} - -/// Parse a sanctify report from JSON bytes. -pub fn parse_sanctify_report(bytes: &[u8]) -> Result<SanctifyReport, SemanticError> { - serde_json::from_slice(bytes) - .map_err(|e| SemanticError::SerializationError(format!("Failed to parse sanctify report: {}", e))) -} - -/// Validate a sanctify contract's structural integrity. -/// -/// Checks: -/// - Contract has a non-empty ID and octad_id -/// - Summary counts match the actual issue counts -/// - Resolution flags are consistent with issue counts -pub fn validate_contract(contract: &SanctifyContract) -> Result<bool, SemanticError> { - if contract.contract_id.is_empty() { - return Err(SemanticError::ConstraintViolation("Contract ID must not be empty".to_string())); - } - if contract.octad_id.is_empty() { - return Err(SemanticError::ConstraintViolation("Octad ID must not be empty".to_string())); - } - - // Verify summary is consistent - let computed = IssueSummary::from_issues(&contract.report.issues); - if computed.total() != contract.report.summary.total() { - return Err(SemanticError::ConstraintViolation(format!( - "Summary mismatch: computed {} issues but summary says {}", - computed.total(), - contract.report.summary.total() - ))); - } - - // Verify resolution flags are consistent - if contract.all_critical_resolved && contract.report.summary.critical > 0 { - return Ok(false); // Claims resolved but has critical issues - } - if contract.all_high_resolved && contract.report.summary.high > 0 { - return Ok(false); // Claims resolved but has high issues - } - - Ok(true) -} - -/// Convert a sanctify contract into a VeriSimDB ProofBlob for storage. -pub fn contract_to_proof_blob(contract: &SanctifyContract) -> Result<ProofBlob, SemanticError> { - let data = serde_json::to_vec(contract) - .map_err(|e| SemanticError::SerializationError(e.to_string()))?; - - let claim = format!( - "security-audit:{} octad:{} issues:{}", - contract.contract_id, - contract.octad_id, - contract.report.summary.total() - ); - - Ok(ProofBlob { - claim, - proof_type: ProofType::Attestation, - data, - timestamp: contract.created_at.clone(), - }) -} - -/// Create a SanctifyContract from a report and octad binding. -pub fn bind_contract_to_octad( - contract_id: String, - octad_id: String, - report: SanctifyReport, -) -> SanctifyContract { - let all_critical_resolved = report.summary.critical == 0; - let all_high_resolved = report.summary.high == 0; - - SanctifyContract { - contract_id, - octad_id, - report, - all_critical_resolved, - all_high_resolved, - created_at: chrono::Utc::now().to_rfc3339(), - } -} - -#[cfg(test)] -mod tests { - use super::*; - - fn sample_report() -> SanctifyReport { - let issues = vec![ - SecurityIssue { - issue_type: IssueType::SqlInjection, - severity: Severity::Critical, - file: "login.php".to_string(), - line: 42, - column: Some(15), - description: "Unsanitized user input in SQL query".to_string(), - remedy: "Use prepared statements with PDO".to_string(), - code: Some("$query = \"SELECT * FROM users WHERE id = \" . $_GET['id']".to_string()), - }, - SecurityIssue { - issue_type: IssueType::CrossSiteScripting, - severity: Severity::High, - file: "profile.php".to_string(), - line: 88, - column: None, - description: "Reflected XSS via unescaped output".to_string(), - remedy: "Use htmlspecialchars() or esc_html()".to_string(), - code: None, - }, - SecurityIssue { - issue_type: IssueType::MissingStrictTypes, - severity: Severity::Info, - file: "utils.php".to_string(), - line: 1, - column: None, - description: "Missing declare(strict_types=1)".to_string(), - remedy: "Add strict types declaration".to_string(), - code: None, - }, - ]; - let summary = IssueSummary::from_issues(&issues); - SanctifyReport { - timestamp: chrono::Utc::now().to_rfc3339(), - version: "0.1.0".to_string(), - files_analysed: 3, - issues, - summary, - } - } - - #[test] - fn test_parse_report_json() { - let report = sample_report(); - let json = serde_json::to_vec(&report).expect("TODO: handle error"); - let parsed = parse_sanctify_report(&json).expect("TODO: handle error"); - assert_eq!(parsed.issues.len(), 3); - assert_eq!(parsed.summary.critical, 1); - assert_eq!(parsed.summary.high, 1); - assert_eq!(parsed.summary.info, 1); - } - - #[test] - fn test_issue_summary() { - let report = sample_report(); - assert_eq!(report.summary.total(), 3); - assert_eq!(report.summary.critical, 1); - assert_eq!(report.summary.high, 1); - assert_eq!(report.summary.info, 1); - assert_eq!(report.summary.medium, 0); - assert_eq!(report.summary.low, 0); - } - - #[test] - fn test_bind_contract() { - let report = sample_report(); - let contract = bind_contract_to_octad( - "audit-001".to_string(), - "octad-abc".to_string(), - report, - ); - assert!(!contract.all_critical_resolved); // Has 1 critical - assert!(!contract.all_high_resolved); // Has 1 high - assert_eq!(contract.report.issues.len(), 3); - } - - #[test] - fn test_validate_contract() { - let report = sample_report(); - let contract = bind_contract_to_octad( - "audit-001".to_string(), - "octad-abc".to_string(), - report, - ); - // Valid contract (resolution flags match issue counts) - assert!(validate_contract(&contract).expect("TODO: handle error")); - } - - #[test] - fn test_validate_contract_empty_id() { - let report = sample_report(); - let contract = SanctifyContract { - contract_id: "".to_string(), - octad_id: "octad-abc".to_string(), - report, - all_critical_resolved: false, - all_high_resolved: false, - created_at: chrono::Utc::now().to_rfc3339(), - }; - assert!(validate_contract(&contract).is_err()); - } - - #[test] - fn test_validate_inconsistent_resolution() { - let report = sample_report(); - let contract = SanctifyContract { - contract_id: "audit-002".to_string(), - octad_id: "octad-abc".to_string(), - report, - all_critical_resolved: true, // Claims resolved but has 1 critical - all_high_resolved: false, - created_at: chrono::Utc::now().to_rfc3339(), - }; - // Should return false (inconsistent) - assert!(!validate_contract(&contract).expect("TODO: handle error")); - } - - #[test] - fn test_contract_to_proof_blob() { - let report = sample_report(); - let contract = bind_contract_to_octad( - "audit-003".to_string(), - "octad-xyz".to_string(), - report, - ); - let blob = contract_to_proof_blob(&contract).expect("TODO: handle error"); - assert!(blob.claim.contains("security-audit:audit-003")); - assert!(blob.claim.contains("octad:octad-xyz")); - assert!(!blob.data.is_empty()); - } - - #[test] - fn test_severity_ordering() { - assert!(Severity::Critical > Severity::High); - assert!(Severity::High > Severity::Medium); - assert!(Severity::Medium > Severity::Low); - assert!(Severity::Low > Severity::Info); - } - - #[test] - fn test_clean_report_contract() { - // A clean report (no issues) - let report = SanctifyReport { - timestamp: chrono::Utc::now().to_rfc3339(), - version: "0.1.0".to_string(), - files_analysed: 5, - issues: vec![], - summary: IssueSummary::default(), - }; - let contract = bind_contract_to_octad( - "clean-001".to_string(), - "octad-clean".to_string(), - report, - ); - assert!(contract.all_critical_resolved); - assert!(contract.all_high_resolved); - assert!(validate_contract(&contract).expect("TODO: handle error")); - } -} diff --git a/verisimdb/rust-core/verisim-semantic/src/verification_keys.rs b/verisimdb/rust-core/verisim-semantic/src/verification_keys.rs deleted file mode 100644 index 78488f11..00000000 --- a/verisimdb/rust-core/verisim-semantic/src/verification_keys.rs +++ /dev/null @@ -1,246 +0,0 @@ -// SPDX-License-Identifier: MPL-2.0 -//! Verification Key Management for VeriSimDB Custom Circuits -//! -//! Stores verification keys per circuit with support for key rotation -//! and federation key export/import (peers need matching keys to verify -//! proofs from other instances). - -use serde::{Deserialize, Serialize}; -use sha2::{Digest, Sha256}; -use std::collections::HashMap; -use std::sync::RwLock; - -use super::circuit_registry::CircuitError; - -/// A verification key entry with rotation support -#[derive(Debug, Clone, Serialize, Deserialize)] -pub struct VerificationKeyEntry { - /// Circuit name this key belongs to - pub circuit_name: String, - /// Current active key - pub active_key: Vec<u8>, - /// Key version (monotonically increasing) - pub version: u64, - /// SHA-256 fingerprint of the key (for quick comparison) - pub fingerprint: String, - /// Previous key (for graceful rotation — accept proofs from both during transition) - pub previous_key: Option<Vec<u8>>, - /// When this key version was created (ISO 8601) - pub created_at: String, -} - -impl VerificationKeyEntry { - /// Create a new key entry - pub fn new(circuit_name: &str, key: Vec<u8>) -> Self { - let fingerprint = key_fingerprint(&key); - Self { - circuit_name: circuit_name.to_string(), - active_key: key, - version: 1, - fingerprint, - previous_key: None, - created_at: chrono::Utc::now().to_rfc3339(), - } - } - - /// Rotate to a new key, keeping the old one as previous - pub fn rotate(&mut self, new_key: Vec<u8>) { - self.previous_key = Some(self.active_key.clone()); - self.active_key = new_key; - self.fingerprint = key_fingerprint(&self.active_key); - self.version += 1; - self.created_at = chrono::Utc::now().to_rfc3339(); - } - - /// Check if a given key matches the active or previous key - pub fn matches(&self, key: &[u8]) -> bool { - self.active_key == key - || self - .previous_key - .as_ref() - .is_some_and(|prev| prev == key) - } -} - -/// Compute SHA-256 fingerprint of a key -fn key_fingerprint(key: &[u8]) -> String { - let mut hasher = Sha256::new(); - hasher.update(key); - let result = hasher.finalize(); - result.iter().map(|b| format!("{:02x}", b)).collect() -} - -/// Exportable key bundle for federation -#[derive(Debug, Clone, Serialize, Deserialize)] -pub struct KeyExportBundle { - /// Source instance identifier - pub source_instance: String, - /// Map of circuit_name → (key_bytes, version, fingerprint) - pub keys: Vec<ExportedKey>, -} - -/// A single exported key -#[derive(Debug, Clone, Serialize, Deserialize)] -pub struct ExportedKey { - pub circuit_name: String, - pub key: Vec<u8>, - pub version: u64, - pub fingerprint: String, -} - -/// The verification key store -pub struct VerificationKeyStore { - keys: RwLock<HashMap<String, VerificationKeyEntry>>, - instance_id: String, -} - -impl VerificationKeyStore { - /// Create a new key store - pub fn new(instance_id: &str) -> Self { - Self { - keys: RwLock::new(HashMap::new()), - instance_id: instance_id.to_string(), - } - } - - /// Store a verification key for a circuit - pub fn store_key( - &self, - circuit_name: &str, - key: Vec<u8>, - ) -> Result<(), CircuitError> { - let mut keys = self.keys.write().map_err(|_| CircuitError::LockPoisoned)?; - - if let Some(entry) = keys.get_mut(circuit_name) { - entry.rotate(key); - } else { - keys.insert( - circuit_name.to_string(), - VerificationKeyEntry::new(circuit_name, key), - ); - } - - Ok(()) - } - - /// Get the active key for a circuit - pub fn get_key(&self, circuit_name: &str) -> Result<Option<Vec<u8>>, CircuitError> { - let keys = self.keys.read().map_err(|_| CircuitError::LockPoisoned)?; - Ok(keys.get(circuit_name).map(|e| e.active_key.clone())) - } - - /// Get full key entry (including version and previous key) - pub fn get_entry( - &self, - circuit_name: &str, - ) -> Result<Option<VerificationKeyEntry>, CircuitError> { - let keys = self.keys.read().map_err(|_| CircuitError::LockPoisoned)?; - Ok(keys.get(circuit_name).cloned()) - } - - /// Export all keys for federation sharing - pub fn export_keys(&self) -> Result<KeyExportBundle, CircuitError> { - let keys = self.keys.read().map_err(|_| CircuitError::LockPoisoned)?; - - let exported = keys - .values() - .map(|entry| ExportedKey { - circuit_name: entry.circuit_name.clone(), - key: entry.active_key.clone(), - version: entry.version, - fingerprint: entry.fingerprint.clone(), - }) - .collect(); - - Ok(KeyExportBundle { - source_instance: self.instance_id.clone(), - keys: exported, - }) - } - - /// Import keys from a federation peer - pub fn import_keys(&self, bundle: &KeyExportBundle) -> Result<usize, CircuitError> { - let mut keys = self.keys.write().map_err(|_| CircuitError::LockPoisoned)?; - let mut imported = 0; - - for exported in &bundle.keys { - let federated_name = format!("{}:{}", bundle.source_instance, exported.circuit_name); - - match keys.get_mut(&federated_name) { - Some(entry) if entry.version < exported.version => { - entry.rotate(exported.key.clone()); - imported += 1; - } - None => { - keys.insert( - federated_name, - VerificationKeyEntry::new(&exported.circuit_name, exported.key.clone()), - ); - imported += 1; - } - _ => { - // Already have same or newer version — skip - } - } - } - - Ok(imported) - } - - /// List all stored circuit names - pub fn list_circuits(&self) -> Result<Vec<String>, CircuitError> { - let keys = self.keys.read().map_err(|_| CircuitError::LockPoisoned)?; - Ok(keys.keys().cloned().collect()) - } -} - -#[cfg(test)] -mod tests { - use super::*; - - #[test] - fn test_store_and_retrieve() { - let store = VerificationKeyStore::new("test-instance"); - let key = vec![1, 2, 3, 4]; - - store.store_key("my-circuit", key.clone()).expect("TODO: handle error"); - - let retrieved = store.get_key("my-circuit").expect("TODO: handle error").expect("TODO: handle error"); - assert_eq!(retrieved, key); - } - - #[test] - fn test_key_rotation() { - let store = VerificationKeyStore::new("test"); - let key1 = vec![1, 2, 3]; - let key2 = vec![4, 5, 6]; - - store.store_key("circuit", key1.clone()).expect("TODO: handle error"); - store.store_key("circuit", key2.clone()).expect("TODO: handle error"); - - let entry = store.get_entry("circuit").expect("TODO: handle error").expect("TODO: handle error"); - assert_eq!(entry.active_key, key2); - assert_eq!(entry.previous_key, Some(key1.clone())); - assert_eq!(entry.version, 2); - assert!(entry.matches(&key1)); - assert!(entry.matches(&key2)); - } - - #[test] - fn test_export_import() { - let store_a = VerificationKeyStore::new("instance-a"); - store_a.store_key("circuit-1", vec![10, 20]).expect("TODO: handle error"); - store_a.store_key("circuit-2", vec![30, 40]).expect("TODO: handle error"); - - let bundle = store_a.export_keys().expect("TODO: handle error"); - assert_eq!(bundle.keys.len(), 2); - - let store_b = VerificationKeyStore::new("instance-b"); - let imported = store_b.import_keys(&bundle).expect("TODO: handle error"); - assert_eq!(imported, 2); - - // Keys are stored with federated prefix - let circuits = store_b.list_circuits().expect("TODO: handle error"); - assert!(circuits.iter().any(|c| c.starts_with("instance-a:"))); - } -} diff --git a/verisimdb/rust-core/verisim-semantic/src/zkp.rs b/verisimdb/rust-core/verisim-semantic/src/zkp.rs deleted file mode 100644 index 6ff90eae..00000000 --- a/verisimdb/rust-core/verisim-semantic/src/zkp.rs +++ /dev/null @@ -1,367 +0,0 @@ -// SPDX-License-Identifier: MPL-2.0 -//! Zero-Knowledge Proof primitives for the semantic store. -//! -//! Provides cryptographic proof mechanisms that allow verification of -//! semantic claims without revealing underlying data: -//! -//! - **Hash Commitments**: Commit to a value, reveal later to prove knowledge. -//! - **Merkle Proofs**: Prove set membership without revealing other members. -//! - **Proof Verification**: Verify stored proofs against their claims. - -use serde::{Deserialize, Serialize}; -use sha2::{Digest, Sha256}; - -// --------------------------------------------------------------------------- -// Hash Commitment Scheme -// --------------------------------------------------------------------------- - -/// A hash commitment: SHA-256(claim || secret). -/// The committer can later reveal the secret to prove they knew the value -/// at commitment time, without having revealed it earlier. -#[derive(Debug, Clone, Serialize, Deserialize)] -pub struct HashCommitment { - /// The commitment hash. - pub commitment: [u8; 32], -} - -/// Create a hash commitment for a claim using a secret. -pub fn commit(claim: &[u8], secret: &[u8]) -> HashCommitment { - let mut hasher = Sha256::new(); - hasher.update(claim); - hasher.update(secret); - let result = hasher.finalize(); - HashCommitment { - commitment: result.into(), - } -} - -/// Verify a hash commitment by checking SHA-256(claim || secret) == commitment. -pub fn verify_commitment(commitment: &HashCommitment, claim: &[u8], secret: &[u8]) -> bool { - let expected = commit(claim, secret); - constant_time_eq(&commitment.commitment, &expected.commitment) -} - -// --------------------------------------------------------------------------- -// Merkle Tree -// --------------------------------------------------------------------------- - -/// An element in a Merkle proof path. -#[derive(Debug, Clone, Serialize, Deserialize)] -pub struct MerklePathElement { - /// Sibling hash at this level. - pub hash: [u8; 32], - /// Whether the sibling is on the left (true) or right (false). - pub is_left: bool, -} - -/// A complete Merkle inclusion proof. -#[derive(Debug, Clone, Serialize, Deserialize)] -pub struct MerkleProof { - /// The leaf value being proven. - pub leaf: Vec<u8>, - /// Path from leaf to root. - pub path: Vec<MerklePathElement>, - /// The Merkle root hash. - pub root: [u8; 32], -} - -/// Compute SHA-256 hash of data. -pub fn hash(data: &[u8]) -> [u8; 32] { - Sha256::digest(data).into() -} - -/// Hash two children to form a parent node. -fn hash_pair(left: &[u8; 32], right: &[u8; 32]) -> [u8; 32] { - let mut hasher = Sha256::new(); - hasher.update(left); - hasher.update(right); - hasher.finalize().into() -} - -/// Build a Merkle tree from leaf data and return the root hash. -/// Leaves are hashed before building the tree. -pub fn merkle_root(leaves: &[Vec<u8>]) -> [u8; 32] { - if leaves.is_empty() { - return [0u8; 32]; - } - - let mut current_level: Vec<[u8; 32]> = leaves.iter().map(|l| hash(l)).collect(); - - // Pad to even number if necessary - while current_level.len() > 1 { - if current_level.len() % 2 != 0 { - let last = *current_level.last().expect("TODO: handle error"); - current_level.push(last); - } - - let mut next_level = Vec::with_capacity(current_level.len() / 2); - for chunk in current_level.chunks(2) { - next_level.push(hash_pair(&chunk[0], &chunk[1])); - } - current_level = next_level; - } - - current_level[0] -} - -/// Generate a Merkle inclusion proof for the leaf at `index`. -pub fn merkle_proof(leaves: &[Vec<u8>], index: usize) -> Option<MerkleProof> { - if index >= leaves.len() || leaves.is_empty() { - return None; - } - - let root = merkle_root(leaves); - let mut hashed: Vec<[u8; 32]> = leaves.iter().map(|l| hash(l)).collect(); - let mut path = Vec::new(); - let mut idx = index; - - while hashed.len() > 1 { - // Pad to even - if hashed.len() % 2 != 0 { - let last = *hashed.last().expect("TODO: handle error"); - hashed.push(last); - } - - // Find sibling - let sibling_idx = if idx % 2 == 0 { idx + 1 } else { idx - 1 }; - let is_left = idx % 2 != 0; // sibling is on left if we're on the right - - path.push(MerklePathElement { - hash: hashed[sibling_idx], - is_left, - }); - - // Move up one level - let mut next_level = Vec::with_capacity(hashed.len() / 2); - for chunk in hashed.chunks(2) { - next_level.push(hash_pair(&chunk[0], &chunk[1])); - } - hashed = next_level; - idx /= 2; - } - - Some(MerkleProof { - leaf: leaves[index].clone(), - path, - root, - }) -} - -/// Verify a Merkle inclusion proof. -pub fn verify_merkle_proof(proof: &MerkleProof) -> bool { - let mut current = hash(&proof.leaf); - - for element in &proof.path { - current = if element.is_left { - hash_pair(&element.hash, ¤t) - } else { - hash_pair(¤t, &element.hash) - }; - } - - constant_time_eq(¤t, &proof.root) -} - -// --------------------------------------------------------------------------- -// Verifiable Proof Types (integrated with ProofBlob) -// --------------------------------------------------------------------------- - -/// A verifiable proof with cryptographic backing. -#[derive(Debug, Clone, Serialize, Deserialize)] -pub enum VerifiableProofData { - /// Hash commitment: prover committed to a value. - Commitment { - commitment: [u8; 32], - }, - /// Hash reveal: prover reveals the secret for a prior commitment. - Reveal { - commitment: [u8; 32], - secret: Vec<u8>, - }, - /// Merkle inclusion: value is a member of a committed set. - MerkleInclusion(MerkleProof), - /// Content integrity: SHA-256 hash of the original content. - ContentIntegrity { - content_hash: [u8; 32], - }, -} - -/// Verify a verifiable proof against its claim. -pub fn verify_proof(data: &VerifiableProofData, claim: &[u8]) -> bool { - match data { - VerifiableProofData::Commitment { .. } => { - // Commitments are valid by construction — they're verified at reveal time. - true - } - VerifiableProofData::Reveal { - commitment, - secret, - } => { - let expected = commit(claim, secret); - constant_time_eq(commitment, &expected.commitment) - } - VerifiableProofData::MerkleInclusion(proof) => verify_merkle_proof(proof), - VerifiableProofData::ContentIntegrity { content_hash } => { - let actual = hash(claim); - constant_time_eq(content_hash, &actual) - } - } -} - -// --------------------------------------------------------------------------- -// Utility -// --------------------------------------------------------------------------- - -/// Constant-time byte comparison to prevent timing side-channels. -fn constant_time_eq(a: &[u8], b: &[u8]) -> bool { - if a.len() != b.len() { - return false; - } - let mut diff = 0u8; - for (x, y) in a.iter().zip(b.iter()) { - diff |= x ^ y; - } - diff == 0 -} - -// --------------------------------------------------------------------------- -// Tests -// --------------------------------------------------------------------------- - -#[cfg(test)] -mod tests { - use super::*; - - #[test] - fn test_hash_commitment_roundtrip() { - let claim = b"entity:123 is-a Person"; - let secret = b"my-secret-nonce-42"; - - let commitment = commit(claim, secret); - assert!(verify_commitment(&commitment, claim, secret)); - } - - #[test] - fn test_hash_commitment_wrong_secret() { - let claim = b"entity:123 is-a Person"; - let commitment = commit(claim, b"correct-secret"); - - assert!(!verify_commitment(&commitment, claim, b"wrong-secret")); - } - - #[test] - fn test_hash_commitment_wrong_claim() { - let secret = b"my-secret"; - let commitment = commit(b"real claim", secret); - - assert!(!verify_commitment(&commitment, b"fake claim", secret)); - } - - #[test] - fn test_merkle_root_single() { - let leaves = vec![b"leaf0".to_vec()]; - let root = merkle_root(&leaves); - assert_eq!(root, hash(b"leaf0")); - } - - #[test] - fn test_merkle_root_deterministic() { - let leaves = vec![b"a".to_vec(), b"b".to_vec(), b"c".to_vec()]; - let root1 = merkle_root(&leaves); - let root2 = merkle_root(&leaves); - assert_eq!(root1, root2); - } - - #[test] - fn test_merkle_proof_verify() { - let leaves = vec![ - b"alpha".to_vec(), - b"beta".to_vec(), - b"gamma".to_vec(), - b"delta".to_vec(), - ]; - - // Prove each leaf - for i in 0..leaves.len() { - let proof = merkle_proof(&leaves, i).expect("TODO: handle error"); - assert!( - verify_merkle_proof(&proof), - "Merkle proof failed for leaf {i}" - ); - } - } - - #[test] - fn test_merkle_proof_odd_leaves() { - let leaves = vec![b"a".to_vec(), b"b".to_vec(), b"c".to_vec()]; - - for i in 0..leaves.len() { - let proof = merkle_proof(&leaves, i).expect("TODO: handle error"); - assert!(verify_merkle_proof(&proof)); - } - } - - #[test] - fn test_merkle_proof_tampered() { - let leaves = vec![b"a".to_vec(), b"b".to_vec(), b"c".to_vec(), b"d".to_vec()]; - let mut proof = merkle_proof(&leaves, 0).expect("TODO: handle error"); - - // Tamper with the leaf - proof.leaf = b"tampered".to_vec(); - assert!(!verify_merkle_proof(&proof)); - } - - #[test] - fn test_verify_proof_content_integrity() { - let content = b"This is the original document content"; - let proof_data = VerifiableProofData::ContentIntegrity { - content_hash: hash(content), - }; - - assert!(verify_proof(&proof_data, content)); - assert!(!verify_proof(&proof_data, b"modified content")); - } - - #[test] - fn test_verify_proof_reveal() { - let claim = b"entity:456 satisfies constraint X"; - let secret = b"witness-data"; - - let commitment = commit(claim, secret); - let proof_data = VerifiableProofData::Reveal { - commitment: commitment.commitment, - secret: secret.to_vec(), - }; - - assert!(verify_proof(&proof_data, claim)); - assert!(!verify_proof(&proof_data, b"wrong claim")); - } - - #[test] - fn test_verify_proof_merkle_inclusion() { - let leaves = vec![ - b"claim-1".to_vec(), - b"claim-2".to_vec(), - b"claim-3".to_vec(), - ]; - - let proof = merkle_proof(&leaves, 1).expect("TODO: handle error"); - let proof_data = VerifiableProofData::MerkleInclusion(proof); - - // Merkle proof verification doesn't use the claim parameter directly - // (the leaf is embedded in the proof), but the proof itself must be valid. - assert!(verify_proof(&proof_data, b"claim-2")); - } - - #[test] - fn test_empty_merkle_root() { - let root = merkle_root(&[]); - assert_eq!(root, [0u8; 32]); - } - - #[test] - fn test_merkle_proof_out_of_bounds() { - let leaves = vec![b"a".to_vec()]; - assert!(merkle_proof(&leaves, 5).is_none()); - } -} diff --git a/verisimdb/rust-core/verisim-semantic/src/zkp_bridge.rs b/verisimdb/rust-core/verisim-semantic/src/zkp_bridge.rs deleted file mode 100644 index 00f218fc..00000000 --- a/verisimdb/rust-core/verisim-semantic/src/zkp_bridge.rs +++ /dev/null @@ -1,628 +0,0 @@ -// SPDX-License-Identifier: MPL-2.0 -//! ZKP Bridge — Privacy-aware proof generation and verification. -//! -//! Wraps the existing ZKP primitives (hash commitments, Merkle proofs, -//! content integrity, circuit verification) with privacy-level routing: -//! -//! - **Public**: Standard proof — data and proof are both visible. -//! - **Private**: Hash commitment hides data; Merkle inclusion proves membership. -//! - **ZeroKnowledge**: Blinded Merkle proof with committed witnesses. -//! (Full ZK-SNARK via sanctify is designed but not yet compiled in.) -//! -//! # Architecture -//! -//! ```text -//! VCL PROOF clause → Elixir executor → Rust API → ZkpBridge -//! │ -//! ┌───────────────────────────────┘ -//! │ -//! ┌───────┴───────┐ -//! │ PrivacyLevel │ -//! └───┬───┬───┬───┘ -//! │ │ │ -//! Public Private ZeroKnowledge -//! │ │ │ -//! ┌───┘ │ └───┐ -//! ▼ ▼ ▼ -//! Standard Committed Blinded -//! proof + Merkle + Nonce -//! ``` - -use serde::{Deserialize, Serialize}; - -use super::circuit_registry::{CircuitError, CircuitRegistry}; -use super::zkp::{ - commit, hash, merkle_proof, merkle_root, verify_merkle_proof, - verify_proof, VerifiableProofData, -}; - -/// Privacy level for proof generation -#[derive(Debug, Clone, Copy, PartialEq, Eq, Serialize, Deserialize)] -pub enum PrivacyLevel { - /// Data and proof are both visible to the verifier. - Public, - /// Data is hidden behind a hash commitment; proof proves - /// the committer knows the value without revealing it. - Private, - /// Zero-knowledge: verifier learns nothing beyond the - /// statement's truth. Uses blinded Merkle proofs with - /// committed witnesses. (Full ZK-SNARK integration pending.) - ZeroKnowledge, -} - -impl Default for PrivacyLevel { - fn default() -> Self { - Self::Public - } -} - -impl std::fmt::Display for PrivacyLevel { - fn fmt(&self, f: &mut std::fmt::Formatter<'_>) -> std::fmt::Result { - match self { - Self::Public => write!(f, "Public"), - Self::Private => write!(f, "Private"), - Self::ZeroKnowledge => write!(f, "ZeroKnowledge"), - } - } -} - -/// A request to generate a privacy-aware proof -#[derive(Debug, Clone, Serialize, Deserialize)] -pub struct ZkpProofRequest { - /// The entity or claim being proven - pub claim: Vec<u8>, - /// Privacy level requested - pub privacy_level: PrivacyLevel, - /// Optional circuit name for CUSTOM proofs - pub circuit_name: Option<String>, - /// Optional witness data (private inputs for circuit proofs) - pub witness: Option<Vec<f64>>, - /// Optional public inputs (for circuit proofs) - pub public_inputs: Option<Vec<f64>>, - /// Optional set of sibling claims (for Merkle membership proofs) - pub membership_set: Option<Vec<Vec<u8>>>, - /// Index of the claim in the membership set - pub membership_index: Option<usize>, -} - -/// A generated proof with privacy metadata -#[derive(Debug, Clone, Serialize, Deserialize)] -pub struct ZkpProof { - /// Privacy level of this proof - pub privacy_level: PrivacyLevel, - /// The underlying verifiable proof data - pub proof_data: VerifiableProofData, - /// Optional blinding nonce (present for Private and ZeroKnowledge) - pub blinding_nonce: Option<Vec<u8>>, - /// Optional commitment (for Private/ZK proofs, the committed value) - pub commitment: Option<[u8; 32]>, - /// Merkle root of the membership set (if applicable) - pub merkle_root: Option<[u8; 32]>, - /// Circuit verification result (for CUSTOM proofs) - pub circuit_result: Option<CircuitVerificationResult>, - /// Timestamp of proof generation - pub generated_at: String, -} - -/// Result of circuit-based verification -#[derive(Debug, Clone, Serialize, Deserialize)] -pub struct CircuitVerificationResult { - /// Circuit name - pub circuit_name: String, - /// Whether the circuit constraints were satisfied - pub satisfied: bool, - /// Number of constraints checked - pub constraints_checked: usize, -} - -/// Generate a privacy-aware proof. -/// -/// Routes to the appropriate proof generation strategy based on the -/// requested privacy level. -pub fn generate_zkp(request: &ZkpProofRequest) -> Result<ZkpProof, CircuitError> { - match request.privacy_level { - PrivacyLevel::Public => generate_public_proof(request), - PrivacyLevel::Private => generate_private_proof(request), - PrivacyLevel::ZeroKnowledge => generate_zk_proof(request), - } -} - -/// Verify a previously generated ZKP proof. -/// -/// For Public proofs, verifies the underlying proof data directly. -/// For Private proofs, verifies the commitment and Merkle inclusion. -/// For ZeroKnowledge proofs, verifies the blinded proof without -/// requiring knowledge of the original data. -pub fn verify_zkp(proof: &ZkpProof, claim: &[u8]) -> bool { - match proof.privacy_level { - PrivacyLevel::Public => verify_proof(&proof.proof_data, claim), - PrivacyLevel::Private => verify_private_proof(proof, claim), - PrivacyLevel::ZeroKnowledge => verify_zk_proof(proof), - } -} - -/// Generate a privacy-aware proof with circuit verification. -/// -/// Combines the ZKP bridge with the circuit registry to verify -/// custom circuit constraints before generating the proof. -pub fn generate_zkp_with_circuit( - request: &ZkpProofRequest, - registry: &CircuitRegistry, -) -> Result<ZkpProof, CircuitError> { - let mut proof = generate_zkp(request)?; - - // If a circuit name is specified, verify against the registry - if let Some(ref circuit_name) = request.circuit_name { - let witness = request.witness.as_deref().unwrap_or(&[]); - let public_inputs = request.public_inputs.as_deref().unwrap_or(&[]); - - let satisfied = registry.verify_with_circuit(circuit_name, witness, public_inputs)?; - - let circuit = registry.get_circuit(circuit_name)?; - let constraints_checked = circuit - .map(|c| c.ir.constraints.len()) - .unwrap_or(0); - - proof.circuit_result = Some(CircuitVerificationResult { - circuit_name: circuit_name.clone(), - satisfied, - constraints_checked, - }); - } - - Ok(proof) -} - -// --------------------------------------------------------------------------- -// Public proof: standard proof with data visible -// --------------------------------------------------------------------------- - -fn generate_public_proof(request: &ZkpProofRequest) -> Result<ZkpProof, CircuitError> { - let proof_data = VerifiableProofData::ContentIntegrity { - content_hash: hash(&request.claim), - }; - - Ok(ZkpProof { - privacy_level: PrivacyLevel::Public, - proof_data, - blinding_nonce: None, - commitment: None, - merkle_root: None, - circuit_result: None, - generated_at: chrono::Utc::now().to_rfc3339(), - }) -} - -// --------------------------------------------------------------------------- -// Private proof: hash commitment hides the data -// --------------------------------------------------------------------------- - -fn generate_private_proof(request: &ZkpProofRequest) -> Result<ZkpProof, CircuitError> { - // Generate a random-ish nonce from claim hash (deterministic for testing; - // in production this would use a CSPRNG) - let nonce = generate_nonce(&request.claim); - - let commitment = commit(&request.claim, &nonce); - - // If a membership set is provided, generate a Merkle inclusion proof - let (proof_data, root) = if let (Some(ref set), Some(index)) = - (&request.membership_set, request.membership_index) - { - let root = merkle_root(set); - match merkle_proof(set, index) { - Some(mp) => (VerifiableProofData::MerkleInclusion(mp), Some(root)), - None => { - return Err(CircuitError::InvalidWitness(format!( - "Membership index {} out of bounds for set of size {}", - index, - set.len() - ))); - } - } - } else { - // No membership set: commitment-only proof - ( - VerifiableProofData::Commitment { - commitment: commitment.commitment, - }, - None, - ) - }; - - Ok(ZkpProof { - privacy_level: PrivacyLevel::Private, - proof_data, - blinding_nonce: Some(nonce), - commitment: Some(commitment.commitment), - merkle_root: root, - circuit_result: None, - generated_at: chrono::Utc::now().to_rfc3339(), - }) -} - -// --------------------------------------------------------------------------- -// Zero-Knowledge proof: blinded Merkle proof with committed witnesses -// --------------------------------------------------------------------------- - -fn generate_zk_proof(request: &ZkpProofRequest) -> Result<ZkpProof, CircuitError> { - let nonce = generate_nonce(&request.claim); - - // Blind the claim: commitment = H(claim || nonce) - let commitment = commit(&request.claim, &nonce); - - // Build a blinded membership set: - // Each leaf is H(original_leaf || shared_nonce) so the verifier - // cannot recover the original leaves, only verify structure. - let (proof_data, root) = if let (Some(ref set), Some(index)) = - (&request.membership_set, request.membership_index) - { - // Blind all leaves - let blinded_leaves: Vec<Vec<u8>> = set - .iter() - .map(|leaf| { - let blinded = commit(leaf, &nonce); - blinded.commitment.to_vec() - }) - .collect(); - - let root = merkle_root(&blinded_leaves); - match merkle_proof(&blinded_leaves, index) { - Some(mp) => (VerifiableProofData::MerkleInclusion(mp), Some(root)), - None => { - return Err(CircuitError::InvalidWitness(format!( - "Membership index {} out of bounds for set of size {}", - index, - set.len() - ))); - } - } - } else { - // No membership set: commitment with reveal proof - // The verifier can check H(claim || nonce) == commitment - // without seeing the claim (they only see the commitment) - ( - VerifiableProofData::Commitment { - commitment: commitment.commitment, - }, - None, - ) - }; - - Ok(ZkpProof { - privacy_level: PrivacyLevel::ZeroKnowledge, - proof_data, - blinding_nonce: Some(nonce), - commitment: Some(commitment.commitment), - merkle_root: root, - circuit_result: None, - generated_at: chrono::Utc::now().to_rfc3339(), - }) -} - -// --------------------------------------------------------------------------- -// Verification helpers -// --------------------------------------------------------------------------- - -fn verify_private_proof(proof: &ZkpProof, claim: &[u8]) -> bool { - // Verify the commitment matches the claim - if let (Some(ref nonce), Some(commitment_hash)) = (&proof.blinding_nonce, proof.commitment) { - let expected = commit(claim, nonce); - if expected.commitment != commitment_hash { - return false; - } - } - - // Verify underlying proof data - match &proof.proof_data { - VerifiableProofData::MerkleInclusion(mp) => verify_merkle_proof(mp), - VerifiableProofData::Commitment { .. } => { - // Commitment is valid by construction (verified above) - true - } - other => verify_proof(other, claim), - } -} - -fn verify_zk_proof(proof: &ZkpProof) -> bool { - // For ZK proofs, we verify the Merkle structure WITHOUT the original data. - // The verifier only checks that the proof path is internally consistent. - match &proof.proof_data { - VerifiableProofData::MerkleInclusion(mp) => { - // Verify the blinded Merkle proof - verify_merkle_proof(mp) - } - VerifiableProofData::Commitment { .. } => { - // Commitment exists — valid at this level. - // Full ZK verification would invoke a ZK-SNARK verifier here - // (sanctify integration, not yet compiled in). - true - } - _ => false, - } -} - -// --------------------------------------------------------------------------- -// Utility -// --------------------------------------------------------------------------- - -/// Generate a deterministic nonce from claim data. -/// In production, replace with a CSPRNG (e.g., `rand::thread_rng().fill_bytes`). -fn generate_nonce(claim: &[u8]) -> Vec<u8> { - let mut h = hash(claim); - // Mix in a domain separator to distinguish from content hashes - let separator = b"verisimdb-zkp-nonce-v1"; - let mut hasher = sha2::Sha256::new(); - use sha2::Digest; - hasher.update(&h); - hasher.update(separator); - h = hasher.finalize().into(); - h.to_vec() -} - -// --------------------------------------------------------------------------- -// Tests -// --------------------------------------------------------------------------- - -#[cfg(test)] -mod tests { - use super::*; - - #[test] - fn test_public_proof_roundtrip() { - let claim = b"entity:123 has-type Person"; - - let request = ZkpProofRequest { - claim: claim.to_vec(), - privacy_level: PrivacyLevel::Public, - circuit_name: None, - witness: None, - public_inputs: None, - membership_set: None, - membership_index: None, - }; - - let proof = generate_zkp(&request).expect("TODO: handle error"); - assert_eq!(proof.privacy_level, PrivacyLevel::Public); - assert!(proof.blinding_nonce.is_none()); - assert!(verify_zkp(&proof, claim)); - } - - #[test] - fn test_public_proof_rejects_wrong_claim() { - let claim = b"entity:123 has-type Person"; - - let request = ZkpProofRequest { - claim: claim.to_vec(), - privacy_level: PrivacyLevel::Public, - circuit_name: None, - witness: None, - public_inputs: None, - membership_set: None, - membership_index: None, - }; - - let proof = generate_zkp(&request).expect("TODO: handle error"); - assert!(!verify_zkp(&proof, b"entity:456 has-type Robot")); - } - - #[test] - fn test_private_proof_commitment() { - let claim = b"confidential-data-hash"; - - let request = ZkpProofRequest { - claim: claim.to_vec(), - privacy_level: PrivacyLevel::Private, - circuit_name: None, - witness: None, - public_inputs: None, - membership_set: None, - membership_index: None, - }; - - let proof = generate_zkp(&request).expect("TODO: handle error"); - assert_eq!(proof.privacy_level, PrivacyLevel::Private); - assert!(proof.blinding_nonce.is_some()); - assert!(proof.commitment.is_some()); - assert!(verify_zkp(&proof, claim)); - } - - #[test] - fn test_private_proof_wrong_claim_fails() { - let claim = b"real-claim"; - - let request = ZkpProofRequest { - claim: claim.to_vec(), - privacy_level: PrivacyLevel::Private, - circuit_name: None, - witness: None, - public_inputs: None, - membership_set: None, - membership_index: None, - }; - - let proof = generate_zkp(&request).expect("TODO: handle error"); - assert!(!verify_zkp(&proof, b"fake-claim")); - } - - #[test] - fn test_private_proof_with_membership_set() { - let claims = vec![ - b"claim-a".to_vec(), - b"claim-b".to_vec(), - b"claim-c".to_vec(), - b"claim-d".to_vec(), - ]; - - let request = ZkpProofRequest { - claim: claims[1].clone(), - privacy_level: PrivacyLevel::Private, - circuit_name: None, - witness: None, - public_inputs: None, - membership_set: Some(claims.clone()), - membership_index: Some(1), - }; - - let proof = generate_zkp(&request).expect("TODO: handle error"); - assert!(proof.merkle_root.is_some()); - assert!(verify_zkp(&proof, &claims[1])); - } - - #[test] - fn test_zk_proof_generation() { - let claim = b"zero-knowledge-secret"; - - let request = ZkpProofRequest { - claim: claim.to_vec(), - privacy_level: PrivacyLevel::ZeroKnowledge, - circuit_name: None, - witness: None, - public_inputs: None, - membership_set: None, - membership_index: None, - }; - - let proof = generate_zkp(&request).expect("TODO: handle error"); - assert_eq!(proof.privacy_level, PrivacyLevel::ZeroKnowledge); - assert!(proof.blinding_nonce.is_some()); - assert!(proof.commitment.is_some()); - // ZK proofs verify without the original claim - assert!(verify_zkp(&proof, claim)); - } - - #[test] - fn test_zk_proof_with_blinded_membership() { - let claims = vec![ - b"secret-1".to_vec(), - b"secret-2".to_vec(), - b"secret-3".to_vec(), - ]; - - let request = ZkpProofRequest { - claim: claims[2].clone(), - privacy_level: PrivacyLevel::ZeroKnowledge, - circuit_name: None, - witness: None, - public_inputs: None, - membership_set: Some(claims.clone()), - membership_index: Some(2), - }; - - let proof = generate_zkp(&request).expect("TODO: handle error"); - assert!(proof.merkle_root.is_some()); - // Blinded proof verifies via Merkle structure - assert!(verify_zkp(&proof, &claims[2])); - } - - #[test] - fn test_zk_proof_blinded_root_differs_from_plain() { - let claims = vec![ - b"a".to_vec(), - b"b".to_vec(), - b"c".to_vec(), - b"d".to_vec(), - ]; - - // Private proof: plain Merkle root - let private_req = ZkpProofRequest { - claim: claims[0].clone(), - privacy_level: PrivacyLevel::Private, - circuit_name: None, - witness: None, - public_inputs: None, - membership_set: Some(claims.clone()), - membership_index: Some(0), - }; - let private_proof = generate_zkp(&private_req).expect("TODO: handle error"); - - // ZK proof: blinded Merkle root - let zk_req = ZkpProofRequest { - claim: claims[0].clone(), - privacy_level: PrivacyLevel::ZeroKnowledge, - circuit_name: None, - witness: None, - public_inputs: None, - membership_set: Some(claims.clone()), - membership_index: Some(0), - }; - let zk_proof = generate_zkp(&zk_req).expect("TODO: handle error"); - - // Roots should differ (one is blinded) - assert_ne!(private_proof.merkle_root, zk_proof.merkle_root); - } - - #[test] - fn test_membership_index_out_of_bounds() { - let claims = vec![b"only-one".to_vec()]; - - let request = ZkpProofRequest { - claim: claims[0].clone(), - privacy_level: PrivacyLevel::Private, - circuit_name: None, - witness: None, - public_inputs: None, - membership_set: Some(claims), - membership_index: Some(5), - }; - - let result = generate_zkp(&request); - assert!(result.is_err()); - } - - #[test] - fn test_generate_with_circuit_registry() { - use super::super::circuit_registry::{ - CircuitIR, CompiledCircuit, R1CSConstraint, sha256_hex, - }; - use std::collections::HashMap; - - let registry = CircuitRegistry::new(); - - // Register a simple multiply circuit: x * y = z - let constraint = R1CSConstraint { - a: HashMap::from([(0, 1.0)]), - b: HashMap::from([(2, 1.0)]), - c: HashMap::from([(1, 1.0)]), - }; - let ir = CircuitIR { - name: "test-mul".to_string(), - num_public_inputs: 2, - num_witness_wires: 1, - num_wires: 3, - constraints: vec![constraint], - parameter_map: HashMap::new(), - }; - let circuit_bytes = serde_json::to_vec(&ir).expect("TODO: handle error"); - let compiled = CompiledCircuit { - ir, - circuit_hash: sha256_hex(&circuit_bytes), - verification_key: vec![0u8; 32], - }; - registry.register_circuit("test-mul", compiled).expect("TODO: handle error"); - - let request = ZkpProofRequest { - claim: b"verified-computation".to_vec(), - privacy_level: PrivacyLevel::Public, - circuit_name: Some("test-mul".to_string()), - witness: Some(vec![4.0]), // y = 4 - public_inputs: Some(vec![3.0, 12.0]), // x = 3, z = 12 - membership_set: None, - membership_index: None, - }; - - let proof = generate_zkp_with_circuit(&request, ®istry).expect("TODO: handle error"); - assert!(proof.circuit_result.is_some()); - - let cr = proof.circuit_result.expect("TODO: handle error"); - assert!(cr.satisfied); - assert_eq!(cr.circuit_name, "test-mul"); - assert_eq!(cr.constraints_checked, 1); - } - - #[test] - fn test_privacy_level_display() { - assert_eq!(PrivacyLevel::Public.to_string(), "Public"); - assert_eq!(PrivacyLevel::Private.to_string(), "Private"); - assert_eq!(PrivacyLevel::ZeroKnowledge.to_string(), "ZeroKnowledge"); - } -} diff --git a/verisimdb/rust-core/verisim-spatial/Cargo.toml b/verisimdb/rust-core/verisim-spatial/Cargo.toml deleted file mode 100644 index 12c71777..00000000 --- a/verisimdb/rust-core/verisim-spatial/Cargo.toml +++ /dev/null @@ -1,26 +0,0 @@ -# SPDX-License-Identifier: MPL-2.0 - -[package] -name = "verisim-spatial" -description = "Spatial modality - geospatial coordinates, geometry, and proximity queries" -version.workspace = true -edition.workspace = true -authors.workspace = true -license.workspace = true - -[dependencies] -serde.workspace = true -serde_json.workspace = true -thiserror.workspace = true -tracing.workspace = true -async-trait.workspace = true -verisim-storage = { path = "../verisim-storage", optional = true } -tokio.workspace = true - -[dev-dependencies] -tempfile = "3" -proptest.workspace = true - -[features] -default = [] -redb-backend = ["verisim-storage/redb-backend"] diff --git a/verisimdb/rust-core/verisim-spatial/src/lib.rs b/verisimdb/rust-core/verisim-spatial/src/lib.rs deleted file mode 100644 index 541cc08d..00000000 --- a/verisimdb/rust-core/verisim-spatial/src/lib.rs +++ /dev/null @@ -1,610 +0,0 @@ -// SPDX-License-Identifier: MPL-2.0 -//! VeriSim Spatial Modality -//! -//! Provides geospatial indexing and querying capabilities for octad entities. -//! Supports WGS84 (EPSG:4326) coordinates by default, with configurable SRID -//! for other coordinate reference systems. -//! -//! # Architecture -//! -//! - **Coordinates**: Latitude/longitude/altitude tuple in WGS84. -//! - **GeometryType**: Point, LineString, Polygon, MultiPoint, MultiPolygon. -//! - **SpatialData**: Full spatial description of an entity including -//! coordinates, geometry type, SRID, and arbitrary properties. -//! - **SpatialStore** trait: Async storage with radius search, bounding box -//! search, and k-nearest-neighbour queries. -//! - **InMemorySpatialStore**: Reference implementation using brute-force -//! distance computation. A production deployment would use an R-tree or -//! similar spatial index. - -#![forbid(unsafe_code)] -#[cfg(feature = "redb-backend")] -pub mod persistent; -#[cfg(feature = "redb-backend")] -pub use persistent::*; -use async_trait::async_trait; -use serde::{Deserialize, Serialize}; -use std::collections::HashMap; -use std::sync::Arc; -use thiserror::Error; -use tokio::sync::RwLock; -use tracing::{debug, instrument}; - -/// Spatial-specific errors -#[derive(Error, Debug)] -pub enum SpatialError { - /// Entity spatial data not found - #[error("Spatial data not found for entity: {0}")] - NotFound(String), - - /// Invalid coordinate values (out of WGS84 range, NaN, etc.) - #[error("Invalid coordinates: {0}")] - InvalidCoordinates(String), - - /// Spatial index error (R-tree corruption, etc.) - #[error("Spatial index error: {0}")] - IndexError(String), - - /// Generic I/O or storage error - #[error("Spatial I/O error: {0}")] - IoError(String), -} - -/// A geographic coordinate in WGS84. -/// -/// Latitude ranges from -90.0 to +90.0 (north positive). -/// Longitude ranges from -180.0 to +180.0 (east positive). -/// Altitude is optional, in metres above the WGS84 ellipsoid. -#[derive(Debug, Clone, Serialize, Deserialize, PartialEq)] -pub struct Coordinates { - /// Latitude in decimal degrees (-90.0 to +90.0) - pub latitude: f64, - /// Longitude in decimal degrees (-180.0 to +180.0) - pub longitude: f64, - /// Altitude in metres above the WGS84 ellipsoid (optional) - pub altitude: Option<f64>, -} - -impl Coordinates { - /// Create new coordinates, validating WGS84 range. - pub fn new(latitude: f64, longitude: f64, altitude: Option<f64>) -> Result<Self, SpatialError> { - if !(-90.0..=90.0).contains(&latitude) { - return Err(SpatialError::InvalidCoordinates(format!( - "Latitude {} out of range [-90, 90]", - latitude - ))); - } - if !(-180.0..=180.0).contains(&longitude) { - return Err(SpatialError::InvalidCoordinates(format!( - "Longitude {} out of range [-180, 180]", - longitude - ))); - } - if latitude.is_nan() || longitude.is_nan() { - return Err(SpatialError::InvalidCoordinates( - "Coordinates must not be NaN".to_string(), - )); - } - Ok(Self { - latitude, - longitude, - altitude, - }) - } - - /// Create coordinates without validation (for internal / trusted use). - pub fn new_unchecked(latitude: f64, longitude: f64, altitude: Option<f64>) -> Self { - Self { - latitude, - longitude, - altitude, - } - } -} - -/// Supported geometry types. -/// -/// Follows the OGC Simple Features specification naming. -#[derive(Debug, Clone, Serialize, Deserialize, PartialEq, Eq, Hash)] -pub enum GeometryType { - /// A single point in 2D or 3D space - Point, - /// An ordered sequence of points forming a line - LineString, - /// A closed ring of points forming a polygon - Polygon, - /// A collection of points - MultiPoint, - /// A collection of polygons - MultiPolygon, -} - -impl std::fmt::Display for GeometryType { - fn fmt(&self, f: &mut std::fmt::Formatter<'_>) -> std::fmt::Result { - match self { - GeometryType::Point => write!(f, "Point"), - GeometryType::LineString => write!(f, "LineString"), - GeometryType::Polygon => write!(f, "Polygon"), - GeometryType::MultiPoint => write!(f, "MultiPoint"), - GeometryType::MultiPolygon => write!(f, "MultiPolygon"), - } - } -} - -/// Full spatial description of an entity. -/// -/// The `coordinates` field holds the representative point (centroid for -/// complex geometries). The `geometry_type` and `srid` describe the -/// coordinate reference context. Arbitrary `properties` can hold extra -/// spatial metadata (e.g., address, region name, accuracy). -#[derive(Debug, Clone, Serialize, Deserialize)] -pub struct SpatialData { - /// Representative coordinates (centroid for complex geometries) - pub coordinates: Coordinates, - /// Type of geometry this entity represents - pub geometry_type: GeometryType, - /// Spatial Reference System Identifier (default: 4326 = WGS84) - pub srid: u32, - /// Arbitrary spatial properties (address, region, accuracy, etc.) - pub properties: HashMap<String, String>, -} - -impl SpatialData { - /// Create spatial data for a point with default WGS84 SRID. - pub fn point(latitude: f64, longitude: f64, altitude: Option<f64>) -> Result<Self, SpatialError> { - Ok(Self { - coordinates: Coordinates::new(latitude, longitude, altitude)?, - geometry_type: GeometryType::Point, - srid: 4326, - properties: HashMap::new(), - }) - } - - /// Create spatial data with a specified geometry type and SRID. - pub fn with_geometry( - coordinates: Coordinates, - geometry_type: GeometryType, - srid: u32, - ) -> Self { - Self { - coordinates, - geometry_type, - srid, - properties: HashMap::new(), - } - } - - /// Add a property to this spatial data. - pub fn with_property(mut self, key: impl Into<String>, value: impl Into<String>) -> Self { - self.properties.insert(key.into(), value.into()); - self - } -} - -/// A bounding box for spatial queries. -/// -/// Defined by the south-west (min) and north-east (max) corners. -#[derive(Debug, Clone, Serialize, Deserialize)] -pub struct BoundingBox { - /// Minimum latitude (south) - pub min_lat: f64, - /// Minimum longitude (west) - pub min_lon: f64, - /// Maximum latitude (north) - pub max_lat: f64, - /// Maximum longitude (east) - pub max_lon: f64, -} - -/// Result of a spatial search, pairing entity ID with its data and distance. -#[derive(Debug, Clone, Serialize, Deserialize)] -pub struct SpatialSearchResult { - /// Entity ID - pub entity_id: String, - /// The entity's spatial data - pub data: SpatialData, - /// Distance from the query point in kilometres - pub distance_km: f64, -} - -/// Async trait for spatial storage backends. -/// -/// Implementations must be `Send + Sync` for safe sharing across Tokio tasks. -#[async_trait] -pub trait SpatialStore: Send + Sync { - /// Index (upsert) spatial data for an entity. - async fn index(&self, entity_id: &str, data: SpatialData) -> Result<(), SpatialError>; - - /// Get spatial data for an entity. - async fn get(&self, entity_id: &str) -> Result<Option<SpatialData>, SpatialError>; - - /// Delete spatial data for an entity. - async fn delete(&self, entity_id: &str) -> Result<(), SpatialError>; - - /// Search for entities within a given radius (km) of a point. - async fn search_radius( - &self, - center: &Coordinates, - radius_km: f64, - limit: usize, - ) -> Result<Vec<SpatialSearchResult>, SpatialError>; - - /// Search for entities within a bounding box. - async fn search_within( - &self, - bounds: &BoundingBox, - limit: usize, - ) -> Result<Vec<SpatialSearchResult>, SpatialError>; - - /// Find the k nearest entities to a given point. - async fn nearest( - &self, - point: &Coordinates, - k: usize, - ) -> Result<Vec<SpatialSearchResult>, SpatialError>; -} - -/// Approximate distance in kilometres between two WGS84 points using the -/// Haversine formula. -/// -/// This is accurate to within ~0.5% for most distances on Earth. -pub fn haversine_distance(a: &Coordinates, b: &Coordinates) -> f64 { - const EARTH_RADIUS_KM: f64 = 6371.0; - - let lat1 = a.latitude.to_radians(); - let lat2 = b.latitude.to_radians(); - let dlat = (b.latitude - a.latitude).to_radians(); - let dlon = (b.longitude - a.longitude).to_radians(); - - let h = (dlat / 2.0).sin().powi(2) - + lat1.cos() * lat2.cos() * (dlon / 2.0).sin().powi(2); - - let c = 2.0 * h.sqrt().asin(); - EARTH_RADIUS_KM * c -} - -/// In-memory implementation of [`SpatialStore`]. -/// -/// Uses brute-force distance computation for searches. Suitable for -/// development, testing, and small-to-medium datasets. A production -/// deployment should use an R-tree or similar spatial index. -pub struct InMemorySpatialStore { - data: Arc<RwLock<HashMap<String, SpatialData>>>, -} - -impl InMemorySpatialStore { - /// Create a new empty in-memory spatial store. - pub fn new() -> Self { - Self { - data: Arc::new(RwLock::new(HashMap::new())), - } - } -} - -impl Default for InMemorySpatialStore { - fn default() -> Self { - Self::new() - } -} - -#[async_trait] -impl SpatialStore for InMemorySpatialStore { - #[instrument(skip(self, data))] - async fn index(&self, entity_id: &str, data: SpatialData) -> Result<(), SpatialError> { - // Validate coordinates even if SpatialData was constructed directly - if !(-90.0..=90.0).contains(&data.coordinates.latitude) - || !(-180.0..=180.0).contains(&data.coordinates.longitude) - { - return Err(SpatialError::InvalidCoordinates(format!( - "lat={}, lon={} out of WGS84 range", - data.coordinates.latitude, data.coordinates.longitude - ))); - } - - let mut store = self.data.write().await; - store.insert(entity_id.to_string(), data); - debug!(entity_id = %entity_id, "Spatial data indexed"); - Ok(()) - } - - async fn get(&self, entity_id: &str) -> Result<Option<SpatialData>, SpatialError> { - let store = self.data.read().await; - Ok(store.get(entity_id).cloned()) - } - - async fn delete(&self, entity_id: &str) -> Result<(), SpatialError> { - let mut store = self.data.write().await; - store.remove(entity_id); - Ok(()) - } - - async fn search_radius( - &self, - center: &Coordinates, - radius_km: f64, - limit: usize, - ) -> Result<Vec<SpatialSearchResult>, SpatialError> { - let store = self.data.read().await; - let mut results: Vec<SpatialSearchResult> = store - .iter() - .filter_map(|(id, data)| { - let dist = haversine_distance(center, &data.coordinates); - if dist <= radius_km { - Some(SpatialSearchResult { - entity_id: id.clone(), - data: data.clone(), - distance_km: dist, - }) - } else { - None - } - }) - .collect(); - - // Sort by distance ascending - results.sort_by(|a, b| { - a.distance_km - .partial_cmp(&b.distance_km) - .unwrap_or(std::cmp::Ordering::Equal) - }); - results.truncate(limit); - Ok(results) - } - - async fn search_within( - &self, - bounds: &BoundingBox, - limit: usize, - ) -> Result<Vec<SpatialSearchResult>, SpatialError> { - let store = self.data.read().await; - // Compute bounding box center for distance calculation - let center = Coordinates::new_unchecked( - (bounds.min_lat + bounds.max_lat) / 2.0, - (bounds.min_lon + bounds.max_lon) / 2.0, - None, - ); - - let mut results: Vec<SpatialSearchResult> = store - .iter() - .filter_map(|(id, data)| { - let lat = data.coordinates.latitude; - let lon = data.coordinates.longitude; - if lat >= bounds.min_lat - && lat <= bounds.max_lat - && lon >= bounds.min_lon - && lon <= bounds.max_lon - { - Some(SpatialSearchResult { - entity_id: id.clone(), - data: data.clone(), - distance_km: haversine_distance(¢er, &data.coordinates), - }) - } else { - None - } - }) - .collect(); - - results.sort_by(|a, b| { - a.distance_km - .partial_cmp(&b.distance_km) - .unwrap_or(std::cmp::Ordering::Equal) - }); - results.truncate(limit); - Ok(results) - } - - async fn nearest( - &self, - point: &Coordinates, - k: usize, - ) -> Result<Vec<SpatialSearchResult>, SpatialError> { - let store = self.data.read().await; - let mut results: Vec<SpatialSearchResult> = store - .iter() - .map(|(id, data)| SpatialSearchResult { - entity_id: id.clone(), - data: data.clone(), - distance_km: haversine_distance(point, &data.coordinates), - }) - .collect(); - - results.sort_by(|a, b| { - a.distance_km - .partial_cmp(&b.distance_km) - .unwrap_or(std::cmp::Ordering::Equal) - }); - results.truncate(k); - Ok(results) - } -} - -#[cfg(test)] -mod tests { - use super::*; - - #[test] - fn test_coordinates_valid() { - let coords = Coordinates::new(51.5074, -0.1278, None); - assert!(coords.is_ok(), "London coordinates should be valid"); - } - - #[test] - fn test_coordinates_invalid_latitude() { - let coords = Coordinates::new(91.0, 0.0, None); - assert!(matches!(coords, Err(SpatialError::InvalidCoordinates(_)))); - } - - #[test] - fn test_coordinates_invalid_longitude() { - let coords = Coordinates::new(0.0, 181.0, None); - assert!(matches!(coords, Err(SpatialError::InvalidCoordinates(_)))); - } - - #[test] - fn test_coordinates_nan() { - let coords = Coordinates::new(f64::NAN, 0.0, None); - assert!(matches!(coords, Err(SpatialError::InvalidCoordinates(_)))); - } - - #[test] - fn test_haversine_same_point() { - let a = Coordinates::new_unchecked(51.5074, -0.1278, None); - let dist = haversine_distance(&a, &a); - assert!(dist < 0.001, "Same point distance should be ~0, got {}", dist); - } - - #[test] - fn test_haversine_london_to_paris() { - let london = Coordinates::new_unchecked(51.5074, -0.1278, None); - let paris = Coordinates::new_unchecked(48.8566, 2.3522, None); - let dist = haversine_distance(&london, &paris); - // London to Paris is ~344 km - assert!( - (330.0..360.0).contains(&dist), - "London-Paris distance should be ~344 km, got {} km", - dist - ); - } - - #[test] - fn test_haversine_antipodal() { - let north = Coordinates::new_unchecked(0.0, 0.0, None); - let south = Coordinates::new_unchecked(0.0, 180.0, None); - let dist = haversine_distance(&north, &south); - // Half circumference ~20015 km - assert!( - dist > 19000.0 && dist < 21000.0, - "Antipodal distance should be ~20015 km, got {} km", - dist - ); - } - - #[test] - fn test_spatial_data_point() { - let data = SpatialData::point(51.5074, -0.1278, None).expect("TODO: handle error"); - assert_eq!(data.geometry_type, GeometryType::Point); - assert_eq!(data.srid, 4326); - } - - #[tokio::test] - async fn test_in_memory_store_index_and_get() { - let store = InMemorySpatialStore::new(); - let data = SpatialData::point(51.5074, -0.1278, None).expect("TODO: handle error"); - - store.index("entity-1", data.clone()).await.expect("TODO: handle error"); - let retrieved = store.get("entity-1").await.expect("TODO: handle error"); - assert!(retrieved.is_some()); - assert_eq!(retrieved.expect("TODO: handle error").coordinates.latitude, 51.5074); - } - - #[tokio::test] - async fn test_in_memory_store_delete() { - let store = InMemorySpatialStore::new(); - let data = SpatialData::point(51.5074, -0.1278, None).expect("TODO: handle error"); - - store.index("entity-1", data).await.expect("TODO: handle error"); - store.delete("entity-1").await.expect("TODO: handle error"); - assert!(store.get("entity-1").await.expect("TODO: handle error").is_none()); - } - - #[tokio::test] - async fn test_in_memory_store_radius_search() { - let store = InMemorySpatialStore::new(); - - // London - store - .index("london", SpatialData::point(51.5074, -0.1278, None).expect("TODO: handle error")) - .await - .expect("TODO: handle error"); - // Paris - store - .index("paris", SpatialData::point(48.8566, 2.3522, None).expect("TODO: handle error")) - .await - .expect("TODO: handle error"); - // New York - store - .index("nyc", SpatialData::point(40.7128, -74.0060, None).expect("TODO: handle error")) - .await - .expect("TODO: handle error"); - - // Search within 500 km of London — should find London and Paris - let center = Coordinates::new(51.5074, -0.1278, None).expect("TODO: handle error"); - let results = store.search_radius(¢er, 500.0, 10).await.expect("TODO: handle error"); - - assert_eq!(results.len(), 2, "Should find London and Paris within 500km"); - assert_eq!(results[0].entity_id, "london", "London should be closest"); - } - - #[tokio::test] - async fn test_in_memory_store_bounding_box() { - let store = InMemorySpatialStore::new(); - - // London - store - .index("london", SpatialData::point(51.5074, -0.1278, None).expect("TODO: handle error")) - .await - .expect("TODO: handle error"); - // Paris - store - .index("paris", SpatialData::point(48.8566, 2.3522, None).expect("TODO: handle error")) - .await - .expect("TODO: handle error"); - // New York - store - .index("nyc", SpatialData::point(40.7128, -74.0060, None).expect("TODO: handle error")) - .await - .expect("TODO: handle error"); - - // Bounding box around Western Europe - let bounds = BoundingBox { - min_lat: 45.0, - min_lon: -5.0, - max_lat: 55.0, - max_lon: 10.0, - }; - - let results = store.search_within(&bounds, 10).await.expect("TODO: handle error"); - assert_eq!(results.len(), 2, "Should find London and Paris in W. Europe box"); - } - - #[tokio::test] - async fn test_in_memory_store_nearest() { - let store = InMemorySpatialStore::new(); - - store - .index("london", SpatialData::point(51.5074, -0.1278, None).expect("TODO: handle error")) - .await - .expect("TODO: handle error"); - store - .index("paris", SpatialData::point(48.8566, 2.3522, None).expect("TODO: handle error")) - .await - .expect("TODO: handle error"); - store - .index("nyc", SpatialData::point(40.7128, -74.0060, None).expect("TODO: handle error")) - .await - .expect("TODO: handle error"); - - // k-nearest to London (k=2) — should return London then Paris - let point = Coordinates::new(51.5074, -0.1278, None).expect("TODO: handle error"); - let results = store.nearest(&point, 2).await.expect("TODO: handle error"); - - assert_eq!(results.len(), 2); - assert_eq!(results[0].entity_id, "london"); - assert_eq!(results[1].entity_id, "paris"); - } - - #[tokio::test] - async fn test_in_memory_store_invalid_coordinates() { - let store = InMemorySpatialStore::new(); - let data = SpatialData { - coordinates: Coordinates::new_unchecked(999.0, 0.0, None), - geometry_type: GeometryType::Point, - srid: 4326, - properties: HashMap::new(), - }; - - let result = store.index("bad", data).await; - assert!(matches!(result, Err(SpatialError::InvalidCoordinates(_)))); - } -} diff --git a/verisimdb/rust-core/verisim-spatial/src/persistent.rs b/verisimdb/rust-core/verisim-spatial/src/persistent.rs deleted file mode 100644 index 08dba5ba..00000000 --- a/verisimdb/rust-core/verisim-spatial/src/persistent.rs +++ /dev/null @@ -1,292 +0,0 @@ -// SPDX-License-Identifier: MPL-2.0 -// Copyright (c) 2026 Jonathan D.A. Jewell (hyperpolymath) <j.d.a.jewell@open.ac.uk> -// -// Persistent spatial store backed by redb via verisim-storage. -// -// Stores spatial data in redb for durability. On open(), all entries are -// scanned into an in-memory HashMap cache for fast brute-force queries. -// Writes go to redb first (durable), then update the cache. -// -// Uses tokio::sync::RwLock to match the async locking pattern of the -// InMemorySpatialStore. - -use std::collections::HashMap; -use std::path::Path; -use std::sync::Arc; - -use async_trait::async_trait; -use tracing::{debug, info, instrument}; -use verisim_storage::redb_backend::RedbBackend; -use verisim_storage::typed::TypedStore; - -use crate::{ - haversine_distance, BoundingBox, Coordinates, SpatialData, SpatialError, SpatialSearchResult, - SpatialStore, -}; - -/// Persistent spatial store: redb for durability, async RwLock cache for fast -/// brute-force spatial queries. -/// -/// A production deployment would rebuild an R-tree from the cached data. -/// Currently uses the same brute-force approach as InMemorySpatialStore. -pub struct RedbSpatialStore { - /// Typed store for spatial data, keyed by entity_id. - store: TypedStore<RedbBackend>, - /// In-memory cache of all spatial data. - data: Arc<tokio::sync::RwLock<HashMap<String, SpatialData>>>, -} - -impl RedbSpatialStore { - /// Open (or create) a persistent spatial store at the given path. - /// - /// On open, all existing spatial data is scanned from redb into the - /// in-memory cache so that reads and spatial queries never hit disk. - pub async fn open(path: impl AsRef<Path>) -> Result<Self, SpatialError> { - let backend = RedbBackend::open(path.as_ref()) - .map_err(|e| SpatialError::IoError(format!("redb open: {}", e)))?; - let store = TypedStore::new(backend, "spatial"); - - let entries: Vec<(String, SpatialData)> = store - .scan_prefix("", 1_000_000) - .await - .map_err(|e| SpatialError::IoError(format!("scan: {}", e)))?; - - let mut cache = HashMap::new(); - for (id, data) in entries { - cache.insert(id, data); - } - - info!(count = cache.len(), "Loaded spatial store from redb"); - Ok(Self { - store, - data: Arc::new(tokio::sync::RwLock::new(cache)), - }) - } -} - -#[async_trait] -impl SpatialStore for RedbSpatialStore { - #[instrument(skip(self, data))] - async fn index(&self, entity_id: &str, data: SpatialData) -> Result<(), SpatialError> { - // Validate coordinates even if SpatialData was constructed directly. - if !(-90.0..=90.0).contains(&data.coordinates.latitude) - || !(-180.0..=180.0).contains(&data.coordinates.longitude) - { - return Err(SpatialError::InvalidCoordinates(format!( - "lat={}, lon={} out of WGS84 range", - data.coordinates.latitude, data.coordinates.longitude - ))); - } - - // Write to redb first (durable). - self.store - .put(entity_id, &data) - .await - .map_err(|e| SpatialError::IoError(format!("put: {}", e)))?; - - // Update cache. - let mut cache = self.data.write().await; - cache.insert(entity_id.to_string(), data); - debug!(entity_id = %entity_id, "Spatial data indexed (persistent)"); - Ok(()) - } - - async fn get(&self, entity_id: &str) -> Result<Option<SpatialData>, SpatialError> { - let cache = self.data.read().await; - Ok(cache.get(entity_id).cloned()) - } - - async fn delete(&self, entity_id: &str) -> Result<(), SpatialError> { - // Delete from redb first. - self.store - .delete(entity_id) - .await - .map_err(|e| SpatialError::IoError(format!("delete: {}", e)))?; - // Then remove from cache. - let mut cache = self.data.write().await; - cache.remove(entity_id); - Ok(()) - } - - async fn search_radius( - &self, - center: &Coordinates, - radius_km: f64, - limit: usize, - ) -> Result<Vec<SpatialSearchResult>, SpatialError> { - let cache = self.data.read().await; - let mut results: Vec<SpatialSearchResult> = cache - .iter() - .filter_map(|(id, data)| { - let dist = haversine_distance(center, &data.coordinates); - if dist <= radius_km { - Some(SpatialSearchResult { - entity_id: id.clone(), - data: data.clone(), - distance_km: dist, - }) - } else { - None - } - }) - .collect(); - - results.sort_by(|a, b| { - a.distance_km - .partial_cmp(&b.distance_km) - .unwrap_or(std::cmp::Ordering::Equal) - }); - results.truncate(limit); - Ok(results) - } - - async fn search_within( - &self, - bounds: &BoundingBox, - limit: usize, - ) -> Result<Vec<SpatialSearchResult>, SpatialError> { - let cache = self.data.read().await; - let center = Coordinates::new_unchecked( - (bounds.min_lat + bounds.max_lat) / 2.0, - (bounds.min_lon + bounds.max_lon) / 2.0, - None, - ); - - let mut results: Vec<SpatialSearchResult> = cache - .iter() - .filter_map(|(id, data)| { - let lat = data.coordinates.latitude; - let lon = data.coordinates.longitude; - if lat >= bounds.min_lat - && lat <= bounds.max_lat - && lon >= bounds.min_lon - && lon <= bounds.max_lon - { - Some(SpatialSearchResult { - entity_id: id.clone(), - data: data.clone(), - distance_km: haversine_distance(¢er, &data.coordinates), - }) - } else { - None - } - }) - .collect(); - - results.sort_by(|a, b| { - a.distance_km - .partial_cmp(&b.distance_km) - .unwrap_or(std::cmp::Ordering::Equal) - }); - results.truncate(limit); - Ok(results) - } - - async fn nearest( - &self, - point: &Coordinates, - k: usize, - ) -> Result<Vec<SpatialSearchResult>, SpatialError> { - let cache = self.data.read().await; - let mut results: Vec<SpatialSearchResult> = cache - .iter() - .map(|(id, data)| SpatialSearchResult { - entity_id: id.clone(), - data: data.clone(), - distance_km: haversine_distance(point, &data.coordinates), - }) - .collect(); - - results.sort_by(|a, b| { - a.distance_km - .partial_cmp(&b.distance_km) - .unwrap_or(std::cmp::Ordering::Equal) - }); - results.truncate(k); - Ok(results) - } -} - -#[cfg(test)] -mod tests { - use super::*; - - #[tokio::test] - async fn test_persistent_spatial_roundtrip() { - let dir = tempfile::tempdir().expect("TODO: handle error"); - let path = dir.path().join("spatial.redb"); - - // Write data in one session. - { - let store = RedbSpatialStore::open(&path).await.expect("TODO: handle error"); - let london = SpatialData::point(51.5074, -0.1278, Some(11.0)).expect("TODO: handle error"); - store.index("london", london).await.expect("TODO: handle error"); - - let paris = SpatialData::point(48.8566, 2.3522, None).expect("TODO: handle error"); - store.index("paris", paris).await.expect("TODO: handle error"); - } - - // Reopen and verify data survived. - { - let store = RedbSpatialStore::open(&path).await.expect("TODO: handle error"); - - let london = store.get("london").await.expect("TODO: handle error").expect("TODO: handle error"); - assert!((london.coordinates.latitude - 51.5074).abs() < 0.001); - assert!((london.coordinates.longitude - (-0.1278)).abs() < 0.001); - - let paris = store.get("paris").await.expect("TODO: handle error").expect("TODO: handle error"); - assert!((paris.coordinates.latitude - 48.8566).abs() < 0.001); - - // Test radius search — 500 km from London should find both cities. - let center = Coordinates::new(51.5074, -0.1278, None).expect("TODO: handle error"); - let results = store.search_radius(¢er, 500.0, 10).await.expect("TODO: handle error"); - assert_eq!(results.len(), 2); - assert_eq!(results[0].entity_id, "london"); - - // Test bounding box — Western Europe. - let bounds = BoundingBox { - min_lat: 45.0, - min_lon: -5.0, - max_lat: 55.0, - max_lon: 10.0, - }; - let bbox_results = store.search_within(&bounds, 10).await.expect("TODO: handle error"); - assert_eq!(bbox_results.len(), 2); - - // Test nearest. - let nearest = store.nearest(¢er, 1).await.expect("TODO: handle error"); - assert_eq!(nearest.len(), 1); - assert_eq!(nearest[0].entity_id, "london"); - } - } - - #[tokio::test] - async fn test_persistent_spatial_delete() { - let dir = tempfile::tempdir().expect("TODO: handle error"); - let path = dir.path().join("spatial-del.redb"); - - let store = RedbSpatialStore::open(&path).await.expect("TODO: handle error"); - let data = SpatialData::point(51.5074, -0.1278, None).expect("TODO: handle error"); - store.index("london", data).await.expect("TODO: handle error"); - - store.delete("london").await.expect("TODO: handle error"); - assert!(store.get("london").await.expect("TODO: handle error").is_none()); - } - - #[tokio::test] - async fn test_persistent_spatial_invalid_coordinates() { - let dir = tempfile::tempdir().expect("TODO: handle error"); - let path = dir.path().join("spatial-invalid.redb"); - - let store = RedbSpatialStore::open(&path).await.expect("TODO: handle error"); - let bad = SpatialData { - coordinates: Coordinates::new_unchecked(999.0, 0.0, None), - geometry_type: crate::GeometryType::Point, - srid: 4326, - properties: HashMap::new(), - }; - - let result = store.index("bad", bad).await; - assert!(matches!(result, Err(SpatialError::InvalidCoordinates(_)))); - } -} diff --git a/verisimdb/rust-core/verisim-storage/Cargo.toml b/verisimdb/rust-core/verisim-storage/Cargo.toml deleted file mode 100644 index ef0cf81d..00000000 --- a/verisimdb/rust-core/verisim-storage/Cargo.toml +++ /dev/null @@ -1,30 +0,0 @@ -# SPDX-License-Identifier: MPL-2.0 - -[package] -name = "verisim-storage" -description = "Pluggable storage backend abstraction for VeriSimDB" -version.workspace = true -edition.workspace = true -authors.workspace = true -license.workspace = true - -[features] -default = [] -# Enable the redb persistent backend (pure Rust, B-tree, ACID, single-file). -redb-backend = ["dep:redb"] - -[dependencies] -serde = { workspace = true } -serde_json = { workspace = true } -thiserror = { workspace = true } -tracing = { workspace = true } -async-trait = { workspace = true } -tokio = { workspace = true } - -# Optional: redb for persistent on-disk storage (pure Rust, no C/C++) -redb = { workspace = true, optional = true } - -[dev-dependencies] -tempfile = "3" -tokio = { workspace = true, features = ["test-util"] } -tokio-test = "0.4" diff --git a/verisimdb/rust-core/verisim-storage/src/backend.rs b/verisimdb/rust-core/verisim-storage/src/backend.rs deleted file mode 100644 index 93e2f924..00000000 --- a/verisimdb/rust-core/verisim-storage/src/backend.rs +++ /dev/null @@ -1,74 +0,0 @@ -// SPDX-License-Identifier: MPL-2.0 -// Copyright (c) 2026 Jonathan D.A. Jewell (hyperpolymath) <j.d.a.jewell@open.ac.uk> -// -// Core storage backend trait for VeriSimDB. -// -// Defines the `StorageBackend` trait that all storage implementations must -// satisfy. The trait provides a key-value interface with support for batch -// operations, prefix scanning, and flush semantics. Backends are expected -// to be thread-safe (`Send + Sync`) and fully asynchronous. - -use async_trait::async_trait; - -use crate::error::StorageError; - -/// A pluggable key-value storage backend. -/// -/// All keys and values are opaque byte slices. Higher-level typed access -/// is provided by [`crate::typed::TypedStore`], which wraps a backend with -/// serde-based serialization and namespace prefixing. -/// -/// Implementations must be safe to share across threads and tokio tasks. -#[async_trait] -pub trait StorageBackend: Send + Sync { - /// Retrieve the value associated with `key`. - /// - /// Returns `Ok(None)` if the key does not exist, rather than an error. - async fn get(&self, key: &[u8]) -> Result<Option<Vec<u8>>, StorageError>; - - /// Store a key-value pair, overwriting any previous value for `key`. - async fn put(&self, key: &[u8], value: &[u8]) -> Result<(), StorageError>; - - /// Delete the value associated with `key`. - /// - /// Returns `Ok(true)` if the key existed and was removed, `Ok(false)` if - /// the key was not present. - async fn delete(&self, key: &[u8]) -> Result<bool, StorageError>; - - /// Check whether `key` exists in the store without retrieving its value. - async fn exists(&self, key: &[u8]) -> Result<bool, StorageError>; - - /// Scan all keys that start with `prefix`, returning up to `limit` - /// (key, value) pairs in lexicographic order. - async fn scan_prefix( - &self, - prefix: &[u8], - limit: usize, - ) -> Result<Vec<(Vec<u8>, Vec<u8>)>, StorageError>; - - /// Retrieve multiple keys in a single call. - /// - /// The returned vector has the same length as `keys`, with `None` for any - /// key that was not found. - async fn multi_get(&self, keys: &[&[u8]]) -> Result<Vec<Option<Vec<u8>>>, StorageError>; - - /// Write multiple key-value pairs atomically. - /// - /// Either all entries are written or none are. Implementations that cannot - /// guarantee atomicity should document this limitation. - async fn batch_put(&self, entries: &[(&[u8], &[u8])]) -> Result<(), StorageError>; - - /// Flush any buffered writes to durable storage. - /// - /// For in-memory backends this is a no-op. For disk-backed backends this - /// should ensure that all previously written data survives a process crash. - async fn flush(&self) -> Result<(), StorageError>; - - /// A human-readable name for this backend, used in logging and metrics. - fn name(&self) -> &str; - - /// Return the approximate total size of stored data in bytes, if known. - /// - /// Backends that cannot cheaply compute this may return `Ok(None)`. - async fn approximate_size(&self) -> Result<Option<u64>, StorageError>; -} diff --git a/verisimdb/rust-core/verisim-storage/src/error.rs b/verisimdb/rust-core/verisim-storage/src/error.rs deleted file mode 100644 index c284ff46..00000000 --- a/verisimdb/rust-core/verisim-storage/src/error.rs +++ /dev/null @@ -1,104 +0,0 @@ -// SPDX-License-Identifier: MPL-2.0 -// Copyright (c) 2026 Jonathan D.A. Jewell (hyperpolymath) <j.d.a.jewell@open.ac.uk> -// -// Storage error types for VeriSimDB backend abstraction. -// -// Provides a unified error enum covering all failure modes that a storage -// backend may encounter: I/O errors, missing keys, serialization failures, -// data corruption, backend unavailability, and size limit violations. - -use thiserror::Error; - -/// Errors that can occur when interacting with a storage backend. -#[derive(Debug, Error)] -pub enum StorageError { - /// An I/O error occurred in the underlying storage layer. - #[error("I/O error: {0}")] - Io(#[from] std::io::Error), - - /// The requested key was not found. - #[error("key not found: {0}")] - NotFound(String), - - /// Failed to serialize or deserialize a value. - #[error("serialization error: {0}")] - SerializationError(String), - - /// The stored data is corrupted or in an unexpected format. - #[error("corrupted data: {0}")] - CorruptedData(String), - - /// The storage backend is not available (e.g., connection lost). - #[error("backend unavailable: {0}")] - BackendUnavailable(String), - - /// The key exceeds the maximum allowed size. - #[error("key too large: {size} bytes (max: {max})")] - KeyTooLarge { - /// Actual key size in bytes. - size: usize, - /// Maximum allowed key size in bytes. - max: usize, - }, - - /// The value exceeds the maximum allowed size. - #[error("value too large: {size} bytes (max: {max})")] - ValueTooLarge { - /// Actual value size in bytes. - size: usize, - /// Maximum allowed value size in bytes. - max: usize, - }, -} - -#[cfg(test)] -mod tests { - use super::*; - - #[test] - fn test_io_error_display() { - let io_err = std::io::Error::new(std::io::ErrorKind::NotFound, "file gone"); - let err = StorageError::Io(io_err); - assert!(err.to_string().contains("I/O error")); - } - - #[test] - fn test_not_found_display() { - let err = StorageError::NotFound("my-key".to_string()); - assert_eq!(err.to_string(), "key not found: my-key"); - } - - #[test] - fn test_serialization_error_display() { - let err = StorageError::SerializationError("bad json".to_string()); - assert!(err.to_string().contains("serialization error")); - } - - #[test] - fn test_corrupted_data_display() { - let err = StorageError::CorruptedData("checksum mismatch".to_string()); - assert!(err.to_string().contains("corrupted data")); - } - - #[test] - fn test_backend_unavailable_display() { - let err = StorageError::BackendUnavailable("connection refused".to_string()); - assert!(err.to_string().contains("backend unavailable")); - } - - #[test] - fn test_key_too_large_display() { - let err = StorageError::KeyTooLarge { size: 2048, max: 1024 }; - assert!(err.to_string().contains("key too large")); - assert!(err.to_string().contains("2048")); - assert!(err.to_string().contains("1024")); - } - - #[test] - fn test_value_too_large_display() { - let err = StorageError::ValueTooLarge { size: 4096, max: 2048 }; - assert!(err.to_string().contains("value too large")); - assert!(err.to_string().contains("4096")); - assert!(err.to_string().contains("2048")); - } -} diff --git a/verisimdb/rust-core/verisim-storage/src/lib.rs b/verisimdb/rust-core/verisim-storage/src/lib.rs deleted file mode 100644 index 86e50250..00000000 --- a/verisimdb/rust-core/verisim-storage/src/lib.rs +++ /dev/null @@ -1,61 +0,0 @@ -// SPDX-License-Identifier: MPL-2.0 -// Copyright (c) 2026 Jonathan D.A. Jewell (hyperpolymath) <j.d.a.jewell@open.ac.uk> -// -// VeriSimDB Storage Backend Abstraction -// -// This crate provides a pluggable key-value storage interface for VeriSimDB. -// The core `StorageBackend` trait defines the contract that all backends must -// implement, enabling the database engine to swap storage implementations -// without changing application logic. -// -// # Modules -// -// - [`backend`] -- The `StorageBackend` trait defining the key-value interface. -// - [`error`] -- The `StorageError` enum covering all backend failure modes. -// - [`memory`] -- An in-memory `BTreeMap`-based backend for testing and -// ephemeral workloads. -// - [`typed`] -- A serde-based typed wrapper with namespace prefixing. -// - [`metrics`] -- A transparent wrapper that collects operation statistics. -// -// # Example -// -// ```rust -// use verisim_storage::backend::StorageBackend; -// use verisim_storage::memory::InMemoryBackend; -// use verisim_storage::typed::TypedStore; -// use verisim_storage::metrics::MetricsBackend; -// -// # tokio_test::block_on(async { -// // Create an in-memory backend with metrics collection. -// let raw = InMemoryBackend::new(); -// let metered = MetricsBackend::new(raw); -// -// // Use a typed store for structured data. -// let store = TypedStore::new(metered, "entities"); -// store.put("e1", &serde_json::json!({"name": "test"})).await.unwrap(); -// -// let val: serde_json::Value = store.get("e1").await.unwrap().unwrap(); -// assert_eq!(val["name"], "test"); -// # }); -// ``` - -#![forbid(unsafe_code)] -pub mod backend; -pub mod error; -pub mod memory; -pub mod metrics; -pub mod typed; - -// Optional persistent backends — feature-gated to keep the default build lean. -#[cfg(feature = "redb-backend")] -pub mod redb_backend; - -// Re-export the most commonly used types at the crate root for convenience. -pub use backend::StorageBackend; -pub use error::StorageError; -pub use memory::InMemoryBackend; -pub use metrics::{BackendStats, MetricsBackend}; -pub use typed::TypedStore; - -#[cfg(feature = "redb-backend")] -pub use redb_backend::RedbBackend; diff --git a/verisimdb/rust-core/verisim-storage/src/memory.rs b/verisimdb/rust-core/verisim-storage/src/memory.rs deleted file mode 100644 index 266082a1..00000000 --- a/verisimdb/rust-core/verisim-storage/src/memory.rs +++ /dev/null @@ -1,282 +0,0 @@ -// SPDX-License-Identifier: MPL-2.0 -// Copyright (c) 2026 Jonathan D.A. Jewell (hyperpolymath) <j.d.a.jewell@open.ac.uk> -// -// In-memory storage backend for VeriSimDB. -// -// Uses a `BTreeMap` wrapped in a tokio `RwLock` for thread-safe, ordered -// key-value storage. The BTreeMap ordering enables efficient prefix scanning. -// Intended for testing, development, and small ephemeral datasets. - -use std::collections::BTreeMap; -use std::sync::Arc; - -use async_trait::async_trait; -use tokio::sync::RwLock; - -use crate::backend::StorageBackend; -use crate::error::StorageError; - -/// An in-memory storage backend backed by a sorted `BTreeMap`. -/// -/// All data lives in process memory and is lost on drop. Thread-safe via -/// `Arc<RwLock<...>>`, making it suitable for concurrent tokio tasks. -/// -/// # Example -/// -/// ```rust -/// use verisim_storage::memory::InMemoryBackend; -/// use verisim_storage::backend::StorageBackend; -/// -/// # tokio_test::block_on(async { -/// let store = InMemoryBackend::new(); -/// store.put(b"hello", b"world").await.unwrap(); -/// let val = store.get(b"hello").await.unwrap(); -/// assert_eq!(val, Some(b"world".to_vec())); -/// # }); -/// ``` -#[derive(Debug, Clone)] -pub struct InMemoryBackend { - /// The underlying sorted map, protected by a read-write lock. - data: Arc<RwLock<BTreeMap<Vec<u8>, Vec<u8>>>>, -} - -impl InMemoryBackend { - /// Create a new, empty in-memory backend. - pub fn new() -> Self { - Self { - data: Arc::new(RwLock::new(BTreeMap::new())), - } - } - - /// Return the number of keys currently stored. - pub async fn len(&self) -> usize { - self.data.read().await.len() - } - - /// Return true if the store contains no keys. - pub async fn is_empty(&self) -> bool { - self.data.read().await.is_empty() - } -} - -impl Default for InMemoryBackend { - fn default() -> Self { - Self::new() - } -} - -#[async_trait] -impl StorageBackend for InMemoryBackend { - async fn get(&self, key: &[u8]) -> Result<Option<Vec<u8>>, StorageError> { - let map = self.data.read().await; - Ok(map.get(key).cloned()) - } - - async fn put(&self, key: &[u8], value: &[u8]) -> Result<(), StorageError> { - let mut map = self.data.write().await; - map.insert(key.to_vec(), value.to_vec()); - Ok(()) - } - - async fn delete(&self, key: &[u8]) -> Result<bool, StorageError> { - let mut map = self.data.write().await; - Ok(map.remove(key).is_some()) - } - - async fn exists(&self, key: &[u8]) -> Result<bool, StorageError> { - let map = self.data.read().await; - Ok(map.contains_key(key)) - } - - async fn scan_prefix( - &self, - prefix: &[u8], - limit: usize, - ) -> Result<Vec<(Vec<u8>, Vec<u8>)>, StorageError> { - let map = self.data.read().await; - let results = map - .range(prefix.to_vec()..) - .take_while(|(k, _)| k.starts_with(prefix)) - .take(limit) - .map(|(k, v)| (k.clone(), v.clone())) - .collect(); - Ok(results) - } - - async fn multi_get(&self, keys: &[&[u8]]) -> Result<Vec<Option<Vec<u8>>>, StorageError> { - let map = self.data.read().await; - let results = keys - .iter() - .map(|key| map.get(*key).cloned()) - .collect(); - Ok(results) - } - - async fn batch_put(&self, entries: &[(&[u8], &[u8])]) -> Result<(), StorageError> { - let mut map = self.data.write().await; - for (key, value) in entries { - map.insert(key.to_vec(), value.to_vec()); - } - Ok(()) - } - - async fn flush(&self) -> Result<(), StorageError> { - // No-op for in-memory backend: all writes are immediately visible. - Ok(()) - } - - fn name(&self) -> &str { - "in-memory" - } - - async fn approximate_size(&self) -> Result<Option<u64>, StorageError> { - let map = self.data.read().await; - let size: u64 = map - .iter() - .map(|(k, v)| (k.len() + v.len()) as u64) - .sum(); - Ok(Some(size)) - } -} - -#[cfg(test)] -mod tests { - use super::*; - - #[tokio::test] - async fn test_basic_crud() { - let backend = InMemoryBackend::new(); - - // Initially empty. - assert!(backend.is_empty().await); - assert_eq!(backend.get(b"key1").await.expect("TODO: handle error"), None); - assert!(!backend.exists(b"key1").await.expect("TODO: handle error")); - - // Put and get. - backend.put(b"key1", b"value1").await.expect("TODO: handle error"); - assert_eq!(backend.get(b"key1").await.expect("TODO: handle error"), Some(b"value1".to_vec())); - assert!(backend.exists(b"key1").await.expect("TODO: handle error")); - assert_eq!(backend.len().await, 1); - - // Overwrite. - backend.put(b"key1", b"updated").await.expect("TODO: handle error"); - assert_eq!(backend.get(b"key1").await.expect("TODO: handle error"), Some(b"updated".to_vec())); - assert_eq!(backend.len().await, 1); - - // Delete existing key. - assert!(backend.delete(b"key1").await.expect("TODO: handle error")); - assert_eq!(backend.get(b"key1").await.expect("TODO: handle error"), None); - assert!(backend.is_empty().await); - - // Delete non-existent key. - assert!(!backend.delete(b"nonexistent").await.expect("TODO: handle error")); - } - - #[tokio::test] - async fn test_scan_prefix() { - let backend = InMemoryBackend::new(); - - // Insert keys with different prefixes. - backend.put(b"user:1:name", b"Alice").await.expect("TODO: handle error"); - backend.put(b"user:1:age", b"30").await.expect("TODO: handle error"); - backend.put(b"user:2:name", b"Bob").await.expect("TODO: handle error"); - backend.put(b"post:1:title", b"Hello").await.expect("TODO: handle error"); - - // Scan with prefix "user:1:". - let results = backend.scan_prefix(b"user:1:", 10).await.expect("TODO: handle error"); - assert_eq!(results.len(), 2); - // BTreeMap ordering: "user:1:age" < "user:1:name". - assert_eq!(results[0].0, b"user:1:age".to_vec()); - assert_eq!(results[1].0, b"user:1:name".to_vec()); - - // Scan with prefix "user:" — should return all user keys. - let results = backend.scan_prefix(b"user:", 10).await.expect("TODO: handle error"); - assert_eq!(results.len(), 3); - - // Scan with limit. - let results = backend.scan_prefix(b"user:", 2).await.expect("TODO: handle error"); - assert_eq!(results.len(), 2); - - // Scan with no matching prefix. - let results = backend.scan_prefix(b"missing:", 10).await.expect("TODO: handle error"); - assert!(results.is_empty()); - } - - #[tokio::test] - async fn test_multi_get() { - let backend = InMemoryBackend::new(); - - backend.put(b"a", b"1").await.expect("TODO: handle error"); - backend.put(b"b", b"2").await.expect("TODO: handle error"); - backend.put(b"c", b"3").await.expect("TODO: handle error"); - - let results = backend - .multi_get(&[b"a" as &[u8], b"missing", b"c"]) - .await - .expect("TODO: handle error"); - - assert_eq!(results.len(), 3); - assert_eq!(results[0], Some(b"1".to_vec())); - assert_eq!(results[1], None); - assert_eq!(results[2], Some(b"3".to_vec())); - } - - #[tokio::test] - async fn test_batch_put() { - let backend = InMemoryBackend::new(); - - backend - .batch_put(&[ - (b"x" as &[u8], b"10" as &[u8]), - (b"y", b"20"), - (b"z", b"30"), - ]) - .await - .expect("TODO: handle error"); - - assert_eq!(backend.len().await, 3); - assert_eq!(backend.get(b"x").await.expect("TODO: handle error"), Some(b"10".to_vec())); - assert_eq!(backend.get(b"y").await.expect("TODO: handle error"), Some(b"20".to_vec())); - assert_eq!(backend.get(b"z").await.expect("TODO: handle error"), Some(b"30".to_vec())); - } - - #[tokio::test] - async fn test_flush_is_noop() { - let backend = InMemoryBackend::new(); - backend.put(b"key", b"val").await.expect("TODO: handle error"); - // Flush should succeed without error. - backend.flush().await.expect("TODO: handle error"); - // Data should still be there. - assert_eq!(backend.get(b"key").await.expect("TODO: handle error"), Some(b"val".to_vec())); - } - - #[tokio::test] - async fn test_name() { - let backend = InMemoryBackend::new(); - assert_eq!(backend.name(), "in-memory"); - } - - #[tokio::test] - async fn test_approximate_size() { - let backend = InMemoryBackend::new(); - - // Empty store has zero size. - assert_eq!(backend.approximate_size().await.expect("TODO: handle error"), Some(0)); - - // Size accounts for both keys and values. - backend.put(b"abc", b"defgh").await.expect("TODO: handle error"); // 3 + 5 = 8 - assert_eq!(backend.approximate_size().await.expect("TODO: handle error"), Some(8)); - - backend.put(b"xy", b"z").await.expect("TODO: handle error"); // 2 + 1 = 3, total = 11 - assert_eq!(backend.approximate_size().await.expect("TODO: handle error"), Some(11)); - } - - #[tokio::test] - async fn test_clone_shares_state() { - let backend = InMemoryBackend::new(); - let clone = backend.clone(); - - backend.put(b"shared", b"data").await.expect("TODO: handle error"); - assert_eq!(clone.get(b"shared").await.expect("TODO: handle error"), Some(b"data".to_vec())); - } -} diff --git a/verisimdb/rust-core/verisim-storage/src/metrics.rs b/verisimdb/rust-core/verisim-storage/src/metrics.rs deleted file mode 100644 index 38be56a9..00000000 --- a/verisimdb/rust-core/verisim-storage/src/metrics.rs +++ /dev/null @@ -1,380 +0,0 @@ -// SPDX-License-Identifier: MPL-2.0 -// Copyright (c) 2026 Jonathan D.A. Jewell (hyperpolymath) <j.d.a.jewell@open.ac.uk> -// -// Metrics-collecting wrapper for VeriSimDB storage backends. -// -// Wraps any `StorageBackend` and transparently collects operation counts, -// latency sums, and byte transfer totals. Useful for profiling, dashboards, -// and adaptive query planning within the VeriSimDB normalizer and drift -// detection systems. - -use std::sync::Arc; -use std::time::Instant; - -use async_trait::async_trait; -use tokio::sync::RwLock; - -use crate::backend::StorageBackend; -use crate::error::StorageError; - -/// Accumulated statistics for a storage backend. -/// -/// All counters are monotonically increasing for the lifetime of the -/// [`MetricsBackend`] that owns them. -#[derive(Debug, Clone, Default)] -pub struct BackendStats { - /// Number of `get` operations performed. - pub get_count: u64, - /// Number of `put` operations performed. - pub put_count: u64, - /// Number of `delete` operations performed. - pub delete_count: u64, - /// Number of `scan_prefix` operations performed. - pub scan_count: u64, - /// Cumulative wall-clock latency of all `get` calls, in milliseconds. - pub get_latency_sum_ms: f64, - /// Cumulative wall-clock latency of all `put` calls, in milliseconds. - pub put_latency_sum_ms: f64, - /// Total bytes read across all `get` and `multi_get` operations. - pub total_bytes_read: u64, - /// Total bytes written across all `put` and `batch_put` operations. - pub total_bytes_written: u64, -} - -/// A storage backend wrapper that collects operation metrics. -/// -/// Delegates every operation to an inner backend while measuring wall-clock -/// latency and counting invocations. Statistics are available via -/// [`MetricsBackend::stats`]. -/// -/// # Example -/// -/// ```rust -/// use verisim_storage::memory::InMemoryBackend; -/// use verisim_storage::metrics::MetricsBackend; -/// use verisim_storage::backend::StorageBackend; -/// -/// # tokio_test::block_on(async { -/// let inner = InMemoryBackend::new(); -/// let metered = MetricsBackend::new(inner); -/// -/// metered.put(b"key", b"value").await.unwrap(); -/// metered.get(b"key").await.unwrap(); -/// -/// let stats = metered.stats().await; -/// assert_eq!(stats.put_count, 1); -/// assert_eq!(stats.get_count, 1); -/// # }); -/// ``` -pub struct MetricsBackend<B: StorageBackend> { - /// The wrapped backend that performs the actual storage operations. - inner: B, - /// Shared, mutable statistics accumulator. - stats: Arc<RwLock<BackendStats>>, -} - -impl<B: StorageBackend> MetricsBackend<B> { - /// Wrap `inner` with metrics collection. - pub fn new(inner: B) -> Self { - Self { - inner, - stats: Arc::new(RwLock::new(BackendStats::default())), - } - } - - /// Return a snapshot of the current statistics. - pub async fn stats(&self) -> BackendStats { - self.stats.read().await.clone() - } - - /// Reset all statistics to zero. - pub async fn reset_stats(&self) { - let mut s = self.stats.write().await; - *s = BackendStats::default(); - } - - /// Return a reference to the inner backend. - pub fn inner(&self) -> &B { - &self.inner - } -} - -#[async_trait] -impl<B: StorageBackend> StorageBackend for MetricsBackend<B> { - async fn get(&self, key: &[u8]) -> Result<Option<Vec<u8>>, StorageError> { - let start = Instant::now(); - let result = self.inner.get(key).await; - let elapsed_ms = start.elapsed().as_secs_f64() * 1000.0; - - let mut s = self.stats.write().await; - s.get_count += 1; - s.get_latency_sum_ms += elapsed_ms; - if let Ok(Some(ref val)) = result { - s.total_bytes_read += val.len() as u64; - } - - result - } - - async fn put(&self, key: &[u8], value: &[u8]) -> Result<(), StorageError> { - let start = Instant::now(); - let result = self.inner.put(key, value).await; - let elapsed_ms = start.elapsed().as_secs_f64() * 1000.0; - - let mut s = self.stats.write().await; - s.put_count += 1; - s.put_latency_sum_ms += elapsed_ms; - if result.is_ok() { - s.total_bytes_written += value.len() as u64; - } - - result - } - - async fn delete(&self, key: &[u8]) -> Result<bool, StorageError> { - let mut s = self.stats.write().await; - s.delete_count += 1; - drop(s); // Release lock before the potentially slow operation. - self.inner.delete(key).await - } - - async fn exists(&self, key: &[u8]) -> Result<bool, StorageError> { - self.inner.exists(key).await - } - - async fn scan_prefix( - &self, - prefix: &[u8], - limit: usize, - ) -> Result<Vec<(Vec<u8>, Vec<u8>)>, StorageError> { - let result = self.inner.scan_prefix(prefix, limit).await; - - let mut s = self.stats.write().await; - s.scan_count += 1; - if let Ok(ref entries) = result { - let bytes: u64 = entries - .iter() - .map(|(k, v)| (k.len() + v.len()) as u64) - .sum(); - s.total_bytes_read += bytes; - } - - result - } - - async fn multi_get(&self, keys: &[&[u8]]) -> Result<Vec<Option<Vec<u8>>>, StorageError> { - let start = Instant::now(); - let result = self.inner.multi_get(keys).await; - let elapsed_ms = start.elapsed().as_secs_f64() * 1000.0; - - let mut s = self.stats.write().await; - s.get_count += keys.len() as u64; - s.get_latency_sum_ms += elapsed_ms; - if let Ok(ref vals) = result { - let bytes: u64 = vals - .iter() - .filter_map(|v| v.as_ref()) - .map(|v| v.len() as u64) - .sum(); - s.total_bytes_read += bytes; - } - - result - } - - async fn batch_put(&self, entries: &[(&[u8], &[u8])]) -> Result<(), StorageError> { - let start = Instant::now(); - let result = self.inner.batch_put(entries).await; - let elapsed_ms = start.elapsed().as_secs_f64() * 1000.0; - - let mut s = self.stats.write().await; - s.put_count += entries.len() as u64; - s.put_latency_sum_ms += elapsed_ms; - if result.is_ok() { - let bytes: u64 = entries.iter().map(|(_, v)| v.len() as u64).sum(); - s.total_bytes_written += bytes; - } - - result - } - - async fn flush(&self) -> Result<(), StorageError> { - self.inner.flush().await - } - - fn name(&self) -> &str { - self.inner.name() - } - - async fn approximate_size(&self) -> Result<Option<u64>, StorageError> { - self.inner.approximate_size().await - } -} - -#[cfg(test)] -mod tests { - use super::*; - use crate::memory::InMemoryBackend; - - #[tokio::test] - async fn test_get_increments_count() { - let inner = InMemoryBackend::new(); - let metered = MetricsBackend::new(inner); - - metered.put(b"k", b"v").await.expect("TODO: handle error"); - metered.get(b"k").await.expect("TODO: handle error"); - metered.get(b"k").await.expect("TODO: handle error"); - metered.get(b"missing").await.expect("TODO: handle error"); - - let stats = metered.stats().await; - assert_eq!(stats.get_count, 3); - assert_eq!(stats.put_count, 1); - } - - #[tokio::test] - async fn test_put_increments_count_and_bytes() { - let inner = InMemoryBackend::new(); - let metered = MetricsBackend::new(inner); - - metered.put(b"a", b"hello").await.expect("TODO: handle error"); // 5 bytes - metered.put(b"b", b"world!").await.expect("TODO: handle error"); // 6 bytes - - let stats = metered.stats().await; - assert_eq!(stats.put_count, 2); - assert_eq!(stats.total_bytes_written, 11); - } - - #[tokio::test] - async fn test_delete_increments_count() { - let inner = InMemoryBackend::new(); - let metered = MetricsBackend::new(inner); - - metered.put(b"k", b"v").await.expect("TODO: handle error"); - metered.delete(b"k").await.expect("TODO: handle error"); - metered.delete(b"nope").await.expect("TODO: handle error"); - - let stats = metered.stats().await; - assert_eq!(stats.delete_count, 2); - } - - #[tokio::test] - async fn test_scan_increments_count_and_bytes() { - let inner = InMemoryBackend::new(); - let metered = MetricsBackend::new(inner); - - metered.put(b"pfx:a", b"11").await.expect("TODO: handle error"); // key=5 + val=2 = 7 - metered.put(b"pfx:b", b"22").await.expect("TODO: handle error"); // key=5 + val=2 = 7 - metered.put(b"other", b"xx").await.expect("TODO: handle error"); - - let results = metered.scan_prefix(b"pfx:", 10).await.expect("TODO: handle error"); - assert_eq!(results.len(), 2); - - let stats = metered.stats().await; - assert_eq!(stats.scan_count, 1); - // Bytes read from scan: 2 * (5 + 2) = 14. - assert_eq!(stats.total_bytes_read, 14); - } - - #[tokio::test] - async fn test_multi_get_increments_count() { - let inner = InMemoryBackend::new(); - let metered = MetricsBackend::new(inner); - - metered.put(b"a", b"1").await.expect("TODO: handle error"); - metered.put(b"b", b"22").await.expect("TODO: handle error"); - - metered - .multi_get(&[b"a" as &[u8], b"b", b"missing"]) - .await - .expect("TODO: handle error"); - - let stats = metered.stats().await; - // multi_get adds keys.len() to get_count. - assert_eq!(stats.get_count, 3); - // Bytes read: 1 + 2 = 3 (missing key contributes 0). - assert_eq!(stats.total_bytes_read, 3); - } - - #[tokio::test] - async fn test_batch_put_increments_count_and_bytes() { - let inner = InMemoryBackend::new(); - let metered = MetricsBackend::new(inner); - - metered - .batch_put(&[ - (b"a" as &[u8], b"111" as &[u8]), - (b"b", b"2222"), - ]) - .await - .expect("TODO: handle error"); - - let stats = metered.stats().await; - assert_eq!(stats.put_count, 2); - // Bytes written: 3 + 4 = 7. - assert_eq!(stats.total_bytes_written, 7); - } - - #[tokio::test] - async fn test_latency_is_recorded() { - let inner = InMemoryBackend::new(); - let metered = MetricsBackend::new(inner); - - metered.put(b"k", b"v").await.expect("TODO: handle error"); - metered.get(b"k").await.expect("TODO: handle error"); - - let stats = metered.stats().await; - // Latency should be non-negative (it might be very small). - assert!(stats.get_latency_sum_ms >= 0.0); - assert!(stats.put_latency_sum_ms >= 0.0); - } - - #[tokio::test] - async fn test_reset_stats() { - let inner = InMemoryBackend::new(); - let metered = MetricsBackend::new(inner); - - metered.put(b"a", b"1").await.expect("TODO: handle error"); - metered.get(b"a").await.expect("TODO: handle error"); - - let before = metered.stats().await; - assert_eq!(before.get_count, 1); - assert_eq!(before.put_count, 1); - - metered.reset_stats().await; - - let after = metered.stats().await; - assert_eq!(after.get_count, 0); - assert_eq!(after.put_count, 0); - assert_eq!(after.total_bytes_read, 0); - assert_eq!(after.total_bytes_written, 0); - } - - #[tokio::test] - async fn test_name_delegates_to_inner() { - let inner = InMemoryBackend::new(); - let metered = MetricsBackend::new(inner); - assert_eq!(metered.name(), "in-memory"); - } - - #[tokio::test] - async fn test_flush_delegates_to_inner() { - let inner = InMemoryBackend::new(); - let metered = MetricsBackend::new(inner); - metered.put(b"k", b"v").await.expect("TODO: handle error"); - metered.flush().await.expect("TODO: handle error"); - // Data should survive flush. - assert_eq!( - metered.get(b"k").await.expect("TODO: handle error"), - Some(b"v".to_vec()) - ); - } - - #[tokio::test] - async fn test_approximate_size_delegates() { - let inner = InMemoryBackend::new(); - let metered = MetricsBackend::new(inner); - metered.put(b"abc", b"defgh").await.expect("TODO: handle error"); - let size = metered.approximate_size().await.expect("TODO: handle error"); - assert_eq!(size, Some(8)); // 3 + 5 - } -} diff --git a/verisimdb/rust-core/verisim-storage/src/redb_backend.rs b/verisimdb/rust-core/verisim-storage/src/redb_backend.rs deleted file mode 100644 index 28127ad5..00000000 --- a/verisimdb/rust-core/verisim-storage/src/redb_backend.rs +++ /dev/null @@ -1,513 +0,0 @@ -// SPDX-License-Identifier: MPL-2.0 -// Copyright (c) 2026 Jonathan D.A. Jewell (hyperpolymath) <j.d.a.jewell@open.ac.uk> -// -// redb-backed persistent storage backend for VeriSimDB. -// -// Uses redb (pure Rust, B-tree, ACID, single-file database) to provide -// durable key-value storage. No C/C++ dependencies — builds on any platform -// with a Rust toolchain. -// -// # Design -// -// - Single redb `Database` file containing one main table. -// - Read transactions for all read operations (concurrent, lock-free). -// - Write transactions for put/delete/batch (serialised by redb internally). -// - `flush()` maps to `compact()` which reclaims free space. -// - `scan_prefix` uses redb's `range()` with a computed upper bound to -// efficiently iterate keys sharing a common prefix. - -use std::path::{Path, PathBuf}; -use std::sync::Arc; - -use async_trait::async_trait; -use redb::{Database, ReadableDatabase, TableDefinition}; -use tracing::debug; - -use crate::backend::StorageBackend; -use crate::error::StorageError; - -/// Table definition for the main key-value store. -/// -/// Keys and values are byte slices, matching the `StorageBackend` trait's -/// opaque byte interface. -const MAIN_TABLE: TableDefinition<&[u8], &[u8]> = TableDefinition::new("main"); - -/// A persistent storage backend powered by redb. -/// -/// redb is a pure-Rust embedded database with ACID transactions, copy-on-write -/// B-tree storage, and zero external dependencies. Each `RedbBackend` wraps a -/// single database file. -/// -/// Thread-safe: `Database` is `Send + Sync` and handles internal locking. -/// -/// # Example -/// -/// ```rust,no_run -/// use verisim_storage::redb_backend::RedbBackend; -/// use verisim_storage::backend::StorageBackend; -/// -/// # tokio_test::block_on(async { -/// let store = RedbBackend::open("/tmp/verisim-test.redb").unwrap(); -/// store.put(b"hello", b"world").await.unwrap(); -/// let val = store.get(b"hello").await.unwrap(); -/// assert_eq!(val, Some(b"world".to_vec())); -/// # }); -/// ``` -pub struct RedbBackend { - /// The redb database handle. - db: Arc<Database>, - /// Path to the database file (for diagnostics and approximate_size). - path: PathBuf, -} - -impl RedbBackend { - /// Open or create a redb database at the given path. - /// - /// Creates the file and parent directories if they don't exist. The main - /// table is created on first write. - pub fn open(path: impl AsRef<Path>) -> Result<Self, StorageError> { - let path = path.as_ref().to_path_buf(); - - // Ensure parent directory exists - if let Some(parent) = path.parent() { - std::fs::create_dir_all(parent).map_err(StorageError::Io)?; - } - - let db = Database::create(&path).map_err(|e| { - StorageError::BackendUnavailable(format!("failed to open redb at {}: {}", path.display(), e)) - })?; - - debug!(path = %path.display(), "opened redb backend"); - - Ok(Self { - db: Arc::new(db), - path, - }) - } - - /// Return the filesystem path of the database file. - pub fn path(&self) -> &Path { - &self.path - } - - /// Compute the upper bound for a prefix scan. - /// - /// Given a prefix like `[0x61, 0x62]` ("ab"), returns the next key - /// after all keys starting with that prefix: `[0x61, 0x63]` ("ac"). - /// Returns `None` if the prefix is all 0xFF bytes (no upper bound). - #[cfg(test)] - fn prefix_upper_bound(prefix: &[u8]) -> Option<Vec<u8>> { - let mut upper = prefix.to_vec(); - // Increment the last non-0xFF byte - while let Some(last) = upper.last_mut() { - if *last < 0xFF { - *last += 1; - return Some(upper); - } - upper.pop(); - } - None // All bytes were 0xFF — no upper bound - } -} - -impl std::fmt::Debug for RedbBackend { - fn fmt(&self, f: &mut std::fmt::Formatter<'_>) -> std::fmt::Result { - f.debug_struct("RedbBackend") - .field("path", &self.path) - .finish() - } -} - -#[async_trait] -impl StorageBackend for RedbBackend { - async fn get(&self, key: &[u8]) -> Result<Option<Vec<u8>>, StorageError> { - let db = Arc::clone(&self.db); - let key = key.to_vec(); - - tokio::task::spawn_blocking(move || -> Result<Option<Vec<u8>>, StorageError> { - let txn = db.begin_read().map_err(|e| { - StorageError::BackendUnavailable(format!("read txn: {e}")) - })?; - - let table = match txn.open_table(MAIN_TABLE) { - Ok(t) => t, - // Table doesn't exist yet — no data has been written - Err(_) => return Ok(None), - }; - - match table.get(key.as_slice()) { - Ok(Some(value)) => Ok(Some(value.value().to_vec())), - Ok(None) => Ok(None), - Err(e) => Err(StorageError::CorruptedData(format!("get: {e}"))), - } - }) - .await - .map_err(|e| StorageError::BackendUnavailable(format!("task join: {e}")))? - } - - async fn put(&self, key: &[u8], value: &[u8]) -> Result<(), StorageError> { - let db = Arc::clone(&self.db); - let key = key.to_vec(); - let value = value.to_vec(); - - tokio::task::spawn_blocking(move || -> Result<(), StorageError> { - let txn = db.begin_write().map_err(|e| { - StorageError::BackendUnavailable(format!("write txn: {e}")) - })?; - { - let mut table = txn.open_table(MAIN_TABLE).map_err(|e| { - StorageError::BackendUnavailable(format!("open table: {e}")) - })?; - table.insert(key.as_slice(), value.as_slice()).map_err(|e| { - StorageError::CorruptedData(format!("insert: {e}")) - })?; - } - txn.commit().map_err(|e| { - StorageError::CorruptedData(format!("commit: {e}")) - })?; - Ok(()) - }) - .await - .map_err(|e| StorageError::BackendUnavailable(format!("task join: {e}")))? - } - - async fn delete(&self, key: &[u8]) -> Result<bool, StorageError> { - let db = Arc::clone(&self.db); - let key = key.to_vec(); - - tokio::task::spawn_blocking(move || -> Result<bool, StorageError> { - let txn = db.begin_write().map_err(|e| { - StorageError::BackendUnavailable(format!("write txn: {e}")) - })?; - let existed; - { - let mut table = txn.open_table(MAIN_TABLE).map_err(|e| { - StorageError::BackendUnavailable(format!("open table: {e}")) - })?; - existed = table.remove(key.as_slice()).map_err(|e| { - StorageError::CorruptedData(format!("remove: {e}")) - })?.is_some(); - } - txn.commit().map_err(|e| { - StorageError::CorruptedData(format!("commit: {e}")) - })?; - Ok(existed) - }) - .await - .map_err(|e| StorageError::BackendUnavailable(format!("task join: {e}")))? - } - - async fn exists(&self, key: &[u8]) -> Result<bool, StorageError> { - // Delegate to get — redb has no separate "exists" check. - Ok(self.get(key).await?.is_some()) - } - - async fn scan_prefix( - &self, - prefix: &[u8], - limit: usize, - ) -> Result<Vec<(Vec<u8>, Vec<u8>)>, StorageError> { - let db = Arc::clone(&self.db); - let prefix = prefix.to_vec(); - - tokio::task::spawn_blocking(move || -> Result<Vec<(Vec<u8>, Vec<u8>)>, StorageError> { - let txn = db.begin_read().map_err(|e| { - StorageError::BackendUnavailable(format!("read txn: {e}")) - })?; - let table = match txn.open_table(MAIN_TABLE) { - Ok(t) => t, - Err(_) => return Ok(Vec::new()), // Table doesn't exist yet - }; - - let mut results = Vec::new(); - - // Scan from the prefix key onward; stop when keys no longer match - let iter = table.range(prefix.as_slice()..).map_err(|e| { - StorageError::CorruptedData(format!("range scan: {e}")) - })?; - - for entry in iter { - let entry = entry.map_err(|e| { - StorageError::CorruptedData(format!("scan entry: {e}")) - })?; - let k = entry.0.value().to_vec(); - let v = entry.1.value().to_vec(); - - if !k.starts_with(&prefix) { - break; - } - - results.push((k, v)); - if results.len() >= limit { - break; - } - } - - Ok(results) - }) - .await - .map_err(|e| StorageError::BackendUnavailable(format!("task join: {e}")))? - } - - async fn multi_get(&self, keys: &[&[u8]]) -> Result<Vec<Option<Vec<u8>>>, StorageError> { - let db = Arc::clone(&self.db); - let owned_keys: Vec<Vec<u8>> = keys.iter().map(|k| k.to_vec()).collect(); - - tokio::task::spawn_blocking(move || -> Result<Vec<Option<Vec<u8>>>, StorageError> { - let txn = db.begin_read().map_err(|e| { - StorageError::BackendUnavailable(format!("read txn: {e}")) - })?; - let table = match txn.open_table(MAIN_TABLE) { - Ok(t) => t, - Err(_) => return Ok(owned_keys.iter().map(|_| None).collect()), - }; - - let mut results = Vec::with_capacity(owned_keys.len()); - for key in &owned_keys { - match table.get(key.as_slice()) { - Ok(Some(v)) => results.push(Some(v.value().to_vec())), - Ok(None) => results.push(None), - Err(e) => { - return Err(StorageError::CorruptedData(format!("multi_get: {e}"))) - } - } - } - Ok(results) - }) - .await - .map_err(|e| StorageError::BackendUnavailable(format!("task join: {e}")))? - } - - async fn batch_put(&self, entries: &[(&[u8], &[u8])]) -> Result<(), StorageError> { - let db = Arc::clone(&self.db); - let owned: Vec<(Vec<u8>, Vec<u8>)> = entries - .iter() - .map(|(k, v)| (k.to_vec(), v.to_vec())) - .collect(); - - tokio::task::spawn_blocking(move || -> Result<(), StorageError> { - let txn = db.begin_write().map_err(|e| { - StorageError::BackendUnavailable(format!("write txn: {e}")) - })?; - { - let mut table = txn.open_table(MAIN_TABLE).map_err(|e| { - StorageError::BackendUnavailable(format!("open table: {e}")) - })?; - for (k, v) in &owned { - table.insert(k.as_slice(), v.as_slice()).map_err(|e| { - StorageError::CorruptedData(format!("batch insert: {e}")) - })?; - } - } - txn.commit().map_err(|e| { - StorageError::CorruptedData(format!("batch commit: {e}")) - })?; - Ok(()) - }) - .await - .map_err(|e| StorageError::BackendUnavailable(format!("task join: {e}")))? - } - - async fn flush(&self) -> Result<(), StorageError> { - // redb commits are durable by default — each write transaction is - // fsynced on commit. No additional flush needed. compact() requires - // &mut self and is a space-reclamation optimisation, not a durability - // operation. - Ok(()) - } - - fn name(&self) -> &str { - "redb" - } - - async fn approximate_size(&self) -> Result<Option<u64>, StorageError> { - match std::fs::metadata(&self.path) { - Ok(meta) => Ok(Some(meta.len())), - Err(_) => Ok(None), - } - } -} - -#[cfg(test)] -mod tests { - use super::*; - use tempfile::tempdir; - - /// Create a temporary RedbBackend for testing. - /// - /// Uses `tempdir()` rather than `NamedTempFile` so the directory persists - /// for the lifetime of the test (NamedTempFile's Drop would unlink the - /// path while redb still holds it open, breaking `approximate_size`). - fn temp_backend() -> (RedbBackend, tempfile::TempDir) { - let dir = tempdir().expect("TODO: handle error"); - let path = dir.path().join("test.redb"); - let backend = RedbBackend::open(&path).expect("TODO: handle error"); - (backend, dir) - } - - #[tokio::test] - async fn test_basic_crud() { - let (backend, _dir) = temp_backend(); - - // Get on empty store returns None - assert_eq!(backend.get(b"key1").await.expect("TODO: handle error"), None); - assert!(!backend.exists(b"key1").await.expect("TODO: handle error")); - - // Put and get - backend.put(b"key1", b"value1").await.expect("TODO: handle error"); - assert_eq!(backend.get(b"key1").await.expect("TODO: handle error"), Some(b"value1".to_vec())); - assert!(backend.exists(b"key1").await.expect("TODO: handle error")); - - // Overwrite - backend.put(b"key1", b"updated").await.expect("TODO: handle error"); - assert_eq!(backend.get(b"key1").await.expect("TODO: handle error"), Some(b"updated".to_vec())); - - // Delete existing key - assert!(backend.delete(b"key1").await.expect("TODO: handle error")); - assert_eq!(backend.get(b"key1").await.expect("TODO: handle error"), None); - - // Delete non-existent key - assert!(!backend.delete(b"nonexistent").await.expect("TODO: handle error")); - } - - #[tokio::test] - async fn test_scan_prefix() { - let (backend, _dir) = temp_backend(); - - backend.put(b"user:1:name", b"Alice").await.expect("TODO: handle error"); - backend.put(b"user:1:age", b"30").await.expect("TODO: handle error"); - backend.put(b"user:2:name", b"Bob").await.expect("TODO: handle error"); - backend.put(b"post:1:title", b"Hello").await.expect("TODO: handle error"); - - // Scan "user:1:" prefix - let results = backend.scan_prefix(b"user:1:", 10).await.expect("TODO: handle error"); - assert_eq!(results.len(), 2); - assert_eq!(results[0].0, b"user:1:age".to_vec()); - assert_eq!(results[1].0, b"user:1:name".to_vec()); - - // Scan "user:" prefix — all user keys - let results = backend.scan_prefix(b"user:", 10).await.expect("TODO: handle error"); - assert_eq!(results.len(), 3); - - // Scan with limit - let results = backend.scan_prefix(b"user:", 2).await.expect("TODO: handle error"); - assert_eq!(results.len(), 2); - - // Scan with no matching prefix - let results = backend.scan_prefix(b"missing:", 10).await.expect("TODO: handle error"); - assert!(results.is_empty()); - } - - #[tokio::test] - async fn test_multi_get() { - let (backend, _dir) = temp_backend(); - - backend.put(b"a", b"1").await.expect("TODO: handle error"); - backend.put(b"b", b"2").await.expect("TODO: handle error"); - backend.put(b"c", b"3").await.expect("TODO: handle error"); - - let results = backend - .multi_get(&[b"a" as &[u8], b"missing", b"c"]) - .await - .expect("TODO: handle error"); - - assert_eq!(results.len(), 3); - assert_eq!(results[0], Some(b"1".to_vec())); - assert_eq!(results[1], None); - assert_eq!(results[2], Some(b"3".to_vec())); - } - - #[tokio::test] - async fn test_batch_put() { - let (backend, _dir) = temp_backend(); - - backend - .batch_put(&[ - (b"x" as &[u8], b"10" as &[u8]), - (b"y", b"20"), - (b"z", b"30"), - ]) - .await - .expect("TODO: handle error"); - - assert_eq!(backend.get(b"x").await.expect("TODO: handle error"), Some(b"10".to_vec())); - assert_eq!(backend.get(b"y").await.expect("TODO: handle error"), Some(b"20".to_vec())); - assert_eq!(backend.get(b"z").await.expect("TODO: handle error"), Some(b"30".to_vec())); - } - - #[tokio::test] - async fn test_flush_compacts() { - let (backend, _dir) = temp_backend(); - backend.put(b"key", b"val").await.expect("TODO: handle error"); - backend.flush().await.expect("TODO: handle error"); - assert_eq!(backend.get(b"key").await.expect("TODO: handle error"), Some(b"val".to_vec())); - } - - #[tokio::test] - async fn test_name() { - let (backend, _dir) = temp_backend(); - assert_eq!(backend.name(), "redb"); - } - - #[tokio::test] - async fn test_approximate_size() { - let (backend, _dir) = temp_backend(); - // After writing data, file size should be non-zero - backend.put(b"key", b"value").await.expect("TODO: handle error"); - let size = backend.approximate_size().await.expect("TODO: handle error"); - assert!(size.is_some()); - assert!(size.expect("TODO: handle error") > 0); - } - - #[tokio::test] - async fn test_persistence_across_reopen() { - let dir = tempdir().expect("TODO: handle error"); - let path = dir.path().join("persist-test.redb"); - - // Write data and drop - { - let backend = RedbBackend::open(&path).expect("TODO: handle error"); - backend.put(b"persistent-key", b"persistent-value").await.expect("TODO: handle error"); - } - - // Reopen and verify data survived - { - let backend = RedbBackend::open(&path).expect("TODO: handle error"); - let val = backend.get(b"persistent-key").await.expect("TODO: handle error"); - assert_eq!(val, Some(b"persistent-value".to_vec())); - } - } - - #[test] - fn test_prefix_upper_bound() { - // Normal case - assert_eq!( - RedbBackend::prefix_upper_bound(b"abc"), - Some(b"abd".to_vec()) - ); - - // Trailing 0xFF byte - assert_eq!( - RedbBackend::prefix_upper_bound(b"ab\xff"), - Some(b"ac".to_vec()) - ); - - // All 0xFF bytes — no upper bound - assert_eq!( - RedbBackend::prefix_upper_bound(b"\xff\xff"), - None - ); - - // Empty prefix — no upper bound - assert_eq!( - RedbBackend::prefix_upper_bound(b""), - None - ); - - // Single byte - assert_eq!( - RedbBackend::prefix_upper_bound(b"z"), - Some(b"{".to_vec()) // 'z' + 1 = '{' - ); - } -} diff --git a/verisimdb/rust-core/verisim-storage/src/typed.rs b/verisimdb/rust-core/verisim-storage/src/typed.rs deleted file mode 100644 index 2cf30d98..00000000 --- a/verisimdb/rust-core/verisim-storage/src/typed.rs +++ /dev/null @@ -1,291 +0,0 @@ -// SPDX-License-Identifier: MPL-2.0 -// Copyright (c) 2026 Jonathan D.A. Jewell (hyperpolymath) <j.d.a.jewell@open.ac.uk> -// -// Typed storage wrapper for VeriSimDB. -// -// Provides a higher-level, serde-based interface on top of any `StorageBackend`. -// Values are serialized as JSON and all keys are automatically prefixed with a -// configurable namespace, enabling multiple logical stores to share a single -// physical backend without key collisions. - -use serde::de::DeserializeOwned; -use serde::Serialize; - -use crate::backend::StorageBackend; -use crate::error::StorageError; - -/// A typed wrapper around a [`StorageBackend`] that handles serialization -/// and namespace prefixing automatically. -/// -/// Keys are prefixed with `"{namespace}:"` before being passed to the -/// underlying backend. Values are serialized to JSON on write and -/// deserialized on read. -/// -/// # Example -/// -/// ```rust -/// use verisim_storage::memory::InMemoryBackend; -/// use verisim_storage::typed::TypedStore; -/// use serde::{Serialize, Deserialize}; -/// -/// #[derive(Debug, Serialize, Deserialize, PartialEq)] -/// struct User { name: String, age: u32 } -/// -/// # tokio_test::block_on(async { -/// let backend = InMemoryBackend::new(); -/// let store = TypedStore::new(backend, "users"); -/// -/// let alice = User { name: "Alice".into(), age: 30 }; -/// store.put("alice", &alice).await.unwrap(); -/// -/// let retrieved: User = store.get("alice").await.unwrap().unwrap(); -/// assert_eq!(retrieved, alice); -/// # }); -/// ``` -pub struct TypedStore<B: StorageBackend> { - /// The underlying raw key-value backend. - backend: B, - /// Namespace prefix applied to all keys. - namespace: String, -} - -impl<B: StorageBackend> TypedStore<B> { - /// Create a new typed store wrapping `backend` with the given namespace. - /// - /// All keys will be prefixed with `"{namespace}:"`. - pub fn new(backend: B, namespace: &str) -> Self { - Self { - backend, - namespace: namespace.to_string(), - } - } - - /// Return a reference to the underlying backend. - pub fn backend(&self) -> &B { - &self.backend - } - - /// Return the namespace prefix used by this store. - pub fn namespace(&self) -> &str { - &self.namespace - } - - /// Build the full namespaced key from a logical key string. - fn prefixed_key(&self, key: &str) -> Vec<u8> { - format!("{}:{}", self.namespace, key).into_bytes() - } - - /// Build the namespace prefix (for scanning). - fn prefix_bytes(&self) -> Vec<u8> { - format!("{}:", self.namespace).into_bytes() - } - - /// Retrieve and deserialize a value by its logical key. - /// - /// Returns `Ok(None)` if the key does not exist. - pub async fn get<T: DeserializeOwned>(&self, key: &str) -> Result<Option<T>, StorageError> { - let full_key = self.prefixed_key(key); - match self.backend.get(&full_key).await? { - Some(bytes) => { - let value: T = serde_json::from_slice(&bytes).map_err(|err| { - StorageError::SerializationError(format!( - "failed to deserialize value for key '{}': {}", - key, err - )) - })?; - Ok(Some(value)) - } - None => Ok(None), - } - } - - /// Serialize and store a value under the given logical key. - pub async fn put<T: Serialize>(&self, key: &str, value: &T) -> Result<(), StorageError> { - let full_key = self.prefixed_key(key); - let bytes = serde_json::to_vec(value).map_err(|err| { - StorageError::SerializationError(format!( - "failed to serialize value for key '{}': {}", - key, err - )) - })?; - self.backend.put(&full_key, &bytes).await - } - - /// Delete a value by its logical key. - /// - /// Returns `Ok(true)` if the key existed and was removed. - pub async fn delete(&self, key: &str) -> Result<bool, StorageError> { - let full_key = self.prefixed_key(key); - self.backend.delete(&full_key).await - } - - /// Scan all entries in this namespace whose keys (after the namespace - /// prefix) start with the given `key_prefix`, returning up to `limit` - /// deserialized (key-suffix, value) pairs. - pub async fn scan_prefix<T: DeserializeOwned>( - &self, - key_prefix: &str, - limit: usize, - ) -> Result<Vec<(String, T)>, StorageError> { - let full_prefix = format!("{}:{}", self.namespace, key_prefix).into_bytes(); - let ns_prefix = self.prefix_bytes(); - let ns_prefix_len = ns_prefix.len(); - - let raw_results = self.backend.scan_prefix(&full_prefix, limit).await?; - - let mut results = Vec::with_capacity(raw_results.len()); - for (raw_key, raw_value) in raw_results { - // Strip the namespace prefix to recover the logical key. - let logical_key = if raw_key.len() >= ns_prefix_len { - String::from_utf8_lossy(&raw_key[ns_prefix_len..]).to_string() - } else { - String::from_utf8_lossy(&raw_key).to_string() - }; - - let value: T = serde_json::from_slice(&raw_value).map_err(|err| { - StorageError::SerializationError(format!( - "failed to deserialize scanned value for key '{}': {}", - logical_key, err - )) - })?; - - results.push((logical_key, value)); - } - - Ok(results) - } -} - -#[cfg(test)] -mod tests { - use super::*; - use crate::memory::InMemoryBackend; - use serde::{Deserialize, Serialize}; - - #[derive(Debug, Clone, Serialize, Deserialize, PartialEq)] - struct TestRecord { - name: String, - score: f64, - } - - #[tokio::test] - async fn test_typed_round_trip() { - let backend = InMemoryBackend::new(); - let store = TypedStore::new(backend, "test"); - - let record = TestRecord { - name: "Alice".to_string(), - score: 95.5, - }; - - // Put and get. - store.put("rec1", &record).await.expect("TODO: handle error"); - let retrieved: TestRecord = store.get("rec1").await.expect("TODO: handle error").expect("TODO: handle error"); - assert_eq!(retrieved, record); - - // Missing key. - let missing: Option<TestRecord> = store.get("nonexistent").await.expect("TODO: handle error"); - assert!(missing.is_none()); - - // Delete. - assert!(store.delete("rec1").await.expect("TODO: handle error")); - assert!(store.get::<TestRecord>("rec1").await.expect("TODO: handle error").is_none()); - - // Delete non-existent. - assert!(!store.delete("rec1").await.expect("TODO: handle error")); - } - - #[tokio::test] - async fn test_namespace_isolation() { - let backend = InMemoryBackend::new(); - let store_a = TypedStore::new(backend.clone(), "ns_a"); - let store_b = TypedStore::new(backend.clone(), "ns_b"); - - store_a.put("key", &"value_a".to_string()).await.expect("TODO: handle error"); - store_b.put("key", &"value_b".to_string()).await.expect("TODO: handle error"); - - // Each namespace sees its own value. - let val_a: String = store_a.get("key").await.expect("TODO: handle error").expect("TODO: handle error"); - let val_b: String = store_b.get("key").await.expect("TODO: handle error").expect("TODO: handle error"); - assert_eq!(val_a, "value_a"); - assert_eq!(val_b, "value_b"); - - // Deleting from one namespace does not affect the other. - store_a.delete("key").await.expect("TODO: handle error"); - assert!(store_a.get::<String>("key").await.expect("TODO: handle error").is_none()); - assert_eq!( - store_b.get::<String>("key").await.expect("TODO: handle error").expect("TODO: handle error"), - "value_b" - ); - } - - #[tokio::test] - async fn test_typed_scan_prefix() { - let backend = InMemoryBackend::new(); - let store = TypedStore::new(backend, "items"); - - store.put("fruit:apple", &10u32).await.expect("TODO: handle error"); - store.put("fruit:banana", &20u32).await.expect("TODO: handle error"); - store.put("vegetable:carrot", &30u32).await.expect("TODO: handle error"); - - let fruits: Vec<(String, u32)> = store.scan_prefix("fruit:", 10).await.expect("TODO: handle error"); - assert_eq!(fruits.len(), 2); - assert_eq!(fruits[0].0, "fruit:apple"); - assert_eq!(fruits[0].1, 10); - assert_eq!(fruits[1].0, "fruit:banana"); - assert_eq!(fruits[1].1, 20); - - // Scan with limit. - let limited: Vec<(String, u32)> = store.scan_prefix("fruit:", 1).await.expect("TODO: handle error"); - assert_eq!(limited.len(), 1); - } - - #[tokio::test] - async fn test_typed_primitive_types() { - let backend = InMemoryBackend::new(); - let store = TypedStore::new(backend, "prims"); - - // Integer. - store.put("int", &42i64).await.expect("TODO: handle error"); - assert_eq!(store.get::<i64>("int").await.expect("TODO: handle error").expect("TODO: handle error"), 42); - - // Boolean. - store.put("flag", &true).await.expect("TODO: handle error"); - assert_eq!(store.get::<bool>("flag").await.expect("TODO: handle error").expect("TODO: handle error"), true); - - // Vec. - store.put("list", &vec![1, 2, 3]).await.expect("TODO: handle error"); - assert_eq!( - store.get::<Vec<i32>>("list").await.expect("TODO: handle error").expect("TODO: handle error"), - vec![1, 2, 3] - ); - } - - #[tokio::test] - async fn test_deserialization_error() { - let backend = InMemoryBackend::new(); - let store = TypedStore::new(backend.clone(), "bad"); - - // Write raw invalid JSON bytes directly via the backend. - let key = b"bad:broken"; - backend.put(key, b"not-valid-json!!!").await.expect("TODO: handle error"); - - // Attempt to deserialize as a struct should fail. - let result = store.get::<TestRecord>("broken").await; - assert!(result.is_err()); - match result.unwrap_err() { - StorageError::SerializationError(msg) => { - assert!(msg.contains("failed to deserialize")); - } - other => panic!("expected SerializationError, got: {:?}", other), - } - } - - #[tokio::test] - async fn test_namespace_and_backend_accessors() { - let backend = InMemoryBackend::new(); - let store = TypedStore::new(backend, "myns"); - assert_eq!(store.namespace(), "myns"); - assert_eq!(store.backend().name(), "in-memory"); - } -} diff --git a/verisimdb/rust-core/verisim-temporal/Cargo.toml b/verisimdb/rust-core/verisim-temporal/Cargo.toml deleted file mode 100644 index a1e6c720..00000000 --- a/verisimdb/rust-core/verisim-temporal/Cargo.toml +++ /dev/null @@ -1,27 +0,0 @@ -# SPDX-License-Identifier: MPL-2.0 - -[package] -name = "verisim-temporal" -description = "Temporal modality - time-series and versioning" -version.workspace = true -edition.workspace = true -authors.workspace = true -license.workspace = true - -[dependencies] -chrono.workspace = true -serde.workspace = true -thiserror.workspace = true -tracing.workspace = true -async-trait.workspace = true -serde_json.workspace = true -verisim-storage = { path = "../verisim-storage", optional = true } -tokio.workspace = true - -[dev-dependencies] -tempfile = "3" -proptest.workspace = true - -[features] -default = [] -redb-backend = ["verisim-storage/redb-backend"] diff --git a/verisimdb/rust-core/verisim-temporal/src/diff.rs b/verisimdb/rust-core/verisim-temporal/src/diff.rs deleted file mode 100644 index 91e367fc..00000000 --- a/verisimdb/rust-core/verisim-temporal/src/diff.rs +++ /dev/null @@ -1,197 +0,0 @@ -// SPDX-License-Identifier: MPL-2.0 -//! Diff functionality for comparing versions - -use serde::{Deserialize, Serialize}; -use std::fmt; - -/// Represents a difference between two versions -#[derive(Debug, Clone, Serialize, Deserialize, PartialEq)] -pub enum Diff<T> { - /// No changes between versions - NoChange { - value: T, - }, - /// Value changed from old to new - Changed { - old: T, - new: T, - }, - /// Value was added (didn't exist before) - Added { - value: T, - }, - /// Value was removed (existed before, doesn't now) - Removed { - value: T, - }, -} - -impl<T> Diff<T> { - /// Create a diff for no change - pub fn no_change(value: T) -> Self { - Diff::NoChange { value } - } - - /// Create a diff for changed value - pub fn changed(old: T, new: T) -> Self { - Diff::Changed { old, new } - } - - /// Create a diff for added value - pub fn added(value: T) -> Self { - Diff::Added { value } - } - - /// Create a diff for removed value - pub fn removed(value: T) -> Self { - Diff::Removed { value } - } - - /// Check if there's a change - pub fn has_change(&self) -> bool { - !matches!(self, Diff::NoChange { .. }) - } - - /// Get the new value if it exists - pub fn new_value(&self) -> Option<&T> { - match self { - Diff::NoChange { value } => Some(value), - Diff::Changed { new, .. } => Some(new), - Diff::Added { value } => Some(value), - Diff::Removed { .. } => None, - } - } - - /// Get the old value if it exists - pub fn old_value(&self) -> Option<&T> { - match self { - Diff::NoChange { value } => Some(value), - Diff::Changed { old, .. } => Some(old), - Diff::Added { .. } => None, - Diff::Removed { value } => Some(value), - } - } -} - -impl<T: fmt::Display> fmt::Display for Diff<T> { - fn fmt(&self, f: &mut fmt::Formatter<'_>) -> fmt::Result { - match self { - Diff::NoChange { value } => write!(f, "= {}", value), - Diff::Changed { old, new } => write!(f, "- {}\n+ {}", old, new), - Diff::Added { value } => write!(f, "+ {}", value), - Diff::Removed { value } => write!(f, "- {}", value), - } - } -} - -/// Error type for diff comparison failures -#[derive(Debug, Clone, PartialEq, Eq)] -pub enum DiffError { - /// Both old and new values are absent — nothing to compare. - IncomparableValues, -} - -impl fmt::Display for DiffError { - fn fmt(&self, f: &mut fmt::Formatter<'_>) -> fmt::Result { - match self { - DiffError::IncomparableValues => write!(f, "Cannot compare two None values"), - } - } -} - -impl std::error::Error for DiffError {} - -/// Compare two optional values and produce a diff. -/// -/// Returns `Err(DiffError::IncomparableValues)` when both values are `None`. -pub fn compare_values<T: Clone + PartialEq>( - old: Option<&T>, - new: Option<&T>, -) -> Result<Diff<T>, DiffError> { - match (old, new) { - (Some(old_val), Some(new_val)) => { - if old_val == new_val { - Ok(Diff::no_change(old_val.clone())) - } else { - Ok(Diff::changed(old_val.clone(), new_val.clone())) - } - } - (Some(old_val), None) => Ok(Diff::removed(old_val.clone())), - (None, Some(new_val)) => Ok(Diff::added(new_val.clone())), - (None, None) => Err(DiffError::IncomparableValues), - } -} - -#[cfg(test)] -mod tests { - use super::*; - - #[test] - fn test_diff_no_change() { - let diff = Diff::no_change("value"); - assert!(!diff.has_change()); - assert_eq!(diff.new_value(), Some(&"value")); - assert_eq!(diff.old_value(), Some(&"value")); - } - - #[test] - fn test_diff_changed() { - let diff = Diff::changed("old", "new"); - assert!(diff.has_change()); - assert_eq!(diff.new_value(), Some(&"new")); - assert_eq!(diff.old_value(), Some(&"old")); - } - - #[test] - fn test_diff_added() { - let diff: Diff<&str> = Diff::added("new"); - assert!(diff.has_change()); - assert_eq!(diff.new_value(), Some(&"new")); - assert_eq!(diff.old_value(), None); - } - - #[test] - fn test_diff_removed() { - let diff: Diff<&str> = Diff::removed("old"); - assert!(diff.has_change()); - assert_eq!(diff.new_value(), None); - assert_eq!(diff.old_value(), Some(&"old")); - } - - #[test] - fn test_compare_values_same() { - let old_val = "value".to_string(); - let new_val = "value".to_string(); - let diff = compare_values(Some(&old_val), Some(&new_val)).expect("TODO: handle error"); - assert!(!diff.has_change()); - } - - #[test] - fn test_compare_values_changed() { - let old_val = "old".to_string(); - let new_val = "new".to_string(); - let diff = compare_values(Some(&old_val), Some(&new_val)).expect("TODO: handle error"); - assert!(diff.has_change()); - assert_eq!(diff, Diff::Changed { old: "old".to_string(), new: "new".to_string() }); - } - - #[test] - fn test_compare_values_added() { - let new_val = "new".to_string(); - let diff = compare_values(None, Some(&new_val)).expect("TODO: handle error"); - assert_eq!(diff, Diff::Added { value: "new".to_string() }); - } - - #[test] - fn test_compare_values_removed() { - let old_val = "old".to_string(); - let diff = compare_values(Some(&old_val), None).expect("TODO: handle error"); - assert_eq!(diff, Diff::Removed { value: "old".to_string() }); - } - - #[test] - fn test_compare_values_both_none() { - let result = compare_values::<String>(None, None); - assert_eq!(result, Err(DiffError::IncomparableValues)); - } -} diff --git a/verisimdb/rust-core/verisim-temporal/src/lib.rs b/verisimdb/rust-core/verisim-temporal/src/lib.rs deleted file mode 100644 index d01d2eb4..00000000 --- a/verisimdb/rust-core/verisim-temporal/src/lib.rs +++ /dev/null @@ -1,390 +0,0 @@ -// SPDX-License-Identifier: MPL-2.0 -//! VeriSim Temporal Modality -//! -//! Time-series and versioning for audit-grade history. -//! Implements Marr's Computational Level: "What happened when?" - -#![forbid(unsafe_code)] -#[cfg(feature = "redb-backend")] -pub mod persistent; -#[cfg(feature = "redb-backend")] -pub use persistent::*; -pub mod diff; - -use async_trait::async_trait; -use chrono::{DateTime, Utc}; -use serde::{Deserialize, Serialize}; -use std::collections::{BTreeMap, HashMap}; -use std::sync::{Arc, RwLock}; -use thiserror::Error; - -/// Temporal modality errors -#[derive(Error, Debug)] -pub enum TemporalError { - #[error("Entity not found: {0}")] - NotFound(String), - - #[error("Version not found: {entity_id} @ {version}")] - VersionNotFound { entity_id: String, version: u64 }, - - #[error("Invalid time range: {0}")] - InvalidTimeRange(String), - - #[error("Conflict: {0}")] - Conflict(String), - - #[error("Lock poisoned: internal concurrency error")] - LockPoisoned, -} - -/// A timestamped version of an entity -#[derive(Debug, Clone, Serialize, Deserialize)] -pub struct Version<T> { - /// Version number (monotonically increasing) - pub version: u64, - /// When this version was created - pub timestamp: DateTime<Utc>, - /// The data at this version - pub data: T, - /// Who/what created this version - pub author: String, - /// Optional commit message - pub message: Option<String>, -} - -impl<T> Version<T> { - /// Create a new version - pub fn new(version: u64, data: T, author: impl Into<String>) -> Self { - Self { - version, - timestamp: Utc::now(), - data, - author: author.into(), - message: None, - } - } - - /// Add a message - pub fn with_message(mut self, message: impl Into<String>) -> Self { - self.message = Some(message.into()); - self - } -} - -/// A time-series data point -#[derive(Debug, Clone, Serialize, Deserialize)] -pub struct TimePoint<T> { - /// Timestamp - pub time: DateTime<Utc>, - /// Value at this time - pub value: T, - /// Optional labels/tags - pub labels: HashMap<String, String>, -} - -impl<T> TimePoint<T> { - /// Create a new time point - pub fn new(time: DateTime<Utc>, value: T) -> Self { - Self { - time, - value, - labels: HashMap::new(), - } - } - - /// Create with current time - pub fn now(value: T) -> Self { - Self::new(Utc::now(), value) - } - - /// Add a label - pub fn with_label(mut self, key: impl Into<String>, value: impl Into<String>) -> Self { - self.labels.insert(key.into(), value.into()); - self - } -} - -/// Time range for queries -#[derive(Debug, Clone, Serialize, Deserialize)] -pub struct TimeRange { - /// Start time (inclusive) - pub start: DateTime<Utc>, - /// End time (exclusive) - pub end: DateTime<Utc>, -} - -impl TimeRange { - /// Create a time range - pub fn new(start: DateTime<Utc>, end: DateTime<Utc>) -> Result<Self, TemporalError> { - if start >= end { - return Err(TemporalError::InvalidTimeRange( - "start must be before end".to_string(), - )); - } - Ok(Self { start, end }) - } - - /// Last N duration from now - pub fn last(duration: chrono::Duration) -> Self { - let now = Utc::now(); - Self { - start: now - duration, - end: now, - } - } - - /// Check if a timestamp is within the range - pub fn contains(&self, time: &DateTime<Utc>) -> bool { - *time >= self.start && *time < self.end - } -} - -/// Temporal store trait for cross-modal consistency -#[async_trait] -pub trait TemporalStore: Send + Sync { - /// Type of data being versioned - type Data: Clone + Send + Sync; - - /// Append a new version - async fn append(&self, entity_id: &str, data: Self::Data, author: &str, message: Option<&str>) -> Result<u64, TemporalError>; - - /// Get the latest version - async fn latest(&self, entity_id: &str) -> Result<Option<Version<Self::Data>>, TemporalError>; - - /// Get a specific version - async fn at_version(&self, entity_id: &str, version: u64) -> Result<Option<Version<Self::Data>>, TemporalError>; - - /// Get version at a specific time - async fn at_time(&self, entity_id: &str, time: DateTime<Utc>) -> Result<Option<Version<Self::Data>>, TemporalError>; - - /// Get all versions in a time range - async fn in_range(&self, entity_id: &str, range: &TimeRange) -> Result<Vec<Version<Self::Data>>, TemporalError>; - - /// Get version history - async fn history(&self, entity_id: &str, limit: usize) -> Result<Vec<Version<Self::Data>>, TemporalError>; - - /// Diff two versions - async fn diff(&self, entity_id: &str, v1: u64, v2: u64) -> Result<diff::Diff<Self::Data>, TemporalError> - where - Self::Data: PartialEq, - { - let version1 = self.at_version(entity_id, v1).await?; - let version2 = self.at_version(entity_id, v2).await?; - - diff::compare_values( - version1.as_ref().map(|v| &v.data), - version2.as_ref().map(|v| &v.data), - ) - .map_err(|_| TemporalError::NotFound(format!("No versions to compare for entity {}", entity_id))) - } - - /// Diff two timestamps - async fn diff_time(&self, entity_id: &str, t1: DateTime<Utc>, t2: DateTime<Utc>) -> Result<diff::Diff<Self::Data>, TemporalError> - where - Self::Data: PartialEq, - { - let version1 = self.at_time(entity_id, t1).await?; - let version2 = self.at_time(entity_id, t2).await?; - - diff::compare_values( - version1.as_ref().map(|v| &v.data), - version2.as_ref().map(|v| &v.data), - ) - .map_err(|_| TemporalError::NotFound(format!("No versions to compare for entity {}", entity_id))) - } -} - -/// Type alias for version history map -type VersionHistory<T> = HashMap<String, BTreeMap<u64, Version<T>>>; - -/// In-memory versioned store -pub struct InMemoryVersionStore<T> { - /// Map of entity_id -> (version -> Version<T>) - versions: Arc<RwLock<VersionHistory<T>>>, -} - -impl<T: Clone + Send + Sync + 'static> InMemoryVersionStore<T> { - pub fn new() -> Self { - Self { - versions: Arc::new(RwLock::new(HashMap::new())), - } - } -} - -impl<T: Clone + Send + Sync + 'static> Default for InMemoryVersionStore<T> { - fn default() -> Self { - Self::new() - } -} - -#[async_trait] -impl<T: Clone + Send + Sync + 'static> TemporalStore for InMemoryVersionStore<T> { - type Data = T; - - async fn append(&self, entity_id: &str, data: Self::Data, author: &str, message: Option<&str>) -> Result<u64, TemporalError> { - let mut store = self.versions.write().map_err(|_| TemporalError::LockPoisoned)?; - let versions = store.entry(entity_id.to_string()).or_default(); - - let next_version = versions.keys().last().map(|v| v + 1).unwrap_or(1); - let mut version = Version::new(next_version, data, author); - if let Some(msg) = message { - version = version.with_message(msg); - } - - versions.insert(next_version, version); - Ok(next_version) - } - - async fn latest(&self, entity_id: &str) -> Result<Option<Version<Self::Data>>, TemporalError> { - let store = self.versions.read().map_err(|_| TemporalError::LockPoisoned)?; - Ok(store - .get(entity_id) - .and_then(|versions| versions.values().last().cloned())) - } - - async fn at_version(&self, entity_id: &str, version: u64) -> Result<Option<Version<Self::Data>>, TemporalError> { - let store = self.versions.read().map_err(|_| TemporalError::LockPoisoned)?; - Ok(store - .get(entity_id) - .and_then(|versions| versions.get(&version).cloned())) - } - - async fn at_time(&self, entity_id: &str, time: DateTime<Utc>) -> Result<Option<Version<Self::Data>>, TemporalError> { - let store = self.versions.read().map_err(|_| TemporalError::LockPoisoned)?; - Ok(store.get(entity_id).and_then(|versions| { - versions - .values() - .filter(|v| v.timestamp <= time) - .last() - .cloned() - })) - } - - async fn in_range(&self, entity_id: &str, range: &TimeRange) -> Result<Vec<Version<Self::Data>>, TemporalError> { - let store = self.versions.read().map_err(|_| TemporalError::LockPoisoned)?; - Ok(store - .get(entity_id) - .map(|versions| { - versions - .values() - .filter(|v| range.contains(&v.timestamp)) - .cloned() - .collect() - }) - .unwrap_or_default()) - } - - async fn history(&self, entity_id: &str, limit: usize) -> Result<Vec<Version<Self::Data>>, TemporalError> { - let store = self.versions.read().map_err(|_| TemporalError::LockPoisoned)?; - Ok(store - .get(entity_id) - .map(|versions| { - versions - .values() - .rev() - .take(limit) - .cloned() - .collect() - }) - .unwrap_or_default()) - } -} - -/// Time-series store for metrics -#[async_trait] -pub trait TimeSeriesStore: Send + Sync { - type Value: Clone + Send + Sync; - - /// Append a time point - async fn append(&self, series_id: &str, point: TimePoint<Self::Value>) -> Result<(), TemporalError>; - - /// Query points in a time range - async fn query(&self, series_id: &str, range: &TimeRange) -> Result<Vec<TimePoint<Self::Value>>, TemporalError>; - - /// Get the latest point - async fn latest(&self, series_id: &str) -> Result<Option<TimePoint<Self::Value>>, TemporalError>; -} - -/// In-memory time series store -pub struct InMemoryTimeSeriesStore<T> { - series: Arc<RwLock<HashMap<String, Vec<TimePoint<T>>>>>, -} - -impl<T: Clone + Send + Sync + 'static> InMemoryTimeSeriesStore<T> { - pub fn new() -> Self { - Self { - series: Arc::new(RwLock::new(HashMap::new())), - } - } -} - -impl<T: Clone + Send + Sync + 'static> Default for InMemoryTimeSeriesStore<T> { - fn default() -> Self { - Self::new() - } -} - -#[async_trait] -impl<T: Clone + Send + Sync + 'static> TimeSeriesStore for InMemoryTimeSeriesStore<T> { - type Value = T; - - async fn append(&self, series_id: &str, point: TimePoint<Self::Value>) -> Result<(), TemporalError> { - let mut store = self.series.write().map_err(|_| TemporalError::LockPoisoned)?; - store.entry(series_id.to_string()).or_default().push(point); - Ok(()) - } - - async fn query(&self, series_id: &str, range: &TimeRange) -> Result<Vec<TimePoint<Self::Value>>, TemporalError> { - let store = self.series.read().map_err(|_| TemporalError::LockPoisoned)?; - Ok(store - .get(series_id) - .map(|points| { - points - .iter() - .filter(|p| range.contains(&p.time)) - .cloned() - .collect() - }) - .unwrap_or_default()) - } - - async fn latest(&self, series_id: &str) -> Result<Option<TimePoint<Self::Value>>, TemporalError> { - let store = self.series.read().map_err(|_| TemporalError::LockPoisoned)?; - Ok(store.get(series_id).and_then(|points| points.last().cloned())) - } -} - -#[cfg(test)] -mod tests { - use super::*; - - #[tokio::test] - async fn test_version_store() { - let store: InMemoryVersionStore<String> = InMemoryVersionStore::new(); - - let v1 = store.append("entity1", "data v1".to_string(), "alice", Some("initial")).await.expect("TODO: handle error"); - let v2 = store.append("entity1", "data v2".to_string(), "bob", Some("update")).await.expect("TODO: handle error"); - - assert_eq!(v1, 1); - assert_eq!(v2, 2); - - let latest = store.latest("entity1").await.expect("TODO: handle error").expect("TODO: handle error"); - assert_eq!(latest.version, 2); - assert_eq!(latest.data, "data v2"); - - let v1_data = store.at_version("entity1", 1).await.expect("TODO: handle error").expect("TODO: handle error"); - assert_eq!(v1_data.data, "data v1"); - } - - #[tokio::test] - async fn test_time_series() { - let store: InMemoryTimeSeriesStore<f64> = InMemoryTimeSeriesStore::new(); - - store.append("cpu", TimePoint::now(0.5)).await.expect("TODO: handle error"); - store.append("cpu", TimePoint::now(0.7)).await.expect("TODO: handle error"); - store.append("cpu", TimePoint::now(0.6)).await.expect("TODO: handle error"); - - let latest = store.latest("cpu").await.expect("TODO: handle error").expect("TODO: handle error"); - assert!((latest.value - 0.6).abs() < f64::EPSILON); - } -} diff --git a/verisimdb/rust-core/verisim-temporal/src/persistent.rs b/verisimdb/rust-core/verisim-temporal/src/persistent.rs deleted file mode 100644 index 1b748f3c..00000000 --- a/verisimdb/rust-core/verisim-temporal/src/persistent.rs +++ /dev/null @@ -1,261 +0,0 @@ -// SPDX-License-Identifier: MPL-2.0 -// Copyright (c) 2026 Jonathan D.A. Jewell (hyperpolymath) <j.d.a.jewell@open.ac.uk> -// -// Persistent temporal version store backed by redb via verisim-storage. -// -// Each entity's full version history is stored as a single JSON blob keyed by -// entity_id. On open(), all histories are scanned into an in-memory BTreeMap -// cache for fast reads. Writes go to redb first (durable), then update the -// cache. -// -// The associated type `Data` is `serde_json::Value`, making this store a -// universal versioned key-value store. Higher layers can convert to/from -// concrete types using serde. - -use std::collections::{BTreeMap, HashMap}; -use std::path::Path; -use std::sync::{Arc, RwLock}; - -use async_trait::async_trait; -use chrono::{DateTime, Utc}; -use tracing::info; -use verisim_storage::redb_backend::RedbBackend; -use verisim_storage::typed::TypedStore; - -use crate::{TemporalError, TemporalStore, TimeRange, Version}; - -/// Type alias matching the InMemory store's internal structure. -type VersionHistory = HashMap<String, BTreeMap<u64, Version<serde_json::Value>>>; - -/// Persistent version store: redb for durability, in-memory BTreeMap cache for -/// fast reads and range queries. -/// -/// Each entity's entire version history is stored as a serialized -/// `BTreeMap<u64, Version<serde_json::Value>>` under the entity_id key. -pub struct RedbVersionStore { - /// Typed store for version histories, keyed by entity_id. - store: TypedStore<RedbBackend>, - /// In-memory cache of all version histories. - versions: Arc<RwLock<VersionHistory>>, -} - -impl RedbVersionStore { - /// Open (or create) a persistent version store at the given path. - /// - /// On open, all existing version histories are scanned from redb into the - /// in-memory cache so that reads never hit disk. - pub async fn open(path: impl AsRef<Path>) -> Result<Self, TemporalError> { - let backend = RedbBackend::open(path.as_ref()) - .map_err(|e| TemporalError::Conflict(format!("redb open: {}", e)))?; - let store = TypedStore::new(backend, "ver"); - - let entries: Vec<(String, BTreeMap<u64, Version<serde_json::Value>>)> = store - .scan_prefix("", 1_000_000) - .await - .map_err(|e| TemporalError::Conflict(format!("scan: {}", e)))?; - - let mut cache: VersionHistory = HashMap::new(); - for (id, history) in entries { - cache.insert(id, history); - } - - info!( - entities = cache.len(), - "Loaded temporal version store from redb" - ); - Ok(Self { - store, - versions: Arc::new(RwLock::new(cache)), - }) - } - - /// Persist a single entity's version history to redb. - async fn persist_entity(&self, entity_id: &str) -> Result<(), TemporalError> { - let history = { - let cache = self - .versions - .read() - .map_err(|_| TemporalError::LockPoisoned)?; - cache.get(entity_id).cloned() - }; - if let Some(history) = history { - self.store - .put(entity_id, &history) - .await - .map_err(|e| TemporalError::Conflict(format!("put: {}", e)))?; - } - Ok(()) - } -} - -#[async_trait] -impl TemporalStore for RedbVersionStore { - type Data = serde_json::Value; - - async fn append( - &self, - entity_id: &str, - data: Self::Data, - author: &str, - message: Option<&str>, - ) -> Result<u64, TemporalError> { - let next_version = { - let mut store = self - .versions - .write() - .map_err(|_| TemporalError::LockPoisoned)?; - let versions = store.entry(entity_id.to_string()).or_default(); - - let next_version = versions.keys().last().map(|v| v + 1).unwrap_or(1); - let mut version = Version::new(next_version, data, author); - if let Some(msg) = message { - version = version.with_message(msg); - } - - versions.insert(next_version, version); - next_version - }; - - // Persist to redb after updating cache. - self.persist_entity(entity_id).await?; - Ok(next_version) - } - - async fn latest( - &self, - entity_id: &str, - ) -> Result<Option<Version<Self::Data>>, TemporalError> { - let store = self - .versions - .read() - .map_err(|_| TemporalError::LockPoisoned)?; - Ok(store - .get(entity_id) - .and_then(|versions| versions.values().last().cloned())) - } - - async fn at_version( - &self, - entity_id: &str, - version: u64, - ) -> Result<Option<Version<Self::Data>>, TemporalError> { - let store = self - .versions - .read() - .map_err(|_| TemporalError::LockPoisoned)?; - Ok(store - .get(entity_id) - .and_then(|versions| versions.get(&version).cloned())) - } - - async fn at_time( - &self, - entity_id: &str, - time: DateTime<Utc>, - ) -> Result<Option<Version<Self::Data>>, TemporalError> { - let store = self - .versions - .read() - .map_err(|_| TemporalError::LockPoisoned)?; - Ok(store.get(entity_id).and_then(|versions| { - versions - .values() - .filter(|v| v.timestamp <= time) - .last() - .cloned() - })) - } - - async fn in_range( - &self, - entity_id: &str, - range: &TimeRange, - ) -> Result<Vec<Version<Self::Data>>, TemporalError> { - let store = self - .versions - .read() - .map_err(|_| TemporalError::LockPoisoned)?; - Ok(store - .get(entity_id) - .map(|versions| { - versions - .values() - .filter(|v| range.contains(&v.timestamp)) - .cloned() - .collect() - }) - .unwrap_or_default()) - } - - async fn history( - &self, - entity_id: &str, - limit: usize, - ) -> Result<Vec<Version<Self::Data>>, TemporalError> { - let store = self - .versions - .read() - .map_err(|_| TemporalError::LockPoisoned)?; - Ok(store - .get(entity_id) - .map(|versions| versions.values().rev().take(limit).cloned().collect()) - .unwrap_or_default()) - } -} - -#[cfg(test)] -mod tests { - use super::*; - - #[tokio::test] - async fn test_persistent_temporal_roundtrip() { - let dir = tempfile::tempdir().expect("TODO: handle error"); - let path = dir.path().join("temporal.redb"); - - // Write data in one session. - { - let store = RedbVersionStore::open(&path).await.expect("TODO: handle error"); - let v1 = store - .append( - "entity-1", - serde_json::json!({"name": "Alice", "version": 1}), - "alice", - Some("initial creation"), - ) - .await - .expect("TODO: handle error"); - assert_eq!(v1, 1); - - let v2 = store - .append( - "entity-1", - serde_json::json!({"name": "Alice Updated", "version": 2}), - "bob", - Some("updated name"), - ) - .await - .expect("TODO: handle error"); - assert_eq!(v2, 2); - } - - // Reopen and verify data survived. - { - let store = RedbVersionStore::open(&path).await.expect("TODO: handle error"); - - let latest = store.latest("entity-1").await.expect("TODO: handle error").expect("TODO: handle error"); - assert_eq!(latest.version, 2); - assert_eq!(latest.data["name"], "Alice Updated"); - assert_eq!(latest.author, "bob"); - - let v1 = store.at_version("entity-1", 1).await.expect("TODO: handle error").expect("TODO: handle error"); - assert_eq!(v1.data["name"], "Alice"); - assert_eq!(v1.author, "alice"); - - let history = store.history("entity-1", 10).await.expect("TODO: handle error"); - assert_eq!(history.len(), 2); - // History is most recent first. - assert_eq!(history[0].version, 2); - assert_eq!(history[1].version, 1); - } - } -} diff --git a/verisimdb/rust-core/verisim-temporal/tests/property_tests.rs b/verisimdb/rust-core/verisim-temporal/tests/property_tests.rs deleted file mode 100644 index ebf21026..00000000 --- a/verisimdb/rust-core/verisim-temporal/tests/property_tests.rs +++ /dev/null @@ -1,319 +0,0 @@ -// SPDX-License-Identifier: MPL-2.0 -//! Property-based tests for temporal modality - -use chrono::{Duration, Utc}; -use proptest::prelude::*; -use verisim_temporal::{diff, InMemoryTimeSeriesStore, InMemoryVersionStore, TemporalStore, TimePoint, TimeRange, TimeSeriesStore}; - -/// Generate arbitrary entity IDs -fn arb_entity_id() -> impl Strategy<Value = String> { - "[a-z]{3,8}-[0-9]{1,4}" -} - -/// Generate arbitrary data -fn arb_data() -> impl Strategy<Value = String> { - "[A-Za-z0-9 ]{10,50}" -} - -/// Generate arbitrary author names -fn arb_author() -> impl Strategy<Value = String> { - "[a-z]{4,10}" -} - -proptest! { - #[test] - fn test_version_append_increases_version_number( - entity_id in arb_entity_id(), - data1 in arb_data(), - data2 in arb_data(), - author in arb_author() - ) { - let runtime = tokio::runtime::Runtime::new().unwrap(); - runtime.block_on(async { - let store: InMemoryVersionStore<String> = InMemoryVersionStore::new(); - - let v1 = store.append(&entity_id, data1, &author, None).await.unwrap(); - let v2 = store.append(&entity_id, data2, &author, None).await.unwrap(); - - prop_assert_eq!(v1, 1); - prop_assert_eq!(v2, 2); - - Ok(()) - })?; - } - - #[test] - fn test_latest_returns_most_recent_version( - entity_id in arb_entity_id(), - versions in prop::collection::vec(arb_data(), 1..10), - author in arb_author() - ) { - let runtime = tokio::runtime::Runtime::new().unwrap(); - runtime.block_on(async { - let store: InMemoryVersionStore<String> = InMemoryVersionStore::new(); - - // Append all versions - for data in &versions { - store.append(&entity_id, data.clone(), &author, None).await.unwrap(); - } - - // Latest should be the last one - let latest = store.latest(&entity_id).await.unwrap(); - prop_assert!(latest.is_some()); - - let latest = latest.unwrap(); - prop_assert_eq!(latest.version, versions.len() as u64); - prop_assert_eq!(&latest.data, versions.last().unwrap()); - - Ok(()) - })?; - } - - #[test] - fn test_at_version_retrieves_specific_version( - entity_id in arb_entity_id(), - data1 in arb_data(), - data2 in arb_data(), - data3 in arb_data(), - author in arb_author() - ) { - let runtime = tokio::runtime::Runtime::new().unwrap(); - runtime.block_on(async { - let store: InMemoryVersionStore<String> = InMemoryVersionStore::new(); - - store.append(&entity_id, data1.clone(), &author, None).await.unwrap(); - store.append(&entity_id, data2.clone(), &author, None).await.unwrap(); - store.append(&entity_id, data3.clone(), &author, None).await.unwrap(); - - // Check each version - let v1 = store.at_version(&entity_id, 1).await.unwrap().unwrap(); - let v2 = store.at_version(&entity_id, 2).await.unwrap().unwrap(); - let v3 = store.at_version(&entity_id, 3).await.unwrap().unwrap(); - - prop_assert_eq!(v1.data, data1); - prop_assert_eq!(v2.data, data2); - prop_assert_eq!(v3.data, data3); - - Ok(()) - })?; - } - - #[test] - fn test_history_returns_limited_versions( - entity_id in arb_entity_id(), - versions in prop::collection::vec(arb_data(), 5..15), - author in arb_author(), - limit in 1usize..10 - ) { - let runtime = tokio::runtime::Runtime::new().unwrap(); - runtime.block_on(async { - let store: InMemoryVersionStore<String> = InMemoryVersionStore::new(); - - // Append all versions - for data in &versions { - store.append(&entity_id, data.clone(), &author, None).await.unwrap(); - } - - // Get history with limit - let history = store.history(&entity_id, limit).await.unwrap(); - - let expected_count = std::cmp::min(limit, versions.len()); - prop_assert_eq!(history.len(), expected_count); - - // History should be in reverse order (newest first) - if !history.is_empty() { - prop_assert_eq!(history[0].version, versions.len() as u64); - } - - Ok(()) - })?; - } -} - -/// Integration test: time-travel queries -#[tokio::test] -async fn test_time_travel_query() { - let store: InMemoryVersionStore<String> = InMemoryVersionStore::new(); - - let now = Utc::now(); - let entity_id = "doc-123"; - - // Create version 1 - tokio::time::sleep(tokio::time::Duration::from_millis(10)).await; - let v1_time = Utc::now(); - store.append(entity_id, "version 1".to_string(), "alice", Some("initial")).await.unwrap(); - - // Create version 2 - tokio::time::sleep(tokio::time::Duration::from_millis(10)).await; - let v2_time = Utc::now(); - store.append(entity_id, "version 2".to_string(), "bob", Some("update")).await.unwrap(); - - // Create version 3 - tokio::time::sleep(tokio::time::Duration::from_millis(10)).await; - let v3_time = Utc::now(); - store.append(entity_id, "version 3".to_string(), "charlie", Some("final")).await.unwrap(); - - // Query before any versions - let before = store.at_time(entity_id, now).await.unwrap(); - assert!(before.is_none(), "Should have no version before creation"); - - // Query at v1 time - let at_v1 = store.at_time(entity_id, v1_time + Duration::milliseconds(1)).await.unwrap(); - assert_eq!(at_v1.unwrap().data, "version 1"); - - // Query at v2 time - let at_v2 = store.at_time(entity_id, v2_time + Duration::milliseconds(1)).await.unwrap(); - assert_eq!(at_v2.unwrap().data, "version 2"); - - // Query at v3 time - let at_v3 = store.at_time(entity_id, v3_time + Duration::milliseconds(1)).await.unwrap(); - assert_eq!(at_v3.unwrap().data, "version 3"); - - // Query after all versions - let after = store.at_time(entity_id, Utc::now()).await.unwrap(); - assert_eq!(after.unwrap().data, "version 3"); -} - -/// Integration test: time range queries -#[tokio::test] -async fn test_time_range_query() { - let store: InMemoryVersionStore<String> = InMemoryVersionStore::new(); - let entity_id = "doc-456"; - - let start_time = Utc::now(); - - // Create 5 versions with small delays - for i in 1..=5 { - tokio::time::sleep(tokio::time::Duration::from_millis(10)).await; - store.append( - entity_id, - format!("version {}", i), - "author", - Some(&format!("v{}", i)) - ).await.unwrap(); - } - - let end_time = Utc::now(); - - // Query all versions in range - let range = TimeRange::new(start_time, end_time).unwrap(); - let versions = store.in_range(entity_id, &range).await.unwrap(); - - assert_eq!(versions.len(), 5, "Should find all 5 versions"); - - // Verify they're in order - for (i, version) in versions.iter().enumerate() { - assert_eq!(version.version, (i + 1) as u64); - assert_eq!(version.data, format!("version {}", i + 1)); - } -} - -/// Integration test: diff between versions -#[tokio::test] -async fn test_diff_versions() { - let store: InMemoryVersionStore<String> = InMemoryVersionStore::new(); - let entity_id = "doc-789"; - - // Create versions - store.append(entity_id, "first".to_string(), "alice", None).await.unwrap(); - store.append(entity_id, "second".to_string(), "bob", None).await.unwrap(); - store.append(entity_id, "third".to_string(), "charlie", None).await.unwrap(); - - // Diff v1 and v2 - let diff_1_2 = store.diff(entity_id, 1, 2).await.unwrap(); - assert!(diff_1_2.has_change()); - assert_eq!(diff_1_2.old_value(), Some(&"first".to_string())); - assert_eq!(diff_1_2.new_value(), Some(&"second".to_string())); - - // Diff v1 and v1 (same version) - let diff_same = store.diff(entity_id, 1, 1).await.unwrap(); - assert!(!diff_same.has_change()); - - // Diff v2 and v3 - let diff_2_3 = store.diff(entity_id, 2, 3).await.unwrap(); - assert!(diff_2_3.has_change()); - assert_eq!(diff_2_3.old_value(), Some(&"second".to_string())); - assert_eq!(diff_2_3.new_value(), Some(&"third".to_string())); -} - -/// Integration test: time series -#[tokio::test] -async fn test_time_series_store() { - let store: InMemoryTimeSeriesStore<f64> = InMemoryTimeSeriesStore::new(); - let series_id = "cpu_usage"; - - let start = Utc::now(); - - // Append data points - for i in 0..10 { - let value = 0.1 * i as f64; - let point = TimePoint::now(value); - store.append(series_id, point).await.unwrap(); - tokio::time::sleep(tokio::time::Duration::from_millis(5)).await; - } - - let end = Utc::now(); - - // Query all points - let range = TimeRange::new(start, end).unwrap(); - let points = store.query(series_id, &range).await.unwrap(); - - assert_eq!(points.len(), 10); - - // Verify values - for (i, point) in points.iter().enumerate() { - assert!((point.value - 0.1 * i as f64).abs() < f64::EPSILON); - } - - // Latest should be 0.9 - let latest = store.latest(series_id).await.unwrap().unwrap(); - assert!((latest.value - 0.9).abs() < f64::EPSILON); -} - -/// Integration test: time series with labels -#[tokio::test] -async fn test_time_series_with_labels() { - let store: InMemoryTimeSeriesStore<String> = InMemoryTimeSeriesStore::new(); - - let point1 = TimePoint::now("event_1".to_string()) - .with_label("severity", "high") - .with_label("source", "server-1"); - - let point2 = TimePoint::now("event_2".to_string()) - .with_label("severity", "low") - .with_label("source", "server-2"); - - store.append("events", point1).await.unwrap(); - store.append("events", point2).await.unwrap(); - - let latest = store.latest("events").await.unwrap().unwrap(); - assert_eq!(latest.value, "event_2"); - assert_eq!(latest.labels.get("severity"), Some(&"low".to_string())); - assert_eq!(latest.labels.get("source"), Some(&"server-2".to_string())); -} - -/// Test diff operations -#[test] -fn test_diff_operations() { - // No change - let diff = diff::Diff::no_change("value"); - assert!(!diff.has_change()); - - // Changed - let diff = diff::Diff::changed("old", "new"); - assert!(diff.has_change()); - assert_eq!(diff.old_value(), Some(&"old")); - assert_eq!(diff.new_value(), Some(&"new")); - - // Added - let diff: diff::Diff<&str> = diff::Diff::added("new"); - assert!(diff.has_change()); - assert_eq!(diff.old_value(), None); - assert_eq!(diff.new_value(), Some(&"new")); - - // Removed - let diff: diff::Diff<&str> = diff::Diff::removed("old"); - assert!(diff.has_change()); - assert_eq!(diff.old_value(), Some(&"old")); - assert_eq!(diff.new_value(), None); -} diff --git a/verisimdb/rust-core/verisim-tensor/Cargo.toml b/verisimdb/rust-core/verisim-tensor/Cargo.toml deleted file mode 100644 index b6078d6a..00000000 --- a/verisimdb/rust-core/verisim-tensor/Cargo.toml +++ /dev/null @@ -1,27 +0,0 @@ -# SPDX-License-Identifier: MPL-2.0 - -[package] -name = "verisim-tensor" -description = "Tensor modality - multi-dimensional array operations via Burn" -version.workspace = true -edition.workspace = true -authors.workspace = true -license.workspace = true - -[dependencies] -ndarray.workspace = true -serde.workspace = true -thiserror.workspace = true -tracing.workspace = true -async-trait.workspace = true -serde_json.workspace = true -verisim-storage = { path = "../verisim-storage", optional = true } -tokio.workspace = true - -[dev-dependencies] -tempfile = "3" -proptest.workspace = true - -[features] -default = [] -redb-backend = ["verisim-storage/redb-backend"] diff --git a/verisimdb/rust-core/verisim-tensor/src/lib.rs b/verisimdb/rust-core/verisim-tensor/src/lib.rs deleted file mode 100644 index 042bc6e5..00000000 --- a/verisimdb/rust-core/verisim-tensor/src/lib.rs +++ /dev/null @@ -1,330 +0,0 @@ -// SPDX-License-Identifier: MPL-2.0 -//! VeriSim Tensor Modality -//! -//! Multi-dimensional array operations via ndarray and Burn. -//! Implements Marr's Computational Level: "What transformations apply?" - -#![forbid(unsafe_code)] -#[cfg(feature = "redb-backend")] -pub mod persistent; -#[cfg(feature = "redb-backend")] -pub use persistent::*; -use async_trait::async_trait; -use ndarray::{Array, ArrayD, IxDyn}; -use serde::{Deserialize, Serialize}; -use std::collections::HashMap; -use std::sync::{Arc, RwLock}; -use thiserror::Error; - -/// Tensor modality errors -#[derive(Error, Debug)] -pub enum TensorError { - #[error("Shape mismatch: expected {expected:?}, got {actual:?}")] - ShapeMismatch { expected: Vec<usize>, actual: Vec<usize> }, - - #[error("Tensor not found: {0}")] - NotFound(String), - - #[error("Invalid operation: {0}")] - InvalidOperation(String), - - #[error("Serialization error: {0}")] - SerializationError(String), - - #[error("Lock poisoned: internal concurrency error")] - LockPoisoned, -} - -/// Data type for tensor elements -#[derive(Debug, Clone, Copy, Serialize, Deserialize, PartialEq, Eq)] -pub enum DType { - Float32, - Float64, - Int32, - Int64, - Bool, -} - -/// A named tensor with metadata -#[derive(Debug, Clone, Serialize, Deserialize)] -pub struct Tensor { - /// Unique identifier (matches Octad entity ID) - pub id: String, - /// Shape of the tensor - pub shape: Vec<usize>, - /// Data type - pub dtype: DType, - /// Flattened data (row-major order) - pub data: Vec<f64>, - /// Optional metadata - pub metadata: HashMap<String, String>, -} - -impl Tensor { - /// Create a new tensor from shape and data - pub fn new(id: impl Into<String>, shape: Vec<usize>, data: Vec<f64>) -> Result<Self, TensorError> { - let expected_len: usize = shape.iter().product(); - if data.len() != expected_len { - return Err(TensorError::InvalidOperation(format!( - "Data length {} doesn't match shape {:?} (expected {})", - data.len(), - shape, - expected_len - ))); - } - Ok(Self { - id: id.into(), - shape, - dtype: DType::Float64, - data, - metadata: HashMap::new(), - }) - } - - /// Create a zeros tensor - pub fn zeros(id: impl Into<String>, shape: Vec<usize>) -> Self { - let len: usize = shape.iter().product(); - Self { - id: id.into(), - shape, - dtype: DType::Float64, - data: vec![0.0; len], - metadata: HashMap::new(), - } - } - - /// Create a ones tensor - pub fn ones(id: impl Into<String>, shape: Vec<usize>) -> Self { - let len: usize = shape.iter().product(); - Self { - id: id.into(), - shape, - dtype: DType::Float64, - data: vec![1.0; len], - metadata: HashMap::new(), - } - } - - /// Convert to ndarray - pub fn to_ndarray(&self) -> ArrayD<f64> { - let shape = IxDyn(&self.shape); - // Use C order (row-major) to match data layout documented on line 49 - Array::from_shape_vec(shape, self.data.clone()) - .expect("Shape should match data length") - } - - /// Create from ndarray - pub fn from_ndarray(id: impl Into<String>, arr: &ArrayD<f64>) -> Self { - Self { - id: id.into(), - shape: arr.shape().to_vec(), - dtype: DType::Float64, - data: arr.iter().copied().collect(), - metadata: HashMap::new(), - } - } - - /// Get number of dimensions - pub fn ndim(&self) -> usize { - self.shape.len() - } - - /// Get total number of elements - pub fn numel(&self) -> usize { - self.shape.iter().product() - } - - /// Add metadata - pub fn with_metadata(mut self, key: impl Into<String>, value: impl Into<String>) -> Self { - self.metadata.insert(key.into(), value.into()); - self - } -} - -/// Tensor store trait for cross-modal consistency -#[async_trait] -pub trait TensorStore: Send + Sync { - /// Store a tensor - async fn put(&self, tensor: &Tensor) -> Result<(), TensorError>; - - /// Retrieve a tensor by ID - async fn get(&self, id: &str) -> Result<Option<Tensor>, TensorError>; - - /// Delete a tensor - async fn delete(&self, id: &str) -> Result<(), TensorError>; - - /// List all tensor IDs - async fn list(&self) -> Result<Vec<String>, TensorError>; - - /// Apply element-wise operation - async fn map(&self, id: &str, op: fn(f64) -> f64) -> Result<Tensor, TensorError>; - - /// Reduce along an axis - async fn reduce(&self, id: &str, axis: usize, op: ReduceOp) -> Result<Tensor, TensorError>; -} - -/// Reduction operations -#[derive(Debug, Clone, Copy, Serialize, Deserialize)] -pub enum ReduceOp { - Sum, - Mean, - Max, - Min, - Prod, -} - -/// In-memory tensor store -pub struct InMemoryTensorStore { - tensors: Arc<RwLock<HashMap<String, Tensor>>>, -} - -impl InMemoryTensorStore { - pub fn new() -> Self { - Self { - tensors: Arc::new(RwLock::new(HashMap::new())), - } - } -} - -impl Default for InMemoryTensorStore { - fn default() -> Self { - Self::new() - } -} - -#[async_trait] -impl TensorStore for InMemoryTensorStore { - async fn put(&self, tensor: &Tensor) -> Result<(), TensorError> { - self.tensors.write().map_err(|_| TensorError::LockPoisoned)?.insert(tensor.id.clone(), tensor.clone()); - Ok(()) - } - - async fn get(&self, id: &str) -> Result<Option<Tensor>, TensorError> { - Ok(self.tensors.read().map_err(|_| TensorError::LockPoisoned)?.get(id).cloned()) - } - - async fn delete(&self, id: &str) -> Result<(), TensorError> { - self.tensors.write().map_err(|_| TensorError::LockPoisoned)?.remove(id); - Ok(()) - } - - async fn list(&self) -> Result<Vec<String>, TensorError> { - Ok(self.tensors.read().map_err(|_| TensorError::LockPoisoned)?.keys().cloned().collect()) - } - - async fn map(&self, id: &str, op: fn(f64) -> f64) -> Result<Tensor, TensorError> { - let tensor = self.tensors.read().map_err(|_| TensorError::LockPoisoned)? - .get(id) - .cloned() - .ok_or_else(|| TensorError::NotFound(id.to_string()))?; - - let new_data: Vec<f64> = tensor.data.iter().map(|&x| op(x)).collect(); - Ok(Tensor { - id: format!("{}_mapped", tensor.id), - shape: tensor.shape, - dtype: tensor.dtype, - data: new_data, - metadata: tensor.metadata, - }) - } - - async fn reduce(&self, id: &str, axis: usize, op: ReduceOp) -> Result<Tensor, TensorError> { - let tensor = self.tensors.read().map_err(|_| TensorError::LockPoisoned)? - .get(id) - .cloned() - .ok_or_else(|| TensorError::NotFound(id.to_string()))?; - - if axis >= tensor.shape.len() { - return Err(TensorError::InvalidOperation(format!( - "Axis {} out of bounds for tensor with {} dimensions", - axis, - tensor.shape.len() - ))); - } - - let arr = tensor.to_ndarray(); - let reduced = match op { - ReduceOp::Sum => arr.sum_axis(ndarray::Axis(axis)), - ReduceOp::Mean => arr.mean_axis(ndarray::Axis(axis)).expect("non-empty axis"), - ReduceOp::Max => { - arr.map_axis(ndarray::Axis(axis), |lane| { - lane.iter().copied().fold(f64::NEG_INFINITY, f64::max) - }) - } - ReduceOp::Min => { - arr.map_axis(ndarray::Axis(axis), |lane| { - lane.iter().copied().fold(f64::INFINITY, f64::min) - }) - } - ReduceOp::Prod => { - arr.map_axis(ndarray::Axis(axis), |lane| { - lane.iter().copied().product() - }) - } - }; - - Ok(Tensor::from_ndarray(format!("{}_reduced", tensor.id), &reduced)) - } -} - -#[cfg(test)] -mod tests { - use super::*; - - #[tokio::test] - async fn test_tensor_store() { - let store = InMemoryTensorStore::new(); - - let tensor = Tensor::new("t1", vec![2, 3], vec![1.0, 2.0, 3.0, 4.0, 5.0, 6.0]).expect("TODO: handle error"); - store.put(&tensor).await.expect("TODO: handle error"); - - let retrieved = store.get("t1").await.expect("TODO: handle error").expect("TODO: handle error"); - assert_eq!(retrieved.shape, vec![2, 3]); - assert_eq!(retrieved.data, vec![1.0, 2.0, 3.0, 4.0, 5.0, 6.0]); - } - - #[test] - fn test_tensor_to_ndarray() { - let tensor = Tensor::new("t", vec![2, 2], vec![1.0, 2.0, 3.0, 4.0]).expect("TODO: handle error"); - let arr = tensor.to_ndarray(); - assert_eq!(arr.shape(), &[2, 2]); - } - - #[tokio::test] - async fn test_reduce_max() { - let store = InMemoryTensorStore::new(); - // 2x3 tensor: [[1, 5, 3], [4, 2, 6]] - let tensor = Tensor::new("t_max", vec![2, 3], vec![1.0, 5.0, 3.0, 4.0, 2.0, 6.0]).expect("TODO: handle error"); - store.put(&tensor).await.expect("TODO: handle error"); - - // Max along axis 0 → [4, 5, 6] - let result = store.reduce("t_max", 0, ReduceOp::Max).await.expect("TODO: handle error"); - assert_eq!(result.data, vec![4.0, 5.0, 6.0]); - - // Max along axis 1 → [5, 6] - let result = store.reduce("t_max", 1, ReduceOp::Max).await.expect("TODO: handle error"); - assert_eq!(result.data, vec![5.0, 6.0]); - } - - #[tokio::test] - async fn test_reduce_min() { - let store = InMemoryTensorStore::new(); - let tensor = Tensor::new("t_min", vec![2, 3], vec![1.0, 5.0, 3.0, 4.0, 2.0, 6.0]).expect("TODO: handle error"); - store.put(&tensor).await.expect("TODO: handle error"); - - // Min along axis 0 → [1, 2, 3] - let result = store.reduce("t_min", 0, ReduceOp::Min).await.expect("TODO: handle error"); - assert_eq!(result.data, vec![1.0, 2.0, 3.0]); - } - - #[tokio::test] - async fn test_reduce_prod() { - let store = InMemoryTensorStore::new(); - let tensor = Tensor::new("t_prod", vec![2, 3], vec![1.0, 2.0, 3.0, 4.0, 5.0, 6.0]).expect("TODO: handle error"); - store.put(&tensor).await.expect("TODO: handle error"); - - // Prod along axis 0 → [1*4, 2*5, 3*6] = [4, 10, 18] - let result = store.reduce("t_prod", 0, ReduceOp::Prod).await.expect("TODO: handle error"); - assert_eq!(result.data, vec![4.0, 10.0, 18.0]); - } -} diff --git a/verisimdb/rust-core/verisim-tensor/src/persistent.rs b/verisimdb/rust-core/verisim-tensor/src/persistent.rs deleted file mode 100644 index 176784d8..00000000 --- a/verisimdb/rust-core/verisim-tensor/src/persistent.rs +++ /dev/null @@ -1,123 +0,0 @@ -// SPDX-License-Identifier: MPL-2.0 -// Copyright (c) 2026 Jonathan D.A. Jewell (hyperpolymath) <j.d.a.jewell@open.ac.uk> -// -// Persistent tensor store backed by redb via verisim-storage. - -use std::collections::HashMap; -use std::path::Path; -use std::sync::{Arc, RwLock}; - -use async_trait::async_trait; -use tracing::info; -use verisim_storage::redb_backend::RedbBackend; -use verisim_storage::typed::TypedStore; - -use crate::{ReduceOp, Tensor, TensorError, TensorStore}; - -/// Persistent tensor store: redb for durability, in-memory cache for compute. -pub struct RedbTensorStore { - store: TypedStore<RedbBackend>, - cache: Arc<RwLock<HashMap<String, Tensor>>>, -} - -impl RedbTensorStore { - pub async fn open(path: impl AsRef<Path>) -> Result<Self, TensorError> { - let backend = RedbBackend::open(path.as_ref()) - .map_err(|e| TensorError::SerializationError(format!("redb open: {}", e)))?; - let store = TypedStore::new(backend, "tensor"); - - let entries: Vec<(String, Tensor)> = store - .scan_prefix("", 1_000_000) - .await - .map_err(|e| TensorError::SerializationError(format!("scan: {}", e)))?; - - let mut cache = HashMap::new(); - for (id, tensor) in entries { - cache.insert(id, tensor); - } - - info!(count = cache.len(), "Loaded tensor store from redb"); - Ok(Self { store, cache: Arc::new(RwLock::new(cache)) }) - } -} - -#[async_trait] -impl TensorStore for RedbTensorStore { - async fn put(&self, tensor: &Tensor) -> Result<(), TensorError> { - self.store.put(&tensor.id, tensor).await - .map_err(|e| TensorError::SerializationError(format!("put: {}", e)))?; - let mut c = self.cache.write().map_err(|_| TensorError::LockPoisoned)?; - c.insert(tensor.id.clone(), tensor.clone()); - Ok(()) - } - - async fn get(&self, id: &str) -> Result<Option<Tensor>, TensorError> { - let c = self.cache.read().map_err(|_| TensorError::LockPoisoned)?; - Ok(c.get(id).cloned()) - } - - async fn delete(&self, id: &str) -> Result<(), TensorError> { - self.store.delete(id).await - .map_err(|e| TensorError::SerializationError(format!("delete: {}", e)))?; - let mut c = self.cache.write().map_err(|_| TensorError::LockPoisoned)?; - c.remove(id); - Ok(()) - } - - async fn list(&self) -> Result<Vec<String>, TensorError> { - let c = self.cache.read().map_err(|_| TensorError::LockPoisoned)?; - Ok(c.keys().cloned().collect()) - } - - async fn map(&self, id: &str, op: fn(f64) -> f64) -> Result<Tensor, TensorError> { - let c = self.cache.read().map_err(|_| TensorError::LockPoisoned)?; - let tensor = c.get(id).ok_or_else(|| TensorError::NotFound(id.to_string()))?; - let new_data: Vec<f64> = tensor.data.iter().map(|&v| op(v)).collect(); - Tensor::new(format!("{}_mapped", id), tensor.shape.clone(), new_data) - } - - async fn reduce(&self, id: &str, axis: usize, op: ReduceOp) -> Result<Tensor, TensorError> { - let c = self.cache.read().map_err(|_| TensorError::LockPoisoned)?; - let tensor = c.get(id).ok_or_else(|| TensorError::NotFound(id.to_string()))?; - let arr = tensor.to_ndarray(); - let reduced = match op { - ReduceOp::Sum => arr.sum_axis(ndarray::Axis(axis)), - ReduceOp::Mean => arr.mean_axis(ndarray::Axis(axis)) - .ok_or_else(|| TensorError::InvalidOperation("mean on empty axis".into()))?, - ReduceOp::Max => arr.map_axis(ndarray::Axis(axis), |lane| { - lane.iter().copied().fold(f64::NEG_INFINITY, f64::max) - }), - ReduceOp::Min => arr.map_axis(ndarray::Axis(axis), |lane| { - lane.iter().copied().fold(f64::INFINITY, f64::min) - }), - ReduceOp::Prod => arr.map_axis(ndarray::Axis(axis), |lane| { - lane.iter().copied().product() - }), - }; - Ok(Tensor::from_ndarray(format!("{}_reduced", id), &reduced.into_dyn())) - } -} - -#[cfg(test)] -mod tests { - use super::*; - - #[tokio::test] - async fn test_persistent_tensor_roundtrip() { - let dir = tempfile::tempdir().expect("TODO: handle error"); - let path = dir.path().join("tensor.redb"); - - { - let store = RedbTensorStore::open(&path).await.expect("TODO: handle error"); - let t = Tensor::new("t1", vec![2, 3], vec![1.0, 2.0, 3.0, 4.0, 5.0, 6.0]).expect("TODO: handle error"); - store.put(&t).await.expect("TODO: handle error"); - } - - { - let store = RedbTensorStore::open(&path).await.expect("TODO: handle error"); - let t = store.get("t1").await.expect("TODO: handle error").expect("TODO: handle error"); - assert_eq!(t.shape, vec![2, 3]); - assert_eq!(t.data, vec![1.0, 2.0, 3.0, 4.0, 5.0, 6.0]); - } - } -} diff --git a/verisimdb/rust-core/verisim-vector/Cargo.toml b/verisimdb/rust-core/verisim-vector/Cargo.toml deleted file mode 100644 index b455808a..00000000 --- a/verisimdb/rust-core/verisim-vector/Cargo.toml +++ /dev/null @@ -1,34 +0,0 @@ -# SPDX-License-Identifier: MPL-2.0 - -[package] -name = "verisim-vector" -description = "Vector modality - HNSW-based similarity search" -version.workspace = true -edition.workspace = true -authors.workspace = true -license.workspace = true - -[dependencies] -# hnsw_rs removed — we implement HNSW from scratch to avoid -# the 'b lifetime parameter issue in hnsw_rs 0.3. -ndarray.workspace = true -serde.workspace = true -thiserror.workspace = true -tracing.workspace = true -async-trait.workspace = true -tokio.workspace = true -serde_json.workspace = true - -# Optional: persistent storage via redb -verisim-storage = { path = "../verisim-storage", optional = true } - -[features] -default = [] -redb-backend = ["verisim-storage/redb-backend"] - -[dev-dependencies] -proptest.workspace = true -criterion.workspace = true -tempfile = "3" -tokio = { workspace = true, features = ["macros", "rt-multi-thread"] } - diff --git a/verisimdb/rust-core/verisim-vector/src/hnsw.rs b/verisimdb/rust-core/verisim-vector/src/hnsw.rs deleted file mode 100644 index 38c8a488..00000000 --- a/verisimdb/rust-core/verisim-vector/src/hnsw.rs +++ /dev/null @@ -1,669 +0,0 @@ -// SPDX-License-Identifier: MPL-2.0 -//! HNSW (Hierarchical Navigable Small World) vector index -//! -//! Pure Rust implementation with proper lifetime management. -//! Avoids the `'b` lifetime parameter issue in hnsw_rs 0.3 by owning -//! all graph data directly — no self-referential structs needed. -//! -//! Algorithm: Malkov & Yashunin, "Efficient and robust approximate -//! nearest neighbor search using Hierarchical Navigable Small World graphs" - -use crate::{DistanceMetric, Embedding, SearchResult, VectorError, VectorStore}; -use async_trait::async_trait; -use serde::{Deserialize, Serialize}; -use std::cmp::{Ordering, Reverse}; -use std::collections::{BinaryHeap, HashMap, HashSet}; -use std::hash::{Hash, Hasher}; -use std::sync::atomic::{AtomicU64, Ordering as AtomicOrdering}; -use std::sync::{Arc, RwLock}; - -/// Maximum supported layers in the HNSW graph. -const MAX_LEVELS: usize = 16; - -/// Monotonic counter for level assignment entropy. -static INSERT_COUNTER: AtomicU64 = AtomicU64::new(0); - -// --------------------------------------------------------------------------- -// Ordered f32 for BinaryHeap (f32 doesn't implement Ord) -// --------------------------------------------------------------------------- - -#[derive(Clone, Copy, PartialEq)] -struct Dist(f32); - -impl Eq for Dist {} - -impl PartialOrd for Dist { - fn partial_cmp(&self, other: &Self) -> Option<Ordering> { - Some(self.cmp(other)) - } -} - -impl Ord for Dist { - fn cmp(&self, other: &Self) -> Ordering { - self.0.partial_cmp(&other.0).unwrap_or(Ordering::Equal) - } -} - -// --------------------------------------------------------------------------- -// Configuration -// --------------------------------------------------------------------------- - -/// HNSW index configuration parameters. -#[derive(Debug, Clone, Serialize, Deserialize)] -pub struct HnswConfig { - /// Max bidirectional connections per node per layer (M parameter). - pub max_connections: usize, - /// Max connections for layer 0 (typically 2*M). - pub max_connections_layer0: usize, - /// Size of dynamic candidate list during construction. - pub ef_construction: usize, - /// Size of dynamic candidate list during search. - pub ef_search: usize, -} - -impl Default for HnswConfig { - fn default() -> Self { - Self { - max_connections: 16, - max_connections_layer0: 32, - ef_construction: 200, - ef_search: 64, - } - } -} - -// --------------------------------------------------------------------------- -// Internal graph structures (no public lifetime parameters) -// --------------------------------------------------------------------------- - -/// Internal node — owns its vector data. -struct Node { - id: String, - vector: Vec<f32>, - metadata: HashMap<String, String>, - /// Neighbors per layer (layer index -> vec of node indices). - neighbors: Vec<Vec<usize>>, - /// Assigned level for this node. - level: usize, - /// Soft-delete flag. - deleted: bool, -} - -/// Internal graph state — fully owned, no lifetimes. -struct Graph { - nodes: Vec<Node>, - id_map: HashMap<String, usize>, - entry_point: Option<usize>, - current_max_level: usize, -} - -impl Graph { - fn new() -> Self { - Self { - nodes: Vec::new(), - id_map: HashMap::new(), - entry_point: None, - current_max_level: 0, - } - } - - /// Compute distance between two vectors (lower = closer). - fn distance(metric: DistanceMetric, a: &[f32], b: &[f32]) -> f32 { - match metric { - DistanceMetric::Cosine => { - let mut dot = 0.0f32; - let mut norm_a = 0.0f32; - let mut norm_b = 0.0f32; - for (x, y) in a.iter().zip(b.iter()) { - dot += x * y; - norm_a += x * x; - norm_b += y * y; - } - let denom = norm_a.sqrt() * norm_b.sqrt(); - if denom > 0.0 { - 1.0 - dot / denom - } else { - 1.0 - } - } - DistanceMetric::Euclidean => a - .iter() - .zip(b.iter()) - .map(|(x, y)| (x - y).powi(2)) - .sum::<f32>() - .sqrt(), - DistanceMetric::DotProduct => { - // Negate so lower value = higher dot product = more similar - -a.iter().zip(b.iter()).map(|(x, y)| x * y).sum::<f32>() - } - } - } - - /// Convert HNSW distance back to similarity score for results. - fn distance_to_score(metric: DistanceMetric, distance: f32) -> f32 { - match metric { - DistanceMetric::Cosine => 1.0 - distance, - DistanceMetric::Euclidean => 1.0 / (1.0 + distance), - DistanceMetric::DotProduct => -distance, - } - } - - /// Assign a random level using hash-based PRNG (no `rand` dependency). - fn assign_level(id: &str, max_connections: usize) -> usize { - let count = INSERT_COUNTER.fetch_add(1, AtomicOrdering::Relaxed); - let mut hasher = std::collections::hash_map::DefaultHasher::new(); - id.hash(&mut hasher); - count.hash(&mut hasher); - let hash = hasher.finish(); - - // Convert to uniform float in (0, 1), avoiding ln(0) - let uniform = ((hash >> 11) as f64 + 1.0) / ((1u64 << 53) as f64 + 1.0); - let ml = 1.0 / (max_connections as f64).ln(); - let level = (-uniform.ln() * ml).floor() as usize; - level.min(MAX_LEVELS - 1) - } - - /// Search a single layer, returning up to `ef` nearest neighbors. - /// Returns Vec<(distance, node_index)> sorted by distance ascending. - fn search_layer( - &self, - query: &[f32], - entry_points: &[usize], - ef: usize, - layer: usize, - metric: DistanceMetric, - ) -> Vec<(f32, usize)> { - let mut visited = HashSet::new(); - // Min-heap for candidates (closest first) - let mut candidates: BinaryHeap<Reverse<(Dist, usize)>> = BinaryHeap::new(); - // Max-heap for results (furthest first, for pruning) - let mut results: BinaryHeap<(Dist, usize)> = BinaryHeap::new(); - - for &ep in entry_points { - if ep >= self.nodes.len() { - continue; - } - let dist = Self::distance(metric, query, &self.nodes[ep].vector); - visited.insert(ep); - // Always add to candidates (for navigation), only add live nodes to results - candidates.push(Reverse((Dist(dist), ep))); - if !self.nodes[ep].deleted { - results.push((Dist(dist), ep)); - } - } - - while let Some(Reverse((Dist(c_dist), c_idx))) = candidates.pop() { - let furthest_dist = results.peek().map(|(Dist(d), _)| *d).unwrap_or(f32::MAX); - // For deleted-heavy graphs, only break when we have enough results - if c_dist > furthest_dist && results.len() >= ef { - break; - } - - if layer < self.nodes[c_idx].neighbors.len() { - for &neighbor_idx in &self.nodes[c_idx].neighbors[layer] { - if neighbor_idx >= self.nodes.len() || visited.contains(&neighbor_idx) { - continue; - } - visited.insert(neighbor_idx); - - let dist = Self::distance(metric, query, &self.nodes[neighbor_idx].vector); - - // Always add to candidates for graph traversal - candidates.push(Reverse((Dist(dist), neighbor_idx))); - - // Only add live nodes to results - if !self.nodes[neighbor_idx].deleted { - let furthest_dist = - results.peek().map(|(Dist(d), _)| *d).unwrap_or(f32::MAX); - if dist < furthest_dist || results.len() < ef { - results.push((Dist(dist), neighbor_idx)); - if results.len() > ef { - results.pop(); - } - } - } - } - } - } - - let mut result_vec: Vec<(f32, usize)> = - results.into_iter().map(|(Dist(d), idx)| (d, idx)).collect(); - result_vec.sort_by(|a, b| a.0.partial_cmp(&b.0).unwrap_or(Ordering::Equal)); - result_vec - } - - /// Select M nearest neighbors from sorted candidates. - fn select_neighbors(candidates: &[(f32, usize)], m: usize) -> Vec<usize> { - candidates.iter().take(m).map(|(_, idx)| *idx).collect() - } - - /// Insert a node into the HNSW graph. - fn insert( - &mut self, - id: String, - vector: Vec<f32>, - metadata: HashMap<String, String>, - config: &HnswConfig, - metric: DistanceMetric, - ) { - // Upsert: if ID exists, update vector in place (connections stay valid - // for approximate search — small perturbations don't break HNSW). - if let Some(&existing_idx) = self.id_map.get(&id) { - self.nodes[existing_idx].vector = vector; - self.nodes[existing_idx].metadata = metadata; - self.nodes[existing_idx].deleted = false; - return; - } - - let level = Self::assign_level(&id, config.max_connections); - let node_idx = self.nodes.len(); - - let node = Node { - id: id.clone(), - vector, - metadata, - neighbors: (0..=level).map(|_| Vec::new()).collect(), - level, - deleted: false, - }; - self.nodes.push(node); - self.id_map.insert(id, node_idx); - - // First node — just set as entry point. - if self.entry_point.is_none() { - self.entry_point = Some(node_idx); - self.current_max_level = level; - return; - } - - let ep = self.entry_point.expect("TODO: handle error"); - let mut current_ep = vec![ep]; - - // Phase 1: Greedy descent from top layer to (node level + 1) - let top = self.current_max_level; - if top > level { - for l in (level + 1..=top).rev() { - let nearest = self.search_layer( - &self.nodes[node_idx].vector, - ¤t_ep, - 1, - l, - metric, - ); - if let Some(&(_, idx)) = nearest.first() { - current_ep = vec![idx]; - } - } - } - - // Phase 2: Insert at each layer from min(level, top) down to 0 - let insert_top = level.min(top); - for l in (0..=insert_top).rev() { - let max_conn = if l == 0 { - config.max_connections_layer0 - } else { - config.max_connections - }; - - let nearest = self.search_layer( - &self.nodes[node_idx].vector, - ¤t_ep, - config.ef_construction, - l, - metric, - ); - - let selected = Self::select_neighbors(&nearest, max_conn); - - // Bidirectional connections - for &neighbor_idx in &selected { - // node -> neighbor - if l < self.nodes[node_idx].neighbors.len() { - self.nodes[node_idx].neighbors[l].push(neighbor_idx); - } - - // neighbor -> node (ensure neighbor has layer allocated) - while self.nodes[neighbor_idx].neighbors.len() <= l { - self.nodes[neighbor_idx].neighbors.push(Vec::new()); - } - self.nodes[neighbor_idx].neighbors[l].push(node_idx); - - // Prune neighbor if over capacity - if self.nodes[neighbor_idx].neighbors[l].len() > max_conn { - let neighbor_vec = self.nodes[neighbor_idx].vector.clone(); - let mut scored: Vec<(f32, usize)> = self.nodes[neighbor_idx].neighbors[l] - .iter() - .map(|&n| { - let d = Self::distance(metric, &neighbor_vec, &self.nodes[n].vector); - (d, n) - }) - .collect(); - scored.sort_by(|a, b| a.0.partial_cmp(&b.0).unwrap_or(Ordering::Equal)); - self.nodes[neighbor_idx].neighbors[l] = - scored.iter().take(max_conn).map(|(_, idx)| *idx).collect(); - } - } - - current_ep = nearest.iter().map(|(_, idx)| *idx).collect(); - if current_ep.is_empty() { - current_ep = vec![ep]; - } - } - - // Update entry point if new node has higher level - if level > self.current_max_level { - self.entry_point = Some(node_idx); - self.current_max_level = level; - } - } - - /// Search the graph for k nearest neighbors. - fn search( - &self, - query: &[f32], - k: usize, - ef_search: usize, - metric: DistanceMetric, - ) -> Vec<(f32, usize)> { - let ep = match self.entry_point { - Some(ep) => ep, - None => return Vec::new(), - }; - - let mut current_ep = vec![ep]; - - // Greedy descent from top layer to layer 1 - for l in (1..=self.current_max_level).rev() { - let nearest = self.search_layer(query, ¤t_ep, 1, l, metric); - if let Some(&(_, idx)) = nearest.first() { - current_ep = vec![idx]; - } - } - - // Beam search on layer 0 - let mut results = self.search_layer(query, ¤t_ep, ef_search.max(k), 0, metric); - results.truncate(k); - results - } -} - -// --------------------------------------------------------------------------- -// Public API -// --------------------------------------------------------------------------- - -/// HNSW-indexed vector store. -/// -/// Provides O(log n) approximate nearest neighbor search with configurable -/// recall/speed tradeoff via `ef_search`. Thread-safe: concurrent reads, -/// exclusive writes via `RwLock`. -pub struct HnswVectorStore { - config: HnswConfig, - dimension: usize, - metric: DistanceMetric, - graph: Arc<RwLock<Graph>>, -} - -impl HnswVectorStore { - /// Create a new HNSW vector store with custom configuration. - pub fn new(dimension: usize, metric: DistanceMetric, config: HnswConfig) -> Self { - Self { - config, - dimension, - metric, - graph: Arc::new(RwLock::new(Graph::new())), - } - } - - /// Create with default HNSW parameters (M=16, ef_construction=200, ef_search=64). - pub fn with_defaults(dimension: usize, metric: DistanceMetric) -> Self { - Self::new(dimension, metric, HnswConfig::default()) - } - - /// Get the number of non-deleted vectors in the index. - pub fn len(&self) -> usize { - let graph = self.graph.read().unwrap_or_else(|e| e.into_inner()); - graph.nodes.iter().filter(|n| !n.deleted).count() - } - - /// Check if the index is empty. - pub fn is_empty(&self) -> bool { - self.len() == 0 - } - - /// Get the current HNSW configuration. - pub fn config(&self) -> &HnswConfig { - &self.config - } -} - -#[async_trait] -impl VectorStore for HnswVectorStore { - async fn upsert(&self, embedding: &Embedding) -> Result<(), VectorError> { - if embedding.dim() != self.dimension { - return Err(VectorError::DimensionMismatch { - expected: self.dimension, - actual: embedding.dim(), - }); - } - - let mut graph = self.graph.write().map_err(|_| VectorError::LockPoisoned)?; - graph.insert( - embedding.id.clone(), - embedding.vector.clone(), - embedding.metadata.clone(), - &self.config, - self.metric, - ); - Ok(()) - } - - async fn search(&self, query: &[f32], k: usize) -> Result<Vec<SearchResult>, VectorError> { - if query.len() != self.dimension { - return Err(VectorError::DimensionMismatch { - expected: self.dimension, - actual: query.len(), - }); - } - - let graph = self.graph.read().map_err(|_| VectorError::LockPoisoned)?; - let results = graph.search(query, k, self.config.ef_search, self.metric); - - Ok(results - .into_iter() - .map(|(dist, idx)| SearchResult { - id: graph.nodes[idx].id.clone(), - score: Graph::distance_to_score(self.metric, dist), - }) - .collect()) - } - - async fn get(&self, id: &str) -> Result<Option<Embedding>, VectorError> { - let graph = self.graph.read().map_err(|_| VectorError::LockPoisoned)?; - Ok(graph.id_map.get(id).and_then(|&idx| { - let node = &graph.nodes[idx]; - if node.deleted { - None - } else { - Some(Embedding { - id: node.id.clone(), - vector: node.vector.clone(), - metadata: node.metadata.clone(), - }) - } - })) - } - - async fn delete(&self, id: &str) -> Result<(), VectorError> { - let mut graph = self.graph.write().map_err(|_| VectorError::LockPoisoned)?; - if let Some(&idx) = graph.id_map.get(id) { - graph.nodes[idx].deleted = true; - } - Ok(()) - } - - fn dimension(&self) -> usize { - self.dimension - } -} - -// --------------------------------------------------------------------------- -// Tests -// --------------------------------------------------------------------------- - -#[cfg(test)] -mod tests { - use super::*; - - #[tokio::test] - async fn test_hnsw_basic_insert_and_search() { - let store = HnswVectorStore::with_defaults(3, DistanceMetric::Cosine); - - let e1 = Embedding::new("e1", vec![1.0, 0.0, 0.0]); - let e2 = Embedding::new("e2", vec![0.9, 0.1, 0.0]); - let e3 = Embedding::new("e3", vec![0.0, 1.0, 0.0]); - - store.upsert(&e1).await.expect("TODO: handle error"); - store.upsert(&e2).await.expect("TODO: handle error"); - store.upsert(&e3).await.expect("TODO: handle error"); - - let results = store.search(&[1.0, 0.0, 0.0], 2).await.expect("TODO: handle error"); - assert_eq!(results.len(), 2); - assert_eq!(results[0].id, "e1"); - assert_eq!(results[1].id, "e2"); - } - - #[tokio::test] - async fn test_hnsw_upsert_updates_vector() { - let store = HnswVectorStore::with_defaults(3, DistanceMetric::Cosine); - - store - .upsert(&Embedding::new("e1", vec![1.0, 0.0, 0.0])) - .await - .expect("TODO: handle error"); - store - .upsert(&Embedding::new("e1", vec![0.0, 1.0, 0.0])) - .await - .expect("TODO: handle error"); - - let emb = store.get("e1").await.expect("TODO: handle error").expect("TODO: handle error"); - assert_eq!(emb.vector, vec![0.0, 1.0, 0.0]); - } - - #[tokio::test] - async fn test_hnsw_delete() { - let store = HnswVectorStore::with_defaults(3, DistanceMetric::Cosine); - - store - .upsert(&Embedding::new("e1", vec![1.0, 0.0, 0.0])) - .await - .expect("TODO: handle error"); - store.delete("e1").await.expect("TODO: handle error"); - - assert!(store.get("e1").await.expect("TODO: handle error").is_none()); - assert_eq!(store.len(), 0); - } - - #[tokio::test] - async fn test_hnsw_dimension_mismatch() { - let store = HnswVectorStore::with_defaults(3, DistanceMetric::Cosine); - let result = store - .upsert(&Embedding::new("e1", vec![1.0, 0.0])) - .await; - assert!(result.is_err()); - } - - #[tokio::test] - async fn test_hnsw_euclidean() { - let store = HnswVectorStore::with_defaults(2, DistanceMetric::Euclidean); - - store - .upsert(&Embedding::new("origin", vec![0.0, 0.0])) - .await - .expect("TODO: handle error"); - store - .upsert(&Embedding::new("near", vec![1.0, 0.0])) - .await - .expect("TODO: handle error"); - store - .upsert(&Embedding::new("far", vec![10.0, 10.0])) - .await - .expect("TODO: handle error"); - - let results = store.search(&[0.0, 0.0], 2).await.expect("TODO: handle error"); - assert_eq!(results[0].id, "origin"); - assert_eq!(results[1].id, "near"); - } - - #[tokio::test] - async fn test_hnsw_dot_product() { - let store = HnswVectorStore::with_defaults(3, DistanceMetric::DotProduct); - - store - .upsert(&Embedding::new("high", vec![1.0, 1.0, 1.0])) - .await - .expect("TODO: handle error"); - store - .upsert(&Embedding::new("low", vec![0.1, 0.1, 0.1])) - .await - .expect("TODO: handle error"); - - let results = store.search(&[1.0, 1.0, 1.0], 2).await.expect("TODO: handle error"); - assert_eq!(results[0].id, "high"); - assert!(results[0].score > results[1].score); - } - - #[tokio::test] - async fn test_hnsw_many_vectors() { - let dim = 32; - let store = HnswVectorStore::with_defaults(dim, DistanceMetric::Cosine); - - for i in 0..200 { - let mut vec = vec![0.0f32; dim]; - vec[i % dim] = 1.0; - vec[(i * 7) % dim] += 0.5; - store - .upsert(&Embedding::new(format!("v{i}"), vec)) - .await - .expect("TODO: handle error"); - } - - let results = store.search(&vec![1.0; dim], 10).await.expect("TODO: handle error"); - assert_eq!(results.len(), 10); - - // Scores should be monotonically non-increasing - for w in results.windows(2) { - assert!(w[0].score >= w[1].score); - } - } - - #[tokio::test] - async fn test_hnsw_empty_search() { - let store = HnswVectorStore::with_defaults(3, DistanceMetric::Cosine); - let results = store.search(&[1.0, 0.0, 0.0], 5).await.expect("TODO: handle error"); - assert!(results.is_empty()); - } - - #[tokio::test] - async fn test_hnsw_deleted_not_in_results() { - let store = HnswVectorStore::with_defaults(3, DistanceMetric::Cosine); - - store - .upsert(&Embedding::new("a", vec![1.0, 0.0, 0.0])) - .await - .expect("TODO: handle error"); - store - .upsert(&Embedding::new("b", vec![0.9, 0.1, 0.0])) - .await - .expect("TODO: handle error"); - store - .upsert(&Embedding::new("c", vec![0.0, 1.0, 0.0])) - .await - .expect("TODO: handle error"); - - store.delete("a").await.expect("TODO: handle error"); - - let results = store.search(&[1.0, 0.0, 0.0], 3).await.expect("TODO: handle error"); - assert!(!results.iter().any(|r| r.id == "a")); - assert_eq!(results[0].id, "b"); - } -} diff --git a/verisimdb/rust-core/verisim-vector/src/lib.rs b/verisimdb/rust-core/verisim-vector/src/lib.rs deleted file mode 100644 index d08ebd02..00000000 --- a/verisimdb/rust-core/verisim-vector/src/lib.rs +++ /dev/null @@ -1,260 +0,0 @@ -// SPDX-License-Identifier: MPL-2.0 -//! VeriSim Vector Modality -//! -//! HNSW-based similarity search for embeddings. -//! Implements Marr's Computational Level: "What is similar to what?" - -#![forbid(unsafe_code)] -mod hnsw; -#[cfg(feature = "redb-backend")] -pub mod persistent; - -pub use hnsw::{HnswConfig, HnswVectorStore}; -#[cfg(feature = "redb-backend")] -pub use persistent::RedbVectorStore; - -use async_trait::async_trait; -use ndarray::{Array1, ArrayView1}; -use serde::{Deserialize, Serialize}; -use std::collections::HashMap; -use std::sync::{Arc, RwLock}; -use thiserror::Error; - -/// Vector modality errors -#[derive(Error, Debug)] -pub enum VectorError { - #[error("Dimension mismatch: expected {expected}, got {actual}")] - DimensionMismatch { expected: usize, actual: usize }, - - #[error("Vector not found: {0}")] - NotFound(String), - - #[error("Index error: {0}")] - IndexError(String), - - #[error("Serialization error: {0}")] - SerializationError(String), - - #[error("Lock poisoned: internal concurrency error")] - LockPoisoned, -} - -/// A vector embedding with metadata -#[derive(Debug, Clone, Serialize, Deserialize)] -pub struct Embedding { - /// Unique identifier (matches Octad entity ID) - pub id: String, - /// The embedding vector - pub vector: Vec<f32>, - /// Optional metadata - pub metadata: HashMap<String, String>, -} - -impl Embedding { - /// Create a new embedding - pub fn new(id: impl Into<String>, vector: Vec<f32>) -> Self { - Self { - id: id.into(), - vector, - metadata: HashMap::new(), - } - } - - /// Add metadata - pub fn with_metadata(mut self, key: impl Into<String>, value: impl Into<String>) -> Self { - self.metadata.insert(key.into(), value.into()); - self - } - - /// Get dimensionality - pub fn dim(&self) -> usize { - self.vector.len() - } - - /// Convert to ndarray - pub fn as_array(&self) -> Array1<f32> { - Array1::from_vec(self.vector.clone()) - } -} - -/// Search result with score -#[derive(Debug, Clone, Serialize, Deserialize)] -pub struct SearchResult { - /// Entity ID - pub id: String, - /// Similarity score (higher is more similar for cosine) - pub score: f32, -} - -/// Distance metric for similarity -#[derive(Debug, Clone, Copy, Serialize, Deserialize, Default)] -pub enum DistanceMetric { - #[default] - Cosine, - Euclidean, - DotProduct, -} - -/// Vector store trait for cross-modal consistency -#[async_trait] -pub trait VectorStore: Send + Sync { - /// Insert or update an embedding - async fn upsert(&self, embedding: &Embedding) -> Result<(), VectorError>; - - /// Search for similar vectors - async fn search(&self, query: &[f32], k: usize) -> Result<Vec<SearchResult>, VectorError>; - - /// Get embedding by ID - async fn get(&self, id: &str) -> Result<Option<Embedding>, VectorError>; - - /// Delete embedding by ID - async fn delete(&self, id: &str) -> Result<(), VectorError>; - - /// Get the dimensionality of the index - fn dimension(&self) -> usize; -} - -/// In-memory vector store with brute-force search -/// -/// Note: This is a simple implementation for correctness. For production -/// workloads with >10k vectors, integrate HNSW with proper lifetime management. -pub struct BruteForceVectorStore { - dimension: usize, - metric: DistanceMetric, - embeddings: Arc<RwLock<HashMap<String, Embedding>>>, -} - -impl BruteForceVectorStore { - /// Create a new vector store - pub fn new(dimension: usize, metric: DistanceMetric) -> Self { - Self { - dimension, - metric, - embeddings: Arc::new(RwLock::new(HashMap::new())), - } - } - - /// Normalize vector for cosine similarity - fn normalize(v: &[f32]) -> Vec<f32> { - let norm: f32 = v.iter().map(|x| x * x).sum::<f32>().sqrt(); - if norm > 0.0 { - v.iter().map(|x| x / norm).collect() - } else { - v.to_vec() - } - } - - /// Compute similarity between two vectors based on metric - fn similarity(&self, a: &[f32], b: &[f32]) -> f32 { - match self.metric { - DistanceMetric::Cosine => { - let a_norm = Self::normalize(a); - let b_norm = Self::normalize(b); - a_norm.iter().zip(b_norm.iter()).map(|(x, y)| x * y).sum() - } - DistanceMetric::DotProduct => { - a.iter().zip(b.iter()).map(|(x, y)| x * y).sum() - } - DistanceMetric::Euclidean => { - let dist_sq: f32 = a.iter().zip(b.iter()).map(|(x, y)| (x - y).powi(2)).sum(); - 1.0 / (1.0 + dist_sq.sqrt()) // Convert distance to similarity - } - } - } -} - -#[async_trait] -impl VectorStore for BruteForceVectorStore { - async fn upsert(&self, embedding: &Embedding) -> Result<(), VectorError> { - if embedding.dim() != self.dimension { - return Err(VectorError::DimensionMismatch { - expected: self.dimension, - actual: embedding.dim(), - }); - } - - self.embeddings - .write() - .map_err(|_| VectorError::LockPoisoned)? - .insert(embedding.id.clone(), embedding.clone()); - - Ok(()) - } - - async fn search(&self, query: &[f32], k: usize) -> Result<Vec<SearchResult>, VectorError> { - if query.len() != self.dimension { - return Err(VectorError::DimensionMismatch { - expected: self.dimension, - actual: query.len(), - }); - } - - let embeddings = self.embeddings.read().map_err(|_| VectorError::LockPoisoned)?; - - // Compute similarities for all embeddings (brute-force) - let mut scored: Vec<_> = embeddings - .iter() - .map(|(id, emb)| { - let score = self.similarity(query, &emb.vector); - SearchResult { - id: id.clone(), - score, - } - }) - .collect(); - - // Sort by similarity descending - scored.sort_by(|a, b| b.score.partial_cmp(&a.score).unwrap_or(std::cmp::Ordering::Equal)); - - // Return top k - scored.truncate(k); - Ok(scored) - } - - async fn get(&self, id: &str) -> Result<Option<Embedding>, VectorError> { - Ok(self.embeddings.read().map_err(|_| VectorError::LockPoisoned)?.get(id).cloned()) - } - - async fn delete(&self, id: &str) -> Result<(), VectorError> { - self.embeddings.write().map_err(|_| VectorError::LockPoisoned)?.remove(id); - Ok(()) - } - - fn dimension(&self) -> usize { - self.dimension - } -} - -/// Compute cosine similarity between two vectors -pub fn cosine_similarity(a: ArrayView1<f32>, b: ArrayView1<f32>) -> f32 { - let dot: f32 = a.iter().zip(b.iter()).map(|(x, y)| x * y).sum(); - let norm_a: f32 = a.iter().map(|x| x * x).sum::<f32>().sqrt(); - let norm_b: f32 = b.iter().map(|x| x * x).sum::<f32>().sqrt(); - if norm_a > 0.0 && norm_b > 0.0 { - dot / (norm_a * norm_b) - } else { - 0.0 - } -} - -#[cfg(test)] -mod tests { - use super::*; - - #[tokio::test] - async fn test_upsert_and_search() { - let store = BruteForceVectorStore::new(3, DistanceMetric::Cosine); - - let e1 = Embedding::new("e1", vec![1.0, 0.0, 0.0]); - let e2 = Embedding::new("e2", vec![0.9, 0.1, 0.0]); - let e3 = Embedding::new("e3", vec![0.0, 1.0, 0.0]); - - store.upsert(&e1).await.expect("TODO: handle error"); - store.upsert(&e2).await.expect("TODO: handle error"); - store.upsert(&e3).await.expect("TODO: handle error"); - - let results = store.search(&[1.0, 0.0, 0.0], 2).await.expect("TODO: handle error"); - assert_eq!(results.len(), 2); - assert_eq!(results[0].id, "e1"); - } -} diff --git a/verisimdb/rust-core/verisim-vector/src/persistent.rs b/verisimdb/rust-core/verisim-vector/src/persistent.rs deleted file mode 100644 index c561847a..00000000 --- a/verisimdb/rust-core/verisim-vector/src/persistent.rs +++ /dev/null @@ -1,285 +0,0 @@ -// SPDX-License-Identifier: MPL-2.0 -// Copyright (c) 2026 Jonathan D.A. Jewell (hyperpolymath) <j.d.a.jewell@open.ac.uk> -// -// Persistent vector store backed by redb via verisim-storage. -// -// Durable storage of embeddings with an ephemeral in-memory index for fast -// similarity search. On startup, all embeddings are loaded from redb and the -// index is rebuilt. Writes go to both redb (durable) and the in-memory index -// (fast search). -// -// Design: -// - TypedStore<RedbBackend> with namespace "vec" handles serialisation + persistence -// - In-memory HashMap + brute-force search for queries (same as BruteForceVectorStore) -// - Startup: load all embeddings from redb, populate in-memory index -// - Upsert: write to redb first (durable), then update in-memory index -// - Delete: remove from redb first, then remove from in-memory index -// - Search: in-memory only (fast, no disk I/O) - -use std::collections::HashMap; -use std::path::Path; -use std::sync::{Arc, RwLock}; - -use async_trait::async_trait; -use tracing::{debug, info}; -use verisim_storage::redb_backend::RedbBackend; -use verisim_storage::typed::TypedStore; - -use crate::{DistanceMetric, Embedding, SearchResult, VectorError, VectorStore}; - -/// Persistent vector store: redb for durability, in-memory index for search. -pub struct RedbVectorStore { - /// Dimensionality of stored vectors. - dimension: usize, - /// Distance metric for similarity computation. - metric: DistanceMetric, - /// Durable storage: TypedStore<RedbBackend> with namespace "vec". - store: TypedStore<RedbBackend>, - /// Ephemeral in-memory index for fast similarity search. - /// Rebuilt from redb on startup. - index: Arc<RwLock<HashMap<String, Embedding>>>, -} - -impl RedbVectorStore { - /// Open or create a persistent vector store at the given path. - /// - /// On first open, creates an empty redb database. - /// On subsequent opens, loads all embeddings from redb and rebuilds the - /// in-memory index. Returns the number of embeddings loaded. - pub async fn open( - path: impl AsRef<Path>, - dimension: usize, - metric: DistanceMetric, - ) -> Result<Self, VectorError> { - let backend = RedbBackend::open(path.as_ref()).map_err(|e| { - VectorError::IndexError(format!("Failed to open redb: {}", e)) - })?; - let store = TypedStore::new(backend, "vec"); - - let mut index = HashMap::new(); - - // Load all existing embeddings from redb into memory - let entries: Vec<(String, Embedding)> = store - .scan_prefix("", 1_000_000) - .await - .map_err(|e| VectorError::IndexError(format!("Failed to scan redb: {}", e)))?; - - for (id, embedding) in &entries { - // Validate dimensionality - if embedding.dim() != dimension { - debug!( - id = %id, - expected = dimension, - actual = embedding.dim(), - "Skipping embedding with wrong dimensionality" - ); - continue; - } - index.insert(id.clone(), embedding.clone()); - } - - info!( - count = index.len(), - dimension = dimension, - path = %path.as_ref().display(), - "Loaded vector store from redb" - ); - - Ok(Self { - dimension, - metric, - store, - index: Arc::new(RwLock::new(index)), - }) - } - - /// Normalise a vector for cosine similarity. - fn normalize(v: &[f32]) -> Vec<f32> { - let norm: f32 = v.iter().map(|x| x * x).sum::<f32>().sqrt(); - if norm > 0.0 { - v.iter().map(|x| x / norm).collect() - } else { - v.to_vec() - } - } - - /// Compute similarity between two vectors. - fn similarity(&self, a: &[f32], b: &[f32]) -> f32 { - match self.metric { - DistanceMetric::Cosine => { - let a_norm = Self::normalize(a); - let b_norm = Self::normalize(b); - a_norm.iter().zip(b_norm.iter()).map(|(x, y)| x * y).sum() - } - DistanceMetric::DotProduct => { - a.iter().zip(b.iter()).map(|(x, y)| x * y).sum() - } - DistanceMetric::Euclidean => { - let dist_sq: f32 = a - .iter() - .zip(b.iter()) - .map(|(x, y)| (x - y).powi(2)) - .sum(); - 1.0 / (1.0 + dist_sq.sqrt()) - } - } - } -} - -#[async_trait] -impl VectorStore for RedbVectorStore { - async fn upsert(&self, embedding: &Embedding) -> Result<(), VectorError> { - if embedding.dim() != self.dimension { - return Err(VectorError::DimensionMismatch { - expected: self.dimension, - actual: embedding.dim(), - }); - } - - // Write to redb first (durable) - self.store - .put(&embedding.id, embedding) - .await - .map_err(|e| VectorError::IndexError(format!("redb put: {}", e)))?; - - // Then update in-memory index - let mut idx = self.index.write().map_err(|_| VectorError::LockPoisoned)?; - idx.insert(embedding.id.clone(), embedding.clone()); - - Ok(()) - } - - async fn search(&self, query: &[f32], k: usize) -> Result<Vec<SearchResult>, VectorError> { - if query.len() != self.dimension { - return Err(VectorError::DimensionMismatch { - expected: self.dimension, - actual: query.len(), - }); - } - - let idx = self.index.read().map_err(|_| VectorError::LockPoisoned)?; - - let mut results: Vec<SearchResult> = idx - .values() - .map(|emb| SearchResult { - id: emb.id.clone(), - score: self.similarity(query, &emb.vector), - }) - .collect(); - - results.sort_by(|a, b| b.score.partial_cmp(&a.score).unwrap_or(std::cmp::Ordering::Equal)); - results.truncate(k); - - Ok(results) - } - - async fn get(&self, id: &str) -> Result<Option<Embedding>, VectorError> { - // Read from in-memory index (fast path) - let idx = self.index.read().map_err(|_| VectorError::LockPoisoned)?; - Ok(idx.get(id).cloned()) - } - - async fn delete(&self, id: &str) -> Result<(), VectorError> { - // Delete from redb first (durable) - self.store - .delete(id) - .await - .map_err(|e| VectorError::IndexError(format!("redb delete: {}", e)))?; - - // Then remove from in-memory index - let mut idx = self.index.write().map_err(|_| VectorError::LockPoisoned)?; - idx.remove(id); - - Ok(()) - } - - fn dimension(&self) -> usize { - self.dimension - } -} - -#[cfg(test)] -mod tests { - use super::*; - - #[tokio::test] - async fn test_persistent_vector_roundtrip() { - let dir = tempfile::tempdir().expect("TODO: handle error"); - let path = dir.path().join("vector.redb"); - - // Create store and insert embeddings - { - let store = RedbVectorStore::open(&path, 3, DistanceMetric::Cosine) - .await - .expect("TODO: handle error"); - - store - .upsert(&Embedding::new("a", vec![1.0, 0.0, 0.0])) - .await - .expect("TODO: handle error"); - store - .upsert(&Embedding::new("b", vec![0.0, 1.0, 0.0])) - .await - .expect("TODO: handle error"); - store - .upsert(&Embedding::new("c", vec![0.9, 0.1, 0.0])) - .await - .expect("TODO: handle error"); - - // Verify search works - let results = store.search(&[1.0, 0.0, 0.0], 2).await.expect("TODO: handle error"); - assert_eq!(results.len(), 2); - assert_eq!(results[0].id, "a"); // Most similar to [1,0,0] - } - - // Reopen store — data should survive - { - let store = RedbVectorStore::open(&path, 3, DistanceMetric::Cosine) - .await - .expect("TODO: handle error"); - - // Verify data persisted - let a = store.get("a").await.expect("TODO: handle error"); - assert!(a.is_some()); - assert_eq!(a.expect("TODO: handle error").vector, vec![1.0, 0.0, 0.0]); - - let b = store.get("b").await.expect("TODO: handle error"); - assert!(b.is_some()); - - // Verify search still works after reload - let results = store.search(&[1.0, 0.0, 0.0], 2).await.expect("TODO: handle error"); - assert_eq!(results.len(), 2); - assert_eq!(results[0].id, "a"); - } - } - - #[tokio::test] - async fn test_persistent_vector_delete() { - let dir = tempfile::tempdir().expect("TODO: handle error"); - let path = dir.path().join("vector-del.redb"); - - { - let store = RedbVectorStore::open(&path, 3, DistanceMetric::Cosine) - .await - .expect("TODO: handle error"); - - store - .upsert(&Embedding::new("x", vec![1.0, 0.0, 0.0])) - .await - .expect("TODO: handle error"); - store.delete("x").await.expect("TODO: handle error"); - - let result: Option<Embedding> = store.get("x").await.expect("TODO: handle error"); - assert!(result.is_none()); - } - - // Reopen — deletion should persist - { - let store = RedbVectorStore::open(&path, 3, DistanceMetric::Cosine) - .await - .expect("TODO: handle error"); - let result: Option<Embedding> = store.get("x").await.expect("TODO: handle error"); - assert!(result.is_none()); - } - } -} diff --git a/verisimdb/rust-core/verisim-wal/Cargo.toml b/verisimdb/rust-core/verisim-wal/Cargo.toml deleted file mode 100644 index b79e3550..00000000 --- a/verisimdb/rust-core/verisim-wal/Cargo.toml +++ /dev/null @@ -1,22 +0,0 @@ -# SPDX-License-Identifier: MPL-2.0 - -[package] -name = "verisim-wal" -description = "Write-ahead log for VeriSimDB crash recovery" -version.workspace = true -edition.workspace = true -authors.workspace = true -license.workspace = true - -[dependencies] -serde = { workspace = true } -serde_json = { workspace = true } -chrono = { workspace = true } -thiserror = { workspace = true } -tracing = { workspace = true } -uuid = { workspace = true } -crc32fast = { workspace = true } - -[dev-dependencies] -proptest = { workspace = true } -tempfile = "3.14" diff --git a/verisimdb/rust-core/verisim-wal/src/entry.rs b/verisimdb/rust-core/verisim-wal/src/entry.rs deleted file mode 100644 index b0911c51..00000000 --- a/verisimdb/rust-core/verisim-wal/src/entry.rs +++ /dev/null @@ -1,466 +0,0 @@ -// SPDX-License-Identifier: MPL-2.0 -// -// VeriSimDB Write-Ahead Log - Entry types -// Copyright (c) 2026 Jonathan D.A. Jewell (hyperpolymath) <j.d.a.jewell@open.ac.uk> -// -// Defines the WAL entry struct and its constituent enums (operation type, -// modality). Provides binary serialization/deserialization for the on-disk -// format with CRC32 integrity checking. -// -// On-disk binary format (all integers little-endian): -// [4 bytes: entry_length (u32)] -- length of everything after this field -// [4 bytes: crc32 checksum] -- CRC32 of all bytes after this field -// [8 bytes: sequence (u64)] -// [8 bytes: timestamp (i64)] -- Unix milliseconds UTC -// [1 byte: operation] -- 0=Insert, 1=Update, 2=Delete, 3=Checkpoint -// [1 byte: modality] -- 0-7 for modalities (octad), 255=All -// [4 bytes: entity_id_len (u32)] -- length of entity_id UTF-8 bytes -// [N bytes: entity_id] -// [4 bytes: payload_len (u32)] -- length of payload bytes -// [M bytes: payload] - -use chrono::{DateTime, TimeZone, Utc}; -use crc32fast::Hasher as Crc32Hasher; -use serde::{Deserialize, Serialize}; - -use crate::error::{WalError, WalResult}; - -/// Maximum allowed entry size: 64 MiB. Any entry declaring a larger size -/// is treated as corrupted. -pub const MAX_ENTRY_SIZE: u32 = 64 * 1024 * 1024; - -/// Size of the fixed-length entry header prefix (entry_length + crc32). -pub const HEADER_PREFIX_SIZE: usize = 4 + 4; - -/// Size of the fixed fields after the header prefix (sequence + timestamp + -/// operation + modality). -pub const FIXED_FIELDS_SIZE: usize = 8 + 8 + 1 + 1; - -// --------------------------------------------------------------------------- -// WalOperation -// --------------------------------------------------------------------------- - -/// The type of mutation recorded by this WAL entry. -#[derive(Debug, Clone, Copy, PartialEq, Eq, Serialize, Deserialize)] -pub enum WalOperation { - /// A new entity was inserted. - Insert = 0, - /// An existing entity was updated. - Update = 1, - /// An entity was deleted. - Delete = 2, - /// A checkpoint marker (used for recovery truncation points). - Checkpoint = 3, -} - -impl WalOperation { - /// Decode a single byte into a `WalOperation`. - pub fn from_byte(byte: u8) -> WalResult<Self> { - match byte { - 0 => Ok(Self::Insert), - 1 => Ok(Self::Update), - 2 => Ok(Self::Delete), - 3 => Ok(Self::Checkpoint), - other => Err(WalError::InvalidOperation(other)), - } - } - - /// Encode this operation as a single byte. - pub fn to_byte(self) -> u8 { - self as u8 - } -} - -// --------------------------------------------------------------------------- -// WalModality -// --------------------------------------------------------------------------- - -/// Which VeriSimDB modality this entry targets (octad: 8 modalities). -#[derive(Debug, Clone, Copy, PartialEq, Eq, Serialize, Deserialize)] -pub enum WalModality { - /// RDF / property graph store. - Graph = 0, - /// HNSW vector similarity store. - Vector = 1, - /// Tensor (ndarray/Burn) store. - Tensor = 2, - /// Semantic proof-blob store. - Semantic = 3, - /// Tantivy full-text document store. - Document = 4, - /// Temporal versioning / time-series store. - Temporal = 5, - /// Origin/lineage tracking store. - Provenance = 6, - /// Geospatial/R-tree indexing store. - Spatial = 7, - /// Applies to all modalities (used in checkpoints). - All = 255, -} - -impl WalModality { - /// Decode a single byte into a `WalModality`. - pub fn from_byte(byte: u8) -> WalResult<Self> { - match byte { - 0 => Ok(Self::Graph), - 1 => Ok(Self::Vector), - 2 => Ok(Self::Tensor), - 3 => Ok(Self::Semantic), - 4 => Ok(Self::Document), - 5 => Ok(Self::Temporal), - 6 => Ok(Self::Provenance), - 7 => Ok(Self::Spatial), - 255 => Ok(Self::All), - other => Err(WalError::InvalidModality(other)), - } - } - - /// Encode this modality as a single byte. - pub fn to_byte(self) -> u8 { - self as u8 - } -} - -// --------------------------------------------------------------------------- -// WalEntry -// --------------------------------------------------------------------------- - -/// A single entry in the write-ahead log. -/// -/// Each entry records one mutation operation against one entity in one -/// modality. Checkpoint entries use `WalModality::All` and an empty payload -/// to mark a recovery point. -#[derive(Debug, Clone, PartialEq, Serialize, Deserialize)] -pub struct WalEntry { - /// Monotonically increasing sequence number assigned by the WAL writer. - pub sequence: u64, - - /// UTC timestamp of when the entry was created. - pub timestamp: DateTime<Utc>, - - /// The mutation operation type. - pub operation: WalOperation, - - /// The target modality for this mutation. - pub modality: WalModality, - - /// The entity identifier this mutation applies to. - pub entity_id: String, - - /// Opaque payload bytes (typically JSON-encoded modality data). - pub payload: Vec<u8>, -} - -impl WalEntry { - /// Serialize this entry to the on-disk binary format. - /// - /// Returns the complete byte buffer including the length prefix and CRC. - pub fn serialize(&self) -> Vec<u8> { - // Build the inner content (everything after entry_length and crc32). - let entity_id_bytes = self.entity_id.as_bytes(); - let inner_size = FIXED_FIELDS_SIZE - + 4 - + entity_id_bytes.len() - + 4 - + self.payload.len(); - - let mut inner = Vec::with_capacity(inner_size); - - // Sequence number (u64 LE). - inner.extend_from_slice(&self.sequence.to_le_bytes()); - - // Timestamp as Unix milliseconds (i64 LE). - let timestamp_millis = self.timestamp.timestamp_millis(); - inner.extend_from_slice(×tamp_millis.to_le_bytes()); - - // Operation (1 byte). - inner.push(self.operation.to_byte()); - - // Modality (1 byte). - inner.push(self.modality.to_byte()); - - // Entity ID (length-prefixed). - inner.extend_from_slice(&(entity_id_bytes.len() as u32).to_le_bytes()); - inner.extend_from_slice(entity_id_bytes); - - // Payload (length-prefixed). - inner.extend_from_slice(&(self.payload.len() as u32).to_le_bytes()); - inner.extend_from_slice(&self.payload); - - // Compute CRC32 over the inner content. - let crc = compute_crc32(&inner); - - // Build the final buffer: [entry_length][crc32][inner...]. - let entry_length = (4 + inner.len()) as u32; // crc32 field + inner content - let mut buffer = Vec::with_capacity(4 + entry_length as usize); - buffer.extend_from_slice(&entry_length.to_le_bytes()); - buffer.extend_from_slice(&crc.to_le_bytes()); - buffer.extend_from_slice(&inner); - - buffer - } - - /// Deserialize a WAL entry from a byte slice that starts immediately - /// after the entry_length field (i.e., begins with the CRC32 bytes). - /// - /// The `entry_length` is provided separately so the caller can validate - /// size bounds before allocating. - pub fn deserialize(data: &[u8], entry_length: u32) -> WalResult<Self> { - if (data.len() as u32) < entry_length { - return Err(WalError::UnexpectedEof(data.len() as u64)); - } - - let data = &data[..entry_length as usize]; - - // First 4 bytes: stored CRC32. - if data.len() < 4 { - return Err(WalError::UnexpectedEof(0)); - } - let stored_crc = u32::from_le_bytes([data[0], data[1], data[2], data[3]]); - let inner = &data[4..]; - - // Verify CRC32 over the inner content. - let computed_crc = compute_crc32(inner); - if stored_crc != computed_crc { - // Try to extract sequence for a better error message. - let sequence = if inner.len() >= 8 { - u64::from_le_bytes(inner[0..8].try_into().expect("TODO: handle error")) - } else { - 0 - }; - return Err(WalError::CrcMismatch { - sequence, - expected: stored_crc, - actual: computed_crc, - }); - } - - // Parse the inner content. - Self::parse_inner(inner) - } - - /// Parse the inner content bytes (after CRC verification). - fn parse_inner(inner: &[u8]) -> WalResult<Self> { - let mut offset = 0; - - // Sequence (u64 LE). - if inner.len() < offset + 8 { - return Err(WalError::UnexpectedEof(offset as u64)); - } - let sequence = u64::from_le_bytes(inner[offset..offset + 8].try_into().expect("TODO: handle error")); - offset += 8; - - // Timestamp (i64 LE, Unix millis). - if inner.len() < offset + 8 { - return Err(WalError::UnexpectedEof(offset as u64)); - } - let timestamp_millis = i64::from_le_bytes(inner[offset..offset + 8].try_into().expect("TODO: handle error")); - offset += 8; - - let timestamp = Utc - .timestamp_millis_opt(timestamp_millis) - .single() - .unwrap_or_else(Utc::now); - - // Operation (1 byte). - if inner.len() < offset + 1 { - return Err(WalError::UnexpectedEof(offset as u64)); - } - let operation = WalOperation::from_byte(inner[offset])?; - offset += 1; - - // Modality (1 byte). - if inner.len() < offset + 1 { - return Err(WalError::UnexpectedEof(offset as u64)); - } - let modality = WalModality::from_byte(inner[offset])?; - offset += 1; - - // Entity ID (length-prefixed u32 LE + UTF-8 bytes). - if inner.len() < offset + 4 { - return Err(WalError::UnexpectedEof(offset as u64)); - } - let entity_id_len = - u32::from_le_bytes(inner[offset..offset + 4].try_into().expect("TODO: handle error")) as usize; - offset += 4; - - if inner.len() < offset + entity_id_len { - return Err(WalError::UnexpectedEof(offset as u64)); - } - let entity_id = String::from_utf8(inner[offset..offset + entity_id_len].to_vec())?; - offset += entity_id_len; - - // Payload (length-prefixed u32 LE + bytes). - if inner.len() < offset + 4 { - return Err(WalError::UnexpectedEof(offset as u64)); - } - let payload_len = - u32::from_le_bytes(inner[offset..offset + 4].try_into().expect("TODO: handle error")) as usize; - offset += 4; - - if inner.len() < offset + payload_len { - return Err(WalError::UnexpectedEof(offset as u64)); - } - let payload = inner[offset..offset + payload_len].to_vec(); - - Ok(Self { - sequence, - timestamp, - operation, - modality, - entity_id, - payload, - }) - } -} - -/// Compute a CRC32 checksum over the given byte slice using the IEEE -/// polynomial (same as zlib/gzip). -pub fn compute_crc32(data: &[u8]) -> u32 { - let mut hasher = Crc32Hasher::new(); - hasher.update(data); - hasher.finalize() -} - -#[cfg(test)] -mod tests { - use super::*; - - /// Helper: create a sample entry for testing. - fn sample_entry(seq: u64) -> WalEntry { - WalEntry { - sequence: seq, - timestamp: Utc::now(), - operation: WalOperation::Insert, - modality: WalModality::Graph, - entity_id: format!("entity-{seq}"), - payload: serde_json::to_vec(&serde_json::json!({ - "type": "test", - "value": seq - })) - .expect("TODO: handle error"), - } - } - - #[test] - fn test_roundtrip_serialize_deserialize() { - let entry = sample_entry(1); - let bytes = entry.serialize(); - - // Read entry_length from the first 4 bytes. - let entry_length = u32::from_le_bytes(bytes[0..4].try_into().expect("TODO: handle error")); - - // Deserialize from the bytes after the length prefix. - let recovered = WalEntry::deserialize(&bytes[4..], entry_length).expect("TODO: handle error"); - - assert_eq!(entry.sequence, recovered.sequence); - assert_eq!(entry.operation, recovered.operation); - assert_eq!(entry.modality, recovered.modality); - assert_eq!(entry.entity_id, recovered.entity_id); - assert_eq!(entry.payload, recovered.payload); - // Timestamps may lose sub-millisecond precision; compare millis. - assert_eq!( - entry.timestamp.timestamp_millis(), - recovered.timestamp.timestamp_millis() - ); - } - - #[test] - fn test_crc_mismatch_detection() { - let entry = sample_entry(42); - let mut bytes = entry.serialize(); - - // Tamper with one byte in the payload area (after the 8-byte header prefix). - let tamper_offset = bytes.len() - 1; - bytes[tamper_offset] ^= 0xFF; - - let entry_length = u32::from_le_bytes(bytes[0..4].try_into().expect("TODO: handle error")); - let result = WalEntry::deserialize(&bytes[4..], entry_length); - - assert!(result.is_err()); - match result.unwrap_err() { - WalError::CrcMismatch { - sequence, - expected, - actual, - } => { - assert_eq!(sequence, 42); - assert_ne!(expected, actual); - } - other => panic!("Expected CrcMismatch, got: {other:?}"), - } - } - - #[test] - fn test_all_operations_roundtrip() { - for op in [ - WalOperation::Insert, - WalOperation::Update, - WalOperation::Delete, - WalOperation::Checkpoint, - ] { - assert_eq!(WalOperation::from_byte(op.to_byte()).expect("TODO: handle error"), op); - } - } - - #[test] - fn test_all_modalities_roundtrip() { - for modality in [ - WalModality::Graph, - WalModality::Vector, - WalModality::Tensor, - WalModality::Semantic, - WalModality::Document, - WalModality::Temporal, - WalModality::Provenance, - WalModality::Spatial, - WalModality::All, - ] { - assert_eq!(WalModality::from_byte(modality.to_byte()).expect("TODO: handle error"), modality); - } - } - - #[test] - fn test_invalid_operation_byte() { - assert!(WalOperation::from_byte(99).is_err()); - } - - #[test] - fn test_invalid_modality_byte() { - assert!(WalModality::from_byte(128).is_err()); - } - - #[test] - fn test_empty_payload_and_entity_id() { - let entry = WalEntry { - sequence: 0, - timestamp: Utc::now(), - operation: WalOperation::Checkpoint, - modality: WalModality::All, - entity_id: String::new(), - payload: Vec::new(), - }; - let bytes = entry.serialize(); - let entry_length = u32::from_le_bytes(bytes[0..4].try_into().expect("TODO: handle error")); - let recovered = WalEntry::deserialize(&bytes[4..], entry_length).expect("TODO: handle error"); - assert_eq!(recovered.entity_id, ""); - assert!(recovered.payload.is_empty()); - } - - #[test] - fn test_large_entity_id() { - let long_id = "x".repeat(10_000); - let entry = WalEntry { - sequence: 99, - timestamp: Utc::now(), - operation: WalOperation::Insert, - modality: WalModality::Document, - entity_id: long_id.clone(), - payload: vec![1, 2, 3], - }; - let bytes = entry.serialize(); - let entry_length = u32::from_le_bytes(bytes[0..4].try_into().expect("TODO: handle error")); - let recovered = WalEntry::deserialize(&bytes[4..], entry_length).expect("TODO: handle error"); - assert_eq!(recovered.entity_id, long_id); - } -} diff --git a/verisimdb/rust-core/verisim-wal/src/error.rs b/verisimdb/rust-core/verisim-wal/src/error.rs deleted file mode 100644 index 50b7a39c..00000000 --- a/verisimdb/rust-core/verisim-wal/src/error.rs +++ /dev/null @@ -1,116 +0,0 @@ -// SPDX-License-Identifier: MPL-2.0 -// -// VeriSimDB Write-Ahead Log - Error types -// Copyright (c) 2026 Jonathan D.A. Jewell (hyperpolymath) <j.d.a.jewell@open.ac.uk> -// -// Defines all error conditions that can arise during WAL operations including -// I/O failures, data corruption, and invalid state transitions. - -use thiserror::Error; - -/// Errors that can occur during WAL operations. -#[derive(Debug, Error)] -pub enum WalError { - /// An I/O error occurred while reading or writing a WAL segment file. - #[error("WAL I/O error: {0}")] - Io(#[from] std::io::Error), - - /// CRC32 checksum mismatch detected during entry validation. - /// This indicates data corruption, either from disk failure or - /// an incomplete write (crash mid-flush). - #[error("CRC mismatch at sequence {sequence}: expected {expected:#010x}, got {actual:#010x}")] - CrcMismatch { - /// The sequence number of the corrupted entry. - sequence: u64, - /// The CRC32 value stored in the entry header. - expected: u32, - /// The CRC32 value computed from the entry payload. - actual: u32, - }, - - /// The entry header declares a length that exceeds the maximum allowed - /// entry size, indicating corruption or a malformed write. - #[error("Entry at sequence {sequence} declares length {length} bytes, exceeding maximum {max_length}")] - EntryTooLarge { - /// The sequence number (if recoverable from the header). - sequence: u64, - /// The declared length in the entry header. - length: u32, - /// The maximum allowed entry length. - max_length: u32, - }, - - /// An invalid operation byte was encountered while deserializing an entry. - #[error("Invalid operation byte: {0}")] - InvalidOperation(u8), - - /// An invalid modality byte was encountered while deserializing an entry. - #[error("Invalid modality byte: {0}")] - InvalidModality(u8), - - /// The WAL segment file is truncated or contains an incomplete entry. - /// This typically happens when a crash occurs mid-write. - #[error("Truncated entry at offset {offset} in segment {segment}")] - TruncatedEntry { - /// The byte offset where the truncation was detected. - offset: u64, - /// The segment file path or identifier. - segment: String, - }, - - /// JSON serialization or deserialization failed for an entry payload. - #[error("JSON error in WAL payload: {0}")] - Json(#[from] serde_json::Error), - - /// UTF-8 decoding failed for the entity ID string. - #[error("Invalid UTF-8 in entity ID: {0}")] - InvalidEntityId(#[from] std::string::FromUtf8Error), - - /// The WAL directory does not exist or is not accessible. - #[error("WAL directory not found or inaccessible: {0}")] - DirectoryNotFound(String), - - /// Attempted to read past the end of a segment file. - #[error("Unexpected end of segment at offset {0}")] - UnexpectedEof(u64), -} - -/// Convenience type alias for WAL results. -pub type WalResult<T> = Result<T, WalError>; - -#[cfg(test)] -mod tests { - use super::*; - - #[test] - fn test_error_display_crc_mismatch() { - let error = WalError::CrcMismatch { - sequence: 42, - expected: 0xDEADBEEF, - actual: 0xCAFEBABE, - }; - let message = format!("{error}"); - assert!(message.contains("42")); - assert!(message.contains("0xdeadbeef")); - assert!(message.contains("0xcafebabe")); - } - - #[test] - fn test_error_display_io() { - let io_error = std::io::Error::new(std::io::ErrorKind::NotFound, "file gone"); - let error = WalError::Io(io_error); - let message = format!("{error}"); - assert!(message.contains("file gone")); - } - - #[test] - fn test_error_display_entry_too_large() { - let error = WalError::EntryTooLarge { - sequence: 7, - length: 999_999_999, - max_length: 67_108_864, - }; - let message = format!("{error}"); - assert!(message.contains("999999999")); - } -} diff --git a/verisimdb/rust-core/verisim-wal/src/lib.rs b/verisimdb/rust-core/verisim-wal/src/lib.rs deleted file mode 100644 index 78e98d0a..00000000 --- a/verisimdb/rust-core/verisim-wal/src/lib.rs +++ /dev/null @@ -1,75 +0,0 @@ -// SPDX-License-Identifier: MPL-2.0 -// -// VeriSimDB Write-Ahead Log (WAL) crate -// Copyright (c) 2026 Jonathan D.A. Jewell (hyperpolymath) <j.d.a.jewell@open.ac.uk> -// -// Provides crash recovery for VeriSimDB by recording all mutations to an -// append-only log before they are applied to the modality stores. On crash, -// the WAL is replayed from the last checkpoint to bring the database back -// to a consistent state. -// -// # Architecture -// -// The WAL is organized as a sequence of **segment files** in a dedicated -// directory. Each segment is an append-only binary file containing -// length-prefixed, CRC32-protected entries. Segments are rotated when they -// exceed a configurable maximum size (default 64 MiB). -// -// ## On-disk entry format (all integers little-endian) -// -// ```text -// [4 bytes: entry_length (u32)] -- length of everything after this field -// [4 bytes: crc32 checksum] -- CRC32 of all bytes after this field -// [8 bytes: sequence (u64)] -// [8 bytes: timestamp (i64)] -- Unix milliseconds UTC -// [1 byte: operation] -- 0=Insert, 1=Update, 2=Delete, 3=Checkpoint -// [1 byte: modality] -- 0-7 for modalities (octad), 255=All -// [4 bytes: entity_id_len (u32)] -- length of entity_id UTF-8 bytes -// [N bytes: entity_id] -// [4 bytes: payload_len (u32)] -- length of payload bytes -// [M bytes: payload] -// ``` -// -// ## Usage -// -// ```no_run -// use verisim_wal::{WalWriter, WalReader, WalEntry, WalOperation, WalModality, SyncMode}; -// use chrono::Utc; -// -// // Open a WAL for writing. -// let mut writer = WalWriter::open("/tmp/verisim-wal", SyncMode::Fsync).unwrap(); -// -// // Append an entry. -// let entry = WalEntry { -// sequence: 0, // assigned by the writer -// timestamp: Utc::now(), -// operation: WalOperation::Insert, -// modality: WalModality::Graph, -// entity_id: "entity-123".to_string(), -// payload: b"{}".to_vec(), -// }; -// let seq = writer.append(entry).unwrap(); -// -// // Write a checkpoint. -// writer.checkpoint().unwrap(); -// -// // Read back. -// let reader = WalReader::open("/tmp/verisim-wal").unwrap(); -// for entry in reader.replay_all().unwrap() { -// println!("seq={} op={:?} entity={}", entry.sequence, entry.operation, entry.entity_id); -// } -// ``` - -#![forbid(unsafe_code)] -pub mod entry; -pub mod error; -pub mod reader; -pub mod segment; -pub mod writer; - -// Re-export the primary public API for ergonomic imports. -pub use entry::{WalEntry, WalModality, WalOperation}; -pub use error::{WalError, WalResult}; -pub use reader::{WalEntryIterator, WalReader}; -pub use segment::{SegmentInfo, DEFAULT_MAX_SEGMENT_SIZE}; -pub use writer::{SyncMode, WalWriter}; diff --git a/verisimdb/rust-core/verisim-wal/src/reader.rs b/verisimdb/rust-core/verisim-wal/src/reader.rs deleted file mode 100644 index 2d098461..00000000 --- a/verisimdb/rust-core/verisim-wal/src/reader.rs +++ /dev/null @@ -1,594 +0,0 @@ -// SPDX-License-Identifier: MPL-2.0 -// -// VeriSimDB Write-Ahead Log - Reader for crash recovery -// Copyright (c) 2026 Jonathan D.A. Jewell (hyperpolymath) <j.d.a.jewell@open.ac.uk> -// -// The `WalReader` reads WAL segment files and replays entries in sequence -// order. It verifies CRC32 checksums on each entry and gracefully handles -// corrupted or truncated entries (which are expected after a crash). - -use std::fs; -use std::path::{Path, PathBuf}; - -use tracing::{debug, warn}; - -use crate::entry::{WalEntry, WalOperation, MAX_ENTRY_SIZE}; -use crate::error::{WalError, WalResult}; -use crate::segment::list_segments; - -// --------------------------------------------------------------------------- -// WalReader -// --------------------------------------------------------------------------- - -/// A reader that replays WAL entries from segment files on disk. -/// -/// The reader scans all segment files in the WAL directory and presents -/// their entries as a sequential stream, ordered by sequence number. -pub struct WalReader { - /// The WAL directory containing segment files. - wal_dir: PathBuf, -} - -impl WalReader { - /// Open a WAL directory for reading. - /// - /// Does not read any data until `replay_from()` or - /// `find_last_checkpoint()` is called. - pub fn open(wal_dir: impl AsRef<Path>) -> WalResult<Self> { - let wal_dir = wal_dir.as_ref().to_path_buf(); - if !wal_dir.is_dir() { - return Err(WalError::DirectoryNotFound( - wal_dir.display().to_string(), - )); - } - Ok(Self { wal_dir }) - } - - /// Return an iterator that replays all WAL entries with sequence number - /// >= `from_sequence`, across all segment files, in order. - /// - /// Corrupted entries are skipped with a warning log. Truncated entries - /// at the end of a segment (indicating a crash during write) are silently - /// ignored. - pub fn replay_from(&self, from_sequence: u64) -> WalResult<WalEntryIterator> { - let segments = list_segments(&self.wal_dir)?; - - // Collect all entries from relevant segments. - let mut all_entries: Vec<WalEntry> = Vec::new(); - - for segment in &segments { - // Skip segments that are entirely before our starting point. - // We cannot skip based on start_sequence alone because the last - // entry in a segment may have a sequence >= from_sequence even - // if the segment's start_sequence is less. - let entries = read_segment_entries(&segment.path)?; - for entry in entries { - if entry.sequence >= from_sequence { - all_entries.push(entry); - } - } - } - - // Sort by sequence number (segments should be in order, but be safe). - all_entries.sort_by_key(|e| e.sequence); - - debug!( - count = all_entries.len(), - from_sequence, - "Replaying WAL entries" - ); - - Ok(WalEntryIterator { - entries: all_entries, - position: 0, - }) - } - - /// Replay all entries from the beginning of the WAL. - pub fn replay_all(&self) -> WalResult<WalEntryIterator> { - self.replay_from(0) - } - - /// Find the sequence number of the last checkpoint entry in the WAL. - /// - /// Returns `None` if no checkpoint entries exist. This is used during - /// recovery to determine the safe starting point for replay. - pub fn find_last_checkpoint(&self) -> WalResult<Option<u64>> { - let segments = list_segments(&self.wal_dir)?; - let mut last_checkpoint: Option<u64> = None; - - // Scan segments in reverse order for efficiency (most likely to find - // the last checkpoint in the latest segments). - for segment in segments.iter().rev() { - let entries = read_segment_entries(&segment.path)?; - for entry in entries.iter().rev() { - if entry.operation == WalOperation::Checkpoint { - match last_checkpoint { - Some(existing) if entry.sequence > existing => { - last_checkpoint = Some(entry.sequence); - } - None => { - last_checkpoint = Some(entry.sequence); - } - _ => {} - } - // Once we find a checkpoint in this segment, we can - // stop scanning earlier segments (they will have lower - // sequence numbers). - return Ok(last_checkpoint); - } - } - } - - Ok(last_checkpoint) - } - - /// Count the total number of valid entries across all segments. - /// - /// Useful for diagnostics and testing. - pub fn entry_count(&self) -> WalResult<usize> { - let segments = list_segments(&self.wal_dir)?; - let mut count = 0; - for segment in &segments { - count += read_segment_entries(&segment.path)?.len(); - } - Ok(count) - } -} - -// --------------------------------------------------------------------------- -// WalEntryIterator -// --------------------------------------------------------------------------- - -/// An iterator over WAL entries, yielded in sequence order. -pub struct WalEntryIterator { - /// Pre-loaded and sorted entries. - entries: Vec<WalEntry>, - /// Current position in the entries vector. - position: usize, -} - -impl Iterator for WalEntryIterator { - type Item = WalEntry; - - fn next(&mut self) -> Option<Self::Item> { - if self.position < self.entries.len() { - let entry = self.entries[self.position].clone(); - self.position += 1; - Some(entry) - } else { - None - } - } - - fn size_hint(&self) -> (usize, Option<usize>) { - let remaining = self.entries.len() - self.position; - (remaining, Some(remaining)) - } -} - -impl ExactSizeIterator for WalEntryIterator {} - -// --------------------------------------------------------------------------- -// Segment reading helpers -// --------------------------------------------------------------------------- - -/// Read all valid entries from a single segment file. -/// -/// Corrupted entries (CRC mismatch) are logged and skipped. Truncated -/// entries at the end of the file are silently ignored (they indicate a -/// crash during write). -fn read_segment_entries(path: &Path) -> WalResult<Vec<WalEntry>> { - let data = fs::read(path)?; - let mut entries = Vec::new(); - let mut offset = 0usize; - let segment_name = path - .file_name() - .map(|n| n.to_string_lossy().to_string()) - .unwrap_or_else(|| "<unknown>".to_string()); - - while offset + 4 <= data.len() { - // Read entry_length (u32 LE). - let entry_length = u32::from_le_bytes( - data[offset..offset + 4] - .try_into() - .map_err(|_| WalError::TruncatedEntry { - segment: segment_name.clone(), - offset: offset as u64, - })?, - ); - - // Validate entry_length. - if entry_length == 0 { - // Zero-length entry is a padding sentinel; stop reading. - break; - } - - if entry_length > MAX_ENTRY_SIZE { - warn!( - offset, - entry_length, - segment = %segment_name, - "Entry declares unreasonable length, stopping segment read" - ); - break; - } - - // Check if the full entry fits in the remaining data. - let entry_end = offset + 4 + entry_length as usize; - if entry_end > data.len() { - // Truncated entry at end of segment (crash during write). - debug!( - offset, - entry_length, - available = data.len() - offset - 4, - segment = %segment_name, - "Truncated entry at end of segment (expected after crash)" - ); - break; - } - - // Try to deserialize the entry. - let entry_data = &data[offset + 4..entry_end]; - match WalEntry::deserialize(entry_data, entry_length) { - Ok(entry) => { - entries.push(entry); - } - Err(WalError::CrcMismatch { - sequence, - expected, - actual, - }) => { - warn!( - sequence, - expected = format!("{expected:#010x}"), - actual = format!("{actual:#010x}"), - offset, - segment = %segment_name, - "Skipping corrupted WAL entry (CRC mismatch)" - ); - // Continue to next entry. - } - Err(other) => { - warn!( - error = %other, - offset, - segment = %segment_name, - "Skipping unreadable WAL entry" - ); - } - } - - offset = entry_end; - } - - Ok(entries) -} - -#[cfg(test)] -mod tests { - use super::*; - use crate::entry::{WalModality, WalOperation}; - use crate::writer::{SyncMode, WalWriter}; - use tempfile::TempDir; - - /// Helper: create a test WAL entry. - fn test_entry(entity_id: &str, modality: WalModality) -> WalEntry { - WalEntry { - sequence: 0, - timestamp: chrono::Utc::now(), - operation: WalOperation::Insert, - modality, - entity_id: entity_id.to_string(), - payload: serde_json::to_vec(&serde_json::json!({"test": true})).expect("TODO: handle error"), - } - } - - #[test] - fn test_write_and_read_back() { - let dir = TempDir::new().expect("TODO: handle error"); - - // Write entries. - { - let mut writer = WalWriter::open(dir.path(), SyncMode::Fsync).expect("TODO: handle error"); - writer - .append(test_entry("entity-1", WalModality::Graph)) - .expect("TODO: handle error"); - writer - .append(test_entry("entity-2", WalModality::Vector)) - .expect("TODO: handle error"); - writer - .append(test_entry("entity-3", WalModality::Tensor)) - .expect("TODO: handle error"); - } - - // Read entries. - let reader = WalReader::open(dir.path()).expect("TODO: handle error"); - let entries: Vec<WalEntry> = reader.replay_all().expect("TODO: handle error").collect(); - - assert_eq!(entries.len(), 3); - assert_eq!(entries[0].sequence, 1); - assert_eq!(entries[0].entity_id, "entity-1"); - assert_eq!(entries[0].modality, WalModality::Graph); - assert_eq!(entries[1].sequence, 2); - assert_eq!(entries[1].entity_id, "entity-2"); - assert_eq!(entries[2].sequence, 3); - assert_eq!(entries[2].entity_id, "entity-3"); - } - - #[test] - fn test_replay_from_sequence() { - let dir = TempDir::new().expect("TODO: handle error"); - - { - let mut writer = WalWriter::open(dir.path(), SyncMode::Fsync).expect("TODO: handle error"); - for i in 0..10 { - writer - .append(test_entry(&format!("e-{i}"), WalModality::Document)) - .expect("TODO: handle error"); - } - } - - let reader = WalReader::open(dir.path()).expect("TODO: handle error"); - - // Replay from sequence 5 onward. - let entries: Vec<WalEntry> = reader.replay_from(5).expect("TODO: handle error").collect(); - assert_eq!(entries.len(), 6); // sequences 5, 6, 7, 8, 9, 10 - assert_eq!(entries[0].sequence, 5); - assert_eq!(entries[5].sequence, 10); - } - - #[test] - fn test_find_last_checkpoint() { - let dir = TempDir::new().expect("TODO: handle error"); - - { - let mut writer = WalWriter::open(dir.path(), SyncMode::Fsync).expect("TODO: handle error"); - writer - .append(test_entry("e-1", WalModality::Graph)) - .expect("TODO: handle error"); - writer - .append(test_entry("e-2", WalModality::Vector)) - .expect("TODO: handle error"); - let cp1 = writer.checkpoint().expect("TODO: handle error"); - assert_eq!(cp1, 3); - - writer - .append(test_entry("e-3", WalModality::Tensor)) - .expect("TODO: handle error"); - writer - .append(test_entry("e-4", WalModality::Semantic)) - .expect("TODO: handle error"); - let cp2 = writer.checkpoint().expect("TODO: handle error"); - assert_eq!(cp2, 6); - - writer - .append(test_entry("e-5", WalModality::Document)) - .expect("TODO: handle error"); - } - - let reader = WalReader::open(dir.path()).expect("TODO: handle error"); - let last_cp = reader.find_last_checkpoint().expect("TODO: handle error"); - assert_eq!(last_cp, Some(6)); - } - - #[test] - fn test_no_checkpoint_returns_none() { - let dir = TempDir::new().expect("TODO: handle error"); - - { - let mut writer = WalWriter::open(dir.path(), SyncMode::Fsync).expect("TODO: handle error"); - writer - .append(test_entry("e-1", WalModality::Graph)) - .expect("TODO: handle error"); - } - - let reader = WalReader::open(dir.path()).expect("TODO: handle error"); - assert_eq!(reader.find_last_checkpoint().expect("TODO: handle error"), None); - } - - #[test] - fn test_empty_wal_produces_empty_iterator() { - let dir = TempDir::new().expect("TODO: handle error"); - - // Create the WAL directory with a writer (creates empty segment). - { - let _writer = WalWriter::open(dir.path(), SyncMode::Async).expect("TODO: handle error"); - } - - let reader = WalReader::open(dir.path()).expect("TODO: handle error"); - let entries: Vec<WalEntry> = reader.replay_all().expect("TODO: handle error").collect(); - assert!(entries.is_empty()); - } - - #[test] - fn test_corrupted_entry_skipped() { - let dir = TempDir::new().expect("TODO: handle error"); - - // Write some entries. - { - let mut writer = WalWriter::open(dir.path(), SyncMode::Fsync).expect("TODO: handle error"); - writer - .append(test_entry("good-1", WalModality::Graph)) - .expect("TODO: handle error"); - writer - .append(test_entry("will-corrupt", WalModality::Vector)) - .expect("TODO: handle error"); - writer - .append(test_entry("good-3", WalModality::Tensor)) - .expect("TODO: handle error"); - } - - // Tamper with the second entry's CRC in the segment file. - let segments = list_segments(dir.path()).expect("TODO: handle error"); - assert_eq!(segments.len(), 1); - - let mut data = fs::read(&segments[0].path).expect("TODO: handle error"); - - // Find the second entry. The first entry starts at offset 0. - // Read the first entry's length to find the second entry's offset. - let first_len = - u32::from_le_bytes(data[0..4].try_into().expect("TODO: handle error")) as usize; - let second_entry_offset = 4 + first_len; - - // The CRC is at bytes [offset+4..offset+8] (after entry_length). - let crc_offset = second_entry_offset + 4; - data[crc_offset] ^= 0xFF; // Flip some bits in the CRC. - - fs::write(&segments[0].path, &data).expect("TODO: handle error"); - - // Read back: should get entries 1 and 3, but skip 2. - let reader = WalReader::open(dir.path()).expect("TODO: handle error"); - let entries: Vec<WalEntry> = reader.replay_all().expect("TODO: handle error").collect(); - - assert_eq!(entries.len(), 2); - assert_eq!(entries[0].entity_id, "good-1"); - assert_eq!(entries[1].entity_id, "good-3"); - } - - #[test] - fn test_multiple_modalities_in_same_wal() { - let dir = TempDir::new().expect("TODO: handle error"); - - { - let mut writer = WalWriter::open(dir.path(), SyncMode::Fsync).expect("TODO: handle error"); - writer - .append(test_entry("e-1", WalModality::Graph)) - .expect("TODO: handle error"); - writer - .append(test_entry("e-2", WalModality::Vector)) - .expect("TODO: handle error"); - writer - .append(test_entry("e-3", WalModality::Tensor)) - .expect("TODO: handle error"); - writer - .append(test_entry("e-4", WalModality::Semantic)) - .expect("TODO: handle error"); - writer - .append(test_entry("e-5", WalModality::Document)) - .expect("TODO: handle error"); - writer - .append(test_entry("e-6", WalModality::Temporal)) - .expect("TODO: handle error"); - } - - let reader = WalReader::open(dir.path()).expect("TODO: handle error"); - let entries: Vec<WalEntry> = reader.replay_all().expect("TODO: handle error").collect(); - - assert_eq!(entries.len(), 6); - assert_eq!(entries[0].modality, WalModality::Graph); - assert_eq!(entries[1].modality, WalModality::Vector); - assert_eq!(entries[2].modality, WalModality::Tensor); - assert_eq!(entries[3].modality, WalModality::Semantic); - assert_eq!(entries[4].modality, WalModality::Document); - assert_eq!(entries[5].modality, WalModality::Temporal); - } - - #[test] - fn test_checkpoint_and_replay_from_checkpoint() { - let dir = TempDir::new().expect("TODO: handle error"); - - { - let mut writer = WalWriter::open(dir.path(), SyncMode::Fsync).expect("TODO: handle error"); - - // Phase 1: some data + checkpoint. - writer - .append(test_entry("old-1", WalModality::Graph)) - .expect("TODO: handle error"); - writer - .append(test_entry("old-2", WalModality::Vector)) - .expect("TODO: handle error"); - let cp = writer.checkpoint().expect("TODO: handle error"); - assert_eq!(cp, 3); - - // Phase 2: more data after checkpoint. - writer - .append(test_entry("new-1", WalModality::Tensor)) - .expect("TODO: handle error"); - writer - .append(test_entry("new-2", WalModality::Semantic)) - .expect("TODO: handle error"); - } - - let reader = WalReader::open(dir.path()).expect("TODO: handle error"); - - // Find the checkpoint. - let cp_seq = reader.find_last_checkpoint().expect("TODO: handle error").expect("TODO: handle error"); - assert_eq!(cp_seq, 3); - - // Replay only from checkpoint onward. - let entries: Vec<WalEntry> = reader.replay_from(cp_seq + 1).expect("TODO: handle error").collect(); - assert_eq!(entries.len(), 2); - assert_eq!(entries[0].entity_id, "new-1"); - assert_eq!(entries[1].entity_id, "new-2"); - } - - #[test] - fn test_entry_count() { - let dir = TempDir::new().expect("TODO: handle error"); - - { - let mut writer = WalWriter::open(dir.path(), SyncMode::Fsync).expect("TODO: handle error"); - for _ in 0..7 { - writer - .append(test_entry("e", WalModality::Graph)) - .expect("TODO: handle error"); - } - } - - let reader = WalReader::open(dir.path()).expect("TODO: handle error"); - assert_eq!(reader.entry_count().expect("TODO: handle error"), 7); - } - - #[test] - fn test_segment_rotation_read_across_segments() { - let dir = TempDir::new().expect("TODO: handle error"); - - // Write with tiny segments to force rotation. - { - let mut writer = - WalWriter::open_with_max_size(dir.path(), SyncMode::Fsync, 100).expect("TODO: handle error"); - for i in 0..20 { - writer - .append(test_entry( - &format!("entity-{i}"), - WalModality::Graph, - )) - .expect("TODO: handle error"); - } - } - - // Verify multiple segments were created. - let segments = list_segments(dir.path()).expect("TODO: handle error"); - assert!(segments.len() > 1); - - // Read all entries across segments. - let reader = WalReader::open(dir.path()).expect("TODO: handle error"); - let entries: Vec<WalEntry> = reader.replay_all().expect("TODO: handle error").collect(); - assert_eq!(entries.len(), 20); - - // Verify sequence continuity. - for (i, entry) in entries.iter().enumerate() { - assert_eq!(entry.sequence, (i + 1) as u64); - assert_eq!(entry.entity_id, format!("entity-{i}")); - } - } - - #[test] - fn test_exact_size_iterator() { - let dir = TempDir::new().expect("TODO: handle error"); - - { - let mut writer = WalWriter::open(dir.path(), SyncMode::Fsync).expect("TODO: handle error"); - for _ in 0..5 { - writer - .append(test_entry("e", WalModality::Graph)) - .expect("TODO: handle error"); - } - } - - let reader = WalReader::open(dir.path()).expect("TODO: handle error"); - let iter = reader.replay_all().expect("TODO: handle error"); - assert_eq!(iter.len(), 5); - } -} diff --git a/verisimdb/rust-core/verisim-wal/src/segment.rs b/verisimdb/rust-core/verisim-wal/src/segment.rs deleted file mode 100644 index fd0317f6..00000000 --- a/verisimdb/rust-core/verisim-wal/src/segment.rs +++ /dev/null @@ -1,291 +0,0 @@ -// SPDX-License-Identifier: MPL-2.0 -// -// VeriSimDB Write-Ahead Log - Segment management -// Copyright (c) 2026 Jonathan D.A. Jewell (hyperpolymath) <j.d.a.jewell@open.ac.uk> -// -// Each WAL segment is a single append-only file named `wal-{sequence:016}.log`. -// Segments are rotated when they exceed the configured maximum size. Old -// segments can be pruned after a checkpoint confirms all their entries have -// been durably applied to the modality stores. - -use std::fs; -use std::path::{Path, PathBuf}; - -use tracing::debug; - -use crate::error::{WalError, WalResult}; - -/// Default maximum segment size in bytes (64 MiB). -pub const DEFAULT_MAX_SEGMENT_SIZE: u64 = 64 * 1024 * 1024; - -/// The file extension used for WAL segment files. -pub const SEGMENT_EXTENSION: &str = "log"; - -/// The prefix used for WAL segment file names. -pub const SEGMENT_PREFIX: &str = "wal-"; - -/// Metadata about a single WAL segment file. -#[derive(Debug, Clone, PartialEq, Eq)] -pub struct SegmentInfo { - /// The full path to the segment file on disk. - pub path: PathBuf, - - /// The starting sequence number encoded in the file name. - /// All entries in this segment have sequence >= this value. - pub start_sequence: u64, - - /// Current file size in bytes. - pub file_size: u64, -} - -impl SegmentInfo { - /// Returns `true` if the segment file has reached or exceeded the given - /// maximum size in bytes. - pub fn is_full(&self, max_size: u64) -> bool { - self.file_size >= max_size - } -} - -impl PartialOrd for SegmentInfo { - fn partial_cmp(&self, other: &Self) -> Option<std::cmp::Ordering> { - Some(self.cmp(other)) - } -} - -impl Ord for SegmentInfo { - fn cmp(&self, other: &Self) -> std::cmp::Ordering { - self.start_sequence.cmp(&other.start_sequence) - } -} - -/// Build the canonical file name for a segment starting at the given -/// sequence number. -/// -/// Format: `wal-0000000000000001.log` -pub fn segment_filename(start_sequence: u64) -> String { - format!("{SEGMENT_PREFIX}{start_sequence:016}.{SEGMENT_EXTENSION}") -} - -/// Build the full path for a segment file in the given WAL directory. -pub fn segment_path(wal_dir: &Path, start_sequence: u64) -> PathBuf { - wal_dir.join(segment_filename(start_sequence)) -} - -/// Parse the starting sequence number from a segment file name. -/// -/// Returns `None` if the name does not match the expected pattern. -pub fn parse_segment_filename(name: &str) -> Option<u64> { - let stripped = name.strip_prefix(SEGMENT_PREFIX)?; - let num_str = stripped.strip_suffix(&format!(".{SEGMENT_EXTENSION}"))?; - num_str.parse::<u64>().ok() -} - -/// Scan a WAL directory and return metadata for all segment files, sorted -/// by starting sequence number (ascending). -/// -/// Non-segment files in the directory are silently ignored. -pub fn list_segments(wal_dir: &Path) -> WalResult<Vec<SegmentInfo>> { - if !wal_dir.is_dir() { - return Err(WalError::DirectoryNotFound( - wal_dir.display().to_string(), - )); - } - - let mut segments = Vec::new(); - - for dir_entry in fs::read_dir(wal_dir)? { - let dir_entry = dir_entry?; - let file_name = dir_entry.file_name(); - let name = file_name.to_string_lossy(); - - if let Some(start_sequence) = parse_segment_filename(&name) { - let metadata = dir_entry.metadata()?; - segments.push(SegmentInfo { - path: dir_entry.path(), - start_sequence, - file_size: metadata.len(), - }); - } - } - - segments.sort(); - - debug!( - count = segments.len(), - dir = %wal_dir.display(), - "Discovered WAL segments" - ); - - Ok(segments) -} - -/// Remove segment files whose starting sequence is strictly less than the -/// given checkpoint sequence. These segments are safe to delete because all -/// their entries have been durably applied. -/// -/// Returns the number of segments removed. -pub fn prune_segments_before(wal_dir: &Path, checkpoint_sequence: u64) -> WalResult<usize> { - let segments = list_segments(wal_dir)?; - let mut removed = 0; - - for segment in &segments { - // Only remove segments that are entirely before the checkpoint. - // A segment starting at sequence N may contain entries up to the - // start of the next segment, so we need to check against the next - // segment's start_sequence. For safety, we only remove segments - // whose start_sequence is strictly less than the checkpoint AND - // there exists a later segment (so we never remove the only segment). - if segment.start_sequence < checkpoint_sequence { - // Check if there is a later segment that covers the checkpoint. - let has_later_segment = segments - .iter() - .any(|s| s.start_sequence >= checkpoint_sequence); - - if has_later_segment { - debug!( - path = %segment.path.display(), - start_sequence = segment.start_sequence, - "Pruning WAL segment (before checkpoint {checkpoint_sequence})" - ); - fs::remove_file(&segment.path)?; - removed += 1; - } - } - } - - Ok(removed) -} - -#[cfg(test)] -mod tests { - use super::*; - use std::fs::File; - use std::io::Write; - use tempfile::TempDir; - - /// Helper to create a temporary WAL directory. We use a local helper - /// instead of depending on tempfile in the library crate. - struct TestDir { - _inner: TempDir, - path: PathBuf, - } - - impl TestDir { - fn new() -> Self { - let inner = TempDir::new().expect("TODO: handle error"); - let path = inner.path().to_path_buf(); - Self { - _inner: inner, - path, - } - } - - fn create_segment(&self, start_seq: u64, size_bytes: usize) { - let file_path = segment_path(&self.path, start_seq); - let mut file = File::create(file_path).expect("TODO: handle error"); - file.write_all(&vec![0u8; size_bytes]).expect("TODO: handle error"); - } - } - - #[test] - fn test_segment_filename_format() { - assert_eq!(segment_filename(0), "wal-0000000000000000.log"); - assert_eq!(segment_filename(1), "wal-0000000000000001.log"); - assert_eq!( - segment_filename(9_999_999_999_999_999), - "wal-9999999999999999.log" - ); - } - - #[test] - fn test_parse_segment_filename_valid() { - assert_eq!( - parse_segment_filename("wal-0000000000000042.log"), - Some(42) - ); - assert_eq!( - parse_segment_filename("wal-0000000000000000.log"), - Some(0) - ); - } - - #[test] - fn test_parse_segment_filename_invalid() { - assert_eq!(parse_segment_filename("not-a-segment.txt"), None); - assert_eq!(parse_segment_filename("wal-.log"), None); - assert_eq!(parse_segment_filename("wal-abc.log"), None); - assert_eq!(parse_segment_filename(""), None); - } - - #[test] - fn test_list_segments_sorted() { - let dir = TestDir::new(); - dir.create_segment(100, 1024); - dir.create_segment(1, 512); - dir.create_segment(50, 2048); - - // Create a non-segment file that should be ignored. - File::create(dir.path.join("readme.txt")).expect("TODO: handle error"); - - let segments = list_segments(&dir.path).expect("TODO: handle error"); - assert_eq!(segments.len(), 3); - assert_eq!(segments[0].start_sequence, 1); - assert_eq!(segments[0].file_size, 512); - assert_eq!(segments[1].start_sequence, 50); - assert_eq!(segments[2].start_sequence, 100); - } - - #[test] - fn test_list_segments_empty_dir() { - let dir = TestDir::new(); - let segments = list_segments(&dir.path).expect("TODO: handle error"); - assert!(segments.is_empty()); - } - - #[test] - fn test_list_segments_nonexistent_dir() { - let result = list_segments(Path::new("/nonexistent/wal/dir")); - assert!(result.is_err()); - } - - #[test] - fn test_segment_info_is_full() { - let info = SegmentInfo { - path: PathBuf::from("test.log"), - start_sequence: 0, - file_size: DEFAULT_MAX_SEGMENT_SIZE, - }; - assert!(info.is_full(DEFAULT_MAX_SEGMENT_SIZE)); - assert!(!info.is_full(DEFAULT_MAX_SEGMENT_SIZE + 1)); - } - - #[test] - fn test_prune_segments_before_checkpoint() { - let dir = TestDir::new(); - dir.create_segment(1, 100); - dir.create_segment(50, 100); - dir.create_segment(100, 100); - - // Prune everything before sequence 100. - let removed = prune_segments_before(&dir.path, 100).expect("TODO: handle error"); - assert_eq!(removed, 2); - - let remaining = list_segments(&dir.path).expect("TODO: handle error"); - assert_eq!(remaining.len(), 1); - assert_eq!(remaining[0].start_sequence, 100); - } - - #[test] - fn test_prune_does_not_remove_only_segment() { - let dir = TestDir::new(); - dir.create_segment(1, 100); - - // Even though sequence 1 < checkpoint 999, we should not remove it - // because there is no later segment covering the checkpoint. - let removed = prune_segments_before(&dir.path, 999).expect("TODO: handle error"); - assert_eq!(removed, 0); - - let remaining = list_segments(&dir.path).expect("TODO: handle error"); - assert_eq!(remaining.len(), 1); - } -} diff --git a/verisimdb/rust-core/verisim-wal/src/writer.rs b/verisimdb/rust-core/verisim-wal/src/writer.rs deleted file mode 100644 index 3de30b64..00000000 --- a/verisimdb/rust-core/verisim-wal/src/writer.rs +++ /dev/null @@ -1,472 +0,0 @@ -// SPDX-License-Identifier: MPL-2.0 -// -// VeriSimDB Write-Ahead Log - Append-only writer -// Copyright (c) 2026 Jonathan D.A. Jewell (hyperpolymath) <j.d.a.jewell@open.ac.uk> -// -// The `WalWriter` is responsible for appending entries to the current WAL -// segment file, managing segment rotation, and controlling fsync behavior -// according to the configured `SyncMode`. - -use std::fs::{self, File, OpenOptions}; -use std::io::Write; -use std::path::{Path, PathBuf}; -use std::time::{Duration, Instant}; - -use chrono::Utc; -use tracing::{debug, info}; - -use crate::entry::{WalEntry, WalModality, WalOperation}; -use crate::error::{WalError, WalResult}; -use crate::segment::{ - list_segments, segment_path, DEFAULT_MAX_SEGMENT_SIZE, SegmentInfo, -}; - -// --------------------------------------------------------------------------- -// SyncMode -// --------------------------------------------------------------------------- - -/// Controls how aggressively the WAL writer calls `fsync` to flush data to -/// stable storage. -#[derive(Debug, Clone)] -pub enum SyncMode { - /// Call `fsync` after every single `append()`. This is the safest mode - /// and guarantees that acknowledged writes survive a crash, but it is - /// also the slowest. - Fsync, - - /// Call `fsync` at most once per the specified duration. Writes between - /// syncs may be lost on crash. This is a good balance between safety - /// and throughput. - Periodic(Duration), - - /// Never explicitly call `fsync`; rely on the OS page cache to flush - /// data to disk eventually. This is the fastest mode but data loss is - /// possible on crash. - Async, -} - -// --------------------------------------------------------------------------- -// WalWriter -// --------------------------------------------------------------------------- - -/// An append-only writer for WAL segment files. -/// -/// The writer maintains a monotonically increasing sequence counter and -/// automatically rotates to a new segment file when the current one exceeds -/// `max_segment_size`. -pub struct WalWriter { - /// The directory containing all WAL segment files. - wal_dir: PathBuf, - - /// The currently open segment file handle. - current_file: File, - - /// Metadata about the current segment. - current_segment: SegmentInfo, - - /// Monotonically increasing sequence number counter. - next_sequence: u64, - - /// Maximum size (in bytes) of a single segment before rotation. - max_segment_size: u64, - - /// How fsync is managed. - sync_mode: SyncMode, - - /// Timestamp of the last fsync call (for `SyncMode::Periodic`). - last_sync: Instant, -} - -impl WalWriter { - /// Open an existing WAL directory or initialize a new one. - /// - /// If the directory already contains segment files, the writer resumes - /// from the end of the last segment (highest sequence number). If the - /// directory is empty or does not exist, a fresh segment is created - /// starting at sequence 1. - /// - /// # Arguments - /// - /// * `wal_dir` - Path to the WAL directory. Created if it does not exist. - /// * `sync_mode` - Controls fsync behavior. - pub fn open(wal_dir: impl AsRef<Path>, sync_mode: SyncMode) -> WalResult<Self> { - Self::open_with_max_size(wal_dir, sync_mode, DEFAULT_MAX_SEGMENT_SIZE) - } - - /// Open the WAL directory with a custom maximum segment size. - /// - /// This is primarily useful for testing with small segment sizes. - pub fn open_with_max_size( - wal_dir: impl AsRef<Path>, - sync_mode: SyncMode, - max_segment_size: u64, - ) -> WalResult<Self> { - let wal_dir = wal_dir.as_ref().to_path_buf(); - - // Ensure the WAL directory exists. - if !wal_dir.exists() { - fs::create_dir_all(&wal_dir)?; - info!(dir = %wal_dir.display(), "Created WAL directory"); - } - - // Discover existing segments. - let segments = list_segments(&wal_dir)?; - - let (current_segment, current_file, next_sequence) = if segments.is_empty() { - // Fresh WAL: create the first segment. - let start_seq = 0; - let path = segment_path(&wal_dir, start_seq); - let file = File::create(&path)?; - let segment = SegmentInfo { - path, - start_sequence: start_seq, - file_size: 0, - }; - info!("Initialized fresh WAL at sequence 0"); - (segment, file, 1u64) - } else { - // Resume from the last segment (safe: is_empty() checked above). - let last = segments.last().expect("segments is non-empty").clone(); - let next_seq = Self::scan_last_sequence(&last)?; - let file = OpenOptions::new().append(true).open(&last.path)?; - info!( - segment = %last.path.display(), - next_sequence = next_seq, - "Resuming WAL" - ); - (last, file, next_seq) - }; - - Ok(Self { - wal_dir, - current_file, - current_segment, - next_sequence, - max_segment_size, - sync_mode, - last_sync: Instant::now(), - }) - } - - /// Append a new entry to the WAL. - /// - /// The entry's `sequence` field is overwritten with the next sequence - /// number assigned by the writer. Returns the assigned sequence number. - pub fn append(&mut self, mut entry: WalEntry) -> WalResult<u64> { - // Assign the next sequence number. - let sequence = self.next_sequence; - entry.sequence = sequence; - self.next_sequence += 1; - - let bytes = entry.serialize(); - - // Check if we need to rotate before writing. - if self.current_segment.file_size + bytes.len() as u64 > self.max_segment_size { - self.rotate()?; - } - - self.current_file.write_all(&bytes)?; - self.current_segment.file_size += bytes.len() as u64; - - // Handle sync according to the configured mode. - self.maybe_sync()?; - - debug!(sequence, entity_id = %entry.entity_id, "Appended WAL entry"); - - Ok(sequence) - } - - /// Force an immediate `fsync` of the current segment file, regardless - /// of the configured `SyncMode`. - pub fn sync(&mut self) -> WalResult<()> { - self.current_file.sync_all()?; - self.last_sync = Instant::now(); - Ok(()) - } - - /// Write a checkpoint entry to the WAL. - /// - /// A checkpoint marks a point in the log where all preceding entries - /// have been durably applied to the modality stores. During recovery, - /// replay can start from the last checkpoint instead of the beginning. - /// - /// Returns the sequence number of the checkpoint entry. - pub fn checkpoint(&mut self) -> WalResult<u64> { - let entry = WalEntry { - sequence: 0, // Will be overwritten by append(). - timestamp: Utc::now(), - operation: WalOperation::Checkpoint, - modality: WalModality::All, - entity_id: String::new(), - payload: Vec::new(), - }; - - let sequence = self.append(entry)?; - - // Always fsync after a checkpoint for crash safety. - self.sync()?; - - info!(sequence, "WAL checkpoint written"); - - Ok(sequence) - } - - /// Rotate to a new segment file. - /// - /// The current segment is fsynced and closed, and a new segment file is - /// created starting at the current `next_sequence` value. - pub fn rotate(&mut self) -> WalResult<()> { - // Sync the current segment before closing. - self.sync()?; - - let new_start = self.next_sequence; - let new_path = segment_path(&self.wal_dir, new_start); - let new_file = File::create(&new_path)?; - - info!( - old_segment = %self.current_segment.path.display(), - new_segment = %new_path.display(), - start_sequence = new_start, - "Rotated WAL segment" - ); - - self.current_file = new_file; - self.current_segment = SegmentInfo { - path: new_path, - start_sequence: new_start, - file_size: 0, - }; - - Ok(()) - } - - /// Returns the sequence number that will be assigned to the next entry. - pub fn next_sequence(&self) -> u64 { - self.next_sequence - } - - /// Returns the path to the WAL directory. - pub fn wal_dir(&self) -> &Path { - &self.wal_dir - } - - /// Returns a reference to the current segment's metadata. - pub fn current_segment(&self) -> &SegmentInfo { - &self.current_segment - } - - // ----------------------------------------------------------------------- - // Private helpers - // ----------------------------------------------------------------------- - - /// Conditionally call fsync based on the configured sync mode. - fn maybe_sync(&mut self) -> WalResult<()> { - match &self.sync_mode { - SyncMode::Fsync => { - self.current_file.sync_all()?; - self.last_sync = Instant::now(); - } - SyncMode::Periodic(interval) => { - if self.last_sync.elapsed() >= *interval { - self.current_file.sync_all()?; - self.last_sync = Instant::now(); - } - } - SyncMode::Async => { - // No-op: rely on OS page cache. - } - } - Ok(()) - } - - /// Scan the last segment file to determine the next sequence number. - /// - /// Reads through all valid entries in the segment and returns the - /// sequence number one past the last valid entry. If the segment is - /// empty, returns `start_sequence + 1`. - fn scan_last_sequence(segment: &SegmentInfo) -> WalResult<u64> { - if segment.file_size == 0 { - return Ok(segment.start_sequence + 1); - } - - let data = fs::read(&segment.path)?; - let mut offset = 0usize; - let mut last_sequence = segment.start_sequence; - - let segment_name = segment - .path - .file_name() - .map(|n| n.to_string_lossy().to_string()) - .unwrap_or_else(|| "<unknown>".to_string()); - - while offset + 4 <= data.len() { - let entry_length = u32::from_le_bytes( - data[offset..offset + 4] - .try_into() - .map_err(|_| WalError::TruncatedEntry { - segment: segment_name.clone(), - offset: offset as u64, - })?, - ); - - // Sanity check: the entry must fit within the remaining data. - if offset + 4 + entry_length as usize > data.len() { - // Truncated entry at end of file (crash during write). - break; - } - - // Try to read just the sequence number from the inner content. - // Layout: [4 bytes crc][8 bytes sequence][...] - let inner_start = offset + 4 + 4; // skip entry_length + crc - if inner_start + 8 <= data.len() { - let seq = u64::from_le_bytes( - data[inner_start..inner_start + 8] - .try_into() - .map_err(|_| WalError::TruncatedEntry { - segment: segment_name.clone(), - offset: inner_start as u64, - })?, - ); - if seq >= last_sequence { - last_sequence = seq; - } - } - - offset += 4 + entry_length as usize; - } - - Ok(last_sequence + 1) - } -} - -#[cfg(test)] -mod tests { - use super::*; - use crate::entry::{WalModality, WalOperation}; - use tempfile::TempDir; - - /// Helper: create a test WAL entry. - fn test_entry(modality: WalModality) -> WalEntry { - WalEntry { - sequence: 0, - timestamp: Utc::now(), - operation: WalOperation::Insert, - modality, - entity_id: "test-entity".to_string(), - payload: b"{}".to_vec(), - } - } - - #[test] - fn test_open_fresh_directory() { - let dir = TempDir::new().expect("TODO: handle error"); - let writer = WalWriter::open(dir.path(), SyncMode::Async).expect("TODO: handle error"); - assert_eq!(writer.next_sequence(), 1); - } - - #[test] - fn test_append_increments_sequence() { - let dir = TempDir::new().expect("TODO: handle error"); - let mut writer = WalWriter::open(dir.path(), SyncMode::Async).expect("TODO: handle error"); - - let seq1 = writer.append(test_entry(WalModality::Graph)).expect("TODO: handle error"); - let seq2 = writer.append(test_entry(WalModality::Vector)).expect("TODO: handle error"); - let seq3 = writer.append(test_entry(WalModality::Tensor)).expect("TODO: handle error"); - - assert_eq!(seq1, 1); - assert_eq!(seq2, 2); - assert_eq!(seq3, 3); - assert_eq!(writer.next_sequence(), 4); - } - - #[test] - fn test_checkpoint_writes_entry() { - let dir = TempDir::new().expect("TODO: handle error"); - let mut writer = WalWriter::open(dir.path(), SyncMode::Async).expect("TODO: handle error"); - - writer.append(test_entry(WalModality::Graph)).expect("TODO: handle error"); - writer.append(test_entry(WalModality::Vector)).expect("TODO: handle error"); - let cp_seq = writer.checkpoint().expect("TODO: handle error"); - - assert_eq!(cp_seq, 3); - assert_eq!(writer.next_sequence(), 4); - } - - #[test] - fn test_segment_rotation() { - let dir = TempDir::new().expect("TODO: handle error"); - // Use a tiny max segment size to force rotation. - let mut writer = - WalWriter::open_with_max_size(dir.path(), SyncMode::Async, 100).expect("TODO: handle error"); - - // Write entries until rotation occurs. - for _ in 0..10 { - writer.append(test_entry(WalModality::Document)).expect("TODO: handle error"); - } - - let segments = list_segments(dir.path()).expect("TODO: handle error"); - assert!( - segments.len() > 1, - "Expected multiple segments after rotation, got {}", - segments.len() - ); - } - - #[test] - fn test_resume_after_close() { - let dir = TempDir::new().expect("TODO: handle error"); - - // Write some entries. - { - let mut writer = WalWriter::open(dir.path(), SyncMode::Fsync).expect("TODO: handle error"); - writer.append(test_entry(WalModality::Graph)).expect("TODO: handle error"); - writer.append(test_entry(WalModality::Vector)).expect("TODO: handle error"); - writer.append(test_entry(WalModality::Tensor)).expect("TODO: handle error"); - } - - // Re-open and verify sequence continues. - { - let mut writer = WalWriter::open(dir.path(), SyncMode::Fsync).expect("TODO: handle error"); - let seq = writer.append(test_entry(WalModality::Semantic)).expect("TODO: handle error"); - assert_eq!(seq, 4, "Expected sequence 4 after resuming, got {seq}"); - } - } - - #[test] - fn test_open_creates_directory() { - let dir = TempDir::new().expect("TODO: handle error"); - let wal_path = dir.path().join("subdir").join("wal"); - assert!(!wal_path.exists()); - - let _writer = WalWriter::open(&wal_path, SyncMode::Async).expect("TODO: handle error"); - assert!(wal_path.exists()); - } - - #[test] - fn test_fsync_mode() { - let dir = TempDir::new().expect("TODO: handle error"); - let mut writer = WalWriter::open(dir.path(), SyncMode::Fsync).expect("TODO: handle error"); - - // Should not panic or error even with fsync on every write. - for _ in 0..5 { - writer.append(test_entry(WalModality::Temporal)).expect("TODO: handle error"); - } - } - - #[test] - fn test_periodic_sync_mode() { - let dir = TempDir::new().expect("TODO: handle error"); - let mut writer = WalWriter::open( - dir.path(), - SyncMode::Periodic(Duration::from_millis(10)), - ) - .expect("TODO: handle error"); - - for _ in 0..5 { - writer.append(test_entry(WalModality::Temporal)).expect("TODO: handle error"); - } - - // Explicit sync should always work. - writer.sync().expect("TODO: handle error"); - } -} diff --git a/verisimdb/scripts/post-commit-hook.sh b/verisimdb/scripts/post-commit-hook.sh deleted file mode 100755 index 0132334b..00000000 --- a/verisimdb/scripts/post-commit-hook.sh +++ /dev/null @@ -1,8 +0,0 @@ -#!/usr/bin/env bash -# SPDX-License-Identifier: MPL-2.0 -# Post-commit hook: auto-ingest the latest commit into .verisimdb/ -# -# Install: cp scripts/post-commit-hook.sh .git/hooks/post-commit - -REPO_ROOT="$(git rev-parse --show-toplevel)" -"${REPO_ROOT}/scripts/self-ingest.sh" --recent 1 2>/dev/null || true diff --git a/verisimdb/scripts/self-ingest.sh b/verisimdb/scripts/self-ingest.sh deleted file mode 100755 index f1f1b1aa..00000000 --- a/verisimdb/scripts/self-ingest.sh +++ /dev/null @@ -1,401 +0,0 @@ -#!/usr/bin/env bash -# SPDX-License-Identifier: MPL-2.0 -# Copyright (c) 2026 Jonathan D.A. Jewell (hyperpolymath) <j.d.a.jewell@open.ac.uk> -# -# self-ingest.sh — Ingest repository metadata into the self-hosted .verisimdb/ instance. -# -# Converts git commits, known issues, and scan results into hexad JSON files. -# Each hexad has all 6 modalities populated where data is available. -# -# Usage: -# ./scripts/self-ingest.sh # Full ingest (all commits) -# ./scripts/self-ingest.sh --recent 10 # Last N commits only -# ./scripts/self-ingest.sh --issues # Known issues only - -set -euo pipefail - -REPO_ROOT="$(git rev-parse --show-toplevel)" -HEXAD_DIR="${REPO_ROOT}/.verisimdb/hexads" -INDEX_FILE="${REPO_ROOT}/.verisimdb/index.json" -BASE_IRI="https://verisim.db/self" - -mkdir -p "${HEXAD_DIR}" - -# ============================================================================ -# Helpers -# ============================================================================ - -# Simple hash-based embedding: deterministic 64-dim vector from text. -# Not a real embedding model — serves as placeholder until one is integrated. -text_to_embedding() { - local text="$1" - local hash - hash=$(printf '%s' "$text" | sha256sum | cut -d' ' -f1) - local dims=() - for i in $(seq 0 2 126); do - local byte_hex="${hash:$((i % 64)):2}" - local byte_val=$((16#${byte_hex})) - # Normalize to [-1.0, 1.0] - local norm - norm=$(echo "scale=6; ($byte_val - 128) / 128" | bc -l | sed 's/^\./0./; s/^-\./-0./') - dims+=("$norm") - done - # Pad to 64 dimensions - while [ ${#dims[@]} -lt 64 ]; do - dims+=("0.0") - done - local result="[" - for i in "${!dims[@]}"; do - [ "$i" -gt 0 ] && result+="," - result+="${dims[$i]}" - done - result+="]" - echo "$result" -} - -# Classify commit type from conventional commit prefix -classify_commit() { - local msg="$1" - case "$msg" in - feat:*|feat\(*) echo "feature" ;; - fix:*|fix\(*) echo "bugfix" ;; - docs:*|docs\(*) echo "documentation" ;; - chore:*|chore\(*) echo "chore" ;; - refactor:*|refactor\(*) echo "refactor" ;; - test:*|test\(*) echo "test" ;; - perf:*|perf\(*) echo "performance" ;; - ci:*|ci\(*) echo "ci" ;; - *) echo "other" ;; - esac -} - -# ============================================================================ -# Commit ingestion -# ============================================================================ - -ingest_commits() { - local limit="${1:-0}" - local log_args=(--no-merges --format='%H|%aI|%an|%ae|%s') - - if [ "$limit" -gt 0 ]; then - log_args+=("-n" "$limit") - fi - - local count=0 - local co_change_map="" - - echo "Ingesting commits..." - - while IFS='|' read -r hash date author email subject; do - local hexad_id="commit-${hash:0:12}" - local hexad_file="${HEXAD_DIR}/${hexad_id}.json" - - # Skip if already ingested - if [ -f "$hexad_file" ]; then - continue - fi - - # Get diff stats - local stats - stats=$(git diff-tree --no-commit-id --numstat "$hash" 2>/dev/null || echo "") - local insertions=0 deletions=0 files_changed=0 - local changed_files=() - - while IFS=$'\t' read -r ins del file; do - [ -z "$ins" ] && continue - [ "$ins" = "-" ] && ins=0 - [ "$del" = "-" ] && del=0 - insertions=$((insertions + ins)) - deletions=$((deletions + del)) - files_changed=$((files_changed + 1)) - changed_files+=("$file") - done <<< "$stats" - - # Build graph relationships: files changed in same commit are co-changed - local graph_rels="[]" - if [ ${#changed_files[@]} -gt 0 ] && [ ${#changed_files[@]} -le 20 ]; then - graph_rels="[" - local first=true - for f in "${changed_files[@]}"; do - $first || graph_rels+="," - first=false - # Escape the filename for JSON - local escaped_f - escaped_f=$(printf '%s' "$f" | sed 's/"/\\"/g') - graph_rels+="{\"predicate\":\"modifies\",\"target\":\"file:${escaped_f}\"}" - done - graph_rels+="]" - fi - - # Commit type classification - local commit_type - commit_type=$(classify_commit "$subject") - - # Embedding (hash-based placeholder) - local embedding - embedding=$(text_to_embedding "$subject") - - # Escape subject for JSON - local escaped_subject - escaped_subject=$(printf '%s' "$subject" | sed 's/\\/\\\\/g; s/"/\\"/g; s/\t/\\t/g') - - # Write hexad JSON - cat > "$hexad_file" << HEXAD_EOF -{ - "id": "${hexad_id}", - "source": "git-log", - "created_at": "${date}", - "document": { - "title": "${escaped_subject}", - "body": "Commit ${hash:0:8} by ${author}: ${escaped_subject}", - "fields": { - "type": "commit", - "hash": "${hash}", - "author": "${author}", - "email": "${email}", - "commit_type": "${commit_type}" - } - }, - "graph": { - "relationships": ${graph_rels} - }, - "vector": { - "embedding": ${embedding}, - "model": "sha256-hash-64d" - }, - "tensor": { - "shape": [1, 3], - "data": [${insertions}.0, ${deletions}.0, ${files_changed}.0] - }, - "semantic": { - "types": ["${BASE_IRI}/type/Commit", "${BASE_IRI}/type/${commit_type}"], - "properties": { - "conventional_commit_type": "${commit_type}", - "files_changed": "${files_changed}" - } - }, - "temporal": { - "timestamp": "${date}", - "version": 1, - "author": "${author}" - } -} -HEXAD_EOF - - count=$((count + 1)) - done < <(git log "${log_args[@]}") - - echo " Ingested ${count} commits." -} - -# ============================================================================ -# Known issues ingestion -# ============================================================================ - -ingest_known_issues() { - local issues_file="${REPO_ROOT}/KNOWN-ISSUES.adoc" - [ -f "$issues_file" ] || { echo "No KNOWN-ISSUES.adoc found."; return; } - - echo "Ingesting known issues..." - local count=0 - local current_id="" current_title="" current_status="" current_body="" - local in_issue=false - - while IFS= read -r line; do - if [[ "$line" =~ ^===\ ([0-9]+)\.\ (.+) ]]; then - # Flush previous issue - if [ -n "$current_id" ]; then - write_issue_hexad "$current_id" "$current_title" "$current_status" "$current_body" - count=$((count + 1)) - fi - - local num="${BASH_REMATCH[1]}" - local title_raw="${BASH_REMATCH[2]}" - current_id="issue-$(printf '%03d' "$num")" - - if [[ "$title_raw" == *"RESOLVED"* ]]; then - current_status="resolved" - elif [[ "$title_raw" == *"OPEN"* ]]; then - current_status="open" - else - current_status="unknown" - fi - - # Strip status markers from title - current_title=$(printf '%s' "$title_raw" | sed 's/ — ✅ RESOLVED//; s/ (OPEN)//; s/ (OPEN — LOW)//') - current_body="" - in_issue=true - elif $in_issue; then - current_body+="${line}\n" - fi - done < "$issues_file" - - # Flush last issue - if [ -n "$current_id" ]; then - write_issue_hexad "$current_id" "$current_title" "$current_status" "$current_body" - count=$((count + 1)) - fi - - echo " Ingested ${count} known issues." -} - -write_issue_hexad() { - local id="$1" title="$2" status="$3" body="$4" - local hexad_file="${HEXAD_DIR}/${id}.json" - - local escaped_title - escaped_title=$(printf '%s' "$title" | sed 's/\\/\\\\/g; s/"/\\"/g') - local escaped_body - escaped_body=$(printf '%s' "$body" | sed 's/\\/\\\\/g; s/"/\\"/g; s/\n/\\n/g' | head -c 2000) - - local embedding - embedding=$(text_to_embedding "$title") - - local severity="medium" - [[ "$status" == "resolved" ]] && severity="resolved" - - cat > "$hexad_file" << ISSUE_EOF -{ - "id": "${id}", - "source": "known-issues", - "created_at": "$(date -u +%Y-%m-%dT%H:%M:%SZ)", - "document": { - "title": "${escaped_title}", - "body": "${escaped_body}", - "fields": { - "type": "known_issue", - "status": "${status}", - "severity": "${severity}" - } - }, - "graph": { - "relationships": [{"predicate": "documented_in", "target": "file:KNOWN-ISSUES.adoc"}] - }, - "vector": { - "embedding": ${embedding}, - "model": "sha256-hash-64d" - }, - "tensor": { - "shape": [1, 2], - "data": [$([ "$status" = "resolved" ] && echo "1.0, 0.0" || echo "0.0, 1.0")] - }, - "semantic": { - "types": ["${BASE_IRI}/type/KnownIssue", "${BASE_IRI}/type/${status}"], - "properties": { - "status": "${status}", - "severity": "${severity}" - } - }, - "temporal": { - "timestamp": "$(date -u +%Y-%m-%dT%H:%M:%SZ)", - "version": 1, - "author": "self-ingest" - } -} -ISSUE_EOF -} - -# ============================================================================ -# Index builder -# ============================================================================ - -build_index() { - echo "Building index..." - local hexad_count - hexad_count=$(find "${HEXAD_DIR}" -name "*.json" | wc -l) - - local commits issues - commits=$(find "${HEXAD_DIR}" -name "commit-*.json" | wc -l) - issues=$(find "${HEXAD_DIR}" -name "issue-*.json" | wc -l) - - # Collect all hexad IDs - local ids="[" - local first=true - for f in "${HEXAD_DIR}"/*.json; do - [ -f "$f" ] || continue - local basename - basename=$(basename "$f" .json) - $first || ids+="," - first=false - ids+="\"${basename}\"" - done - ids+="]" - - cat > "$INDEX_FILE" << INDEX_EOF -{ - "instance": "verisimdb-self", - "version": "0.1.0-alpha", - "base_iri": "${BASE_IRI}", - "generated_at": "$(date -u +%Y-%m-%dT%H:%M:%SZ)", - "stats": { - "total_hexads": ${hexad_count}, - "commits": ${commits}, - "known_issues": ${issues} - }, - "hexad_ids": ${ids} -} -INDEX_EOF - - echo " Index: ${hexad_count} hexads (${commits} commits, ${issues} issues)." -} - -# ============================================================================ -# Main -# ============================================================================ - -main() { - local mode="full" - local limit=0 - - while [ $# -gt 0 ]; do - case "$1" in - --recent) - mode="recent" - limit="${2:-10}" - shift 2 - ;; - --issues) - mode="issues" - shift - ;; - --index) - mode="index" - shift - ;; - *) - echo "Usage: $0 [--recent N] [--issues] [--index]" - exit 1 - ;; - esac - done - - echo "VeriSimDB Self-Ingest" - echo "=====================" - echo "Instance: verisimdb-self" - echo "Store: ${HEXAD_DIR}" - echo "" - - case "$mode" in - full) - ingest_commits 0 - ingest_known_issues - build_index - ;; - recent) - ingest_commits "$limit" - build_index - ;; - issues) - ingest_known_issues - build_index - ;; - index) - build_index - ;; - esac - - echo "" - echo "Done. Hexads stored in ${HEXAD_DIR}" -} - -main "$@" diff --git a/verisimdb/scripts/self-query.sh b/verisimdb/scripts/self-query.sh deleted file mode 100755 index 32fdc1a6..00000000 --- a/verisimdb/scripts/self-query.sh +++ /dev/null @@ -1,197 +0,0 @@ -#!/usr/bin/env bash -# SPDX-License-Identifier: MPL-2.0 -# Copyright (c) 2026 Jonathan D.A. Jewell (hyperpolymath) <j.d.a.jewell@open.ac.uk> -# -# self-query.sh — Query the self-hosted .verisimdb/ instance. -# -# Usage: -# ./scripts/self-query.sh search "drift" # Full-text search -# ./scripts/self-query.sh type feature # Filter by commit type -# ./scripts/self-query.sh stats # Show statistics -# ./scripts/self-query.sh issues [open|resolved]# List known issues -# ./scripts/self-query.sh big # Largest commits by insertions -# ./scripts/self-query.sh recent [N] # Most recent N hexads - -set -euo pipefail - -REPO_ROOT="$(git rev-parse --show-toplevel)" -HEXAD_DIR="${REPO_ROOT}/.verisimdb/hexads" -INDEX_FILE="${REPO_ROOT}/.verisimdb/index.json" - -# ============================================================================ -# Query functions -# ============================================================================ - -cmd_search() { - local query="$1" - echo "Searching for: ${query}" - echo "---" - local count=0 - for f in "${HEXAD_DIR}"/*.json; do - if grep -iq "$query" "$f" 2>/dev/null; then - local id title type - id=$(python3 -c "import json,sys; d=json.load(open(sys.argv[1])); print(d['id'])" "$f") - title=$(python3 -c "import json,sys; d=json.load(open(sys.argv[1])); print(d['document']['title'])" "$f") - type=$(python3 -c "import json,sys; d=json.load(open(sys.argv[1])); print(d.get('document',{}).get('fields',{}).get('type','?'))" "$f") - local date - date=$(python3 -c "import json,sys; d=json.load(open(sys.argv[1])); print(d.get('temporal',{}).get('timestamp','?')[:10])" "$f") - printf " [%s] %-14s %s %s\n" "$date" "$id" "($type)" "$title" - count=$((count + 1)) - fi - done - echo "---" - echo "${count} results." -} - -cmd_type() { - local commit_type="$1" - echo "Commits of type: ${commit_type}" - echo "---" - local count=0 - for f in "${HEXAD_DIR}"/commit-*.json; do - local ctype - ctype=$(python3 -c "import json,sys; d=json.load(open(sys.argv[1])); print(d.get('semantic',{}).get('properties',{}).get('conventional_commit_type',''))" "$f" 2>/dev/null) - if [ "$ctype" = "$commit_type" ]; then - local id title date - id=$(python3 -c "import json,sys; d=json.load(open(sys.argv[1])); print(d['id'])" "$f") - title=$(python3 -c "import json,sys; d=json.load(open(sys.argv[1])); print(d['document']['title'])" "$f") - date=$(python3 -c "import json,sys; d=json.load(open(sys.argv[1])); print(d.get('temporal',{}).get('timestamp','?')[:10])" "$f") - printf " [%s] %-20s %s\n" "$date" "$id" "$title" - count=$((count + 1)) - fi - done - echo "---" - echo "${count} ${commit_type} commits." -} - -cmd_stats() { - echo "VeriSimDB Self-Hosted Instance Statistics" - echo "=========================================" - - if [ -f "$INDEX_FILE" ]; then - python3 -c " -import json -d = json.load(open('$INDEX_FILE')) -s = d['stats'] -print(f\" Total hexads: {s['total_hexads']}\") -print(f\" Commits: {s['commits']}\") -print(f\" Known issues: {s['known_issues']}\") -print(f\" Generated: {d['generated_at']}\") -" - fi - - echo "" - echo "Commit type breakdown:" - python3 -c " -import json, os, collections -types = collections.Counter() -total_ins = 0 -total_del = 0 -for f in sorted(os.listdir('$HEXAD_DIR')): - if not f.startswith('commit-'): continue - d = json.load(open(os.path.join('$HEXAD_DIR', f))) - ct = d.get('semantic',{}).get('properties',{}).get('conventional_commit_type','other') - types[ct] += 1 - tensor = d.get('tensor',{}).get('data',[0,0,0]) - total_ins += tensor[0] - total_del += tensor[1] -for t, c in types.most_common(): - print(f' {t:15s} {c:4d}') -print(f'') -print(f' Total insertions: {int(total_ins):,}') -print(f' Total deletions: {int(total_del):,}') -print(f' Net lines: {int(total_ins - total_del):+,}') -" - - echo "" - echo "Known issues:" - python3 -c " -import json, os, collections -statuses = collections.Counter() -for f in sorted(os.listdir('$HEXAD_DIR')): - if not f.startswith('issue-'): continue - d = json.load(open(os.path.join('$HEXAD_DIR', f))) - s = d.get('document',{}).get('fields',{}).get('status','unknown') - statuses[s] += 1 -for s, c in statuses.most_common(): - print(f' {s:15s} {c:4d}') -" -} - -cmd_issues() { - local filter="${1:-all}" - echo "Known Issues (filter: ${filter})" - echo "---" - for f in "${HEXAD_DIR}"/issue-*.json; do - [ -f "$f" ] || continue - local status title id - id=$(python3 -c "import json,sys; d=json.load(open(sys.argv[1])); print(d['id'])" "$f") - title=$(python3 -c "import json,sys; d=json.load(open(sys.argv[1])); print(d['document']['title'])" "$f") - status=$(python3 -c "import json,sys; d=json.load(open(sys.argv[1])); print(d.get('document',{}).get('fields',{}).get('status','?'))" "$f") - - if [ "$filter" = "all" ] || [ "$filter" = "$status" ]; then - local marker=" " - [ "$status" = "resolved" ] && marker="ok" - [ "$status" = "open" ] && marker="!!" - printf " [%s] %-12s %s\n" "$marker" "$id" "$title" - fi - done -} - -cmd_big() { - echo "Largest commits by insertions:" - echo "---" - python3 -c " -import json, os -commits = [] -for f in sorted(os.listdir('$HEXAD_DIR')): - if not f.startswith('commit-'): continue - d = json.load(open(os.path.join('$HEXAD_DIR', f))) - tensor = d.get('tensor',{}).get('data',[0,0,0]) - commits.append((tensor[0], tensor[1], tensor[2], d['id'], d['document']['title'][:60])) -commits.sort(reverse=True) -for ins, dels, files, cid, title in commits[:15]: - print(f' +{int(ins):5d} -{int(dels):5d} ({int(files):2d} files) {title}') -" -} - -cmd_recent() { - local n="${1:-10}" - echo "Most recent ${n} hexads:" - echo "---" - python3 -c " -import json, os -hexads = [] -for f in sorted(os.listdir('$HEXAD_DIR')): - d = json.load(open(os.path.join('$HEXAD_DIR', f))) - ts = d.get('temporal',{}).get('timestamp','1970-01-01') - hexads.append((ts, d['id'], d['document']['title'][:70], d.get('document',{}).get('fields',{}).get('type','?'))) -hexads.sort(reverse=True) -for ts, hid, title, htype in hexads[:${n}]: - print(f' [{ts[:10]}] {hid:20s} ({htype:12s}) {title}') -" -} - -# ============================================================================ -# Main -# ============================================================================ - -case "${1:-help}" in - search) cmd_search "${2:?Usage: self-query.sh search <query>}" ;; - type) cmd_type "${2:?Usage: self-query.sh type <feat|fix|docs|chore>}" ;; - stats) cmd_stats ;; - issues) cmd_issues "${2:-all}" ;; - big) cmd_big ;; - recent) cmd_recent "${2:-10}" ;; - *) - echo "Usage: self-query.sh <command> [args]" - echo "" - echo "Commands:" - echo " search <query> Full-text search across all hexads" - echo " type <commit-type> Filter commits by type (feat/fix/docs/chore)" - echo " stats Show instance statistics" - echo " issues [status] List known issues (all/open/resolved)" - echo " big Largest commits by insertions" - echo " recent [N] Most recent N hexads" - ;; -esac diff --git a/verisimdb/scripts/smoke-test.sh b/verisimdb/scripts/smoke-test.sh deleted file mode 100755 index b64a5ecf..00000000 --- a/verisimdb/scripts/smoke-test.sh +++ /dev/null @@ -1,149 +0,0 @@ -#!/usr/bin/env bash -# SPDX-License-Identifier: MPL-2.0 -# VeriSimDB single-node production smoke test. -# Validates: startup, create, read, shutdown, restart, verify persistence. -# -# Usage: ./scripts/smoke-test.sh [--persistent] -# -# Requires: cargo, curl, jq - -set -euo pipefail - -PERSIST="" -DATA_DIR="" -PORT=18080 -GRPC_PORT=18051 - -if [[ "${1:-}" == "--persistent" ]]; then - PERSIST="yes" - DATA_DIR=$(mktemp -d /tmp/verisimdb-smoke-XXXXXX) - echo "=== VeriSimDB Smoke Test (PERSISTENT mode) ===" - echo " Data dir: $DATA_DIR" -else - echo "=== VeriSimDB Smoke Test (in-memory mode) ===" -fi - -cleanup() { - echo "" - echo "Cleaning up..." - if [[ -n "${SERVER_PID:-}" ]] && kill -0 "$SERVER_PID" 2>/dev/null; then - kill -SIGTERM "$SERVER_PID" 2>/dev/null || true - wait "$SERVER_PID" 2>/dev/null || true - fi - if [[ -n "$DATA_DIR" ]]; then - rm -rf "$DATA_DIR" - fi -} -trap cleanup EXIT - -# Build -echo "" -echo "[1/6] Building VeriSimDB..." -if [[ -n "$PERSIST" ]]; then - cargo build -p verisim-api --features persistent --release 2>&1 | tail -1 - BINARY="target/release/verisim-api" -else - cargo build -p verisim-api --release 2>&1 | tail -1 - BINARY="target/release/verisim-api" -fi - -# Start server -echo "[2/6] Starting server on port $PORT..." -export VERISIM_HOST="127.0.0.1" -export VERISIM_PORT="$PORT" -export VERISIM_GRPC_PORT="$GRPC_PORT" -if [[ -n "$PERSIST" ]]; then - export VERISIM_PERSISTENCE_DIR="$DATA_DIR" -fi - -$BINARY & -SERVER_PID=$! -sleep 2 - -# Check it's running -if ! kill -0 "$SERVER_PID" 2>/dev/null; then - echo "FAIL: Server did not start" - exit 1 -fi -echo " Server started (PID: $SERVER_PID)" - -# Health check -echo "[3/6] Health check..." -HEALTH=$(curl -sf "http://127.0.0.1:$PORT/health" 2>/dev/null || echo "FAIL") -if echo "$HEALTH" | grep -q "healthy\|ok"; then - echo " Health: OK" -else - echo " FAIL: Health check returned: $HEALTH" - exit 1 -fi - -# Create entity -echo "[4/6] Creating entity..." -CREATE_RESP=$(curl -sf -X POST "http://127.0.0.1:$PORT/octads" \ - -H "Content-Type: application/json" \ - -d '{ - "document": { - "title": "Smoke Test Entity", - "body": "This entity tests single-node production readiness." - }, - "vector": { - "embedding": [0.1, 0.2, 0.3, 0.4, 0.5] - } - }' 2>/dev/null || echo "FAIL") - -if echo "$CREATE_RESP" | grep -q "id"; then - ENTITY_ID=$(echo "$CREATE_RESP" | python3 -c "import sys,json; print(json.load(sys.stdin).get('id',''))" 2>/dev/null || echo "") - if [[ -z "$ENTITY_ID" ]]; then - # Try jq - ENTITY_ID=$(echo "$CREATE_RESP" | jq -r '.id' 2>/dev/null || echo "") - fi - echo " Created: $ENTITY_ID" -else - echo " FAIL: Create returned: $CREATE_RESP" - exit 1 -fi - -# Read entity back -echo "[5/6] Reading entity back..." -GET_RESP=$(curl -sf "http://127.0.0.1:$PORT/octads/$ENTITY_ID" 2>/dev/null || echo "FAIL") -if echo "$GET_RESP" | grep -q "$ENTITY_ID"; then - echo " Read: OK" -else - echo " FAIL: Get returned: $GET_RESP" - exit 1 -fi - -# Graceful shutdown -echo "[6/6] Graceful shutdown..." -kill -SIGTERM "$SERVER_PID" -wait "$SERVER_PID" 2>/dev/null || true -echo " Server stopped cleanly" - -# If persistent, restart and verify data survived -if [[ -n "$PERSIST" ]]; then - echo "" - echo "[BONUS] Persistence verification..." - echo " Restarting server..." - $BINARY & - SERVER_PID=$! - sleep 2 - - if ! kill -0 "$SERVER_PID" 2>/dev/null; then - echo " FAIL: Server did not restart" - exit 1 - fi - - GET_RESP2=$(curl -sf "http://127.0.0.1:$PORT/octads/$ENTITY_ID" 2>/dev/null || echo "FAIL") - if echo "$GET_RESP2" | grep -q "$ENTITY_ID"; then - echo " Persistence: VERIFIED — entity survived restart" - else - echo " WARN: Entity not found after restart (WAL replay may need octad registry rebuild)" - echo " Response: $GET_RESP2" - fi - - kill -SIGTERM "$SERVER_PID" - wait "$SERVER_PID" 2>/dev/null || true -fi - -echo "" -echo "=== SMOKE TEST PASSED ===" diff --git a/verisimdb/scripts/two-node-test.sh b/verisimdb/scripts/two-node-test.sh deleted file mode 100755 index 67f67356..00000000 --- a/verisimdb/scripts/two-node-test.sh +++ /dev/null @@ -1,139 +0,0 @@ -#!/usr/bin/env bash -# SPDX-License-Identifier: MPL-2.0 -# VeriSimDB Phase 4.B: Two-node federation test. -# -# Starts two VeriSimDB instances (primary + replica), creates data on the -# primary, and verifies it can be queried via the replica's federation endpoint. -# -# This proves VeriSimDB can coordinate across multiple nodes. -# -# Usage: ./scripts/two-node-test.sh - -set -euo pipefail - -PRIMARY_PORT=18080 -PRIMARY_GRPC=18051 -REPLICA_PORT=18090 -REPLICA_GRPC=18052 -PIDS=() - -echo "=== VeriSimDB Two-Node Federation Test ===" - -cleanup() { - echo "" - echo "Cleaning up..." - for pid in "${PIDS[@]}"; do - if kill -0 "$pid" 2>/dev/null; then - kill -SIGTERM "$pid" 2>/dev/null || true - wait "$pid" 2>/dev/null || true - fi - done - rm -rf /tmp/verisimdb-node-{a,b} 2>/dev/null || true -} -trap cleanup EXIT - -# Build persistent version -echo "[1/7] Building VeriSimDB (persistent)..." -cargo build -p verisim-api --features persistent --release 2>&1 | tail -1 -BINARY="target/release/verisim-api" - -# Start Node A (primary) -echo "[2/7] Starting Node A (primary) on port $PRIMARY_PORT..." -mkdir -p /tmp/verisimdb-node-a -VERISIM_HOST=127.0.0.1 VERISIM_PORT=$PRIMARY_PORT VERISIM_GRPC_PORT=$PRIMARY_GRPC \ - VERISIM_PERSISTENCE_DIR=/tmp/verisimdb-node-a \ - $BINARY & -PIDS+=($!) -sleep 2 - -if ! kill -0 "${PIDS[0]}" 2>/dev/null; then - echo "FAIL: Node A did not start" - exit 1 -fi -echo " Node A running (PID: ${PIDS[0]})" - -# Start Node B (replica) -echo "[3/7] Starting Node B (replica) on port $REPLICA_PORT..." -mkdir -p /tmp/verisimdb-node-b -VERISIM_HOST=127.0.0.1 VERISIM_PORT=$REPLICA_PORT VERISIM_GRPC_PORT=$REPLICA_GRPC \ - VERISIM_PERSISTENCE_DIR=/tmp/verisimdb-node-b \ - $BINARY & -PIDS+=($!) -sleep 2 - -if ! kill -0 "${PIDS[1]}" 2>/dev/null; then - echo "FAIL: Node B did not start" - exit 1 -fi -echo " Node B running (PID: ${PIDS[1]})" - -# Health check both nodes -echo "[4/7] Health check..." -HA=$(curl -sf "http://127.0.0.1:$PRIMARY_PORT/health" 2>/dev/null || echo "FAIL") -HB=$(curl -sf "http://127.0.0.1:$REPLICA_PORT/health" 2>/dev/null || echo "FAIL") -if echo "$HA" | grep -q "healthy" && echo "$HB" | grep -q "healthy"; then - echo " Both nodes healthy" -else - echo " FAIL: Node A: $HA, Node B: $HB" - exit 1 -fi - -# Create entity on Node A -echo "[5/7] Creating entity on Node A..." -CREATE_RESP=$(curl -sf -X POST "http://127.0.0.1:$PRIMARY_PORT/octads" \ - -H "Content-Type: application/json" \ - -d '{ - "document": { - "title": "Federation Test Entity", - "body": "Created on Node A, should be queryable from Node B." - } - }' 2>/dev/null || echo "FAIL") - -ENTITY_ID=$(echo "$CREATE_RESP" | python3 -c "import sys,json; print(json.load(sys.stdin).get('id',''))" 2>/dev/null || echo "") -if [[ -z "$ENTITY_ID" ]]; then - ENTITY_ID=$(echo "$CREATE_RESP" | jq -r '.id' 2>/dev/null || echo "") -fi - -if [[ -n "$ENTITY_ID" ]]; then - echo " Created on Node A: $ENTITY_ID" -else - echo " FAIL: Create returned: $CREATE_RESP" - exit 1 -fi - -# Verify entity exists on Node A -echo "[6/7] Verifying entity on Node A..." -GET_A=$(curl -sf "http://127.0.0.1:$PRIMARY_PORT/octads/$ENTITY_ID" 2>/dev/null || echo "FAIL") -if echo "$GET_A" | grep -q "$ENTITY_ID"; then - echo " Node A: entity found" -else - echo " FAIL: Node A doesn't have the entity" - exit 1 -fi - -# Verify entity does NOT exist on Node B (separate instance, no replication yet) -echo "[7/7] Verifying Node B is independent..." -GET_B=$(curl -sf "http://127.0.0.1:$REPLICA_PORT/octads/$ENTITY_ID" 2>/dev/null || echo "NOT_FOUND") -if echo "$GET_B" | grep -q "$ENTITY_ID"; then - echo " Node B: entity found (unexpected — replication working?)" -else - echo " Node B: entity NOT found (expected — nodes are independent)" - echo " Federation replication is the next step (Phase 4.C)" -fi - -echo "" -echo "=== TWO-NODE TEST PASSED ===" -echo "" -echo "Results:" -echo " - Two VeriSimDB instances run simultaneously on different ports" -echo " - Both respond to health checks" -echo " - Entity created on Node A is retrievable from Node A" -echo " - Nodes are independent (no automatic replication)" -echo " - Federation replication (Phase 4.C) will enable cross-node queries" -echo "" -echo "This validates Phase 4.B: two nodes can coexist." -echo "Phase 4.C (full federation) will add:" -echo " - Peer registration between nodes" -echo " - Cross-node query routing via Elixir Resolver" -echo " - Drift detection across federated nodes" -echo " - Write replication policies" diff --git a/verisimdb/selur-compose.yml b/verisimdb/selur-compose.yml deleted file mode 100644 index a9922675..00000000 --- a/verisimdb/selur-compose.yml +++ /dev/null @@ -1,262 +0,0 @@ -# SPDX-License-Identifier: MPL-2.0 -# SPDX-FileCopyrightText: 2026 Jonathan D.A. Jewell (hyperpolymath) -# -# selur-compose.yml — VeriSimDB deployment stack -# Run with: selur seal && podman-compose -f selur-compose.yml up -d -# -# Stack: verisimdb (Rust API + Elixir OTP) with full security tooling -# Vordr: runtime verification, Cerro-Torre: image signing, Rokur: secret rotation - -version: "1.0" - -x-security-defaults: &security-defaults - restart: unless-stopped - security_opt: - - no-new-privileges:true - read_only: true - tmpfs: - - /tmp:noexec,nosuid,size=64M - cap_drop: - - ALL - -x-svalinn: - policy: strict - verification: - require_manifest: true - require_attestation: true - require_signature: true - attestations: - require-sbom: true - require-signature: true - require-provenance: true - slsa-level: 3 - crypto: - signature-algorithm: ML-DSA-87 - hash-algorithm: SHA-3-256 - key-exchange: ML-KEM-1024 - -x-vordr-config: - enable-formal-proofs: true - proof-systems: [idris2, lean4] - memory-model: linear-types - concurrency-model: capability-safe - syscall-policy: deny-by-default - network-policy: deny-by-default - formal-verification: true - runtime-checks: true - memory-safety: proven - -x-cerro-torre: - signing: - algorithm: ML-DSA-87 - key-source: rokur - key-path: verisimdb/signing/key - attestations: - sbom-format: spdx-json - sbom-path: /app/sbom.spdx.json - provenance-log: verisimdb-data/provenance/ - verification: - enforce: true - trust-root: cerro-torre-root-2026 - -x-rokur: - secrets-backend: rokur - rotation-policy: - interval: 30d - algorithm: argon2id - lifecycle: - retention: 90d - auto-rotate: true - -services: - # ── VeriSimDB Rust Core ────────────────────────────────────── - verisimdb-api: - image: ghcr.io/hyperpolymath/verisimdb:latest - build: - context: . - dockerfile: container/Containerfile - target: "" - container_name: verisimdb-api - hostname: verisimdb-api - <<: *security-defaults - read_only: false - ports: - - "[::1]:8080:8080" - environment: - RUST_LOG: info - VERISIM_HOST: "[::]" - VERISIM_PORT: "8080" - VERISIM_LOG_FORMAT: json - VERISIM_DATA_DIR: /app/data - volumes: - - verisimdb_data:/app/data:rw - networks: - - verisimdb-internal - healthcheck: - test: ["CMD", "curl", "-sf", "http://localhost:8080/health"] - interval: 30s - timeout: 5s - start_period: 10s - retries: 3 - user: "verisim" - x-svalinn: - policy: strict - verify: true - labels: - io.hyperpolymath.service: "verisimdb-api" - io.hyperpolymath.modalities: "graph,vector,tensor,semantic,document,temporal" - - # ── Elixir OTP Orchestration ───────────────────────────────── - verisimdb-otp: - image: ghcr.io/hyperpolymath/verisimdb:latest - container_name: verisimdb-otp - hostname: verisimdb-otp - <<: *security-defaults - read_only: false - command: ["/app/elixir/bin/verisim", "start"] - environment: - VERISIM_RUST_CORE_URL: "http://verisimdb-api:8080/api/v1" - MIX_ENV: prod - RELEASE_NODE: "verisim@verisimdb-otp" - ERL_AFLAGS: "+JPperf true" - depends_on: - verisimdb-api: - condition: service_healthy - networks: - - verisimdb-internal - healthcheck: - test: ["CMD", "/app/elixir/bin/verisim", "pid"] - interval: 30s - timeout: 5s - retries: 3 - user: "verisim" - x-svalinn: - policy: strict - verify: true - labels: - io.hyperpolymath.service: "verisimdb-otp" - io.hyperpolymath.role: "orchestration" - - # ── Svalinn Gateway ────────────────────────────────────────── - svalinn: - image: ghcr.io/hyperpolymath/svalinn:latest - container_name: verisimdb-svalinn - hostname: svalinn - <<: *security-defaults - ports: - - "[::]:8443:8443" - environment: - SVALINN_UPSTREAM: "http://verisimdb-api:8080" - SVALINN_TLS_CERT: /etc/svalinn/tls/cert.pem - SVALINN_TLS_KEY: /etc/svalinn/tls/key.pem - SVALINN_POLICY: strict - volumes: - - svalinn_tls:/etc/svalinn/tls:ro - depends_on: - verisimdb-api: - condition: service_healthy - networks: - - verisimdb-internal - - verisimdb-external - x-svalinn: - policy: strict - verify: true - labels: - io.hyperpolymath.service: "svalinn" - io.hyperpolymath.role: "gateway" - - # ── Vordr Runtime Verifier ─────────────────────────────────── - vordr: - image: ghcr.io/hyperpolymath/vordr:latest - container_name: verisimdb-vordr - hostname: vordr - <<: *security-defaults - environment: - VORDR_TARGET: "verisimdb-api:8080" - VORDR_POLICY: strict - VORDR_FORMAL_PROOFS: "true" - VORDR_REPORT_FORMAT: json - depends_on: - verisimdb-api: - condition: service_healthy - networks: - - verisimdb-internal - x-svalinn: - policy: strict - verify: true - labels: - io.hyperpolymath.service: "vordr" - io.hyperpolymath.role: "verification" - - # ── Cerro-Torre Image Signer ───────────────────────────────── - cerro-torre: - image: ghcr.io/hyperpolymath/cerro-torre:latest - container_name: verisimdb-cerro-torre - hostname: cerro-torre - <<: *security-defaults - environment: - CERRO_TORRE_REGISTRY: "ghcr.io/hyperpolymath" - CERRO_TORRE_ALGORITHM: ML-DSA-87 - CERRO_TORRE_SBOM_FORMAT: spdx-json - volumes: - - cerro_torre_keys:/etc/cerro-torre/keys:ro - networks: - - verisimdb-internal - x-svalinn: - policy: strict - verify: true - labels: - io.hyperpolymath.service: "cerro-torre" - io.hyperpolymath.role: "signing" - - # ── Rokur Secret Manager ───────────────────────────────────── - rokur: - image: ghcr.io/hyperpolymath/rokur:latest - container_name: verisimdb-rokur - hostname: rokur - <<: *security-defaults - environment: - ROKUR_STORE: /var/rokur/secrets - ROKUR_ROTATION_INTERVAL: 30d - ROKUR_ALGORITHM: argon2id - volumes: - - rokur_secrets:/var/rokur/secrets:rw - networks: - - verisimdb-internal - x-svalinn: - policy: strict - verify: true - labels: - io.hyperpolymath.service: "rokur" - io.hyperpolymath.role: "secrets" - -networks: - verisimdb-internal: - driver: bridge - internal: true - ipam: - config: - - subnet: "fd00:verisim:1::/48" - verisimdb-external: - driver: bridge - -volumes: - verisimdb_data: - driver: local - svalinn_tls: - driver: local - cerro_torre_keys: - driver: local - rokur_secrets: - driver: local - -secrets: - db-signing-key: - provider: rokur - key: verisimdb/signing/key - tls-cert: - provider: rokur - key: verisimdb/tls/cert - tls-key: - provider: rokur - key: verisimdb/tls/key diff --git a/verisimdb/site/index.md b/verisimdb/site/index.md deleted file mode 100644 index 4b2f8237..00000000 --- a/verisimdb/site/index.md +++ /dev/null @@ -1,21 +0,0 @@ ---- -title: VeriSimDB -date: 2026-03-31 ---- - -# VeriSimDB - -The public web home for this project is [verisimdb.org](https://verisimdb.org). - -Cross-system data consistency that catches drift before it causes damage. - -Traditional tools detect drift after it causes downstream failures. VeriSimDB detects and repairs it continuously, before anyone notices. - -## Project Links - -- Website: [verisimdb.org](https://verisimdb.org) -- Source: [https://github.com/hyperpolymath/nextgen-databases/tree/main/verisimdb](https://github.com/hyperpolymath/nextgen-databases/tree/main/verisimdb) -- README: [project overview](https://github.com/hyperpolymath/nextgen-databases/blob/main/verisimdb/README.adoc) -- Docs: [documentation directory](https://github.com/hyperpolymath/nextgen-databases/tree/main/verisimdb/docs) - -This page is a lightweight landing point for the repository and will grow with the project. diff --git a/verisimdb/spec/README.adoc b/verisimdb/spec/README.adoc deleted file mode 100644 index b28297d6..00000000 --- a/verisimdb/spec/README.adoc +++ /dev/null @@ -1,23 +0,0 @@ -// SPDX-License-Identifier: CC-BY-SA-4.0 -// @taxonomy: spec/index -= verisimdb — Specification Directory -:toc: - -== Overview - -This directory contains the canonical language/query specification files for verisimdb. - -== Contents - -=== grammar.ebnf -The canonical EBNF grammar for verisimdb. -Copied from the original source location; the original is preserved. - -=== SPEC.core.scm -Not yet available. - -== Conventions - -- All specification files use the `@taxonomy: spec/*` annotation prefix -- Grammar changes must update both the original and this canonical copy -- SPEC.core.scm follows the RSR META-FORMAT-SPEC diff --git a/verisimdb/spec/grammar.ebnf b/verisimdb/spec/grammar.ebnf deleted file mode 100644 index d19d69cc..00000000 --- a/verisimdb/spec/grammar.ebnf +++ /dev/null @@ -1,318 +0,0 @@ -(* @taxonomy: spec/grammar *) -(* SPDX-License-Identifier: MPL-2.0 *) -(* Copyright (c) 2026 Jonathan D.A. Jewell (hyperpolymath) <j.d.a.jewell@open.ac.uk> *) -(* VeriSim Consonance Language (VCL) Grammar *) -(* Format: Extended Backus-Naur Form (EBNF) *) -(* Version: 3.0 — Octad (8 modalities), provenance conditions, spatial conditions *) -(* Date: 2026-02-27 *) - -(* ============================================================================ - 1. TOP-LEVEL STRUCTURE - ============================================================================ *) -statement = query | mutation ; - -query = select_clause, - from_clause, - [where_clause], - [group_by_clause], - [having_clause], - [proof_clause], - [order_by_clause], - [limit_clause], - [offset_clause] ; - -(* ============================================================================ - 2. SELECT CLAUSE - Modality Selection, Column Projections & Aggregates - ============================================================================ *) -select_clause = 'SELECT', select_item_list ; - -select_item_list = select_item, { ',', select_item } ; - -(* A select item can be an aggregate, a column projection, or a full modality *) -select_item = aggregate_expr | field_ref | modality_spec ; - -(* Column projection: MODALITY.field_name *) -field_ref = modality_name, '.', identifier ; - -(* Aggregate functions *) -aggregate_expr = count_all | aggregate_field ; -count_all = 'COUNT', '(', '*', ')' ; -aggregate_field = aggregate_func, '(', field_ref, ')' ; -aggregate_func = 'COUNT' | 'SUM' | 'AVG' | 'MIN' | 'MAX' ; - -(* Bare modality selection *) -modality_spec = 'GRAPH', [graph_projection] | 'VECTOR', [vector_projection] | 'TENSOR', [tensor_projection] | 'SEMANTIC', [semantic_projection] | 'DOCUMENT', [document_projection] | 'TEMPORAL', [temporal_projection] | 'PROVENANCE', [provenance_projection] | 'SPATIAL', [spatial_projection] | '*' (* All available modalities *) ; - -modality_name = 'GRAPH' | 'VECTOR' | 'TENSOR' | 'SEMANTIC' | 'DOCUMENT' | 'TEMPORAL' | 'PROVENANCE' | 'SPATIAL' ; - -(* Projection specifics per modality *) -graph_projection = '(', sparql_pattern, ')' ; -vector_projection = '(', vector_fields, ')' ; -tensor_projection = '(', tensor_slice, ')' ; -semantic_projection = '(', contract_names, ')' ; -document_projection = '(', document_fields, ')' ; -temporal_projection = '(', version_spec, ')' ; -provenance_projection = '(', provenance_fields, ')' ; -spatial_projection = '(', spatial_fields, ')' ; - -(* ============================================================================ - 3. FROM CLAUSE - Data Sources - ============================================================================ *) -from_clause = 'FROM', source_spec ; - -source_spec = hexad_source | federation_source | store_source ; - -(* Direct hexad reference *) -hexad_source = 'HEXAD', uuid ; - -(* Federation pattern (multiple nodes) *) -federation_source = 'FEDERATION', node_pattern, [drift_policy] ; - -node_pattern = glob_pattern (* e.g., '/universities/*' *) | node_list (* e.g., '[node1, node2, node3]' *) ; - -drift_policy = 'WITH', 'DRIFT', drift_mode ; -drift_mode = 'STRICT' (* Fail on any drift *) | 'REPAIR' (* Auto-repair detected drift *) | 'TOLERATE' (* Return data despite drift *) | 'LATEST' (* Use most recent version *) ; - -(* Specific store reference *) -store_source = 'STORE', store_id ; - -(* ============================================================================ - 4. WHERE CLAUSE - Filtering Conditions, - ============================================================================ *) -where_clause = 'WHERE', condition ; - -condition = simple_condition | compound_condition | '(', condition, ')' ; - -simple_condition = graph_condition | vector_condition | tensor_condition | semantic_condition | document_condition | temporal_condition | provenance_condition | spatial_condition | cross_modal_condition ; - -(* 4.9. Cross-Modal Conditions — relationships BETWEEN modalities *) -cross_modal_condition = cross_modal_field_compare | drift_condition | consistency_condition | exists_condition | not_exists_condition ; - -(* Compare fields across modalities: WHERE DOCUMENT.severity > GRAPH.centrality *) -cross_modal_field_compare = field_ref, comparison_op, field_ref ; - -(* Drift between modalities: WHERE DRIFT(VECTOR, DOCUMENT) > 0.3 *) -drift_condition = 'DRIFT', '(', modality_name, ',', modality_name, ')', comparison_op, float ; - -(* Consistency check: WHERE CONSISTENT(VECTOR, SEMANTIC) USING COSINE *) -consistency_condition = 'CONSISTENT', '(', modality_name, ',', modality_name, ')', 'USING', metric_name ; -metric_name = 'COSINE' | 'EUCLIDEAN' | 'DOT_PRODUCT' | 'JACCARD' ; - -(* Modality existence: WHERE VECTOR EXISTS *) -exists_condition = modality_name, 'EXISTS' ; - -(* Modality absence: WHERE TENSOR NOT EXISTS *) -not_exists_condition = modality_name, 'NOT', 'EXISTS' ; - -compound_condition = condition, 'AND', condition | condition, 'OR', condition | 'NOT', condition ; - -(* 4.1. Graph Conditions (SPARQL-like) *) -graph_condition = sparql_pattern | path_pattern ; - -sparql_pattern = '(', node_var, ')', edge_pattern, '(', node_var, ')' ; -edge_pattern = '-[', edge_type, ']->' | '-[', edge_type, ']-' | '<-[', edge_type, ']-' ; -node_var = identifier | ('?', identifier) ; -edge_type = ':', identifier ; - -path_pattern = node_var, path_quantifier, node_var ; -path_quantifier = '-[', edge_type, ('*' | '+' | '{', integer, ',', integer, '}'), ']->' ; - -(* 4.2. Vector Conditions (Similarity Search) *) -vector_condition = vector_field, 'SIMILAR', 'TO', vector_literal, [similarity_threshold] | vector_field, 'NEAREST', integer, [metric_type] ; - -vector_field = identifier, '.', 'embedding' ; -vector_literal = '[', float, { ',', float }, ']' ; -similarity_threshold = 'WITHIN', float ; -metric_type = 'USING', ('COSINE' | 'EUCLIDEAN' | 'DOT_PRODUCT') ; - -(* 4.3. Tensor Conditions (Multi-dimensional) *) -tensor_condition = tensor_field, tensor_op, tensor_literal ; -tensor_op = '==' | '>' | '<' | '>=' | '<=' | 'SHAPE' | 'RANK' ; -tensor_literal = array_literal | scalar_literal ; - -(* 4.4. Semantic Conditions (ZKP Contracts) *) -semantic_condition = 'SATISFIES', contract_name, [contract_params] | 'HAS', 'PROOF', proof_type | 'VERIFIED', 'BY', verifier_id ; - -contract_name = identifier ; -contract_params = '(', param_list, ')' ; -param_list = identifier, '=', literal, { ',', identifier, '=', literal } ; - -(* 4.5. Document Conditions (Full-text Search) *) -document_condition = 'FULLTEXT', 'CONTAINS', string_literal | 'FULLTEXT', 'MATCHES', regex_literal | 'FIELD', identifier, comparison_op, literal ; - -comparison_op = '==' | '!=' | '>' | '<' | '>=' | '<=' | 'LIKE' ; - -(* 4.6. Temporal Conditions (Versioning) *) -temporal_condition = 'AS', 'OF', timestamp | 'BETWEEN', timestamp, 'AND', timestamp | 'VERSION', version_id | 'MODIFIED', 'BY', actor_id ; - -(* 4.7. Provenance Conditions (Lineage Tracking) *) -provenance_condition = provenance_actor | provenance_origin | provenance_chain | provenance_event ; - -(* Filter by actor: WHERE PROVENANCE.actor = 'system' *) -provenance_actor = 'PROVENANCE', '.', 'actor', comparison_op, string_literal ; - -(* Filter by origin: WHERE PROVENANCE.origin = 'import' *) -provenance_origin = 'PROVENANCE', '.', 'origin', comparison_op, string_literal ; - -(* Filter by chain integrity: WHERE PROVENANCE.chain_valid = true *) -provenance_chain = 'PROVENANCE', '.', 'chain_valid', '=', boolean - | 'PROVENANCE', '.', 'chain_length', comparison_op, integer ; - -(* Filter by event type: WHERE PROVENANCE.event_type = 'Modified' *) -provenance_event = 'PROVENANCE', '.', 'event_type', comparison_op, string_literal ; - -(* Provenance-specific fields *) -provenance_fields = identifier, { ',', identifier } ; - -(* 4.8. Spatial Conditions (Geospatial Queries) *) -spatial_condition = spatial_radius | spatial_bounds | spatial_nearest | spatial_field ; - -(* Radius search: WHERE WITHIN RADIUS(51.5074, -0.1278, 100.0) *) -spatial_radius = 'WITHIN', 'RADIUS', '(', float, ',', float, ',', float, ')' ; - -(* Bounding box: WHERE WITHIN BOUNDS(51.0, -1.0, 52.0, 0.5) *) -spatial_bounds = 'WITHIN', 'BOUNDS', '(', float, ',', float, ',', float, ',', float, ')' ; - -(* K-nearest: WHERE NEAREST(51.5074, -0.1278, 10) *) -spatial_nearest = 'NEAREST', '(', float, ',', float, ',', integer, ')' ; - -(* Field access: WHERE SPATIAL.latitude > 51.0 *) -spatial_field = 'SPATIAL', '.', identifier, comparison_op, literal ; - -(* Spatial-specific fields *) -spatial_fields = identifier, { ',', identifier } ; - -(* ============================================================================ - 5. PROOF CLAUSE - Dependent-Type Safety (Multi-Proof Composition) - ============================================================================ *) -proof_clause = 'PROOF', proof_spec_list ; - -proof_spec_list = proof_spec, { 'AND', proof_spec } ; - -proof_spec = proof_type, '(', contract_name, ')', [proof_params] ; - -proof_type = 'EXISTENCE' (* Hexad exists and is accessible *) | 'CITATION' (* Citation chain is valid *) | 'ACCESS' (* User has access rights *) | 'INTEGRITY' (* Data has not been tampered with *) | 'PROVENANCE' (* Lineage is verifiable *) | 'CUSTOM' (* Custom ZKP contract *) ; - -proof_params = 'WITH', param_list ; - -(* ============================================================================ - 6. GROUP BY / HAVING CLAUSES - Aggregation - ============================================================================ *) -group_by_clause = 'GROUP', 'BY', field_ref_list ; -field_ref_list = field_ref, { ',', field_ref } ; - -having_clause = 'HAVING', condition ; - -(* ============================================================================ - 7. ORDER BY CLAUSE - Sorting - ============================================================================ *) -order_by_clause = 'ORDER', 'BY', order_by_list ; -order_by_list = order_by_item, { ',', order_by_item } ; -order_by_item = field_ref, [sort_direction] ; -sort_direction = 'ASC' | 'DESC' ; - -(* ============================================================================ - 8. PAGINATION CLAUSES - ============================================================================ *) -limit_clause = 'LIMIT', integer ; -offset_clause = 'OFFSET', integer ; - -(* ============================================================================ - 9. MUTATIONS - Write Path (INSERT / UPDATE / DELETE) - ============================================================================ *) -mutation = insert_mutation | update_mutation | delete_mutation ; - -(* INSERT: Create a new hexad with modality data *) -insert_mutation = 'INSERT', 'HEXAD', 'WITH', modality_data_list, [proof_clause] ; - -modality_data_list = modality_data, { ',', modality_data } ; - -modality_data = document_data | vector_data | graph_data | tensor_data | semantic_data | temporal_data | provenance_data | spatial_data ; - -document_data = 'DOCUMENT', '(', field_assignment_list, ')' ; -vector_data = 'VECTOR', '(', vector_literal, ')' ; -graph_data = 'GRAPH', '(', identifier, ',', identifier, ')' ; (* edge_type, target_hexad_id *) -tensor_data = 'TENSOR', '(', array_literal, ')' ; -semantic_data = 'SEMANTIC', '(', identifier, ')' ; (* contract name *) -temporal_data = 'TEMPORAL', '(', timestamp, ')' ; -provenance_data = 'PROVENANCE', '(', field_assignment_list, ')' ; (* actor, event_type, description, source *) -spatial_data = 'SPATIAL', '(', field_assignment_list, ')' ; (* latitude, longitude, altitude, geometry_type, srid *) - -field_assignment_list = field_assignment, { ',', field_assignment } ; -field_assignment = identifier, '=', literal ; - -(* UPDATE: Modify fields of an existing hexad *) -update_mutation = 'UPDATE', 'HEXAD', uuid, 'SET', set_list, [proof_clause] ; - -set_list = set_assignment, { ',', set_assignment } ; -set_assignment = field_ref, '=', literal ; - -(* DELETE: Remove a hexad *) -delete_mutation = 'DELETE', 'HEXAD', uuid, [proof_clause] ; - -(* ============================================================================ - 10. LEXICAL ELEMENTS - ============================================================================ *) -uuid = hex_digit, hex_digit, hex_digit, hex_digit, hex_digit, hex_digit, hex_digit, hex_digit, '-', - hex_digit, hex_digit, hex_digit, hex_digit, '-', - hex_digit, hex_digit, hex_digit, hex_digit, '-', - hex_digit, hex_digit, hex_digit, hex_digit, '-', - hex_digit, hex_digit, hex_digit, hex_digit, hex_digit, hex_digit, hex_digit, hex_digit, hex_digit, hex_digit, hex_digit, hex_digit ; -hex_digit = '0' | '1' | '2' | '3' | '4' | '5' | '6' | '7' | '8' | '9' - | 'a' | 'b' | 'c' | 'd' | 'e' | 'f' - | 'A' | 'B' | 'C' | 'D' | 'E' | 'F' ; - -identifier = ? letter or underscore ?, { ? letter, digit, or underscore ? } ; -store_id = identifier ; -verifier_id = identifier ; -actor_id = identifier ; -version_id = identifier ; - -integer = ? digit ?, { ? digit ? } ; -float = ? digit ?, { ? digit ? }, '.', ? digit ?, { ? digit ? }, [('e' | 'E'), ['+' | '-'], ? digit ?, { ? digit ? }] ; - -string_literal = "'", { ? any character except ' and \ ? | "\'", | "\\" }, "'" ; -regex_literal = '/', { ? any character except / and \ ? | '\/' | '\\' }, '/' ; - -timestamp = iso8601_datetime ; -iso8601_datetime = year, '-', month, '-', day, 'T', hour, ':', minute, ':', second, ['.', fraction], timezone ; - -glob_pattern = '/', path_segment, { '/', path_segment }, ('/*' | '/**') ; -path_segment = ? alphanumeric, underscore, or hyphen ?, { ? alphanumeric, underscore, or hyphen ? } ; - -array_literal = '[', literal, { ',', literal }, ']' ; -scalar_literal = integer | float | string_literal | boolean ; -boolean = 'true' | 'false' ; -literal = scalar_literal | array_literal ; - -(* ============================================================================ - 11. COMMENTS - ============================================================================ *) -comment = '--', { ? any character except newline ? }, '\n' (* Line comment *) - | '/*', ? any characters ?, '*/' (* Block comment *) ; - -(* ============================================================================ - 12. RESERVED KEYWORDS - ============================================================================ *) -(* Keywords are case-insensitive in VCL *) -keywords = 'SELECT' | 'FROM' | 'WHERE' | 'PROOF' | 'LIMIT' | 'OFFSET' - | 'GRAPH' | 'VECTOR' | 'TENSOR' | 'SEMANTIC' | 'DOCUMENT' | 'TEMPORAL' | 'PROVENANCE' | 'SPATIAL' - | 'HEXAD' | 'FEDERATION' | 'STORE' - | 'WITH' | 'DRIFT' | 'STRICT' | 'REPAIR' | 'TOLERATE' | 'LATEST' - | 'AND' | 'OR' | 'NOT' - | 'SIMILAR' | 'TO' | 'WITHIN' | 'NEAREST' | 'USING' - | 'SATISFIES' | 'HAS' | 'VERIFIED' | 'BY' - | 'FULLTEXT' | 'CONTAINS' | 'MATCHES' | 'FIELD' | 'LIKE' - | 'AS' | 'OF' | 'BETWEEN' | 'VERSION' | 'MODIFIED' - | 'EXISTENCE' | 'CITATION' | 'ACCESS' | 'INTEGRITY' | 'CUSTOM' - | 'COSINE' | 'EUCLIDEAN' | 'DOT_PRODUCT' | 'JACCARD' | 'SHAPE' | 'RANK' - | 'RADIUS' | 'BOUNDS' - | 'ORDER' | 'GROUP' | 'HAVING' | 'ASC' | 'DESC' - | 'COUNT' | 'SUM' | 'AVG' | 'MIN' | 'MAX' - | 'INSERT' | 'UPDATE' | 'DELETE' | 'SET' - | 'EXISTS' | 'CONSISTENT' - | 'true' | 'false' ; - -(* ============================================================================ - END OF GRAMMAR - ============================================================================ *) diff --git a/verisimdb/spec/system-specs.md b/verisimdb/spec/system-specs.md deleted file mode 100644 index e139517f..00000000 --- a/verisimdb/spec/system-specs.md +++ /dev/null @@ -1,179 +0,0 @@ -# SPDX-License-Identifier: CC-BY-SA-4.0 -# Copyright (c) 2026 Jonathan D.A. Jewell (hyperpolymath) <j.d.a.jewell@open.ac.uk> - -# VeriSimDB System Specifications - -VeriSimDB is a multi-modal verification database with 8 modality stores and -built-in proof verification. Implementation stack: Rust core storage engine, -Elixir/OTP API layer, ReScript frontend. - ---- - -## Memory Model - -VeriSimDB's memory model spans three layers, each with distinct ownership -and allocation strategies. - -### Rust Storage Engine - -- **B-tree indices**: Rust-owned `BTreeMap` variants with custom page sizes. - Pages are allocated via a slab allocator for predictable latency. -- **Vector indices**: Dense f32/f64 arrays for similarity search, allocated - as contiguous `Vec<f32>` buffers. SIMD-aligned to 32-byte boundaries. -- **Write-ahead log (WAL)**: Memory-mapped file (`mmap`) with append-only - writes. Rust owns the mapping lifetime via `MmapMut`. -- **Buffer pool**: Fixed-size page cache (configurable, default 256 MB). - Pages are reference-counted (`Arc<Page>`) with LRU eviction. - -### 8 Modality Stores - -Each modality store manages its own memory independently: - -| Modality | Storage Type | Memory Strategy | -|----------------|-------------------------|------------------------------| -| Textual | B-tree + inverted index | Slab-allocated postings | -| Numeric | B-tree | Inline leaf values | -| Temporal | Interval tree | Arena-allocated nodes | -| Spatial | R-tree | Page-based with bulk loading | -| Vector | HNSW graph | Contiguous f32 buffers | -| Graph | Adjacency lists | CSR format, arena-allocated | -| Provenance | Merkle DAG | Hash-addressed content store | -| Categorical | Bitmap index | Roaring bitmaps | - -### Elixir API Layer - -- Elixir/BEAM manages all API-layer memory via its per-process heap GC. -- Each Elixir process has an isolated heap — no shared mutable state. -- Query results crossing from Rust to Elixir are serialised as Erlang terms - via NIF (Native Implemented Function) calls. -- Large result sets use resource objects (`enif_alloc_resource`) to avoid - copying — Rust retains ownership, Elixir holds a reference. - -### ReScript Frontend - -- ReScript compiles to JavaScript; browser GC manages all frontend memory. -- Query results are received as JSON over HTTP/WebSocket. -- No direct memory sharing between frontend and backend. - ---- - -## Concurrency Model - -VeriSimDB uses a hybrid concurrency model combining Elixir's actor model -with Rust's async runtime. - -### Elixir/OTP Actor Model (Query Processing) - -- Each incoming query spawns a dedicated Elixir process (lightweight, ~2 KB). -- Query planning, optimisation, and result assembly happen in Elixir processes. -- OTP supervisors manage process lifecycles and restart on failure. -- GenServer processes manage connection pools to the Rust storage engine. - -### Rust Tokio Runtime (Storage I/O) - -- The Rust storage engine runs on a multi-threaded `tokio` runtime. -- Read operations use shared locks (`RwLock<T>`) on B-tree pages. -- Write operations acquire exclusive locks with WAL-based crash recovery. -- Background tasks (compaction, index rebuilding) run as spawned tokio tasks - with lower priority. - -### Cross-Modal Query Coordination - -- Queries spanning multiple modalities are decomposed by the Elixir query - planner into per-modality sub-queries. -- Sub-queries execute concurrently (one tokio task per modality store). -- Results are collected via `tokio::sync::mpsc` channels. -- The Elixir process assembles final results from all modality responses. -- Join operations across modalities use hash-join or merge-join depending - on the query planner's cost estimate. - -### Consistency Model - -- Single-modality operations are serialisable (WAL + exclusive write locks). -- Cross-modal transactions use a two-phase commit protocol coordinated by - the Elixir transaction manager. -- Read snapshots use MVCC — readers never block writers. - ---- - -## Effect System - -VeriSimDB's effect system centres on proof verification. Every query result -carries proof metadata. - -### Proof Effects - -Proof verification occurs during query execution, not after: - -| Proof Type | Verification | Cost | -|---------------|------------------------------------------------|----------| -| `EXISTENCE` | Merkle proof that the record exists in the store| O(log n) | -| `INTEGRITY` | Hash chain verification of record contents | O(1) | -| `CITATION` | Provenance DAG traversal to source records | O(d) | -| `TEMPORAL` | Timestamp ordering proof via interval tree | O(log n) | -| `SPATIAL` | Bounding-box containment proof via R-tree | O(log n) | - -### Query Proof Composition - -- A cross-modal query produces a **composite proof**: one sub-proof per - modality involved. -- Composite proofs are serialised as a Merkle tree of sub-proofs. -- The root hash of the composite proof is returned with the query result. - -### Proof Verification Modes - -| Mode | Behaviour | -|-------------|--------------------------------------------------| -| `Strict` | All proof types verified; query fails on any failure | -| `Optimistic`| Proofs computed but verification deferred to client | -| `None` | No proofs computed (performance mode) | - -### Effect Tracking in the Elixir Layer - -- Each query carries an effect context (`%ProofContext{}` struct). -- The context accumulates proof obligations as the query plan executes. -- NIF calls to the Rust engine return proof artifacts alongside data. -- The Elixir layer assembles the final proof tree before returning results. - ---- - -## Module System - -VeriSimDB's module system reflects its polyglot architecture. - -### Rust Crate Workspace - -| Crate | Responsibility | -|--------------------|------------------------------------------------| -| `verisimdb-core` | Storage engine, page management, WAL | -| `verisimdb-index` | B-tree, R-tree, HNSW, bitmap index implementations | -| `verisimdb-proof` | Merkle proofs, hash chains, proof composition | -| `verisimdb-nif` | Erlang NIF bindings (Rustler) | -| `verisimdb-query` | Query IR, physical operators, execution engine | - -### Elixir OTP Application - -| Module | Responsibility | -|---------------------------|-----------------------------------------| -| `VeriSimDB.Application` | OTP application entry, supervisor tree | -| `VeriSimDB.QueryPlanner` | SQL-like query parsing and planning | -| `VeriSimDB.ModalRouter` | Routes sub-queries to modality stores | -| `VeriSimDB.ProofAssembler`| Collects and composes modality proofs | -| `VeriSimDB.Connection` | NIF connection pool to Rust engine | -| `VeriSimDB.Transaction` | Two-phase commit coordinator | - -### ReScript Frontend - -| Module | Responsibility | -|----------------------|--------------------------------------------| -| `QueryBuilder` | Type-safe query construction | -| `ProofViewer` | Proof tree visualisation component | -| `ModalitySelector` | UI for selecting query modalities | -| `ResultTable` | Tabular result display with proof badges | - -### Inter-Layer Communication - -- **Rust <-> Elixir**: Erlang NIFs via Rustler. Binary protocol with - zero-copy where possible (resource objects). -- **Elixir <-> ReScript**: JSON over HTTP (REST) or WebSocket (streaming - results). Phoenix Channels for live query subscriptions. diff --git a/verisimdb/src/abi/Foreign.idr b/verisimdb/src/abi/Foreign.idr deleted file mode 100644 index 8cdb633e..00000000 --- a/verisimdb/src/abi/Foreign.idr +++ /dev/null @@ -1,305 +0,0 @@ -||| SPDX-License-Identifier: MPL-2.0 -||| VeriSimDB Foreign Function Interface Declarations -||| -||| All C-compatible functions implemented in ffi/zig/. -||| These are the canonical FFI entry points for VeriSimDB. -||| -||| Naming convention: verisimdb_<operation> -||| All functions return VResult codes (Bits32). - -module VeriSimDB.ABI.Foreign - -import VeriSimDB.ABI.Types -import VeriSimDB.ABI.Layout - -%default total - --------------------------------------------------------------------------------- --- Library Lifecycle --------------------------------------------------------------------------------- - -||| Initialize a VeriSimDB instance with configuration -||| config_ptr: pointer to VDBConfig struct -||| Returns: handle to VeriSimDB instance (0 on failure) -export -%foreign "C:verisimdb_init, libverisimdb" -prim__init : Bits64 -> PrimIO Bits64 - -||| Safe wrapper for initialization -export -init : Bits64 -> IO (Maybe VDBHandle) -init configPtr = do - ptr <- primIO (prim__init configPtr) - pure (createVDBHandle ptr) - -||| Shut down and free all resources -export -%foreign "C:verisimdb_free, libverisimdb" -prim__free : Bits64 -> PrimIO () - -||| Safe shutdown -export -free : VDBHandle -> IO () -free h = primIO (prim__free (vdbPtr h)) - --------------------------------------------------------------------------------- --- Entity Operations --------------------------------------------------------------------------------- - -||| Create a new octad entity -||| db: VDBHandle, id_high/id_low: EntityId parts, mask: active modalities -||| Returns: EntityHandle (0 on failure) -export -%foreign "C:verisimdb_entity_create, libverisimdb" -prim__entityCreate : Bits64 -> Bits64 -> Bits64 -> Bits8 -> PrimIO Bits64 - -||| Safe entity creation -export -entityCreate : VDBHandle -> EntityId -> ModalityMask -> IO (Maybe EntityHandle) -entityCreate db eid mask = do - ptr <- primIO (prim__entityCreate (vdbPtr db) eid.high eid.low mask) - pure (createEntityHandle ptr) - -||| Look up an entity by ID -export -%foreign "C:verisimdb_entity_get, libverisimdb" -prim__entityGet : Bits64 -> Bits64 -> Bits64 -> PrimIO Bits64 - -||| Safe entity lookup -export -entityGet : VDBHandle -> EntityId -> IO (Maybe EntityHandle) -entityGet db eid = do - ptr <- primIO (prim__entityGet (vdbPtr db) eid.high eid.low) - pure (createEntityHandle ptr) - -||| Delete an entity and all its modality data -export -%foreign "C:verisimdb_entity_delete, libverisimdb" -prim__entityDelete : Bits64 -> Bits64 -> Bits64 -> PrimIO Bits32 - -||| Safe entity deletion -export -entityDelete : VDBHandle -> EntityId -> IO VResult -entityDelete db eid = do - code <- primIO (prim__entityDelete (vdbPtr db) eid.high eid.low) - pure $ case vresultFromInt code of - Just r => r - Nothing => VError - -||| Release an entity handle (does NOT delete the entity) -export -%foreign "C:verisimdb_entity_handle_free, libverisimdb" -prim__entityHandleFree : Bits64 -> PrimIO () - -export -entityHandleFree : EntityHandle -> IO () -entityHandleFree h = primIO (prim__entityHandleFree (entityPtr h)) - --------------------------------------------------------------------------------- --- Modality Data Operations --------------------------------------------------------------------------------- - -||| Write modality data for an entity -||| entity: EntityHandle, slice_ptr: pointer to ModalitySlice struct -||| Returns: VResult code -export -%foreign "C:verisimdb_modality_write, libverisimdb" -prim__modalityWrite : Bits64 -> Bits64 -> PrimIO Bits32 - -||| Safe modality write -export -modalityWrite : EntityHandle -> Bits64 -> IO VResult -modalityWrite entity slicePtr = do - code <- primIO (prim__modalityWrite (entityPtr entity) slicePtr) - pure $ case vresultFromInt code of - Just r => r - Nothing => VError - -||| Read modality data for an entity -||| entity: EntityHandle, modality: Bits32, out_ptr/out_len: output buffer -||| Returns: bytes written (0 on error) -export -%foreign "C:verisimdb_modality_read, libverisimdb" -prim__modalityRead : Bits64 -> Bits32 -> Bits64 -> Bits64 -> PrimIO Bits64 - -||| Safe modality read -export -modalityRead : EntityHandle -> Modality -> Bits64 -> Bits64 -> IO Bits64 -modalityRead entity mod outPtr outLen = - primIO (prim__modalityRead (entityPtr entity) (modalityToInt mod) outPtr outLen) - -||| Get active modality mask for an entity -export -%foreign "C:verisimdb_entity_modalities, libverisimdb" -prim__entityModalities : Bits64 -> PrimIO Bits8 - -export -entityModalities : EntityHandle -> IO ModalityMask -entityModalities entity = primIO (prim__entityModalities (entityPtr entity)) - --------------------------------------------------------------------------------- --- Drift Detection --------------------------------------------------------------------------------- - -||| Check drift between two modalities of an entity -||| Returns: DriftScore (fixed-point * 10000), or 0xFFFFFFFF on error -export -%foreign "C:verisimdb_drift_check, libverisimdb" -prim__driftCheck : Bits64 -> Bits32 -> Bits32 -> Bits32 -> PrimIO Bits32 - -||| Safe drift check -export -driftCheck : EntityHandle -> Modality -> Modality -> DriftMethod -> IO (Maybe DriftScore) -driftCheck entity src tgt method = do - score <- primIO (prim__driftCheck (entityPtr entity) - (modalityToInt src) (modalityToInt tgt) (driftMethodToInt method)) - pure $ if score == 0xFFFFFFFF then Nothing else Just score - -||| Run drift detection sweep on all entities -||| db: VDBHandle, report_buf: pointer to array of DriftReport structs -||| report_max: max reports to write -||| Returns: number of drift reports written -export -%foreign "C:verisimdb_drift_sweep, libverisimdb" -prim__driftSweep : Bits64 -> Bits64 -> Bits32 -> PrimIO Bits32 - -export -driftSweep : VDBHandle -> Bits64 -> Bits32 -> IO Bits32 -driftSweep db reportBuf maxReports = - primIO (prim__driftSweep (vdbPtr db) reportBuf maxReports) - -||| Trigger normalization for a drifted entity -export -%foreign "C:verisimdb_normalize, libverisimdb" -prim__normalize : Bits64 -> Bits64 -> Bits64 -> PrimIO Bits32 - -export -normalize : VDBHandle -> EntityId -> IO VResult -normalize db eid = do - code <- primIO (prim__normalize (vdbPtr db) eid.high eid.low) - pure $ case vresultFromInt code of - Just r => r - Nothing => VError - --------------------------------------------------------------------------------- --- VCL Query Execution --------------------------------------------------------------------------------- - -||| Parse and execute a VCL query -||| db: VDBHandle, req_ptr: pointer to QueryRequest struct -||| Returns: ResultSetHandle (0 on failure) -export -%foreign "C:verisimdb_query, libverisimdb" -prim__query : Bits64 -> Bits64 -> PrimIO Bits64 - -||| Safe query execution -export -query : VDBHandle -> Bits64 -> IO (Maybe ResultSetHandle) -query db reqPtr = do - ptr <- primIO (prim__query (vdbPtr db) reqPtr) - if ptr == 0 - then pure Nothing - else pure (Just (MkResultSetHandle ptr)) - -||| Get number of results in a result set -export -%foreign "C:verisimdb_resultset_count, libverisimdb" -prim__resultSetCount : Bits64 -> PrimIO Bits64 - -export -resultSetCount : ResultSetHandle -> IO Bits64 -resultSetCount rs = primIO (prim__resultSetCount (resultSetPtr rs)) - -||| Read result at index as JSON bytes into buffer -||| Returns: bytes written (0 on error or out of bounds) -export -%foreign "C:verisimdb_resultset_get, libverisimdb" -prim__resultSetGet : Bits64 -> Bits64 -> Bits64 -> Bits64 -> PrimIO Bits64 - -export -resultSetGet : ResultSetHandle -> Bits64 -> Bits64 -> Bits64 -> IO Bits64 -resultSetGet rs idx outPtr outLen = - primIO (prim__resultSetGet (resultSetPtr rs) idx outPtr outLen) - -||| Free a result set -export -%foreign "C:verisimdb_resultset_free, libverisimdb" -prim__resultSetFree : Bits64 -> PrimIO () - -export -resultSetFree : ResultSetHandle -> IO () -resultSetFree rs = primIO (prim__resultSetFree (resultSetPtr rs)) - --------------------------------------------------------------------------------- --- Transaction Support --------------------------------------------------------------------------------- - -||| Begin a transaction -export -%foreign "C:verisimdb_txn_begin, libverisimdb" -prim__txnBegin : Bits64 -> PrimIO Bits64 - -export -txnBegin : VDBHandle -> IO (Maybe TxnHandle) -txnBegin db = do - ptr <- primIO (prim__txnBegin (vdbPtr db)) - if ptr == 0 - then pure Nothing - else pure (Just (MkTxnHandle ptr)) - -||| Commit a transaction -export -%foreign "C:verisimdb_txn_commit, libverisimdb" -prim__txnCommit : Bits64 -> PrimIO Bits32 - -export -txnCommit : TxnHandle -> IO VResult -txnCommit txn = do - code <- primIO (prim__txnCommit (txnPtr txn)) - pure $ case vresultFromInt code of - Just r => r - Nothing => VError - -||| Rollback a transaction -export -%foreign "C:verisimdb_txn_rollback, libverisimdb" -prim__txnRollback : Bits64 -> PrimIO Bits32 - -export -txnRollback : TxnHandle -> IO VResult -txnRollback txn = do - code <- primIO (prim__txnRollback (txnPtr txn)) - pure $ case vresultFromInt code of - Just r => r - Nothing => VError - --------------------------------------------------------------------------------- --- Version Information --------------------------------------------------------------------------------- - -||| Get VeriSimDB version string -export -%foreign "C:verisimdb_version, libverisimdb" -prim__version : PrimIO Bits64 - -||| Get version as string (caller must not free) -export -%foreign "support:idris2_getString, libidris2_support" -prim__getString : Bits64 -> String - -export -version : IO String -version = do - ptr <- primIO prim__version - pure (prim__getString ptr) - -||| Get build info (features, platform) -export -%foreign "C:verisimdb_build_info, libverisimdb" -prim__buildInfo : PrimIO Bits64 - -export -buildInfo : IO String -buildInfo = do - ptr <- primIO prim__buildInfo - pure (prim__getString ptr) diff --git a/verisimdb/src/abi/Layout.idr b/verisimdb/src/abi/Layout.idr deleted file mode 100644 index ddcc5c0c..00000000 --- a/verisimdb/src/abi/Layout.idr +++ /dev/null @@ -1,220 +0,0 @@ -||| SPDX-License-Identifier: MPL-2.0 -||| VeriSimDB Memory Layout Proofs -||| -||| Formal proofs about memory layout, alignment, and padding for -||| VeriSimDB's C-compatible structs passed across the FFI boundary. -||| -||| Key structs: EntityId (16B), DriftReport (40B), ModalitySlice (24B), -||| VDBConfig (32B). - -module VeriSimDB.ABI.Layout - -import VeriSimDB.ABI.Types -import Data.Vect -import Data.So - -%default total - --------------------------------------------------------------------------------- --- Alignment Utilities --------------------------------------------------------------------------------- - -||| Calculate padding needed for alignment -public export -paddingFor : (offset : Nat) -> (alignment : Nat) -> Nat -paddingFor offset alignment = - if offset `mod` alignment == 0 - then 0 - else alignment - (offset `mod` alignment) - -||| Round up to next alignment boundary -public export -alignUp : (size : Nat) -> (alignment : Nat) -> Nat -alignUp size alignment = size + paddingFor size alignment - -||| Proof that alignment divides aligned size -public export -data Divides : Nat -> Nat -> Type where - DivideBy : (k : Nat) -> {n : Nat} -> {m : Nat} -> (m = k * n) -> Divides n m - --------------------------------------------------------------------------------- --- Struct Field Layout --------------------------------------------------------------------------------- - -||| A field in a struct with its offset and size -public export -record Field where - constructor MkField - name : String - offset : Nat - size : Nat - alignment : Nat - -||| Calculate the offset of the next field -public export -nextFieldOffset : Field -> Nat -> Nat -nextFieldOffset f nextAlign = alignUp (f.offset + f.size) nextAlign - -||| A struct layout is a vector of fields with size and alignment -public export -record StructLayout where - constructor MkStructLayout - layoutName : String - fields : List Field - totalSize : Nat - alignment : Nat - -||| Proof that field offsets are correctly aligned -public export -data FieldAligned : Field -> Type where - IsAligned : (f : Field) -> (0 _ : Divides f.alignment f.offset) -> FieldAligned f - --------------------------------------------------------------------------------- --- EntityId Layout (16 bytes, 8-byte aligned) --------------------------------------------------------------------------------- - -||| EntityId: two Bits64 fields = 16 bytes, no padding needed -public export -entityIdLayout : StructLayout -entityIdLayout = MkStructLayout "EntityId" - [ MkField "high" 0 8 8 -- Bits64 at offset 0 - , MkField "low" 8 8 8 -- Bits64 at offset 8 - ] - 16 -- total size - 8 -- alignment - -||| Proof: EntityId high field is aligned (offset 0 divides by 8) -public export -entityIdHighAligned : FieldAligned (MkField "high" 0 8 8) -entityIdHighAligned = IsAligned _ (DivideBy 0 Refl) - -||| Proof: EntityId low field is aligned (offset 8 divides by 8) -public export -entityIdLowAligned : FieldAligned (MkField "low" 8 8 8) -entityIdLowAligned = IsAligned _ (DivideBy 1 Refl) - --------------------------------------------------------------------------------- --- DriftReport Layout (40 bytes, 8-byte aligned) --------------------------------------------------------------------------------- - -||| DriftReport: sent from Rust drift detector to Elixir/Zig consumers -||| Fields: -||| entity_id : EntityId (16 bytes at offset 0) -||| source_mod : Bits32 (4 bytes at offset 16) -||| target_mod : Bits32 (4 bytes at offset 20) -||| drift_score : Bits32 (4 bytes at offset 24, fixed-point * 10000) -||| method : Bits32 (4 bytes at offset 28) -||| timestamp : Bits64 (8 bytes at offset 32) -public export -driftReportLayout : StructLayout -driftReportLayout = MkStructLayout "DriftReport" - [ MkField "entity_id_high" 0 8 8 - , MkField "entity_id_low" 8 8 8 - , MkField "source_mod" 16 4 4 - , MkField "target_mod" 20 4 4 - , MkField "drift_score" 24 4 4 - , MkField "method" 28 4 4 - , MkField "timestamp" 32 8 8 - ] - 40 - 8 - --------------------------------------------------------------------------------- --- ModalitySlice Layout (24 bytes, 8-byte aligned) --------------------------------------------------------------------------------- - -||| ModalitySlice: pointer + length to modality-specific data buffer -||| Used when reading/writing individual modality data across FFI -||| data_ptr : Bits64 (8 bytes at offset 0) -||| data_len : Bits64 (8 bytes at offset 8) -||| modality : Bits32 (4 bytes at offset 16) -||| flags : Bits32 (4 bytes at offset 20) -public export -modalitySliceLayout : StructLayout -modalitySliceLayout = MkStructLayout "ModalitySlice" - [ MkField "data_ptr" 0 8 8 - , MkField "data_len" 8 8 8 - , MkField "modality" 16 4 4 - , MkField "flags" 20 4 4 - ] - 24 - 8 - --------------------------------------------------------------------------------- --- VDBConfig Layout (32 bytes, 8-byte aligned) --------------------------------------------------------------------------------- - -||| VDBConfig: configuration passed to verisimdb_init -||| max_entities : Bits64 (8 bytes at offset 0) -||| drift_threshold : Bits32 (4 bytes at offset 8, fixed-point * 10000) -||| modality_mask : Bits8 (1 byte at offset 12) -||| enable_wal : Bits8 (1 byte at offset 13) -||| enable_telemetry : Bits8 (1 byte at offset 14) -||| _pad1 : Bits8 (1 byte at offset 15) -||| data_dir_ptr : Bits64 (8 bytes at offset 16, pointer to C string) -||| data_dir_len : Bits64 (8 bytes at offset 24) -public export -vdbConfigLayout : StructLayout -vdbConfigLayout = MkStructLayout "VDBConfig" - [ MkField "max_entities" 0 8 8 - , MkField "drift_threshold" 8 4 4 - , MkField "modality_mask" 12 1 1 - , MkField "enable_wal" 13 1 1 - , MkField "enable_telemetry" 14 1 1 - , MkField "_pad1" 15 1 1 - , MkField "data_dir_ptr" 16 8 8 - , MkField "data_dir_len" 24 8 8 - ] - 32 - 8 - --------------------------------------------------------------------------------- --- QueryRequest Layout (32 bytes, 8-byte aligned) --------------------------------------------------------------------------------- - -||| QueryRequest: VCL query submitted across FFI -||| vcl_ptr : Bits64 (8 bytes at offset 0, pointer to UTF-8 VCL string) -||| vcl_len : Bits64 (8 bytes at offset 8) -||| timeout_ms : Bits32 (4 bytes at offset 16) -||| proof_type : Bits32 (4 bytes at offset 20, 0xFF = no proof requested) -||| txn_handle : Bits64 (8 bytes at offset 24, 0 = auto-commit) -public export -queryRequestLayout : StructLayout -queryRequestLayout = MkStructLayout "QueryRequest" - [ MkField "vcl_ptr" 0 8 8 - , MkField "vcl_len" 8 8 8 - , MkField "timeout_ms" 16 4 4 - , MkField "proof_type" 20 4 4 - , MkField "txn_handle" 24 8 8 - ] - 32 - 8 - --------------------------------------------------------------------------------- --- Layout Verification --------------------------------------------------------------------------------- - -||| Verify that a field offset + size fits within the struct -public export -fieldFitsInStruct : (layout : StructLayout) -> (f : Field) -> - So (f.offset + f.size <= layout.totalSize) -> - () -fieldFitsInStruct _ _ _ = () - -||| Verify no field overlap: field2 starts at or after field1 ends -public export -data NoOverlap : Field -> Field -> Type where - FieldsDisjoint : (f1 : Field) -> (f2 : Field) -> - {auto 0 prf : So (f1.offset + f1.size <= f2.offset)} -> - NoOverlap f1 f2 - -||| All VeriSimDB layouts collected for batch verification -public export -allLayouts : List StructLayout -allLayouts = - [ entityIdLayout - , driftReportLayout - , modalitySliceLayout - , vdbConfigLayout - , queryRequestLayout - ] diff --git a/verisimdb/src/abi/Types.idr b/verisimdb/src/abi/Types.idr deleted file mode 100644 index 9e0bdec5..00000000 --- a/verisimdb/src/abi/Types.idr +++ /dev/null @@ -1,348 +0,0 @@ -||| SPDX-License-Identifier: MPL-2.0 -||| VeriSimDB ABI Type Definitions -||| -||| Formal type definitions for the VeriSimDB cross-modal entity engine. -||| All types include dependent-type proofs of correctness for C ABI compatibility. -||| -||| The octad model: each entity exists simultaneously across 8 modalities -||| (Graph, Vector, Tensor, Semantic, Document, Temporal, Provenance, Spatial). - -module VeriSimDB.ABI.Types - -import Data.Bits -import Data.So -import Data.Vect - -%default total - --------------------------------------------------------------------------------- --- Platform Detection --------------------------------------------------------------------------------- - -||| Supported platforms for the VeriSimDB ABI -public export -data Platform = Linux | Windows | MacOS | BSD | WASM - -||| Compile-time platform detection -public export -thisPlatform : Platform -thisPlatform = Linux -- Default; override with compiler flags - --------------------------------------------------------------------------------- --- Result Codes --------------------------------------------------------------------------------- - -||| Result codes for FFI operations (C-compatible integers) -public export -data VResult : Type where - ||| Operation succeeded - VOk : VResult - ||| Generic error - VError : VResult - ||| Invalid parameter - VInvalidParam : VResult - ||| Out of memory - VOutOfMemory : VResult - ||| Null pointer encountered - VNullPointer : VResult - ||| Entity not found - VNotFound : VResult - ||| Modality not available for entity - VModalityUnavailable : VResult - ||| Drift threshold exceeded - VDriftExceeded : VResult - ||| Query parse error - VQueryParseError : VResult - ||| Transaction conflict - VTxnConflict : VResult - -||| Convert VResult to C integer -public export -vresultToInt : VResult -> Bits32 -vresultToInt VOk = 0 -vresultToInt VError = 1 -vresultToInt VInvalidParam = 2 -vresultToInt VOutOfMemory = 3 -vresultToInt VNullPointer = 4 -vresultToInt VNotFound = 5 -vresultToInt VModalityUnavailable = 6 -vresultToInt VDriftExceeded = 7 -vresultToInt VQueryParseError = 8 -vresultToInt VTxnConflict = 9 - -||| Convert C integer back to VResult -public export -vresultFromInt : Bits32 -> Maybe VResult -vresultFromInt 0 = Just VOk -vresultFromInt 1 = Just VError -vresultFromInt 2 = Just VInvalidParam -vresultFromInt 3 = Just VOutOfMemory -vresultFromInt 4 = Just VNullPointer -vresultFromInt 5 = Just VNotFound -vresultFromInt 6 = Just VModalityUnavailable -vresultFromInt 7 = Just VDriftExceeded -vresultFromInt 8 = Just VQueryParseError -vresultFromInt 9 = Just VTxnConflict -vresultFromInt _ = Nothing - -||| VResults are decidably equal -public export -DecEq VResult where - decEq VOk VOk = Yes Refl - decEq VError VError = Yes Refl - decEq VInvalidParam VInvalidParam = Yes Refl - decEq VOutOfMemory VOutOfMemory = Yes Refl - decEq VNullPointer VNullPointer = Yes Refl - decEq VNotFound VNotFound = Yes Refl - decEq VModalityUnavailable VModalityUnavailable = Yes Refl - decEq VDriftExceeded VDriftExceeded = Yes Refl - decEq VQueryParseError VQueryParseError = Yes Refl - decEq VTxnConflict VTxnConflict = Yes Refl - decEq _ _ = No absurd - --------------------------------------------------------------------------------- --- Opaque Handles --------------------------------------------------------------------------------- - -||| Opaque handle to a VeriSimDB instance -public export -data VDBHandle : Type where - MkVDBHandle : (ptr : Bits64) -> {auto 0 nonNull : So (ptr /= 0)} -> VDBHandle - -||| Opaque handle to an octad entity -public export -data EntityHandle : Type where - MkEntityHandle : (ptr : Bits64) -> {auto 0 nonNull : So (ptr /= 0)} -> EntityHandle - -||| Opaque handle to a VCL query -public export -data QueryHandle : Type where - MkQueryHandle : (ptr : Bits64) -> {auto 0 nonNull : So (ptr /= 0)} -> QueryHandle - -||| Opaque handle to a transaction -public export -data TxnHandle : Type where - MkTxnHandle : (ptr : Bits64) -> {auto 0 nonNull : So (ptr /= 0)} -> TxnHandle - -||| Opaque handle to a query result set -public export -data ResultSetHandle : Type where - MkResultSetHandle : (ptr : Bits64) -> {auto 0 nonNull : So (ptr /= 0)} -> ResultSetHandle - -||| Safely create a handle from a raw pointer -public export -createVDBHandle : Bits64 -> Maybe VDBHandle -createVDBHandle 0 = Nothing -createVDBHandle ptr = Just (MkVDBHandle ptr) - -public export -createEntityHandle : Bits64 -> Maybe EntityHandle -createEntityHandle 0 = Nothing -createEntityHandle ptr = Just (MkEntityHandle ptr) - -public export -createQueryHandle : Bits64 -> Maybe QueryHandle -createQueryHandle 0 = Nothing -createQueryHandle ptr = Just (MkQueryHandle ptr) - -||| Extract pointer from handle -public export -vdbPtr : VDBHandle -> Bits64 -vdbPtr (MkVDBHandle ptr) = ptr - -public export -entityPtr : EntityHandle -> Bits64 -entityPtr (MkEntityHandle ptr) = ptr - -public export -queryPtr : QueryHandle -> Bits64 -queryPtr (MkQueryHandle ptr) = ptr - -public export -txnPtr : TxnHandle -> Bits64 -txnPtr (MkTxnHandle ptr) = ptr - -public export -resultSetPtr : ResultSetHandle -> Bits64 -resultSetPtr (MkResultSetHandle ptr) = ptr - --------------------------------------------------------------------------------- --- Modality Enumeration --------------------------------------------------------------------------------- - -||| The 8 modalities of an octad entity -public export -data Modality - = Graph - | Vector - | Tensor - | Semantic - | Document - | Temporal - | Provenance - | Spatial - -||| Modality to C integer mapping (stable ABI) -public export -modalityToInt : Modality -> Bits32 -modalityToInt Graph = 0 -modalityToInt Vector = 1 -modalityToInt Tensor = 2 -modalityToInt Semantic = 3 -modalityToInt Document = 4 -modalityToInt Temporal = 5 -modalityToInt Provenance = 6 -modalityToInt Spatial = 7 - -||| C integer to modality -public export -modalityFromInt : Bits32 -> Maybe Modality -modalityFromInt 0 = Just Graph -modalityFromInt 1 = Just Vector -modalityFromInt 2 = Just Tensor -modalityFromInt 3 = Just Semantic -modalityFromInt 4 = Just Document -modalityFromInt 5 = Just Temporal -modalityFromInt 6 = Just Provenance -modalityFromInt 7 = Just Spatial -modalityFromInt _ = Nothing - -||| Total number of modalities (compile-time constant) -public export -modalityCount : Nat -modalityCount = 8 - -||| Proof that modalityCount equals the number of constructors -public export -modalityCountCorrect : modalityCount = length [Graph, Vector, Tensor, Semantic, Document, Temporal, Provenance, Spatial] -modalityCountCorrect = Refl - --------------------------------------------------------------------------------- --- Modality Bitmask --------------------------------------------------------------------------------- - -||| Bitmask for selecting which modalities are active on an entity -||| Bit 0 = Graph, Bit 1 = Vector, ..., Bit 7 = Spatial -public export -ModalityMask : Type -ModalityMask = Bits8 - -||| Set a modality bit in the mask -public export -setModality : ModalityMask -> Modality -> ModalityMask -setModality mask mod = mask .|. (1 `shiftL` (cast (modalityToInt mod))) - -||| Check if a modality is active in the mask -public export -hasModality : ModalityMask -> Modality -> Bool -hasModality mask mod = (mask .&. (1 `shiftL` (cast (modalityToInt mod)))) /= 0 - -||| All modalities active (0xFF) -public export -allModalities : ModalityMask -allModalities = 0xFF - -||| No modalities active -public export -noModalities : ModalityMask -noModalities = 0x00 - --------------------------------------------------------------------------------- --- Drift Types --------------------------------------------------------------------------------- - -||| Drift measurement between modalities -||| Stored as fixed-point: value * 10000 (4 decimal places) -public export -DriftScore : Type -DriftScore = Bits32 - -||| Drift detection method -public export -data DriftMethod - = Cosine - | Euclidean - | DotProduct - | Jaccard - | Hamming - | Custom - -||| Drift method to C integer -public export -driftMethodToInt : DriftMethod -> Bits32 -driftMethodToInt Cosine = 0 -driftMethodToInt Euclidean = 1 -driftMethodToInt DotProduct = 2 -driftMethodToInt Jaccard = 3 -driftMethodToInt Hamming = 4 -driftMethodToInt Custom = 5 - --------------------------------------------------------------------------------- --- Entity ID --------------------------------------------------------------------------------- - -||| Entity ID is a 128-bit UUID stored as two 64-bit halves -public export -record EntityId where - constructor MkEntityId - high : Bits64 - low : Bits64 - -||| Proof that EntityId has fixed size (16 bytes) -public export -data HasSize : Type -> Nat -> Type where - SizeProof : {0 t : Type} -> {n : Nat} -> HasSize t n - -public export -entityIdSize : HasSize EntityId 16 -entityIdSize = SizeProof - --------------------------------------------------------------------------------- --- Platform-Specific Types --------------------------------------------------------------------------------- - -||| C size_t varies by platform (64-bit on most, 32-bit on WASM) -public export -CSize : Platform -> Type -CSize WASM = Bits32 -CSize _ = Bits64 - -||| Pointer size by platform -public export -ptrSize : Platform -> Nat -ptrSize WASM = 32 -ptrSize _ = 64 - --------------------------------------------------------------------------------- --- Proof Type Enumeration --------------------------------------------------------------------------------- - -||| VCL proof types supported by VeriSimDB -public export -data ProofType - = Existence - | Integrity - | Consistency - | ProvenanceProof - | Freshness - | Access - | Citation - | CustomProof - | ZKP - | Proven - | Sanctify - -||| Proof type to C integer -public export -proofTypeToInt : ProofType -> Bits32 -proofTypeToInt Existence = 0 -proofTypeToInt Integrity = 1 -proofTypeToInt Consistency = 2 -proofTypeToInt ProvenanceProof = 3 -proofTypeToInt Freshness = 4 -proofTypeToInt Access = 5 -proofTypeToInt Citation = 6 -proofTypeToInt CustomProof = 7 -proofTypeToInt ZKP = 8 -proofTypeToInt Proven = 9 -proofTypeToInt Sanctify = 10 diff --git a/verisimdb/src/registry/KRaftCluster.res b/verisimdb/src/registry/KRaftCluster.res deleted file mode 100644 index a8036add..00000000 --- a/verisimdb/src/registry/KRaftCluster.res +++ /dev/null @@ -1,549 +0,0 @@ -// SPDX-License-Identifier: MPL-2.0 -// KRaft Cluster Manager -// Drives the Raft consensus lifecycle: elections, heartbeats, client requests, -// and applies committed commands to the Registry state machine. - -// ============================================================================ -// Cluster Configuration -// ============================================================================ - -type clusterConfig = { - nodeId: MetadataLog.nodeId, - peers: array<MetadataLog.nodeId>, - electionTimeoutMinMs: int, - electionTimeoutMaxMs: int, - heartbeatIntervalMs: int, - maxBatchSize: int, -} - -let defaultConfig = (~nodeId: MetadataLog.nodeId): clusterConfig => { - { - nodeId: nodeId, - peers: [], - electionTimeoutMinMs: 150, - electionTimeoutMaxMs: 300, - heartbeatIntervalMs: 50, - maxBatchSize: 100, - } -} - -// ============================================================================ -// Cluster State -// ============================================================================ - -type electionTimer = { - timeoutMs: int, - elapsedMs: int, -} - -type pendingRequest = { - command: MetadataLog.command, - index: MetadataLog.index, - timestamp: float, -} - -type clusterState = { - config: clusterConfig, - raft: MetadataLog.nodeState, - registry: Registry.registryState, - electionTimer: electionTimer, - votesReceived: array<MetadataLog.nodeId>, - pendingRequests: array<pendingRequest>, - leaderId: option<MetadataLog.nodeId>, - // Metrics - totalCommitted: int, - totalApplied: int, - electionCount: int, -} - -// ============================================================================ -// Initialization -// ============================================================================ - -let createCluster = (~config: clusterConfig): clusterState => { - let registryConfig = Registry.defaultConfig() - - { - config: config, - raft: MetadataLog.create(~nodeId=config.nodeId), - registry: Registry.create(~config=registryConfig), - electionTimer: { - timeoutMs: config.electionTimeoutMinMs + - mod( - Belt.Int.fromFloat(Js.Date.now()), - config.electionTimeoutMaxMs - config.electionTimeoutMinMs, - ), - elapsedMs: 0, - }, - votesReceived: [], - pendingRequests: [], - leaderId: None, - totalCommitted: 0, - totalApplied: 0, - electionCount: 0, - } -} - -// ============================================================================ -// Election Management -// ============================================================================ - -/// Reset the election timer with a new random timeout. -let resetElectionTimer = (state: clusterState): clusterState => { - let range = state.config.electionTimeoutMaxMs - state.config.electionTimeoutMinMs - let jitter = mod(Belt.Int.fromFloat(Js.Date.now()), Js.Math.max_int(range, 1)) - let timeout = state.config.electionTimeoutMinMs + jitter - - { - ...state, - electionTimer: {timeoutMs: timeout, elapsedMs: 0}, - } -} - -/// Advance the election timer by deltaMs. Returns true if timed out. -let tickElectionTimer = (state: clusterState, deltaMs: int): (clusterState, bool) => { - let newElapsed = state.electionTimer.elapsedMs + deltaMs - let timedOut = newElapsed >= state.electionTimer.timeoutMs - - let newState = { - ...state, - electionTimer: {...state.electionTimer, elapsedMs: newElapsed}, - } - - (newState, timedOut) -} - -/// Start an election: become candidate, vote for self, prepare vote requests. -let startElection = (state: clusterState): (clusterState, array<(MetadataLog.nodeId, MetadataLog.voteRequest)>) => { - let raft = MetadataLog.toCandidate(state.raft) - - // Vote for self - let raft = {...raft, votedFor: Some(state.config.nodeId)} - - let voteRequest: MetadataLog.voteRequest = { - term: raft.currentTerm, - candidateId: state.config.nodeId, - lastLogIndex: MetadataLog.getLastLogIndex(raft), - lastLogTerm: MetadataLog.getLastLogTerm(raft), - } - - // Prepare requests for all peers - let requests = state.config.peers->Belt.Array.map(peer => (peer, voteRequest)) - - let newState = resetElectionTimer({ - ...state, - raft: raft, - votesReceived: [state.config.nodeId], // Self-vote - electionCount: state.electionCount + 1, - }) - - (newState, requests) -} - -/// Handle a vote response from a peer. -let handleVoteResponse = ( - state: clusterState, - fromPeer: MetadataLog.nodeId, - response: MetadataLog.voteResponse, -): (clusterState, bool) => { - // If response term is higher, step down - if response.term > state.raft.currentTerm { - let newState = { - ...state, - raft: MetadataLog.toFollower(state.raft, response.term), - votesReceived: [], - leaderId: None, - } - (resetElectionTimer(newState), false) - } else if response.voteGranted && state.raft.role == Candidate { - // Count vote - let newVotes = Belt.Array.concat(state.votesReceived, [fromPeer]) - let totalNodes = Belt.Array.length(state.config.peers) + 1 // +1 for self - let quorum = totalNodes / 2 + 1 - let wonElection = Belt.Array.length(newVotes) >= quorum - - let newState = if wonElection { - // Become leader - let raft = MetadataLog.toLeader(state.raft, state.config.peers) - // Append NoOp to commit entries from previous terms - let raft = MetadataLog.append(raft, MetadataLog.NoOp) - - { - ...state, - raft: raft, - votesReceived: newVotes, - leaderId: Some(state.config.nodeId), - } - } else { - { - ...state, - votesReceived: newVotes, - } - } - - (newState, wonElection) - } else { - (state, false) - } -} - -/// Handle a vote request from a candidate. -let handleVoteRequest = ( - state: clusterState, - request: MetadataLog.voteRequest, -): (clusterState, MetadataLog.voteResponse) => { - let (newRaft, response) = MetadataLog.requestVote(state.raft, request) - - let newState = if response.voteGranted { - resetElectionTimer({...state, raft: newRaft}) - } else { - {...state, raft: newRaft} - } - - (newState, response) -} - -// ============================================================================ -// Log Replication -// ============================================================================ - -/// Leader creates AppendEntries requests for all followers. -let createHeartbeats = ( - state: clusterState, -): array<(MetadataLog.nodeId, MetadataLog.appendEntriesRequest)> => { - if state.raft.role != Leader { - [] - } else { - state.config.peers->Belt.Array.keepMap(peer => { - switch MetadataLog.createAppendEntriesRequest(state.raft, peer) { - | Some(request) => Some((peer, request)) - | None => None - } - }) - } -} - -/// Handle AppendEntries from a leader. -let handleAppendEntries = ( - state: clusterState, - request: MetadataLog.appendEntriesRequest, -): (clusterState, MetadataLog.appendEntriesResponse) => { - let (newRaft, response) = MetadataLog.appendEntries(state.raft, request) - - let newState = if response.success { - resetElectionTimer({ - ...state, - raft: newRaft, - leaderId: Some(request.leaderId), - }) - } else if request.term >= state.raft.currentTerm { - resetElectionTimer({ - ...state, - raft: newRaft, - leaderId: Some(request.leaderId), - }) - } else { - {...state, raft: newRaft} - } - - (newState, response) -} - -/// Leader handles AppendEntries response from a follower. -let handleAppendEntriesResponse = ( - state: clusterState, - fromPeer: MetadataLog.nodeId, - response: MetadataLog.appendEntriesResponse, -): clusterState => { - if response.term > state.raft.currentTerm { - // Step down - resetElectionTimer({ - ...state, - raft: MetadataLog.toFollower(state.raft, response.term), - leaderId: None, - }) - } else if state.raft.role == Leader { - if response.success { - // Update nextIndex and matchIndex for the follower - let nextIndex = Js.Dict.fromArray(Js.Dict.entries(state.raft.nextIndex)) - let matchIndex = Js.Dict.fromArray(Js.Dict.entries(state.raft.matchIndex)) - - Js.Dict.set(nextIndex, fromPeer, response.matchIndex + 1) - Js.Dict.set(matchIndex, fromPeer, response.matchIndex) - - let raft = {...state.raft, nextIndex: nextIndex, matchIndex: matchIndex} - // Try to advance commit index - let raft = MetadataLog.updateCommit(raft, state.config.peers) - - {...state, raft: raft} - } else { - // Decrement nextIndex for the follower and retry - let nextIndex = Js.Dict.fromArray(Js.Dict.entries(state.raft.nextIndex)) - let currentNext = - Js.Dict.get(nextIndex, fromPeer)->Belt.Option.getWithDefault(1) - Js.Dict.set(nextIndex, fromPeer, Js.Math.max_int(currentNext - 1, 1)) - - {...state, raft: {...state.raft, nextIndex: nextIndex}} - } - } else { - state - } -} - -// ============================================================================ -// Client Request Handling -// ============================================================================ - -type clientResult = - | Accepted({index: MetadataLog.index}) - | NotLeader({leaderId: option<MetadataLog.nodeId>}) - | Error({message: string}) - -/// Propose a command (only succeeds on the leader). -let propose = (state: clusterState, command: MetadataLog.command): (clusterState, clientResult) => { - switch state.raft.role { - | Leader => { - let raft = MetadataLog.append(state.raft, command) - let index = MetadataLog.getLastLogIndex(raft) - - let pending: pendingRequest = { - command: command, - index: index, - timestamp: Js.Date.now(), - } - - let newState = { - ...state, - raft: raft, - pendingRequests: Belt.Array.concat(state.pendingRequests, [pending]), - } - - (newState, Accepted({index: index})) - } - - | _ => (state, NotLeader({leaderId: state.leaderId})) - } -} - -/// Convenience: propose a store registration. -let proposeRegisterStore = ( - state: clusterState, - storeId: string, - endpoint: string, - modalities: array<string>, -): (clusterState, clientResult) => { - propose( - state, - MetadataLog.RegisterStore({storeId, endpoint, modalities}), - ) -} - -/// Convenience: propose updating a store's trust level. -let proposeUpdateTrust = ( - state: clusterState, - storeId: string, - newTrust: float, -): (clusterState, clientResult) => { - propose( - state, - MetadataLog.UpdateTrust({storeId, newTrust}), - ) -} - -/// Convenience: propose a hexad mapping. -let proposeMapHexad = ( - state: clusterState, - hexadId: string, - locations: Js.Dict.t<Js.Json.t>, -): (clusterState, clientResult) => { - propose( - state, - MetadataLog.MapHexad({hexadId, locations}), - ) -} - -// ============================================================================ -// State Machine Application -// ============================================================================ - -/// Apply a single committed command to the Registry state machine. -let applyCommand = (registry: Registry.registryState, command: MetadataLog.command): Registry.registryState => { - switch command { - | RegisterStore({storeId, endpoint, modalities}) => { - let modalityTypes = - modalities->Belt.Array.keepMap(m => Registry.modalityFromString(m)) - Registry.register(registry, storeId, endpoint, modalityTypes) - } - - | UnregisterStore({storeId}) => { - // Remove store from registry - let newStores = Js.Dict.fromArray( - Js.Dict.entries(registry.stores)->Belt.Array.keep(((id, _)) => id != storeId), - ) - {...registry, stores: newStores} - } - - | MapHexad({hexadId, locations}) => { - // Convert JSON locations to storeLocation dict - // In production, would deserialize properly - Registry.map(registry, hexadId, Js.Dict.empty()) - } - - | UnmapHexad({hexadId}) => { - let newMappings = Js.Dict.fromArray( - Js.Dict.entries(registry.mappings)->Belt.Array.keep(((id, _)) => id != hexadId), - ) - {...registry, mappings: newMappings} - } - - | UpdateTrust({storeId, newTrust: _}) => { - // Trust updates go through health update mechanism - // The newTrust is applied during health checks - registry - } - - | NoOp => registry - } -} - -/// Apply all committed but unapplied entries to the Registry. -let applyCommitted = (state: clusterState): clusterState => { - let (newRaft, commands) = MetadataLog.applyCommitted(state.raft) - - let newRegistry = commands->Belt.Array.reduce(state.registry, (reg, cmd) => { - applyCommand(reg, cmd) - }) - - // Remove fulfilled pending requests - let newPending = state.pendingRequests->Belt.Array.keep(req => { - req.index > newRaft.lastApplied - }) - - { - ...state, - raft: newRaft, - registry: newRegistry, - pendingRequests: newPending, - totalApplied: state.totalApplied + Belt.Array.length(commands), - totalCommitted: Js.Math.max_int(state.totalCommitted, newRaft.commitIndex), - } -} - -// ============================================================================ -// Tick — Main Loop Driver -// ============================================================================ - -type tickAction = - | SendVoteRequests(array<(MetadataLog.nodeId, MetadataLog.voteRequest)>) - | SendAppendEntries(array<(MetadataLog.nodeId, MetadataLog.appendEntriesRequest)>) - | BecameLeader - | AppliedEntries({count: int}) - | NoAction - -/// Advance the cluster by deltaMs. Returns the new state and any actions to perform. -let tick = (state: clusterState, deltaMs: int): (clusterState, array<tickAction>) => { - let actions = [] - - // 1. Advance election timer (followers and candidates only) - let (state, actions) = switch state.raft.role { - | Follower | Candidate => { - let (state, timedOut) = tickElectionTimer(state, deltaMs) - - if timedOut { - let (state, voteRequests) = startElection(state) - (state, Belt.Array.concat(actions, [SendVoteRequests(voteRequests)])) - } else { - (state, actions) - } - } - - | Leader => { - // Leaders don't use election timers; they send heartbeats - let heartbeats = createHeartbeats(state) - - if Belt.Array.length(heartbeats) > 0 { - (state, Belt.Array.concat(actions, [SendAppendEntries(heartbeats)])) - } else { - (state, actions) - } - } - } - - // 2. Apply committed entries - let prevApplied = state.raft.lastApplied - let state = applyCommitted(state) - let appliedCount = state.raft.lastApplied - prevApplied - - let actions = if appliedCount > 0 { - Belt.Array.concat(actions, [AppliedEntries({count: appliedCount})]) - } else { - actions - } - - (state, actions) -} - -// ============================================================================ -// Cluster Diagnostics -// ============================================================================ - -type clusterDiagnostics = { - nodeId: MetadataLog.nodeId, - role: string, - currentTerm: MetadataLog.term, - commitIndex: MetadataLog.index, - lastApplied: MetadataLog.index, - logLength: int, - peerCount: int, - leaderId: option<MetadataLog.nodeId>, - registeredStores: int, - mappedHexads: int, - totalCommitted: int, - totalApplied: int, - electionCount: int, - pendingRequests: int, -} - -let diagnostics = (state: clusterState): clusterDiagnostics => { - let roleStr = switch state.raft.role { - | Leader => "leader" - | Follower => "follower" - | Candidate => "candidate" - } - - { - nodeId: state.config.nodeId, - role: roleStr, - currentTerm: state.raft.currentTerm, - commitIndex: state.raft.commitIndex, - lastApplied: state.raft.lastApplied, - logLength: Belt.Array.length(state.raft.log), - peerCount: Belt.Array.length(state.config.peers), - leaderId: state.leaderId, - registeredStores: Belt.Array.length(Js.Dict.keys(state.registry.stores)), - mappedHexads: Belt.Array.length(Js.Dict.keys(state.registry.mappings)), - totalCommitted: state.totalCommitted, - totalApplied: state.totalApplied, - electionCount: state.electionCount, - pendingRequests: Belt.Array.length(state.pendingRequests), - } -} - -// ============================================================================ -// Public API -// ============================================================================ - -let create = createCluster -let election = startElection -let vote = handleVoteRequest -let voteResult = handleVoteResponse -let replicate = handleAppendEntries -let replicateResult = handleAppendEntriesResponse -let heartbeats = createHeartbeats -let submit = propose -let submitRegister = proposeRegisterStore -let submitTrust = proposeUpdateTrust -let submitMap = proposeMapHexad -let apply = applyCommitted -let advance = tick -let status = diagnostics diff --git a/verisimdb/src/registry/KRaftSerializer.res b/verisimdb/src/registry/KRaftSerializer.res deleted file mode 100644 index 52f279c3..00000000 --- a/verisimdb/src/registry/KRaftSerializer.res +++ /dev/null @@ -1,524 +0,0 @@ -// SPDX-License-Identifier: MPL-2.0 -// KRaft Serializer — JSON persistence for Raft log and cluster state. -// -// Provides encode/decode functions for all Raft types so the log -// can be persisted to disk (or transmitted over the wire for RPC). - -// ============================================================================ -// Command Serialization -// ============================================================================ - -let commandToJson = (cmd: MetadataLog.command): Js.Json.t => { - open Js.Json - - switch cmd { - | RegisterStore({storeId, endpoint, modalities}) => - Js.Dict.fromArray([ - ("type", string("RegisterStore")), - ("storeId", string(storeId)), - ("endpoint", string(endpoint)), - ( - "modalities", - array(modalities->Belt.Array.map(m => string(m))), - ), - ])->object_ - - | UnregisterStore({storeId}) => - Js.Dict.fromArray([ - ("type", string("UnregisterStore")), - ("storeId", string(storeId)), - ])->object_ - - | MapHexad({hexadId, locations}) => - Js.Dict.fromArray([ - ("type", string("MapHexad")), - ("hexadId", string(hexadId)), - ("locations", locations->object_), - ])->object_ - - | UnmapHexad({hexadId}) => - Js.Dict.fromArray([ - ("type", string("UnmapHexad")), - ("hexadId", string(hexadId)), - ])->object_ - - | UpdateTrust({storeId, newTrust}) => - Js.Dict.fromArray([ - ("type", string("UpdateTrust")), - ("storeId", string(storeId)), - ("newTrust", number(newTrust)), - ])->object_ - - | NoOp => - Js.Dict.fromArray([("type", string("NoOp"))])->object_ - } -} - -let commandFromJson = (json: Js.Json.t): option<MetadataLog.command> => { - open Belt.Option - - let dict = Js.Json.decodeObject(json) - - dict->flatMap(d => { - let cmdType = - Js.Dict.get(d, "type") - ->flatMap(Js.Json.decodeString) - - switch cmdType { - | Some("RegisterStore") => { - let storeId = Js.Dict.get(d, "storeId")->flatMap(Js.Json.decodeString) - let endpoint = Js.Dict.get(d, "endpoint")->flatMap(Js.Json.decodeString) - let modalities = - Js.Dict.get(d, "modalities") - ->flatMap(Js.Json.decodeArray) - ->map(arr => arr->Belt.Array.keepMap(Js.Json.decodeString)) - - switch (storeId, endpoint, modalities) { - | (Some(s), Some(e), Some(m)) => - Some(MetadataLog.RegisterStore({storeId: s, endpoint: e, modalities: m})) - | _ => None - } - } - - | Some("UnregisterStore") => { - let storeId = Js.Dict.get(d, "storeId")->flatMap(Js.Json.decodeString) - storeId->map(s => MetadataLog.UnregisterStore({storeId: s})) - } - - | Some("MapHexad") => { - let hexadId = Js.Dict.get(d, "hexadId")->flatMap(Js.Json.decodeString) - let locations = - Js.Dict.get(d, "locations") - ->flatMap(Js.Json.decodeObject) - - switch (hexadId, locations) { - | (Some(h), Some(l)) => - Some(MetadataLog.MapHexad({hexadId: h, locations: l})) - | _ => None - } - } - - | Some("UnmapHexad") => { - let hexadId = Js.Dict.get(d, "hexadId")->flatMap(Js.Json.decodeString) - hexadId->map(h => MetadataLog.UnmapHexad({hexadId: h})) - } - - | Some("UpdateTrust") => { - let storeId = Js.Dict.get(d, "storeId")->flatMap(Js.Json.decodeString) - let newTrust = Js.Dict.get(d, "newTrust")->flatMap(Js.Json.decodeNumber) - - switch (storeId, newTrust) { - | (Some(s), Some(t)) => - Some(MetadataLog.UpdateTrust({storeId: s, newTrust: t})) - | _ => None - } - } - - | Some("NoOp") => Some(MetadataLog.NoOp) - | _ => None - } - }) -} - -// ============================================================================ -// Log Entry Serialization -// ============================================================================ - -let logEntryToJson = (entry: MetadataLog.logEntry): Js.Json.t => { - Js.Dict.fromArray([ - ("term", Js.Json.number(Belt.Int.toFloat(entry.term))), - ("index", Js.Json.number(Belt.Int.toFloat(entry.index))), - ("command", commandToJson(entry.command)), - ("timestamp", Js.Json.number(entry.timestamp)), - ])->Js.Json.object_ -} - -let logEntryFromJson = (json: Js.Json.t): option<MetadataLog.logEntry> => { - open Belt.Option - - Js.Json.decodeObject(json)->flatMap(d => { - let term = - Js.Dict.get(d, "term") - ->flatMap(Js.Json.decodeNumber) - ->map(Belt.Float.toInt) - - let index = - Js.Dict.get(d, "index") - ->flatMap(Js.Json.decodeNumber) - ->map(Belt.Float.toInt) - - let command = - Js.Dict.get(d, "command") - ->flatMap(commandFromJson) - - let timestamp = - Js.Dict.get(d, "timestamp") - ->flatMap(Js.Json.decodeNumber) - - switch (term, index, command, timestamp) { - | (Some(t), Some(i), Some(c), Some(ts)) => - Some({ - term: t, - index: i, - command: c, - timestamp: ts, - }: MetadataLog.logEntry) - | _ => None - } - }) -} - -// ============================================================================ -// Node State Serialization -// ============================================================================ - -let roleToString = (role: MetadataLog.nodeRole): string => { - switch role { - | Leader => "leader" - | Follower => "follower" - | Candidate => "candidate" - } -} - -let roleFromString = (s: string): option<MetadataLog.nodeRole> => { - switch s { - | "leader" => Some(MetadataLog.Leader) - | "follower" => Some(MetadataLog.Follower) - | "candidate" => Some(MetadataLog.Candidate) - | _ => None - } -} - -let dictToJsonNumbers = (d: Js.Dict.t<MetadataLog.index>): Js.Json.t => { - let entries = Js.Dict.entries(d)->Belt.Array.map(((k, v)) => { - (k, Js.Json.number(Belt.Int.toFloat(v))) - }) - Js.Dict.fromArray(entries)->Js.Json.object_ -} - -let jsonToDictNumbers = (json: Js.Json.t): Js.Dict.t<MetadataLog.index> => { - switch Js.Json.decodeObject(json) { - | None => Js.Dict.empty() - | Some(d) => { - let entries = Js.Dict.entries(d)->Belt.Array.keepMap(((k, v)) => { - Js.Json.decodeNumber(v)->Belt.Option.map(n => (k, Belt.Float.toInt(n))) - }) - Js.Dict.fromArray(entries) - } - } -} - -let nodeStateToJson = (state: MetadataLog.nodeState): Js.Json.t => { - Js.Dict.fromArray([ - ("role", Js.Json.string(roleToString(state.role))), - ("currentTerm", Js.Json.number(Belt.Int.toFloat(state.currentTerm))), - ( - "votedFor", - switch state.votedFor { - | Some(id) => Js.Json.string(id) - | None => Js.Json.null - }, - ), - ("log", Js.Json.array(state.log->Belt.Array.map(logEntryToJson))), - ("commitIndex", Js.Json.number(Belt.Int.toFloat(state.commitIndex))), - ("lastApplied", Js.Json.number(Belt.Int.toFloat(state.lastApplied))), - ("nextIndex", dictToJsonNumbers(state.nextIndex)), - ("matchIndex", dictToJsonNumbers(state.matchIndex)), - ])->Js.Json.object_ -} - -let nodeStateFromJson = (json: Js.Json.t): option<MetadataLog.nodeState> => { - open Belt.Option - - Js.Json.decodeObject(json)->flatMap(d => { - let role = - Js.Dict.get(d, "role") - ->flatMap(Js.Json.decodeString) - ->flatMap(roleFromString) - - let currentTerm = - Js.Dict.get(d, "currentTerm") - ->flatMap(Js.Json.decodeNumber) - ->map(Belt.Float.toInt) - - let votedFor = - Js.Dict.get(d, "votedFor") - ->flatMap(v => - if v == Js.Json.null { - Some(None) - } else { - Js.Json.decodeString(v)->map(s => Some(s)) - } - ) - - let log = - Js.Dict.get(d, "log") - ->flatMap(Js.Json.decodeArray) - ->map(arr => arr->Belt.Array.keepMap(logEntryFromJson)) - - let commitIndex = - Js.Dict.get(d, "commitIndex") - ->flatMap(Js.Json.decodeNumber) - ->map(Belt.Float.toInt) - - let lastApplied = - Js.Dict.get(d, "lastApplied") - ->flatMap(Js.Json.decodeNumber) - ->map(Belt.Float.toInt) - - let nextIndex = - Js.Dict.get(d, "nextIndex") - ->map(jsonToDictNumbers) - ->getWithDefault(Js.Dict.empty()) - - let matchIndex = - Js.Dict.get(d, "matchIndex") - ->map(jsonToDictNumbers) - ->getWithDefault(Js.Dict.empty()) - - switch (role, currentTerm, votedFor, log, commitIndex, lastApplied) { - | (Some(r), Some(ct), Some(vf), Some(l), Some(ci), Some(la)) => - Some({ - role: r, - currentTerm: ct, - votedFor: vf, - log: l, - commitIndex: ci, - lastApplied: la, - nextIndex: nextIndex, - matchIndex: matchIndex, - }: MetadataLog.nodeState) - | _ => None - } - }) -} - -// ============================================================================ -// Vote Request/Response Serialization -// ============================================================================ - -let voteRequestToJson = (req: MetadataLog.voteRequest): Js.Json.t => { - Js.Dict.fromArray([ - ("term", Js.Json.number(Belt.Int.toFloat(req.term))), - ("candidateId", Js.Json.string(req.candidateId)), - ("lastLogIndex", Js.Json.number(Belt.Int.toFloat(req.lastLogIndex))), - ("lastLogTerm", Js.Json.number(Belt.Int.toFloat(req.lastLogTerm))), - ])->Js.Json.object_ -} - -let voteRequestFromJson = (json: Js.Json.t): option<MetadataLog.voteRequest> => { - open Belt.Option - - Js.Json.decodeObject(json)->flatMap(d => { - let term = Js.Dict.get(d, "term")->flatMap(Js.Json.decodeNumber)->map(Belt.Float.toInt) - let candidateId = Js.Dict.get(d, "candidateId")->flatMap(Js.Json.decodeString) - let lastLogIndex = Js.Dict.get(d, "lastLogIndex")->flatMap(Js.Json.decodeNumber)->map(Belt.Float.toInt) - let lastLogTerm = Js.Dict.get(d, "lastLogTerm")->flatMap(Js.Json.decodeNumber)->map(Belt.Float.toInt) - - switch (term, candidateId, lastLogIndex, lastLogTerm) { - | (Some(t), Some(c), Some(li), Some(lt)) => - Some({term: t, candidateId: c, lastLogIndex: li, lastLogTerm: lt}: MetadataLog.voteRequest) - | _ => None - } - }) -} - -let voteResponseToJson = (res: MetadataLog.voteResponse): Js.Json.t => { - Js.Dict.fromArray([ - ("term", Js.Json.number(Belt.Int.toFloat(res.term))), - ("voteGranted", Js.Json.boolean(res.voteGranted)), - ])->Js.Json.object_ -} - -let voteResponseFromJson = (json: Js.Json.t): option<MetadataLog.voteResponse> => { - open Belt.Option - - Js.Json.decodeObject(json)->flatMap(d => { - let term = Js.Dict.get(d, "term")->flatMap(Js.Json.decodeNumber)->map(Belt.Float.toInt) - let voteGranted = Js.Dict.get(d, "voteGranted")->flatMap(Js.Json.decodeBoolean) - - switch (term, voteGranted) { - | (Some(t), Some(v)) => - Some({term: t, voteGranted: v}: MetadataLog.voteResponse) - | _ => None - } - }) -} - -// ============================================================================ -// AppendEntries Request/Response Serialization -// ============================================================================ - -let appendEntriesRequestToJson = (req: MetadataLog.appendEntriesRequest): Js.Json.t => { - Js.Dict.fromArray([ - ("term", Js.Json.number(Belt.Int.toFloat(req.term))), - ("leaderId", Js.Json.string(req.leaderId)), - ("prevLogIndex", Js.Json.number(Belt.Int.toFloat(req.prevLogIndex))), - ("prevLogTerm", Js.Json.number(Belt.Int.toFloat(req.prevLogTerm))), - ("entries", Js.Json.array(req.entries->Belt.Array.map(logEntryToJson))), - ("leaderCommit", Js.Json.number(Belt.Int.toFloat(req.leaderCommit))), - ])->Js.Json.object_ -} - -let appendEntriesRequestFromJson = (json: Js.Json.t): option<MetadataLog.appendEntriesRequest> => { - open Belt.Option - - Js.Json.decodeObject(json)->flatMap(d => { - let term = Js.Dict.get(d, "term")->flatMap(Js.Json.decodeNumber)->map(Belt.Float.toInt) - let leaderId = Js.Dict.get(d, "leaderId")->flatMap(Js.Json.decodeString) - let prevLogIndex = Js.Dict.get(d, "prevLogIndex")->flatMap(Js.Json.decodeNumber)->map(Belt.Float.toInt) - let prevLogTerm = Js.Dict.get(d, "prevLogTerm")->flatMap(Js.Json.decodeNumber)->map(Belt.Float.toInt) - let entries = - Js.Dict.get(d, "entries") - ->flatMap(Js.Json.decodeArray) - ->map(arr => arr->Belt.Array.keepMap(logEntryFromJson)) - let leaderCommit = Js.Dict.get(d, "leaderCommit")->flatMap(Js.Json.decodeNumber)->map(Belt.Float.toInt) - - switch (term, leaderId, prevLogIndex, prevLogTerm, entries, leaderCommit) { - | (Some(t), Some(l), Some(pi), Some(pt), Some(e), Some(lc)) => - Some({ - term: t, - leaderId: l, - prevLogIndex: pi, - prevLogTerm: pt, - entries: e, - leaderCommit: lc, - }: MetadataLog.appendEntriesRequest) - | _ => None - } - }) -} - -let appendEntriesResponseToJson = (res: MetadataLog.appendEntriesResponse): Js.Json.t => { - Js.Dict.fromArray([ - ("term", Js.Json.number(Belt.Int.toFloat(res.term))), - ("success", Js.Json.boolean(res.success)), - ("matchIndex", Js.Json.number(Belt.Int.toFloat(res.matchIndex))), - ])->Js.Json.object_ -} - -let appendEntriesResponseFromJson = (json: Js.Json.t): option<MetadataLog.appendEntriesResponse> => { - open Belt.Option - - Js.Json.decodeObject(json)->flatMap(d => { - let term = Js.Dict.get(d, "term")->flatMap(Js.Json.decodeNumber)->map(Belt.Float.toInt) - let success = Js.Dict.get(d, "success")->flatMap(Js.Json.decodeBoolean) - let matchIndex = Js.Dict.get(d, "matchIndex")->flatMap(Js.Json.decodeNumber)->map(Belt.Float.toInt) - - switch (term, success, matchIndex) { - | (Some(t), Some(s), Some(mi)) => - Some({term: t, success: s, matchIndex: mi}: MetadataLog.appendEntriesResponse) - | _ => None - } - }) -} - -// ============================================================================ -// Full Snapshot Serialization (for persistence) -// ============================================================================ - -type persistedSnapshot = { - version: int, - nodeState: Js.Json.t, - snapshotTimestamp: float, -} - -let snapshotToJson = (state: MetadataLog.nodeState): Js.Json.t => { - Js.Dict.fromArray([ - ("version", Js.Json.number(1.0)), - ("nodeState", nodeStateToJson(state)), - ("snapshotTimestamp", Js.Json.number(Js.Date.now())), - ])->Js.Json.object_ -} - -let snapshotFromJson = (json: Js.Json.t): option<MetadataLog.nodeState> => { - open Belt.Option - - Js.Json.decodeObject(json)->flatMap(d => { - let version = - Js.Dict.get(d, "version") - ->flatMap(Js.Json.decodeNumber) - ->map(Belt.Float.toInt) - - switch version { - | Some(1) => - Js.Dict.get(d, "nodeState")->flatMap(nodeStateFromJson) - | _ => None // Unknown version - } - }) -} - -// ============================================================================ -// Write-Ahead Log (WAL) Entry Format -// ============================================================================ - -/// Encode a single log entry as a line for append-only WAL file. -let walEncode = (entry: MetadataLog.logEntry): string => { - Js.Json.stringify(logEntryToJson(entry)) -} - -/// Decode a WAL line back to a log entry. -let walDecode = (line: string): option<MetadataLog.logEntry> => { - try { - let json = Js.Json.parseExn(line) - logEntryFromJson(json) - } catch { - | _ => None - } -} - -/// Encode multiple WAL entries (newline-delimited JSON). -let walEncodeAll = (entries: array<MetadataLog.logEntry>): string => { - entries - ->Belt.Array.map(walEncode) - ->Belt.Array.joinWith("\n", s => s) -} - -/// Decode all entries from a WAL string. -let walDecodeAll = (data: string): array<MetadataLog.logEntry> => { - Js.String2.split(data, "\n") - ->Belt.Array.keepMap(line => { - let trimmed = Js.String2.trim(line) - if Js.String2.length(trimmed) > 0 { - walDecode(trimmed) - } else { - None - } - }) -} - -// ============================================================================ -// Public API -// ============================================================================ - -// Commands -let encodeCommand = commandToJson -let decodeCommand = commandFromJson - -// Log entries -let encodeEntry = logEntryToJson -let decodeEntry = logEntryFromJson - -// Node state -let encodeState = nodeStateToJson -let decodeState = nodeStateFromJson - -// RPC messages -let encodeVoteReq = voteRequestToJson -let decodeVoteReq = voteRequestFromJson -let encodeVoteRes = voteResponseToJson -let decodeVoteRes = voteResponseFromJson -let encodeAppendReq = appendEntriesRequestToJson -let decodeAppendReq = appendEntriesRequestFromJson -let encodeAppendRes = appendEntriesResponseToJson -let decodeAppendRes = appendEntriesResponseFromJson - -// Snapshots -let snapshot = snapshotToJson -let restore = snapshotFromJson - -// WAL -let wal = walEncode -let unwal = walDecode -let walAll = walEncodeAll -let unwalAll = walDecodeAll diff --git a/verisimdb/src/registry/MetadataLog.res b/verisimdb/src/registry/MetadataLog.res deleted file mode 100644 index 7d6e76e7..00000000 --- a/verisimdb/src/registry/MetadataLog.res +++ /dev/null @@ -1,375 +0,0 @@ -// SPDX-License-Identifier: MPL-2.0 -// KRaft-Inspired Metadata Log -// Replicated state machine with Raft consensus - -// ============================================================================ -// Types -// ============================================================================ - -type term = int -type index = int -type nodeId = string - -type logEntry = { - term: term, - index: index, - command: command, - timestamp: float, -} - -and command = - | RegisterStore({storeId: string, endpoint: string, modalities: array<string>}) - | UnregisterStore({storeId: string}) - | MapHexad({hexadId: string, locations: Js.Dict.t<Js.Json.t>}) - | UnmapHexad({hexadId: string}) - | UpdateTrust({storeId: string, newTrust: float}) - | NoOp - -type nodeRole = - | Leader - | Follower - | Candidate - -type nodeState = { - role: nodeRole, - currentTerm: term, - votedFor: option<nodeId>, - log: array<logEntry>, - commitIndex: index, - lastApplied: index, - // Leader state - nextIndex: Js.Dict.t<index>, - matchIndex: Js.Dict.t<index>, -} - -type voteRequest = { - term: term, - candidateId: nodeId, - lastLogIndex: index, - lastLogTerm: term, -} - -type voteResponse = { - term: term, - voteGranted: bool, -} - -type appendEntriesRequest = { - term: term, - leaderId: nodeId, - prevLogIndex: index, - prevLogTerm: term, - entries: array<logEntry>, - leaderCommit: index, -} - -type appendEntriesResponse = { - term: term, - success: bool, - matchIndex: index, -} - -// ============================================================================ -// Node Operations -// ============================================================================ - -let createNode = (~nodeId: nodeId): nodeState => { - { - role: Follower, - currentTerm: 0, - votedFor: None, - log: [], - commitIndex: 0, - lastApplied: 0, - nextIndex: Js.Dict.empty(), - matchIndex: Js.Dict.empty(), - } -} - -// Append entry to log -let appendEntry = (state: nodeState, command: command): nodeState => { - let newIndex = Belt.Array.length(state.log) + 1 - let entry: logEntry = { - term: state.currentTerm, - index: newIndex, - command: command, - timestamp: Js.Date.now(), - } - - {...state, log: Belt.Array.concat(state.log, [entry])} -} - -// Get last log entry -let getLastLogEntry = (state: nodeState): option<logEntry> => { - Belt.Array.get(state.log, Belt.Array.length(state.log) - 1) -} - -// Get last log term -let getLastLogTerm = (state: nodeState): term => { - switch getLastLogEntry(state) { - | None => 0 - | Some(entry) => entry.term - } -} - -// Get last log index -let getLastLogIndex = (state: nodeState): index => { - Belt.Array.length(state.log) -} - -// ============================================================================ -// Leader Election -// ============================================================================ - -let becomeCandidate = (state: nodeState): nodeState => { - { - ...state, - role: Candidate, - currentTerm: state.currentTerm + 1, - votedFor: None, // Will vote for self - } -} - -let becomeLeader = (state: nodeState, peers: array<nodeId>): nodeState => { - // Initialize nextIndex and matchIndex for all peers - let nextIndex = Js.Dict.empty() - let matchIndex = Js.Dict.empty() - - peers->Belt.Array.forEach(peer => { - Js.Dict.set(nextIndex, peer, getLastLogIndex(state) + 1) - Js.Dict.set(matchIndex, peer, 0) - }) - - { - ...state, - role: Leader, - nextIndex: nextIndex, - matchIndex: matchIndex, - } -} - -let becomeFollower = (state: nodeState, newTerm: term): nodeState => { - { - ...state, - role: Follower, - currentTerm: newTerm, - votedFor: None, - } -} - -// Request vote from a follower -let handleVoteRequest = ( - state: nodeState, - request: voteRequest, -): (nodeState, voteResponse) => { - let grantVote = if request.term < state.currentTerm { - false - } else if request.term > state.currentTerm { - // Higher term, become follower and grant vote - true - } else { - // Same term - switch state.votedFor { - | Some(_) => false // Already voted - | None => { - // Check if candidate's log is at least as up-to-date - let lastLogTerm = getLastLogTerm(state) - let lastLogIndex = getLastLogIndex(state) - - if request.lastLogTerm > lastLogTerm { - true - } else if request.lastLogTerm == lastLogTerm && request.lastLogIndex >= lastLogIndex { - true - } else { - false - } - } - } - } - - let newState = if grantVote && request.term >= state.currentTerm { - {...state, currentTerm: request.term, votedFor: Some(request.candidateId)} - } else if request.term > state.currentTerm { - becomeFollower(state, request.term) - } else { - state - } - - let response: voteResponse = { - term: newState.currentTerm, - voteGranted: grantVote, - } - - (newState, response) -} - -// ============================================================================ -// Log Replication -// ============================================================================ - -let handleAppendEntries = ( - state: nodeState, - request: appendEntriesRequest, -): (nodeState, appendEntriesResponse) => { - // Check term - if request.term < state.currentTerm { - let response: appendEntriesResponse = { - term: state.currentTerm, - success: false, - matchIndex: 0, - } - (state, response) - } else { - // Become follower if we were candidate - let newState = if request.term > state.currentTerm { - becomeFollower(state, request.term) - } else { - state - } - - // Check if log contains entry at prevLogIndex with prevLogTerm - let prevEntry = if request.prevLogIndex == 0 { - Some({term: 0, index: 0, command: NoOp, timestamp: 0.0}) - } else { - Belt.Array.get(newState.log, request.prevLogIndex - 1) - } - - switch prevEntry { - | None => { - // Log doesn't have entry at prevLogIndex - let response: appendEntriesResponse = { - term: newState.currentTerm, - success: false, - matchIndex: 0, - } - (newState, response) - } - | Some(entry) => - if entry.term != request.prevLogTerm { - // Log entry doesn't match - let response: appendEntriesResponse = { - term: newState.currentTerm, - success: false, - matchIndex: entry.index, - } - (newState, response) - } else { - // Append new entries - let logBeforePrev = Belt.Array.slice(newState.log, ~offset=0, ~len=request.prevLogIndex) - let newLog = Belt.Array.concat(logBeforePrev, request.entries) - - let finalState = { - ...newState, - log: newLog, - commitIndex: Js.Math.min_int(request.leaderCommit, getLastLogIndex({...newState, log: newLog})), - } - - let response: appendEntriesResponse = { - term: finalState.currentTerm, - success: true, - matchIndex: getLastLogIndex(finalState), - } - - (finalState, response) - } - } - } -} - -// Leader sends AppendEntries to follower -let createAppendEntriesRequest = ( - state: nodeState, - followerId: nodeId, -): option<appendEntriesRequest> => { - switch state.role { - | Leader => { - let nextIdx = Js.Dict.get(state.nextIndex, followerId)->Belt.Option.getWithDefault(1) - - let prevLogIndex = nextIdx - 1 - let prevLogTerm = if prevLogIndex == 0 { - 0 - } else { - Belt.Array.get(state.log, prevLogIndex - 1) - ->Belt.Option.map(e => e.term) - ->Belt.Option.getWithDefault(0) - } - - let entries = Belt.Array.sliceToEnd(state.log, nextIdx - 1) - - Some({ - term: state.currentTerm, - leaderId: "self", // Would be actual node ID - prevLogIndex: prevLogIndex, - prevLogTerm: prevLogTerm, - entries: entries, - leaderCommit: state.commitIndex, - }) - } - | _ => None - } -} - -// ============================================================================ -// Commit & Apply -// ============================================================================ - -let updateCommitIndex = (state: nodeState, peers: array<nodeId>): nodeState => { - switch state.role { - | Leader => { - // Find highest N where majority of matchIndex[i] >= N - let matchIndices = peers - ->Belt.Array.map(peer => { - Js.Dict.get(state.matchIndex, peer)->Belt.Option.getWithDefault(0) - }) - ->Belt.Array.concat([getLastLogIndex(state)]) - ->Belt.SortArray.stableSortBy((a, b) => b - a) - - let quorumIndex = (Belt.Array.length(peers) + 1) / 2 - let newCommitIndex = Belt.Array.get(matchIndices, quorumIndex)->Belt.Option.getWithDefault(state.commitIndex) - - // Only commit entries from current term - let canCommit = switch Belt.Array.get(state.log, newCommitIndex - 1) { - | None => false - | Some(entry) => entry.term == state.currentTerm - } - - if canCommit && newCommitIndex > state.commitIndex { - {...state, commitIndex: newCommitIndex} - } else { - state - } - } - | _ => state - } -} - -// Apply committed entries to state machine -let applyCommittedEntries = (state: nodeState): (nodeState, array<command>) => { - if state.lastApplied >= state.commitIndex { - (state, []) - } else { - let toApply = Belt.Array.slice( - state.log, - ~offset=state.lastApplied, - ~len=state.commitIndex - state.lastApplied, - ) - - let commands = toApply->Belt.Array.map(entry => entry.command) - - ({...state, lastApplied: state.commitIndex}, commands) - } -} - -// ============================================================================ -// Public API -// ============================================================================ - -let create = createNode -let append = appendEntry -let requestVote = handleVoteRequest -let appendEntries = handleAppendEntries -let toCandidate = becomeCandidate -let toLeader = becomeLeader -let toFollower = becomeFollower -let updateCommit = updateCommitIndex -let applyCommitted = applyCommittedEntries diff --git a/verisimdb/src/registry/Registry.res b/verisimdb/src/registry/Registry.res deleted file mode 100644 index 24bbec57..00000000 --- a/verisimdb/src/registry/Registry.res +++ /dev/null @@ -1,868 +0,0 @@ -// SPDX-License-Identifier: MPL-2.0 -// ReScript Federation Registry -// The "tiny core" (<5k LOC) for universal federated knowledge - -// ============================================================================ -// Types -// ============================================================================ - -type hexadId = string - -type storeId = string - -type modalityType = - | Graph - | Vector - | Tensor - | Semantic - | Document - | Temporal - | Provenance - | Spatial - -type storeLocation = { - storeId: storeId, - endpoint: string, - modalities: array<modalityType>, - trustLevel: float, // 0.0-1.0 - lastSeen: Js.Date.t, - responseTimeMs: option<int>, -} - -type hexadMapping = { - hexadId: hexadId, - locations: Js.Dict.t<storeLocation>, // modalityType -> storeLocation - primaryStore: option<storeId>, - created: Js.Date.t, - modified: Js.Date.t, -} - -type registryState = { - mappings: Js.Dict.t<hexadMapping>, // hexadId -> hexadMapping - stores: Js.Dict.t<storeLocation>, // storeId -> storeLocation - config: registryConfig, -} - -and registryConfig = { - minTrustLevel: float, - maxStoreDowntimeMs: int, - replicationFactor: int, - consistencyMode: consistencyMode, -} - -and consistencyMode = - | Strong // All replicas must agree - | Eventual // Accept temporary inconsistency - | Quorum // Majority must agree - -// ============================================================================ -// Registry Operations -// ============================================================================ - -let createRegistry = (~config: registryConfig): registryState => { - { - mappings: Js.Dict.empty(), - stores: Js.Dict.empty(), - config: config, - } -} - -let defaultConfig = (): registryConfig => { - { - minTrustLevel: 0.5, - maxStoreDowntimeMs: 300_000, // 5 minutes - replicationFactor: 3, - consistencyMode: Quorum, - } -} - -// Register a new store -let registerStore = ( - registry: registryState, - storeId: storeId, - endpoint: string, - modalities: array<modalityType>, -): registryState => { - let location: storeLocation = { - storeId: storeId, - endpoint: endpoint, - modalities: modalities, - trustLevel: 1.0, - lastSeen: Js.Date.make(), - responseTimeMs: None, - } - - let newStores = Js.Dict.fromArray(Js.Dict.entries(registry.stores)) - Js.Dict.set(newStores, storeId, location) - - {...registry, stores: newStores} -} - -// Map a hexad to store locations -let mapHexad = ( - registry: registryState, - hexadId: hexadId, - locations: Js.Dict.t<storeLocation>, -): registryState => { - let mapping: hexadMapping = { - hexadId: hexadId, - locations: locations, - primaryStore: None, - created: Js.Date.make(), - modified: Js.Date.make(), - } - - let newMappings = Js.Dict.fromArray(Js.Dict.entries(registry.mappings)) - Js.Dict.set(newMappings, hexadId, mapping) - - {...registry, mappings: newMappings} -} - -// Get store locations for a hexad -let getHexadLocations = ( - registry: registryState, - hexadId: hexadId, -): option<hexadMapping> => { - Js.Dict.get(registry.mappings, hexadId) -} - -// Find stores that have a specific modality -let findStoresByModality = ( - registry: registryState, - modality: modalityType, -): array<storeLocation> => { - registry.stores - ->Js.Dict.values - ->Belt.Array.keep(store => { - store.modalities->Belt.Array.some(m => m == modality) - }) -} - -// Select best store for a modality based on trust and response time -let selectBestStore = ( - registry: registryState, - modality: modalityType, -): option<storeLocation> => { - let candidates = findStoresByModality(registry, modality) - - if Belt.Array.length(candidates) == 0 { - None - } else { - // Score stores by trust level and response time - let scored = candidates->Belt.Array.map(store => { - let trustScore = store.trustLevel - let responseScore = switch store.responseTimeMs { - | None => 0.5 - | Some(ms) => 1.0 -. (Belt.Int.toFloat(ms) /. 1000.0)->Js.Math.min_float(1.0) - } - let score = trustScore *. 0.7 +. responseScore *. 0.3 - (store, score) - }) - - // Sort by score descending - let sorted = scored->Belt.Array.reverse->Belt.SortArray.stableSortBy(((_, scoreA), (_, scoreB)) => { - Belt.Float.toInt((scoreB -. scoreA) *. 1000.0) - }) - - sorted->Belt.Array.get(0)->Belt.Option.map(((store, _)) => store) - } -} - -// Update store health metrics -let updateStoreHealth = ( - registry: registryState, - storeId: storeId, - responseTimeMs: int, - success: bool, -): registryState => { - switch Js.Dict.get(registry.stores, storeId) { - | None => registry - | Some(store) => { - let newTrust = if success { - Js.Math.min_float(store.trustLevel +. 0.05, 1.0) - } else { - Js.Math.max_float(store.trustLevel -. 0.1, 0.0) - } - - let updatedStore = { - ...store, - trustLevel: newTrust, - lastSeen: Js.Date.make(), - responseTimeMs: Some(responseTimeMs), - } - - let newStores = Js.Dict.fromArray(Js.Dict.entries(registry.stores)) - Js.Dict.set(newStores, storeId, updatedStore) - - {...registry, stores: newStores} - } - } -} - -// Remove stores that haven't been seen recently -let pruneDeadStores = (registry: registryState): registryState => { - let now = Js.Date.now() - let maxDowntime = Belt.Int.toFloat(registry.config.maxStoreDowntimeMs) - - let liveStores = registry.stores - ->Js.Dict.entries - ->Belt.Array.keep(((_, store)) => { - let timeSinceLastSeen = now -. Js.Date.getTime(store.lastSeen) - timeSinceLastSeen < maxDowntime - }) - ->Js.Dict.fromArray - - {...registry, stores: liveStores} -} - -// ============================================================================ -// Federation Queries -// ============================================================================ - -type federationQuery = { - pattern: string, // e.g., "/universities/*" - modalities: array<modalityType>, - limit: int, -} - -type queryResult = { - storeId: storeId, - hexadId: hexadId, - modality: modalityType, - data: Js.Json.t, -} - -// Resolve a federation pattern to list of stores -let resolvePattern = ( - registry: registryState, - pattern: string, -): array<storeLocation> => { - // Simple pattern matching - in production would use regex - if Js.String2.endsWith(pattern, "/*") { - let prefix = Js.String2.slice(pattern, ~from=0, ~to_=Js.String2.length(pattern) - 2) - - registry.stores - ->Js.Dict.values - ->Belt.Array.keep(store => { - Js.String2.startsWith(store.storeId, prefix) - }) - } else { - // Exact match - switch Js.Dict.get(registry.stores, pattern) { - | None => [] - | Some(store) => [store] - } - } -} - -// Execute a federated query across multiple stores via HTTP fan-out -let executeFederatedQuery = async ( - registry: registryState, - query: federationQuery, -): Promise.t<array<queryResult>> => { - let stores = resolvePattern(registry, query.pattern) - - // Filter stores by required modalities and minimum trust level - let eligibleStores = stores->Belt.Array.keep(store => { - store.trustLevel >= registry.config.minTrustLevel && - query.modalities->Belt.Array.every(modality => { - store.modalities->Belt.Array.some(m => m == modality) - }) - }) - - // Fan out HTTP requests to each eligible store - let fetchPromises = eligibleStores->Belt.Array.map(store => { - let url = store.endpoint ++ "/hexads?limit=" ++ Belt.Int.toString(query.limit) - Fetch.fetch(url, {method: #GET}) - ->Promise.then(response => { - if Fetch.Response.ok(response) { - Fetch.Response.json(response) - ->Promise.then(json => { - // Map response items to queryResult - let items = switch Js.Json.classify(json) { - | Js.Json.JSONArray(arr) => - arr->Belt.Array.flatMap(item => { - query.modalities->Belt.Array.map(modality => { - { - storeId: store.storeId, - hexadId: switch Js.Json.classify(item) { - | Js.Json.JSONObject(obj) => - switch Js.Dict.get(obj, "id") { - | Some(id) => - switch Js.Json.classify(id) { - | Js.Json.JSONString(s) => s - | _ => "unknown" - } - | None => "unknown" - } - | _ => "unknown" - }, - modality: modality, - data: item, - } - }) - }) - | _ => [] - } - Promise.resolve(items) - }) - } else { - Promise.resolve([]) - } - }) - ->Promise.catch(_err => { - Promise.resolve([]) - }) - }) - - // Collect results from all stores - let allResults = await Promise.all(fetchPromises) - let combined = allResults->Belt.Array.flatMap(r => r) - - // Apply limit - let limited = combined->Belt.Array.slice(~offset=0, ~len=query.limit) - limited -} - -// ============================================================================ -// Consistency & Replication -// ============================================================================ - -type replicationStatus = - | UpToDate - | Stale({lagMs: int}) - | Diverged({conflictCount: int}) - -let checkReplicationStatus = ( - registry: registryState, - hexadId: hexadId, -): replicationStatus => { - // Check if hexad replicas are consistent by examining mapping and store health - switch Js.Dict.get(registry.mappings, hexadId) { - | None => UpToDate // No mapping = nothing to replicate - | Some(mapping) => { - let locationEntries = Js.Dict.values(mapping.locations) - let storeCount = Belt.Array.length(locationEntries) - - if storeCount <= 1 { - UpToDate - } else { - // Check how many stores are alive and responsive - let aliveCount = locationEntries->Belt.Array.keep(loc => { - switch Js.Dict.get(registry.stores, loc.storeId) { - | None => false - | Some(store) => { - let now = Js.Date.now() - let age = now -. Js.Date.getTime(store.lastSeen) - age < Belt.Int.toFloat(registry.config.maxStoreDowntimeMs) - } - } - })->Belt.Array.length - - if aliveCount < storeCount { - let lagMs = storeCount - aliveCount - Stale({lagMs: lagMs * 1000}) - } else { - // All stores alive — check for trust divergence as proxy for data divergence - let trusts = locationEntries->Belt.Array.map(loc => { - switch Js.Dict.get(registry.stores, loc.storeId) { - | None => 0.0 - | Some(store) => store.trustLevel - } - }) - let minTrust = trusts->Belt.Array.reduce(1.0, (a, b) => Js.Math.min_float(a, b)) - let maxTrust = trusts->Belt.Array.reduce(0.0, (a, b) => Js.Math.max_float(a, b)) - - if maxTrust -. minTrust > 0.3 { - Diverged({conflictCount: 1}) - } else { - UpToDate - } - } - } - } - } -} - -// Trigger replication for a hexad: fetch from source, push to targets -let replicateHexad = async ( - registry: registryState, - hexadId: hexadId, - sourceStore: storeId, - targetStores: array<storeId>, -): Promise.t<Result.t<unit, string>> => { - // Look up source store endpoint - let sourceEndpoint = switch Js.Dict.get(registry.stores, sourceStore) { - | None => None - | Some(store) => Some(store.endpoint) - } - - switch sourceEndpoint { - | None => Error("Source store '" ++ sourceStore ++ "' not found in registry") - | Some(endpoint) => { - // Fetch hexad from source - let fetchUrl = endpoint ++ "/hexads/" ++ hexadId - let fetchResult = try { - let response = await Fetch.fetch(fetchUrl, {method: #GET}) - if Fetch.Response.ok(response) { - let json = await Fetch.Response.json(response) - Ok(json) - } else { - Error("Source store returned " ++ Belt.Int.toString(Fetch.Response.status(response))) - } - } catch { - | exn => { - // Type-safe exception message extraction (no Obj.magic) - let msg = switch Exn.asJsExn(exn) { - | Some(jsExn) => Js.Exn.message(jsExn)->Belt.Option.getWithDefault("unknown") - | None => Exn.message(exn)->Option.getOr("unknown") - } - Error("Failed to fetch from source: " ++ msg) - } - } - - switch fetchResult { - | Error(msg) => Error(msg) - | Ok(hexadData) => { - // Push to each target store - let errors = ref([]) - - let pushPromises = targetStores->Belt.Array.map(targetId => { - switch Js.Dict.get(registry.stores, targetId) { - | None => { - errors := Belt.Array.concat(errors.contents, ["Target '" ++ targetId ++ "' not found"]) - Promise.resolve() - } - | Some(target) => { - let pushUrl = target.endpoint ++ "/hexads/" ++ hexadId - Fetch.fetch(pushUrl, { - method: #PUT, - body: Fetch.BodyInit.make(Js.Json.stringify(hexadData)), - headers: Fetch.HeadersInit.make({"Content-Type": "application/json"}), - }) - ->Promise.then(resp => { - if !Fetch.Response.ok(resp) { - errors := Belt.Array.concat(errors.contents, [ - "Push to '" ++ targetId ++ "' failed: " ++ Belt.Int.toString(Fetch.Response.status(resp)) - ]) - } - Promise.resolve() - }) - ->Promise.catch(_err => { - errors := Belt.Array.concat(errors.contents, ["Push to '" ++ targetId ++ "' failed: network error"]) - Promise.resolve() - }) - } - } - }) - - let _ = await Promise.all(pushPromises) - - if Belt.Array.length(errors.contents) > 0 { - Error(Belt.Array.joinWith(errors.contents, "; ")) - } else { - Ok() - } - } - } - } - } -} - -// ============================================================================ -// Trust & Byzantine Fault Tolerance -// ============================================================================ - -type consensusResult<'a> = { - value: 'a, - agreement: float, // 0.0-1.0 - participants: array<storeId>, -} - -// Achieve consensus across stores using quorum voting -let achieveConsensus = async ( - registry: registryState, - stores: array<storeId>, - getValue: storeId => Promise.t<option<'a>>, -): Promise.t<option<consensusResult<'a>>> => { - // Fetch values from all stores concurrently - let fetchPromises = stores->Belt.Array.map(id => { - getValue(id)->Promise.then(result => Promise.resolve((id, result))) - }) - - let results = await Promise.all(fetchPromises) - - // Filter to stores that returned values - let successful = results->Belt.Array.keepMap(((id, result)) => { - switch result { - | Some(val) => Some((id, val)) - | None => None - } - }) - - let totalResponders = Belt.Array.length(successful) - if totalResponders == 0 { - None - } else { - // Determine quorum threshold based on consistency mode - let quorumThreshold = switch registry.config.consistencyMode { - | Strong => Belt.Array.length(stores) // All must agree - | Quorum => Belt.Array.length(stores) / 2 + 1 // Majority - | Eventual => 1 // Any response suffices - } - - if totalResponders >= quorumThreshold { - // Return the first value (in a full implementation, would compare values - // and select the one with the most votes) - let (_, firstValue) = successful->Belt.Array.getExn(0) - let participants = successful->Belt.Array.map(((id, _)) => id) - let agreement = Belt.Int.toFloat(totalResponders) /. Belt.Int.toFloat(Belt.Array.length(stores)) - - Some({ - value: firstValue, - agreement: agreement, - participants: participants, - }) - } else { - None // Quorum not met - } - } -} - -// Detect Byzantine faults by identifying stores with anomalous trust levels -// (proxy for divergent data — a full implementation would compare actual responses) -let detectByzantineFaults = ( - registry: registryState, - hexadId: hexadId, -): array<storeId> => { - switch Js.Dict.get(registry.mappings, hexadId) { - | None => [] - | Some(mapping) => { - let locations = Js.Dict.values(mapping.locations) - let storeCount = Belt.Array.length(locations) - - if storeCount < 2 { - [] // Need at least 2 stores to detect divergence - } else { - // Compute median trust level - let trusts = locations->Belt.Array.map(loc => { - switch Js.Dict.get(registry.stores, loc.storeId) { - | None => 0.0 - | Some(store) => store.trustLevel - } - }) - let sorted = trusts->Belt.SortArray.stableSortBy((a, b) => Belt.Float.toInt((a -. b) *. 1000.0)) - let median = switch Belt.Array.get(sorted, storeCount / 2) { - | Some(m) => m - | None => 0.5 - } - - // Flag stores that deviate significantly from the median (> 0.3 difference) - locations->Belt.Array.keepMap(loc => { - let trust = switch Js.Dict.get(registry.stores, loc.storeId) { - | None => 0.0 - | Some(store) => store.trustLevel - } - if Js.Math.abs_float(trust -. median) > 0.3 { - Some(loc.storeId) - } else { - None - } - }) - } - } - } -} - -// ============================================================================ -// Serialization -// ============================================================================ - -let modalityToString = (m: modalityType): string => { - switch m { - | Graph => "graph" - | Vector => "vector" - | Tensor => "tensor" - | Semantic => "semantic" - | Document => "document" - | Temporal => "temporal" - | Provenance => "provenance" - | Spatial => "spatial" - } -} - -let modalityFromString = (s: string): option<modalityType> => { - switch s { - | "graph" => Some(Graph) - | "vector" => Some(Vector) - | "tensor" => Some(Tensor) - | "semantic" => Some(Semantic) - | "document" => Some(Document) - | "temporal" => Some(Temporal) - | "provenance" => Some(Provenance) - | "spatial" => Some(Spatial) - | _ => None - } -} - -let serializeStoreLocation = (store: storeLocation): Js.Json.t => { - Js.Dict.fromArray([ - ("storeId", Js.Json.string(store.storeId)), - ("endpoint", Js.Json.string(store.endpoint)), - ("modalities", Js.Json.array(store.modalities->Belt.Array.map(m => Js.Json.string(modalityToString(m))))), - ("trustLevel", Js.Json.number(store.trustLevel)), - ("lastSeen", Js.Json.string(Js.Date.toISOString(store.lastSeen))), - ("responseTimeMs", switch store.responseTimeMs { - | None => Js.Json.null - | Some(ms) => Js.Json.number(Belt.Int.toFloat(ms)) - }), - ])->Js.Json.object_ -} - -let serializeRegistry = (registry: registryState): Js.Json.t => { - // Serialize stores - let storesJson = Js.Dict.empty() - registry.stores->Js.Dict.entries->Belt.Array.forEach(((id, store)) => { - Js.Dict.set(storesJson, id, serializeStoreLocation(store)) - }) - - // Serialize mappings - let mappingsJson = Js.Dict.empty() - registry.mappings->Js.Dict.entries->Belt.Array.forEach(((id, mapping)) => { - let locsJson = Js.Dict.empty() - mapping.locations->Js.Dict.entries->Belt.Array.forEach(((key, loc)) => { - Js.Dict.set(locsJson, key, serializeStoreLocation(loc)) - }) - Js.Dict.set(mappingsJson, id, Js.Dict.fromArray([ - ("hexadId", Js.Json.string(mapping.hexadId)), - ("locations", Js.Json.object_(locsJson)), - ("primaryStore", switch mapping.primaryStore { - | None => Js.Json.null - | Some(s) => Js.Json.string(s) - }), - ("created", Js.Json.string(Js.Date.toISOString(mapping.created))), - ("modified", Js.Json.string(Js.Date.toISOString(mapping.modified))), - ])->Js.Json.object_) - }) - - // Serialize config - let configJson = Js.Dict.fromArray([ - ("minTrustLevel", Js.Json.number(registry.config.minTrustLevel)), - ("maxStoreDowntimeMs", Js.Json.number(Belt.Int.toFloat(registry.config.maxStoreDowntimeMs))), - ("replicationFactor", Js.Json.number(Belt.Int.toFloat(registry.config.replicationFactor))), - ("consistencyMode", Js.Json.string(switch registry.config.consistencyMode { - | Strong => "strong" - | Eventual => "eventual" - | Quorum => "quorum" - })), - ])->Js.Json.object_ - - Js.Dict.fromArray([ - ("stores", Js.Json.object_(storesJson)), - ("mappings", Js.Json.object_(mappingsJson)), - ("config", configJson), - ])->Js.Json.object_ -} - -let deserializeStoreLocation = (json: Js.Json.t): option<storeLocation> => { - switch Js.Json.classify(json) { - | Js.Json.JSONObject(obj) => { - let getString = key => switch Js.Dict.get(obj, key) { - | Some(v) => switch Js.Json.classify(v) { - | Js.Json.JSONString(s) => Some(s) - | _ => None - } - | None => None - } - let getFloat = key => switch Js.Dict.get(obj, key) { - | Some(v) => switch Js.Json.classify(v) { - | Js.Json.JSONNumber(n) => Some(n) - | _ => None - } - | None => None - } - - switch (getString("storeId"), getString("endpoint")) { - | (Some(sid), Some(ep)) => { - let modalities = switch Js.Dict.get(obj, "modalities") { - | Some(arr) => switch Js.Json.classify(arr) { - | Js.Json.JSONArray(items) => - items->Belt.Array.keepMap(item => { - switch Js.Json.classify(item) { - | Js.Json.JSONString(s) => modalityFromString(s) - | _ => None - } - }) - | _ => [] - } - | None => [] - } - - let responseTimeMs = switch getFloat("responseTimeMs") { - | Some(n) => Some(Belt.Float.toInt(n)) - | None => None - } - - Some({ - storeId: sid, - endpoint: ep, - modalities: modalities, - trustLevel: getFloat("trustLevel")->Belt.Option.getWithDefault(1.0), - lastSeen: switch getString("lastSeen") { - | Some(s) => Js.Date.fromString(s) - | None => Js.Date.make() - }, - responseTimeMs: responseTimeMs, - }) - } - | _ => None - } - } - | _ => None - } -} - -let deserializeRegistry = (json: Js.Json.t): option<registryState> => { - switch Js.Json.classify(json) { - | Js.Json.JSONObject(root) => { - // Deserialize config - let config = switch Js.Dict.get(root, "config") { - | Some(configJson) => switch Js.Json.classify(configJson) { - | Js.Json.JSONObject(obj) => { - let getFloat = key => switch Js.Dict.get(obj, key) { - | Some(v) => switch Js.Json.classify(v) { - | Js.Json.JSONNumber(n) => Some(n) - | _ => None - } - | None => None - } - let getString = key => switch Js.Dict.get(obj, key) { - | Some(v) => switch Js.Json.classify(v) { - | Js.Json.JSONString(s) => Some(s) - | _ => None - } - | None => None - } - - { - minTrustLevel: getFloat("minTrustLevel")->Belt.Option.getWithDefault(0.5), - maxStoreDowntimeMs: getFloat("maxStoreDowntimeMs") - ->Belt.Option.map(Belt.Float.toInt) - ->Belt.Option.getWithDefault(300_000), - replicationFactor: getFloat("replicationFactor") - ->Belt.Option.map(Belt.Float.toInt) - ->Belt.Option.getWithDefault(3), - consistencyMode: switch getString("consistencyMode") { - | Some("strong") => Strong - | Some("eventual") => Eventual - | _ => Quorum - }, - } - } - | _ => defaultConfig() - } - | None => defaultConfig() - } - - // Deserialize stores - let stores = Js.Dict.empty() - switch Js.Dict.get(root, "stores") { - | Some(storesJson) => switch Js.Json.classify(storesJson) { - | Js.Json.JSONObject(storesObj) => - storesObj->Js.Dict.entries->Belt.Array.forEach(((id, storeJson)) => { - switch deserializeStoreLocation(storeJson) { - | Some(store) => Js.Dict.set(stores, id, store) - | None => () - } - }) - | _ => () - } - | None => () - } - - // Deserialize mappings - let mappings = Js.Dict.empty() - switch Js.Dict.get(root, "mappings") { - | Some(mappingsJson) => switch Js.Json.classify(mappingsJson) { - | Js.Json.JSONObject(mappingsObj) => - mappingsObj->Js.Dict.entries->Belt.Array.forEach(((id, mapJson)) => { - switch Js.Json.classify(mapJson) { - | Js.Json.JSONObject(mapObj) => { - let getString = key => switch Js.Dict.get(mapObj, key) { - | Some(v) => switch Js.Json.classify(v) { - | Js.Json.JSONString(s) => Some(s) - | _ => None - } - | None => None - } - - let locations = Js.Dict.empty() - switch Js.Dict.get(mapObj, "locations") { - | Some(locsJson) => switch Js.Json.classify(locsJson) { - | Js.Json.JSONObject(locsObj) => - locsObj->Js.Dict.entries->Belt.Array.forEach(((key, locJson)) => { - switch deserializeStoreLocation(locJson) { - | Some(loc) => Js.Dict.set(locations, key, loc) - | None => () - } - }) - | _ => () - } - | None => () - } - - let primaryStore = switch Js.Dict.get(mapObj, "primaryStore") { - | Some(v) => switch Js.Json.classify(v) { - | Js.Json.JSONString(s) => Some(s) - | Js.Json.JSONNull => None - | _ => None - } - | None => None - } - - Js.Dict.set(mappings, id, { - hexadId: getString("hexadId")->Belt.Option.getWithDefault(id), - locations: locations, - primaryStore: primaryStore, - created: switch getString("created") { - | Some(s) => Js.Date.fromString(s) - | None => Js.Date.make() - }, - modified: switch getString("modified") { - | Some(s) => Js.Date.fromString(s) - | None => Js.Date.make() - }, - }) - } - | _ => () - } - }) - | _ => () - } - | None => () - } - - Some({ - mappings: mappings, - stores: stores, - config: config, - }) - } - | _ => None - } -} - -// ============================================================================ -// Public API -// ============================================================================ - -let create = createRegistry -let register = registerStore -let map = mapHexad -let lookup = getHexadLocations -let selectStore = selectBestStore -let updateHealth = updateStoreHealth -let prune = pruneDeadStores -let query = executeFederatedQuery -let replicate = replicateHexad -let consensus = achieveConsensus diff --git a/verisimdb/src/vcl/VCLBidir.res b/verisimdb/src/vcl/VCLBidir.res deleted file mode 100644 index 427a5091..00000000 --- a/verisimdb/src/vcl/VCLBidir.res +++ /dev/null @@ -1,852 +0,0 @@ -// SPDX-License-Identifier: MPL-2.0 -// VCL Bidirectional Type Inference -// -// Implements bidirectional type checking for VCL queries: -// - synthesize: infer the type of a query from its structure -// - check: verify a query expression has an expected type -// -// The synthesizer walks the query AST and produces a typed result, -// verifying that all field references, operators, aggregates, and -// proof obligations are well-typed. - -module AST = VCLParser.AST -module Types = VCLTypes -module Ctx = VCLContext -module Sub = VCLSubtyping - -// ============================================================================ -// Type Error -// ============================================================================ - -type typeError = - | SubtypingFailed({expected: Types.vclType, got: Types.vclType, reason: string}) - | FieldTypeMismatch({field: string, expected: Types.primitiveType, got: Types.primitiveType}) - | OperatorTypeMismatch({op: string, leftType: Types.primitiveType, rightType: Types.primitiveType}) - | VectorDimensionMismatch({expected: int, got: int}) - | ProofObligationFailed({proofKind: Types.proofKind, reason: string}) - | AggregateTypeMismatch({func: string, fieldType: Types.primitiveType}) - | MultiProofConflict({proof1: string, proof2: string, reason: string}) - | UnknownField({modality: Types.modalityType, fieldName: string}) - | UnknownContract(string) - | UnknownModality(string) - | MissingProof - | InvalidSource(string) - // Phase 2: Cross-modal errors - | CrossModalTypeMismatch({ - mod1: Types.modalityType, - field1: string, - mod2: Types.modalityType, - field2: string, - reason: string, - }) - | DriftRequiresNumeric({mod1: Types.modalityType, mod2: Types.modalityType}) - | ConsistencyMetricInvalid({mod1: Types.modalityType, mod2: Types.modalityType, metric: string}) - // Phase 3: Mutation errors - | InsertModalityMismatch({modality: string, reason: string}) - | UpdateFieldNotFound({hexadId: string, field: string}) - | MutationProofFailed({operation: string, reason: string}) - -let formatTypeError = (err: typeError): string => { - switch err { - | SubtypingFailed({expected, got, reason}) => - `Subtyping failed: expected ${Types.vclTypeToString(expected)}, got ${Types.vclTypeToString(got)}: ${reason}` - | FieldTypeMismatch({field, expected, got}) => - `Field '${field}' type mismatch: expected ${Types.primitiveTypeToString(expected)}, got ${Types.primitiveTypeToString(got)}` - | OperatorTypeMismatch({op, leftType, rightType}) => - `Operator '${op}' cannot compare ${Types.primitiveTypeToString(leftType)} with ${Types.primitiveTypeToString(rightType)}` - | VectorDimensionMismatch({expected, got}) => - `Vector dimension mismatch: expected ${Belt.Int.toString(expected)}, got ${Belt.Int.toString(got)}` - | ProofObligationFailed({proofKind, reason}) => - `Proof obligation failed for ${Types.proofKindToString(proofKind)}: ${reason}` - | AggregateTypeMismatch({func, fieldType}) => - `Aggregate function ${func} cannot operate on ${Types.primitiveTypeToString(fieldType)}` - | MultiProofConflict({proof1, proof2, reason}) => - `Proofs '${proof1}' and '${proof2}' conflict: ${reason}` - | UnknownField({modality, fieldName}) => - `Unknown field '${fieldName}' for modality ${Types.modalityTypeToString(modality)}` - | UnknownContract(name) => `Unknown contract: '${name}'` - | UnknownModality(name) => `Unknown modality: '${name}'` - | MissingProof => "Dependent-type query requires PROOF clause" - | InvalidSource(reason) => `Invalid source: ${reason}` - | CrossModalTypeMismatch({mod1, field1, mod2, field2, reason}) => - `Cross-modal type mismatch: ${Types.modalityTypeToString(mod1)}.${field1} vs ${Types.modalityTypeToString(mod2)}.${field2}: ${reason}` - | DriftRequiresNumeric({mod1, mod2}) => - `DRIFT requires numeric/vector modalities, got ${Types.modalityTypeToString(mod1)} and ${Types.modalityTypeToString(mod2)}` - | ConsistencyMetricInvalid({mod1, mod2, metric}) => - `Metric '${metric}' not supported for ${Types.modalityTypeToString(mod1)} and ${Types.modalityTypeToString(mod2)}` - | InsertModalityMismatch({modality, reason}) => - `INSERT modality '${modality}' error: ${reason}` - | UpdateFieldNotFound({hexadId, field}) => - `UPDATE field '${field}' not found in hexad '${hexadId}'` - | MutationProofFailed({operation, reason}) => - `${operation} proof failed: ${reason}` - } -} - -// ============================================================================ -// Synthesize: Infer type of a complete query -// ============================================================================ - -type synthesizeResult = Result<Types.vclType, typeError> - -let synthesizeQuery = (ctx: Ctx.context, query: AST.query): synthesizeResult => { - // 1. Resolve modalities to type-level representations - let resolvedMods = Types.resolveModalities(query.modalities) - if Js.Array2.length(resolvedMods) == 0 { - Error(UnknownModality("No valid modalities in SELECT")) - } else { - // 2. Check source validity - switch checkSource(ctx, query.source) { - | Error(e) => Error(e) - | Ok() => - // 3. Check WHERE conditions against available modalities - switch checkWhereClause(ctx, query.where, resolvedMods) { - | Error(e) => Error(e) - | Ok() => - // 4. Check projections - switch checkProjections(ctx, query.projections, resolvedMods) { - | Error(e) => Error(e) - | Ok(projTypeInfos) => - // 5. Check aggregates - switch checkAggregates(ctx, query.aggregates, resolvedMods) { - | Error(e) => Error(e) - | Ok(aggTypeInfos) => - // 6. Check GROUP BY fields - switch checkGroupBy(ctx, query.groupBy, resolvedMods) { - | Error(e) => Error(e) - | Ok() => - // 7. Check ORDER BY fields - switch checkOrderBy(ctx, query.orderBy, resolvedMods) { - | Error(e) => Error(e) - | Ok() => - // 8. Build the query result type - let resultInfo: Types.queryResultInfo = { - modalities: resolvedMods, - projections: projTypeInfos, - aggregates: aggTypeInfos, - } - // 9. Handle proof clause - switch query.proof { - | None => - // Slipstream path: just the query result type - Ok(Types.QueryResultType(resultInfo)) - | Some(proofSpecs) => - // Dependent-type path: synthesize proved result - switch checkMultiProof(ctx, proofSpecs, resolvedMods) { - | Error(e) => Error(e) - | Ok(proofKinds) => - // For multi-proof, the result is a Sigma type pairing result with first proof - // (each additional proof adds another layer) - switch proofKinds[0] { - | Some((kind, contract)) => - Ok(Types.ProvedResultType(resultInfo, kind, contract)) - | None => Error(MissingProof) - } - } - } - } - } - } - } - } - } - } -} - -// ============================================================================ -// Check: Verify a query has expected type -// ============================================================================ - -let checkQuery = ( - ctx: Ctx.context, - query: AST.query, - expectedType: Types.vclType, -): Result<unit, typeError> => { - switch synthesizeQuery(ctx, query) { - | Error(e) => Error(e) - | Ok(inferredType) => - switch Sub.isSubtype(inferredType, expectedType) { - | Ok() => Ok() - | Error({expected, got, reason}) => Error(SubtypingFailed({expected, got, reason})) - } - } -} - -// ============================================================================ -// Source checking -// ============================================================================ - -let checkSource = (_ctx: Ctx.context, source: AST.source): Result<unit, typeError> => { - switch source { - | Hexad(id) => - // UUID format validation is done by parser; just verify non-empty - if Js.String2.length(id) > 0 { - Ok() - } else { - Error(InvalidSource("Empty hexad ID")) - } - | Federation(pattern, _drift) => - if Js.String2.length(pattern) > 0 { - Ok() - } else { - Error(InvalidSource("Empty federation pattern")) - } - | Store(storeId) => - if Js.String2.length(storeId) > 0 { - Ok() - } else { - Error(InvalidSource("Empty store ID")) - } - } -} - -// ============================================================================ -// WHERE clause checking -// ============================================================================ - -let checkWhereClause = ( - ctx: Ctx.context, - where: option<AST.condition>, - availableMods: array<Types.modalityType>, -): Result<unit, typeError> => { - switch where { - | None => Ok() - | Some(condition) => checkCondition(ctx, condition, availableMods) - } -} - -and checkCondition = ( - ctx: Ctx.context, - condition: AST.condition, - availableMods: array<Types.modalityType>, -): Result<unit, typeError> => { - switch condition { - | Simple(sc) => checkSimpleCondition(ctx, sc, availableMods) - | And(left, right) => - switch checkCondition(ctx, left, availableMods) { - | Error(e) => Error(e) - | Ok() => checkCondition(ctx, right, availableMods) - } - | Or(left, right) => - switch checkCondition(ctx, left, availableMods) { - | Error(e) => Error(e) - | Ok() => checkCondition(ctx, right, availableMods) - } - | Not(inner) => checkCondition(ctx, inner, availableMods) - } -} - -and checkSimpleCondition = ( - ctx: Ctx.context, - sc: AST.simpleCondition, - _availableMods: array<Types.modalityType>, -): Result<unit, typeError> => { - switch sc { - | FulltextContains(_text) => - // Full-text search is always valid if Document modality is available - Ok() - | FulltextMatches(_pattern) => - Ok() - | FieldCondition(fieldName, op, literal) => - // Infer the literal type - let litType = inferLiteralType(literal) - // Check operator validity for this type - if Types.isOperatorValidForType(op, litType) { - Ok() - } else { - let opStr = operatorToString(op) - Error(OperatorTypeMismatch({ - op: opStr, - leftType: litType, - rightType: litType, - })) - } - | VectorSimilar(embedding, _threshold) => - // Check that embedding is non-empty - if Js.Array2.length(embedding) == 0 { - Error(VectorDimensionMismatch({expected: 1, got: 0})) - } else { - Ok() - } - | GraphPattern(_pattern) => - // Graph pattern validation is deferred to the graph engine - Ok() - // Phase 2: Cross-modal conditions - | CrossModalFieldCompare(mod1, field1, _op, mod2, field2) => - checkCrossModalFieldCompare(ctx, mod1, field1, mod2, field2) - | ModalityDrift(mod1, mod2, _threshold) => - checkDriftTypes(mod1, mod2) - | ModalityExists(_modality) => Ok() - | ModalityNotExists(_modality) => Ok() - | ModalityConsistency(mod1, mod2, metric) => - checkConsistencyTypes(mod1, mod2, metric) - } -} - -// ============================================================================ -// Phase 2: Cross-modal type checking -// ============================================================================ - -and checkCrossModalFieldCompare = ( - ctx: Ctx.context, - mod1: AST.modality, - field1: string, - mod2: AST.modality, - field2: string, -): Result<unit, typeError> => { - switch (Types.modalityTypeOfAstModality(mod1), Types.modalityTypeOfAstModality(mod2)) { - | (Some(mt1), Some(mt2)) => - switch (Ctx.lookupField(ctx, mt1, field1), Ctx.lookupField(ctx, mt2, field2)) { - | (Some(f1), Some(f2)) => - // Both fields must have compatible types for comparison - if Types.eqPrimitiveType(f1.fieldType, f2.fieldType) || - Sub.isSubPrimitive(f1.fieldType, f2.fieldType) || - Sub.isSubPrimitive(f2.fieldType, f1.fieldType) { - Ok() - } else { - Error(CrossModalTypeMismatch({ - mod1: mt1, - field1, - mod2: mt2, - field2, - reason: `${Types.primitiveTypeToString(f1.fieldType)} vs ${Types.primitiveTypeToString(f2.fieldType)}`, - })) - } - | (None, _) => Error(UnknownField({modality: mt1, fieldName: field1})) - | (_, None) => Error(UnknownField({modality: mt2, fieldName: field2})) - } - | (None, _) => Error(UnknownModality("All")) - | (_, None) => Error(UnknownModality("All")) - } -} - -and checkDriftTypes = (mod1: AST.modality, mod2: AST.modality): Result<unit, typeError> => { - // DRIFT requires both modalities to have numeric/vector representations - switch (Types.modalityTypeOfAstModality(mod1), Types.modalityTypeOfAstModality(mod2)) { - | (Some(mt1), Some(mt2)) => - // All modality pairs support drift computation via their canonical embeddings - let _ = (mt1, mt2) - Ok() - | _ => Error(DriftRequiresNumeric({ - mod1: Types.modalityTypeOfAstModality(mod1)->Belt.Option.getWithDefault(Types.GraphModality), - mod2: Types.modalityTypeOfAstModality(mod2)->Belt.Option.getWithDefault(Types.GraphModality), - })) - } -} - -and checkConsistencyTypes = ( - mod1: AST.modality, - mod2: AST.modality, - metric: string, -): Result<unit, typeError> => { - let validMetrics = ["COSINE", "EUCLIDEAN", "DOT_PRODUCT", "JACCARD"] - if !(validMetrics->Js.Array2.includes(Js.String2.toUpperCase(metric))) { - switch (Types.modalityTypeOfAstModality(mod1), Types.modalityTypeOfAstModality(mod2)) { - | (Some(mt1), Some(mt2)) => - Error(ConsistencyMetricInvalid({mod1: mt1, mod2: mt2, metric})) - | _ => - Error(ConsistencyMetricInvalid({ - mod1: Types.GraphModality, - mod2: Types.GraphModality, - metric, - })) - } - } else { - Ok() - } -} - -// ============================================================================ -// Projection checking -// ============================================================================ - -let checkProjections = ( - ctx: Ctx.context, - projections: option<array<AST.fieldRef>>, - availableMods: array<Types.modalityType>, -): Result<array<Types.fieldTypeInfo>, typeError> => { - switch projections { - | None => Ok([]) - | Some(projs) => - projs->Belt.Array.reduce(Ok([]), (acc, proj) => { - switch acc { - | Error(e) => Error(e) - | Ok(infos) => - switch Types.modalityTypeOfAstModality(proj.modality) { - | None => - // 'All' modality — skip projection type check - Ok(infos) - | Some(modType) => - // Verify modality is in SELECT - if !(availableMods->Js.Array2.some(m => Types.eqModalityType(m, modType))) { - Error(UnknownModality(Types.modalityTypeToString(modType))) - } else { - // Look up field type - switch Ctx.lookupField(ctx, modType, proj.field) { - | None => - // Field not in registry — allow it (dynamic schema) but type as String - let info: Types.fieldTypeInfo = { - modality: modType, - fieldName: proj.field, - fieldType: Types.StringType, - } - Ok(infos->Js.Array2.concat([info])) - | Some(fieldEntry) => - let info: Types.fieldTypeInfo = { - modality: modType, - fieldName: proj.field, - fieldType: fieldEntry.fieldType, - } - Ok(infos->Js.Array2.concat([info])) - } - } - } - } - }) - } -} - -// ============================================================================ -// Aggregate checking -// ============================================================================ - -let checkAggregates = ( - ctx: Ctx.context, - aggregates: option<array<AST.aggregateExpr>>, - availableMods: array<Types.modalityType>, -): Result<array<Types.aggregateTypeInfo>, typeError> => { - switch aggregates { - | None => Ok([]) - | Some(aggs) => - aggs->Belt.Array.reduce(Ok([]), (acc, agg) => { - switch acc { - | Error(e) => Error(e) - | Ok(infos) => - switch agg { - | CountAll => - let info: Types.aggregateTypeInfo = { - func: AST.Count, - resultType: Types.IntType, - sourceField: None, - } - Ok(infos->Js.Array2.concat([info])) - | AggregateField(func, fieldRef) => - switch Types.modalityTypeOfAstModality(fieldRef.modality) { - | None => Ok(infos) // All modality — skip - | Some(modType) => - if !(availableMods->Js.Array2.some(m => Types.eqModalityType(m, modType))) { - Error(UnknownModality(Types.modalityTypeToString(modType))) - } else { - let fieldType = switch Ctx.lookupField(ctx, modType, fieldRef.field) { - | Some(f) => f.fieldType - | None => Types.FloatType // default for unknown fields - } - // SUM, AVG require numeric types - switch func { - | Sum | Avg => - if !Types.isNumericPrimitive(fieldType) { - let funcStr = switch func { - | Sum => "SUM" - | Avg => "AVG" - | Count => "COUNT" - | Min => "MIN" - | Max => "MAX" - } - Error(AggregateTypeMismatch({func: funcStr, fieldType})) - } else { - let resultType = switch func { - | Avg => Types.FloatType - | _ => fieldType - } - let sourceInfo: Types.fieldTypeInfo = { - modality: modType, - fieldName: fieldRef.field, - fieldType, - } - let info: Types.aggregateTypeInfo = { - func, - resultType, - sourceField: Some(sourceInfo), - } - Ok(infos->Js.Array2.concat([info])) - } - | Count => - let sourceInfo: Types.fieldTypeInfo = { - modality: modType, - fieldName: fieldRef.field, - fieldType, - } - let info: Types.aggregateTypeInfo = { - func, - resultType: Types.IntType, - sourceField: Some(sourceInfo), - } - Ok(infos->Js.Array2.concat([info])) - | Min | Max => - if !Types.isComparablePrimitive(fieldType) { - let funcStr = switch func { - | Min => "MIN" - | Max => "MAX" - | _ => "?" - } - Error(AggregateTypeMismatch({func: funcStr, fieldType})) - } else { - let sourceInfo: Types.fieldTypeInfo = { - modality: modType, - fieldName: fieldRef.field, - fieldType, - } - let info: Types.aggregateTypeInfo = { - func, - resultType: fieldType, - sourceField: Some(sourceInfo), - } - Ok(infos->Js.Array2.concat([info])) - } - } - } - } - } - } - }) - } -} - -// ============================================================================ -// GROUP BY / ORDER BY checking -// ============================================================================ - -let checkGroupBy = ( - ctx: Ctx.context, - groupBy: option<array<AST.fieldRef>>, - availableMods: array<Types.modalityType>, -): Result<unit, typeError> => { - switch groupBy { - | None => Ok() - | Some(fields) => - fields->Belt.Array.reduce(Ok(), (acc, field) => { - switch acc { - | Error(e) => Error(e) - | Ok() => - switch Types.modalityTypeOfAstModality(field.modality) { - | None => Ok() - | Some(modType) => - if !(availableMods->Js.Array2.some(m => Types.eqModalityType(m, modType))) { - Error(UnknownModality(Types.modalityTypeToString(modType))) - } else { - // Verify field exists (or accept dynamic) - let _ = Ctx.lookupField(ctx, modType, field.field) - Ok() - } - } - } - }) - } -} - -let checkOrderBy = ( - _ctx: Ctx.context, - orderBy: option<array<AST.orderByItem>>, - availableMods: array<Types.modalityType>, -): Result<unit, typeError> => { - switch orderBy { - | None => Ok() - | Some(items) => - items->Belt.Array.reduce(Ok(), (acc, item) => { - switch acc { - | Error(e) => Error(e) - | Ok() => - switch Types.modalityTypeOfAstModality(item.field.modality) { - | None => Ok() - | Some(modType) => - if !(availableMods->Js.Array2.some(m => Types.eqModalityType(m, modType))) { - Error(UnknownModality(Types.modalityTypeToString(modType))) - } else { - Ok() - } - } - } - }) - } -} - -// ============================================================================ -// Multi-proof checking -// ============================================================================ - -let checkMultiProof = ( - ctx: Ctx.context, - proofSpecs: array<AST.proofSpec>, - _availableMods: array<Types.modalityType>, -): Result<array<(Types.proofKind, string)>, typeError> => { - if Js.Array2.length(proofSpecs) == 0 { - Error(MissingProof) - } else { - // Check each proof spec individually - let results = proofSpecs->Belt.Array.map(spec => { - let kind = Types.proofKindOfAstProofType(spec.proofType) - // If contract registry has this contract, verify compatibility - switch Ctx.lookupContract(ctx, spec.contractName) { - | None => - // Contract not in registry — accept it (registry may not be populated) - Ok((kind, spec.contractName)) - | Some(contractSpec) => - // Verify proof kind matches contract - if contractSpec.proofKind != kind { - Error(ProofObligationFailed({ - proofKind: kind, - reason: `Contract '${spec.contractName}' expects ${Types.proofKindToString(contractSpec.proofKind)}, got ${Types.proofKindToString(kind)}`, - })) - } else { - Ok((kind, spec.contractName)) - } - } - }) - - // Check for errors - let firstError = results->Belt.Array.getBy(r => { - switch r { - | Error(_) => true - | Ok(_) => false - } - }) - - switch firstError { - | Some(Error(e)) => Error(e) - | _ => - // Extract successful results - let kinds = results->Belt.Array.keepMap(r => { - switch r { - | Ok(v) => Some(v) - | Error(_) => None - } - }) - - // Check mutual composability - if Js.Array2.length(kinds) > 1 { - let contractNames = kinds->Belt.Array.map(((_, c)) => c) - if !Ctx.areProofsComposable(ctx, contractNames) { - // Find the first conflicting pair - let len = Js.Array2.length(contractNames) - let conflict = ref(None) - for i in 0 to len - 2 { - for j in i + 1 to len - 1 { - switch (contractNames[i], contractNames[j]) { - | (Some(c1), Some(c2)) => - if !Ctx.canComposeProofs(ctx, c1, c2) && conflict.contents->Belt.Option.isNone { - conflict := Some((c1, c2)) - } - | _ => () - } - } - } - switch conflict.contents { - | Some((c1, c2)) => - Error(MultiProofConflict({proof1: c1, proof2: c2, reason: "Contracts are not composable"})) - | None => - // If no specific conflict found but composability check failed, - // the contracts may not have composability info — allow it - Ok(kinds) - } - } else { - Ok(kinds) - } - } else { - Ok(kinds) - } - } - } -} - -// ============================================================================ -// Phase 3: Mutation type checking -// ============================================================================ - -let synthesizeMutation = ( - ctx: Ctx.context, - mutation: AST.mutation, -): synthesizeResult => { - switch mutation { - | Insert({modalities: modalityData, proof}) => - // Check each modality data entry is well-formed - switch checkModalityDataArray(ctx, modalityData) { - | Error(e) => Error(e) - | Ok() => - switch proof { - | None => Ok(Types.UnitType) - | Some(proofSpecs) => - let allMods = Types.allModalityTypes - switch checkMultiProof(ctx, proofSpecs, allMods) { - | Error(e) => Error(e) - | Ok(_) => Ok(Types.UnitType) - } - } - } - | Update({hexadId, sets, proof}) => - // Check hexad ID is non-empty - if Js.String2.length(hexadId) == 0 { - Error(InvalidSource("Empty hexad ID in UPDATE")) - } else { - // Check each SET assignment - switch checkSetAssignments(ctx, sets) { - | Error(e) => Error(e) - | Ok() => - switch proof { - | None => Ok(Types.UnitType) - | Some(proofSpecs) => - let allMods = Types.allModalityTypes - switch checkMultiProof(ctx, proofSpecs, allMods) { - | Error(e) => Error(e) - | Ok(_) => Ok(Types.UnitType) - } - } - } - } - | Delete({hexadId, proof}) => - if Js.String2.length(hexadId) == 0 { - Error(InvalidSource("Empty hexad ID in DELETE")) - } else { - switch proof { - | None => Ok(Types.UnitType) - | Some(proofSpecs) => - let allMods = Types.allModalityTypes - switch checkMultiProof(ctx, proofSpecs, allMods) { - | Error(e) => Error(e) - | Ok(_) => Ok(Types.UnitType) - } - } - } - } -} - -and checkModalityDataArray = ( - _ctx: Ctx.context, - data: array<AST.modalityData>, -): Result<unit, typeError> => { - if Js.Array2.length(data) == 0 { - Error(InsertModalityMismatch({modality: "none", reason: "INSERT requires at least one modality data"})) - } else { - data->Belt.Array.reduce(Ok(), (acc, d) => { - switch acc { - | Error(e) => Error(e) - | Ok() => - switch d { - | DocumentData(fields) => - if Js.Array2.length(fields) == 0 { - Error(InsertModalityMismatch({modality: "DOCUMENT", reason: "Empty document data"})) - } else { - Ok() - } - | VectorData(embedding) => - if Js.Array2.length(embedding) == 0 { - Error(InsertModalityMismatch({modality: "VECTOR", reason: "Empty embedding vector"})) - } else { - Ok() - } - | GraphData(edgeType, targetId) => - if Js.String2.length(edgeType) == 0 || Js.String2.length(targetId) == 0 { - Error(InsertModalityMismatch({modality: "GRAPH", reason: "Edge type and target ID required"})) - } else { - Ok() - } - | TensorData(values) => - if Js.Array2.length(values) == 0 { - Error(InsertModalityMismatch({modality: "TENSOR", reason: "Empty tensor data"})) - } else { - Ok() - } - | SemanticData(contractName) => - if Js.String2.length(contractName) == 0 { - Error(InsertModalityMismatch({modality: "SEMANTIC", reason: "Contract name required"})) - } else { - Ok() - } - | TemporalData(timestamp) => - if Js.String2.length(timestamp) == 0 { - Error(InsertModalityMismatch({modality: "TEMPORAL", reason: "Timestamp required"})) - } else { - Ok() - } - | ProvenanceData(fields) => - if Js.Array2.length(fields) == 0 { - Error(InsertModalityMismatch({modality: "PROVENANCE", reason: "Empty provenance data"})) - } else { - Ok() - } - | SpatialData(fields) => - if Js.Array2.length(fields) == 0 { - Error(InsertModalityMismatch({modality: "SPATIAL", reason: "Empty spatial data"})) - } else { - Ok() - } - } - } - }) - } -} - -and checkSetAssignments = ( - ctx: Ctx.context, - sets: array<(AST.fieldRef, AST.literal)>, -): Result<unit, typeError> => { - sets->Belt.Array.reduce(Ok(), (acc, (fieldRef, literal)) => { - switch acc { - | Error(e) => Error(e) - | Ok() => - switch Types.modalityTypeOfAstModality(fieldRef.modality) { - | None => Ok() // All modality — skip validation - | Some(modType) => - switch Ctx.lookupField(ctx, modType, fieldRef.field) { - | None => - // Dynamic schema — accept any field - Ok() - | Some(fieldEntry) => - let litType = inferLiteralType(literal) - if Types.eqPrimitiveType(fieldEntry.fieldType, litType) || - Sub.isSubPrimitive(litType, fieldEntry.fieldType) { - Ok() - } else { - Error(FieldTypeMismatch({ - field: `${Types.modalityTypeToString(modType)}.${fieldRef.field}`, - expected: fieldEntry.fieldType, - got: litType, - })) - } - } - } - } - }) -} - -// ============================================================================ -// Utility functions -// ============================================================================ - -let inferLiteralType = (lit: AST.literal): Types.primitiveType => { - switch lit { - | String(_) => Types.StringType - | Int(_) => Types.IntType - | Float(_) => Types.FloatType - | Bool(_) => Types.BoolType - | Array(arr) => - // Infer element type from first element - switch arr[0] { - | Some(Float(_)) => Types.VectorType(Js.Array2.length(arr)) - | _ => Types.StringType // default - } - } -} - -let operatorToString = (op: AST.operator): string => { - switch op { - | Eq => "==" - | Neq => "!=" - | Gt => ">" - | Lt => "<" - | Gte => ">=" - | Lte => "<=" - | Like => "LIKE" - | Contains => "CONTAINS" - | Matches => "MATCHES" - } -} diff --git a/verisimdb/src/vcl/VCLCircuit.res b/verisimdb/src/vcl/VCLCircuit.res deleted file mode 100644 index 2fc2e836..00000000 --- a/verisimdb/src/vcl/VCLCircuit.res +++ /dev/null @@ -1,71 +0,0 @@ -// SPDX-License-Identifier: MPL-2.0 -// VCL Circuit DSL — Defines types for custom ZKP circuits in VCL. -// -// Usage in VCL: -// PROOF CUSTOM "circuit-name" WITH (threshold=0.5, min_score=0.1) - -/// Gate types available in custom circuits -type gateType = - | AND - | OR - | XOR - | NOT - | LinearCombination - -/// A wire in the circuit (carries a signal) -type wire = { - name: string, - isPublic: bool, - isOutput: bool, -} - -/// A gate connecting input wires to an output wire -type gate = { - gateType: gateType, - inputs: array<string>, - output: string, -} - -/// A constraint in the circuit (R1CS: A * B = C) -type constraint = { - description: string, - a: array<(int, float)>, - b: array<(int, float)>, - c: array<(int, float)>, -} - -/// A circuit definition from VCL PROOF CUSTOM clause -type circuitDef = { - name: string, - wires: array<wire>, - gates: array<gate>, - parameters: array<string>, -} - -/// Parameters passed via VCL WITH clause -type circuitParams = { - values: Js.Dict.t<string>, -} - -/// Result of a custom circuit verification -type verificationResult = { - circuitName: string, - verified: bool, - publicInputs: array<float>, - constraintsSatisfied: int, - totalConstraints: int, -} - -/// Parse a PROOF CUSTOM clause from VCL -let parseCustomProof = (circuitName: string, withParams: array<(string, string)>): (string, circuitParams) => { - let dict = Js.Dict.empty() - withParams->Array.forEach(((key, value)) => { - Js.Dict.set(dict, key, value) - }) - (circuitName, {values: dict}) -} - -/// Serialize a circuit definition to JSON for the Rust bridge -let serializeCircuitDef = (def: circuitDef): string => { - Js.Json.stringifyAny(def)->Option.getOr("{}") -} diff --git a/verisimdb/src/vcl/VCLContext.res b/verisimdb/src/vcl/VCLContext.res deleted file mode 100644 index 937ec4f8..00000000 --- a/verisimdb/src/vcl/VCLContext.res +++ /dev/null @@ -1,247 +0,0 @@ -// SPDX-License-Identifier: MPL-2.0 -// VCL Context — Typing environment for bidirectional type checking -// -// Maintains bindings, contract registry, modality field registries, -// and store capabilities. - -module Types = VCLTypes - -// ============================================================================ -// Contract Specification -// ============================================================================ - -type contractSpec = { - name: string, - proofKind: Types.proofKind, - requiredModalities: array<Types.modalityType>, - requiredFields: array<(Types.modalityType, string, Types.primitiveType)>, - composableWith: array<Types.proofKind>, // proof kinds this can compose with -} - -// ============================================================================ -// Field Registry — known fields per modality -// ============================================================================ - -type fieldEntry = { - fieldName: string, - fieldType: Types.primitiveType, -} - -// ============================================================================ -// Context -// ============================================================================ - -type context = { - bindings: Js.Dict.t<Types.vclType>, - contracts: Js.Dict.t<contractSpec>, - modalityFields: Js.Dict.t<array<fieldEntry>>, - storeModalities: Js.Dict.t<array<Types.modalityType>>, -} - -// ============================================================================ -// Construction -// ============================================================================ - -let empty = (): context => { - { - bindings: Js.Dict.empty(), - contracts: Js.Dict.empty(), - modalityFields: Js.Dict.empty(), - storeModalities: Js.Dict.empty(), - } -} - -// Default context with standard modality field registries -let defaultContext = (): context => { - let fields = Js.Dict.empty() - - // Graph modality fields - Js.Dict.set(fields, "GRAPH", [ - {fieldName: "predicate", fieldType: Types.StringType}, - {fieldName: "subject", fieldType: Types.StringType}, - {fieldName: "object", fieldType: Types.StringType}, - {fieldName: "centrality", fieldType: Types.FloatType}, - {fieldName: "degree", fieldType: Types.IntType}, - {fieldName: "edge_type", fieldType: Types.StringType}, - ]) - - // Vector modality fields - Js.Dict.set(fields, "VECTOR", [ - {fieldName: "embedding", fieldType: Types.VectorType(768)}, - {fieldName: "magnitude", fieldType: Types.FloatType}, - {fieldName: "dimension", fieldType: Types.IntType}, - ]) - - // Tensor modality fields - Js.Dict.set(fields, "TENSOR", [ - {fieldName: "rank", fieldType: Types.IntType}, - {fieldName: "dtype", fieldType: Types.StringType}, - {fieldName: "mean", fieldType: Types.FloatType}, - {fieldName: "std", fieldType: Types.FloatType}, - ]) - - // Semantic modality fields - Js.Dict.set(fields, "SEMANTIC", [ - {fieldName: "contract", fieldType: Types.StringType}, - {fieldName: "verified", fieldType: Types.BoolType}, - {fieldName: "verifier", fieldType: Types.StringType}, - ]) - - // Document modality fields - Js.Dict.set(fields, "DOCUMENT", [ - {fieldName: "name", fieldType: Types.StringType}, - {fieldName: "title", fieldType: Types.StringType}, - {fieldName: "severity", fieldType: Types.IntType}, - {fieldName: "author", fieldType: Types.StringType}, - {fieldName: "year", fieldType: Types.IntType}, - {fieldName: "doi", fieldType: Types.StringType}, - {fieldName: "impact_factor", fieldType: Types.FloatType}, - {fieldName: "count", fieldType: Types.IntType}, - {fieldName: "total", fieldType: Types.IntType}, - ]) - - // Temporal modality fields - Js.Dict.set(fields, "TEMPORAL", [ - {fieldName: "timestamp", fieldType: Types.TimestampType}, - {fieldName: "version", fieldType: Types.StringType}, - {fieldName: "actor", fieldType: Types.StringType}, - ]) - - // Provenance modality fields - Js.Dict.set(fields, "PROVENANCE", [ - {fieldName: "origin", fieldType: Types.StringType}, - {fieldName: "actor", fieldType: Types.StringType}, - {fieldName: "event_type", fieldType: Types.StringType}, - {fieldName: "chain_length", fieldType: Types.IntType}, - {fieldName: "chain_valid", fieldType: Types.BoolType}, - {fieldName: "content_hash", fieldType: Types.StringType}, - {fieldName: "description", fieldType: Types.StringType}, - ]) - - // Spatial modality fields - Js.Dict.set(fields, "SPATIAL", [ - {fieldName: "latitude", fieldType: Types.FloatType}, - {fieldName: "longitude", fieldType: Types.FloatType}, - {fieldName: "altitude", fieldType: Types.FloatType}, - {fieldName: "geometry_type", fieldType: Types.StringType}, - {fieldName: "srid", fieldType: Types.IntType}, - ]) - - { - bindings: Js.Dict.empty(), - contracts: Js.Dict.empty(), - modalityFields: fields, - storeModalities: Js.Dict.empty(), - } -} - -// ============================================================================ -// Lookup operations -// ============================================================================ - -let bind = (ctx: context, name: string, ty: Types.vclType): context => { - let newBindings = Js.Dict.fromArray(Js.Dict.entries(ctx.bindings)) - Js.Dict.set(newBindings, name, ty) - {...ctx, bindings: newBindings} -} - -let lookup = (ctx: context, name: string): option<Types.vclType> => { - Js.Dict.get(ctx.bindings, name) -} - -let lookupContract = (ctx: context, name: string): option<contractSpec> => { - Js.Dict.get(ctx.contracts, name) -} - -let lookupModalityFields = (ctx: context, modality: Types.modalityType): array<fieldEntry> => { - let key = Types.modalityTypeToString(modality) - Js.Dict.get(ctx.modalityFields, key)->Belt.Option.getWithDefault([]) -} - -let lookupField = ( - ctx: context, - modality: Types.modalityType, - fieldName: string, -): option<fieldEntry> => { - let fields = lookupModalityFields(ctx, modality) - fields->Belt.Array.getBy(f => f.fieldName == fieldName) -} - -let lookupStoreModalities = (ctx: context, storeId: string): option<array<Types.modalityType>> => { - Js.Dict.get(ctx.storeModalities, storeId) -} - -// ============================================================================ -// Registration operations -// ============================================================================ - -let registerContract = (ctx: context, spec: contractSpec): context => { - let newContracts = Js.Dict.fromArray(Js.Dict.entries(ctx.contracts)) - Js.Dict.set(newContracts, spec.name, spec) - {...ctx, contracts: newContracts} -} - -let registerField = ( - ctx: context, - modality: Types.modalityType, - entry: fieldEntry, -): context => { - let key = Types.modalityTypeToString(modality) - let existing = lookupModalityFields(ctx, modality) - // Only add if not already present - let alreadyExists = existing->Js.Array2.some(f => f.fieldName == entry.fieldName) - if alreadyExists { - ctx - } else { - let newFields = Js.Dict.fromArray(Js.Dict.entries(ctx.modalityFields)) - Js.Dict.set(newFields, key, existing->Js.Array2.concat([entry])) - {...ctx, modalityFields: newFields} - } -} - -let registerStoreModalities = ( - ctx: context, - storeId: string, - modalities: array<Types.modalityType>, -): context => { - let newStores = Js.Dict.fromArray(Js.Dict.entries(ctx.storeModalities)) - Js.Dict.set(newStores, storeId, modalities) - {...ctx, storeModalities: newStores} -} - -// ============================================================================ -// Contract composition checks -// ============================================================================ - -// Check if two proof kinds can be composed together -let canComposeProofs = (ctx: context, contract1: string, contract2: string): bool => { - switch (lookupContract(ctx, contract1), lookupContract(ctx, contract2)) { - | (Some(spec1), Some(spec2)) => - spec1.composableWith->Js.Array2.some(k => k == spec2.proofKind) && - spec2.composableWith->Js.Array2.some(k => k == spec1.proofKind) - | _ => false - } -} - -// Check if a list of proof specs are all mutually composable -let areProofsComposable = (ctx: context, contractNames: array<string>): bool => { - let len = Js.Array2.length(contractNames) - if len <= 1 { - true - } else { - // Check all pairs - let allOk = ref(true) - for i in 0 to len - 2 { - for j in i + 1 to len - 1 { - switch (contractNames[i], contractNames[j]) { - | (Some(c1), Some(c2)) => - if !canComposeProofs(ctx, c1, c2) { - allOk := false - } - | _ => allOk := false - } - } - } - allOk.contents - } -} diff --git a/verisimdb/src/vcl/VCLError.res b/verisimdb/src/vcl/VCLError.res deleted file mode 100644 index 63836932..00000000 --- a/verisimdb/src/vcl/VCLError.res +++ /dev/null @@ -1,458 +0,0 @@ -// SPDX-License-Identifier: MPL-2.0 - -/** - * VCL Error Types - Structured error representation - * - * Provides comprehensive error types for all VCL failure modes: - * - Parse errors (syntax) - * - Type errors (dependent-type verification) - * - Runtime errors (execution) - * - Modality-specific errors - * - Federation errors - */ - -type position = { - line: int, - column: int, - offset: int, -} - -type span = { - start: position, - end_: position, -} - -// ============================================================================ -// Parse Errors -// ============================================================================ - -type parseErrorKind = - | UnexpectedToken({expected: array<string>, found: string}) - | UnterminatedString - | InvalidNumber(string) - | InvalidModality(string) - | InvalidDriftPolicy(string) - | InvalidProofType(string) - | MissingFromClause - | MissingSelectClause - | InvalidGraphPattern(string) - | InvalidVectorExpression(string) - | InvalidSemanticContract(string) - | InvalidAggregateExpression(string) - | InvalidOrderByField(string) - | InvalidGroupByField(string) - | HavingWithoutGroupBy - | AggregateWithoutGroupBy(string) - -type parseError = { - kind: parseErrorKind, - span: span, - source: string, // The original query string - hint: option<string>, -} - -// ============================================================================ -// Type Errors (Dependent-Type Path) -// ============================================================================ - -type typeErrorKind = - | ContractNotFound(string) - | ContractViolation({contract: string, reason: string}) - | ProofGenerationFailed({contract: string, error: string}) - | ProofVerificationFailed({contract: string, reason: string}) - | TypeMismatch({expected: string, found: string}) - | MissingTypeAnnotation(string) - | CircularDependency(array<string>) - // Phase 1: Dependent type errors - | SubtypingFailed({expected: string, got: string}) - | FieldTypeMismatch({field: string, expected: string, got: string}) - | OperatorTypeMismatch({op: string, leftType: string, rightType: string}) - | VectorDimensionMismatch({expected: int, got: int}) - | ProofObligationFailed({proofType: string, reason: string}) - | AggregateTypeMismatch({func: string, fieldType: string}) - | MultiProofConflict({proof1: string, proof2: string, reason: string}) - | UnknownField({modality: string, fieldName: string}) - // Phase 2: Cross-modal errors - | CrossModalTypeMismatch({mod1: string, field1: string, mod2: string, field2: string, reason: string}) - | DriftRequiresNumeric({mod1: string, mod2: string}) - | ConsistencyMetricInvalid({mod1: string, mod2: string, metric: string}) - // Phase 3: Write path errors - | InsertConflict(string) - | UpdateNotFound(string) - | DeleteNotFound(string) - | ConstraintViolation({field: string, constraint: string, value: string}) - | WriteProofFailed({proofType: string, reason: string}) - | ReadOnlyStore(string) - -type typeError = { - kind: typeErrorKind, - hexad_id: option<string>, - modality: option<string>, - context: string, -} - -// ============================================================================ -// Runtime Errors -// ============================================================================ - -type runtimeErrorKind = - | StoreUnavailable({store_id: string, reason: string}) - | QueryTimeout({duration_ms: int, limit_ms: int}) - | DriftDetected({hexad_id: string, details: string}) - | PermissionDenied({user_id: option<string>, resource: string}) - | ResourceExhausted({resource: string, limit: string}) - | InvalidHexadId(string) - | NetworkError({endpoint: string, status: option<int>}) - | InternalError(string) - -type runtimeError = { - kind: runtimeErrorKind, - query_id: option<string>, - timestamp: Js.Date.t, - recoverable: bool, -} - -// ============================================================================ -// Modality-Specific Errors -// ============================================================================ - -type graphError = - | MalformedRDF(string) - | InvalidTriplePattern(string) - | CycleDetected(array<string>) - | PredicateNotFound(string) - | TraversalDepthExceeded(int) - -type vectorError = - | DimensionMismatch({expected: int, found: int}) - | InvalidDistanceMetric(string) - | EmbeddingNotFound(string) - | ANNIndexUnavailable(string) - -type tensorError = - | ShapeMismatch({expected: array<int>, found: array<int>}) - | NumericOverflow(string) - | InvalidOperation(string) - | UnsupportedDtype(string) - -type semanticError = - | InvalidContract(string) - | ZKPVerificationFailed(string) - | WitnessGenerationFailed(string) - | ContractExpired(string) - -type documentError = - | InvalidFullTextQuery(string) - | UnsupportedLanguage(string) - | IndexCorrupted(string) - -type temporalError = - | InvalidTimestamp(string) - | VersionNotFound({hexad_id: string, timestamp: string}) - | MerkleVerificationFailed(string) - | TemporalConflict(string) - -type provenanceError = - | ChainCorrupted({hexad_id: string, broken_at: int}) - | ChainNotFound(string) - | InvalidProvenanceEvent(string) - -type spatialError = - | InvalidCoordinates({latitude: float, longitude: float}) - | InvalidBounds(string) - | SpatialIndexError(string) - -type modalityError = - | GraphError(graphError) - | VectorError(vectorError) - | TensorError(tensorError) - | SemanticError(semanticError) - | DocumentError(documentError) - | TemporalError(temporalError) - | ProvenanceError(provenanceError) - | SpatialError(spatialError) - -// ============================================================================ -// Federation Errors -// ============================================================================ - -type federationErrorKind = - | RemoteStoreUnreachable({endpoint: string, timeout_ms: int}) - | PartialResults({succeeded: array<string>, failed: array<string>}) - | CrossOrgAccessDenied({org_id: string, resource: string}) - | ByzantineFaultDetected({suspicious_nodes: array<string>}) - | ConsensusTimeout({participants: int, duration_ms: int}) - | FederationPolicyViolation(string) - -type federationError = { - kind: federationErrorKind, - federation_pattern: string, - affected_stores: array<string>, -} - -// ============================================================================ -// Composite Error Type -// ============================================================================ - -type vclError = - | ParseError(parseError) - | TypeError(typeError) - | RuntimeError(runtimeError) - | ModalityError(modalityError) - | FederationError(federationError) - | MultipleErrors(array<vclError>) - -// ============================================================================ -// Error Formatting -// ============================================================================ - -let formatPosition = (pos: position): string => { - `${pos.line->Int.toString}:${pos.column->Int.toString}` -} - -let formatSpan = (span: span): string => { - `${formatPosition(span.start)}-${formatPosition(span.end_)}` -} - -let formatParseError = (err: parseError): string => { - let kindStr = switch err.kind { - | UnexpectedToken({expected, found}) => - `Expected ${expected->Array.joinWith(", ", x => `'${x}'`)}, found '${found}'` - | UnterminatedString => "Unterminated string literal" - | InvalidNumber(num) => `Invalid number: '${num}'` - | InvalidModality(mod) => `Invalid modality: '${mod}'. Valid: GRAPH, VECTOR, TENSOR, SEMANTIC, DOCUMENT, TEMPORAL` - | InvalidDriftPolicy(policy) => `Invalid drift policy: '${policy}'. Valid: STRICT, REPAIR, TOLERATE, LATEST` - | InvalidProofType(proof) => `Invalid proof type: '${proof}'. Valid: EXISTENCE, CITATION, ACCESS, INTEGRITY, PROVENANCE` - | MissingFromClause => "Missing FROM clause" - | MissingSelectClause => "Missing SELECT clause" - | InvalidGraphPattern(pattern) => `Invalid graph pattern: '${pattern}'` - | InvalidVectorExpression(expr) => `Invalid vector expression: '${expr}'` - | InvalidSemanticContract(contract) => `Invalid semantic contract: '${contract}'` - | InvalidAggregateExpression(expr) => `Invalid aggregate expression: '${expr}'. Valid: COUNT(*), SUM(M.field), AVG(M.field), MIN(M.field), MAX(M.field)` - | InvalidOrderByField(field) => `Invalid ORDER BY field: '${field}'. Use MODALITY.field format` - | InvalidGroupByField(field) => `Invalid GROUP BY field: '${field}'. Use MODALITY.field format` - | HavingWithoutGroupBy => "HAVING clause requires GROUP BY" - | AggregateWithoutGroupBy(func) => `Aggregate function ${func} used without GROUP BY clause` - } - - let hintStr = switch err.hint { - | Some(hint) => `\n Hint: ${hint}` - | None => "" - } - - `Parse Error at ${formatSpan(err.span)}: ${kindStr}${hintStr}` -} - -let formatTypeError = (err: typeError): string => { - let kindStr = switch err.kind { - | ContractNotFound(contract) => `Contract not found: '${contract}'` - | ContractViolation({contract, reason}) => `Contract '${contract}' violated: ${reason}` - | ProofGenerationFailed({contract, error}) => `Failed to generate proof for '${contract}': ${error}` - | ProofVerificationFailed({contract, reason}) => `Proof verification failed for '${contract}': ${reason}` - | TypeMismatch({expected, found}) => `Type mismatch: expected ${expected}, found ${found}` - | MissingTypeAnnotation(field) => `Missing type annotation for field: '${field}'` - | CircularDependency(cycle) => `Circular dependency detected: ${cycle->Array.joinWith(" → ", x => x)}` - // Phase 1 - | SubtypingFailed({expected, got}) => - `Subtyping failed: expected ${expected}, got ${got}` - | FieldTypeMismatch({field, expected, got}) => - `Field '${field}' type mismatch: expected ${expected}, got ${got}` - | OperatorTypeMismatch({op, leftType, rightType}) => - `Operator '${op}' cannot compare ${leftType} with ${rightType}` - | VectorDimensionMismatch({expected, got}) => - `Vector dimension mismatch: expected ${expected->Int.toString}, got ${got->Int.toString}` - | ProofObligationFailed({proofType, reason}) => - `Proof obligation '${proofType}' failed: ${reason}` - | AggregateTypeMismatch({func, fieldType}) => - `Aggregate function ${func} cannot operate on ${fieldType}` - | MultiProofConflict({proof1, proof2, reason}) => - `Proofs '${proof1}' and '${proof2}' conflict: ${reason}` - | UnknownField({modality, fieldName}) => - `Unknown field '${fieldName}' for modality ${modality}` - // Phase 2 - | CrossModalTypeMismatch({mod1, field1, mod2, field2, reason}) => - `Cross-modal type mismatch: ${mod1}.${field1} vs ${mod2}.${field2}: ${reason}` - | DriftRequiresNumeric({mod1, mod2}) => - `DRIFT requires numeric/vector modalities: ${mod1}, ${mod2}` - | ConsistencyMetricInvalid({mod1, mod2, metric}) => - `Metric '${metric}' not supported for ${mod1} and ${mod2}` - // Phase 3 - | InsertConflict(hexadId) => `INSERT conflict: hexad '${hexadId}' already exists` - | UpdateNotFound(hexadId) => `UPDATE failed: hexad '${hexadId}' not found` - | DeleteNotFound(hexadId) => `DELETE failed: hexad '${hexadId}' not found` - | ConstraintViolation({field, constraint, value}) => - `Constraint violation on '${field}': ${constraint} (value: ${value})` - | WriteProofFailed({proofType, reason}) => - `Write proof '${proofType}' failed: ${reason}` - | ReadOnlyStore(storeId) => `Store '${storeId}' is read-only` - } - - let contextStr = switch (err.hexad_id, err.modality) { - | (Some(hexad), Some(mod)) => ` [hexad: ${hexad}, modality: ${mod}]` - | (Some(hexad), None) => ` [hexad: ${hexad}]` - | (None, Some(mod)) => ` [modality: ${mod}]` - | (None, None) => "" - } - - `Type Error${contextStr}: ${kindStr}\n Context: ${err.context}` -} - -let formatRuntimeError = (err: runtimeError): string => { - let kindStr = switch err.kind { - | StoreUnavailable({store_id, reason}) => `Store '${store_id}' unavailable: ${reason}` - | QueryTimeout({duration_ms, limit_ms}) => `Query timeout: exceeded ${limit_ms}ms (ran for ${duration_ms}ms)` - | DriftDetected({hexad_id, details}) => `Drift detected for hexad '${hexad_id}': ${details}` - | PermissionDenied({user_id, resource}) => { - let user = switch user_id { - | Some(id) => `user '${id}'` - | None => "user" - } - `Permission denied: ${user} cannot access '${resource}'` - } - | ResourceExhausted({resource, limit}) => `Resource exhausted: ${resource} (limit: ${limit})` - | InvalidHexadId(id) => `Invalid hexad ID: '${id}'` - | NetworkError({endpoint, status}) => { - let statusStr = switch status { - | Some(code) => ` (HTTP ${code->Int.toString})` - | None => "" - } - `Network error connecting to '${endpoint}'${statusStr}` - } - | InternalError(msg) => `Internal error: ${msg}` - } - - let recoverable = if err.recoverable { - " [recoverable]" - } else { - " [non-recoverable]" - } - - `Runtime Error${recoverable}: ${kindStr}` -} - -let formatModalityError = (err: modalityError): string => { - switch err { - | GraphError(ge) => - switch ge { - | MalformedRDF(msg) => `Graph Error: Malformed RDF: ${msg}` - | InvalidTriplePattern(pattern) => `Graph Error: Invalid triple pattern: '${pattern}'` - | CycleDetected(path) => `Graph Error: Cycle detected: ${path->Array.joinWith(" → ", x => x)}` - | PredicateNotFound(pred) => `Graph Error: Predicate not found: '${pred}'` - | TraversalDepthExceeded(depth) => `Graph Error: Traversal depth exceeded: ${depth->Int.toString}` - } - | VectorError(ve) => - switch ve { - | DimensionMismatch({expected, found}) => `Vector Error: Dimension mismatch: expected ${expected->Int.toString}, found ${found->Int.toString}` - | InvalidDistanceMetric(metric) => `Vector Error: Invalid distance metric: '${metric}'` - | EmbeddingNotFound(id) => `Vector Error: Embedding not found: '${id}'` - | ANNIndexUnavailable(reason) => `Vector Error: ANN index unavailable: ${reason}` - } - | TensorError(te) => - switch te { - | ShapeMismatch({expected, found}) => { - let expStr = expected->Array.map(Int.toString)->Array.joinWith("×", x => x) - let foundStr = found->Array.map(Int.toString)->Array.joinWith("×", x => x) - `Tensor Error: Shape mismatch: expected [${expStr}], found [${foundStr}]` - } - | NumericOverflow(msg) => `Tensor Error: Numeric overflow: ${msg}` - | InvalidOperation(op) => `Tensor Error: Invalid operation: '${op}'` - | UnsupportedDtype(dtype) => `Tensor Error: Unsupported dtype: '${dtype}'` - } - | SemanticError(se) => - switch se { - | InvalidContract(contract) => `Semantic Error: Invalid contract: '${contract}'` - | ZKPVerificationFailed(reason) => `Semantic Error: ZKP verification failed: ${reason}` - | WitnessGenerationFailed(reason) => `Semantic Error: Witness generation failed: ${reason}` - | ContractExpired(contract) => `Semantic Error: Contract expired: '${contract}'` - } - | DocumentError(de) => - switch de { - | InvalidFullTextQuery(query) => `Document Error: Invalid full-text query: '${query}'` - | UnsupportedLanguage(lang) => `Document Error: Unsupported language: '${lang}'` - | IndexCorrupted(index) => `Document Error: Index corrupted: '${index}'` - } - | TemporalError(te) => - switch te { - | InvalidTimestamp(ts) => `Temporal Error: Invalid timestamp: '${ts}'` - | VersionNotFound({hexad_id, timestamp}) => `Temporal Error: Version not found for hexad '${hexad_id}' at '${timestamp}'` - | MerkleVerificationFailed(reason) => `Temporal Error: Merkle verification failed: ${reason}` - | TemporalConflict(msg) => `Temporal Error: Temporal conflict: ${msg}` - } - } -} - -let formatFederationError = (err: federationError): string => { - let kindStr = switch err.kind { - | RemoteStoreUnreachable({endpoint, timeout_ms}) => `Remote store unreachable: '${endpoint}' (timeout: ${timeout_ms->Int.toString}ms)` - | PartialResults({succeeded, failed}) => { - let succStr = succeeded->Array.joinWith(", ", x => x) - let failStr = failed->Array.joinWith(", ", x => x) - `Partial results: succeeded=[${succStr}], failed=[${failStr}]` - } - | CrossOrgAccessDenied({org_id, resource}) => `Cross-org access denied: org '${org_id}' cannot access '${resource}'` - | ByzantineFaultDetected({suspicious_nodes}) => `Byzantine fault detected: suspicious nodes=[${suspicious_nodes->Array.joinWith(", ", x => x)}]` - | ConsensusTimeout({participants, duration_ms}) => `Consensus timeout: ${participants->Int.toString} participants, ${duration_ms->Int.toString}ms` - | FederationPolicyViolation(msg) => `Federation policy violation: ${msg}` - } - - `Federation Error [${err.federation_pattern}]: ${kindStr}` -} - -let format = (err: vclError): string => { - switch err { - | ParseError(e) => formatParseError(e) - | TypeError(e) => formatTypeError(e) - | RuntimeError(e) => formatRuntimeError(e) - | ModalityError(e) => formatModalityError(e) - | FederationError(e) => formatFederationError(e) - | MultipleErrors(errors) => { - let header = `Multiple Errors (${errors->Array.length->Int.toString}):` - let formatted = errors->Array.mapWithIndex((err, idx) => { - ` ${(idx + 1)->Int.toString}. ${format(err)}` - }) - [header]->Array.concat(formatted)->Array.joinWith("\n", x => x) - } - } -} - -// ============================================================================ -// Error Helpers -// ============================================================================ - -let isRecoverable = (err: vclError): bool => { - switch err { - | RuntimeError(e) => e.recoverable - | FederationError({kind: PartialResults(_)}) => true - | FederationError({kind: RemoteStoreUnreachable(_)}) => true - | _ => false - } -} - -let getErrorCode = (err: vclError): string => { - switch err { - | ParseError(_) => "VCL_PARSE_ERROR" - | TypeError(_) => "VCL_TYPE_ERROR" - | RuntimeError({kind: StoreUnavailable(_)}) => "VCL_STORE_UNAVAILABLE" - | RuntimeError({kind: QueryTimeout(_)}) => "VCL_QUERY_TIMEOUT" - | RuntimeError({kind: DriftDetected(_)}) => "VCL_DRIFT_DETECTED" - | RuntimeError({kind: PermissionDenied(_)}) => "VCL_PERMISSION_DENIED" - | RuntimeError({kind: ResourceExhausted(_)}) => "VCL_RESOURCE_EXHAUSTED" - | RuntimeError(_) => "VCL_RUNTIME_ERROR" - | ModalityError(GraphError(_)) => "VCL_GRAPH_ERROR" - | ModalityError(VectorError(_)) => "VCL_VECTOR_ERROR" - | ModalityError(TensorError(_)) => "VCL_TENSOR_ERROR" - | ModalityError(SemanticError(_)) => "VCL_SEMANTIC_ERROR" - | ModalityError(DocumentError(_)) => "VCL_DOCUMENT_ERROR" - | ModalityError(TemporalError(_)) => "VCL_TEMPORAL_ERROR" - | FederationError(_) => "VCL_FEDERATION_ERROR" - | MultipleErrors(_) => "VCL_MULTIPLE_ERRORS" - } -} - -let toJson = (err: vclError): Js.Json.t => { - Js.Dict.fromArray([ - ("error_code", Js.Json.string(getErrorCode(err))), - ("message", Js.Json.string(format(err))), - ("recoverable", Js.Json.boolean(isRecoverable(err))), - ])->Js.Json.object_ -} diff --git a/verisimdb/src/vcl/VCLExplain.res b/verisimdb/src/vcl/VCLExplain.res deleted file mode 100644 index 8414538b..00000000 --- a/verisimdb/src/vcl/VCLExplain.res +++ /dev/null @@ -1,436 +0,0 @@ -// SPDX-License-Identifier: MPL-2.0 -// VCL EXPLAIN - Query Plan Visualization - -module AST = VCLParser.AST - -type planNode = { - step: int, - operation: string, - modality: string, - estimatedCost: int, - estimatedSelectivity: float, - optimizationHint: option<string>, - pushedPredicates: array<string>, -} - -type proofPlanNode = { - proofType: string, - contractName: string, - circuit: string, - estimatedTimeMs: int, -} - -type executionPlan = { - strategy: [#Sequential | #Parallel], - totalCost: int, - optimizationMode: string, - nodes: array<planNode>, - bidirectionalOptimization: bool, - proofObligations: array<proofPlanNode>, -} - -// Parse EXPLAIN query -let parseExplain = (query: string): Result<(bool, string), string> => { - let trimmed = Js.String2.trim(query) - if Js.String2.startsWith(trimmed, "EXPLAIN") { - let queryWithoutExplain = Js.String2.sliceToEnd(trimmed, ~from=7) |> Js.String2.trim - Ok((true, queryWithoutExplain)) - } else { - Ok((false, query)) - } -} - -// Format execution plan for display -let formatPlan = (plan: executionPlan): string => { - let lines = [] - - // Header - lines->Js.Array2.push("╔════════════════════════════════════════════════════════════════╗") - lines->Js.Array2.push("║ VCL QUERY EXECUTION PLAN ║") - lines->Js.Array2.push("╚════════════════════════════════════════════════════════════════╝") - lines->Js.Array2.push("") - - // Strategy - let strategyStr = switch plan.strategy { - | #Sequential => "Sequential Pipeline (operations run in series)" - | #Parallel => "Parallel Execution (operations run concurrently)" - } - lines->Js.Array2.push(`Strategy: ${strategyStr}`) - lines->Js.Array2.push(`Optimization Mode: ${plan.optimizationMode}`) - lines->Js.Array2.push(`Bidirectional Optimization: ${plan.bidirectionalOptimization ? "Enabled" : "Disabled"}`) - lines->Js.Array2.push(`Estimated Total Cost: ${Belt.Int.toString(plan.totalCost)}ms`) - lines->Js.Array2.push("") - lines->Js.Array2.push("─────────────────────────────────────────────────────────────────") - lines->Js.Array2.push("") - - // Steps - plan.nodes->Js.Array2.forEach(node => { - lines->Js.Array2.push(`Step ${Belt.Int.toString(node.step)}: ${node.operation} (${node.modality})`) - lines->Js.Array2.push(` Cost: ${Belt.Int.toString(node.estimatedCost)}ms`) - lines->Js.Array2.push(` Selectivity: ${Belt.Float.toString(node.estimatedSelectivity *. 100.0)}% of data`) - - // Optimization hints - switch node.optimizationHint { - | Some(hint) => lines->Js.Array2.push(` Optimization: ${hint}`) - | None => () - } - - // Pushed predicates - if Js.Array2.length(node.pushedPredicates) > 0 { - lines->Js.Array2.push(` Pushed predicates:`) - node.pushedPredicates->Js.Array2.forEach(pred => { - lines->Js.Array2.push(` - ${pred}`) - }) - } - - lines->Js.Array2.push("") - }) - - lines->Js.Array2.push("─────────────────────────────────────────────────────────────────") - lines->Js.Array2.push("") - - // Cost breakdown - let costByModality = plan.nodes->Belt.Array.reduce(Js.Dict.empty(), (acc, node) => { - let current = Js.Dict.get(acc, node.modality)->Belt.Option.getWithDefault(0) - Js.Dict.set(acc, node.modality, current + node.estimatedCost) - acc - }) - - lines->Js.Array2.push("Cost Breakdown by Modality:") - costByModality - ->Js.Dict.entries - ->Js.Array2.forEach(((modality, cost)) => { - let percentage = Belt.Float.fromInt(cost) /. Belt.Float.fromInt(plan.totalCost) *. 100.0 - lines->Js.Array2.push(` ${modality}: ${Belt.Int.toString(cost)}ms (${Belt.Float.toString(percentage)}%)`) - }) - - lines->Js.Array2.push("") - - // Proof obligations - if Js.Array2.length(plan.proofObligations) > 0 { - lines->Js.Array2.push("Proof Obligations:") - plan.proofObligations->Js.Array2.forEach(proof => { - lines->Js.Array2.push(` ${proof.proofType}(${proof.contractName}) circuit=${proof.circuit} est=${Belt.Int.toString(proof.estimatedTimeMs)}ms`) - }) - lines->Js.Array2.push("") - } - - // Performance hints - lines->Js.Array2.push("Performance Hints:") - let hints = generatePerformanceHints(plan) - if Js.Array2.length(hints) == 0 { - lines->Js.Array2.push(" ✓ Query plan is optimal") - } else { - hints->Js.Array2.forEach(hint => { - lines->Js.Array2.push(` • ${hint}`) - }) - } - - lines->Js.Array2.joinWith("\n") -} - -// Generate performance improvement hints -let generatePerformanceHints = (plan: executionPlan): array<string> => { - let hints = [] - - // Hint 1: Sequential with low selectivity first step - switch plan.strategy { - | #Sequential => { - switch plan.nodes[0] { - | Some(firstNode) => - if firstNode.estimatedSelectivity > 0.1 { - hints->Js.Array2.push( - "First step has low selectivity (>10%). Consider reordering or using more selective conditions." - ) - } - | None => () - } - } - | #Parallel => () - } - - // Hint 2: Expensive operation without index - plan.nodes->Js.Array2.forEach(node => { - if node.estimatedCost > 200 && node.optimizationHint == None { - hints->Js.Array2.push( - `${node.modality} operation is expensive (${Belt.Int.toString(node.estimatedCost)}ms) and not using indexes. Consider adding predicates.` - ) - } - }) - - // Hint 3: Parallel execution opportunity - if plan.strategy == #Sequential && Js.Array2.length(plan.nodes) > 2 { - let firstSelectivity = plan.nodes[0]->Belt.Option.map(n => n.estimatedSelectivity)->Belt.Option.getWithDefault(0.0) - if firstSelectivity > 0.2 { - hints->Js.Array2.push( - "Query might benefit from parallel execution. First step is not highly selective." - ) - } - } - - // Hint 4: Missing LIMIT - if plan.totalCost > 500 { - hints->Js.Array2.push( - "Query is expensive. Consider adding LIMIT clause to reduce result size." - ) - } - - hints -} - -// Example usage in client -let explainQuery = (query: string): Result<string, string> => { - switch parseExplain(query) { - | Ok((true, actualQuery)) => { - // Parse the query - switch VCLParser.parse(actualQuery) { - | Ok(ast) => { - // Generate plan (this would call Elixir QueryPlanner) - let plan = generatePlanFromAst(ast) - Ok(formatPlan(plan)) - } - | Error(e) => Error(`Parse error: ${e.message}`) - } - } - | Ok((false, _)) => Error("Not an EXPLAIN query") - | Error(msg) => Error(msg) - } -} - -// Generate plan based on actual AST analysis (replaces hardcoded mock) -let generatePlanFromAst = (ast: VCLParser.query): executionPlan => { - let nodes = ast.modalities->Belt.Array.mapWithIndex((idx, modality) => { - let modalityStr = switch modality { - | Graph => "GRAPH" - | Vector => "VECTOR" - | Tensor => "TENSOR" - | Semantic => "SEMANTIC" - | Document => "DOCUMENT" - | Temporal => "TEMPORAL" - | Provenance => "PROVENANCE" - | Spatial => "SPATIAL" - | All => "ALL" - } - - // Estimate costs based on modality type - let (cost, selectivity, hint) = switch modality { - | Graph => (150, 0.2, Some("Graph traversal — O(E) scan")) - | Vector => (50, 0.01, Some("HNSW approximate nearest neighbor")) - | Tensor => (200, 0.5, Some("Tensor reduction — shape dependent")) - | Semantic => (300, 0.8, Some("ZKP verification — expensive")) - | Document => (80, 0.05, Some("Tantivy inverted index lookup")) - | Temporal => (30, 0.1, Some("Version tree lookup — cached")) - | Provenance => (60, 0.3, Some("Hash-chain traversal — O(n) chain length")) - | Spatial => (70, 0.1, Some("R-tree spatial index lookup")) - | All => (500, 1.0, Some("Full hexad scan across all modalities")) - } - - // Adjust for LIMIT clause - let adjustedSelectivity = switch ast.limit { - | Some(limit) => - let limitF = Belt.Float.fromInt(limit) - Js.Math.min_float(selectivity, limitF /. 1000.0) - | None => selectivity - } - - { - step: idx + 1, - operation: "Query", - modality: modalityStr, - estimatedCost: cost, - estimatedSelectivity: adjustedSelectivity, - optimizationHint: hint, - pushedPredicates: [], - } - }) - - // Add GROUP BY / Aggregate step if present - let aggregateNode = switch (ast.groupBy, ast.aggregates) { - | (Some(groupFields), Some(_aggs)) => { - let groupFieldStrs = groupFields->Belt.Array.map(f => { - let modStr = switch f.modality { - | Graph => "GRAPH" - | Vector => "VECTOR" - | Tensor => "TENSOR" - | Semantic => "SEMANTIC" - | Document => "DOCUMENT" - | Temporal => "TEMPORAL" - | Provenance => "PROVENANCE" - | Spatial => "SPATIAL" - | All => "ALL" - } - `${modStr}.${f.field}` - }) - Some({ - step: Js.Array2.length(nodes) + 1, - operation: "Group & Aggregate", - modality: "AGGREGATE", - estimatedCost: 20, - estimatedSelectivity: 0.3, - optimizationHint: Some(`Group by: ${groupFieldStrs->Js.Array2.joinWith(", ")}`), - pushedPredicates: [], - }) - } - | (None, Some(_aggs)) => - Some({ - step: Js.Array2.length(nodes) + 1, - operation: "Aggregate (no grouping)", - modality: "AGGREGATE", - estimatedCost: 10, - estimatedSelectivity: 1.0, - optimizationHint: Some("Full-result aggregation — single output row"), - pushedPredicates: [], - }) - | _ => None - } - - switch aggregateNode { - | Some(node) => nodes->Js.Array2.push(node)->ignore - | None => () - } - - // Add ORDER BY / Sort step if present - switch ast.orderBy { - | Some(orderItems) => { - let orderStrs = orderItems->Belt.Array.map(item => { - let modStr = switch item.field.modality { - | Graph => "GRAPH" - | Vector => "VECTOR" - | Tensor => "TENSOR" - | Semantic => "SEMANTIC" - | Document => "DOCUMENT" - | Temporal => "TEMPORAL" - | Provenance => "PROVENANCE" - | Spatial => "SPATIAL" - | All => "ALL" - } - let dirStr = switch item.direction { - | Asc => "ASC" - | Desc => "DESC" - } - `${modStr}.${item.field.field} ${dirStr}` - }) - nodes->Js.Array2.push({ - step: Js.Array2.length(nodes) + 1, - operation: "Sort", - modality: "SORT", - estimatedCost: 15, - estimatedSelectivity: 1.0, - optimizationHint: Some(`Order by: ${orderStrs->Js.Array2.joinWith(", ")}`), - pushedPredicates: [], - })->ignore - } - | None => () - } - - // Determine strategy - let strategy = if Js.Array2.length(nodes) > 1 { - #Parallel - } else { - #Sequential - } - - let totalCost = nodes->Belt.Array.reduce(0, (acc, node) => acc + node.estimatedCost) - - // Generate proof obligation nodes from PROOF clause - let proofNodes = switch ast.proof { - | None => [] - | Some(proofSpecs) => - proofSpecs->Belt.Array.map(spec => { - let typeStr = switch spec.proofType { - | Existence => "EXISTENCE" - | Citation => "CITATION" - | Access => "ACCESS" - | Integrity => "INTEGRITY" - | Provenance => "PROVENANCE" - | Custom => "CUSTOM" - } - let circuit = switch spec.proofType { - | Existence => "existence-proof-v1" - | Citation => "citation-proof-v1" - | Access => "access-control-v1" - | Integrity => "integrity-check-v1" - | Provenance => "provenance-chain-v1" - | Custom => "custom-circuit" - } - let est = switch spec.proofType { - | Existence => 50 - | Citation => 100 - | Access => 150 - | Integrity => 200 - | Provenance => 300 - | Custom => 500 - } - { - proofType: typeStr, - contractName: spec.contractName, - circuit: circuit, - estimatedTimeMs: est, - } - }) - } - - let proofCost = proofNodes->Belt.Array.reduce(0, (acc, p) => acc + p.estimatedTimeMs) - - { - strategy: strategy, - totalCost: totalCost + proofCost, - optimizationMode: "Balanced (client-side estimate)", - nodes: nodes, - bidirectionalOptimization: false, - proofObligations: proofNodes, - } -} - -// Deprecated: Use generatePlanFromAst instead. Kept for test compatibility only. -let generateMockPlan = (ast: VCLParser.query): executionPlan => { - let nodes = ast.modalities->Belt.Array.mapWithIndex((idx, modality) => { - let modalityStr = switch modality { - | Graph => "GRAPH" - | Vector => "VECTOR" - | Tensor => "TENSOR" - | Semantic => "SEMANTIC" - | Document => "DOCUMENT" - | Temporal => "TEMPORAL" - | Provenance => "PROVENANCE" - | Spatial => "SPATIAL" - | All => "ALL" - } - - { - step: idx + 1, - operation: "Query", - modality: modalityStr, - estimatedCost: 100, - estimatedSelectivity: 0.05, - optimizationHint: Some("Using index"), - pushedPredicates: ["LIMIT 10"], - } - }) - - { - strategy: #Sequential, - totalCost: 300, - optimizationMode: "Balanced", - nodes: nodes, - bidirectionalOptimization: true, - proofObligations: [], - } -} - -// Export for testing -let testExplain = () => { - let query = ` - EXPLAIN - SELECT GRAPH, VECTOR - FROM FEDERATION /universities/* - WHERE (h)-[:CITES]->(target) - AND h.embedding SIMILAR TO [0.1, 0.2, 0.3] WITHIN 0.9 - LIMIT 10 - ` - - switch explainQuery(query) { - | Ok(plan) => Js.Console.log(plan) - | Error(e) => Js.Console.error(e) - } -} diff --git a/verisimdb/src/vcl/VCLParser.res b/verisimdb/src/vcl/VCLParser.res deleted file mode 100644 index 94be0d29..00000000 --- a/verisimdb/src/vcl/VCLParser.res +++ /dev/null @@ -1,1195 +0,0 @@ -// SPDX-License-Identifier: MPL-2.0 -// VCL Slipstream Parser - Untyped AST -// Phase 1: Simple parser for slipstream queries (no dependent types) - -// ============================================================================ -// AST Types -// ============================================================================ - -module AST = { - type modality = - | Graph - | Vector - | Tensor - | Semantic - | Document - | Temporal - | Provenance - | Spatial - | All - - type source = - | Hexad(string) // UUID - | Federation(string, option<driftPolicy>) // pattern, drift policy - | Store(string) // store ID - | Reflect // Meta-circular: query the query store itself - - and driftPolicy = - | Strict - | Repair - | Tolerate - | Latest - - type operator = - | Eq - | Neq - | Gt - | Lt - | Gte - | Lte - | Like - | Contains - | Matches - - type condition = - | Simple(simpleCondition) - | And(condition, condition) - | Or(condition, condition) - | Not(condition) - - and simpleCondition = - | FulltextContains(string) - | FulltextMatches(string) - | FieldCondition(string, operator, literal) - | VectorSimilar(array<float>, option<float>) // embedding, threshold - | GraphPattern(string) // SPARQL-like pattern (simplified) - // Phase 2: Cross-modal conditions - | CrossModalFieldCompare(modality, string, operator, modality, string) - // e.g., WHERE DOCUMENT.severity > GRAPH.centrality - | ModalityDrift(modality, modality, float) - // e.g., WHERE DRIFT(VECTOR, DOCUMENT) > 0.3 - | ModalityExists(modality) - // e.g., WHERE VECTOR EXISTS - | ModalityNotExists(modality) - // e.g., WHERE TENSOR NOT EXISTS - | ModalityConsistency(modality, modality, string) - // e.g., WHERE CONSISTENT(VECTOR, SEMANTIC) USING COSINE - - and literal = - | String(string) - | Int(int) - | Float(float) - | Bool(bool) - | Array(array<literal>) - - // Field reference: DOCUMENT.name, GRAPH.predicate, etc. - type fieldRef = { - modality: modality, - field: string, - } - - // Aggregate functions (SQL-compatible) - type aggregateFunc = - | Count - | Sum - | Avg - | Min - | Max - - // Aggregate expression in SELECT - type aggregateExpr = - | CountAll // COUNT(*) - | AggregateField(aggregateFunc, fieldRef) // AVG(DOCUMENT.severity) - - // Sort direction for ORDER BY - type sortDirection = - | Asc - | Desc - - // ORDER BY item - type orderByItem = { - field: fieldRef, - direction: sortDirection, - } - - type query = { - modalities: array<modality>, - projections: option<array<fieldRef>>, // Column selection: DOCUMENT.name, DOCUMENT.severity - aggregates: option<array<aggregateExpr>>, // COUNT(*), SUM(DOCUMENT.severity) - source: source, - where: option<condition>, - groupBy: option<array<fieldRef>>, // GROUP BY DOCUMENT.name - having: option<condition>, // HAVING COUNT(*) > 5 - proof: option<array<proofSpec>>, - orderBy: option<array<orderByItem>>, // ORDER BY DOCUMENT.severity DESC - limit: option<int>, - offset: option<int>, - } - - and proofSpec = { - proofType: proofType, - contractName: string, - customParams: option<array<(string, string)>>, // WITH (key=value, ...) for Custom proofs - } - - and proofType = - | Existence - | Citation - | Access - | Integrity - | Provenance - | Custom - - // Phase 3: Mutation types (INSERT / UPDATE / DELETE) - type modalityData = - | DocumentData(array<(string, literal)>) // field-value pairs - | VectorData(array<float>) // embedding - | GraphData(string, string) // edge_type, target_hexad_id - | TensorData(array<literal>) // tensor values - | SemanticData(string) // contract name - | TemporalData(string) // timestamp - | ProvenanceData(array<(string, literal)>) // event_type, actor, description, source - | SpatialData(array<(string, literal)>) // latitude, longitude, altitude, geometry_type - - type mutation = - | Insert({ - modalities: array<modalityData>, - proof: option<array<proofSpec>>, - }) - | Update({ - hexadId: string, - sets: array<(fieldRef, literal)>, - proof: option<array<proofSpec>>, - }) - | Delete({ - hexadId: string, - proof: option<array<proofSpec>>, - }) - - type statement = - | Query(query) - | Mutation(mutation) -} - -// ============================================================================ -// Parser Combinators -// ============================================================================ - -module Parser = { - type parseError = { - message: string, - position: int, - } - - type parseResult<'a> = Result<('a, int), parseError> - - type parser<'a> = string => parseResult<'a> - - // Basic combinators - let pure = (value: 'a): parser<'a> => { - input => Ok((value, 0)) - } - - let fail = (message: string): parser<'a> => { - _input => Error({message, position: 0}) - } - - let map = (p: parser<'a>, f: 'a => 'b): parser<'b> => { - input => { - switch p(input) { - | Ok((value, consumed)) => Ok((f(value), consumed)) - | Error(e) => Error(e) - } - } - } - - let bind = (p: parser<'a>, f: 'a => parser<'b>): parser<'b> => { - input => { - switch p(input) { - | Ok((value, consumed)) => { - let remaining = Js.String2.sliceToEnd(input, ~from=consumed) - switch f(value)(remaining) { - | Ok((value2, consumed2)) => Ok((value2, consumed + consumed2)) - | Error(e) => Error({...e, position: e.position + consumed}) - } - } - | Error(e) => Error(e) - } - } - } - - let (<|>) = (p1: parser<'a>, p2: parser<'a>): parser<'a> => { - input => { - switch p1(input) { - | Ok(result) => Ok(result) - | Error(_) => p2(input) - } - } - } - - // Whitespace handling - let ws: parser<unit> = input => { - let trimmed = Js.String2.trimStart(input) - let consumed = Js.String2.length(input) - Js.String2.length(trimmed) - Ok(((), consumed)) - } - - let lexeme = (p: parser<'a>): parser<'a> => { - bind(p, value => map(ws, _ => value)) - } - - // String matching - let string = (s: string): parser<string> => { - input => { - if Js.String2.startsWith(input, s) { - Ok((s, Js.String2.length(s))) - } else { - Error({message: `Expected "${s}"`, position: 0}) - } - } - } - - let keyword = (k: string): parser<string> => { - lexeme(string(k)) - } - - // Regex-based parsers - let regex = (pattern: string): parser<string> => { - input => { - let re = Js.Re.fromStringWithFlags(pattern, ~flags="i") - switch Js.Re.exec_(re, input) { - | Some(result) => { - let matched = Js.Re.captures(result)[0] - switch Js.Nullable.toOption(matched) { - | Some(str) => Ok((str, Js.String2.length(str))) - | None => Error({message: `Regex ${pattern} failed`, position: 0}) - } - } - | None => Error({message: `Regex ${pattern} failed`, position: 0}) - } - } - } - - let identifier: parser<string> = lexeme(regex("^[a-zA-Z_][a-zA-Z0-9_]*")) - - let uuid: parser<string> = lexeme( - regex("^[0-9a-fA-F]{8}-[0-9a-fA-F]{4}-[0-9a-fA-F]{4}-[0-9a-fA-F]{4}-[0-9a-fA-F]{12}") - ) - - let integer: parser<int> = { - input => { - let intStr = lexeme(regex("^[0-9]+")) - switch intStr(input) { - | Ok((str, consumed)) => { - switch Belt.Int.fromString(str) { - | Some(n) => Ok((n, consumed)) - | None => Error({message: "Invalid integer", position: 0}) - } - } - | Error(e) => Error(e) - } - } - } - - let float: parser<float> = { - input => { - let floatStr = lexeme(regex("^[0-9]+\\.[0-9]+")) - switch floatStr(input) { - | Ok((str, consumed)) => { - switch Belt.Float.fromString(str) { - | Some(f) => Ok((f, consumed)) - | None => Error({message: "Invalid float", position: 0}) - } - } - | Error(e) => Error(e) - } - } - } - - let stringLiteral: parser<string> = { - input => { - let quoted = lexeme(regex("^\"([^\"\\\\]|\\\\.)*\"")) - switch quoted(input) { - | Ok((str, consumed)) => { - // Remove quotes - let unquoted = Js.String2.slice(str, ~from=1, ~to_=Js.String2.length(str) - 1) - Ok((unquoted, consumed)) - } - | Error(e) => Error(e) - } - } - } - - // Many combinator - let rec many = (p: parser<'a>): parser<array<'a>> => { - input => { - switch p(input) { - | Ok((value, consumed)) => { - let remaining = Js.String2.sliceToEnd(input, ~from=consumed) - switch many(p)(remaining) { - | Ok((values, consumed2)) => Ok(([value]->Js.Array2.concat(values), consumed + consumed2)) - | Error(_) => Ok(([value], consumed)) - } - } - | Error(_) => Ok(([], 0)) - } - } - } - - let sepBy = (p: parser<'a>, sep: parser<'b>): parser<array<'a>> => { - input => { - switch p(input) { - | Ok((first, consumed1)) => { - let remaining = Js.String2.sliceToEnd(input, ~from=consumed1) - let parseRest = bind(sep, _ => p) - switch many(parseRest)(remaining) { - | Ok((rest, consumed2)) => Ok(([first]->Js.Array2.concat(rest), consumed1 + consumed2)) - | Error(_) => Ok(([first], consumed1)) - } - } - | Error(e) => Error(e) - } - } - } - - let optional = (p: parser<'a>): parser<option<'a>> => { - input => { - switch p(input) { - | Ok((value, consumed)) => Ok((Some(value), consumed)) - | Error(_) => Ok((None, 0)) - } - } - } -} - -// ============================================================================ -// VCL Grammar Parsers -// ============================================================================ - -module Grammar = { - open Parser - open AST - - // Modality parser (octad: 8 modalities + All) - let modality: parser<modality> = { - let graph = map(keyword("GRAPH"), _ => Graph) - let vector = map(keyword("VECTOR"), _ => Vector) - let tensor = map(keyword("TENSOR"), _ => Tensor) - let semantic = map(keyword("SEMANTIC"), _ => Semantic) - let document = map(keyword("DOCUMENT"), _ => Document) - let temporal = map(keyword("TEMPORAL"), _ => Temporal) - let provenance = map(keyword("PROVENANCE"), _ => Provenance) - let spatial = map(keyword("SPATIAL"), _ => Spatial) - let all = map(keyword("*"), _ => All) - - graph <|> vector <|> tensor <|> semantic <|> document <|> temporal <|> provenance <|> spatial <|> all - } - - let modalityList: parser<array<modality>> = sepBy(modality, keyword(",")) - - // Field reference parser: DOCUMENT.name, GRAPH.predicate, etc. - let fieldRef: parser<AST.fieldRef> = { - bind(modality, mod => - bind(keyword("."), _ => - map(identifier, field => { - AST.modality: mod, - field: field, - }) - ) - ) - } - - // Aggregate function name parser - let aggregateFunc: parser<AST.aggregateFunc> = { - let count = map(keyword("COUNT"), _ => AST.Count) - let sum = map(keyword("SUM"), _ => AST.Sum) - let avg = map(keyword("AVG"), _ => AST.Avg) - let min_ = map(keyword("MIN"), _ => AST.Min) - let max_ = map(keyword("MAX"), _ => AST.Max) - - count <|> sum <|> avg <|> min_ <|> max_ - } - - // Aggregate expression parser: COUNT(*) or AVG(DOCUMENT.severity) - let aggregateExpr: parser<AST.aggregateExpr> = { - let countAll = { - bind(keyword("COUNT"), _ => - bind(keyword("("), _ => - bind(keyword("*"), _ => - map(keyword(")"), _ => AST.CountAll) - ) - ) - ) - } - - let aggregateField = { - bind(aggregateFunc, func => - bind(keyword("("), _ => - bind(fieldRef, ref => - map(keyword(")"), _ => AST.AggregateField(func, ref)) - ) - ) - ) - } - - countAll <|> aggregateField - } - - // Extended select item: aggregate | field projection | bare modality - type selectItem = - | SelectAggregate(AST.aggregateExpr) - | SelectField(AST.fieldRef) - | SelectModality(AST.modality) - - let selectItem: parser<selectItem> = { - let agg = map(aggregateExpr, a => SelectAggregate(a)) - let field = map(fieldRef, f => SelectField(f)) - let mod = map(modality, m => SelectModality(m)) - - agg <|> field <|> mod - } - - let selectItemList: parser<array<selectItem>> = sepBy(selectItem, keyword(",")) - - // Classify select items into modalities, projections, and aggregates - type classifiedSelect = { - modalities: array<AST.modality>, - projections: option<array<AST.fieldRef>>, - aggregates: option<array<AST.aggregateExpr>>, - } - - let classifySelect = (items: array<selectItem>): classifiedSelect => { - let mods = [] - let projs = [] - let aggs = [] - - items->Js.Array2.forEach(item => { - switch item { - | SelectModality(m) => mods->Js.Array2.push(m)->ignore - | SelectField(f) => { - projs->Js.Array2.push(f)->ignore - // Also add the modality if not already present - if !(mods->Js.Array2.some(m => m == f.modality)) { - mods->Js.Array2.push(f.modality)->ignore - } - } - | SelectAggregate(a) => { - aggs->Js.Array2.push(a)->ignore - // Add modality from aggregate field ref if present - switch a { - | AggregateField(_, ref) => - if !(mods->Js.Array2.some(m => m == ref.modality)) { - mods->Js.Array2.push(ref.modality)->ignore - } - | CountAll => () - } - } - } - }) - - { - modalities: mods, - projections: if Js.Array2.length(projs) > 0 { Some(projs) } else { None }, - aggregates: if Js.Array2.length(aggs) > 0 { Some(aggs) } else { None }, - } - } - - // SELECT clause (extended to support projections and aggregates) - let selectClause: parser<classifiedSelect> = { - map(bind(keyword("SELECT"), _ => selectItemList), items => classifySelect(items)) - } - - // Drift policy - let driftPolicy: parser<driftPolicy> = { - let strict = map(keyword("STRICT"), _ => Strict) - let repair = map(keyword("REPAIR"), _ => Repair) - let tolerate = map(keyword("TOLERATE"), _ => Tolerate) - let latest = map(keyword("LATEST"), _ => Latest) - - bind(keyword("WITH"), _ => - bind(keyword("DRIFT"), _ => - strict <|> repair <|> tolerate <|> latest - ) - ) - } - - // Source parser - let source: parser<source> = { - let hexadSource = { - bind(keyword("HEXAD"), _ => - map(uuid, id => Hexad(id)) - ) - } - - let federationSource = { - bind(keyword("FEDERATION"), _ => - bind(identifier, pattern => - map(optional(driftPolicy), drift => - Federation(pattern, drift) - ) - ) - ) - } - - let storeSource = { - bind(keyword("STORE"), _ => - map(identifier, id => Store(id)) - ) - } - - hexadSource <|> federationSource <|> storeSource - } - - // FROM clause - let fromClause: parser<source> = { - bind(keyword("FROM"), _ => source) - } - - // Operators - let operator: parser<operator> = { - let eq = map(keyword("=="), _ => Eq) - let neq = map(keyword("!="), _ => Neq) - let gte = map(keyword(">="), _ => Gte) - let lte = map(keyword("<="), _ => Lte) - let gt = map(keyword(">"), _ => Gt) - let lt = map(keyword("<"), _ => Lt) - let like = map(keyword("LIKE"), _ => Like) - let contains = map(keyword("CONTAINS"), _ => Contains) - let matches = map(keyword("MATCHES"), _ => Matches) - - eq <|> neq <|> gte <|> lte <|> gt <|> lt <|> like <|> contains <|> matches - } - - // Literals - let rec literal: parser<literal> = { - input => { - let stringLit = map(stringLiteral, s => String(s)) - let intLit = map(integer, i => Int(i)) - let floatLit = map(float, f => Float(f)) - let boolLit = { - let t = map(keyword("true"), _ => Bool(true)) - let f = map(keyword("false"), _ => Bool(false)) - t <|> f - } - - let arrayLit = { - bind(keyword("["), _ => - bind(sepBy(literal, keyword(",")), values => - map(keyword("]"), _ => Array(values)) - ) - ) - } - - let p = arrayLit <|> floatLit <|> intLit <|> stringLit <|> boolLit - p(input) - } - } - - // Simple conditions - let simpleCondition: parser<simpleCondition> = { - let fulltextContains = { - bind(keyword("FULLTEXT"), _ => - bind(keyword("CONTAINS"), _ => - map(stringLiteral, text => FulltextContains(text)) - ) - ) - } - - let fulltextMatches = { - bind(keyword("FULLTEXT"), _ => - bind(keyword("MATCHES"), _ => - map(stringLiteral, pattern => FulltextMatches(pattern)) - ) - ) - } - - let fieldCondition = { - bind(keyword("FIELD"), _ => - bind(identifier, field => - bind(operator, op => - map(literal, value => FieldCondition(field, op, value)) - ) - ) - ) - } - - let vectorSimilar = { - bind(identifier, _field => - bind(keyword("SIMILAR"), _ => - bind(keyword("TO"), _ => - bind(literal, embedding => - map(optional(bind(keyword("WITHIN"), _ => float)), threshold => { - // Extract floats from array literal - let floats = switch embedding { - | Array(arr) => arr->Js.Array2.map(lit => - switch lit { - | Float(f) => f - | Int(i) => Belt.Int.toFloat(i) - | _ => 0.0 - } - ) - | _ => [] - } - VectorSimilar(floats, threshold) - }) - ) - ) - ) - ) - } - - let graphPattern = { - // Simplified: just capture the pattern as string for now - map(stringLiteral, pattern => GraphPattern(pattern)) - } - - // Phase 2: Cross-modal conditions - let driftCondition = { - bind(keyword("DRIFT"), _ => - bind(keyword("("), _ => - bind(modality, mod1 => - bind(keyword(","), _ => - bind(modality, mod2 => - bind(keyword(")"), _ => - bind(operator, _op => - map(float, threshold => - ModalityDrift(mod1, mod2, threshold) - ) - ) - ) - ) - ) - ) - ) - ) - } - - let consistencyCondition = { - bind(keyword("CONSISTENT"), _ => - bind(keyword("("), _ => - bind(modality, mod1 => - bind(keyword(","), _ => - bind(modality, mod2 => - bind(keyword(")"), _ => - bind(keyword("USING"), _ => - map(identifier, metric => - ModalityConsistency(mod1, mod2, metric) - ) - ) - ) - ) - ) - ) - ) - ) - } - - let existsCondition = { - bind(modality, mod => - map(keyword("EXISTS"), _ => - ModalityExists(mod) - ) - ) - } - - let notExistsCondition = { - bind(modality, mod => - bind(keyword("NOT"), _ => - map(keyword("EXISTS"), _ => - ModalityNotExists(mod) - ) - ) - ) - } - - // Cross-modal field compare: MODALITY1.field op MODALITY2.field - let crossModalFieldCompare = { - bind(modality, mod1 => - bind(keyword("."), _ => - bind(identifier, field1 => - bind(operator, op => - bind(modality, mod2 => - bind(keyword("."), _ => - map(identifier, field2 => - CrossModalFieldCompare(mod1, field1, op, mod2, field2) - ) - ) - ) - ) - ) - ) - ) - } - - driftCondition <|> consistencyCondition <|> notExistsCondition <|> existsCondition <|> crossModalFieldCompare <|> fulltextContains <|> fulltextMatches <|> fieldCondition <|> vectorSimilar <|> graphPattern - } - - // Compound conditions - let rec condition: parser<condition> = { - input => { - let simple = map(simpleCondition, c => Simple(c)) - - let andCond = { - bind(condition, left => - bind(keyword("AND"), _ => - map(condition, right => And(left, right)) - ) - ) - } - - let orCond = { - bind(condition, left => - bind(keyword("OR"), _ => - map(condition, right => Or(left, right)) - ) - ) - } - - let notCond = { - bind(keyword("NOT"), _ => - map(condition, c => Not(c)) - ) - } - - let p = andCond <|> orCond <|> notCond <|> simple - p(input) - } - } - - // WHERE clause - let whereClause: parser<condition> = { - bind(keyword("WHERE"), _ => condition) - } - - // PROOF clause - let proofType: parser<proofType> = { - let existence = map(keyword("EXISTENCE"), _ => Existence) - let citation = map(keyword("CITATION"), _ => Citation) - let access = map(keyword("ACCESS"), _ => Access) - let integrity = map(keyword("INTEGRITY"), _ => Integrity) - let provenance = map(keyword("PROVENANCE"), _ => Provenance) - let custom = map(keyword("CUSTOM"), _ => Custom) - - existence <|> citation <|> access <|> integrity <|> provenance <|> custom - } - - let proofSpec: parser<proofSpec> = { - bind(proofType, pType => - bind(keyword("("), _ => - bind(identifier, contract => - map(keyword(")"), _ => { - proofType: pType, - contractName: contract, - }) - ) - ) - ) - } - - // Multi-proof: PROOF spec1 AND spec2 AND spec3 - let proofClause: parser<array<proofSpec>> = { - bind(keyword("PROOF"), _ => - sepBy(proofSpec, keyword("AND")) - ) - } - - // LIMIT clause - let limitClause: parser<int> = { - bind(keyword("LIMIT"), _ => integer) - } - - // OFFSET clause - let offsetClause: parser<int> = { - bind(keyword("OFFSET"), _ => integer) - } - - // GROUP BY clause - let groupByClause: parser<array<AST.fieldRef>> = { - bind(keyword("GROUP"), _ => - bind(keyword("BY"), _ => - sepBy(fieldRef, keyword(",")) - ) - ) - } - - // HAVING clause (reuses condition parser — conditions on aggregates) - let havingClause: parser<AST.condition> = { - bind(keyword("HAVING"), _ => condition) - } - - // Sort direction parser - let sortDirection: parser<AST.sortDirection> = { - let asc = map(keyword("ASC"), _ => AST.Asc) - let desc = map(keyword("DESC"), _ => AST.Desc) - - asc <|> desc - } - - // ORDER BY item: DOCUMENT.severity DESC | DOCUMENT.name (defaults to ASC) - let orderByItem: parser<AST.orderByItem> = { - bind(fieldRef, ref => - map(optional(sortDirection), dir => { - AST.field: ref, - direction: switch dir { - | Some(d) => d - | None => Asc - }, - }) - ) - } - - // ORDER BY clause - let orderByClause: parser<array<AST.orderByItem>> = { - bind(keyword("ORDER"), _ => - bind(keyword("BY"), _ => - sepBy(orderByItem, keyword(",")) - ) - ) - } - - // Full query parser - let query: parser<query> = { - input => { - // Parse in sequence: - // SELECT ... FROM ... [WHERE ...] [GROUP BY ...] [HAVING ...] - // [PROOF ...] [ORDER BY ...] [LIMIT ...] [OFFSET ...] - let parseQuery = { - bind(ws, _ => - bind(selectClause, classified => - bind(fromClause, src => - bind(optional(whereClause), whereCond => - bind(optional(groupByClause), groupBy => - bind(optional(havingClause), having => - bind(optional(proofClause), proof => - bind(optional(orderByClause), orderBy => - bind(optional(limitClause), lim => - map(optional(offsetClause), off => { - modalities: classified.modalities, - projections: classified.projections, - aggregates: classified.aggregates, - source: src, - where: whereCond, - groupBy: groupBy, - having: having, - proof: proof, - orderBy: orderBy, - limit: lim, - offset: off, - }) - ) - ) - ) - ) - ) - ) - ) - ) - ) - } - - parseQuery(input) - } - } -} - -// ============================================================================ -// Public API -// ============================================================================ - -type parseError = Parser.parseError -type query = AST.query - -let parse = (input: string): Result<query, parseError> => { - switch Grammar.query(input) { - | Ok((query, _consumed)) => Ok(query) - | Error(e) => Error(e) - } -} - -let parseSlipstream = (input: string): Result<query, parseError> => { - // Slipstream path: no proof clause allowed - switch parse(input) { - | Ok(query) => { - switch query.proof { - | Some(proofs) if Js.Array2.length(proofs) > 0 => - Error({message: "Slipstream queries cannot have PROOF clause", position: 0}) - | _ => Ok(query) - } - } - | Error(e) => Error(e) - } -} - -let parseDependentType = (input: string): Result<query, parseError> => { - // Dependent-type path: proof clause required - switch parse(input) { - | Ok(query) => { - if query.proof->Belt.Option.isNone { - Error({message: "Dependent-type queries require PROOF clause", position: 0}) - } else { - Ok(query) - } - } - | Error(e) => Error(e) - } -} - -// ============================================================================ -// Phase 3: Mutation Parsers -// ============================================================================ - -type mutation = AST.mutation -type modalityData = AST.modalityData -type statement = AST.statement - -module MutationParser = { - open Parser - open AST - - // Parse a single modality data entry - let documentData: parser<modalityData> = { - bind(keyword("DOCUMENT"), _ => - bind(keyword("("), _ => - bind(sepBy( - bind(Grammar.identifier, field => - bind(keyword("="), _ => - map(Grammar.literal, value => (field, value)) - ) - ), - keyword(","), - ), fields => - map(keyword(")"), _ => DocumentData(fields)) - ) - ) - ) - } - - let vectorData: parser<modalityData> = { - bind(keyword("VECTOR"), _ => - bind(keyword("("), _ => - bind(keyword("["), _ => - bind(sepBy(Grammar.float, keyword(",")), values => - bind(keyword("]"), _ => - map(keyword(")"), _ => VectorData(values)) - ) - ) - ) - ) - ) - } - - let graphData: parser<modalityData> = { - bind(keyword("GRAPH"), _ => - bind(keyword("("), _ => - bind(Grammar.identifier, edgeType => - bind(keyword(","), _ => - bind(Grammar.identifier, targetId => - map(keyword(")"), _ => GraphData(edgeType, targetId)) - ) - ) - ) - ) - ) - } - - let tensorData: parser<modalityData> = { - bind(keyword("TENSOR"), _ => - bind(keyword("("), _ => - bind(sepBy(Grammar.literal, keyword(",")), values => - map(keyword(")"), _ => TensorData(values)) - ) - ) - ) - } - - let semanticData: parser<modalityData> = { - bind(keyword("SEMANTIC"), _ => - bind(keyword("("), _ => - bind(Grammar.identifier, contractName => - map(keyword(")"), _ => SemanticData(contractName)) - ) - ) - ) - } - - let temporalData: parser<modalityData> = { - bind(keyword("TEMPORAL"), _ => - bind(keyword("("), _ => - bind(Grammar.stringLiteral, timestamp => - map(keyword(")"), _ => TemporalData(timestamp)) - ) - ) - ) - } - - // PROVENANCE(field=value, ...) - let provenanceData: parser<modalityData> = { - bind(keyword("PROVENANCE"), _ => - bind(keyword("("), _ => - bind(sepBy( - bind(Grammar.identifier, key => - bind(keyword("="), _ => - map(Grammar.literal, value => (key, value)) - ) - ), - keyword(",") - ), fields => - map(keyword(")"), _ => ProvenanceData(fields)) - ) - ) - ) - } - - // SPATIAL(field=value, ...) - let spatialData: parser<modalityData> = { - bind(keyword("SPATIAL"), _ => - bind(keyword("("), _ => - bind(sepBy( - bind(Grammar.identifier, key => - bind(keyword("="), _ => - map(Grammar.literal, value => (key, value)) - ) - ), - keyword(",") - ), fields => - map(keyword(")"), _ => SpatialData(fields)) - ) - ) - ) - } - - let modalityData: parser<modalityData> = { - documentData <|> vectorData <|> graphData <|> tensorData <|> semanticData <|> temporalData <|> provenanceData <|> spatialData - } - - // INSERT HEXAD WITH modalityData [, modalityData]* [PROOF ...] - let insertMutation: parser<mutation> = { - bind(keyword("INSERT"), _ => - bind(keyword("HEXAD"), _ => - bind(keyword("WITH"), _ => - bind(sepBy(modalityData, keyword(",")), data => - map(optional(Grammar.proofClause), proof => - Insert({ - modalities: data, - proof: proof, - }) - ) - ) - ) - ) - ) - } - - // UPDATE HEXAD uuid SET field = value [, field = value]* [PROOF ...] - let updateMutation: parser<mutation> = { - bind(keyword("UPDATE"), _ => - bind(keyword("HEXAD"), _ => - bind(uuid, id => - bind(keyword("SET"), _ => - bind(sepBy( - bind(Grammar.fieldRef, field => - bind(keyword("="), _ => - map(Grammar.literal, value => (field, value)) - ) - ), - keyword(","), - ), sets => - map(optional(Grammar.proofClause), proof => - Update({ - hexadId: id, - sets: sets, - proof: proof, - }) - ) - ) - ) - ) - ) - ) - } - - // DELETE HEXAD uuid [PROOF ...] - let deleteMutation: parser<mutation> = { - bind(keyword("DELETE"), _ => - bind(keyword("HEXAD"), _ => - bind(uuid, id => - map(optional(Grammar.proofClause), proof => - Delete({ - hexadId: id, - proof: proof, - }) - ) - ) - ) - ) - } - - let mutation: parser<mutation> = { - insertMutation <|> updateMutation <|> deleteMutation - } - - // Top-level statement: query or mutation - let statement: parser<statement> = { - input => { - let queryP = map(Grammar.query, q => Query(q)) - let mutationP = map(mutation, m => Mutation(m)) - - let p = bind(ws, _ => mutationP <|> queryP) - p(input) - } - } -} - -let parseMutation = (input: string): Result<AST.mutation, parseError> => { - switch MutationParser.mutation(input) { - | Ok((m, _consumed)) => Ok(m) - | Error(e) => Error(e) - } -} - -let parseStatement = (input: string): Result<AST.statement, parseError> => { - switch MutationParser.statement(input) { - | Ok((s, _consumed)) => Ok(s) - | Error(e) => Error(e) - } -} - -// ============================================================================ -// Example Usage -// ============================================================================ - -/* -// Slipstream query -let slipstreamQuery = ` - SELECT GRAPH, VECTOR - FROM FEDERATION /universities/* - WHERE FULLTEXT CONTAINS "machine learning" - LIMIT 100 -` - -switch parseSlipstream(slipstreamQuery) { -| Ok(query) => Js.Console.log(query) -| Error(e) => Js.Console.error(e.message) -} - -// Dependent-type query -let dependentQuery = ` - SELECT GRAPH, VECTOR - FROM HEXAD 550e8400-e29b-41d4-a716-446655440000 - WHERE h.embedding SIMILAR TO [0.1, 0.2, 0.3] WITHIN 0.9 - AND FULLTEXT CONTAINS "climate change" - PROOF CITATION(CitationContract) - LIMIT 50 -` - -switch parseDependentType(dependentQuery) { -| Ok(query) => Js.Console.log(query) -| Error(e) => Js.Console.error(e.message) -} - -// SQL-compatible query with column projections, aggregates, ORDER BY, GROUP BY -let sqlCompatQuery = ` - SELECT DOCUMENT.name, DOCUMENT.severity, COUNT(*), AVG(DOCUMENT.severity) - FROM FEDERATION /universities/* - WHERE FIELD severity > 5 - GROUP BY DOCUMENT.name, DOCUMENT.severity - HAVING FIELD count > 3 - ORDER BY DOCUMENT.severity DESC - LIMIT 50 -` - -switch parseSlipstream(sqlCompatQuery) { -| Ok(query) => Js.Console.log(query) -| Error(e) => Js.Console.error(e.message) -} -*/ diff --git a/verisimdb/src/vcl/VCLParser_test.res b/verisimdb/src/vcl/VCLParser_test.res deleted file mode 100644 index ccc5f971..00000000 --- a/verisimdb/src/vcl/VCLParser_test.res +++ /dev/null @@ -1,558 +0,0 @@ -// SPDX-License-Identifier: MPL-2.0 -// VCL Parser Tests - -open VCLParser - -// Test helper -let assertOk = (result: Result<'a, 'b>, testName: string) => { - switch result { - | Ok(_) => Js.Console.log(`✓ ${testName}`) - | Error(e) => Js.Console.error(`✗ ${testName}: ${e.message}`) - } -} - -let assertError = (result: Result<'a, 'b>, testName: string) => { - switch result { - | Ok(_) => Js.Console.error(`✗ ${testName}: Expected error but got Ok`) - | Error(_) => Js.Console.log(`✓ ${testName}`) - } -} - -// ============================================================================ -// Test Suite -// ============================================================================ - -Js.Console.log("\n=== VCL Parser Tests ===\n") - -// Test 1: Simple hexad query -let test1 = ` - SELECT * - FROM HEXAD 550e8400-e29b-41d4-a716-446655440000 -` - -assertOk(parseSlipstream(test1), "Test 1: Simple hexad query") - -// Test 2: Federation query with drift policy -let test2 = ` - SELECT GRAPH, VECTOR - FROM FEDERATION /universities/* WITH DRIFT REPAIR -` - -assertOk(parseSlipstream(test2), "Test 2: Federation with drift policy") - -// Test 3: Full-text search with LIMIT -let test3 = ` - SELECT DOCUMENT - FROM STORE tantivy-node-1 - WHERE FULLTEXT CONTAINS "machine learning" - LIMIT 100 -` - -assertOk(parseSlipstream(test3), "Test 3: Full-text search with LIMIT") - -// Test 4: Vector similarity query -let test4 = ` - SELECT VECTOR - FROM HEXAD abc12345-0000-0000-0000-000000000000 - WHERE h.embedding SIMILAR TO [0.1, 0.2, 0.3] WITHIN 0.9 -` - -assertOk(parseSlipstream(test4), "Test 4: Vector similarity query") - -// Test 5: Multiple modalities -let test5 = ` - SELECT GRAPH, VECTOR, DOCUMENT - FROM FEDERATION /research/* - LIMIT 50 - OFFSET 100 -` - -assertOk(parseSlipstream(test5), "Test 5: Multiple modalities with pagination") - -// Test 6: Dependent-type query (should have PROOF) -let test6 = ` - SELECT GRAPH - FROM HEXAD 550e8400-e29b-41d4-a716-446655440000 - WHERE FULLTEXT CONTAINS "climate change" - PROOF CITATION(CitationContract) -` - -assertOk(parseDependentType(test6), "Test 6: Dependent-type with PROOF") - -// Test 7: Slipstream with PROOF (should fail) -let test7 = ` - SELECT * - FROM HEXAD 550e8400-e29b-41d4-a716-446655440000 - PROOF EXISTENCE(ExistenceContract) -` - -assertError(parseSlipstream(test7), "Test 7: Slipstream rejects PROOF clause") - -// Test 8: Dependent-type without PROOF (should fail) -let test8 = ` - SELECT GRAPH - FROM HEXAD 550e8400-e29b-41d4-a716-446655440000 -` - -assertError(parseDependentType(test8), "Test 8: Dependent-type requires PROOF") - -// Test 9: Field condition -let test9 = ` - SELECT DOCUMENT - FROM STORE archive-1 - WHERE FIELD year >= 2020 - LIMIT 10 -` - -assertOk(parseSlipstream(test9), "Test 9: Field condition with operator") - -// Test 10: Multiple WHERE conditions (simplified - parser needs enhancement) -let test10 = ` - SELECT DOCUMENT - FROM FEDERATION /archives/* - WHERE FULLTEXT CONTAINS "quantum computing" -` - -assertOk(parseSlipstream(test10), "Test 10: WHERE with FULLTEXT") - -// Test 11: Complex dependent-type query -let test11 = ` - SELECT GRAPH, VECTOR, SEMANTIC - FROM FEDERATION /universities/* WITH DRIFT STRICT - WHERE h.embedding SIMILAR TO [0.5, 0.3, 0.2] - PROOF INTEGRITY(DataIntegrityContract) - LIMIT 100 -` - -assertOk(parseDependentType(test11), "Test 11: Complex dependent-type query") - -// Test 12: All modalities -let test12 = ` - SELECT * - FROM HEXAD 550e8400-e29b-41d4-a716-446655440000 - LIMIT 10 -` - -assertOk(parseSlipstream(test12), "Test 12: All modalities (wildcard)") - -// Test 13: Store query -let test13 = ` - SELECT VECTOR - FROM STORE milvus-us-east-1 - WHERE h.embedding SIMILAR TO [0.1, 0.2] - LIMIT 20 -` - -assertOk(parseSlipstream(test13), "Test 13: Store-specific query") - -// Test 14: PROOF with different types -let test14a = `SELECT * FROM HEXAD 550e8400-e29b-41d4-a716-446655440000 PROOF EXISTENCE(ExistenceContract)` -assertOk(parseDependentType(test14a), "Test 14a: EXISTENCE proof") - -let test14b = `SELECT * FROM HEXAD 550e8400-e29b-41d4-a716-446655440000 PROOF ACCESS(AccessContract)` -assertOk(parseDependentType(test14b), "Test 14b: ACCESS proof") - -let test14c = `SELECT * FROM HEXAD 550e8400-e29b-41d4-a716-446655440000 PROOF PROVENANCE(ProvenanceContract)` -assertOk(parseDependentType(test14c), "Test 14c: PROVENANCE proof") - -// Test 15: Invalid UUID (should fail) -let test15 = ` - SELECT * - FROM HEXAD not-a-valid-uuid -` - -assertError(parse(test15), "Test 15: Invalid UUID format") - -// Test 16: Missing FROM (should fail) -let test16 = ` - SELECT GRAPH - WHERE FULLTEXT CONTAINS "test" -` - -assertError(parse(test16), "Test 16: Missing FROM clause") - -// Test 17: Drift policy variations -let test17a = `SELECT * FROM FEDERATION /nodes/* WITH DRIFT STRICT` -assertOk(parseSlipstream(test17a), "Test 17a: DRIFT STRICT") - -let test17b = `SELECT * FROM FEDERATION /nodes/* WITH DRIFT REPAIR` -assertOk(parseSlipstream(test17b), "Test 17b: DRIFT REPAIR") - -let test17c = `SELECT * FROM FEDERATION /nodes/* WITH DRIFT TOLERATE` -assertOk(parseSlipstream(test17c), "Test 17c: DRIFT TOLERATE") - -let test17d = `SELECT * FROM FEDERATION /nodes/* WITH DRIFT LATEST` -assertOk(parseSlipstream(test17d), "Test 17d: DRIFT LATEST") - -// ============================================================================ -// SQL Compatibility Tests -// ============================================================================ - -Js.Console.log("\n=== SQL Compatibility Tests ===\n") - -// Test 18: ORDER BY single field -let test18 = ` - SELECT DOCUMENT - FROM STORE archive-1 - WHERE FULLTEXT CONTAINS "security" - ORDER BY DOCUMENT.severity DESC - LIMIT 50 -` - -assertOk(parseSlipstream(test18), "Test 18: ORDER BY single field DESC") - -// Test 19: ORDER BY multiple fields -let test19 = ` - SELECT DOCUMENT - FROM FEDERATION /archives/* - ORDER BY DOCUMENT.severity DESC, DOCUMENT.name ASC - LIMIT 100 -` - -assertOk(parseSlipstream(test19), "Test 19: ORDER BY multiple fields") - -// Test 20: ORDER BY default direction (ASC) -let test20 = ` - SELECT DOCUMENT - FROM STORE archive-1 - ORDER BY DOCUMENT.name - LIMIT 10 -` - -assertOk(parseSlipstream(test20), "Test 20: ORDER BY default ASC direction") - -// Test 21: Column projection (DOCUMENT.name, DOCUMENT.severity) -let test21 = ` - SELECT DOCUMENT.name, DOCUMENT.severity - FROM HEXAD 550e8400-e29b-41d4-a716-446655440000 -` - -assertOk(parseSlipstream(test21), "Test 21: Column projection within modality") - -// Test 22: Mixed modalities and column projections -let test22 = ` - SELECT GRAPH, DOCUMENT.name, DOCUMENT.severity - FROM FEDERATION /universities/* - LIMIT 50 -` - -assertOk(parseSlipstream(test22), "Test 22: Mixed modalities and column projections") - -// Test 23: COUNT(*) aggregate -let test23 = ` - SELECT COUNT(*) - FROM FEDERATION /archives/* - WHERE FULLTEXT CONTAINS "vulnerability" -` - -assertOk(parseSlipstream(test23), "Test 23: COUNT(*) aggregate") - -// Test 24: AVG aggregate with field ref -let test24 = ` - SELECT AVG(DOCUMENT.severity) - FROM FEDERATION /scans/* -` - -assertOk(parseSlipstream(test24), "Test 24: AVG(DOCUMENT.severity) aggregate") - -// Test 25: GROUP BY with aggregate -let test25 = ` - SELECT DOCUMENT.name, COUNT(*), AVG(DOCUMENT.severity) - FROM FEDERATION /universities/* - GROUP BY DOCUMENT.name - LIMIT 100 -` - -assertOk(parseSlipstream(test25), "Test 25: GROUP BY with aggregates") - -// Test 26: GROUP BY + HAVING -let test26 = ` - SELECT DOCUMENT.name, COUNT(*) - FROM FEDERATION /archives/* - GROUP BY DOCUMENT.name - HAVING FIELD count > 3 - ORDER BY DOCUMENT.name ASC - LIMIT 50 -` - -assertOk(parseSlipstream(test26), "Test 26: GROUP BY + HAVING + ORDER BY") - -// Test 27: Full SQL-compat query (all features combined) -let test27 = ` - SELECT DOCUMENT.name, DOCUMENT.severity, COUNT(*), SUM(DOCUMENT.severity), AVG(DOCUMENT.severity) - FROM FEDERATION /universities/* WITH DRIFT REPAIR - WHERE FIELD severity > 3 - GROUP BY DOCUMENT.name, DOCUMENT.severity - HAVING FIELD total > 10 - ORDER BY DOCUMENT.severity DESC, DOCUMENT.name ASC - LIMIT 100 - OFFSET 20 -` - -assertOk(parseSlipstream(test27), "Test 27: Full SQL-compat query (all features)") - -// Test 28: MIN/MAX aggregates -let test28 = ` - SELECT MIN(DOCUMENT.severity), MAX(DOCUMENT.severity) - FROM STORE tantivy-node-1 -` - -assertOk(parseSlipstream(test28), "Test 28: MIN/MAX aggregates") - -// Test 29: SQL-compat with PROOF (dependent-type path) -let test29 = ` - SELECT DOCUMENT.name, COUNT(*) - FROM FEDERATION /universities/* WITH DRIFT STRICT - GROUP BY DOCUMENT.name - PROOF INTEGRITY(DataIntegrityContract) - ORDER BY DOCUMENT.name ASC - LIMIT 50 -` - -assertOk(parseDependentType(test29), "Test 29: SQL-compat with PROOF clause") - -// Test 30: Verify parsed projections are populated -switch parse(test21) { -| Ok(query) => { - let hasProjections = query.projections->Belt.Option.isSome - if hasProjections { - Js.Console.log("✓ Test 30: Projections populated correctly") - } else { - Js.Console.error("✗ Test 30: Projections should be Some but got None") - } - } -| Error(e) => Js.Console.error(`✗ Test 30: Parse failed: ${e.message}`) -} - -// Test 31: Verify ORDER BY parsed correctly -switch parse(test18) { -| Ok(query) => { - let hasOrderBy = query.orderBy->Belt.Option.isSome - if hasOrderBy { - Js.Console.log("✓ Test 31: ORDER BY parsed correctly") - } else { - Js.Console.error("✗ Test 31: orderBy should be Some but got None") - } - } -| Error(e) => Js.Console.error(`✗ Test 31: Parse failed: ${e.message}`) -} - -// Test 32: Verify GROUP BY parsed correctly -switch parse(test25) { -| Ok(query) => { - let hasGroupBy = query.groupBy->Belt.Option.isSome - let hasAggregates = query.aggregates->Belt.Option.isSome - if hasGroupBy && hasAggregates { - Js.Console.log("✓ Test 32: GROUP BY and aggregates parsed correctly") - } else { - Js.Console.error(`✗ Test 32: groupBy=${hasGroupBy->Belt.Bool.toString}, aggregates=${hasAggregates->Belt.Bool.toString}`) - } - } -| Error(e) => Js.Console.error(`✗ Test 32: Parse failed: ${e.message}`) -} - -Js.Console.log("\n=== Multi-Proof Tests ===\n") - -// Test 33: Multi-proof composition (AND separated) -let test33 = ` - SELECT GRAPH, SEMANTIC - FROM HEXAD 550e8400-e29b-41d4-a716-446655440000 - PROOF EXISTENCE(ExistenceContract) AND INTEGRITY(IntegrityContract) -` - -assertOk(parseDependentType(test33), "Test 33: Multi-proof (EXISTENCE AND INTEGRITY)") - -// Test 34: Triple proof composition -let test34 = ` - SELECT * - FROM FEDERATION /hospitals/* WITH DRIFT STRICT - PROOF ACCESS(AccessContract) AND PROVENANCE(ProvenanceContract) AND INTEGRITY(IntegrityContract) -` - -assertOk(parseDependentType(test34), "Test 34: Triple proof composition") - -// Test 35: Verify multi-proof parsed as array -switch parse(test33) { -| Ok(query) => { - switch query.proof { - | Some(proofs) => - if Js.Array2.length(proofs) == 2 { - Js.Console.log("✓ Test 35: Multi-proof parsed as array of 2") - } else { - Js.Console.error(`✗ Test 35: Expected 2 proofs, got ${Belt.Int.toString(Js.Array2.length(proofs))}`) - } - | None => Js.Console.error("✗ Test 35: Proof should be Some but got None") - } - } -| Error(e) => Js.Console.error(`✗ Test 35: Parse failed: ${e.message}`) -} - -Js.Console.log("\n=== Cross-Modal Condition Tests ===\n") - -// Test 36: DRIFT condition -let test36 = ` - SELECT VECTOR, DOCUMENT - FROM HEXAD 550e8400-e29b-41d4-a716-446655440000 - WHERE DRIFT(VECTOR, DOCUMENT) > 0.3 -` - -assertOk(parseSlipstream(test36), "Test 36: DRIFT(VECTOR, DOCUMENT) condition") - -// Test 37: CONSISTENT condition -let test37 = ` - SELECT VECTOR, SEMANTIC - FROM HEXAD 550e8400-e29b-41d4-a716-446655440000 - WHERE CONSISTENT(VECTOR, SEMANTIC) USING COSINE -` - -assertOk(parseSlipstream(test37), "Test 37: CONSISTENT(VECTOR, SEMANTIC) USING COSINE") - -// Test 38: EXISTS condition -let test38 = ` - SELECT * - FROM HEXAD 550e8400-e29b-41d4-a716-446655440000 - WHERE VECTOR EXISTS -` - -assertOk(parseSlipstream(test38), "Test 38: VECTOR EXISTS condition") - -// Test 39: NOT EXISTS condition -let test39 = ` - SELECT * - FROM HEXAD 550e8400-e29b-41d4-a716-446655440000 - WHERE TENSOR NOT EXISTS -` - -assertOk(parseSlipstream(test39), "Test 39: TENSOR NOT EXISTS condition") - -// Test 40: Cross-modal field compare -let test40 = ` - SELECT DOCUMENT, GRAPH - FROM HEXAD 550e8400-e29b-41d4-a716-446655440000 - WHERE DOCUMENT.severity > GRAPH.centrality -` - -assertOk(parseSlipstream(test40), "Test 40: Cross-modal field compare") - -Js.Console.log("\n=== Mutation Tests ===\n") - -// Test 41: INSERT mutation -let test41 = ` - INSERT HEXAD WITH - DOCUMENT(title = "New Paper", author = "Jane Doe"), - VECTOR([0.1, 0.2, 0.3, 0.4]) -` - -assertOk(parseMutation(test41), "Test 41: INSERT HEXAD with DOCUMENT and VECTOR") - -// Test 42: UPDATE mutation -let test42 = ` - UPDATE HEXAD 550e8400-e29b-41d4-a716-446655440000 - SET DOCUMENT.title = "Updated Title", DOCUMENT.severity = 5 -` - -assertOk(parseMutation(test42), "Test 42: UPDATE HEXAD with SET") - -// Test 43: DELETE mutation -let test43 = ` - DELETE HEXAD 550e8400-e29b-41d4-a716-446655440000 -` - -assertOk(parseMutation(test43), "Test 43: DELETE HEXAD") - -// Test 44: INSERT with PROOF -let test44 = ` - INSERT HEXAD WITH - DOCUMENT(title = "Verified Entry") - PROOF INTEGRITY(WriteContract) -` - -assertOk(parseMutation(test44), "Test 44: INSERT with PROOF clause") - -// Test 45: DELETE with multi-proof -let test45 = ` - DELETE HEXAD 550e8400-e29b-41d4-a716-446655440000 - PROOF ACCESS(AccessContract) AND PROVENANCE(AuditContract) -` - -assertOk(parseMutation(test45), "Test 45: DELETE with multi-proof") - -Js.Console.log("\n=== Statement Tests ===\n") - -// Test 46: parseStatement with query -let test46 = ` - SELECT GRAPH - FROM HEXAD 550e8400-e29b-41d4-a716-446655440000 - LIMIT 10 -` - -assertOk(parseStatement(test46), "Test 46: parseStatement dispatches to query") - -// Test 47: parseStatement with mutation -let test47 = ` - DELETE HEXAD 550e8400-e29b-41d4-a716-446655440000 -` - -assertOk(parseStatement(test47), "Test 47: parseStatement dispatches to mutation") - -// Test 48: parseStatement with INSERT -let test48 = ` - INSERT HEXAD WITH DOCUMENT(title = "Test") -` - -assertOk(parseStatement(test48), "Test 48: parseStatement dispatches INSERT to mutation") - -Js.Console.log("\n=== Tests Complete ===\n") - -// ============================================================================ -// Example: Extracting parsed data -// ============================================================================ - -Js.Console.log("=== Example: Parsing and Inspecting Query ===\n") - -let exampleQuery = ` - SELECT GRAPH, VECTOR - FROM FEDERATION /universities/* WITH DRIFT REPAIR - WHERE FULLTEXT CONTAINS "neural networks" - PROOF CITATION(NeuralNetworkContract) - LIMIT 50 -` - -switch parseDependentType(exampleQuery) { -| Ok(query) => { - Js.Console.log("Parsed query successfully:") - Js.Console.log(` Modalities: ${query.modalities->Js.Array2.length->Belt.Int.toString}`) - Js.Console.log(` Source: ${switch query.source { - | Hexad(id) => `Hexad(${id})` - | Federation(pattern, drift) => { - let driftStr = switch drift { - | Some(Strict) => " WITH DRIFT STRICT" - | Some(Repair) => " WITH DRIFT REPAIR" - | Some(Tolerate) => " WITH DRIFT TOLERATE" - | Some(Latest) => " WITH DRIFT LATEST" - | None => "" - } - `Federation(${pattern}${driftStr})` - } - | Store(id) => `Store(${id})` - }}`) - Js.Console.log(` Has WHERE: ${query.where->Belt.Option.isSome->Belt.Bool.toString}`) - Js.Console.log(` Has PROOF: ${query.proof->Belt.Option.isSome->Belt.Bool.toString}`) - switch query.proof { - | Some(proofs) => - proofs->Js.Array2.forEach(proof => { - Js.Console.log(` Proof contract: ${proof.contractName}`) - }) - | None => () - } - Js.Console.log(` Limit: ${switch query.limit { - | Some(n) => Belt.Int.toString(n) - | None => "None" - }}`) - } -| Error(e) => { - Js.Console.error(`Parse error: ${e.message}`) - } -} - -Js.Console.log("\n") diff --git a/verisimdb/src/vcl/VCLProofObligation.res b/verisimdb/src/vcl/VCLProofObligation.res deleted file mode 100644 index edcaf4ed..00000000 --- a/verisimdb/src/vcl/VCLProofObligation.res +++ /dev/null @@ -1,251 +0,0 @@ -// SPDX-License-Identifier: MPL-2.0 -// VCL Proof Obligation — Generates typed proof obligations from queries -// -// For each PROOF spec in a dependent-type query, generates a structured -// obligation that the executor must satisfy. Handles multi-proof -// composition validation and conflict detection. - -module AST = VCLParser.AST -module Types = VCLTypes -module Ctx = VCLContext - -// ============================================================================ -// Proof Obligation Types -// ============================================================================ - -type obligationKind = - | ExistenceObligation // Hexad must exist and be accessible - | IntegrityObligation // Data integrity (hash/Merkle verification) - | AccessObligation // Access control (ZKP-based permission) - | CitationObligation // Citation chain validity - | ProvenanceObligation // Lineage/provenance chain - | CustomObligation(string) // Custom contract-specific - -type proofObligation = { - kind: obligationKind, - contractName: string, - witnessFields: array<string>, - circuit: string, - estimatedTimeMs: int, - requiredModalities: array<Types.modalityType>, -} - -type composedProofPlan = { - obligations: array<proofObligation>, - totalEstimatedTimeMs: int, - isParallelizable: bool, - compositionStrategy: compositionStrategy, -} - -and compositionStrategy = - | Independent // proofs are independent, can run in parallel - | Sequential(array<int>) // indices of obligations in dependency order - | Nested // proof N requires result of proof N-1 - -// ============================================================================ -// Obligation generation -// ============================================================================ - -let generateObligation = ( - ctx: Ctx.context, - proofSpec: AST.proofSpec, - queryResultType: Types.queryResultInfo, -): Result<proofObligation, string> => { - let proofKind = Types.proofKindOfAstProofType(proofSpec.proofType) - let kind = proofKindToObligationKind(proofKind, proofSpec.contractName) - - // Determine witness fields based on proof type - let witnessFields = switch proofSpec.proofType { - | Existence => ["hexad_id", "timestamp"] - | Citation => ["hexad_id", "citation_chain", "source_ids"] - | Access => ["hexad_id", "user_id", "role", "permissions"] - | Integrity => ["hexad_id", "modality_hashes", "merkle_root"] - | Provenance => ["hexad_id", "lineage_chain", "actors", "timestamps"] - | Custom => ["hexad_id", "contract_params"] - } - - // Get required modalities from contract spec or query - let requiredMods = switch Ctx.lookupContract(ctx, proofSpec.contractName) { - | Some(spec) => spec.requiredModalities - | None => queryResultType.modalities - } - - Ok({ - kind, - contractName: proofSpec.contractName, - witnessFields, - circuit: proofKindToCircuit(proofKind), - estimatedTimeMs: estimateProofTime(proofKind), - requiredModalities: requiredMods, - }) -} - -// Generate obligations for multiple proof specs (multi-proof) -let generateObligations = ( - ctx: Ctx.context, - proofSpecs: array<AST.proofSpec>, - queryResultType: Types.queryResultInfo, -): Result<composedProofPlan, string> => { - // Generate each obligation - let obligationResults = proofSpecs->Belt.Array.map(spec => { - generateObligation(ctx, spec, queryResultType) - }) - - // Check for errors - let firstError = obligationResults->Belt.Array.getBy(r => { - switch r { - | Error(_) => true - | Ok(_) => false - } - }) - - switch firstError { - | Some(Error(e)) => Error(e) - | _ => - let obligations = obligationResults->Belt.Array.keepMap(r => { - switch r { - | Ok(o) => Some(o) - | Error(_) => None - } - }) - - // Determine composition strategy - let strategy = determineCompositionStrategy(obligations) - let totalTime = switch strategy { - | Independent => - // Parallel: max of all obligations - obligations->Belt.Array.reduce(0, (acc, o) => - Js.Math.max_int(acc, o.estimatedTimeMs) - ) - | Sequential(_) | Nested => - // Sequential: sum of all obligations - obligations->Belt.Array.reduce(0, (acc, o) => acc + o.estimatedTimeMs) - } - - Ok({ - obligations, - totalEstimatedTimeMs: totalTime, - isParallelizable: strategy == Independent, - compositionStrategy: strategy, - }) - } -} - -// ============================================================================ -// Composition strategy determination -// ============================================================================ - -let determineCompositionStrategy = (obligations: array<proofObligation>): compositionStrategy => { - let len = Js.Array2.length(obligations) - if len <= 1 { - Independent - } else { - // Check if any obligation depends on another's result - let hasProvenance = obligations->Js.Array2.some(o => { - switch o.kind { - | ProvenanceObligation => true - | _ => false - } - }) - - let hasCitation = obligations->Js.Array2.some(o => { - switch o.kind { - | CitationObligation => true - | _ => false - } - }) - - // Provenance + Citation must be sequential (provenance verifies citation chain) - if hasProvenance && hasCitation { - // Citation first, then provenance - let indices = [] - obligations->Js.Array2.forEachi((o, i) => { - switch o.kind { - | CitationObligation => indices->Js.Array2.push(i)->ignore - | _ => () - } - }) - obligations->Js.Array2.forEachi((o, i) => { - switch o.kind { - | CitationObligation => () // already added - | _ => indices->Js.Array2.push(i)->ignore - } - }) - Sequential(indices) - } else { - // All other combinations are independent - Independent - } - } -} - -// ============================================================================ -// Utility functions -// ============================================================================ - -let proofKindToObligationKind = (pk: Types.proofKind, contractName: string): obligationKind => { - switch pk { - | ExistenceProof => ExistenceObligation - | CitationProof => CitationObligation - | AccessProof => AccessObligation - | IntegrityProof => IntegrityObligation - | ProvenanceProof => ProvenanceObligation - | CustomProof => CustomObligation(contractName) - } -} - -let proofKindToCircuit = (pk: Types.proofKind): string => { - switch pk { - | ExistenceProof => "existence-proof-v1" - | CitationProof => "citation-proof-v1" - | AccessProof => "access-control-v1" - | IntegrityProof => "integrity-check-v1" - | ProvenanceProof => "provenance-chain-v1" - | CustomProof => "custom-circuit" - } -} - -let estimateProofTime = (pk: Types.proofKind): int => { - switch pk { - | ExistenceProof => 50 - | CitationProof => 100 - | AccessProof => 150 - | IntegrityProof => 200 - | ProvenanceProof => 300 - | CustomProof => 500 - } -} - -let obligationKindToString = (k: obligationKind): string => { - switch k { - | ExistenceObligation => "EXISTENCE" - | IntegrityObligation => "INTEGRITY" - | AccessObligation => "ACCESS" - | CitationObligation => "CITATION" - | ProvenanceObligation => "PROVENANCE" - | CustomObligation(name) => `CUSTOM(${name})` - } -} - -let formatObligation = (o: proofObligation): string => { - let kind = obligationKindToString(o.kind) - let witnesses = o.witnessFields->Js.Array2.joinWith(", ") - `${kind}(${o.contractName}) circuit=${o.circuit} witnesses=[${witnesses}] est=${Belt.Int.toString(o.estimatedTimeMs)}ms` -} - -let formatPlan = (plan: composedProofPlan): string => { - let lines = [] - lines->Js.Array2.push("Proof Plan:")->ignore - lines->Js.Array2.push(` Strategy: ${switch plan.compositionStrategy { - | Independent => "Independent (parallel)" - | Sequential(_) => "Sequential (ordered)" - | Nested => "Nested (chained)" - }}`)->ignore - lines->Js.Array2.push(` Parallelizable: ${plan.isParallelizable ? "yes" : "no"}`)->ignore - lines->Js.Array2.push(` Estimated time: ${Belt.Int.toString(plan.totalEstimatedTimeMs)}ms`)->ignore - lines->Js.Array2.push(` Obligations:`)->ignore - plan.obligations->Js.Array2.forEachi((o, i) => { - lines->Js.Array2.push(` ${Belt.Int.toString(i + 1)}. ${formatObligation(o)}`)->ignore - }) - lines->Js.Array2.joinWith("\n") -} diff --git a/verisimdb/src/vcl/VCLSubtyping.res b/verisimdb/src/vcl/VCLSubtyping.res deleted file mode 100644 index cd1092d0..00000000 --- a/verisimdb/src/vcl/VCLSubtyping.res +++ /dev/null @@ -1,247 +0,0 @@ -// SPDX-License-Identifier: MPL-2.0 -// VCL Subtyping — Subtype relation for the VCL type system -// -// Implements the 6 subtyping rules from the formal spec (Section 4): -// 1. Reflexivity: t <: t -// 2. Transitivity: if t1 <: t2 and t2 <: t3 then t1 <: t3 -// 3. List covariance: if t <: s then Array<t> <: Array<s> -// 4. Arrow contra/covariance: if s1 <: t1 and t2 <: s2 then (t1 -> t2) <: (s1 -> s2) -// 5. Hexad modality contravariance: requesting fewer modalities subtypes requesting more -// 6. Refinement subsumption: DEFERRED (needs SMT solver) - -module Types = VCLTypes - -type subtypeResult = Result<unit, subtypeError> - -and subtypeError = { - expected: Types.vclType, - got: Types.vclType, - reason: string, -} - -// ============================================================================ -// Primitive subtyping -// ============================================================================ - -// Numeric widening: Int <: Float (safe promotion) -let isSubPrimitive = (sub: Types.primitiveType, sup: Types.primitiveType): bool => { - Types.eqPrimitiveType(sub, sup) || - switch (sub, sup) { - | (IntType, FloatType) => true // Int widens to Float - | (VectorType(_), VectorType(0)) => true // any vector subtypes unknown-dim vector - | _ => false - } -} - -// ============================================================================ -// Core subtype relation -// ============================================================================ - -let rec isSubtype = (sub: Types.vclType, sup: Types.vclType): subtypeResult => { - // Rule 1: Reflexivity - if Types.eqType(sub, sup) { - Ok() - } else { - checkStructuralSubtype(sub, sup) - } -} - -and checkStructuralSubtype = (sub: Types.vclType, sup: Types.vclType): subtypeResult => { - switch (sub, sup) { - // Primitive widening - | (Primitive(ps), Primitive(pp)) => - if isSubPrimitive(ps, pp) { - Ok() - } else { - Error({ - expected: sup, - got: sub, - reason: `${Types.primitiveTypeToString(ps)} is not a subtype of ${Types.primitiveTypeToString(pp)}`, - }) - } - - // Rule 3: List covariance — Array<t> <: Array<s> if t <: s - | (ArrayType(innerSub), ArrayType(innerSup)) => - switch isSubtype(innerSub, innerSup) { - | Ok() => Ok() - | Error(e) => - Error({ - expected: sup, - got: sub, - reason: `Array element type mismatch: ${e.reason}`, - }) - } - - // Rule 4: Arrow contra/covariance — Pi(x, t1, t2) <: Pi(x, s1, s2) if s1 <: t1 and t2 <: s2 - | (PiType(_, domSub, codSub), PiType(_, domSup, codSup)) => - // Contravariant in domain - switch isSubtype(domSup, domSub) { - | Ok() => - // Covariant in codomain - switch isSubtype(codSub, codSup) { - | Ok() => Ok() - | Error(e) => - Error({ - expected: sup, - got: sub, - reason: `Function codomain: ${e.reason}`, - }) - } - | Error(e) => - Error({ - expected: sup, - got: sub, - reason: `Function domain (contravariant): ${e.reason}`, - }) - } - - // Rule 5: Hexad modality contravariance - // A hexad with MORE modalities is a subtype of one requesting FEWER - // (having more data satisfies a request for less) - | (HexadType(modsSub), HexadType(modsSup)) => - // Check that every modality in sup is present in sub - let missing = modsSup->Belt.Array.keep(supMod => { - !(modsSub->Js.Array2.some(subMod => Types.eqModalityType(subMod, supMod))) - }) - if Js.Array2.length(missing) == 0 { - Ok() - } else { - let missingStrs = missing->Belt.Array.map(Types.modalityTypeToString)->Js.Array2.joinWith(", ") - Error({ - expected: sup, - got: sub, - reason: `Hexad missing required modalities: ${missingStrs}`, - }) - } - - // ModalityType is a subtype of itself only (handled by reflexivity) - | (ModalityType(a), ModalityType(b)) => - if Types.eqModalityType(a, b) { - Ok() - } else { - Error({ - expected: sup, - got: sub, - reason: `${Types.modalityTypeToString(a)} is not ${Types.modalityTypeToString(b)}`, - }) - } - - // NeverType is a subtype of everything (bottom type) - | (NeverType, _) => Ok() - - // Everything is a subtype of UnitType (top for values) - | (_, UnitType) => Ok() - - // ProvedResult subtypes plain QueryResult (can forget proof) - | (ProvedResultType(info, _, _), QueryResultType(infoSup)) => - isSubQueryResult(info, infoSup) - - // QueryResult subtyping: covariant in modalities and projections - | (QueryResultType(infoSub), QueryResultType(infoSup)) => - isSubQueryResult(infoSub, infoSup) - - // No other subtyping relationships - | _ => - Error({ - expected: sup, - got: sub, - reason: `${Types.vclTypeToString(sub)} is not a subtype of ${Types.vclTypeToString(sup)}`, - }) - } -} - -// Query result subtyping: sub result must provide at least what sup requires -and isSubQueryResult = ( - sub: Types.queryResultInfo, - sup: Types.queryResultInfo, -): subtypeResult => { - // Check that all required modalities are present - let missingMods = sup.modalities->Belt.Array.keep(supMod => { - !(sub.modalities->Js.Array2.some(subMod => Types.eqModalityType(subMod, supMod))) - }) - - if Js.Array2.length(missingMods) > 0 { - let missingStrs = missingMods->Belt.Array.map(Types.modalityTypeToString)->Js.Array2.joinWith(", ") - Error({ - expected: QueryResultType(sup), - got: QueryResultType(sub), - reason: `Result missing modalities: ${missingStrs}`, - }) - } else { - Ok() - } -} - -// ============================================================================ -// Rule 2: Transitivity check (explicit) -// If a <: b and b <: c, then a <: c -// ============================================================================ - -let transitiveSubtype = ( - a: Types.vclType, - b: Types.vclType, - c: Types.vclType, -): subtypeResult => { - switch isSubtype(a, b) { - | Ok() => - switch isSubtype(b, c) { - | Ok() => Ok() - | Error(e) => - Error({ - expected: c, - got: a, - reason: `Transitivity failed at second step: ${e.reason}`, - }) - } - | Error(e) => - Error({ - expected: c, - got: a, - reason: `Transitivity failed at first step: ${e.reason}`, - }) - } -} - -// ============================================================================ -// Convenience: check operator type compatibility -// ============================================================================ - -// Given two types and an operator, check if comparison is valid -let checkOperatorTypes = ( - leftType: Types.primitiveType, - op: VCLParser.AST.operator, - rightType: Types.primitiveType, -): subtypeResult => { - // Types must be compatible (one subtype of the other, or same) - let compatible = - Types.eqPrimitiveType(leftType, rightType) || - isSubPrimitive(leftType, rightType) || - isSubPrimitive(rightType, leftType) - - if !compatible { - Error({ - expected: Primitive(leftType), - got: Primitive(rightType), - reason: `Cannot compare ${Types.primitiveTypeToString(leftType)} with ${Types.primitiveTypeToString(rightType)}`, - }) - } else if !Types.isOperatorValidForType(op, leftType) && !Types.isOperatorValidForType(op, rightType) { - let opStr = switch op { - | Eq => "==" - | Neq => "!=" - | Gt => ">" - | Lt => "<" - | Gte => ">=" - | Lte => "<=" - | Like => "LIKE" - | Contains => "CONTAINS" - | Matches => "MATCHES" - } - Error({ - expected: Primitive(leftType), - got: Primitive(rightType), - reason: `Operator ${opStr} is not valid for types ${Types.primitiveTypeToString(leftType)} and ${Types.primitiveTypeToString(rightType)}`, - }) - } else { - Ok() - } -} diff --git a/verisimdb/src/vcl/VCLTypeChecker.res b/verisimdb/src/vcl/VCLTypeChecker.res deleted file mode 100644 index ec03e920..00000000 --- a/verisimdb/src/vcl/VCLTypeChecker.res +++ /dev/null @@ -1,354 +0,0 @@ -// SPDX-License-Identifier: MPL-2.0 -// VCL Type Checker — Thin facade over VCLBidir bidirectional type inference -// -// Maintains backward-compatible public API (checkQuery, planProofGeneration) -// while delegating to the real type system in VCLBidir. - -module AST = VCLParser.AST -module Error = VCLError -module Types = VCLTypes -module Ctx = VCLContext -module Bidir = VCLBidir -module ProofObl = VCLProofObligation - -// ============================================================================ -// Type Context (backward-compatible) -// ============================================================================ - -type contractInfo = { - name: string, - proofType: AST.proofType, - requiredFields: array<string>, - constraints: array<constraint>, -} - -and constraint = - | FieldMustExist(string) - | FieldMustBeType(string, primitiveType) - | ModalityRequired(AST.modality) - | MinimumSelectivity(float) - | MaxDriftThreshold(float) - -and primitiveType = - | TString - | TInt - | TFloat - | TBool - | TArray(primitiveType) - | TVector(int) - | TUuid - -type typeContext = { - contracts: Js.Dict.t<contractInfo>, - availableModalities: array<AST.modality>, - strictMode: bool, -} - -// ============================================================================ -// Context construction -// ============================================================================ - -type typeCheckResult = Result<unit, Error.typeError> - -let makeTypeContext = ( - ~contracts: Js.Dict.t<contractInfo>=Js.Dict.empty(), - ~strictMode=true, - (), -): typeContext => { - { - contracts: contracts, - availableModalities: [Graph, Vector, Tensor, Semantic, Document, Temporal], - strictMode: strictMode, - } -} - -let createDefaultContext = (): typeContext => { - makeTypeContext(~strictMode=true, ()) -} - -// Convert old-style context to new bidirectional context -let toBidirContext = (ctx: typeContext): Ctx.context => { - let bidirCtx = Ctx.defaultContext() - - // Register contracts from old context - Js.Dict.entries(ctx.contracts)->Js.Array2.forEach(((name, info)) => { - let proofKind = Types.proofKindOfAstProofType(info.proofType) - let requiredMods = info.constraints->Belt.Array.keepMap(c => { - switch c { - | ModalityRequired(m) => Types.modalityTypeOfAstModality(m) - | _ => None - } - }) - let spec: Ctx.contractSpec = { - name: name, - proofKind: proofKind, - requiredModalities: requiredMods, - requiredFields: [], - composableWith: [ - Types.ExistenceProof, - Types.CitationProof, - Types.AccessProof, - Types.IntegrityProof, - Types.ProvenanceProof, - Types.CustomProof, - ], - } - Js.Dict.set(bidirCtx.contracts, name, spec) - }) - - bidirCtx -} - -// ============================================================================ -// Public API (backward-compatible) -// ============================================================================ - -let checkQuery = (query: AST.query, context: typeContext): typeCheckResult => { - // First: SQL-compat validation (HAVING requires GROUP BY, etc.) - switch checkSqlCompat(query) { - | Error(e) => Error(e) - | Ok() => - // Delegate to bidirectional type checker - let bidirCtx = toBidirContext(context) - switch Bidir.synthesizeQuery(bidirCtx, query) { - | Ok(_type) => Ok() - | Error(typeErr) => - // Convert Bidir.typeError to VCLError.typeError - Error({ - kind: Error.TypeMismatch({ - expected: "well-typed query", - found: Bidir.formatTypeError(typeErr), - }), - hexad_id: None, - modality: None, - context: Bidir.formatTypeError(typeErr), - }) - } - } -} - -// ============================================================================ -// Proof Generation Verification -// ============================================================================ - -type proofGenerationResult = Result<proofPlan, Error.typeError> - -and proofPlan = { - contract: string, - proofType: AST.proofType, - witnessFields: array<string>, - circuit: string, - estimatedTimeMs: int, -} - -let planProofGeneration = (query: AST.query, context: typeContext): proofGenerationResult => { - switch query.proof { - | None => - Error({ - kind: Error.MissingTypeAnnotation("proof"), - hexad_id: None, - modality: None, - context: "No PROOF clause in query", - }) - | Some(proofSpecs) => { - let bidirCtx = toBidirContext(context) - let resolvedMods = Types.resolveModalities(query.modalities) - let resultInfo: Types.queryResultInfo = { - modalities: resolvedMods, - projections: [], - aggregates: [], - } - switch ProofObl.generateObligations(bidirCtx, proofSpecs, resultInfo) { - | Error(msg) => - Error({ - kind: Error.ProofGenerationFailed({ - contract: switch proofSpecs[0] { - | Some(s) => s.contractName - | None => "unknown" - }, - error: msg, - }), - hexad_id: None, - modality: None, - context: msg, - }) - | Ok(plan) => - // Return the first obligation as the primary plan (backward compat) - switch plan.obligations[0] { - | Some(obl) => - Ok({ - contract: obl.contractName, - proofType: switch proofSpecs[0] { - | Some(s) => s.proofType - | None => Existence - }, - witnessFields: obl.witnessFields, - circuit: obl.circuit, - estimatedTimeMs: plan.totalEstimatedTimeMs, - }) - | None => - Error({ - kind: Error.MissingTypeAnnotation("proof"), - hexad_id: None, - modality: None, - context: "No proof obligations generated", - }) - } - } - } - } -} - -// ============================================================================ -// SQL-Compat Validation -// ============================================================================ - -let validateAggregates = (query: AST.query): typeCheckResult => { - switch (query.having, query.groupBy) { - | (Some(_), None) => - Error({ - kind: Error.ContractViolation({ - contract: "sql-compat", - reason: "HAVING clause requires GROUP BY", - }), - hexad_id: None, - modality: None, - context: "Add a GROUP BY clause or remove the HAVING clause", - }) - | _ => Ok() - } -} - -let validateOrderByModalities = (query: AST.query): typeCheckResult => { - switch query.orderBy { - | None => Ok() - | Some(items) => { - let invalidField = items->Belt.Array.getBy(item => { - !(query.modalities->Belt.Array.some(m => m == item.field.modality || m == All)) - }) - switch invalidField { - | None => Ok() - | Some(item) => { - let modStr = modalityToString(item.field.modality) - Error({ - kind: Error.ContractViolation({ - contract: "order-by-validation", - reason: `ORDER BY references modality ${modStr} which is not in SELECT`, - }), - hexad_id: None, - modality: Some(modStr), - context: `Add ${modStr} to your SELECT clause or remove it from ORDER BY`, - }) - } - } - } - } -} - -let validateGroupByModalities = (query: AST.query): typeCheckResult => { - switch query.groupBy { - | None => Ok() - | Some(fields) => { - let invalidField = fields->Belt.Array.getBy(f => { - !(query.modalities->Belt.Array.some(m => m == f.modality || m == All)) - }) - switch invalidField { - | None => Ok() - | Some(field) => { - let modStr = modalityToString(field.modality) - Error({ - kind: Error.ContractViolation({ - contract: "group-by-validation", - reason: `GROUP BY references modality ${modStr} which is not in SELECT`, - }), - hexad_id: None, - modality: Some(modStr), - context: `Add ${modStr} to your SELECT clause or remove it from GROUP BY`, - }) - } - } - } - } -} - -let checkSqlCompat = (query: AST.query): typeCheckResult => { - switch validateAggregates(query) { - | Error(e) => Error(e) - | Ok() => - switch validateOrderByModalities(query) { - | Error(e) => Error(e) - | Ok() => validateGroupByModalities(query) - } - } -} - -// ============================================================================ -// Utility Functions -// ============================================================================ - -let proofTypeToString = (pt: AST.proofType): string => { - switch pt { - | Existence => "EXISTENCE" - | Citation => "CITATION" - | Access => "ACCESS" - | Integrity => "INTEGRITY" - | Provenance => "PROVENANCE" - | Custom => "CUSTOM" - } -} - -let modalityToString = (m: AST.modality): string => { - switch m { - | Graph => "GRAPH" - | Vector => "VECTOR" - | Tensor => "TENSOR" - | Semantic => "SEMANTIC" - | Document => "DOCUMENT" - | Temporal => "TEMPORAL" - | All => "*" - } -} - -let primitiveTypeToString = (pt: primitiveType): string => { - switch pt { - | TString => "String" - | TInt => "Int" - | TFloat => "Float" - | TBool => "Bool" - | TArray(inner) => `Array<${primitiveTypeToString(inner)}>` - | TVector(dim) => `Vector<${Belt.Int.toString(dim)}>` - | TUuid => "UUID" - } -} - -// ============================================================================ -// Registration (backward-compatible) -// ============================================================================ - -let registerContract = ( - context: typeContext, - name: string, - info: contractInfo, -): typeContext => { - let newContracts = Js.Dict.fromArray(Js.Dict.entries(context.contracts)) - Js.Dict.set(newContracts, name, info) - {...context, contracts: newContracts} -} - -// Export for testing -let testCreateConstraint = ( - constraintType: string, - param: string, -): option<constraint> => { - switch constraintType { - | "FieldMustExist" => Some(FieldMustExist(param)) - | "ModalityRequired" => - switch param { - | "GRAPH" => Some(ModalityRequired(Graph)) - | "VECTOR" => Some(ModalityRequired(Vector)) - | "SEMANTIC" => Some(ModalityRequired(Semantic)) - | _ => None - } - | _ => None - } -} diff --git a/verisimdb/src/vcl/VCLTypes.res b/verisimdb/src/vcl/VCLTypes.res deleted file mode 100644 index 8b2a32a7..00000000 --- a/verisimdb/src/vcl/VCLTypes.res +++ /dev/null @@ -1,305 +0,0 @@ -// SPDX-License-Identifier: MPL-2.0 -// VCL Types — Core type definitions for the bidirectional type checker -// -// Implements the type system from vcl-type-system.adoc: -// - Pi types (dependent function types) -// - Sigma types (dependent pair types) -// - Modality types -// - Hexad types -// - Query result types -// - Proof types - -module AST = VCLParser.AST - -// ============================================================================ -// Modality Types (type-level representation) -// ============================================================================ - -type modalityType = - | GraphModality - | VectorModality - | TensorModality - | SemanticModality - | DocumentModality - | TemporalModality - | ProvenanceModality - | SpatialModality - -// ============================================================================ -// Primitive Types -// ============================================================================ - -type primitiveType = - | IntType - | FloatType - | StringType - | BoolType - | VectorType(int) // fixed-size vector with dimension - | TensorType(array<int>) // shape - | UuidType - | TimestampType - -// ============================================================================ -// Core VCL Type System -// ============================================================================ - -type rec vclType = - | Primitive(primitiveType) - | ArrayType(vclType) - | ModalityType(modalityType) - | HexadType(array<modalityType>) // hexad carrying specific modalities - | QueryResultType(queryResultInfo) // result of a SELECT query - | ProofType(proofKind, string) // Proof<kind, contract> - | ProvedResultType(queryResultInfo, proofKind, string) // Sigma(result, proof) - | PiType(string, vclType, vclType) // Pi(x, domain, codomain) - | SigmaType(string, vclType, vclType) // Sigma(x, fst, snd) - | UnitType - | NeverType - -and queryResultInfo = { - modalities: array<modalityType>, - projections: array<fieldTypeInfo>, - aggregates: array<aggregateTypeInfo>, -} - -and fieldTypeInfo = { - modality: modalityType, - fieldName: string, - fieldType: primitiveType, -} - -and aggregateTypeInfo = { - func: AST.aggregateFunc, - resultType: primitiveType, - sourceField: option<fieldTypeInfo>, -} - -and proofKind = - | ExistenceProof - | CitationProof - | AccessProof - | IntegrityProof - | ProvenanceProof - | CustomProof - -// ============================================================================ -// Conversions from AST types -// ============================================================================ - -let proofKindOfAstProofType = (pt: AST.proofType): proofKind => { - switch pt { - | Existence => ExistenceProof - | Citation => CitationProof - | Access => AccessProof - | Integrity => IntegrityProof - | Provenance => ProvenanceProof - | Custom => CustomProof - } -} - -let astProofTypeOfProofKind = (pk: proofKind): AST.proofType => { - switch pk { - | ExistenceProof => Existence - | CitationProof => Citation - | AccessProof => Access - | IntegrityProof => Integrity - | ProvenanceProof => Provenance - | CustomProof => Custom - } -} - -let modalityTypeOfAstModality = (m: AST.modality): option<modalityType> => { - switch m { - | Graph => Some(GraphModality) - | Vector => Some(VectorModality) - | Tensor => Some(TensorModality) - | Semantic => Some(SemanticModality) - | Document => Some(DocumentModality) - | Temporal => Some(TemporalModality) - | Provenance => Some(ProvenanceModality) - | Spatial => Some(SpatialModality) - | All => None // 'All' expands to all modalities, not a single type - } -} - -let astModalityOfModalityType = (m: modalityType): AST.modality => { - switch m { - | GraphModality => Graph - | VectorModality => Vector - | TensorModality => Tensor - | SemanticModality => Semantic - | DocumentModality => Document - | TemporalModality => Temporal - | ProvenanceModality => Provenance - | SpatialModality => Spatial - } -} - -let allModalityTypes: array<modalityType> = [ - GraphModality, - VectorModality, - TensorModality, - SemanticModality, - DocumentModality, - TemporalModality, - ProvenanceModality, - SpatialModality, -] - -let resolveModalities = (mods: array<AST.modality>): array<modalityType> => { - if mods->Js.Array2.some(m => m == AST.All) { - allModalityTypes - } else { - mods->Belt.Array.keepMap(modalityTypeOfAstModality) - } -} - -// ============================================================================ -// String representations for error messages -// ============================================================================ - -let modalityTypeToString = (m: modalityType): string => { - switch m { - | GraphModality => "GRAPH" - | VectorModality => "VECTOR" - | TensorModality => "TENSOR" - | SemanticModality => "SEMANTIC" - | DocumentModality => "DOCUMENT" - | TemporalModality => "TEMPORAL" - | ProvenanceModality => "PROVENANCE" - | SpatialModality => "SPATIAL" - } -} - -let primitiveTypeToString = (pt: primitiveType): string => { - switch pt { - | IntType => "Int" - | FloatType => "Float" - | StringType => "String" - | BoolType => "Bool" - | VectorType(dim) => `Vector<${Belt.Int.toString(dim)}>` - | TensorType(shape) => { - let shapeStr = shape->Belt.Array.map(Belt.Int.toString)->Js.Array2.joinWith("x") - `Tensor<${shapeStr}>` - } - | UuidType => "UUID" - | TimestampType => "Timestamp" - } -} - -let proofKindToString = (pk: proofKind): string => { - switch pk { - | ExistenceProof => "EXISTENCE" - | CitationProof => "CITATION" - | AccessProof => "ACCESS" - | IntegrityProof => "INTEGRITY" - | ProvenanceProof => "PROVENANCE" - | CustomProof => "CUSTOM" - } -} - -let rec vclTypeToString = (t: vclType): string => { - switch t { - | Primitive(pt) => primitiveTypeToString(pt) - | ArrayType(inner) => `Array<${vclTypeToString(inner)}>` - | ModalityType(m) => modalityTypeToString(m) - | HexadType(mods) => { - let modStrs = mods->Belt.Array.map(modalityTypeToString)->Js.Array2.joinWith(", ") - `Hexad<${modStrs}>` - } - | QueryResultType(info) => { - let modStrs = info.modalities->Belt.Array.map(modalityTypeToString)->Js.Array2.joinWith(", ") - `QueryResult<${modStrs}>` - } - | ProofType(kind, contract) => `Proof<${proofKindToString(kind)}, ${contract}>` - | ProvedResultType(info, kind, contract) => { - let modStrs = info.modalities->Belt.Array.map(modalityTypeToString)->Js.Array2.joinWith(", ") - `Sigma(QueryResult<${modStrs}>, Proof<${proofKindToString(kind)}, ${contract}>)` - } - | PiType(x, domain, codomain) => - `Pi(${x}: ${vclTypeToString(domain)}) -> ${vclTypeToString(codomain)}` - | SigmaType(x, fst, snd) => - `Sigma(${x}: ${vclTypeToString(fst)}, ${vclTypeToString(snd)})` - | UnitType => "Unit" - | NeverType => "Never" - } -} - -// ============================================================================ -// Type equality (structural) -// ============================================================================ - -let rec eqPrimitiveType = (a: primitiveType, b: primitiveType): bool => { - switch (a, b) { - | (IntType, IntType) => true - | (FloatType, FloatType) => true - | (StringType, StringType) => true - | (BoolType, BoolType) => true - | (VectorType(d1), VectorType(d2)) => d1 == d2 - | (TensorType(s1), TensorType(s2)) => - Js.Array2.length(s1) == Js.Array2.length(s2) && - s1->Js.Array2.everyi((v, i) => { - switch s2[i] { - | Some(v2) => v == v2 - | None => false - } - }) - | (UuidType, UuidType) => true - | (TimestampType, TimestampType) => true - | _ => false - } -} - -let eqModalityType = (a: modalityType, b: modalityType): bool => { - switch (a, b) { - | (GraphModality, GraphModality) => true - | (VectorModality, VectorModality) => true - | (TensorModality, TensorModality) => true - | (SemanticModality, SemanticModality) => true - | (DocumentModality, DocumentModality) => true - | (TemporalModality, TemporalModality) => true - | _ => false - } -} - -let rec eqType = (a: vclType, b: vclType): bool => { - switch (a, b) { - | (Primitive(pa), Primitive(pb)) => eqPrimitiveType(pa, pb) - | (ArrayType(ia), ArrayType(ib)) => eqType(ia, ib) - | (ModalityType(ma), ModalityType(mb)) => eqModalityType(ma, mb) - | (HexadType(ma), HexadType(mb)) => - Js.Array2.length(ma) == Js.Array2.length(mb) && - ma->Js.Array2.every(m => mb->Js.Array2.some(m2 => eqModalityType(m, m2))) - | (UnitType, UnitType) => true - | (NeverType, NeverType) => true - | (ProofType(k1, c1), ProofType(k2, c2)) => k1 == k2 && c1 == c2 - | _ => false - } -} - -// ============================================================================ -// Numeric type checks (for aggregates and comparisons) -// ============================================================================ - -let isNumericPrimitive = (pt: primitiveType): bool => { - switch pt { - | IntType | FloatType => true - | _ => false - } -} - -let isComparablePrimitive = (pt: primitiveType): bool => { - switch pt { - | IntType | FloatType | StringType | TimestampType => true - | _ => false - } -} - -// Check if an operator is valid for given primitive types -let isOperatorValidForType = (op: AST.operator, pt: primitiveType): bool => { - switch op { - | Eq | Neq => true // equality works on all types - | Gt | Lt | Gte | Lte => isComparablePrimitive(pt) - | Like | Contains | Matches => pt == StringType - } -} diff --git a/verisimdb/stapeln.toml b/verisimdb/stapeln.toml deleted file mode 100644 index 66d47bde..00000000 --- a/verisimdb/stapeln.toml +++ /dev/null @@ -1,140 +0,0 @@ -# SPDX-License-Identifier: MPL-2.0 -# SPDX-FileCopyrightText: 2026 Jonathan D.A. Jewell (hyperpolymath) -# -# stapeln.toml — Layer-based container build configuration for VeriSimDB -# -# stapeln builds containers as composable layers (German: "to stack"). -# Each layer is independently cacheable, verifiable, and signable. - -[metadata] -name = "verisimdb" -version = "0.1.0-alpha" -description = "6-core multimodal database with self-normalization" -author = "Jonathan D.A. Jewell <j.d.a.jewell@open.ac.uk>" -license = "PMPL-1.0-or-later" -registry = "ghcr.io/hyperpolymath" - -[build] -containerfile = "container/Containerfile" -context = "." -runtime = "podman" - -# ── Layer Definitions ────────────────────────────────────────── -# Layers are built bottom-up. Each layer extends the previous. - -[layers.base] -description = "Chainguard Wolfi minimal base" -from = "cgr.dev/chainguard/wolfi-base:latest" -cache = true -verify = true - -[layers.rust-toolchain] -description = "Rust compiler and build dependencies" -extends = "base" -packages = ["rust", "pkgconf", "build-base", "protoc", "clang-19"] -env = { LIBCLANG_PATH = "/usr/lib", PROTOC = "/usr/bin/protoc" } -cache = true - -[layers.rust-deps] -description = "Cargo dependency fetch and compile (no source)" -extends = "rust-toolchain" -commands = [ - "cargo fetch --locked", - "cargo build --release --workspace --lib 2>/dev/null || true", -] -cache-key = "Cargo.lock" -cache = true - -[layers.rust-build] -description = "VeriSimDB Rust core compilation" -extends = "rust-deps" -commands = ["cargo build --release -p verisim-api"] -artifacts = [ - { src = "target/release/verisim-api", dst = "/app/verisim-api" }, -] - -[layers.elixir-toolchain] -description = "Elixir/OTP 27 runtime and build tools" -extends = "base" -packages = ["erl27-elixir-1.18", "erlang-27", "erlang-27-dev", "git", "build-base"] -cache = true - -[layers.elixir-deps] -description = "Mix dependency fetch and compile" -extends = "elixir-toolchain" -env = { MIX_ENV = "prod" } -commands = [ - "mix local.hex --force", - "mix local.rebar --force", - "mix deps.get --only prod", - "mix compile", -] -cache-key = "elixir-orchestration/mix.lock" -cache = true - -[layers.elixir-build] -description = "Elixir OTP release build" -extends = "elixir-deps" -commands = ["mix release verisim"] -artifacts = [ - { src = "_build/prod/rel/verisim", dst = "/app/elixir/" }, -] - -[layers.runtime] -description = "Minimal runtime with Rust binary + Elixir release" -from = "cgr.dev/chainguard/wolfi-base:latest" -packages = ["ca-certificates", "curl", "libstdc++", "ncurses"] -copy-from = [ - { layer = "rust-build", src = "/app/verisim-api", dst = "/app/verisim-api" }, - { layer = "elixir-build", src = "/app/elixir/", dst = "/app/elixir/" }, -] -entrypoint = ["/app/entrypoint.sh"] -user = "verisim" -expose = [8080] - -# ── Security ─────────────────────────────────────────────────── - -[security] -non-root = true -read-only-root = false -no-new-privileges = true -cap-drop = ["ALL"] -seccomp-profile = "default" - -[security.signing] -algorithm = "ML-DSA-87" -provider = "cerro-torre" - -[security.sbom] -format = "spdx-json" -output = "sbom.spdx.json" -include-deps = true - -[security.attestation] -slsa-level = 3 -provenance = true -reproducible = true - -# ── Verification ─────────────────────────────────────────────── - -[verify] -vordr = true -svalinn = true -scan-on-build = true -fail-on = ["critical", "high"] - -# ── Targets ──────────────────────────────────────────────────── -# Named build targets for different deployment scenarios. - -[targets.development] -layers = ["base", "rust-toolchain", "rust-build", "elixir-toolchain", "elixir-build"] -env = { RUST_LOG = "debug", MIX_ENV = "dev" } - -[targets.production] -layers = ["runtime"] -env = { RUST_LOG = "info", MIX_ENV = "prod" } - -[targets.test] -layers = ["base", "rust-toolchain", "rust-build", "elixir-toolchain", "elixir-build"] -commands = ["cargo test", "mix test"] -env = { RUST_LOG = "debug", MIX_ENV = "test" } diff --git a/verisimdb/tests/integration_test.rs b/verisimdb/tests/integration_test.rs deleted file mode 100644 index 73574d91..00000000 --- a/verisimdb/tests/integration_test.rs +++ /dev/null @@ -1,344 +0,0 @@ -// SPDX-License-Identifier: MPL-2.0 -//! Integration Tests for VeriSimDB -//! -//! Tests the full stack: Rust stores → HTTP API → Elixir orchestration → VCL - -use verisim_api::ConcreteOctadStore; -use verisim_document::TantivyDocumentStore; -use verisim_octad::{ - OctadConfig, OctadDocumentInput, OctadGraphInput, OctadId, OctadInput, OctadSemanticInput, - OctadStore, OctadVectorInput, InMemoryOctadStore, OctadSnapshot, -}; -use verisim_provenance::InMemoryProvenanceStore; -use verisim_semantic::InMemorySemanticStore; -use verisim_spatial::InMemorySpatialStore; -use verisim_temporal::InMemoryVersionStore; -use verisim_tensor::InMemoryTensorStore; -use verisim_graph::SimpleGraphStore; -use verisim_vector::{BruteForceVectorStore, DistanceMetric}; - -use std::collections::HashMap; -use std::sync::Arc; - -/// Helper to create a test octad store with all eight modality stores. -fn create_test_store() -> ConcreteOctadStore { - let graph = Arc::new(SimpleGraphStore::in_memory().unwrap()); - let vector = Arc::new(BruteForceVectorStore::new(384, DistanceMetric::Cosine)); - let document = Arc::new(TantivyDocumentStore::in_memory().unwrap()); - let tensor = Arc::new(InMemoryTensorStore::new()); - let semantic = Arc::new(InMemorySemanticStore::new()); - let temporal = Arc::new(InMemoryVersionStore::<OctadSnapshot>::new()); - let provenance = Arc::new(InMemoryProvenanceStore::new()); - let spatial = Arc::new(InMemorySpatialStore::new()); - - InMemoryOctadStore::new( - OctadConfig::default(), - graph, - vector, - document, - tensor, - semantic, - temporal, - provenance, - spatial, - ) -} - -#[tokio::test] -async fn test_octad_create_and_retrieve() { - let store = create_test_store(); - - let embedding = vec![0.1f32; 384]; - let input = OctadInput { - document: Some(OctadDocumentInput { - title: "Test Document".to_string(), - body: "This is a test document for VeriSimDB integration testing.".to_string(), - fields: HashMap::new(), - }), - vector: Some(OctadVectorInput { - embedding: embedding.clone(), - model: None, - }), - ..Default::default() - }; - - let octad = store.create(input).await.unwrap(); - let octad_id = octad.id; - - // Retrieve the octad - let snapshot = store.get(&octad_id).await.unwrap().unwrap(); - - // Verify document modality - assert_eq!(snapshot.document.as_ref().unwrap().title, "Test Document"); - - // Verify vector modality (embedding field) - assert_eq!(snapshot.embedding.as_ref().unwrap().vector.len(), 384); - - // Verify modality status flags - assert!(snapshot.status.modality_status.document); - assert!(snapshot.status.modality_status.vector); -} - -#[tokio::test] -async fn test_cross_modal_consistency() { - let store = create_test_store(); - - let input = OctadInput { - document: Some(OctadDocumentInput { - title: "Consistency Test".to_string(), - body: "Testing cross-modal consistency.".to_string(), - fields: HashMap::new(), - }), - vector: Some(OctadVectorInput { - embedding: vec![0.2f32; 384], - model: None, - }), - semantic: Some(OctadSemanticInput { - types: vec!["https://example.org/Document".to_string()], - properties: HashMap::new(), - }), - ..Default::default() - }; - - let octad = store.create(input).await.unwrap(); - let octad_id = octad.id; - - // Verify all supplied modalities are present - let snapshot = store.get(&octad_id).await.unwrap().unwrap(); - assert!(snapshot.document.is_some()); - assert!(snapshot.embedding.is_some()); - assert!(snapshot.semantic.is_some()); -} - -#[tokio::test] -async fn test_drift_detection() { - // Drift detection runs through the DriftDetector component, not directly - // on the octad store. This test verifies that creating an octad with - // document + vector succeeds and leaves both modalities populated. - let store = create_test_store(); - - let input = OctadInput { - document: Some(OctadDocumentInput { - title: "Drift Test".to_string(), - body: "Testing drift detection.".to_string(), - fields: HashMap::new(), - }), - vector: Some(OctadVectorInput { - embedding: vec![0.3f32; 384], - model: None, - }), - ..Default::default() - }; - - let octad = store.create(input).await.unwrap(); - let snapshot = store.get(&octad.id).await.unwrap().unwrap(); - assert!(snapshot.status.modality_status.document); - assert!(snapshot.status.modality_status.vector); -} - -#[tokio::test] -async fn test_vector_similarity_search() { - let store = create_test_store(); - - for i in 0..10 { - let mut embedding = vec![0.0f32; 384]; - embedding[0] = i as f32 / 10.0; - - let input = OctadInput { - document: Some(OctadDocumentInput { - title: format!("Document {}", i), - body: format!("Content {}", i), - fields: HashMap::new(), - }), - vector: Some(OctadVectorInput { - embedding, - model: None, - }), - ..Default::default() - }; - - store.create(input).await.unwrap(); - } - - let query_embedding = vec![0.5f32; 384]; - let results = store.search_similar(&query_embedding, 5).await.unwrap(); - - assert!(!results.is_empty()); - assert!(results.len() <= 5); -} - -#[tokio::test] -async fn test_fulltext_search() { - let store = create_test_store(); - - let documents = vec![ - ("Machine Learning Basics", - "Introduction to machine learning algorithms and neural networks."), - ("Deep Learning Tutorial", - "Advanced deep learning techniques including transformers."), - ("AI Safety Research", - "Research on alignment and safety of artificial intelligence systems."), - ]; - - for (title, body) in documents { - let input = OctadInput { - document: Some(OctadDocumentInput { - title: title.to_string(), - body: body.to_string(), - fields: HashMap::new(), - }), - vector: Some(OctadVectorInput { - embedding: vec![0.5f32; 384], - model: None, - }), - ..Default::default() - }; - - store.create(input).await.unwrap(); - } - - let results = store.search_text("machine learning", 10).await.unwrap(); - assert!(!results.is_empty()); -} - -#[tokio::test] -async fn test_temporal_versioning() { - let store = create_test_store(); - - let input = OctadInput { - document: Some(OctadDocumentInput { - title: "Version Test".to_string(), - body: "Initial version".to_string(), - fields: HashMap::new(), - }), - ..Default::default() - }; - - let octad = store.create(input).await.unwrap(); - let octad_id = octad.id; - - let v1 = store.get(&octad_id).await.unwrap().unwrap(); - // Version counter starts at 1 after creation - assert!(v1.status.version >= 1); -} - -#[tokio::test] -async fn test_graph_relationships() { - let store = create_test_store(); - - let input1 = OctadInput { - document: Some(OctadDocumentInput { - title: "Paper 1".to_string(), - body: "First research paper.".to_string(), - fields: HashMap::new(), - }), - ..Default::default() - }; - - let id1 = store.create(input1).await.unwrap().id; - - let input2 = OctadInput { - document: Some(OctadDocumentInput { - title: "Paper 2".to_string(), - body: "Second research paper.".to_string(), - fields: HashMap::new(), - }), - graph: Some(OctadGraphInput { - relationships: vec![("cites".to_string(), id1.0.clone())], - }), - ..Default::default() - }; - - let id2 = store.create(input2).await.unwrap().id; - - // Both entities were created - assert!(store.get(&id1).await.unwrap().is_some()); - assert!(store.get(&id2).await.unwrap().is_some()); -} - -#[tokio::test] -async fn test_normalization() { - let store = create_test_store(); - - let input = OctadInput { - document: Some(OctadDocumentInput { - title: "Normalization Test".to_string(), - body: "Testing self-normalization.".to_string(), - fields: HashMap::new(), - }), - vector: Some(OctadVectorInput { - embedding: vec![0.4f32; 384], - model: None, - }), - ..Default::default() - }; - - let octad = store.create(input).await.unwrap(); - let snapshot = store.get(&octad.id).await.unwrap().unwrap(); - // Both supplied modalities should be present; normalizer can regenerate from either - assert!(snapshot.status.modality_status.document); - assert!(snapshot.status.modality_status.vector); -} - -#[tokio::test] -async fn test_multi_modal_query() { - let store = create_test_store(); - - let input = OctadInput { - document: Some(OctadDocumentInput { - title: "Multi-modal Test".to_string(), - body: "Testing multi-modal queries with semantic types.".to_string(), - fields: HashMap::new(), - }), - vector: Some(OctadVectorInput { - embedding: vec![0.6f32; 384], - model: None, - }), - semantic: Some(OctadSemanticInput { - types: vec!["https://example.org/Document".to_string()], - properties: HashMap::new(), - }), - ..Default::default() - }; - - let octad = store.create(input).await.unwrap(); - let snapshot = store.get(&octad.id).await.unwrap().unwrap(); - assert!(snapshot.status.modality_status.document); - assert!(snapshot.status.modality_status.vector); - assert!(snapshot.status.modality_status.semantic); -} - -#[tokio::test] -async fn test_concurrent_operations() { - let store = create_test_store(); - - let mut handles = vec![]; - - for i in 0..10 { - let store_clone = store.clone(); - let handle = tokio::spawn(async move { - let input = OctadInput { - document: Some(OctadDocumentInput { - title: format!("Concurrent {}", i), - body: format!("Testing concurrency {}", i), - fields: HashMap::new(), - }), - vector: Some(OctadVectorInput { - embedding: vec![i as f32 / 10.0; 384], - model: None, - }), - ..Default::default() - }; - - store_clone.create(input).await - }); - - handles.push(handle); - } - - let results = futures::future::join_all(handles).await; - - for result in results { - assert!(result.unwrap().is_ok()); - } -} diff --git a/verisimdb/vcl-bridge/vcl_parser_port.js b/verisimdb/vcl-bridge/vcl_parser_port.js deleted file mode 100644 index 9f6f53a0..00000000 --- a/verisimdb/vcl-bridge/vcl_parser_port.js +++ /dev/null @@ -1,594 +0,0 @@ -#!/usr/bin/env -S deno run --allow-read -// SPDX-License-Identifier: MPL-2.0 -// VCL Parser Port — stdin/stdout JSON bridge for Elixir -// -// Reads JSON messages (one per line) from stdin, parses VCL queries -// using the compiled ReScript VCLParser, and writes JSON results to stdout. -// -// Protocol: newline-delimited JSON -// Request: {"id": 1, "action": "parse", "query": "SELECT ..."} -// Response: {"id": 1, "ok": {...ast...}} or {"id": 1, "error": "..."} - -// Import compiled ReScript VCL parser and type checker. -// The compiled JS output will be at ../src/vcl/*.res.mjs (ReScript compiler output). -let VCLParser; -let VCLTypeChecker; -let VCLBidir; - -try { - VCLParser = await import("../src/vcl/VCLParser.res.mjs"); -} catch (_e) { - try { - // Legacy .bs.js suffix (older ReScript versions) - VCLParser = await import("../src/vcl/VCLParser.bs.js"); - } catch (_e2) { - VCLParser = null; - } -} - -try { - VCLTypeChecker = await import("../src/vcl/VCLTypeChecker.res.mjs"); -} catch (_e) { - try { - VCLTypeChecker = await import("../src/vcl/VCLTypeChecker.bs.js"); - } catch (_e2) { - VCLTypeChecker = null; - } -} - -try { - VCLBidir = await import("../src/vcl/VCLBidir.res.mjs"); -} catch (_e) { - try { - VCLBidir = await import("../src/vcl/VCLBidir.bs.js"); - } catch (_e2) { - VCLBidir = null; - } -} - -/** - * Minimal VCL parser fallback (when compiled ReScript is unavailable). - * Produces AST compatible with VCLParser.res types. - */ -function fallbackParse(query) { - const tokens = query.trim().split(/\s+/); - let pos = 0; - - function peek() { return tokens[pos] || null; } - function advance() { return tokens[pos++] || null; } - function expect(t) { - const got = advance(); - if (got?.toUpperCase() !== t.toUpperCase()) { - throw new Error(`Expected "${t}", got "${got}" at position ${pos}`); - } - return got; - } - - // SELECT clause (extended: supports MODALITY.field projections and aggregates) - expect("SELECT"); - const modalities = []; - const projections = []; - const aggregates = []; - const MODALITY_NAMES = ["GRAPH","VECTOR","TENSOR","SEMANTIC","DOCUMENT","TEMPORAL","PROVENANCE","SPATIAL"]; - const AGG_FUNCS = ["COUNT","SUM","AVG","MIN","MAX"]; - - while (peek() && peek().toUpperCase() !== "FROM") { - const tok = peek()?.replace(/,/g, ""); - const tokUp = tok?.toUpperCase(); - - // Aggregate: COUNT(*) or FUNC(MODALITY.field) - if (AGG_FUNCS.includes(tokUp)) { - const func = advance().replace(/,/g, "").toUpperCase(); - if (peek() === "(" || peek()?.startsWith("(")) { - let inner = advance().replace(/[()]/g, ""); - if (!inner && peek()) inner = advance().replace(/[()]/g, ""); - // consume closing ) if separate token - if (peek() === ")") advance(); - - if (inner === "*") { - aggregates.push({ TAG: "CountAll" }); - } else if (inner.includes(".")) { - const [mod, field] = inner.split(".", 2); - aggregates.push({ - TAG: "AggregateField", - func: func.charAt(0) + func.slice(1).toLowerCase(), - modality: mod.charAt(0) + mod.slice(1).toLowerCase(), - field: field - }); - const modName = mod.charAt(0) + mod.slice(1).toLowerCase(); - if (!modalities.includes(modName)) modalities.push(modName); - } - } - continue; - } - - // MODALITY.field projection - if (tok?.includes(".")) { - const [mod, field] = tok.split(".", 2); - if (MODALITY_NAMES.includes(mod.toUpperCase())) { - advance(); // consume the token - projections.push({ - modality: mod.charAt(0).toUpperCase() + mod.slice(1).toLowerCase(), - field: field - }); - const modName = mod.charAt(0).toUpperCase() + mod.slice(1).toLowerCase(); - if (!modalities.includes(modName)) modalities.push(modName); - continue; - } - } - - // Bare modality or wildcard - advance(); - if (tokUp === "*") modalities.push("All"); - else if (MODALITY_NAMES.includes(tokUp)) { - modalities.push(tokUp.charAt(0) + tokUp.slice(1).toLowerCase()); - } - } - - // FROM clause - expect("FROM"); - let source; - const sourceType = advance()?.toUpperCase(); - if (sourceType === "HEXAD") { - source = { TAG: "Hexad", _0: advance() }; - } else if (sourceType === "FEDERATION") { - const pattern = advance(); - let driftPolicy = null; - if (peek()?.toUpperCase() === "WITH") { - advance(); // WITH - expect("DRIFT"); - const policy = advance()?.toUpperCase(); - driftPolicy = policy; - } - source = { TAG: "Federation", _0: pattern, _1: driftPolicy }; - } else if (sourceType === "STORE") { - source = { TAG: "Store", _0: advance() }; - } else { - throw new Error(`Unknown source type: ${sourceType}`); - } - - // WHERE clause (simplified) - let whereClause = null; - if (peek()?.toUpperCase() === "WHERE") { - advance(); // WHERE - const condTokens = []; - while (peek() && !["PROOF","LIMIT","OFFSET"].includes(peek()?.toUpperCase())) { - condTokens.push(advance()); - } - whereClause = { TAG: "Raw", _0: condTokens.join(" ") }; - } - - // GROUP BY clause - let groupBy = null; - if (peek()?.toUpperCase() === "GROUP") { - advance(); // GROUP - expect("BY"); - groupBy = []; - while (peek() && !["HAVING","PROOF","ORDER","LIMIT","OFFSET"].includes(peek()?.toUpperCase())) { - const tok = advance().replace(/,/g, ""); - if (tok.includes(".")) { - const [mod, field] = tok.split(".", 2); - groupBy.push({ - modality: mod.charAt(0).toUpperCase() + mod.slice(1).toLowerCase(), - field: field - }); - } - } - } - - // HAVING clause - let having = null; - if (peek()?.toUpperCase() === "HAVING") { - advance(); // HAVING - const havingTokens = []; - while (peek() && !["PROOF","ORDER","LIMIT","OFFSET"].includes(peek()?.toUpperCase())) { - havingTokens.push(advance()); - } - having = { TAG: "Raw", _0: havingTokens.join(" ") }; - } - - // PROOF clause (multi-proof: PROOF spec AND spec AND spec) - let proof = null; - if (peek()?.toUpperCase() === "PROOF") { - advance(); // PROOF - const proofSpecs = []; - while (peek() && !["ORDER","LIMIT","OFFSET"].includes(peek()?.toUpperCase())) { - const proofType = advance(); - const contractRaw = advance()?.replace(/[()]/g, "") || ""; - proofSpecs.push({ proofType, contractName: contractRaw }); - // Check for AND to chain multiple proofs - if (peek()?.toUpperCase() === "AND") { - advance(); // AND - } else { - break; - } - } - proof = proofSpecs.length > 0 ? proofSpecs : null; - } - - // ORDER BY clause - let orderBy = null; - if (peek()?.toUpperCase() === "ORDER") { - advance(); // ORDER - expect("BY"); - orderBy = []; - while (peek() && !["LIMIT","OFFSET"].includes(peek()?.toUpperCase())) { - const tok = advance().replace(/,/g, ""); - if (tok.includes(".")) { - const [mod, field] = tok.split(".", 2); - let direction = "Asc"; - if (peek()?.toUpperCase() === "ASC") { advance(); direction = "Asc"; } - else if (peek()?.toUpperCase() === "DESC") { advance(); direction = "Desc"; } - orderBy.push({ - field: { - modality: mod.charAt(0).toUpperCase() + mod.slice(1).toLowerCase(), - field: field - }, - direction: direction - }); - } - } - } - - // LIMIT - let limit = null; - if (peek()?.toUpperCase() === "LIMIT") { - advance(); - limit = parseInt(advance(), 10); - } - - // OFFSET - let offset = null; - if (peek()?.toUpperCase() === "OFFSET") { - advance(); - offset = parseInt(advance(), 10); - } - - return { - modalities, - projections: projections.length > 0 ? projections : null, - aggregates: aggregates.length > 0 ? aggregates : null, - source, - where: whereClause, - groupBy, - having, - proof, - orderBy, - limit, - offset, - }; -} - -/** - * Minimal VCL mutation parser fallback (INSERT / UPDATE / DELETE). - * Produces AST compatible with VCLParser.res mutation types. - */ -function fallbackParseMutation(input) { - const tokens = input.trim().split(/\s+/); - let pos = 0; - - function peek() { return tokens[pos] || null; } - function advance() { return tokens[pos++] || null; } - function expect(t) { - const got = advance(); - if (got?.toUpperCase() !== t.toUpperCase()) { - throw new Error(`Expected "${t}", got "${got}" at position ${pos}`); - } - return got; - } - - const cmd = advance()?.toUpperCase(); - - if (cmd === "INSERT") { - expect("HEXAD"); - expect("WITH"); - - const MODALITY_NAMES = ["GRAPH","VECTOR","TENSOR","SEMANTIC","DOCUMENT","TEMPORAL","PROVENANCE","SPATIAL"]; - const modalityData = []; - - while (peek() && MODALITY_NAMES.includes(peek()?.toUpperCase())) { - const mod = advance().toUpperCase(); - // Collect everything in parens - if (peek()?.startsWith("(")) { - let raw = advance(); - // Collect until closing paren - while (raw && !raw.endsWith(")")) { - raw += " " + advance(); - } - raw = raw.replace(/^\(/, "").replace(/\)$/, ""); - modalityData.push({ modality: mod, raw }); - } - // Skip comma between modality data - if (peek() === ",") advance(); - } - - // Optional PROOF - let proof = null; - if (peek()?.toUpperCase() === "PROOF") { - advance(); - const proofSpecs = []; - while (peek()) { - const proofType = advance(); - const contractRaw = advance()?.replace(/[()]/g, "") || ""; - proofSpecs.push({ proofType, contractName: contractRaw }); - if (peek()?.toUpperCase() === "AND") { advance(); } else { break; } - } - proof = proofSpecs.length > 0 ? proofSpecs : null; - } - - return { TAG: "Insert", modalities: modalityData, proof }; - - } else if (cmd === "UPDATE") { - expect("HEXAD"); - const hexadId = advance(); - expect("SET"); - - const sets = []; - while (peek() && peek()?.toUpperCase() !== "PROOF") { - const field = advance()?.replace(/,/g, ""); - if (peek() === "=") { - advance(); // = - const value = advance()?.replace(/,/g, ""); - sets.push({ field, value }); - } else { - break; - } - } - - let proof = null; - if (peek()?.toUpperCase() === "PROOF") { - advance(); - const proofSpecs = []; - while (peek()) { - const proofType = advance(); - const contractRaw = advance()?.replace(/[()]/g, "") || ""; - proofSpecs.push({ proofType, contractName: contractRaw }); - if (peek()?.toUpperCase() === "AND") { advance(); } else { break; } - } - proof = proofSpecs.length > 0 ? proofSpecs : null; - } - - return { TAG: "Update", hexadId, sets, proof }; - - } else if (cmd === "DELETE") { - expect("HEXAD"); - const hexadId = advance(); - - let proof = null; - if (peek()?.toUpperCase() === "PROOF") { - advance(); - const proofSpecs = []; - while (peek()) { - const proofType = advance(); - const contractRaw = advance()?.replace(/[()]/g, "") || ""; - proofSpecs.push({ proofType, contractName: contractRaw }); - if (peek()?.toUpperCase() === "AND") { advance(); } else { break; } - } - proof = proofSpecs.length > 0 ? proofSpecs : null; - } - - return { TAG: "Delete", hexadId, proof }; - - } else { - throw new Error(`Expected INSERT, UPDATE, or DELETE, got "${cmd}"`); - } -} - -/** - * Parse a VCL statement (query or mutation) using fallback parser. - */ -function fallbackParseStatement(input) { - const trimmed = input.trim(); - const firstWord = trimmed.split(/\s+/)[0]?.toUpperCase(); - - if (["INSERT", "UPDATE", "DELETE"].includes(firstWord)) { - return { TAG: "Mutation", _0: fallbackParseMutation(trimmed) }; - } else { - return { TAG: "Query", _0: fallbackParse(trimmed) }; - } -} - -/** - * Parse a VCL query using the ReScript parser or fallback. - */ -function parseQuery(query, action) { - if (VCLParser) { - let result; - switch (action) { - case "parse_slipstream": - result = VCLParser.parseSlipstream(query); - break; - case "parse_dependent": - result = VCLParser.parseDependentType(query); - break; - case "parse_mutation": - result = VCLParser.parseMutation(query); - break; - case "parse_statement": - result = VCLParser.parseStatement(query); - break; - default: - result = VCLParser.parse(query); - } - // ReScript Result type: { TAG: "Ok", _0: value } or { TAG: "Error", _0: error } - if (result.TAG === "Ok") return { ok: result._0 }; - return { error: result._0.message || JSON.stringify(result._0) }; - } - - // Fallback parser - switch (action) { - case "parse_mutation": - return { ok: fallbackParseMutation(query) }; - case "parse_statement": - return { ok: fallbackParseStatement(query) }; - default: - return { ok: fallbackParse(query) }; - } -} - -// --------------------------------------------------------------------------- -// Type checking: invoke ReScript VCLBidir or extract obligations from AST -// --------------------------------------------------------------------------- - -/** - * Type-check a parsed VCL-UT AST. - * - * If compiled ReScript is available, delegates to VCLBidir.synthesizeQuery - * for full bidirectional type inference. Otherwise, extracts proof obligations - * directly from the AST's proof clause — a sound but less precise fallback. - * - * @param {object} ast - Parsed VCL query AST - * @returns {{ ok: object } | { error: string }} - Type info or error - */ -function typecheckAST(ast) { - // Try the full ReScript type checker first - if (VCLBidir) { - try { - const result = VCLBidir.synthesizeQuery(ast); - if (result.TAG === "Ok") { - return { ok: result._0 }; - } - return { error: result._0?.message || JSON.stringify(result._0) }; - } catch (e) { - // If VCLBidir crashes, fall through to the lightweight fallback - } - } - - if (VCLTypeChecker) { - try { - const result = VCLTypeChecker.checkQuery(ast); - if (result.TAG === "Ok") { - return { ok: result._0 }; - } - return { error: result._0?.message || JSON.stringify(result._0) }; - } catch (e) { - // Fall through to lightweight fallback - } - } - - // Lightweight fallback: extract proof obligations from the AST directly. - // This does NOT perform type inference — it only parses the PROOF clause - // into a structured list of obligations with composition strategy. - return { ok: fallbackTypecheck(ast) }; -} - -/** - * Fallback type checking: extract proof obligations from an AST's proof clause. - * - * Returns a structure matching what VCLBidir.synthesizeQuery would produce: - * { proof_obligations, composition_strategy, inferred_types } - */ -function fallbackTypecheck(ast) { - const proofSpecs = ast.proof || ast.proofSpecs || null; - const obligations = []; - let compositionStrategy = "conjunction"; - - if (Array.isArray(proofSpecs)) { - for (const spec of proofSpecs) { - const proofType = (spec.proofType || spec.TAG || "unknown").toUpperCase(); - const contractName = spec.contractName || spec.contract || spec._0 || null; - - // Map proof types to obligation records - const PROOF_TIMES = { - "EXISTENCE": 50, "CITATION": 100, "ACCESS": 150, - "INTEGRITY": 200, "PROVENANCE": 300, "CUSTOM": 500, - "ZKP": 400, "PROVEN": 350, "SANCTIFY": 200 - }; - - obligations.push({ - type: proofType.toLowerCase(), - contract: contractName, - estimated_time_ms: PROOF_TIMES[proofType] || 200, - privacy_level: spec.privacy_level || "public" - }); - } - - // Multiple obligations default to conjunction (all must pass) - compositionStrategy = obligations.length > 1 ? "conjunction" : "independent"; - } else if (proofSpecs && typeof proofSpecs === "object" && proofSpecs.raw) { - // Raw proof string from built-in parser: "EXISTENCE(abc-123)" - const raw = proofSpecs.raw; - const matches = raw.matchAll(/(\w+)\(([^)]*)\)/g); - for (const match of matches) { - obligations.push({ - type: match[1].toLowerCase(), - contract: match[2] || null, - estimated_time_ms: 200, - privacy_level: "public" - }); - } - compositionStrategy = obligations.length > 1 ? "conjunction" : "independent"; - } - - // Extract modality information for inferred_types - const modalities = ast.modalities || []; - const inferredTypes = {}; - for (const mod of modalities) { - const modName = typeof mod === "string" ? mod.toLowerCase() : String(mod); - inferredTypes[modName] = { _type: "modality_data" }; - } - - return { - proof_obligations: obligations, - composition_strategy: compositionStrategy, - inferred_types: inferredTypes - }; -} - -// --------------------------------------------------------------------------- -// Main loop: read lines from stdin, process, write to stdout -// --------------------------------------------------------------------------- - -const decoder = new TextDecoder(); -const encoder = new TextEncoder(); - -// Deno stdin reading -const buf = new Uint8Array(1_048_576); -let buffer = ""; - -async function main() { - const stdin = Deno.stdin; - const stdout = Deno.stdout; - - while (true) { - const n = await stdin.read(buf); - if (n === null) break; // EOF - - buffer += decoder.decode(buf.subarray(0, n)); - - // Process complete lines - let newlineIdx; - while ((newlineIdx = buffer.indexOf("\n")) !== -1) { - const line = buffer.slice(0, newlineIdx).trim(); - buffer = buffer.slice(newlineIdx + 1); - - if (!line) continue; - - try { - const request = JSON.parse(line); - const { id, action, query, ast } = request; - - try { - let result; - if (action === "typecheck") { - // Type checking receives an AST (not a query string) - result = typecheckAST(ast || query); - } else { - result = parseQuery(query, action); - } - const response = JSON.stringify({ id, ...result }) + "\n"; - await stdout.write(encoder.encode(response)); - } catch (e) { - const response = JSON.stringify({ id, error: e.message }) + "\n"; - await stdout.write(encoder.encode(response)); - } - } catch (e) { - // Malformed JSON — skip - const response = JSON.stringify({ id: 0, error: `Invalid JSON: ${e.message}` }) + "\n"; - await stdout.write(encoder.encode(response)); - } - } - } -} - -main().catch(console.error); diff --git a/verisimdb/verification/PROOF-STATUS.md b/verisimdb/verification/PROOF-STATUS.md deleted file mode 100644 index 941b15de..00000000 --- a/verisimdb/verification/PROOF-STATUS.md +++ /dev/null @@ -1,55 +0,0 @@ -# Proof Status — VeriSimDB - -<!-- SPDX-License-Identifier: CC-BY-SA-4.0 --> -<!-- Copyright (c) 2026 Jonathan D.A. Jewell (hyperpolymath) --> - -Tracking status of formal verification proofs for the VeriSimDB cross-modality octad database. - -Obligations catalogued in `developer-ecosystem/standards/docs/proofs/spec-templates/T1-critical/verisimdb.md`. - -## Policy - -- **No** `believe_me`, `assert_total`, `postulate`, `sorry`, `Admitted` in fixable positions. -- All proofs are constructive. -- `%default total` in all Idris2 files. -- Idris2 for type-level / constructive proofs; TLA+ for concurrent + adversarial state-machine properties; Lean4 for the remaining numeric/induction-heavy obligations. - -## Proof Inventory - -### Idris2 - -| File | LOC | Properties | Obligation | Status | -|------|-----|------------|------------|--------| -| `verification/proofs/idris2/OctadCoherence.idr` | 327 | 8 modalities + 3 irredundant cross-modality invariants; `opPreservesCoherence` for all 6 Ops; `opsPreserveCoherence` for any Op sequence | **V1** | COMPLETE 2026-04-17 (`e2e5d9b`) | -| `verification/proofs/idris2/DriftMetric.idr` | ~140 | Hamming-distance drift is a metric: reflexivity, symmetry, triangle inequality; threshold-detection soundness + completeness | **V8** | COMPLETE 2026-04-17 (`182cc7c`) | -| `verification/proofs/idris2/ConnectorSafety.idr` | ~200 | Schema + ValidatedValue + total validator; every JSON → typed conversion is schema-checked by construction (Obj.magic elimination) | **V11** | COMPLETE 2026-04-17 (`95c3867`) | -| `verification/proofs/idris2/FFIOwnership.idr` | ~100 | `Owned` ownership-state token, non-null deref witness, and freeability restricted to `Alive` so double-free is uninhabited by type | **V12** | COMPLETE 2026-04-17 (Codex replay-checked) | - -### Lean4 - -| File | LOC | Properties | Obligation | Status | -|------|-----|------------|------------|--------| -| `verification/proofs/lean4/VCLTypeSoundness.lean` | ~300 | Query-core progress + preservation + multi-step soundness + synth/check soundness for `VCLBidir`-shaped typing | **V2** | COMPLETE 2026-04-17 (Codex replay-checked with `lean` + `lake build`) | -| `verification/proofs/lean4/VCLSubtyping.lean` | ~200 | Structural subtype transitivity + decidability for core VCL types | **V3** | COMPLETE 2026-04-17 (Codex replay-checked with `lean` + `lake build`) | -| `verification/proofs/lean4/RaftSafety.lean` | ~260 | Commit monotonicity, append isolation, WF preservation, single-node log matching | **V4** | COMPLETE 2026-04-17 (Codex replay-checked with `lean` + `lake build`) | -| `verification/proofs/lean4/WALIntegrity.lean` | ~200 | Sequence monotonicity, CRC validity model, replay compositionality, checkpoint idempotence | **V6** | COMPLETE 2026-04-17 (Codex replay-checked with `lean` + `lake build`) | - -### TLA+ - -| File | States | Depth | Properties | Obligation | Status | -|------|--------|-------|------------|------------|--------| -| `verification/proofs/tlaplus/OctadAtomicity.tla` | 134,160 | 19 | Atomicity (COMMITTED⇒all 8; ABORTED⇒none), NoObservablePartial, StatusMonotone, EveryTxnResolves (liveness) | **V5** | COMPLETE 2026-04-17 (`9d3dfd8`) | -| `verification/proofs/tlaplus/Normalizer.tla` | 84 | 2 | SourceIsMaximal, NormalizeIdempotent, PostStepNoDrift, FixedPointStable, Convergence (`<>[]~HasDrift`) | **V9** | COMPLETE 2026-04-17 (`81e1d52`) | -| `verification/proofs/tlaplus/Serializability.tla` | 33 | 7 | NoSharedLocks, LocksOnlyWhileActive, ActiveHoldsFullAccessSet, CommitLogInjective/Sound, NoConcurrentConflict, EveryTxnCommits | **V10** | COMPLETE 2026-04-17 (`94aa56f`, hardened `2390e52`) | - -Model-checked via `just verify-tlaplus` (Eclipse Temurin 21 JRE via podman-ephemeral container; no host Java install needed on Fedora Atomic). Regression-gated by `.github/workflows/verify-tlaplus.yml` on every push/PR touching `verisimdb/verification/proofs/tlaplus/**`. - -## Remaining Debt - -None in the tracked V1-V12 set as of 2026-04-17 replay checks. - -## Notes - -- **Lean4 replay check**: `VCLTypeSoundness.lean`, `VCLSubtyping.lean`, `RaftSafety.lean`, and `WALIntegrity.lean` all pass direct `lean` checks and package-level `lake build` under Lean 4.16.0 on this host. -- **V10 scenario**: the committed scenario uses partial conflicts (t1/t2 disjoint, t3 shares with both) to actually exercise concurrency — TLC reaches states with multiple simultaneously-ACTIVE transactions, so `NoConcurrentConflict` is non-vacuous. -- **Podman recipe**: `just verify-tlaplus` fetches `tla2tools.jar` to `~/.local/share/` on first run and executes TLC in an `eclipse-temurin:21-jre` container. Works identically on Fedora Atomic (this host) and any machine with podman. Host Java is used if present. diff --git a/verisimdb/verification/README.adoc b/verisimdb/verification/README.adoc deleted file mode 100644 index 3af1ace5..00000000 --- a/verisimdb/verification/README.adoc +++ /dev/null @@ -1,40 +0,0 @@ -// SPDX-License-Identifier: CC-BY-SA-4.0 -// @taxonomy: verification/index -= verisimdb — Verification Directory -:toc: - -== Overview - -Unified verification gateway for verisimdb. Each subdirectory is a symlink -to the actual location, providing a single entry point for all verification. - -== Structure - -|=== -| Directory | Target | Purpose - -| `proofs/` -| (local) -| Formal proofs and mathematical foundations - -| `tests/` -| ../tests -| Unit and integration tests - -| `conformance/` -| (local) -| Language/spec conformance test suite - -| `benchmarks/` -| ../benches -| Performance benchmarks - -| `fuzzing/` -| ../fuzz -| Fuzz testing targets -|=== - -== Usage - -All verification can be discovered from this directory. Symlinks point to -the actual directories so tools and CI can find everything from one place. diff --git a/verisimdb/verification/benchmarks b/verisimdb/verification/benchmarks deleted file mode 120000 index 34335628..00000000 --- a/verisimdb/verification/benchmarks +++ /dev/null @@ -1 +0,0 @@ -../benches \ No newline at end of file diff --git a/verisimdb/verification/fuzzing b/verisimdb/verification/fuzzing deleted file mode 120000 index ab541972..00000000 --- a/verisimdb/verification/fuzzing +++ /dev/null @@ -1 +0,0 @@ -../fuzz \ No newline at end of file diff --git a/verisimdb/verification/proofs/agda/ProvenanceChain.agda b/verisimdb/verification/proofs/agda/ProvenanceChain.agda deleted file mode 100644 index cd5d0123..00000000 --- a/verisimdb/verification/proofs/agda/ProvenanceChain.agda +++ /dev/null @@ -1,245 +0,0 @@ --- SPDX-License-Identifier: MPL-2.0 --- Copyright (c) 2026 Jonathan D.A. Jewell (hyperpolymath) <j.d.a.jewell@open.ac.uk> --- --- ProvenanceChain.agda — Formal proof of provenance chain immutability (V7). --- --- REQUIREMENTS-MASTER.md: V7 | Provenance chain immutability (hash chain, monotonic timestamps) | SEC | Ag | P1 --- --- Models rust-core/verisim-provenance/src/lib.rs: --- - ProvenanceRecord has: content, timestamp, parentHash, actorId, contentHash --- - Hash chain: parentHash[i+1] = contentHash[i] (each record links to its predecessor) --- - First record has parentHash = genesisHash (the "0"*64 sentinel) --- --- List representation: PREPEND order — the HEAD of the list is the NEWEST record. --- This makes chain extension (hc-cons) structurally natural. --- --- Properties proved: --- 1. HashChain predicate — well-formedness of a provenance chain --- 2. chain-tail-valid — tail of valid chain is valid --- 3. chain-head-consistent — head is self-consistent --- 4. chain-link — parentHash of head links to contentHash of predecessor --- 5. tamper-detected — changing content changes contentHash (by hash injectivity) --- 6. chain-tamper-breaks-link — a tampered record breaks the link from its successor --- 7. timestamps-non-decreasing — timestamps grow along the chain (oldest last) --- 8. genesis-unique — only the tail (oldest) record has genesisHash as parent --- 9. chain-extend — prepending a new record with a valid link extends the chain - -module ProvenanceChain where - -open import Data.Nat using (ℕ; zero; suc; _≤_; _<_; z≤n; s≤s) -open import Data.Nat.Properties using (≤-refl; ≤-trans; n<1+n; <⇒≤) -open import Data.List using (List; []; _∷_; length) -open import Data.Product using (_×_; _,_; proj₁; proj₂) -open import Data.Empty using (⊥; ⊥-elim) -open import Relation.Binary.PropositionalEquality - using (_≡_; refl; sym; trans; cong; _≢_) -open import Relation.Nullary using (¬_) - ------------------------------------------------------------------------- --- Section 1: Abstract hash model ------------------------------------------------------------------------- - --- Hashes are abstracted as ℕ. -Hash : Set -Hash = ℕ - --- Abstract hash function for a provenance record's fields. --- content × timestamp × parentHash × actorId → contentHash -postulate - hashRecord : Hash → ℕ → Hash → Hash → Hash - -- Collision resistance: equal outputs imply equal inputs. - hash-injective : ∀ {c₁ t₁ p₁ a₁ c₂ t₂ p₂ a₂} - → hashRecord c₁ t₁ p₁ a₁ ≡ hashRecord c₂ t₂ p₂ a₂ - → c₁ ≡ c₂ × t₁ ≡ t₂ × p₁ ≡ p₂ × a₁ ≡ a₂ - -- hashRecord never returns the genesis sentinel. - -- Holds in Rust because SHA-256 of non-empty input ≠ "0"*64. - hashRecord-not-genesis : ∀ (c t p a : Hash) → hashRecord c t p a ≢ 0 - --- Genesis parent hash: sentinel for the oldest record's parentHash. -genesisHash : Hash -genesisHash = 0 - ------------------------------------------------------------------------- --- Section 2: Provenance record model ------------------------------------------------------------------------- - -record ProvenanceRecord : Set where - constructor mkRecord - field - content : Hash - timestamp : ℕ - parentHash : Hash - actorId : Hash - contentHash : Hash -- = hashRecord content timestamp parentHash actorId - --- A record is self-consistent if its contentHash is correctly computed. -SelfConsistent : ProvenanceRecord → Set -SelfConsistent r = - ProvenanceRecord.contentHash r ≡ - hashRecord - (ProvenanceRecord.content r) - (ProvenanceRecord.timestamp r) - (ProvenanceRecord.parentHash r) - (ProvenanceRecord.actorId r) - ------------------------------------------------------------------------- --- Section 3: Hash chain predicate --- --- The list is stored NEWEST-FIRST (prepend order). --- HashChain (r_new ∷ r_old ∷ rs) means: --- - r_new is the most recently appended record --- - r_old is the record before r_new --- - r_new.parentHash = r_old.contentHash (link) --- - r_old.timestamp ≤ r_new.timestamp (monotone) ------------------------------------------------------------------------- - -data HashChain : List ProvenanceRecord → Set where - - hc-nil : HashChain [] - - hc-first : ∀ {r} - → SelfConsistent r - → ProvenanceRecord.parentHash r ≡ genesisHash - → HashChain (r ∷ []) - - -- Prepend r_new to an existing chain headed by r_old. - hc-cons : ∀ {r_new r_old rs} - → HashChain (r_old ∷ rs) -- existing chain - → SelfConsistent r_new - → ProvenanceRecord.parentHash r_new ≡ ProvenanceRecord.contentHash r_old - → ProvenanceRecord.timestamp r_old ≤ ProvenanceRecord.timestamp r_new - → HashChain (r_new ∷ r_old ∷ rs) -- extended chain - ------------------------------------------------------------------------- --- Section 4: Structural lemmas ------------------------------------------------------------------------- - --- V7 — LEMMA 1: The tail of a valid chain is valid. --- (Easy: hc-cons's first argument is exactly the tail chain.) -chain-tail-valid : ∀ {r rs} - → HashChain (r ∷ rs) - → HashChain rs -chain-tail-valid (hc-first _ _) = hc-nil -chain-tail-valid (hc-cons hc_tail _ _ _) = hc_tail - --- V7 — LEMMA 2: The head of a valid chain is self-consistent. -chain-head-consistent : ∀ {r rs} - → HashChain (r ∷ rs) - → SelfConsistent r -chain-head-consistent (hc-first sc _) = sc -chain-head-consistent (hc-cons _ sc _ _) = sc - --- V7 — LEMMA 3: The head's parentHash equals its predecessor's contentHash. -chain-link : ∀ {r_new r_old rs} - → HashChain (r_new ∷ r_old ∷ rs) - → ProvenanceRecord.parentHash r_new ≡ ProvenanceRecord.contentHash r_old -chain-link (hc-cons _ _ link _) = link - --- V7 — LEMMA 4: Timestamps are non-decreasing head-to-predecessor. --- Since the list is newest-first, the head timestamp is ≥ the second element's. -chain-ts-head-ge : ∀ {r_new r_old rs} - → HashChain (r_new ∷ r_old ∷ rs) - → ProvenanceRecord.timestamp r_old ≤ ProvenanceRecord.timestamp r_new -chain-ts-head-ge (hc-cons _ _ _ ts) = ts - ------------------------------------------------------------------------- --- Section 5: Immutability theorems ------------------------------------------------------------------------- - --- V7 — THEOREM 1: Two self-consistent records with the same contentHash --- have identical field values (injectivity of hashRecord). -same-hash-same-fields : - ∀ {r₁ r₂ : ProvenanceRecord} - → SelfConsistent r₁ - → SelfConsistent r₂ - → ProvenanceRecord.contentHash r₁ ≡ ProvenanceRecord.contentHash r₂ - → ProvenanceRecord.content r₁ ≡ ProvenanceRecord.content r₂ - × ProvenanceRecord.timestamp r₁ ≡ ProvenanceRecord.timestamp r₂ - × ProvenanceRecord.parentHash r₁ ≡ ProvenanceRecord.parentHash r₂ - × ProvenanceRecord.actorId r₁ ≡ ProvenanceRecord.actorId r₂ -same-hash-same-fields sc₁ sc₂ heq = - hash-injective (trans (sym sc₁) (trans heq sc₂)) - --- V7 — THEOREM 2: Changing the content field changes the contentHash. -tamper-detected : - ∀ {r r' : ProvenanceRecord} - → SelfConsistent r - → SelfConsistent r' - → ProvenanceRecord.content r ≢ ProvenanceRecord.content r' - → ProvenanceRecord.contentHash r ≢ ProvenanceRecord.contentHash r' -tamper-detected sc sc' hne heq = - hne (proj₁ (same-hash-same-fields sc sc' heq)) - --- V7 — THEOREM 3: Replacing the predecessor r_old with a tampered r_old' --- (different content) causes the head r_new's parentHash to no longer match. -chain-tamper-breaks-link : - ∀ {r_new r_old : ProvenanceRecord} {rs : List ProvenanceRecord} - → HashChain (r_new ∷ r_old ∷ rs) - → (r_old' : ProvenanceRecord) - → SelfConsistent r_old' - → ProvenanceRecord.content r_old ≢ ProvenanceRecord.content r_old' - → ProvenanceRecord.parentHash r_new ≢ ProvenanceRecord.contentHash r_old' -chain-tamper-breaks-link hc r_old' sc' hContentNe = - let hc-old : SelfConsistent _ - hc-old = chain-head-consistent (chain-tail-valid hc) - hLink = chain-link hc - hHashNe = tamper-detected hc-old sc' hContentNe - in λ link' → hHashNe (trans (sym hLink) link') - ------------------------------------------------------------------------- --- Section 6: Timestamp monotonicity ------------------------------------------------------------------------- - --- Extract timestamps in newest-first order. -chainTimestamps : List ProvenanceRecord → List ℕ -chainTimestamps [] = [] -chainTimestamps (r ∷ rs) = ProvenanceRecord.timestamp r ∷ chainTimestamps rs - --- V7 — THEOREM 4: Timestamps in a valid chain are non-decreasing from --- oldest (tail) to newest (head). Equivalently, chainTimestamps is a --- non-increasing sequence (newest first → values are ≥ previous). --- --- We prove the pointwise version: for any adjacent pair in the chain, --- the older record's timestamp ≤ the newer record's timestamp. -timestamps-non-decreasing : - ∀ {r_new r_old rs} - → HashChain (r_new ∷ r_old ∷ rs) - → ProvenanceRecord.timestamp r_old ≤ ProvenanceRecord.timestamp r_new -timestamps-non-decreasing = chain-ts-head-ge - ------------------------------------------------------------------------- --- Section 7: Genesis uniqueness ------------------------------------------------------------------------- - --- V7 — THEOREM 5: Only the tail (oldest) record has genesisHash as parentHash. --- Any non-first record's parentHash equals the contentHash of its predecessor, --- and contentHash is produced by hashRecord, which never returns genesisHash. -genesis-unique : - ∀ {r_new r_old rs} - → HashChain (r_new ∷ r_old ∷ rs) - → ProvenanceRecord.parentHash r_new ≢ genesisHash -genesis-unique {r_new} {r_old} hc hEq = - let hLink = chain-link hc - sc_old = chain-head-consistent (chain-tail-valid hc) - in hashRecord-not-genesis - (ProvenanceRecord.content r_old) - (ProvenanceRecord.timestamp r_old) - (ProvenanceRecord.parentHash r_old) - (ProvenanceRecord.actorId r_old) - (trans (sym sc_old) (trans (sym hLink) hEq)) - ------------------------------------------------------------------------- --- Section 8: Chain extension ------------------------------------------------------------------------- - --- V7 — THEOREM 6: Prepending a new well-formed record extends the chain. --- The caller supplies the link proof and the timestamp proof. -chain-extend : - ∀ {r_old rs} - → HashChain (r_old ∷ rs) - → (r_new : ProvenanceRecord) - → SelfConsistent r_new - → ProvenanceRecord.parentHash r_new ≡ ProvenanceRecord.contentHash r_old - → ProvenanceRecord.timestamp r_old ≤ ProvenanceRecord.timestamp r_new - → HashChain (r_new ∷ r_old ∷ rs) -chain-extend hc r_new sc link ts = hc-cons hc sc link ts diff --git a/verisimdb/verification/proofs/agda/ProvenanceChain.agdai b/verisimdb/verification/proofs/agda/ProvenanceChain.agdai deleted file mode 100644 index 212dc3eb5432d8bef63b3c9948876d5345610a09..0000000000000000000000000000000000000000 GIT binary patch literal 0 HcmV?d00001 literal 150906 zcmXuL2UJt(_C9<{8W2J#O6UO;ut5|Rq#jV3B1E0Bj3`kY#Y#~@q#aPOAw_1Cxgx}| zfPWp45i2UeLb*B;gsX@Jm5d^aC{iSl<lArfefO+&nKkQVpIx5)?DC#7^S`>TIG?hR zEoFqoPUZHSVDLY#5jo!G5C{Du#vmEBI&W;^U$MVF8~LtMr)i=_dPVPFT8>m7;GSvZ z|Le5u%!S|2Yh3PnWd;1dbA}(qldjA^cY4o`MIJlmCwg=hJO9C3^fBbwsSgA1F2%jM zq(fPK`*6lFLUW`n2gkY}eSCK3```b(-=fFS<zqV-`c_yiCeNr`GX3x*ie2}(cIJ58 zT*B?7C4^Y8)<|4!Geb4_?Xc#DAtka9b9Ci~p=)?p_m7$r-@fL3@cuBFx;ORF126b# zb{lUA&CBUI*mKELiqF1FP?x4lMfV-(-J{b#SPz|?a_z@VVVjAZKXV5abjfTsn_4yc zr;E>y2W~NQbmf=$D<umkcD9S`X%VqQ@?-6m!&x+I`Q8V#WHYs9|5j|LT&a0|tF!;} z2O_fv%fhqiT#3$ETpYf9=Jl-|<9z=xh;AGIB8m+%s-vc##jFI@`yaHi<e2B04b&w) z(BvZhOY<E~c4tTb<6biJ6|PU_tg)0Edg&%(Qs-Oh+LbO|+m8>C<_Y&ocFmOX8RBKC z#8-pYR`vCTW0|+{ES*roBQy1q3$|(38;i6ZEerE34WgGHtbbEU^SIXc?Lu)6(JsNE zOFB~Adf%_KBpvjwC%acV`o=1n7bGZW3zzepxFS{;hj-r(E4SKo%-c8i$)X|LoF}W@ zMM%ke)s@qb;;;8@t(+vqc@K>7+CnVLmPfXpaJ_b-a?Sb8cXT<qm<g+h*~Q7dUx<rc zmqmQf+uZA=`isuK_sztT*(Lj$+Qx!FaDB74UYaAtb89Q<H}0TH8PWavnCe}AjZfcA zZ>+Wj&%)P}9+{orxLE6E=bDks!o1`z59(ga*U(~2P@l3T%c|0J@Sd%^`4o(aIbE1X zGOXl=jN0v>VT#4dwU+XX+NR2hvzaqxA?{}UV#0|VH(6@S(G?f@VROAK9N*3JGPTDT zO(X+nl);mPyL_aVj7G3b@IdC9snk~i#lBF7b<(`+d@>X_f0O=|puYN3Z?_baf_I@8 z?oKW~WQggC>ARaglbpxI4u)T>kCb$-!}M_duf#0@KZ4^_eE1laz01&9#@sb25OXTt zcMR7RF-)BA`p`VZ+DElZ;m_hnwNGi*F6L7r^AMhOIZj8i|5zutIAhsb27T8&LYL!= znJj(#GtIhgiE~AJ9*Qg~0<qf7SeEP8U#M6!>D!&0i<n8~GtxX$NE5dV#ZI>j<zHD5 z2i$PoV^R7NTDZ&gqCa+}`Ao;?y#pHeL>=L-DPZa+Oy9WZt|g$I7sIB$@+fw3c6J>6 ze^&M+io9jqOE8{cy30CW(YCL%e_oz7BmsRkT_71vIN18t2kzXLy%5v?>U_(FpMr_6 zo7*tBanml|+xdeJwo+`9Z88r^#6XU*my)XwW(R<ozCOAfAC^<Z;XE??hsY;0?~Snl zmtvRktmBnDIsfH;3_|h4L`q(~%byDoxW_l9SwCj_kQ^U9_*D~bo|#LS_~c>PKPGIe zopya(Bu3RI$9ZqNYl`dlJ{Y*JBmOaiaAL$T7*X__^K`C|vyE*ciK5<qTF&mGJw@5i zY2FXpI+1}K>$rantG$M2o&RPf<9EoOeVvs8kTAKk*ApWjy-B8p^qUd-?z68X_0^lY zXJGhJ$4Rz|&VKRL{+5nNcI@l8QGp-ZWa(O6eh!&giX+A62#fT(sK<3vb4ZW;@6+|Z z0~Yi~jrgOw;_DVQEbg}@7wUJ9vfwF8qjdQwWrf#G<rx=hr*2_I(~<9dsIQ*X+ij!P zqWdwGxW!j?4mg&b&R`6oTYnf!0g!k100&7+_itp#Gw$fN6tSb}nl}~n8(Z{4ebN2W zZ<w`AgV3Y;;VGfbymCSS;~LDwTIP$_-lMbAZGX5)F%8bRjOo|fbENj1Obub4*|Lbf zxuJgL(D2=VHM`7&r?ITxPlHi+BOZ}E_RgGvQ4M@~MnvXw0egqA;@*}=?3u#Obp(Hg zj?;|5^?3H_DDS6*&mowX{xg$e$B$bdwVfho_`N6m-{S{!3WjOhKlbRiQM2!Iuq=yU zlCcEE<~i69vFG`6BAR-CWFc;Q9Nqd4KKJkGD}u)#nI<Brw{}l3+v5lo>kOsjjq#T& z?WOqh-?r#;p0mX3a<=FRj{McI_V1!C`wTIUKkW1xj;_;@VlS|RTf)tyWR&kcABc#o z&AW|#95>ImY7C?EPv8M{4bHS35W(N1U;e7bZJ5)Hg$eAvQ&}b}f(J2^+4<gh)_|g! zX03B!h=jIz#NQ_P-p~4+_n5i=U;t+2$L|wRbLD*M)$Y4%K7H}&Kl%`}%_B{g9{H8d zm5OY>V^+2g8#-}*JSHXoIHZHn@Q)~Cm}V)cc)m~}p$h(0=+L&$Z9Y|&SY`r`oNt^# zvH#Y0K%_Wlzn2a_ULb9Vk9U^@m+S|x<;sxP8fH+jZV)d25~7PrsAF6uX9~mZf=K^d z++;;qE|c3>*tkgH6mJRn7cuWB?vej$dWIFI-_Wf1NauIr0Y?irVA+a|>!z;=Eja-G zn)91wNXh&#%tT@?R;*ZOCM8o%KWAee@*&wa0b!oMq3{FdAI1RTVT8$Z1DxA9zi)Zp z^7t8Zm{`El{3Uomu3a*wYtg`!UMk6f2Yu(u=+8kVheG1b3C<iV{v&ec-r@`tJw+%r zH`ar2lC@v~j~(>o8RuGJ9f&Sj>vMPGPVfs;kwxck#~?B{XuI=#S^lWOoI}L{kL}On z;&uPI#Ssb5Hw+g|SYSvsFlmX8)K=lo0K&I*!CT*k*Sz<P8SdRz%HkxMNiQ9_fX`%P z9flZeZg`0E`*gvxh7KZL7j^YM<2`Aff2eUHoxhAYcwyo(B3oexDH(4DF1z5c8MA6I z3Y7qG<8_i|2^F(wyUmT8@wxGHW=qH<E4b$+^%ToIjM>IOcjDJzfm=0PERtOX{r{&} zNawqeAafa>tr$?ukDr*N7l=hGepcxAEv4{Wwn>u33bIRfA!rQ~?7mRHhItk|=GpK+ zIzJKV(o1j~Uo<X&Yu371*ue|-y{x?vI!;@Jy-bs*g`08xh7s`Y1-stb2{(B7USQp+ z*v!2td_>0$`!FUHum*j|SkZ^T;yeugtS`=s1550OIYPxL+Ac~FMcWRGZj?lX`-Cke zr42KO84_w~!EsT@Q9C;SIIdsofXKk#hv_8SL=}V=&Z2pb)cS!ODS6pO##sVZURG?U z^H<^lPmeT_*$P+1_;|fxVue0!_o=}YC;vW}qEo2wVDS5XrJ&KURm|3gD3WRNW&%}E zQNZew6~Bsg2J4H@xSOJ=TPTjx`7tQzV?f>~MMAviFb{$2Y-1wGv$;VQuA4+~hOLFm zInkV9oPgq-Z$s#c2$LI!F5o*PgA>C1F{`56e8)Aqa(c9A6dt~WD!8K90Ja+5D0~@` z78Zt61v}SyOQ=En#Hb??`!N4BoHLt5ZmYnuDhffZw!>`}Q0jA1xye!ErTC3~qvQGU zF46|Ub4+lgBR{e*vTy{*s9<9s(6b*fQVlBj4K}!ypg40Q7n+XZgB!UJqS6ME&SxWn zDaGgZvEzMjl#Vii?*DK|XPQXtse&5hg|-cqZnHuA&)eH9pj~L_XuL<~r(@uTbO2;} zqe_w1=K|EPu<eVX8q%R0<#$A7PMxS>bCMYREj$hrpnnpQX(|1u`3hDS^}_gt9nGsE zFHL(-n2X6x9L_A3)8EDBZ0@<RaX+^*0J(o>{)fR8&Hd-@ggN^7-v8)xu3~%tLxMlW z+i~V~OL1$;6U=J)*XW2nn8zhaWW*j@DSm}+7w|V3Qw7sE*-C=9&im)o`$fm0ve>n> zuBY+!<^N1a&=If1GN<7uuOgR}2hhZ4meZ!qPcUX1Q`~y>G;S_zO#s(+WbK*G>yrJ* zixoLw`sEupB~#^AkniPI_}onvn{=-1ZP&5KxD6&NsH+nel%K$vZ5HCziwMwh4BlLF zhveOI%j^(}+?3e8h!!V<w-zN6*{$1J*KV3V5DngHb#7fvix+_K+lCbRx9JrFsD7I! zIh;#1q{{2D%)hXcgDC3d^|&rf0p+R?5IUM>2`!)OM@KAYJ#H>$iyz`Ac?=V?Q#^uu zu71hG@?TqPp{SLwDtE+KMJIW+6Rx-roSja><qXci1S*KX*Mlprx$~j@<*xVkQM5P+ z)17k89@qC<gG=pO3tE?8{BCH5`nH6ko-{{v0E$5C^wvsRT#q00_|NXil3;H$r~<pq z5d)pgW$X~17ajNq*DsGcw*m9_HWs%ww?e7d=>-&Yg#ajRlcJkGL>VDVH;9{9dw27s z<)_XWgC8gEbcf3bc&D}H+=fMtq^`l@Ai^WRMgplrsSh)Yn9-E?vH#$E&9zUrmIR-e zZxv=0b|TvlJLr++Va4x`rPh!;SwQF>#&X(251P)o-Zq7;lNtQ(Ez~6eznh6F_kcfj z;@04@#EAZj|CTcLdYForhn!${ada)BtikZ4)<a-;3NSCj7&Q#lYY6Ct%1xJ;68^Ut z;#Ows0#J9>(<twO84$YV3*ooaU>##Nd#2R+^lBbwK$lu%DhBTDw%a0Ob+1K@P=<$r zJi~IY&Sxx|x?OzJgx@^_-BbaQTDJIu_=EUZ6m&fStx3>Ui$f)*6xqGF`wA_tAWe4S zJDHN;+u2hc4xM`jcZL0R()r}-g|~4rJ%|4D;_1ia4ynrIkRioxzr8#M8sP4KT7Ra+ zF*y8kV5uyBxA9KiPTmzB_<ta<dn+ye0A;T{5!3y2?mQlFVlL_$#Jz;c(<RR^b8&AV zgeGxuFY?`OtHRb~0N`evS;Thg_V}A6&q(7Qm2kSKbcdr7o)jOzyZ6vy8^UD8f>j2B z#l4Qpj!Zdcf#_yu>pnc|P=h&WHGfM6E+u=Ow-Iwexl^t6*hwFr)89vx3`ub6-k0HD z%C<pg74UQs)S#PC0ii?8oY=;NXH7ne3R?I0?iN~H3XQ^o)bXOzJ^^1GhnAGygP)`x zxjByC{T=1A#q_X?r|w{!fdJ}~J}L#NX5w|pe}k04U$#I^kzc$&C?ID_@63M%5t;>E zOzSg<(CqFht;@<|U_h{&QXWG<n%v3Kb^rrgsd5)IT(u_R%w2ZU^4<cdDWN5MJNXpb zp){{kq$j79_`qDM+!Kn*;zT^Vb?BUDxo5e}8W8rW+=c{UX_+DMRs?5Y33X|ju(b(? z_G^>36tjU)=pk?3=#Zhny;n0_>b$_xWvvzX$utt2bCWGPofh(H;j4wJ@esKcL9g@# zlI(U`D{rTba5<_c;u^Ne8xgfAeFAhP&p*Y^fco_LTIaf%AS%3-q{RjVP)$0Edv4^M zuQ*+*hJNi}=#q3MN{qcBs5Mkt2wQuRUfsPk5y5imwgK9@B$<6N5GYk-uq+Ayt2(4~ z6Sk%!#?z48bLoo>mKZF_Lq2P1O~=Vca><dY0oX|<Ka}*}j56#m@}SZh$X)i;&T=o* z^IHUajkOftnlqcn*-D7l&Dm-sI<3EvY^P`IA>;gx>w|H#$2`B4Cv3Ckw|gOinr$O) zeTtg9JvyI(+(paxr;6Be`s(zp##FQ(bcHRmncR(f(H70K|J;gr;DU534@6yMIy7$N z2&<;ywLDrr9t}8E4}jh=r}o)MoNE2&EWkue{pvOjh_l+NBdpp8n5gQ(=eA$1`fKCL zmIUZq1D47YP~~|6y}?Wdzx^>S38JE_(Lhu63aW&$xGJdNSJ(-Z9IS^T#to~Bnr;g6 z9DAB+`5<IP*h*c&lR;@u*ozZa-L#eLDq{BbJ4F-9Y-zbQ$w?sgUbR7ZO}3O`Z0z7w zU*xwpiETDeWWOeU&bvXgs>o=db}J5)#RR@2IN58f*;cb23Gm7~W{(b_KYE|!NO3mi z?`gvS%;ofipcwRsDEujTFoCg==1sk6VPK2?7#OU99-*hPstj({g&`zirX<+&NP>G$ zZ2NSm6%8|cmdh~Ls!nZVB`?wCBru`uM27>;5~$`9KNw6_Zpo^K9Jt!1pSZ>cgfZH6 zspzJ=8e@hdR+_(^pi4VO)M>qL-%H;g|EQk1<LinWlYJ&`I{4EL{+Z2+3z{19#giv& z5<Px7H2&eA4)HUC<IglyZZ_UJ{^X_H@|{zUM)Z8A?B91m%~H85r}sy7h1YSh!^=<m zGQ~wtN)PH(?!_3YjkMH8KbS%IrL&~?W7bS%$UOFhOWh~hXbk?W-6DhO%;!`om!q6S z2R>>JXL6OlrP-<GDR)1@<og4Zu6WiY)runi0uCh)CmGflNBNPSC*=3L!df7KI#n97 z1q4lwolWG+{@?xQ!4zBhydO4XYA?!u_LEWrLI35hrjx$=(s=8!==)jyrkL&$<&7f# zaXu}r;q%+{^Er&nG;8Vp-|s)YZ~kue{jGH1Cs0u|w9?Cj87%o?u*>Nv=^>BwPoE_0 zvcL})ovPuAx?IkdhWnPP<3Y(Ti>?Zc@hu9Te4bD9?(g|n!7@49zwFOu7YfT%n^W>8 zasrVdt>MWrH_}|C57J}v0$%1~@Mq%7-2c(Mez)ch@W4t8NgpTbiaA|VqTa21AtEXJ z@gHN@ATCkjKh+lV(`=w5iIm)h&`LCxV>Zm0${M~^zHHF9b|;4-UTBP-twD+7-Ji~6 zsBR;tcCZEgi7g!g>82`G5iT`6)DK&rNh+hz{Sz6&uH#sSHwearh8%`!4;@&Y)kQO| zM#}z&r1j>UuY{gLy+#=X_c%!nxA#B9_o!`^zxF@uIzC;ipA8B6<wcF@#}IxMQgUA% z6!ID)Id%;*S55wtG-~_y?{#B5SW^G&UD#A8gz4Gp6>uLr`o5*=JME{Lsanz>-dr(; zpmZD%R8hNkxj&YfHcj|dLdk0wT8TV)bh^WA%&4eCzsw%fc~tS><8u8jN;6rtnU-oS zq=ptBH9SFog}f$IGeh+WS!?=4Ln3kYd)p8oEohmMu&aZC9GI>r$38$vd=Kz~B?A?4 zIDn|z_Z^{C@VGjXmZr1i`{yVV@jYqAszBx8{@yh#I&kWf7!={&|1%jEUUJkrZ#;tI z4)EH`oQH<s!(lV$`tcB5o0fEfqf95wrw3^!N#yDCYJN|fqS`PPic%5)=_0Z{3V`fj zag@`v=#GSKU@9G0vj@fE#rS2mn9i2On24-txK&rpSq+Qd-$5X8RuZ{pUz(>XL8)1* z)hyk>DfDByOzS#z5G8S;<QlE(G_KO-MC3Yw+gn=*ZRdI~#&%Gy=7|QL+JPT#3)Zkf zntx4Unxo1YWuIOMefKwR=>CNqDqo-PZiP}M6!vz6yulC5T&XfKMSAVDaQ25iz{7@r zX`v(FtzZhar`Z0`)ObB@3dc{BiY9)Tjtmqo!--@v^Q94z^io%jxsWsCEf6}o^>=xK z^OwdvcrhIY<er8e+|L;U3t#Jt8n!n61(%gJ8yaVUPPaTA?i&$&g*TS556^t5qZMZc z$-ti3SAzH@R_6|+*QKw1DOSub%K69;i!mIIR2#y%eXAADa9z~8Uy%x#kmRe)goY84 zFkmx=Daww|iFcz&FKc;&_x@CX*@**l`84DG>YTZkbmbW-q8kiOpDi8uIsA)`K(VLc zRpYD1+ep2q1qMu$vx?V6{4yct=fmZ<vH1!g6j8+;n$<Q=-k{nKP<tM?{LsHN<LF=a zez}0{5Zd_p81WT2(l3wT#;3<myK&Tu?~2V_Tj0eJ7ad`rG6wp~&toy35=I$R{yxkc z%A37`J&I=88UuzfBhF)xf)?0jfe(+bFr)*wOt3K!DBdbg!exv!bDogL0L>O@nNL{C z;q>v5nXqFrs&IHsjk@;OgMvJHL+1f~n$Z`q9Nd_1acMu^Mfn`?!ug6qgb*L5s9~p~ z60RjE_w(#aSx%viWn`re&gnCwWofi998r_7420qGhC;2d+t=MZgc-L4Lv>(m_{3*q zW{Lx3YU2i+m83hzeuH8o#G;_Spb4@{@sZjBS?@*icQH^>B{@h7_89`d{VjX#L=g{J z(7;0koQJ)r{3iCywYgqwQN!iN%MF(sOeK&39aM=PsAM|EzA1*n23Pz5C2WAn0Oqj* z7YER!agMtYP*I<e<pCWjIlMUDUf6I92e93-6g3>nw`kB;Oc?{)GD3Gr-_P3WlV8DN zSnF{5W@)3j89#i-r5V==y|eceq`E^#U2~f`=tT|(Qu#3i=#}Cs$**vX47!t>qK8V4 z-SARs&HOw1#%qzUdT1F=3pHj4xwrUoj7^vihe0ANk2A@giy1Aiw60%+%QHqz`(Dw) z4iZ>$xPtb(bLiIoTl-U7#^m`Rc{XG6)>3a4(Ly&I!Q&WV{w9{P_YM9`21`cAu^~qv zF?n|_!BpG6d$iD(h`z5@1mSyb<&0v(cpoZa`*$kVgSa%V@jDZSupy5m*hJ<n0;xAT zL5im_q7`Kbxn<Rz&H%R)3AL*PFaeb{Hes2gQ=u#<K45!p9lUkq){#*!t$InbAnj{t zZ5NQFFO2028(g&LjU;Bu=Fov@e*m>D^);=az}!^l)fs@awBJX~g$?N>xNkI3Cam#a z9ke<q%}SdIqXwYg>O8HyVSVb2gv!^FxUhy9O}l~nN>Sqfv+X(;LVJAQyBVRtBQb>< zCa2F6WmX51+vm_12v_o0=KljkGU_u%BnDVUJ%)S-<?78UI#9hI@$M^n3S|A&RI7`! z$uM3!D$>c!(MibICT!2>^rR_CQ--%|Ij+wXvC{v=_82XU!BG;#uxGy0wmVyNM%K;| zX8fAGCUOgyss7KXzOX?{;A)<zL1Wu1>=kNE$ISEww`u4wRP5D;e0PSt!Pgc;2=*mC zh_aJ#HVJuE1E)wv(Z?igh(*Gr1eE%kFg3|nfbO5e@tJ)t9GOrP_&Vfu2=jj^x<9c^ zhD|3=w+-{dGAC0oia_pff`+7nF>})#!13o;t2_IEG4Hv3^#)3kE=3EMUvvO>vkU~S zdw&U>c511%^esLQgn4D6)@`jzY3X~UTlrQ@Kh{aQ=!KyLrW1Pc{&c7>g_~o!OEC|z zUh6B&uLlw!aFf17{9?(A#V-~o)<dIsJ^v<KDms&JUt5>27z<<Fh2^WX;*uepeyA+G z+g-dDw0fZLu45(7xal?!PD{U_Y`8K-B0ki5yF7W(&;xDL`GxHidTwN_B70-}T~meS zue7q@m^0Q`y2!u|5`WgQ&ZXe7S<l`r-2z?1v_%Fu<BB5$6qY)E-bL+u+V|{s(xUf@ z8>nWTtwq~FEcGUimJTve%kY#ik29<^5N6!i9P17h>*cao_xG4pd5}1d!PtpO4M&y$ zhi^*X<h#kYo6=@hs~>>FL!iy9@*t83!ZJM=D6>3F<X8}4zSA87aN(8<#Kvg(4U4c^ z@V|AC70&7Y2Ebhuj&ZvyTi3u<papL0A`rLRZY%=Lw5${Nv(iJ&krkZMkl5j0f}kT7 z-LT1p#47S)1WA}uawjpcjuHbf*4hvJB$JlRvkJEg_tpo*z4Pz>h~%*x?mIAq<;Hl1 zBbNCel&E4?sDMm)`4W`(|5n7Zz}o+Iyy0>w_9pLvP)gi{XwE}lRKDa~AzT}xPy@2K z4EmXFXd0lkEdPKJL5PyX8=#$ms^opF4GWeG_i8=YSMx3wL9HbG1j&JgPvDl*kD1)= zq}Jzf>7nuhOBx15^}EP+He|yI1CE!W{M==J>mkJG<;NgZuh4;8CqQ8Ta(VCoFt;)? z`%nSRYkhcb8J4+wA`0wzY|n19fyLctTF+~{)pL4S)}#URG5wwq&eGNMZ|8a(Z>LU# zo(QoDKQYfL&}wc^kasqKVc$*D0cz9AO&d0uU73~SI13i(8vB@u=>e~U)%mw_Z)E`8 z9R&*2>9gxI>f5RCgu=m8tWYE=a#DCnQ<LnHoRVhVIGVz};W-=|Z`y||o+@(J6`cQ= zc;p21Y?`iXPtrqb3e<Cz$3D8reFHRhsu{|0{fW&kXKG5)Y*i`B%^#ai*6fFQ=s6m% z%5418?NV)&wf*cTLuY&E)!3v?QNH>Z5E(FHe7iX<V8p&>U)76fH~F@+sRt6gzm0Fd zQgu1{n4EX=gx`yO38Jk*?XRo;S2e3DaAnA7(e*>u4_yEKhwis&?<T)naDU(aDF?pq zH{h2UelSWl{-kri1QpY%c;;>rYO~>S%<r~N{{5BZZpOZ6Uvu2|Uo9C7YY#dQc_h5- zh}Vg$nd<p{PJLNb6~@(ujs@w$J$I%giBj3h)m7xi1&Z@s>?1L!9-k`lyHh&2N;$v1 z_ncK@-l8GLp|d5y#dB}vc@%l<xt{UEHjb%UasP^9pue{;(ecB_XD!WN0&i!J`luq5 zE)6fb(xs|mb)0gdLe&+1=IUwdz;w1MSUF2!ik;Z)Juto-UIGZbtuK$WQit8P5??3F z7WYy@Oy5z0=TjHz;-``m_=)%8x9D_HtLEJ{=C@X3hi8YY^dw@_ckCbdz&6O&uU=Cz zLR$>Pz03)G&!{=$A-!SI#%AoCE^U5)ZUMGOxvV?2b)<YD<m20W%Q42lYQ(3oVPmWf zpIT4iR^>xWYzW4{8tT+POKfm%H`)3tTtR@qve$E0&0odiSn73gx}!MS-*wV~B{@ur z*lCpzfEmG(YZ=N5pO95LobDxb;PJAH1h-qa+ql)SJRBw2*0Vk&<VDlkOc*egmcm#| zVflG&PPky^sT-z3SD5@(4fzl??MohPJz7oAx!{>MVRri7dJNxl^XSd&o7qZTAiCD| z-RB|umF5l{NY`|dk;`U0#fflu-X~7CFTZtQOhXGeR`-@J>bOq`Gyd09w;LRJNf~jW zel`d{+?q~g-qb<Dui<-?<9Cuf$$?04&7$MydNK3W$TC&T@;$@^xh8iObPcWlv<~5Z zDh&O?0DfT|$Btv~L`@Q?vy?iu^#J_hd}Q`On<0#2$D+Iw;4C`LU*#q&FD1;kmMG06 zVqb?Z!Y|UaTr<o(Z4Z-r!=S`9+L11e87$S1n?>wkX-C$q<*Y91m7mH$SRP5DUqS`C zw&>cDYfEZS)o2e@@#N))mqUe7{<S>s6{d3`&EjcPBE{}DXte<qZBV>@an?X`cLH2Z zMY!;s=FFGVY4cPBzcmsu$*QSAVi>ds+^9E=beM|FSfH&s)v%q|Mny5k!+bPnU&|%k zeA!Xs;_ScR6z4<%EinPXW~4}PI3r{FKCt1+AC+8j(Tt-cfph8@{9TiUf|aNr5iNys z;{igZ*68E~LU~Y=Ugi(sMCiQPN+Ppz91^pVWJQgODw*EY#sWUy<Qlb-TJzsFo-u>q z=k(wiWq4+s8-iIK58m(K$_0t_5X5PTwPuJgzZ`VE1gqCd`zS|Bl)#$V?iFP!lf2wh znO&J37Y^xA{!dOauDhhp4k1JUb5seBQ<sWt6jT%D5?AT;FZ&=|D_7-yN4?Tb^ABrg zNy$z;2l~fK3#4B)$?8<+jLi7*u#MCz*oo@lH%Ta;Ly3-J8O8{?0x0=cczQ4soyfcb ziamUMAB{$m&;RTa2n7agq?nCN;`LjlX_QOd9WA=dHbW~C<*Su5$Yw-^Ypt}$4_|cw zk3H(4p#){zI8nKG4!oR|k+TP@br%5Ua)LmiMLDE^EwbLD^?e*iD*Dy(pbm2W8d?@Y zOT5OT`uxqEviAY_lEp-H4seJZuXQ(qB6#d=G)jQJf8DjQ2w!KZ XZiMqj3PmojZ zn!DEWfjiF?n=oJvhXe6-P{vct!M3GoM)j!2Pnwqd02<0Vog7Cp^P35fglmbeAjfff zuEer-mpLRt&d!`NGBeHuJUNd6%6XB-bs^xcv$V{FpHryySQ*s60J?5@9=QLz!CeoS z$iHu{<edkIHLWdXaq2GPhi5mrTZsfs>x%#D{(Bs2OfZ}(EUEjWNCbzPnQ`MKo^|JP zE}I33)`Lg(xel)N%8ZMwn=9BX(A+QDFjXX4Ds~g6mt8!mKA~g`BvFfdPdcA*F1tfU z=X|imYT-b5k(=l-enIN9AoXyZQ<QIALe2)!<I{H@E`PjyotHpv*7WNwI?h>&cV^Kt zuG}yOa02si&Ud`0tJG)x+v$#@>-B{K7uL*T*)AQSVOOxaRJaRo%`P^;)FZyV-_@lb zl16T(oJldd9O`WEb1c4b+U6~X8m9fU^7zoSkcZP|wjMD4_h@`bfAhX41uqxtjGN>B zADurASDm^1@9emU?7yD(ZZM?+@)syJV|>MCeBAzH5@Qc3@on;!yX$3)hL(nBcxH;X zH2)#Iz%U{F!Qtx#UZfsoQXK4zPv%!VaM*%}(kWKM4lviTOyd7Y(Gxbh&3N^Ahi8(* zc+BZ~Q3Uam@bLqO&w54JutVP#xQ#@&c?ugZ7hM#svhB!VV!C0r$FOnxx3=wN6Ao0# z!aI6=#SPluxv-ZtA&D2a?A=90@+FL#*g@U-J;gD}TFM_1!Vk*J-uZASX7`LQdvrm% zPhSd^{{qj7aOf}y?z^pMgo_<J3)8)EcvL8bM}2|i&$U8RLU-E)Uq%s@WJrh}{H$OW z6-<Hkfep(f1%HMSHcXQUYBk~HUNY(=Ar$p|S|{JMz=j<fH*VN)N1yR7Yhgn;;lQ!n zgZ%_uWty(IyvYufDH`P`+oL~><Wpe8YQHpdzT$UeL@7QYX-<+^lG*Tz2z}aro4qf) z=-uSL$&VVw!7O<m&xpeO(yYZW%?~rF8>891pVRK1h$$g5NA09<V&lh&4U43Il9Mwd zCf0CiVGU7ErLa8VU+^RE=eBXd@SpoUsC*5_Nnt)bi#if{7h*K(I29Ae+TGACry}-Q zRWe0*pGThs{8|UyHH|v*%vBpBax6y7pM=}$a@$PdHVZ`}j8=Nm{LP9Q;>hTRvr|Hs zHT15<>C7b!KT~g(QvtB883LS?N`kyh;0K<H8a#6}k+OZ(a2yYL9sGLU>v?KS5BsU_ zlj2%qZ^1Q8R}w?nYW74hf0Ep^8T7*@i(c7`NHGWM^?Ws-ugJr)BJ9(-!F|4s@Tx~+ zsRXC}FBbYT8A-+{YU!{;xmI`$hp46Vz*;?q<lvm)$f5!Zh*}MUF`TMZsRGa1t$Tw} zF596$VheFd(;r!bP_eQvkru+*PcdWIm!6<GKB_-#(U)F{QMT6#u90Vp3N`!8Ze(JN zVK((<n3gi7Pz!#+P5LG&jHE;l&z*ypQ!n(`Ye3!OPmK6gERggUd1UEE<MzF?0hG4a zn%C2l-d9<P?3clgax1~eB^g^Ckxo@7@T`wZA%GbRgF2rBQfFJ`Z8$|bP%nPJY9WkH z#7yE3U-cAjJif77o^Y}UCBlucYI`tmtKNDAi{Mvz;)Dwu!K~U<Dt2*PhL3C8p#zi4 z<VznHTV}9|u%3%mz67IZ8+fc)Ey%B`mJj~f6Y#<UGBo}~LZZxMdwZ!Qa2ssfs@cL= zFJbf$Lo3cyDw;#1)O35h1{8~Twa#XQ(t~AHBSdC`FJ*gd;|Osi!KVH8#-X#XY>+gI z9vqqZmo{7E)*Q|zCX$pal43m4q&*F}E6I#E51X7{TAL9pV?-OA@#gZY_K~_bLX#Q7 zjUz4A{>hvk41n3NwJL}>ThwBmrFA`=GRbsksrPu413eUlzvI#VBd*ACC!Fubfqn@$ zP3qSeONc7pGH+{sl_NMRa72%K!u!)ZchJ2(65#27u|=&bdQhAo8grc^LAWp&J$!8- z;1r&(<5V@b@n<B<F&8j+sfCQ=&c+?6Wb>Yn@NiPxFst1a4@sWfGpVXf&bvndEX%z0 zK;+<-M-BQ6VRRGDSj{bm%9p^9io<Oo8V&p2rGa{V1=nqiM0k6LJQsoy0~B;hWhLC0 zz7US2CiL7`j!E-BRn0+h%`m%>fiZfbDS57zK9a4aSP8?a>O<jqT31n4IS$aa`tV=T zw2pi=<q4+3$!PrH1VXs6zB%`d)q-SONLYnOPXLOAX_Pjy$vVQXMwC2E8`)$tskoQA zNKBSrdxtTkgzBU_W(BH+RHegnfsZaWuVx0fPl<LZyInlE2?b+kdn1yzL}c#<0G<Yr zV={VX(FZAEbOqBy_!@VrYJt>eez`vGm=GIxz4Kn2&U&O^xTh7r$`#yn>NcKV^{n~m z#)nmIW1)D>C|Vnq>^c^TYZ&23WDYLXhT<~D{2D=>ynPtxKziF~1UNIgqE_7n0$=+j zI|`$r0N)5qHWhA+DLb*y^LT>un8mM<jzeQu_B7zb>ygpDV+gd7VuMVR_JnvHamL4! zZ3ppkS^pnx2Tx$e(o}lTkCMadyYYV{*a|oHmYoRqEFH8RgWw@-$)k-^it>T|<vv_a z`w4tp2K*D@^xLP#i4fH3?ifVFdM}v|y7>f-xC^68k)**8>WR<%>b(g%J&VUAh2sEb z=X*I3us%wdBaHPxB+_06jARImM3rO)ft3v{U!!M2T9lqi<?*Y|gWJxo(hqLm5WVcl z#nY?ypyWS1n1nLcbc{+?>2Zn~oc76DVaG7$em7FLXNqF7lt})Z{oRn({T!EAoc9kX zcV~Vzh)ibyR0d~JPo73co=yE@FnkPm6+m(=7wdCg!Zb7po5A0;k?kNLvK*(X7zDh} zY0i%WASFh!+)X_pB}RBXk(l4eGT`*+4(5cO(v2`QOH&QTSOpPO&xn?P<=X)7E42bO zJf}RbH)Yj5sdcApFiyJx#@yGk4gZPZP~K)GHmq71t-)!j%27(hg>!dM%y+l|6)zac zZSc&&DC*>^`6D{Qji0U?x;;rdh<bw4ub|X@&dZnv7N*Us*5_BjVNlP~b{8yU#H^(! z^jZtz-`WRNE1`FATq}tL=U!`0)919uU|9)MdTQ{Hk#Q~K&Q7W7(3Y=gz+naGn~tG4 zt4+ljyUbbE1nBUrN<%Vw3~2xA-+B2}9WC3+UWFMzXg6vCMke+o_5&m?Z@dO|E6IHj z2u_qVb#NX#l=eB}DD2V~M_+s5Dg5|*f@{xur3_fZ=VTMmt(CawqTCJFZRzQ^!kKPQ z5{Xdyjw^H=#dpZ{{i`6;VpxpT`W)pp<i0K_2Ex~*Q@=OLw5!Djra*kRJkOp|=_(=m zr}cZ{nRgtaCii=iN790nU;1BtOkSfUy<P$`t!9cg)B2$ngs~W^JS~!kSSG^9OcP~k zK7@En62mzNLj12VBG(Esqjzo%C_h91WYS`&C%t|(o~keXBih>AqA>wA(;~I3*EK{k z!f=L4tyR+iGfBxx(328{=VRS~UJ5tIxlI5K+08a^LOU1vJ=K?bViVjkA1+illzY`; z0JZX-=u|J1S=C$Z0V5lm4BJ9qBQf2^bLtO5Abk0ef*0R@dJ~pcKMTr7SKb*L4oBIE z=tW~D{Fryr_n_KI>8AyCzb`v<dYP}aYdnc_l*U-b65LO-0P#zI4FLNpuB$qL5@luD z;U;JX+TuNZ_TfdCimxoh{nB-jBj@5F=|(CEY>UdqD65nOj!gF$bL0l>Y*FWKp)ZiN zN<_VxMN54xP<2wxq5X%ap`_O6@aAHZ<uUix6hXV^dqn*VPFVG!Pwqd3h4^NsxvOH8 z2zUP5PQfwc!8xiMW0ms*Et`QdTG>n!?rd>alRY$#g8H|?;LXk~cOxB8bGW|~^GnB( zuueS0_h=edB~W%mX@!+yNKM=I3SENVZ;EH8&D2U;fMqGCAqI?d@nQo5UbiXP0~K98 zL-`OaW+G9ApjUx<v2wvj4XtGx^?VeV^J8pNW6pQi=l3tsCI_$$d|q&}CL3A=NeVXz zd?+Zry9kBlVSfwgW`zVkBaf(mQJ(8x^0Dmybi<Ku<d|-51W`dURR3sk18;E&9}_0Z zm8b#_UT2uvK~y>yN9>0lMzuxx0z`VEs3c-X(#ESom4Eanjzy&jhd|fpjRpD!ErVQ# z<n17xXj5D_)dEE#eKHToOKx!QB+{?<RH*?+nE)-!$6odX)$RW8T_2i#|Ic7QWbjzS zFGt-$i&i<1#NUAFs_r4kRHu&LfRyXm>~35Qk=1ZiP+Wa=p>I}2(WfIrBTrVyeRH+F zng+F^s=_fUgsS-%&@w}d!iP9>e-xgj-iVr*RpIwOq9(?hT^8jADwdv1AJ6Y^!lBO9 z=;rrN`M6A8A&oq&6<NcA)Ke#orCub)cnxoKm}=8`8Hk)l`)g+0&~>6d<4*mv;>U>~ zF%f@{QpGf{T=v-Du8siMQd6Lw1qa(Y(9$TM;UOAbm4|X?|EEe!_S>k(@YcR$#kBqy zOqV}|n<%d#!bwPDt=_96B|MwH*?oqtS@ohn9eBhOwF??%e!oRm<e%vW)D@6?@|qBh zv#{$Kl!+Qc$f;pH<q*E)(Z`N^IcdfKN!XITW}qf1%~gt3+DPRws8MQ5vT|HYfQHxJ z&if3rm&&(Gqr*}Dg&vU3308Wj*!?#0^u6E%)$x8CRI<`D%~Y^+?}9-<wO;GgQiu~9 zZup>$)6~uhp%ABkx4P?~ICV%hM*RLr<W$Wp;l~<jM_|oiwFR(vSI23M8w%95qUIQw z>eZgqn|~l#4YY})AwVdhtI{Ap%#OKxemee><ok(ln~)z!d#u(8be!%$pVnlc7YPC% z0K=;n)**&07?alw;MK~i0}e0YI^|etMb(TeLpVy|g#|#Z`2SZ(z_*opLbKxn*Sefa zGi^bwoC3;4xoJZ#NW>VA;lmjNWN?lrq5I}1YR$_7kcQII40xquT-SzoAE7%GeR<TT z$Iq#Trs2``DMGX48COMDMQV<=Q}5gcy;6>EowgsTbQ6AEN{Q4AL|>iOKsr3JU}yr6 zPIS%vT7YGHOsr_Kgvd#D{X%l;N(m&<MQAp(_Vc~AeR00p4ksrSLTolvua)=>quxxV zMAjS><!?d?9~FoN2HGfpd+n|Z<deHH&V&UbRb~cEM&}yjKx<xbU0M#?)%wvc*!n>W z+lfIg*4a--O{c_5m+x-DaM0k%tGte({Ej6)-ZYMVO)y098H&yR4Bx$pn>0q(XI zH~d;>j5Ypng~b(Vmad(g4u3m_%&xhsRWeQ_{5qQwm1v;|8?1ak=!r$!9vJg!-5h-{ z{5WPTvHUOEB!<I2;_$Z})Qpn5yJmDY=4{Z`K-UtTqP36Rxmx0>yNqRAgI$!RHY)dU z=2uuET84ncOPzo0awfzXR!tAKSCU|I%^FK4r>+CMSogz|nxXzr9HWy{w;j!!&&;+X zlZi2rF!OlAg2)<yu&$COeoU5%a_v#7yk&Wj;6&c1U98z8*j(~$h?W$C&&***wC)!D zEv4?enaf=?pKQ1#033cZ8p94S&}$t|-9to?qSZVnx{1*3oPZd}YVht!#6Z_wcn@H! zBdAz+3qF}z2BTS>yIFoty5v^oH{H5P+S-~k5ovi&OKPIkB<GeMLQa&KxM+1a^5N^; zJ;;YcZ-rZ+rTMnI?jvZ2BklA-yWBYQ&Y5)<WA?@nq0=|qG8;2^C04tf1tViEeETMv zlZUnegA*+?v7aOf-n?2KVET34E>EGL7kcc&-@>RFkJlGJa6c5s8-vR#6x{7&aJgl{ zcQ=I{nbuG!_HPJk%`L6O1cDa`4VtG8JBkf<zA%GxiB1v$(1#OTH=Hvx<MqSqu$+MD zIL;Q*u5<2~oz4W~HEb#OA*K~MvrJ@cEGte2trI@SS{_C^n&buo0DlNNJ84M@fdJ@; zH~w(yw9m;CIU`bdS@**PNt~hG>j;jKxlt*5<CvYdbH>%J$QcraMHM&hVs>81`QRG1 z&a&1$``fy@%Vz8uB@Meooy|F^f|X_~1&JFT394ISHtbMt?eAzwr!+3gV_llC8hFnu zHA_87nbdF4+;qC;c-myu!v14V!q3{wO`i^(1`&DI_fVQh^+$hDvnKP-?`jU5e2qLw zEbBO&;F`LuUB4>*6pRE5+kdZedE$I>m7iy6!t|ak?X%lw^n|uQ+*tM`_ss3zU|1Ej zFNxmQVjYxh)MHy!FaNSInZMPg-KT0*bn+92Dyzm0dH%B_HFhrc6}96OXJEtjfX-|x zEX5d|^Qkvp+WY|2@!>tf&dCUQ-ZgXu?5Zw(c=q+@&rs|;&FenkwLb_b_BosJ#_R3y zzNHH`a`gLKnQzCZJEbG36}>xee_ZDBG4fs9)Vrsj_~~!nK6}giOXL4DWV_GuS<$^; z&K6xd{N8<^IR4Ro@lQt-azEaIqTl9zvj|A~`Xk+^`+4i-b044|F>84c+9A1l<m0Ap zQ!HyY-co`2sq7@DOBUYdZ&>V0wv~RId%-Q@t}-LE^jD9pj~_1FF5UQ`J@ZO(<)F^f z$5B1Q7p(r<{$ly&*5t)D|5Z+EowK?AZ?#GH<`wCaJ~f^3$sJg>#Mdv?W?))tXSs`{ z=jnx41-m&rYuk;NCks9;`#a6*liL!<?#CBOA1dpYk63R$c1gYd>FS5d*B5$gl~K#n zg%WAnLFL5m06NoOHA5mB4s|)1#uP;*c;x<>HZB7*Sr3Ho2O9`ri0+B>-5=yja}n_( zp{}`{Hf&2+V0oA&pp`qj|D~mAq}a(*T%%J|er%D;nbq@E;~{B6>%}$tMQaZ2X6nM~ zzB>_|Flk^uC9c4@t-W^T9i)kLHB%7lGcdA?l+^9u4y^4q;JfRJW8J@CoLg4>*7LYM zH0f%Nh=`jy`w~x(v5PtT5-y;3PaCMGrBG1>38&R2qDATOR{w7Ofek<(Zv-!K{w}t# z+?SmL)j_Ez<J>=gDPt>*m;0Vo+X%~R7$(_P=W2-1)=%eNz~zS6_IB=WdkOVs6-Txo zPOfWkF2^zh`_1%X<ml@?F^SJ3iA?26s`hQ`M2U=Ryp1U^Y@IkTgmJoUFnj*KQ&CnT z;@i3%TQHVXUBKYXz_M=I=}G=n=Yme=@i37w>&3@t!~IN;1cFQ1Ia2wXY5!FPpXdX{ z=&-AB9MC~Z{!IqsHIwr<83;?QZCIeNBsX^@&JsHIji-f)43qAIFIhzmfnO-Npi5!J zPbhq8Cf@J42qYJU)vGx&PN+@2x(I#}6Yz!L^gTt-s3vp=)-xn_F(J1P>^2_OG`xa+ z%fk3XD&G^+-RXIQEfrVT#^?#?TULbV)a%gKtdPfs@P3Q|r<g0?E36D@LwM^ZkuZ^5 z=ZYJuks0pMzH2=R5y7?!-crvMVK+>ZmccC~($GRwkbET3p*QsPiI@7oNE3STqD3(d zy$2?IfL$rPWqHd&G*`6Jwjq_T6+FZOPvbPfMLQSYojQUH(kDFPFqR|3O%0cUU$G=C zNhoaCgIsC=>wRCarFZU^1z#4tKIX+hPcUFUYEU)!1{~QAyX+C1=*1uLA1&qd;o4Zj zTREUQN8I3B=8W6NzA@fkD#YM#;ek>KNFS#9{Ai&K*Qs^fFkd!L%}f#`35LaeeBkH8 z6oDl0OOTJH)+07J=a#8R*brog0VahuOeZR3yttv}PY`?QMa{cFhJ=xjJZ{^4hJa4t z4!eO=JtWl~q|O?S?rT<bHfnv}r$d_!tnZ5{P*}oyX7K8YiHsKJQ5XulkD=H(^{nUW z9W7J7=yCcQ>CAhL_Bc6twg?YoBuyAL!5BA~P70!z2vu}+Mq4iQLC~?BYYj_vq6~NF zb1H@<Ihuts4UPm6fwDs)$ub${{(H)JVZ$ld3g$PlLKUtE_l+L__e3ydTN-4oBf#P9 zTOAScdcO{|?<x4O0jPwV#QUl3C3dCZufkr1-QnwDIf|Bs4f#zIBN<LPd5(FE9NJ~A z6)9mv?A;kk<M3u4)c@fbl&z`4mvrcRpeSr`ET}(H4P&Hco?1^>{sL!~<-4#=EElS$ zNU*+f92utQflp*zyZ=$dQ6)6&z=6Ym0)J@qaRJo);Y~7H#-NG&F!bIzdXteO+g5YG z7ogMD^X!Kee16uD?x=ycs?rJFodI{CDSIZwjE6n?=E0?}2!W2pVK|aD8#eDd92}nu z;vGQV#xY>PsemW~w(>8^7sxo4ebN67gAB6lTmIiL=(_V`96KO7Cr}C@XQvm&OzfCn z5^7DrIDMp-g&;R1Y@7q{xF+KnEf|^UHEv8}9}3TYh-al}!Mta6=t~;F>{oB(CV|vX z%O8c;><mcIHjN=Fun~sa!{r!GP_}N3dBl3f<_3+Q7RY%3WHy)6$7M*+%c+jAd^?0? zc@+Xl_+XsA7H#&F;yJP}c@qOe!6B*x_xD?Y;u{J}AqaD*lNTGdgD)oaO@%9BDkeH& z|A*qSA_<)NFYR9vzKRHx3|%uaX3038)nRMzdO$#kzC*Oo1yRkYk;JZsN;@(-A;~a_ zgXs<@^*yA`ZrBWu6QV?pWHNj>Ko1YbT)6;a3eunf`xW?+!As@*nh43jOLhM>F$Y5Q zurUVSL3EIcmnp7+Mmzf!!}UU8i=vanL@AcW)}~km*aq+ICTne5hC_G&>jvcuz@xo{ z?(p1hBN3K)3n~!1H)M4DCt#voF7K|D!y`5~Y{vc60&)42K1ce%Xk=fIBCnwo=Jx{2 z#tDfIaDXF=ohHqHwgl1Fk5j+NFhRzirjrar$Lh?NJWihm5lp&HQA25Ahht5+S{L#P z4y;&IzAw7R9E#2@t6>WncctQRa34&CaMYAxKxRe|C~ws~G5t1(_E%>hjMnZmTEImQ zs=aUIiLr#5V3P9-ux-E;`ancMjbnPa1*{EiD^4Az&8|kfPDp~OwF*Y%g?(OBZ6)&I z`fIYD>12*<OVCSeV+r=l6LVp-6<F8KlKlDeyWQ`0xARb^d>j^!(CHAmTP+ZEzRI%; zL<PFq_k;jOI;{P1)j~?XO$TVGXDNN)s&SlmocC5!n5UvWpKOCI<wf-TPs<%YRB70p z_HA@Viuo|N56(TS^@r`i)GF^wVI}k_*d<PbNt%?(-y`J`?pVCe08l?3l6i^<mh@Fd zLJ%eY!a$FCOdm){b_k75tVR9k=^h=Z{#Q+p6hj+PJw;gL!$n!WH3kk?paNF2wOU8| zJrZF6M&tcvkkQfU{?(?!=n>?vYB$NBuL=_q+h1)QLCcu5lB+h7>}l<43Tl-9oGdRR z5%KuGIY9Ww%%>ufOL^ue@P*NAm@1^NOhwH+=_@#bYPUtc(#-QT5#1Z#6Q1FP!+FAp z{0o~?%;U6^h|Z-dGbsk2(!uEtqVsRp+WbMI&XBgdz~+Ng`yLD(Gl`c~?ZX_}BjvnP zR)NV*Xew~%S%Z9N5QLI8IuB#0*{E4+x5u+~%|_PkcCZFtQQoF`6u3$?uBQU`OXea` zHF#)MP>XdyvPn<nSTO7H^=*H8Tccc%dgFXJp+ea5v=r0d9@GAwnFw2|UjI1)T#zx+ z@q4@|P=-}GHkZK3K>-j|v)WqJToOFu+1}Ayl8JUXAuotWmjYI5d8(?6AZ%dA>?DH@ z2YrSld<~~Q8x})Gv<)n}t_w0(*&I552|BujnLW=z12zKOcHFpUMb!p*iC?nyn8Cxy zU>8bjum<C#Sfaf@YPAOaDZCCA20f*SQD4S{%)yYNs*HFHDSAx4T9?!Q5c2F#m8;~> z+eL$R?T4c4QCzn@Q8Q6&+$YR|cP|IysW)6ozFNztJ-a|C{HJXO=R@`r2D?1|tp%F) zdb8p;b`*n4jkvX+$NfAFk*b?9<DQ?ZjH4aps}E_dPqso}sj){iU8TVoJ~*fSleVLT zPd>2))Jc^o3enb}xATCnu6mP>YcVdBY}^C8ncGYxawS-QxT-MPvH5kjR^&(wN{ee_ zA`>y96jpev5MRl|VAowjI#ms!7OCyuT@iaBMVu?@1KMyA)??Cs4rCf2nH%W?gR9%C zH};;rtv%K(9K?}K13Rsn1~{<+PO1#GQ7VH&Ls$;9YsO?Qy^b8ZYTF^be8>}<vJ+sB zBl2gek*FDthn#v&Rk@<M*KK1d3jD}SZQ_k^g<mI8^2aP~kXBImfBykA?L%mKb<uWp zF6j9xE%25l5KVhhRTG(+;Dhk#0#LtE)zs`0RIMFX3kL&`o_gB8rTrL|5l&_f{;Jh8 z5;u{KT8pD2PoeA?T66y!#OSk%|8Ca>oF!f5Uv0v#(gr@&O4RIpHud+xpWFW(!+Z<| zm`^p<Vm?)m@2;IS+Q-p%Kd(Vt^EBoz%=(7qn|JZAf~eGEgWi<<A^{a@FjY3M`pBSF zdt$UNI`epWTIGVmZ7`<a|7hu|F#wrUCnNk(wNr^Lv%yy7sW`zl=yC>)+rx3c5jd?} zR`m&>Q(v`t<Fh9h>&IkSB0QREn})5P!r0lu=$bK*PoZ!^D&KY*S=!UN3%X|F#pAk9 zC?Yi{U%_b^ICoWZnqE<K&DmE}b;i~EV*+_d;E+}zob$*wz@zOYP}+CU#C1b)R~8CK zu5Yc@b3gPf^JdAg{sJG8k&2AH!z$_`y}SG($Fy4pf#{!yGO$aZ|HBeKd(%P?PH<#! z@N~ysi|T9l;A=O-wyUPAvmjAA4%~5tPsxyOF*-$EFS@QtJ2KOzpxxR+r8A^-*R?>c zz-of{fMwpXK)Ut6!kK~+Hk=8ejQ+fRT$}@#J2Lv4yD{|6s)YVI!mev<&@YWGn^!YU zHM{?Xym!AgQEH4K7|IjcU?^8%8QXEc4sC3et3W2a-O_*i9Ab%)2lrP1SYD1W-c3dG zQ^%gPaP&nKWQy`OW(+4n)y=(u9c;8PDNIAz*Pl8jG>w)e!53LtzEz?TVjrBj2oXyR z-K2krK+z%PjBjh8G3^>UV*}rJokYgN!BVrqS({+r{50)iaiLO$yKZT=q0xXM*Kqki zyo6opJXEGsF3{GL=s5MW(TLUdVc~r<1nch!@3;uozHqEW$!%Cr_TWhAfLfv?KN_6& zCA6BVk0F=ljm0RBU_53sR2AAvqnRbezOks28hsR{RYS7hk<i^6A`a*ZXtlc%-WE?M z0F(VS%+UV&C$L@Ns$$2&)Pd@QO@>xe_&_KE(`wmB6Tf-GMtBJ2%lm!&5twg8HNKSw zP)xg4_A`ZD;Y<jTdZuh%x-PtoThM<BbsugY!`Z+Q=`1Zrq&vWET7rPn9PHnXw0}Ae z=~rDm4+h~J*{auJmKgw}bZZs7qEdsQ@c3d5%IzPU&iWot<7!KT$_BYIa=%u=NCqdx z5YC6R<7_(o=Xm~ghvS?a1oILDyOK;0T9W5bk_q)DZK?_mdwqFG=tJCCRr(R$ul`ND z>(4*lzYH;iO`cY0IL2^hFjPNl;jYma(|z-)lcm8ZyPFF4!?xT0(ub7?%GT-wI<NG@ zdt>TeZO}<kbL8CgS*ri_YsS*C#v0O5rLRp#)dcFz0x-%DRc_^RjOhd~YpA9ncMjgP z`8E}5!{Bb=x2YFMmJ}BH9dIk^#gv`^&%gTJS~^109mning@W#=Ee`#OBmrk*>eYxK zyUJkJu!KAs(jsYS{QWs7=+&-qJQSqfvUH5!Z-8bP4PWwSX`t`#uw8@qmZ8Jz&qL;y z=4#FN)#ap^@%!^&7$p2O?o#ZY1q5CPVSeP*@7MWI9aLWyM(l?``4*~y^L_OZI$Z&s zAxg7;*RBod#CF@0CKg0qGgTXTjivDG&srrvz<*NZma(|2RNDh}#2l!w1%GO&6n*4Q zy?#CTz|;;QQx9U40*k=!U6E)9{PJ`p1KK?GuiEOZ9tYpkGEu346CgZ$2$IrBt6<GO z`0@u7^!1O(=0VWSRG(;RBo8TQK_5t)q*|+d&>u6_4%9G_z^z(=X-wf)I7_O~(yw|I z$P9<@Ox7m)u**AL2+x=8Gvb)DF;Wie%1BC{&Ot&q$>wQ%RLT9l&Dukq8zFG81u_3+ zOv@<h&04rqSDQi`fIS@H1ZKBZ(f^OAZ;yvM-~XS_%rNdIl9AgKxoj6?OC^j5m0a@O zYF9<+qAheoxlX#!rf5yKt(c|4*>-iArJGBeRq1TVY!_B`E1RZ<O@=kX@A-buc|3lP z{;-d8^y+oLy<e}_r!W|&b8_oyN=@FxAom~BO7|3YkMRi1li=h|@{rYCtD(G;1w=FK zG@P6z+%}iGP3REHjpocHL^1E3d^HC+CZ$a|dKcN`sHeIYPsJR<NVDi-m$x3R2^Yu4 zQLkYL{duNC5j?`yS*1;ov$T`4Y7CvW2Gi904~~l;3ot$Z+<+?e@|8ycT3kz#;9DBH zUb)iHkbA_!(wf3LVhCTk^}YDh7`fKi-*uMtPt)M(&i@XN7M^~yXS90hpen@a__80( zu5{@<pZ<T1$M6thIb~V>k|_(?mhPXhaO+RA@BMxM;Hq2u!;aoRek<(Jy2fWs+4KEa zS&v!1Z&}tznzw2PiLlDj_;)|<B5uvlnc>Tt7yZ$0yY9?rmSnmU2B5E+30o?hA~^vP z=M5bE!DCqd{fNd&@ua^c%KQiJT%l7thmC#1t9DX9AzyNir5j`|4BtoqzVDy4&W_#8 zc&-lJTF-Z7HFh{9^SE6$V*lFW_iRCU(#5Ukx1N8GK@F2Vi*Gna=gi&UvA)lM(P``v zNAKWKc&UsD--Bhc!9_hq%BdWic*Tb47AtH0E(ID3SQ`==y~Wb^`a%6u8+}*-lGKQS z^Y0hkFrL5)dhOh0%lQ4*d2pda6Xqn|^K{bm`I2GIJlNfH+x(T8ULz*WOl#?@9b75r z;=R9i<gz%@Po`ZgQJ4|eUTxOpD<0RiSHcUk5$$o+z?(#9&YoPbNe92f=GGYMG~Q(h z7w5C{#d8GtPI>2}*z=+f_rClLFe|U?uK1p3AntRK__2KFj?eF2AqivxB$*ejg6E#< zi`DQkh6Nn6Cuv~vxHTM|#uy6IN;gP6%$Rg-Cycz5YKNh?dC?u7+k_5cCah4jXR^>K zkYh6PO`m;TPT%gva7l?%Xak#dX#R!yTCVkH`~FYENH=kvp;(Gd2)+)t3tRo)wYV_W zm*aCr#)*wqRvjF1V+_PsrSnbsl2@!l84LGJi1t4nQRurzyDJK5=;To6dz^!(A_|WL z7h*kVkCE`7)uONv?k@O-zEvXjx<FC!{Jad^Zt-s?ciTeJt#!kfZ6N3e5_4A84n3ke zrR)I7n~0F&v;KQ9cShv0knpipqdP9z+;|G^s&2mcbH={$VfRU~(X%;@J%?Dk>KZnU z8>EV3A#rDX_sk{mHi3t&&Mp^jOZxOQJe*48yvdI8aE+Euqk|}0By_4^?7=JlJ=2Mt zE!XVR2gT2}+03Y)SI2T13m8fkg+1aBK_^1cs~YD@b~=%czklJfHO7K<A$vG8_Z(z# zyVxRJ50u{4ZJxx}X*|Q|3hYJ2pLhgjYW!pjJk5)4x|Xx^3L!euIf29Nnk*`g3ZJ2! z>y+^Qs-PSBeMuT5J&B-C3MrC!V0#%xksQ+vs*6q)72iHXm{=We|6)JLUgI<6#eM=Y z)wvfCM~J<^m4v_WyCK=(+*OO<56PS)lJOQ($qWgF|M-c*ybY+FYZ={~6>O6(J8ISG zpZ(+qG~X?~ZuAa1o|fCUT@>~c)p1jhULu+5v?6R2>6v|>WQQ4yu&dzyYYR4C63VbK z3)J;B#k7dAzv<Wv<#C<I#f-AGYs%*X0XzY#>!ZYlmDN{EaQtK&?LTDIbxqQ=tOhwG zjC#m~z?-Pi#gok5Gt}#(ndOUv#QsAQ2^-J+xXl}yV;L~V)(cK2A?zoRoI;pbqbmrn z3Au4%$Mmkql`P@0QVFh$or<X0M@hcH@Ff*&V$L`##{8sh-#GXds#9iL_drL)lTY0? zGXR#75i~ST=;WhK1nmGz;WQqRSUUTlw*gD}G@e+By9T=^W^WM)Swg9;W!pco4yont zsyi6?k_~H~ZS{@5D;L(hr7?l!J2ahWVihzy(#azt{^B1fEeW;OgxWz3_k37vOdWKE zijA<GH);?hs`3<u_ekL({+z_iw>b(q``P%X%+Usxql>>}XYFt#K$D@~wPEq4e}oKK zQghj$1#KMiq}pgE{Ks7sW=b5RoDni_y~O_3kBLFIe_6ki=(lp!%NzFr@M$<aavSjv z-!XQIoVIXJg@6ip9%JBEI)PIBwH7yn`h`Tp{O-^aASU%5=QxRF9aeOxAdKNi<3rA2 z>qrY&l|`}x<c9L!%FcvWA!>qk@rJ;vL;Ey|CO2_DLSGbC#UpOvCP(Pjc8+NCgNkj_ zAlQQ1mn*lj48%8&&Es$rHxa@1LYI?^dxG%Rl}xQe;RWj`zooH`WZ;FZta!?}WS+fn zH77`(O8wTb{-Frm|9CbM{BnFQFzi`ajg6&LLg!86X#b9ao4iD0MGaLJ+~@(RFL)G1 zb)LBHHwti&V=rwt&o3J+C6bUjzm0`-joOax`uqt;Ta?aH?ZvB<DWRFi3XqSK50<>h zOA%5@sC?nL<UEa|K{=wxpb(AO_Clp`kUYDUNaAB_4sS9GKd<HRisAlKTvf*a(VajX zL(H@FO&6%v%KfMfmuIt~(6$GtLZ^Q&+=dQ&HbZpU%PN26*c1};W@lQ?Aa3U4lI_lj z%5uB%KqgB)hwu^i2f}Tacx9z<Rk|9Q9SOX=02=bVb@~6!5|)KBL@et66R}@@CMxzm zLPVpSm~j~TmlXN3Y9+#`>aBbe7G`-cl0-%jo_E2<&+DK%SR<P4gy>D02sLKL|JP7h z)}&2AS^=b~*x&0I0U1|uXFBX|>+sS)vfO}XL0J=rs~RG1q0m4OUbMt(Zwgzbr9m@- zh`Ur1NuzXh6cedNx3b>${b2L)Q|r^QKQT>R!_EK4UIAx5Dd3r2BrHAARJ8#Qc5lT- zIZVnW-Y(Ilz&tjv4_072%}YNYSJa|_e$?Wzo$8{9s6(*HkniK!yF}yrpRw?JUGf#2 zEcHZDaiOml&$YPbo6*_o_Oc4RyX&3|_h8sncp~hjB5mUA5^X(dgS9VF`YctNC@c!{ zbuUX`<(7>Fq+6>J<@%7{FK^{bfy<B0F{yQ|vNsX9FQYgBe(Y1*ipQyzzFz#3yeb+I zB}JHMDs+2x_@=Ha_82FNcm|Jt`DvwbxA@M~%TrVER}--7!<S>LEk#>(4u3UHez@E` zqVs?j>mYfDUe+0uK2_zz97?<w;la$4dg8hvt6SXT^!|tnw%+NN?jkJf<=87M<b#|5 zd2&QlaJrjnCd5~Ea@iR8QeoGt6VzW7?!yY|_bJWbll5MXGy_dYmdZ~)rVTq;b2PY` zJn#mx_F|51L_h_S$1(Zf2nkGHy|{xbz#gQJ*eK{tOe*1hidQAcz*GhDP^jZ(!qn=C zuEiNyDQ2qA6q_!yaBOY1YL&vWJmEsdg!~{+aClzyWtXG|Rx`72<gz<4w-L4~6FY!D z<Zt6y-G5}hYO&(D+yym`|A>`pBHo$V24b}t7Ybz@UaJS6Gct3g>MJZ#Msc_qJROAx z4Jd1PK_c}$VR-E#;<9_5nz$^0#SFg7F<+|#9NXo+*j45b6YSK=E$b{^6IxxU_R{G0 z5u)F{8l%o+>3p)`V=s}I;D{SMtT7PLrO|$s|5@U*ZF=oP?&-L)6is-Z=Bmtzf{&WG z7LQ6&OjQnw3XNTRkU*8-OYLxUBL80wUZk~%(WFcjGMke{cvvPz^E^9e1pe$)`q^j3 z?;w)Vz3M<w*ccc2tfB~xySKb5JfY+T1_cpasyT>OYIEwl6;aGyABttQX}&l<0RX9C zXV*yjSHhv$q-ejr(v4_8dQO`!W5KYeGXZil)eOaI8AF}fj~qzd@EA$pW9>v?y)+4{ zw8<A^L9Asap*XVR&=wL?Qp8(YAgeRYdcs&UWU(@BK}&<I@JZ_vO6%q`%)wKwv#IZH zp3UQqqL+<}Jm{xZa_xmb^6cB5&@|+atm~d6&R0$9IXWqtC5zkK&SX`3K`61FtN>YD zY-CGiu~L&ZwCjS@%5l_n&*pIBd1O0<fb1~}61_JLlC7p$;yUv*E~<g)19gF}(37N4 zJb*}*M&MqXjd)z#iZ-v<L1~9R4SbcPyoht5+-s|5_IlDRTgeF$wXuR^nItYpzcSD9 zA{75SWnHKN#4%#0tN}nC0XV|4YfFzEemro2tg;|gs`LRsB}0R~jK^kYQu!vapfu5< zVy_$`-_<jOM@im(@^FsbF#z=1KY)gl5gYQrtHU_u9c*s%9@}W-o602sqcs@Xp(V~D zIcHlf%OwA(mVMjzgp2P}T(^G?F5a>Uu{r`bOP(J%I&Eu-7Frt#T5J}2?oZND>QL2g zi$P+%Igq+pvQ%;cF9J3crYTGJpY5Vz>!mOk(a`LpfvkDbPpo?}NWf%{t%Yozwg?8* zfTy-=vDcvi|2EUjp;%=>*tgb)38v1Y$%xVdyy1|9GmZ(WKTgog-L%^Z*u6^qYx^md zwkUOiG75Cq5T)6~36Q;tZ94B&Fkq#@4%aaMmYHfaut|&CgQ<F4e1wn-je|s&SceAK zB>W}KJNAr}eCO7Pw&o>@ndXoS)qK9poI%V_ni%cxc_3+~a(&xjTCHq$A?Z!r0%~{F z<N5z$Lkd^nVa%Rbz}RgrY*|B-y4e+VIJm^yy3;8B7nB3B6<8wS%dQccQ_TSLaFPQi zntss0tROJYX_C&=>B7svvy*zA5l<nreJWrE60c)qZej>Lt!Fq)e8PZsFC_}E)$q7e z+bkKowjtTa*f%XN_ES?%btj@;pb>Q@i=7Dt|B=%ELOVzt$}&-|BeW}w<}d(x(><gP z9;eMS_9UKGLJ;{|UT6tgUWDArNgAj0(O`)a$fsIgXzH|7W1$XttDv$TW)6)iiEx6A zekIA%DXRmeLfG@IvjYq^RS+1awq--PGh1eqtPU8~)Vj)AqB<)yM(xeknT$;i8WUuQ z3GbpP!B3L^)FyRqjj2a8ysR=ab_JF&77_Hkfu6VTt#6a2C|ha7m;FHC6=@t|`UHNx zMz`OfP`Sr3QD&0N{aK$s#zuD4ohW_zbw31M6EtihjD5<2pvfb}K&H&4xkfasN8Ok5 z2ckV&G}?n$p!`VLSHqy5LXNpQl!)ZTf$KWpUv+8yLXzevBV`GA@USKr#<Ox`JI}8U zymfS7wMHEWk&|rp&=^aNu}(2yW>c6M+@QKYS)$?&e*)9n#r{Peu=Fl9?Ii(FSt=_6 z**gi)3`WqBv0&H@4-sUW%)jt)21(v;Sx`roq%X{wghvRPj}1hM*lFYQv1@{)G3>=R zuOJyG^H+0rf$)sbVU931lT>KRP<V-CS>)!b$D_Gab#7V1H74;i7U&`tKA}+{l5J9Q zm-xi{2@T_**{=_8=aK+BY#FDe)7A;a{&iwqPLjDYr0v`#Ey}?y#K->3P(WL(|NIL7 zhavfA76bpMvUjmQ()7IpiJ0|kf?}jgplQ`m;e_wcE%H1GliMKM^f;v0>)7m4Yn*FT zirRcy2QG}5Ds5!dn<HD2E~;eCsu^wb6mb}{a!@3cJ&nMG*SDB7J6vvip-_}QP9;!$ zklU4oUi8jY3lt+|#us{gt6dTU6<+12GD?!<0XNP~l$>$SJpV_~kXF}MjmgeM7t#*w z(CTWHEOE9iUKO;*yvwT5lC%Q8+hNi*MY1U@L~6Bg&&00i#&M)^6Y#uXhfiZbqm9H+ zV(sMDII*BMOglK^x7SGyKGA(s9in3Tyc~))^nKxAyLQK{f@Jq0*8=_1w+}y<8~oz_ zdeedxZgY=4*c^OxL!XsH;LN!PAK0C;dh~iU*l68VUcvCp=X)P`o&MaReWtF1X+!4v z+v`5*H`K;_@@nvj`SeADhwJF)ueJy>b3S~L+xg(%9S!EL8J@4MSrj$>yD+zDUE2It z5dxpnAGYTj$E1D!YM;R8yAOrAw*S?6<<*@KbNg7D*(>FPWVbbeqvM7uUMJ*Zo<T75 zRbT3Z75SYDY9yb&>O1f{=6;2PZAiw<Az@$1gTYBYr*7--u~mECPv+S7y}9o$x?G8c zt|=swyQC6x>(_|u+fI?F^p^LJ?|S%yiaK5kixd6h<>kzsuShFo%NFstEv8RqZaMw) zZ;0)IJ@-%Fp1Z}8b?E2c9st~IE&IOBudgw8+D<_4I9WgcNMdMooF(c2mS|m^CvDN~ zw(&k_fS?@m>%C?`kK;W_;&Qim5`X>k5x{T1Rylo3`_<V9DAqUZHue2|Nof#tspH6e z4tGm3@6x8Y$R|C%nU-53ublX@JuA&@%RA<t+lyl#_DZ(g-}G&=bp9FjhQ<hs&c)Sd zKZFW=_WsN${oBmss|B`;Z|xh{)po1AJuD-jToV25>8-#0JGmh_yvA?HtIIGqxnygg z*LOQ6c6~3acFJz~X~QSohLbI`!@L&_8x)kfdY?`*QvILe+Wo+S&idV6uZF*^U2wxi zxuDIVF<y2AyIaQ5JLaog<bf?l7p)(>J`lX-!E3|d8Q)J(q?dJ;zWC+e!@0FfZ|v`5 z+K$N1oSwh`LCmT5zr6Mi_P+l*G}!ys*AIi^qx<yx_P+LeFzAqZ>O=k~$G%(lJ7dxu zUTJmp<~EfWmK$puthBr`K-;~_-)Xz|?9Sfd+V-*P<VQmnUTkq+cv$|0-th}xET}S3 z94}U8hC7efd$50IrS7al;p1=k-M>HYw#WUN4c1?-oBZXyn)TWrhu$yfFw4&`yT#0m zzq6%$kjY&+!kd5mrY9RQ`*%EHOqE11SWrp~CI7jT{HH2Iyx?@iV6j!`IrNbJuknUB zxWm%7+DFJF+KL*c%hR`OMf*Pw>e#1darK0M*#~2%<m2VsDVPXhPm3`WauR(MqhhHI zdVN}|J7PDL@1>f8-;<ZwX3b@p6drLF=HO-5M#qU{!+gA?rA;egW`+%x+MuHTUhKan z%$So~Zd_s$>~%`L?Y~U^HoJ?Mc)GB%p;Bxh_bbDAfjUSu*`+Lxom-w4o#y$fy*#hM zigj8Yrua*aylE@rq_a++t8x<gqD@l5=&sD<3&U;qIZIi|K4-r031sPQ8(ervJ@*95 zqAc$MuKCFkgZB7(8M$7xS>kmt#Xz-ttG1~4Q5oX)*5X1jIa;AOB;GHz(V|G32$^?i zBPdMWa1p<U^QFp^3(&TW|0#;*-WPctNcl`yR>}~F)~*-552eghtyfq)GZ6cq3tzzk z#7Azv*riUNDa_d*3J5FuhBclP{K(_G+))C~2DQPRqE!c2<D*liZ#N47)jP}Hpf=q_ zsBU8AmWP%UF7!IA4%T4#h6#A)?`2Dlk;peOSt~jBbrd`K;w((TnzKwUkD7`7(`Nwa zFUO`SkFv(|ZSFqq5pZUP6@^yo@a6cGJ-*cfAexslNwq+sf5n1zXkH1X2KO(OFL9-q zjuYnmEY3wi4$eYo&O*kLWum%pimjJ(A${dd7ULv5F%a{Xoy%C1xs=pc@#QNB)buIQ z!6S}}4f2O&6~Am`*(-9@P2zF0!(HxiiJJXRg4fI}^;cY}5W~r4V(!*wXv~adIZMSW zuURkY^=Ikj%<{b2A41?(l&mq-QM_Xg-Imrct{2xVto}A-iYiKjZOu5>UmO&(P6^TJ zU!#9tr+5-e3thKPgP^*;(_;T>n+OqevB-W+9(bz$@})`ug@XfO_&BEnL|3yK<W*t# z7Xn8-G(1h<<z7CdHX?UV!V<XL7M6u?x;Z-xxvr{1%)M00&M)h|Gd@hTdn&f6Sz7z6 zuF1?BAgIA*SjbeR6D{sE7DSAlqeUU_GUaIopJ16yc{<`M%rSzX0yToN)v}NF)3R4= zgfZ1WZd!R0tQMZg`1SD=A}ODM5!`JnLe-I#q7R;iDje{~QA2x%9)Vg)b!x;wn8W0c zlVe|z(C=d$RCn+$>yUaO0sEcH<$i_>-hMPY;VpI(a#RoHk*L1)gvZR~73>DNOL#4* z$hT-zwbn?m)>LIld75F$uRj3|$k5klfTfTrh=K}uLIp1ORLrqtoJ0q%g~OgYOI{f1 zC<cj`Mp>>GGft}G6jkM8M}A`k$t_eybvkFteCnzA1I4+*O4uTq4j1H$rx&t=<VDeX zPIZ+rf50T7+jZlA)ydzhc2_h3NiBk86p$=h)?55y;fSq9PPJ?zr!7=YnH-&UQ(;h< z+0>v$ykMp-;HfRlls|9Dt;vwnk6*T##fY>OMdit81(_3r8q<fI{6{z{WBJ%!jh3z; zXlCA}X!eSQmD=mM)2W<d@kn79BPuqdVj7y}WlY30WBuDb5L0u5><hB`4uy&|6to7& zf*;CsXLE$%n`G08*-wwj!uE8fcK*0}<jqj|ws|7<Y%5fs2Z_e-oVmmUVypkIb>&T6 z;MgW=M@^DS45?ku^<v^z7rtnFLPgW^gaeCjr+%qzwpciuH8ZQynrdfUs)$`VhnF8a zSU)~Y9G6%3lrLs)O|`E%I8MmHZL$zC>f!wR`lx_o+I*Umq)Jh{#Th2Q%)M;}OtU5F zS3-20hlt-!fO=A(D1y0{G7W_-0{|-8J!45KVddd3+e`_XBeAKd<2@iB*)XR;<{W3- zW+O}2XuMJ%I;3Q%4teqU^_9>erA9pMwqWHJ=FmVi<>kf`GXh}l7a~{Q6cDI>n~jlx z<4mkeERSiBWj7m1two*!4XTp}szwx5T=wR-!;2)%(V&V$p<Cy0^JDV}s+ab-FY&M= zJhfbT(}I94*&N8unvMJ@D`|XmFy{7j+Q!LBG=^l&A;`K>WUcje)}@M?ZFt1t4eSO} zztj(sy&{^_9{-yHGxh9s?an_W*QWreuFMn@fo#Ho0#;00K-+dyx{fhMaQH2SI)b$a z2JDG~He&;NyAKIUvR0PvvyJQ^nY$|hTE?$aJ<$;{SN^(o>BU6wz4>n$jus7!C&Y(p z<!+uTYu#Mh`l#NDIOcO&0GKqJoQjGK<#4xS)CT+62q3BRgiQN1#zc9V_+>-R4-lSc z>$79OTfbvj#J*`QL?M6*_co>iT`!&sOx0Fe$n;%#h;y;#?0}Wudzq66B74JRJsU?H zopLgvTHC(uR{}3?&3_e2=+e3~55KIoegY&%k0t2}V_n$b<-oeSQv(ay^kZG5)!d!L z(I+`a?>B4PKzWHeTy~iFqvkc*L3Q{n+@S$I$b3#5uWtj$4r$L$v#4aRmsWo>(1eQF z2x)`zcd?uDxa?%>nvhh(zlrNm{wyB1Ja$DiF`=X>!klxMM%ANeCPF#enf%y_d(ij- zpO6<htnq)oZohAN#U_96K!wZ$$4Cr`Y_>q<En`=-{&cZgmm)c^5jl_S9reezZ20xn z7!jRH4HM^iFSQ%x$LbS2hl2X95IkQ5c~95^Rwk5IbNDh#!pdq>*NY>T$~ak7%ZmfZ zZW04hJqY1THH7m36Pn>dVg5#7TIa|15<DYngGn^4E3Z1{a%Y0rKxX{d8-r)r1n3Aa zI(V3+OqUI|c$1VV5o3V_r~_4^*_^Cr7JOMALmX$iRlYcJiP-;^76M$z{OsE<>@cxu z(ZUzzO8_&p#TygCn^=H}Z$(^8|8_vP%`vv<F4d!S8`od-l=>9W^l|+98Dc5E5<%p? zT*UqgUAq+|yZIEgY$8B@9P{7&j16`swt1iv;6T_p#?IXwRg!*Sz^tuGqefK>)F|DO zsu3El&&>F;UW#T1x^^RYGRqJn2YXX}@aleu=(fkQ6PFF{i;%SYiJDPlL+l~YqfS|G zYfb2|n?~$2h=_`)*asp+{9CltCMJaNbGZjq`G_%6o3o?leYmC0JQ5tMXqYMFpn&Z} zt0a<GP%S5Sb8Yd+w*yPsdflkHXR>lHFBrC?CX~tMUh$;zL0p#~5&I8E5t+X;>3ahk zKdV~Y_vQ+a9G<F7z~KpNxXyXNF;iCCJi5Rb$!VE;3PR9=Cdx;1G$^%+E!uKSnoWc) zCt==yemd-oHY)q<*v!VSrXfl=_D^T|Mrp%rYP4N1R$Qzo#-CrQOe#14RY5_)0}pC( z(`h8@G=Y?HE-`&7LG4xP>n$+Evma9XaC>jPw4yYD#Z9*asom}DhW;Pl4q^hY#+tG( zJr(Thx2777Qw=#@EZ<IQ-06-Yw!$jEd$0k@crDMN6A|~%_off3)>-85t<g_kA!;U7 zWLAc4-KiS)kv09PEy?#Nj-BSvn{;EUi-TP56><QALul^Q&lmq(i&papQJjDC@~cRq z{psF=_82oECp}8d!HTmGn$r)5-k;u`+KT&{*;KzfZ2@zHMPcFu#5z4(fnbQkuq6%y zH)b>VM93-pp~=X!OKk#cbRxk?<L_<po9KP6X|^>IwEk}<5T|0{FEHOQZh3bgRh-cv z_+E-*=I6)>8pQrZYmsLhYA+Z>%-Qxb{&h8xvCnrI2Ey(<hLAliI@sr%413Jl#ZaB; z&Lg79oKJM7`yO1RKEMA6NwM*129wLGDPsRjHc<Q{aqbTrp@%&wc0X)fsSlW^NBu?3 zMMSwr#@C(7Fd6YzAL@_Lz$_*9QgoD*7itT!a6p@zZbahE=z4hGc{+HM{T0%Q?0od^ zJdj;gllLK!2)t;ak0rmK78pKo<wc7y6X@K(Qv)*;iR*$)ZHi|Tm%BX^qg*^9p3DhC zCh~=4es>e{R`HpHnEurO>M@qnl1zxN%^zI_M0o_!A+g&CTYU|4fl>RZd@>pMAN8+P zG4Vl%dd%EJu_>mA5eMI~4vm^XGq3*@5U>fE!S%m5xsU2$g}Ai31anf3L^WMFxo=NC zjta)2Y1b)|6JaP2)YqC2Z_$s&;cgbOLUk9s<cOWHC5n`5mQjmeSinm*Oug?(3d<2a z^>OUn>k}pVffa+(trZt*zZ;!L?NI$Yk$45IEk@9K#atTUWdL}YF$u2>gqOE^?*0dN zD73D>^SzWZ9@R&4XK<|%zn_*BwOXzhYY$}@j!aZXYMALkQ8c6|Y6)|A_~b<riQl*| zVRHHR9thnLLQi}9TW$qvm47ZI&#W;)gy`<Q*vZvVrxM3I<{2G~Z$bp%y9R^nXhx8r zO(bMXC4|M|59Cu0v=|<p1XOWozk6&iX!r0r*%Aj;im#Eh#G$4WjX9;8V*Hmxh%v|L zV4Re=+i%?JbC61CYz~o-qsGv+`1Kk1tO5;>BdubtK>{8eHW8~lXi?&beCYaL+<?%A zi6U+Fo3}U*F>;ugP%T%`>JvC-Fyh|tqRCwnw1DZD9>o;vLRWl*-BDDR#<XEZ815O< zCt>HAe<7qk=dYcwok?DhP?8Dt>?L{Y*bLw~0k;?u9@lCtz1Eso`Y_d`F&4jmA2st{ zmVNZQkLXj|5jFbw?$SIG8_y+QR`4%z)=S4uWtiiH12N3e2pEy$NOr&N<IMk2K9^AL zLMb1cLT=!Yz&NItFaFU+JE*^tR1`nvZA_bvFcz7&DQ!BzF*9(?3p2u`pc)OVJ2e`q zu6T=2K4{W{(M`l@-l7E+bS)<v*B&ydt!JW2d$>VVJZ25L<JImZk#NMfOI>ejLF$66 zNlexxG!pKFv0}9QYg@(|Tz?@hGe*D(Sh;rP0b=I|r-=ACXK<HGBJJG?kc|t+{@^TR zSldRv$em5BXEY_i@``r;-nG_#T3kQkgumxB7;LdPfj`<*(+^H(>@K<f!6ABS5XofC z;ZtZSBKDIPOprl4A$fb0!wT5<y^esNNZ?-Uc~HsRxy*s(7z}cAEh;xWtGm4SFIyV8 z1xei8%fDPUC4v_8Xz3hO5o!fV+V;^(rMUj`R4BPs^29yF6t{=kxx^qw{N-(Mi<9+N z4K&B~JtJ^`-*8#~?hJ!es^f0K6hUK(AgY5oI(&mcqB(ai8w;5&wT3|k#&s~qS4orc z5nR1(e-ys{9iA<EH~(&iIQp3_qU=EqE9M|?=KF`2UGDCq)+n%HBlG?1d!jqhxLBIA zkGZprPn^cVedzI284BI3zJcL5Ru`q&L)Zm1{@x~*kc-WuZF7VI8wRfJ?l`gA-L(eU zu2#SN6QtoKM)Fp!-TgVG{U{u!6NOm90z@-FByK45-x)(fFf;UeK^osdY+!J2=KKkK zgFGUezditBjNs_TIp6&?M+5N{A%1H;7lqVaU2!y)p^2>V<hHWPCSMRjB5uJ^dt0<( z{ycAbKRaf+kaM>{fa2h8IpOA6ZGrpzV}IEn#Q#P^Zp<7GjOH%2L&4u-p%I6BH-hBT zuJbWRl{lzvfAAVfG$d8|O@82Jqv!oMYjq)_%~;9J`fC;3ub|LK;C}YEZ4P%&8#L8X zrQGcRyE*ug*v*}7-q3PL^Es7!A>Kr*x%q+_q=ekL{8%jtwm;7%zV2N+Yb=z1mi5T{ zcYW4aT2t?TqkBd=J%`Ti{Os+8+ux=vy}jJ!X~+$O1;?)YyDvUb%>UBwv*WUa#uL6L zbew-Ib}~7>_{`$a;o^}CFWzwGNqq`O6LArV$?u-fkFk~hdo8w0GN+9!LUeE|d|YIJ z&D~7B01|IQ;d?mv!C6n@7|NViCGZu7r&u$$b$jhfwqse$=?~nM9Ahow^(UOqgqK9) zXP3VE8p}`Z{<cZ#vu|{|1}Z-MQ<pm$Czymj5rSwvRvI#mgm#EK@aV4I=K0;72Xuv; zKr2zQ53I&+O-Qm&fY_+J({F4IzK>%}aP8$8yzN{%7NSmmoXJ8R(Ga0F#<$XNfyR!6 zGBoONQ<>}@1hRAmMFv6r7(u>Yx7M~~*ajf){&Zp3hJbWAU#r9KN8jhB*2u5oNh4bQ zZ(521Yl(0^NJdTY;VJbJw*}}EW?I1avG0>uwnM~j*o0xH<An5o1K;-ek)OwR4X)*O zi$`DeW4s05LMo1I&Sbluul;6JPd&Aprj~$+XHIwZI|^}Zk%*6#t1)L$0Ag7Y%sg$f zuq9L5RHSXxUeII16>_qbfyu@MO+<t}7_z2PisBl)^;wc>H*%!kW37?6e)r<BDFo^n z4W4k7sLIC7DAX#p$z`lb`;UDQ4DtNQ+>99~Hn93f*l}g05pYey=)Lb)BEhv&SMWIV z2pPrcuHB1Z_R4u=H9q6jpB%Fvwch9%u`mf&gsK<|<b$1@pt^57k#F1r2QJ@BSTrUg zKlh!_#CQ0_qifW?-~dnTUnL#Ghc3sw)V<A_*`t9P;g1Q3@V#?MDL%&KPg~8UOx=Wx z<KHO=W>%noVBX<)FVLTQ&acYm2}Xy;SlA&UXTgB!n4atTu_*QM+uqW#S;#VTOZ#zV zz_ny-I#tt5p3XW{0ApytLHf4vNgwbkIA#iu@&0G6)sWTYrvtkTw?OyD^hfY8KITE8 zS_7*4$i9AKd_+m3Px!DLpgK~hwkS}hVtCz<Ad5L;Garuq;$OJ=d#h7!+A#O)*1T&` zt+Z@rrbK9qr3TEQc;7GeUC-~5NVyy0;Ex2M&hPNS%kTiI^V2xo4{1a`k|`qg@_2TT zgsRe?e!CI^p-N%jr+V7p!L-k{@upbXPxv^#E`R35OVUfy|I+|71P}YR5gvxwU{2eF z7*AsgSeRxd41aPDc#ZgkB<n!aM-oarx$iKp^oiHyCYEkx!6Y%;?AKW0oL4;mn?B+# zft@96Cw$bD*pB2LYFlf*<OmD04lC3Ns%tc49qO}(2CclN-!Tz#*)nh;a^G=0m5wdk zqbGDS1p~uy*LhiC@bgnw-8PtNV4k5cEh)@WB<mc$WD{#98PrCT{W$*8d`qJ1xAS|Z z0VZ$F!sH1jF!nf_-Df2Vo6UnsZZL}udY;=it?`;uWHeRv=O&>1su__9rs1=mVPZOt z=ww(W$3FV)IkEq91~4m!4d3nt%y!4eA7X)KQ^|WBoyM;TnpjkJZo05F^m*Y`cd9%R zXMFu`gE{48k{0&^8#EgeDPP50!4l+bjfF(Zf17)LmJT4_*e}z@!?S45qzx{jumNpg zNhYbgv?N|msbK?msN?8c2ZL;^p~A+<`#KjY>=EeVJkxw9mphP8+!IIT5!U+XGmhU5 zG?+*^jlPnT?5vp_LbZ`#nbCPG(OmHkR}G3M8AxZ%(Q!y!8Yf|jp5nK|3HZHXnRIP~ zWS;e`dfY<ri_AaTfRA5GRxBek*vE+SSZN8C9hF)I%+2llS|b_EQn;2_9ii~tXvx!s zmjz5~-(yTJ%e3>ID((WsFSZ<;?GMtw*4pfM$*^0zV{(^0#?C3pP{<wZuHB6vB&3R3 zH~+$I8wh5{L@F2@^xlne-vkzk6mubLvxm%~J_n*eBQB4>0Gdf(o!MoEPijlgF5IJ! z-z5IGD-LY@P2#caMbKPJYJrXI_>!a3n`q*J`iw;UVf>guS((f+xx8zpF#N?~#EHlw zL6MWdg8$@K{z6C}jKE@lBaay25WY>g$F|E$G8p!V6!#w_eQSwm0$qtupKHvoKg{^v zU=df*wLru^F$!DETZce#w9{dAF2r-swqOS3r&gDVC~OABb0%na>zdrSP?D$dzc;K1 z?OK%f6$m`$d`TXMDEcZxI7)7iBs3Zz6@$lSz`QT37d-AF-gk#J=BvZJb%{Y5g7mK% zXGyj=6_I2WXl^EG9-$hPrcEn}G$V@UM=^KWIh`Ph3qkY5$z#h2>3#Dae>M%{fI}u- z8%1I93}NR(z;tLdcZ!mxX=&K`BOJP<osvHEnJ{OTSo#MZKe$;tc+jL;jBcMJn*Qp= zd3`>d<!=64Hi3EL@D2~5Qzw_ej08+eNkv3W$k1$!YNoOvqRu-yNVB?bAuFOXzhRy) zzEF#`2_6BM!g+iFGGFd<dY3akFev?$(-$KO>mjP%<p_Y;lHQ1#&>?3Hz#^jSJ@+WU zH^30@NZ4|~JlB#yj!D-^0<i1VqBDf*9j1>r5>EdAY4GVzwHyOyjKF90={J%-6EyP^ zL!4!x##wT?A0Wq30}9j}63@^MlGLsRD?j#3P8YzFf6Nf6T`8LT`ZC4XMM1V9$!7UA zb~<NX_;kF_MPou3*K{rUg<v_0<!gl#$~*${J7VHi`DK2@B-e#LzDP{+S+>lX1hLg; zZW;;8?1&dSiBGfsT*gs|Z(32&C`{ml5~&82O=WzoxH^iv)crn1t@C!PsCWVa7(eAL zF`mlQ`ES42Jx9v<y2Klge1R;J4!p(I<$qP~^>gV@q{4gUtO|#$<tmyIV4^t(zC|q{ zrXx3G>2}+qIsD3%S^sCF);>$fdet7oeU_K4ni-W(c^<|okug+B<GF%m9_PJJs8`~N zm2xtkWP|A@M9<4YwYhGXEOr*dZIqdc*y_2$a1Y|xJ{CQU{)V;w&rcl|!T6veqpB?r z#V=y%rAq3}aXlZuIQ=tKxqK5!6r?Wc6F|7?Io0!cZJl*^wc1Gu_u$z}#{jdc4lt)C zuis{kU^X?X?$3GC;5OnHc1+~U<5)294t#$%e$ncMs~2_{X?Qt6NVleHo<4&VD83yu z6-=BZWJZ=Shi<wf7dSh}+V?w{`Doa5-|s6SosI)HrwZ{vPneljmp|`9kAKIZo1RoD z<zCS8veb3N5vkcCz6(a^sZW<zGrHZiydZfx;pR6k?$;+mh0K!ia|xCphL_>uY9gL& zrk4$K4$93VqWm#&!=f6550yZJY;uSbR`|kl7~(C@F(%EuhC%XuZh&`t?mlZs=WMe5 zKI_ZKE!8^|@cuPyXr$Z^UkZL$wo`F7MW8|PE_`KKAoZ1MT!z|21rvz;Z<muWJn9JX zlevpb5fmdY1%4Ne?`A4m<S{H?uL%U&Q}#jm)3Cf^n}uHgW8w$G#HVWs#VKrI4*ALu z%`y}Z;D4$O^2<tz|M`!8Q0M|!q(bj;0o9++F&8c^WZ=^jmMlNTyYkYo(C-TmX>{@k zf$2eEsyB-G+0e-eR3}NA`Z>>D@tjz<BJ6PhaTnHpZ94b>NLbVNWGp+GO$0NeE?=Fb zXfI1BsnLW81@S*+y}uGe#Dg!X%2XUmx-A!$V}Oxn8X}5(+pHm(CvUz@x@QNVJ`22z z`EoQn;(e=G_^k0M^&~}hnYl(Ot99Tp)Zx@)U^8kvUVTr}#AS_MP~Bscha&=L&{mb( zV<TdYH&4hoNYcwzgjI04DopoBhWQ|byOgd8sLoB%q^R;O1ka*6isx4%KE@|fyhzrj za3n6nFu$ybWRpKWbySkD^X|wZ&loMp=!m=tD)%LY$rCK|Sz5g}g`@gNQ|V>3EZF17 zWNMEd=-o{bmU+;)^MVaLzXYCo1kcp*Z=*>vWD__~K+@h90V5cVh$8-1m4x~5sd#~Q ziRjhTJnBRQ9>qCokQp3){SWvWk|hZ|sxwG0=NRRebv6Le4x`B(hk<C-iu|5m=Yjan z?=#H#_-r$zl;RQ{RGqD!uBa;OB&G8QiJC)z8dGmHf-we+hg2I=ppF1~iesYML?r*t z<ZV0<z2$v&mj_7>M$f8i#ZqsQvMOAlO(S@A&Rje)MWgUhBv541Bu0T{pxYBEz+?b? zW|?1>N&Lp!>E|O+@C5e?JHAOoYH53ZRF^OJ;Sg0ahbjePw7a%+)J)?yI*EufY2;Da zVp)X7TU8MNg0tacq$Iud!v7`@dh}COLC+oJ8lTHpTs$>*WIcgH&y27w_V$0MDRZz- zv{5#Rh$eOmL}Mt7eMeR32pSOEKNGf`AmwO$tz$<v;bLt;&t08nCQd7EiZB*0+mX#S z+JeU?VjtqqJB?CeF#W%=S|pj4>FK<T#j7Zd-2?c9%<|h9VSGn|*o>#gdPvIl@J&$i zI(UpqZ)FP}OnC4>v3ZQa=DcFBZwG?fv^6-|;QQSHsa6z8OLo>A%&6a^`s2kQXLDqm zC=U^_w40r`os96KOnGeV22Y85f)4-`W4zO2LjlyX6<@mcKch*|Q$K@4k258<(>z;? zyL}pYA0?iRJd94oPG_4tTwC<XS?Aj!nfE(f-zMd<teXp$w&Mc>V<d4=&J|Dk-=&mx zO>}Jg4C>Q1f$C%&miFclry=_hhLeK1-^yPh>9`6<Yeb#sc9(CBe%&GB$&{Nt@$|Hc zu(FURcwAW$=huF)(vGT*>~|uZb2LJ<;}!2mTGb{oa2Qw|X*Wa#F*|d@=5*-dnS_w9 zZ2bVQV{<B|2Cy48J=TR~u(_#Jxw!rKzyj*ZH>W}sTLv14apRo7@F?bpe<e;Rc9?Av zW(=AqB79$;E9g0h!1(vLyB)4zqo;TP*NDLpkSwgruk?xYY0HqE)JXdbf%rG2*xEqn zWg_N?t%!c(BM58--j9=<VA+4{ns!$SF8;A==3V6#BPpltFPR-)Pf<oskHyNy=IoX= zzNuPD(qj{ZZVn_c*HD9L+rsiai_7kA6w7;Xzn_p~SZ71ULv6<;Be~kGUi@-~0b|{~ zCeRZXbWDk-^wf1dZylOCUHP-DK$8)iV?%>H&8Svl)oqS0#S!1<X*9+AvNoZ5V1~BO zLR7U%FN2J4CILsw%?X&@{*B}VctwBw==?cf&wT?EG&m;UJMbmdh7{m646R#JfLL(~ z7y}zUiRyIxyzMjO0*NPI{U5>a#Zt^!<NE=Gjp@*w)v{#r^#LD^q>WK^w|LVai}7mT zv)GErrqcF?i?nkT`+M3#+E+d~J@K|>?h8-Y-feqBwtA8#DJNqrTMw?M4SgXc8wXuJ z#_DU3JV}7gpfN`|Q^YT&UkZRPZ^Q7E1Ye9wG9$wIM{p3q`;ULP2QSYC*}u!c^I<L# z4Wg|quN7YkSaeGhZQE2tAwy|?k;D<^Jf&8cJ~3b1QVTv>fR8Ja{jm4~>K$M1B4m39 z4&xIbQ7l6CVIX)-X1UqpalEE_7$6Oyw%@4wbYP93l#*Nnl4s(Tn&|hF4ih{Q0jlTO zLqsS~W2<nkhWd#pp26m*3muDZ9Zz!9FfoD9+e4FGWjyG`>+f`Gf(a<Y-s_2q(|N>X z=4U_SGWev?9Oi8Lf6Q5<Lxr8E-I9B1xBQg)+uTjx>J2u|p3xTdZprP&{{KrUUz6%C zb6lWEaWXM3KYwIdi~pkHg`a&T`^@XO>`>CN#}^O0c;lyC=~M7*zqgdr;1m7HPvkl} z4(ziS6Ets1^Lq785OwcgN6fXmedkugG~dM>s+&LBsQ(6t_-J8#Xd-!}>*!nzOoocq zeOeOCSX_TK5&Piq-`1UL51(n~?5g|sX)Xqq8axmZzWtEEq&L*k`Qhj^eN1QPZp=;4 z>UZ7r$<qMS+^#i-{Q7r<ldhv2__AAUlUF|L^zP#jw#gad%RU9JbFaiq`oB%NV{DxU zJX749^2(Ey-jKPDj7M6G8De84z`Y@pM_?LLn4^wDfu>dU>J}oybIW`uS1kmDO{S%! z!iftxdl2L=F5V@Uu7@vMmLm3^A{aaTq1GuK^YVh<oFZ8%ry{fklhI1O#O7|tKKc&? z#jxdocEcFWplE{O8$RCKES?SRs}EnP)h(V`d`XMDVJZRmgi%YTJ13X?apbBEMX&LR z*yeAE+d<M{!4sAx%Xc$z5@_<Um#;AhlEe}GQuMb^fDOL2uK)77C7KtNl2mNYx)6Xp zhTRc8zBfF(dL&9rFR%N&WUF!a+K}grAL%U_iiII;?uJ03K9XU)-|Y|x!;>1uBcrGl zoh7kw@ZF)0VTjd97v5{@G-6->grtK>Z=4}H=j0ha8!Kt`u>sFH5cgF_iX(a0umCKE z6x@Gcoay*oG#f;a_Oo3M+b03brRIF|bz9IF;6;oY@=?n{)_q~3qCR0G$H`ba-#|n6 zUB>Ib&UtlwQpp|znlvBZGaiygSEQ@u6dR4hl$T48@eO>(kNO46RZhdp@FiNwJqmE8 zsOYLOE+M>-;CsdtNn;>6^Dc$=K`foGM_k#4Ou|OJAc-T<oE}e;USCAmVTYGq0xbSo ziT2DCI*n-s?OM_p!Afr|a0->q-@AkCz2<I|l(He}!h@4Y7pRr-J*%hF=&tF}=?2x| ztK0s2Ftlp0F1r&cG*`Iq*<Z7lTYobhtSnrz$42PXgpS&j4MyD&g?@V`bbToy`v-u! zj|H-MF*&5ig3<=Y-8Y#^#~jAJ!(JG^<{M{pVy%K!k@%WtZhDgP@??Y&!|lRQYt?xx zz3zxL0ef^Y`+rv>@!KXKea9RsA7r|_?QjQxO`Ab@5IkOjZsJ5DpI4>z7(LP+xPI=s z!8_o2^x`LD^q>ekcTAIpJ`dhAtLt~ktN+2A0GNIAh}DE4z<zL`NInIUR~^v^5(CLw z?-yC^1Cl3s<L*VniSL-%g}uhPT11B;phJr!rp}JbLdXtPOKx9)$arVtTWUGCb@>vF z&%LdcFMhy*0xbuE^$ja7D^?O#?jHXG{g9ocDjb8Ajp%*HFWTcM4V_!CL%YkM@s<X~ zvszH}xt3JXG1$1<mbPW*in+lyFr2})diRC1M^As@C}aV&=a-4><)j0)1Z^mZF(LXY zx8J$9^!CZVHH|#16Tx)O64SZCq7aRdMxs|{&L)bU#swp-!N_pJ$hXU0@{z1~rXQWd z{Q=dDcfbuUUxHFMe$mhzY3PkpJI0}3<*30jjR>dDnKE()I?>B1BOfB0+5=nkK1>+- zDHkhGNGGfKj-;iJZ4Z=`-o`{P&7qfXoQl3rJJ`2S<AS2t{jQ~}4p6;3iFO)IRJwjZ z-uwcs#Hf2P?nIAQ!9eic=r`_Hw40f(uNyrW-=PWc=FU%9YFrzr2^gB#)lQvx_$V=) z9n-1ybYfV~0o^B>Lf$aS7yr%?6|c$zj_-Pov?9J<Y8SV{k&A{Vbp@hZBOtOa0!Z61 zw3NKNv^6ODz>SywkwZ~jDi*$gXKA8tGK0ZPl$S6(<AAepU$@V<cW|G=pt#m}pLsDh zgZnIT;Z6DpCGrVLB<QYiu08+h+Z|Iimea|=9^*9zJ_w*t3?+4ynOu8(Jj}ka$W3cK zf9fXJIp#<$mcO|3=c;}PXLRumI1O~i;DwUmh0Gg^oHC0);Y)l74vjv<3ni~3=D5RL zn3K%PqJCo*OBXBYCO_W=Gj9#74I^2fYNfWB7j7)r{A{eFm_RJih?V~JMTf78fBJ>0 z_&hGC{;7d@phUQGPsVpOrqXd7V2Y4Yt`YmMhOGCkR5@!qA~@9l=Lf;r{j*TYy(@c% zmdQjAEq9ZKc6>9h_-L6H#qMju#@SSJ96_*_j$$?;xNMMRFCDeUHh=y2eBb`Owf~sF zXu55i{+!ZH>diAZt#PuJgxH%LX=t^1Vx*!{Zm1Eq2XSi2^#n(iw-_H5A;$?wpK3Kr zH#QG*!|`yuIkWOMZqs7t`(D~Jf8SJBY)f{WyNOB;LyYl@I!v2OL&|S0OTN&9Rc(_J z3snswn0$CG;bBUqh~JCBVYY9&5JY1o)LRgNgs;oTBFzj^3HU*P-4;?BE=YG0mKmcc zo}&6o?5*l7GcNYdQR`6sDa#|A9Q7nBiG>HQYA+=_PupIq)J5}$$~kK7Wt3hy$A7Sx z)`GTax`FBnidQ4T#t~M!A~S68`yRh@n%Iy&Uz$Bv<wNon-ws=TeJU0aCewC60pfPG zO@3J@iHV7&uYQArd~$BDM+p!8pYe*l0EzA26Y^3<Egvhk5qXK#yEHhC8NyejYf~JX z@Pv+DW0Qt*BO#;e!uo*?!owd!TeRO1?b)NhPt^_o^Qat0BaGRE?r2ewI!&>&yy-#> z4i!Qh75dP|a+hDIbbf`z^>q|WsNF>2Vi78mvTURm!kceqJ~IY0Z))uuA-3oxNf8Lk zDp1)>!577g)Ow1C<u0W*do|Jtg5OpA;7UmM(&I}s?Y<~HDJmXf@8WadxOGn9GfTLi zgQkb7kHKQ{zPG_Yyq#!x&4lLCQR77SM0Zu9MuGB>O_dgkiD_dm!S}W;U+zIG?LsUJ zDzGQ1%7l2L?-thE!Ch><Zu>z8TUYXkazr|H0LfTFdHv;RQ(Y_!PB*2NE%)F+DLb<$ zObaL_zIq!>l@fajt~GpY9n?WO?;e$8lCI=C@8Gd{2GF6-g<kr6IjXbxMd=>W&~!(Y zQeG9&xlcp46`?zoiagy`nDY_jo0F74H6JNKCMUm)WT?-sdt4!J-baS<IW^c$VQ4jr zk8cOVcGPRZN?!%9%vQQ4M+4H1DEeW_$`L1F&J?Oa_n<-PX2NinUxDM1A95@RxT*C9 z7U;#A%0tm~b0$=~aQe$_OV=!@-ml)QQJ|BAYM*_SYFt$u5m0cbX)bN$+AH1=E|%8v zvq>Upv0bnn*kv43Pa($vh@c<BX!PY%%U3L`&C+1^1HrC>nxQ&T#J5Fr>}$%!1hz>t z-s`X_0@MB?*>#Q<l;@zAc4@RBT3%z}72vlgk#>f8b7{sIO(#hcwNtE#;SAHLy~bGQ zC5}R5>JXo%wiB`2^qE6yQxdql!wqWovexrUMl8}e#Jn>XSC8Y%&k%os3$>xud(|<D zymB90fkY~x)+eHIp`sZv6XxJP+lV10o=gjmP7{VF+yO5QZ6^(&lD+)bx^+=Rg*w9% z&}|$fh>kMNr7sqbOjlLO+clAUbR2@1+L8LoF;wt=XegOcaT{U2>^4j%lLW?nO{KRC zpz3#4JYIcJ4<;U}I473AAQ_;d7GA@rEVE?vh&HvxuaN4unh^VrwHD?WbF!Y$sQ8X! z;`(Ajek=!sUswC}hmvYrRr<E*n{g=Y@+&9uWo8ig;~wqq*5xgg(m3BpjYh^T0oiRj zB)O3NEEeJ(V4O9fx(y>t4agSTh^PE1ecZ?r_={Dk_9IJ3?EaCQP$n<Fk9&-|TlM0$ zv^|l<X)=orl9pG}z>SYnazBj2ATE6{ue8Rz5hi5Xj3DuU{&V-fF$M<y`Oj~!c_h22 zJXUFkuUVq^GbPSZR8+aQ(km{t?GZKamLvw|t^9^4CC&;6z4eH&EAceDrHW(UW&?QJ z{2hA5=&R)Xc;7P`3)ZsgSj_AyGl*O|Y@17!i`xC${Zm=gv~eN;Otjfg;o)Tfqu`{g zG6`~Cje`c+$kFPiCJ0v`S*f^vdes$nw|IaYA@&`{wTQ^aHd@`SHX*4_N-QsG<m@$( z&}>YlGR)4pE<g{Bx*s&w#XlP}_H7vi$JeUYUT{BFd))UJZ$$+*u)7W460h_@2DCYp z)xX)j?Nw}Ei5ITd!R@vcASd6<izWeVU_C4cI~eq#C6<Ry&A>X`H86x|j(Q(U6T-Ml zLLF^$ND-geS_m$=6aSMsK^mGmPgyQ2eO!2inj_k3Hh;_B(9B|BzKD+{_SrlV7$Z@I z4g82s2od+}&V9Fu|M|`G!Eg*@gMEU-=qR}nJYAUn%`Q@Ff6tr&d@K>gXDob(YiYF= z)gW9~OVhNd!Y~0>$P1FKASKS<?;h}w?><X29b%Eb#ZjrXu1?zz#1K98noF%iJadx- z%HK7@Uc*5MYtN;^LEqAv3OmUF!VX3EXFE}z?{?lFh6mdH+tBL~k_e<)<FIja2>KvX zlq#_|T2nEb?>!E|aY`8WfHu+aMbulIgAtmFrRV4jPts?A*^6Bg4}#;dGxu)+$2mI^ z`er<v*vu%G$~>_7F@Z?@5rk9P8d7pDAXSt0$=neWHz&}v4|91!Z;<0jVq!C_wysSt zNOldL^3Q)p?za<|pBBB2bH%(ss<F}vzoHMrZX0H$<81M!mPf_pa4`}14I-P5)R7Of zggLn7f{$I2Cg(>+H#QKBSie*7nMSb51E<*Y(3oQ17go+YamP@AjN`|nG21w#l)!qp zn=sav<S<GTQPIHKwuQ38u_v(u|KINKa|hPSY?ZIXT$u~t6@ZsHs)UW(rp`k)G<wb; z{@1TzIuB!uw7JS1gm7<8cdPfpTjC@}c@5IRA%S5lVXQg+nbn6_5c}%AlawFP<gH=p zyqt+2bc`p-3?r#H_cpADFn%`Ct@<34{%oRI^|{NlXqPDptx-#2nX3)ETWb~%&s08? zkyDi@>W3%5wJRqQkJ(&fsPnQFTPXDL7pDT)Pn1~D>Lu3soY&HB2#@)^;$oF$rS>k- zu(vV=n+td)Skh+9q_-u?dLP#;86KyxcUuIt)(jfT+t7TNfqy@Ma`p`J%L=%VCX&2G zuKwlEfL$n_e|Tu|2Nq+EVr$Qj=wZEqB&`FM;H7?BQ>;%(M=oqhX)5a^UL^ib{=gnM z<0^L6ztgDWj01PEbEz~Y-2M5f>vfC<3GtnnKT{AdWIm5av0PF_%yU>@8hY#8Kx$j> z=1d$uf~qF*2=xh6RY&kG)^<|aH;<yS-))Pd2PjWCPe1J&y@7PFrlZBj{0H2AyJOF~ zV^30pMZE6T(S^4Pk{k;Nf+ZFOHh@;(O}x^nt&Gmgo#N=Lw2mm4gj{e~T)!TxD%!^q z*8dZXO#Ob9>u!hT0QlwI^-Dyq76jyjuS7);I!;~Mc=fvn%c*s`eaY)y`)IwF1}Fri zCvv4;=@gDyRi;f9x}7GyCg1kv)j4p-Qo~*5>)=5i%nt5hbR638e)sJIFDBg~dDG)l zpF0d{N|XoHu%^>nvx&(&v~#kqoA86NNQZs)tK+bZSY#X&+_Mj~fAs5qgieX8&v!jm zmzZ63>o+;vVr4z}397XIs-?7Py`zKY%;(1*jHmYGa0^m*e@?x;1N!#=vWa*cQW=xW zUC+h!uc=FJ9=FSOVI96%<jA_bJcc7K8cf^d{q*|a?8Kkvyh+=C^-R%(NlS`Tj|b>& zSW<j?*@+2l;U^|9wokpfF|2m_`jM&AWyfYc|ExHtIA_`D8(-}KuY)~@1xSjv`t6w} zbjl-E*!4kN*e7UQF8R@^ohbSGfIaN!{d2nx36P6Q=cDT|oo3Tyc&QqG`>Ger=6#iE z2lfRG-VbiQS+N#7{C%G{8lk&K+KC&irKR(a_l<9ia0(^4%Y)pjb^y4+F_*MO;Qe1> zd}$*sLYwgZS?z~T#x&#wE2YEZ;njc5>51q@LY-7H-;ytRL;F^%M4``*U)A4W-q=IL zd^&u9mFrqM-#CV}>)t>je7;1}bGWUOFIFAJamd3Y6};lnVFEK;mw9ta16pA&-@nHp zPTF$VV<9blzwmj%RmTmsh!dy&>$I~@XJ;duLQH3TS8;V3A5yvpc)$@CFZ-~GX0bA@ zuxBUKd%~OMbfXwD!W6y+Aomh)Va4rsuK3?b6zn>9BaV66{nxqTj(~LQNb;i1?Q$h{ zExC(EM&dYO_)*lNdlPvDA*8|9qx0B#B}VFYe2zPl(%U$Nb3)?5O2NHOJEy>~qvQ_Z zgWMg{h-{8f*`Om-KZsMC`b2RWw<A4Jv(aTDt*|XFlg>fGqm{?cn_j`0oh~lslzjJ) z-np|QJEgJR>0ZRq(zE2mOo2rg4vGCODR2uwwX+BHJSHP?kCVVmMVg*yL`IkZ^WA4} z*OUH;jcA_5RHUmfLWfvFCp=+8hy0*J-i_sw2xr>{4W1JTo^}*Ze0bp{23oweNm_)9 zlI-Q!a*>3RsP<E8!B2sp{awcScN9;9#!|-jOJaw87NXFh@m;2p&@k`A8a#IrJl!dt zbM^W48>kKyK|XieAs=imTU_7euC^nYpgKrx=LW|+G_z}~C=9O(=nAc1yt<Z1T*7KZ zLlZRu+DQZyb&m?jQY=lgA&C|Sh`3vKn+n5gh*6(h<)0~lS}vJU>SI92estk1Pp1*r z>Z;DCnKws_=JmxjRymvFdK*mV2KKhk?ikkqpXOu^y|m&>dNos4c-J`i6oYUxf#tuW z*8M!ZY+SeRHm=u$7EN{Xfr?}kn5%TVoli#0xUg#R5Qz@ZB3n+tsfdg$H+Y{$+^}DT zADX_<Pi*_>-$Pl4cFclBUK)0ITTSZl)uI#^qU;zjn#h-QQu~PnE2ngsOQOPl`kxw2 zC*qk$Ei!|r^O7f)uA&2beLRR~L@(dTgZRmtw}Na*7TjOh&*YM}kpGao(8+_gh0H~v zH;(lMHrhJV=m6DpCom^an4{V{FBel)kA$jMqj@#D7tHjG63s&qUiau^!K1qZsAW=P zEly;55VcH(s{1%sg~yltLC*0(Eq#c%Z@Ch+ya87bNzD=N@h?Z_(0AzY|M~JvcN7`p zdD;Fp>KlxDmYtr*=Kf4t$SNCgVO#jViE=ZIzv$H!tPQb_$tB$u7ZKkF9UL>F;fk+8 zt`F|sk&5h;*S=k(7yo=`koh|Jw3`gZV#-hw+?<l=CkkC%bkio$OA(-f=%Eb{6i=n1 zNN4$e67c0T6Ko@ut+z0|m-Z+}&g_~Bh#fr+9pBhAGl>z35ym$inacEXV{X5EIGm>t zb&%e^gS+BykSph{B{-_1fg{6JeNS6BX!t~{o})LyC7w5l|8^tH@Rs%Bn-*RuMI|`3 zC-v4N{gyyZio&kkxYYaXP0~0I-YCcV$A8NsDa@O*ggHex(y1vLV<48HDd;re-hp<d zclN*TK=Vfz`_Wt^7G<G0@5z@t(u#h)D752*dXC~%Sv%<qtq(}ogPapnsho@O%)d!2 z?P1ZD9jb|?$o~|rB)RMfy@uNg7T1TTUW#354B8zP?r0=g5Jcg8dxBO6t`1Ch(BQcm zhFIS83o*pR6`&pSUC$})O?dF`x*_)Oz#ffu<=Y%GDzWw2x7V?n&y&9iFG5#QpXBN3 zlK|2;eTM3W++2fZhbc@^y^f+eh*u{wa{ZIErEA8cef|YN@&MB4=HZ?Yu#@L_)3XJE zaYh*mq?@)xKrXy~`CQco`N{A>O&CjIfpC=>^=UOG{QBQ1h7}l!R}C|VQm}Yv@IJ-p ztu~QP#+hB56rku=RslKwOf1_*ue+QVVSL^u$m@R+|BFc6jw-m92`*t!^lUJXsR(78 zs8mSME=Nx3*bhHmwRDlk_h{I89ZZ=fEK8tWtEJ*XwYK6N&Rx`4Xf!C2aPc0McFKHV z4*8DVY+{6RY?M6v#~K>M1CFb5GeUu)dFSDb&`YRJybh^DL}8&s%~h+#g(;IY9U)SY z6<DaoBP=RXH8`5_>+w-TG=_-f43lQOwlLf$2ZkNldGfju4ExC$)2D06Oi(s$U5!JM zPFh#E$mk;)VW~<qw&SA>YaE$Fyd+Lla6fFp{Ru6iK`Gc4vvd#@&manP`{({hQk^82 ztEaodm5)qQ?Zn&`W>J?Yz69+Siqi6;i=!?Y4Qe7@vQ|vIq{5u>ML(RhtsWvMrn%$V zV-`Wu`}n_Pv0{3%Lru0DkbGK(trxU?fMhMxT^@CzaH-cB^-`J-ly$OU;2D-AA&_H5 zS&aiM;N(t5_tYos^1g|Id%RPXZwh$(#+HVVq-u=_=jP;xi8@vbnj>nLc;%(os8TgF z*BXQ7YAyx01o&A{{II6+WIhJfd`PdZ)v_dzY^bE~49IJMcf^6qJal<TU0d$M^vi9R zRY=$PjyP)wE@D7%=21%U7h-AZR4SF<@sNYjT15r%kNJ+r0ud67z6`V>Ihxv1k$|Bl zZDN627%$aRdDODZMb$<c<>@8LQ%hr^Dn-P9iq0c_0hL?Aqq(EWT|_|kdkNwWz~%Oz zbZyfkHu(RDdJ}l6)AoP-9LH9+qYc?ksnAB5NWw8Kw4uzjNt7B)p;1xRlN!q;RI2G| zp(uK$LUV>T)MQe6%95GtNG4@MVaO6X$NzoZpT7U!tJiCKp4XV${kgB@zV7RKzpo43 zFQDN0q~NbKTQ0?2eC1YhIFc&Ua6e_V$6;=dS0p0bG;PkT8c5BoD-lc%TEXu=AKYE8 zg?$z|rC#ULnhtppP+dGJyjOK9^#HRJn)}MV#E{!(99&uu(?<=luuM()3QI7NpzRIQ z11UH?y?1Ihm0eB`=q}Es4=;)jNAClWd68_<sD9evVpy|SdW^}XW;M3>bTzvficCaF z&MYr-1>I08fly(E>A!9XpFu3oBO{t3FXNS5pI)RZ3N)lwSJ8;9+=f#{{9u7BhSE#3 zv&X_**8f3%><U_4IFU?bF|d@6GsNeEDKZOC?5lBw6Fa@vCOeI8nLeIOfE<q|q#}LH zcWhTkhiW=pqlD}PrD{;?rzY&0ku8AQ=nAEY%uP6g;4w#!KZ;v4_+HCsoZu5#YK=MB z3ffg5H-To0cyAt;ob7X#y4IR3smGi%+%&Sk;CG}Iest3mX^E0hK94eDJAH$l?lch9 zoTW--KuJQ40WBJ7oNdDvWtgg9HCoh*=h&W1jX7(JIOo{+4Zd9VlnJJuFDySK$>JnK zn5t7RmEC}4!Xq}j?AOh{c+TlJ-2EvA9VyhLXM$|d$C0dBHWNG5l%-0IgR1D|^)RJq z8a!^W9r^I;vv|#Upe-fKM+ZlUm27XQrUzD`*yNIuy7><$>~z+=!()diKTG{3Fso@1 zNxXeT3n8Ig>i73;W|oa6BskNJ+?l>Y7a$I!HXAEOtBY#>f!p1o=;*0OfHpVIYA~3> z$odmSrs!Znp+#_<*u%TJ?_B!=2P#0|IH$pKg7V8r=<VN1L?J*9iOE?=_rr3A(I~({ z88=Ppt9`f4;4QKvCUzzv7C;sV0(D?=M_Ippj&y&J+S61_H^PU9Pr`?{Td)sLRaNc+ zS!g4R&yX!r`YVnvs!65Hxmo^1X9Kw1!_SUnrof}gP6dV{H!_31+gPn7C^W}-fkQ(0 zd`_3pgdBYz6Ubx8Mo2^o4Cr!TFL-x6!_YRT=2bGp?%FPXg_OXZmA4KTS5UqsqUIHP zLs-(9C2G}A-7lV5=#U&fm0Po)t(nNy9OApYWSJ}Eg_Qns2Hww<3Pjmgr{q(*w=VP{ zy2By^bRc{lGC=fHZ1<O(4IkgWK&n&Yl*;a#cJ!J52Ph-b%|6UDabk({z91~PoS}^P z26KbQW<MqivbB~Oy{83J!>5P>11MVQD}xHHlN~9_i6Nv2gri5o@#TcG3?HqkE@P2J zpf4he6-cM%Q^F>(-+xyeE=5zP(>3x!pjl+kf*K*tQ}op_@EK3eu{l=7lu(_W9ixWW zqA!j`0%tVTn>vxR<HX|wEn2VTCe@%D>{?olX&Rj6wW?avG!oAr=yDO?7elG5nfxB) zz4WSQX8HWYefexMpk}tF4<5;@kuPE&j~*s`LxgH_%6-n1>&XkD>NDk5@<Mto_Ib&w zfZ4*lgBd|<_>nS?K<D%Fq&^K!)z^^wh)rhV!PWf|Umwau9;@MYrz8}qrQb;1J07)| z6728gnrdfFcwlK`O{|VO5MTIuIlubp)oh;yw-WkXxkbTsM`DRPQV$#{)M*8ULV<BO z(tF2%-ZH>WpWI4U%p&)3D_${+u1EQyz6cB`WOBC^d))f5Z;i}1XqZ#@J~WlvDQjn6 zFUQ-?neTE-z|!9CVx^7}V-c`j(50PIvk8)U`d7>BdeSBN3BCxgvzFtUBjDFB$`kuc z6S(^<K~sOz6KfVxE<g<?Tx@-rXdicXctAp>R-#!i!AMuVDmUqWJ_g$U+nTVn6|nR> z1wU<?BhvqiG9&ktt8FfC(YHx4_>1##n*`Wo!W`P9-*r1tQXF|=zvoCf9J<Pzl(2#> z$Y<y?Ky|_Vw45@-X_$d9CSjy{f-2lw0>9ekW#T6PghK`$+=KQWItBmFz?NMX_v9d5 z>yj5vKmy$w>6uyEQ(=xn`R5&g`n#j}u_O92_V=$}LL#WT=zdqcWsiPV0Ep1vX`Y%j z--xMuf-7WWGk+9?8JcK<g3|o!l_VQ|3M>DsBN%n>IE9trlNtAH+q3l43OqnHy=~^J z3N`ELvG}@Ba(2_Sgc<rpetV3Vk!z-@rC0|%AIBE;;3@wAJlI*5N_`EW{pERrbp5jw zUv2B3cm_odzxMcDN~?ME{Ps+}_moFWa^T_qQi@DyHz@Ldp1R4HTNLdlu&6Fu)PWy= zU4x{3?+N0UC(4r&0w_DMPv`tCU3~jKo!ncxc<Uqt%=&=CM3;|{hyQLDSAj>Kvg!W6 z)Fk-l|LLDa2dO6o4ir`D_kmFTgnjo}T;*+mQ{5wr^n3T%t&w=39~UfX6azf5$s+;J z^i2c`n|&MNzMb@1MbpafITvdroPc+;xvk0eID{0P6_bw&?C~q{=I1(IV(aN-P1Yd0 z-SO_20v{|AR9u;qnX&?j%G>d;kAZ%Up*S-PHfpsPY<#kVcaN{BMyc+%$;<zm{o>?n z5tMO{5ae;3pBtY-;Mc!Fy~A_|zTeJ!aL4k*k^K93%<F6W&ovTix-$8=iL3q{pRm9{ zd+S`f&lC0=;D)%z$FG{fwr2;K>x+LssNbO8!n^zLVzD(6MLEjq3B?rK{BMVmNFUgw zOQ43c{H--Gwh6_8_zZC5S$RL<DE=#VaiZT`1NfKM#+sc<yYLgUZ-Y8|@05W%n$jlt z8@v+W-}P!@WXsPspPgNHDDBy6Z>9Zd;C5sG>61@VT6?yg)NkyMId)}7$oCmXCgs#E zktcun-Tvjo@$Y{d_u`uEp}+|v`xj=M|JripPk}3gR<4_NB4dX4C+Dhq>+{aqXmgS# zi2n)=C%Q<AEgoq55N#r=$viLuMKAlbju1HO{2P*v5USV2{e(U!)p0<BjO)qdgL%9p zZlXN7CznHe#}V3F9>9S8!}j9z1oQ>*rTe)@1CReJDw77Si27C<Iw?j2%2paWIYwgv z_~E%v<3<R3eDN+)ksE+^t=H<2?R-eFoZ)L*%Lr8*H*PVst))s;Imn*QelQ>DiGexj z_Azq`9$Jc}35k^E`TU!v4gtP<d~?wcqos`Kbb}xH3kE;7RGo89R49E}QiYaec>ESC zC1v?D=Ba(rned-bwB_g|GRp-I|2;xP5zBP8>=CX891o+@qA;TjVZ^u|jCtufK{RHE zo-d>ZFr2}UUc(qSibRPh7|lkeu|klq3hV-DO$CE}e{7iUAjHR?G5zkvK7G_r%8n~g z!oe&9>@<^?(Bn+qQ@$TGv}PnLOiu$pdI|&3W0XNH=;=(kvdYkA;LlDph|MIVzj4{t z7YbYf9iMC-^;Lj7Ah&9X8=*evi>o?b@eF5L)f$>zc1RJ<9Y7}!!0Iwa@C993ISAA2 z;#hS>L5zv{rs*wlX^_7|?6K=aa7fh~)$fPi*Mm_l7)?6pOzazl2Ucz9iKuB@G#tRK zicuOkRx2i{;gRaKa(1l4#_5vB9oT6c9qS_JO>>ch-BDAe?@qY)VXMJ7@WY+qWwUsi zt)1D!YI0`_j&nYl4^GECM!cB~wl#|$XCs38;~kRY#0)PQG%wVGjIP)eYJqcOVOBX3 zSXbIznDd?VYr0zv@MkQ8EoU6!gPQ2i7PG|{`;%T~tibFbU%p$$<5?hB{X4o<L3i`N zRkd<0Gb9WDL_*%>232J(ba<MU-k!GE+ZV6f&*)3d`ylTO6En(-E2&prl)!i_w9+Nt z>R@u$Y9dNhspMe%i{^)&maQ8=YVb5!zE@pjp@UmlrFoNN!q#n-`z>_N&*w-DmA158 zFf$!gk=#IuG+__**|_OjUxUf4JzAWOjc<)&8#l%?sla$3cdZh@EKlBMZq@?s*5ASb zv%xSYWrjtXil$}S#mg5iU-*!M{mSQ6&ZJ*Zv1wtP5q9jucUceMtsPs-9Bmd#F)7QV z9Mp+>5f_IvaHRM8vxYS2Og+xe2cR|=4l(|kISAPMuQ_ausQkk7{XmBF;{PXgOD#nf zLky9)9PgVs;E+DW!TM#ExCPL=?j5i81*yD|cpwf14*@JGBS0x3LvA^=vM<{xi96VJ zw9x?8Xdp}|aytbG>ISyYcQASB3_wsNBSdcA?D>cy#6GRWqoH?4i3EnPmY1UEi?OHn z#oNF;o**^o20OR=>&k<cRu}&VgF5f)dQt$L0(n~g14<+tk>Nc^O6sYNV?MeK^^v7O z*U#<V`gQ@;=Le>CPhcBfymt;8UIo$Zz=~5V&t)vIe9xKa?Qt;CZca@BLaF^;?D>j? zligFmC^Z@3mSIQio<Z}&ap=dZ`d8;N#jL;3+yC~lZd?4QHVhvS(gl^?7TYYhS-QfN zE-Xt0)Yi>gi6c0tTe{*5Ld1z#x7Bb@C@d<Y_{Q-Hx5^<<Wy;QU#{%A8=`icO(IKR^ z1riY;WEnaP03D4Uo-0Z<zOa~e9S-X(hUt}eyxmXCopm1e;|PNd%Zt{u&M{grpx$hj z@rL`VXksyUk6`wKODu>bGb@vkuy;bu$DHh?scP0ss=fZ<``vXzQWbMACzNQ=+b#q! zTkD+YpU?q(QspH^FIwe9^r8VHat;=m7&y0($rLxd&T~!RLP5}S>ACf54H!(SHJy(a z*ZYI86gICJ6C|UN$sCtqgCo+j%w6f8eVQW>Hcul(`IX~t8wzu{lEOBL@_fG$=@#$c zY(yq|oD5e$kWKqe6x1k_NrS8mncJc2>SAwdopajNJ~2$YnY$J&8<pS{H!ylg<s$rE zEee#HH$#djVh_+obJ_W)=ox5Hymoj(_it-)F=O%<E`<$<IyEYr!CoH4bS>Hx9J0aP z8Kyia+toWsVPC=KOqJrlL4JJ{TXZVHew5NJp5a0ib=#~{7*XX_?bpxtgviHmt$7#+ zMokYGC%Piz7t1wXAsi<!%nY||GrJNzLsXS70cy|HrLI6YT^g8#rU};=5@r2pHIjPQ z+iRwK@B5a@6j`4~2{YJ@(7g>xyt6Vmqu!zzZKc4xy&-_O!&r9<3tR}7pm5$yK#efE zKRjR<Q80|_y`IOl;L~0XPC2esI1*(AlS)&R70?@OhD98DDMWa6oPW7!;7c!sGWz@u z#*AssmxvtI0d!?4mCqRbaG@Rb6w@iVDbbmrj8z;I0+`9N2hXTcPOtMs-^4k{kh`pq zWQCKCeuq6<O`Y0C{aprk*G<4F>^4}Ok*zDqLLEPdg=B4bAzWw&sdp5-r3ps=Or?*g zJA}7vxLc_S8Za8&nl%uKnz6Xyz)kR~FVe~(s`{%sU5&je>K=p01Sip;gwwb=-mJ0) zPNG;@xfMiPC)O7LbqfA6_z9W6l6vhu<~pV!5>*_Hwt-qamqgnT63!{<CX(RZNL-=7 z04jFgAz7XdHG_yX?mY*stxM$ma9yGxem(p2VlcZueneMbH#+7IM>Tyx@A><h%Dof* zG;`v>psyq=1o<G1wro&+v-8rjoZ@dfvYV~Wn3#&&*wqoEDRR?MQ{HXPKDz?=pr$e6 z^BwWH%e@^lpA#K<b)v4fzZY_UK1A-`P3=xRGkN-}OkEpJ-n7{r!LIU49?CF<j(xh6 z<UPuBk*Nyna+$~NKtAy}Zj9Yl*yw>@)|LtY#JzsD*@2OWpXMQWzi9<^$n(0Be$x)r z?%Y-V)Z7xZ&s{|OR4^Kn2~WMY2xrP3p1mg}$}i)*Q6kZR;ubc2b379jRpB&u|AFMS zsx6X8VXr-Y7L9<$)x86=GJvJ_W1P6J*CQqK<%rV&Cq*LDt;UoLqpBo-W+I*Xh}hUq zqSl2|IOi~Y@7jI}JkEIef?ht2oRvHj4DamHjKNiH6F6t3L&USvQFOc4t>MS_xMA89 z!>(ac71CiN((~Taq7S?u5CC#+cX;<Ex+h9g2hqf#kx^6)A*XGZ61{YCOR0s~eXq7* z+G%47J5z7$bPVbXETxJ;sQgpWrs*JT<D$?FID%o8s;Qsm?M={y>>Vz30%L4Ja%$Al zUzL_qx_?j`g#s2Img+%M2U-6dG`Ci-RR%ni$thNdsH+dK%0zP2x2>$NQ#o*N0RBeE zU6-~hjjy~LgMMb@Re<fC=IOjpF5E5Pjx>$puC#kZVC;PZ17*z&Q0L%#&qo2zoWi&{ zCv4`A#)5O6S<fF0$6DvH${7Ighh$E%<E?YfE@Y#{Y`EhDFVs5w>`&~=fhBchuDX|c zSG23U)v3S<ID)`$#cf<SeV@O($oEIwm8tf}qQ}x*8Jmo%fM*r)+j^ISq&&WNS6gdc zK6h6}vaSK!Q!HFTcLbT)4akv5)<LT627U#gcX_@Q%EBCD{M#w?cxW4dANV-G`xv69 zzkfb)IKO;6z94iZGMShJhI1~%IhVPVNFw>Yn_&+v&sLTRaC^yTa_(j$n&PXI8}E&x zPe70CM|8+P8$Mj)57*@%GTHYK75AL85MZ?F&}Adn^>dRu$qic?#U-OS2lfk@?Hd1u z_T%wGb8rN(v2Dy5<O?SGJ#7S5*&K1?j{A)U7@WK1@8z0LD)6+aGrVNSkB`Lcx*2)S zDw>z<Jo9rgI*V)S(MC-@>bRyJas!bY;z7V!$p!kt7f7dxK#@l9l%_&ZzY;<EzBI># z&pW5{Xa8I1^U%3Mj>>b~0*wChi~}@HqpcG(>-IW_OA)I&iAo6X)i2R@iJJtRgS`Gy zKk0Ml{iH!P*jyaWPMiHE5ULBLhN#OD6Sw!52hVUF@9RBvB>ygQoONX@hhcWk%<K?O z%BTId(y-TwJLF-CCgpbAAuH5`IX>(Gf&BZ52<1uZiH~ZYA5cJ*pramJ3PvO52o%UL zcxWrB-3%}lPaY3UF4ARMEY&6IF>9g`h$b5B6wnJ?dHW!!bB9P?Z*LiVh{wUcEt~aR z1v)I9EW3zQKHcTAe<TxvN6YR@4+VW`HH7Zrev9jwM?g#M_@24qYBViGk<$@|cA<fv zyqX-s^Zp~D<%mTqVl$%6NO9-L_Vdsx9Igwz;sxt3oOO3R)91%Nw+6lQ{1!)9Ebgh8 zfVzOb$n6#|pe)Xk6_H-?-K4H$YlTsVL+{X{CgM^)0^I5;=<}-_>;O%zHqp|rHCaTO z8XU!~8ct8UD}N*@<^qfcBX8}(L@1U+>RSun6@y|+ICbXma=PnBiP+rJ|MQEZ&MqJ@ z13MXxe4I)<#z3|IfP#U;({PTn-?xxBq$bRHJ0i)^fE_dBitH?zm_d4A6um<eetl59 z8RFJzoYeReZtp;k>~@U8Dkb@>;4hH9`sN#{9G(kzY}7oGPh3h%+5#|Uc_KIDyD+_q zvYP}}3i1H1mF@%E&&kn-bhqv>9k8W~8-P<_C(n|C%0BWI$bP-#-}ZA~H23uzsR=^V zgfC~Ri&d3L)g2HBA2UQ+dlhc2wYRi>Af*0pqnqYYQsX<RJy&s1X&;O4-YPv&z47az zP%H?VbRv`-ZOzFSwjTosmGz1V*9Q;K8S<g<?T#!!XopB^?a+oMRXRv-S4^ZLG$9VM z7Ypu?{SuUav4qcV0nD8Ob3c@$&&J_qMcZobMPGs!{V^Z`+S}>%*tfr&kv&Y7y-aDw zLAq9_qRxVrR>fH*=p|ONzt2`yALz=p972O@i9x#~lTL}AFvlFZ?^rTMX*M3XKG>)Z zDN<!|6Au!f)_71*fLU_w_oQja{jVfFN(FP@^n8q@k1vblKD+|`x@v4$KBI4zIQ=ZR z(iJS9;e>HcEwjbv&)^v*n5KR26)vmQxA{^2WI{&+GqFuC+U6Ne{aSZ0#2wD|1#wFv zV`QCuCZ4VV*(HgneFUZ9`s|^V^+2o-AH;g^E`PT}4YJv3^5Ac65{`Rz3<&1{sW!xi z&+~Khu6FVmn%};R-7B@F)PlxAK-V;YxCBY1D&RU8`3-coF8v7$#yRfaYKPR&r}ulO z-_oTm?hHgbz;L9vW&|>FbM~1P4yXFGxGgqP+;b+4ZBYeVWN+|rz$CpxaEhVykSiG~ z{4f&7*za^|>l#12Qu>T#5lEsdlaVijlK(;FhMH2Y^-lY2GMr{6nix}|RD+!%lGuLD zxoEU+M*29U)wCK5ZYx{y@+Xq@nwqPqEwQG7u{gx5Ir8+r1*jjeCX81F*r5-%Ll2(I zyfJhU0mZ6xKMF|R@37w7w?4a0P(A-`_O26tV^Iy6A@#-vF;R9L=3*sCGiyUIIBUqB z^OL@$;jzi=lTo>lsj&sak+D0i;wEJz`}fO|jPZ=x9-~qypk_#N)AnNYd;ptiob(n< zFf=ftx+6<&!fAK|?jUItdzaD7tYQ|8$uXq?ohqYwOIK%%5IFiawAFw7&!#}(E5pRU zfBCN}|H84qA57eQ;9Kcqjggyk>QCQE`Q_((R{wmfH(K?jZT<+GDdh)7Jp5^)&%!0D zW5#^0dhyRuN%^2JGR;|z3D1$E5=Tue)Bz6U^6;VFGC_PVkvkp^K9-#AHC#8W@jh;o z+?2}DXA2b`cX&Jh3S~X2-UF=LBNlryN3M)k{n7k%3u6?vrG`mYpw=?y<SWmCTAhu) zds)=SZ^;W@i{;597RlG*QA&DauvJ;#&;5le%K{yeh#vD?UNBdC4HO6OB1}k+Le;%I zB%m-1-S=ER^l7Ztc0n>c%&;;QjoXw_tEij`c%<0C>wd-@B-iz;EKb2cSQayfvsDAp zJR@zr01@@EDIVU8|MomVi_kd}ivsSghK+dsWvL>}Ky!j+O9ed{^((=ZUV4&Ha(foP zFt)yP`kEShbQh~BZ`C8GkWm7V%BOka^Dp^c;M+PMnoHSOh3|wXU!wr~B5^h*B)h|c zb6y%$dv&<IICaf#t1q{?Aw-5=NHD@1VkmpM2{K<CVHEWDf-V$b2kZ!DFzi~6`R<w@ zhDSFAJiya4BHKt)y%LvX$P#<yYA@~cMsD`$VLIHJ{V=)Qmm2I0qj0l&q8k{?MN&&U z@6~vEcA}7y<WuFh6$mev4}+su7g0Bpu(-<C$RhDjlgdy7(o$}bxKca&e}oR0pFJN< zdro)zvGfIYS$mHKRwk-=ygUXqn($qli-!us@Faj(JXdS~j-+WbwaCkwKZ;g#<wsV= zd78NIo!e}|?b8*RBQKNxYm*0}ID5z!B42tI8N>>LsVsWmvj&N->w%_&zZ8zZg2@^F ze3z31V7SKX$uRdcaouawyo4#@N=)GObKlFlsaklbD7WGrJD>4Snhj`86}OB4DwyVn zUo0c`ZGK-)86Ts%J<l$%*T!UcI?TN@0RTSLy}Z)KvWgf{T`Aqs(C$hNW_)`mh}?!L z!^OTGqzJ!WZ%D`k61rZHTSj=@AB)lLn-<-;*j^VoGumnQNU_IWOVGDZcy1(o;2VSH zZzcSHGS;w^T=#f!dMc#9TS5uXd_Y1$hCBb+nHFvNW^fzyazoGF!EJ@cNJuWpK&aaN z9WIzvuasIWLvwEL>sGDwrqb^G?AvHtdAd8CiRcwyy%tA*do(*!A2{$AV8suksO|Eh z)UWwR*ka+(uWR>CpfpqGx!dTe*Uzw#v0@MVlieYew=7MmyxEl!HquXG-@Kps+h}!R zjw_c4qqpEws1gyRum8QE=L7-f{`TrQvB?bkJkmW2s99v<mQ|Yb9Gl}Ta#y)uh#rNS z&wZQ9T*T>?nXH{lvkkIB%50CnLPoJ;(=9TIx1|BuP`5ahp{Kqnhp<K8oYNo`pMVcq zGB7(UH+o;UJ!sOt1PO(Vz9ktwV=@|;Cvrn}FKiy&T*ym}c#df#=c5f_nC?R+`{_7O zhvV-;0k#=o830<oBOT`x<(`aA&(cK$<C=#!b&5sg<37M@8WuoZwg!qlpD>pFzz7ll z7jtyW3!nn@tA{<#lmvX0*cz9D=VhgpMSy@Ntka_Wo{UVd*rfxO%^NtKe@FKIgtK=g z*gMjDi<sccfQWsQELnJiO67e+GrNp&fepz|f6}(aJ9xv_H^vAv9?|FDFsk;7T`(}G z`B}y^&NMfSppCNjtmx0D+^UaE7q?Li2RfCh?Uvz8h}aqC*EXT~)!)b06mCl4W25MY zh-^GDXfxgeU!3+OC&nGAIUQ*x%Yh8QYjO{q8z2qTR6*(y<dphi=8lAnicBk3J}38K z(`?E$^dmO6Bbp;NTEu0)B!~bs-WMAG4O1djz;IbgJDjXi)vb;X-*GxlSH1{J`K0W$ zV&Y}Q`P!F&?eQVf+HU2ROOA8AY<3&6BgDKi!T&DRvt}+JDd$x1<L5D#-e8h033dm7 zj9O@9g^2!wsOx5w&kI3T@#g5=V}*5dNG2ekwx!k?=Q((WQYrh5!EOyG`0fB^HW6HS z-@)+Skq*cyegFX(70&&6i3MSr@NIQ!*TOW5z`_l#8-adyWZyT{vT|3q&D(t@GK$+_ z6e2yZeio7L4n?Ggao)zvWq^d|cd8%nP<LpKXX(NZfB<KmcKPrVCTxG{W*0|$)Zz8i zIe6vPVYnoxi^a5JC->gphnFoNuKQW8tP>9SpG5ws-VV~mkEZ&sfVuqppZEA1B5rfc z15vPdBi%i(46D6#&e6QNNTLnsSLO91CzET!ExJh}!Ze^ZR$%h%MB(`#D0nZQ7JXE@ z{W&pCj^8TSWTuLZ+O**b7GsTg?~)F%eQK+8a!uT0EKnOy38+gX1vh2ARNIg}_zyA% z$}B~27(OVLM${44#x5@7wJ+rJ^79V~aerV|c5fi^a4wDv;8($mY+G69y2^A(%ye$i z3brT<(b`5M(ZjXM!%eaMv$1HU9_{k)X~}ej8+}^cw0b*{xwo`=VV7tHHUkY3d+N+r zz5^t-l{?N%J}f3?3f<y)RmcUcW$)~FVCr1}2Dl6BB|VPVvXywVwx;EkBT3#dCYbY{ zRZ6#Got{!8c-$2$)3Qs|nNMpwNJ?qWg8RWyEV%!FOzu1PO(CFW+U2wT$%gdVy*ru~ zvt2H2`vtq4L%Yn!S+4Q)$T{m4GlTo=XY|>=+-JiYB1aN_mR?_ui2dqcjdZKC!zC&D zE7k7;`YtE^=D)7B0JlV&_oYxUXWL2YrGjd&#;;>W_pYcLCSY1u%Ah+eyX_WpmN9mC z&w&lRg?)cNF%^a4^QAxss9xO=u!o0t!?gY_dw3u(c{+<7RgeqAyWz;-1hGeqb8iU- z!8sEl>Wx9K8O*!Ri@6I>rPr1T$*1o|Y7@`(<iQ{4l4#nVFhnKgCOggkKZ>A7nK`i| zg>@ceLW5vT$d!Bj>(W*Z>v1O<Kqs2X<;*@p<&BZgVr~O2<BkjkJo#++jSXE!2}6{X zZ@Qq_e|HM<A^EGyRMp#_(QPhQ&B~4aad2|miM(8Y7QV@+z}9=cnIDQpi39F`(1(4z z4FGRIwB<4M##lA$mxP2F+rRm`9r<5{WN-m;I22#Dr2&6BRM)w(%+-;GZr`_Q^T366 z$(co%y#d(jXk>?21bi0A71^JXg^q;Rz8Yn7?<j(}4gDqw=9|2+)sJ}~07U{3H{%j% z)C~^Q`bDGiY0)xvOZIMtMzL&ZB(9!qYk;sOF2?nLPH1`n60+lQv!=%#LfwvR!;4x$ z{WR*(J)^C;ZrP<1tA=&BT~f?Z*T`ay$T|Pla!~<~{@Wu!z-Q25?tk3{AAI%2+1Fjq z`B*QqO$+N`SjTa6^^0uVUAHaonddMUsRKBR^$a<Y8;aa2SX}T5T>)gB5V;O*u&<Cy zyUX6qj<#Qoz&iBYtC8wpUPCy7mP^Aqua=jb9LH?31}$3cuP|}Xz@dyLpRpc(NcOE{ zVLf3wpD7N#X2Fo%mQU-^^GJfOc*ulMi-FoAl^*F>S-NdbFsI{1##k(vC5B~v5>W{Q zUbfrfzQL<0(1D^rEmQwJ)aT65sp{Wvw&RK@T3nA8c`)fsy;4iiFMx77)P@r!s4Jz2 zu880@bVVSaqujceuI#|~SdUoEgpIYEFL`Z=vkTcX>7G{sBGurp>p?SQ9W-t^X64RV ze{QVoEcc~-sEXybbB*GVLiJ`e;d(U)0yuznx*B%KpBj`>Hn`Wy|C9R%B5_*<JBzD9 z=mPSP2Lzf;@X}MY9h@YBr(t*PSD8{9tE<<31U8dpQ!!2+I|L2)7RA!NYi333;EUgf zZuw+F&L><)=7?7;H5v@+SM8B52vmEjeX&7}sSBXLg8v02ay`?KO0z<`CYF76KSI>@ zqY_ZQp)#KpZVAW|8+?+FHZVC9%#uH9NS2@z2<Y56N1YsSkWcq{vMVQ02XF0MM<bSB zaOK7KOeEVwQ$5(81{-IWzRx-O#A|{;e|}w|nVB#a`GxF2SI~t6X~k4P4_#|gbBvcn zX)+%vrL4+RlSKB|euEEJBFEECWmN;gs9?a*_CFu+xcmXYZ{aQFX;Fl!?wce#ih*Sq z?E-Q*G)JX<t*lU*OWcb+iz$m{-)_JT6Na1vj8)%uwoAGmil9F1VY1cRpo3DK*B0P! z!bg4#tOEB{g?w&_u=?rQWe%y~M(v9@iEpHri36W~Hi+T7l#m`{8RIAe*fX+dVPq{m z?=<b`vD5Q)!YS4WaxyTvStZmdM*55wDVr8(j{Hx_I8HwA;U9wXStN-W>>{~2OVt<x zCg2D#c_={+59V+AnH)~E@x22GqA}LG1k=ij8QL^l>=9+tK1sGj8o|A6hpccfpamdx z?vBzE1#&mdP$=N5CIUBUKZTIT0%p%whgu$cXER^j0a02>8q#9bPK+FE3H<6t*+c2k zAXB1;cct`ApcNC@ibLv9A+!z1W-2sH?MPtQL;B=(!e-r@1-)E$^}D(E;lZUI)lMLN z=aQ2c`YJ0o=5$fcjG#;|)CWqI|BtkasU;F!=^FtH)niHRQgqUn-Nl-tXidfHgj;l$ zpSL93qQW=kx2-`wYdp!^SFJ~C#W?3fs+<x%C0%bUXE{l4iRg^Fx-5cQ6j;+X3>NQm zWCEI@8G&}ZJ6dn)cO1~dfi~vKc|xfv-HCl>Vh@L-;i~PLvNY~FhVF}g?n5Ie$Ov@C zbF>)=9*)7UCsF>eYc3|m5okH!<tfDu#QFjAC5j!?tMkc)GivHG#?tfQwFJGZesb~5 zid#ZX>Wk>}eOZs+zHRt?F5_n?Rtld#Ivree2+8iOxQr-Kq4W3025bjQ=GN(O6;U8= z4eSzhx?3sXBicvFI=RP#+?4cvDsyeI0M2q4TLsVUPU;%Wtfxj|pZ;m>iNYQG=^vfe z{yU2yq4$B$(9V$1kwI}vJ<@GKaXIqCeF2<}#~~@1E9aV48uDEt#{yl;T-+^0Eym91 zTo6aP^eipM7uf25d(Er!H{;F7<<y4A8=d6!NaJclE;jsc65FKz5$O8KU1XDfN3h}? z(a$}kM?Y)gdDk^jy!jeoICD)Glu?bh$NXu&KNg%)Z~a33|CrHJ#8q3$w+N1EEidBU z^$#&8<MhKmw&pc~$St3*e3S+0)-SNlKTm)`GzZ^4_uavsTqx4qclABF{<yT4hdkP% zA^Q0PX0I#*M>=xKsgsvyuiJW!buS259u5Z=5avL~X}wNOx%D@R-zp0It%Kl-OSRH} z{y@msccninG$HeMr6&}c+cm)AZa?{91Y=Nx#r;Zh_Myr0DfrVNVbRr52RJ#K7i#1` zmS4nu#>#z19Z(k<+<>n)g~p%WX*ZZprPFtc9^ZT6678Rj?!28T@w-OtjFxwdY&rR% zYK!lbq1pDFCQTVu)%SbyR>XFC!%+0|OGfKCiWg16psILt7+v6VBaSKR=~Z^7I?k}3 zKhH1BYD&MV@=<ga#6DyAjz^<r+aGv1VYxQ3?7LDPZi0F)R~WO6cBrY>8uISG-9g@S zxd}9?s8boZc3QV*Pi->R)sC6fx^S4awA_7JkN5fWtIjj_!#2Y^-4~=jIXvT4zx?{2 z(wBXI&X}>~h}*>T|H~LPYW+7qjafNq+13?-BWElPv$*-q%^4x5!jDO<?ADLDFf^aK zmw#BgS6gJkT8(#a0LEcEx*Pgf+~RN<F9fLXcF(8J_Ur=?0V5as*kgGwACKkRKy3|C z3F6PhSs%!C#*DeAJq^)s-vtbs0JI5}#}|OP+)hUdFSL@+GRUjQy`w=NOA2Cq8YU-L z?4*w-$W*~Yd&xq3_0b5c?2o>_PvMn)=Ii@Z+`b(jW{r#w7;!4iLhLiF@{xsAw$V=< zYRj5HNWZW!-QGx;16olSdq^Rpro6p-z4-iXRV-?I`hPo!Mv-8nKa>1i_q<AUX|g9g zDHzxZ8P8Uwo;O;w*NkPRkvu|yF_oJp@ECrMQ69pk<HL=BX)30YMnb2}25RzmFIe*# z<>I?-iGFQQqj54uW3EY8#v6-@m9IEFVe=%~<U@8@-o~sDe=adb(64)-HHRQ~{`}v- z1ZwJN`{B-x5vUemRD?UgO6VN!?fFDmlki6uPA|3BV3}^<uz)G-WKm$V92-OrMtib! zi+2E)526P8qym+;uDx@O@Xenl8l{d!Zpbui1ICZAffP5?t#dYr?bE_cnyJo=$tx$Q zrs{Nn;)6;h93xUQx=c39P{?V}a?7AfvtRa_*MYD{!xFx$Kl1yP6&9x0<Y?UHSsVEJ ztDKhJNC`7&KFW#PiK?nMSmiRqLlD1fJn&bno8=AaD2s~H4U)`#hu5B1^3A)~o&dj# z2c|cJ0=R(?CzYsSVH|}j{=5S`hGr;Gxv5!|DP;(h(h*rP)>9v7E#(viGB{RLDEVY- zlnl{))-y^T0)967$L?=I5NA-SXAsS>XUO@N3vU>3M>!0iv9go_PLKw3nTL@xp{#Ha zjYzX*#{c23O2^WpY;BBTA3ejHhXApP?_NOCiN%0|Qr8<lhG{fca33B*Adb(<WZfCy z<L*Kxr<g_kpE04eCKyWsd7ydug%-7>S^IImk%ZDr`4(m99$-gj`5D6m(BRJj<mpap z?#pSeNw*sNas^`7lgJ>kF2~L2<Fxt;>+%Qr4C($rSd(_U3EIEU#+tsCaRTM$nY2v> zdciBczHv0nyxEl7W(UoVc)6d69Au1027RwAfdh`z6s@GHs=&iM6))z0@9a!|L(r8K zb=tx@nx@r;X7#GgjTv!XiQHOiNPSK)EA$zmjzyr1d5<K@(>swD`jhI49ha_;gEHRq z*)apT)V%(4sp#&9v6kZil3LI<zHKW0#BgE`q58DXLE0E_Qlb*H$Y)t!Gz1*6Hr5Ii ze!>WQ51(F3z2DzEB>zU_6U<-$!}bcfo(2FxLfjitZW<spL&d9A#Z&<dVM~6G+?8CC z#nEPw3v1RBjM_~gv$ezeFR7z~JoUQnFL;ERy2S|5i%%7PXSb66s^R@Njdwe=E#L&= zG=}gm)=VI#+Um7002E>|buawR^fJIz(x^fG4B!2?;2*c|RqFWtK0Z6`f)kFg0&n^F zL1nal*C$BTc_&?_ivT<Unc>0Svb)v91?@4+Ph_I8K(_-g`KB4PX8R8ovDBaTp#Q%P z+u?TxcM<10W(vZAJO^Nua#v<zvc+x=7=5XVfQ22b@Yyy-n6nc{&UOE?r%j^We-&XR z9!K1xqLXc-Uh!u1$7wSG+A01fH5t#VQq#P2%PS(gV{rrLdceiT9;4Xv_@<7?jee3T zPFpON%+z$Nlar%3=)kI}GZcp$c9`M`kC_v|uIwY_egbR<zfNzZ@4_?ORTD#_k9xN< z?`Wc|UOT{sOAyM%4+uyP3<;S)3N+B!r2$8AYFMM}4jpCUN?!nfkl-lJDuujFmc~r( zJqE5dvlkQOdKWv6BonNAhL;2%jQ}~w6LP(ea*?y6iE#zJ2i77dUC?Wmyp?Xm(qw3U zUMNYtcZOs_@1eTvORipN+{m6tBiiN_Zks?nso~fiSjuh$zT8*YtuBIYMB8V&giW>e z6^&nw$RFMO%td*f=uM{rqIG)YsNRKb$;be4ASy#BQ6!sIl-CKh{F<OfEOCWtMa^Cy z&u*lvp-qIo${6(U{lo&|_{O$D+0F{c3~ykQF?GM@1+1*%ux^n~v`HnVmd?PssiA?M zh;4EaHrdH{t6K)Rh+o5EHX-EZ(ChDs@9N{}49KD@nNG<VjG;)1oYlLzZ7XLwR%+0p z*fH!V)A;ebX0heW@fM9m-w@|ym}38s@G6>4?kJnKV={Gs+PFmil_+lwaR3Glpyifr zW9y~?a>Njy4TO{^VWmsJj$Gti+VdSsgWCaUdJpXp)<0Rd7cIqHwnte{W9>8QhA+NC zcK8qpJC=_fo{co_o!6$v9b@bWI>yl~P3fzu+=@;lW$elx{1zqtsJS4Ar;ZTqx7fMb z1aDDcvk-qDg9Z&RldFpI8Y!?qfoekUcc6>=X2%+G#u!10Ud*K`eZ=Wb0B!a&dyGSf z_8s8#ZD}}4VZGw?F*O8~<wa%w%`R(405&+S*Bt2}^o$-?hpKm-saKqvd_u1+yJ#;t z({aAy62DeFm-Wc3O%o#sszZ~ff@rk9jUzZhzw?-4O4kjqyskncT*`74e2k1BG{RsU zwi;2_NPwx)9B$5mwi>}6dxHVkDun6FwmPf|*Kra#KnHRiF-EhPwp*6#lYi^H_(}jo z6f!w3ay=lT)Ou&O`P4;mXR098WJ1DhiPlxLLn->+Hb3Ph+tN&nD_6}D;4_K_=glfd zN)00;MhK<HNIr*9mrp+)bh3!>|8FdV$3i26=dkD9HlE9S_FlXXWkl{X(2%a{o2Pa- z;?hEM-bgaS;IF+k)B%XXQVlxRISnU@bc4-gG}afWnD+>L0#C7bU3i>09gTy7Ie^s+ zJaKyFgMm#WW{N)_c}*aekO9HHQ-$!7MOwk$;%drd1dg@Nag-E|4W721vzWh7U9Bl% zp;T}bRz!A!wN(w&QJUm8fNr0s7}VzA4<oS0-`4y+il9Jfy|)QOWz5VTCmmMS4AMF0 z_;a71?2Gepr0H;d7G#9LNvYSmB8$Rx=6VCG2|-FoiNq0R_k_j(J}jB5Bfw+PyoR7I z0O1#7sD+<f)Wr5KX~omCH1z$AW&0(<HW+hRhw&meQ}p8df0(5PBV_C7D4mP91W^;h zOr1?Zl0(Qcg`Zo2Uv5huehcmX18Tz0w=ay+<^UKCP@BLimX8J7I^9WI9k?bu!m9g# zgK9w4vPcj6d?~kH{SsX)|L&-`)}&4CB<fEUVb*~a*-fzTo%B8>*+=4Ef5DyL>HVEl z&UWg+6XtjT=`)I{j{vuKG2X5J!1L~?nzpvf3hun$sPd<{P3V%{H(Tby#!KKx*YFag zCu+Zw8nhE)4PwtHbc`_v;5jD|#V{;8vOda}8}-uz^KsL<T;A}gma?*iVTQ+3D}bw% znqkyLd>=!zY}C-NP0eB%wl9j5PiAP$B^QYu&w9nk(2Z?v|041E5(7l0H4T-uCRnjX ze&m89!!50q{eUgvr8rLd26!9qz>4bBb0u2LU&!i7HM!D8_(YLgCOG1g*mobrclX<4 zQ9_l~l>B(X2c?bk4red;%!2QAIMx?J6UlLy?T!F#jgdJ7b>^H&<;1Q@AfA)#&N-u^ zUV99`o=vQ2m*d;xgrb@!fYiBMmhHRn5sPmA^Xb}oY);6>K%~sWDP<mLia8<gYPsYH zw^KFRsUBxUv#?WOeild&umyrG_m?xWVy0|S`@<Iuaqnz?d~$k&gfMj{YREdWNke16 z+wfUBWKB>m&aDqA!_M1BGNnges`gv-&Ev8z#O!uP)+Ljz`r^oiDn!!d{|R0@I2~oO zK*K;>(BXhh%?Qj?(F|Ji^y+z^MMFlMic<D2@EMCX`u0Lak;&F#TeZWJu@soKgHWcV z<|}LG0ZJk|uCXj&6SBbJquU=Li=;zq`T}6)dp>4gpNeb^=h|XuY-QhWwo~vJTE9o5 zKSs>#TY>8W!)^oFY@tr4Gp~#B{m~JX|9aDCF3_2M@T2rJk@Z9T5rK{8V-EHSI2Cm! z6_vBGuToRIb~rqx2VIkt!b*T!+SdO&0Sf*q#UKFEc(9?dOPZr4^@VZP2RdJnXY@_u z)+?jvT+VsmMy2?^(M*>B&MgQ*p<b>FUFsOHT}Fm^wzF`IK4J}dGDc?vM$jeE7XZ7A zlR8%`#6ud~QI_GMQ=>>=8DJlY$OlfAb?T_@Kzo}8kZ1Zl5Ns^$cRL=A<Xuun;q-Q@ zjmk=-;K>nctp#J#<Ojob+D!%zFzt*{!%fcBzMXMN{E9K_>>Oi}8S3>=n>45k_p!Ox zXL>NIFG+v)7fAVVbUnYctP#Y=(m1k~PJw={)zkbO-zCoX8CE~m3<^JZh<T4VUvc_X z7POV2?NIZ3jQNd)v^~p~d_a9ZD>m`D0h$TRBL{(gilcgPb*Q#}ZGLuRMywb5$*>5k zkz8@?G$!j%E%m|QNO@X{@$w$9k}UMa1)epm%UT^P0?HLM{2;Uk!@!Hl9OZqI2{9KM zkH>ls4s#3`1#$0sbXdh&zMt<hIG<vn!9--dzgKtb(V=mb?N_J#10fd;8O^E?VtaA5 z&KLxoM$S^6jTyv=Pn|Oq=X%n(?^VpA71NjxmcPMlSs{tJ2SKmjpSBp7JBi{yMIh^| zUvKJR6u%V*=ukb35$JdpfVr8AVaXmBL53aO8d1)UP!U8+da#VQ9B+d`2Jsj?15NY= z8FU>POilK>d93fa!?j&g2`-#jiwVd*08@cqiNQzmnLQ3qJC~OkGS%s+q6cjwOIP)v zkKr!NU+!lG*5PQ_wkgIMTM>;arSWbpUAwcK4qiA#O?u@{J%F~$849-%cX{KfSgXO3 z7GF+%<O<rT8jG+LXzSgeu0ANBu%)^Rdo49iN9!+Gjr>h=49u{{V`85R5(IHV`6Eor z#kd#mYEys|%Ti+9opa;3X<R)FHEYUh^@I2D?Mk8+9hOub{vL0MGBvjzLN<$Ur94O3 zNIaxpXpU#{;Irzoal#%3oRz-^X*atiEz3Dx5y&jDwUGkW;l)gc%7CP>Qi@kRW_r{} zFEiZOaYl5X4_a<OUe)jEFFWjM(KoG)s_KKq^wd+}si(NoQ&Eed))ToSWkgu#5foMT zn8RJLLf&8mA0v=fh|fEZg?Y<r$~$nHvQtS7Xg^+^xp>$_Tn(5Cqz+EnV8qkRuj59_ zPqHYzHj|Pq!yCCx>Ii!}S+Xw97C9@xIh^&dUYBIqx*O&QtLOhca&<@YWN6a6r@o-o z($t4|#n;&5{K1{zD-VmfJ?bG1R^!XcgM-U(nBc<dvF~;UVflWQ@+1tU=I2DE?fECB z@7AX^#h}xMi4}42e=GmpUAU3=@5iX&PJ92<yYDnN&$U@PSk~#Vd*ZT#->jNyu;}18 zSyfP@Q-k!C?WK#yzMt4|s{VJ^b*h$QgP;EQyw=#Cq?fj=I-Wi>|GA07X7kJV7WulK z0H|FRe~aHPO>I~7;EZ(8rp5+6?`kvE)WNk*jy03(;asYoX1c_w%2nuG*5wRO#I+<p zVZHe=TB@}`8(Y;t4rd7R!F(jkoG_}v3UIH!8-Ev!zUuUSCNTDj_mY(MLbLGS(O)2I zSXFMv7xpwNZSj*jXd_^=zG<a>4`Wwa%2o^;gOd(=kX=n2w0n5Gd{K{HO&oFXrX1{X zCOmwZvf1i<RPczI?W3eVXv$JcW#~`VH38;LJ?Yl&)q8^PDey`BSf~k`scMPS27wyf z*(rmMpp2QdzKi1w<YT$-DI<rbIH6w{=?Kj!1K!G&i6kG<YP5x@#+N!8K1z6wSr2!v zjeV~72$hoI<BIiQdPDh7wZvfx*>D1Q-oLcw6T<%CM0qE|ftF#3?l0(p<TC7vpJd=; zMHO;tnQ&`y)OEbwBwhm%#`gY&8<L{nFoq4~hj5S%qMFe2Wf;gcJWa0EV^<wYZVhWX z<KAuwy`S{wd!P%4%|$JyjLV}K2e_~YDB1PE9{Ig6{nr{H;3L~No2bu;qOm6eC*>|a z-{=;H2sq9`ex;|lc36!&&j=#W`j!6&9>x<H%4K|)Rti=<XJZ6>01&U#r0mhFpXBfr z)1yrL?yW}3)7$g7AbvKO^@qTyUQju^5{TN%<jWsU+&YbNvVZ9perJO;T`<q(rsn?* zamMQ;1rED193^J!>tm^=qaXfU;El)02k{Lx5!lY7<clla!KIF>T#7$a;V6EdzABuq z><{s}+iY==CsDQ?E%xNN8XqvR@zXkFJoPRDof>TrLIVG2wZkC+o9w^%0W&*dKyw*= zK~2;0*2<10H_R_kXM-(=S3x!wTAU7D8z;;`Ix)<du80gvfa_KZMs20~L-~Jpu8So^ z>E(f`k92-Njzx8{K7ii%3bES47i!iyAC}y3ZdPQl1luNKjzg}-BEW7xae6E;gRE#x zEOEcjM&k4HHNC{Dt53(j!^;+f0?#IH#hPhZMwn=921^^=EJh-JN2LN3>NH~LJe=5E z18<(O-+FT|fJ_!?$YwBKK^<uQK7-se-vld;+nCZ%InO<FUhN0q+~U0jrf+sDL*fhf zPV{rLKM-c!>{q#oQ*|S4xU{{a7GK_&2HfXEpj5fh!<Qf)Urb2OC49o~7d}z*@gJR= z_CF`!t3gHpj%X68B@LWD>u!&FK=V3o#rgE8?KVomXbYsdkqx_v?w#+LN#5qcBmo?p zO|Z&f_dBOZ*>}!<#h{7@Ev%55oeX|-a`!5H_}7&wv$Jn&GQ9+L6?(aGA7?&$CyU(H zAT!J+3}h^Vw!<(hw~@Bj8#n>R=2!3k?@D?^?VZsK0OCe6N}yRaEr%D`{|9|xwYZ($ zruGZnQ@-r<6m|*<Z7~DO+y-5vYa%+fQiZJk<)H{V(6##_{Ajk6zh$Sga`RkvpfI7m zYR$nkEa(v8wWEtFqhOU5?`Lr$>O`$|4fzg>P9PPcG33?>CWQY(2`GD$^nGtkRydiM zfeaU=I<SXcy#v<GX2`hHFf0p0u%Z(!%st&f;oiyHGOjP+wP>97xzngFe_9G4sgaj> zsmI>u4yL!drE0+9j=00C_t*Q7bf&;YWM(nEV89nLJu1=MmFPDX^6NOc_D2=$e>giD z!#IF-Y90eV_kzlY89JU$C<@?0$~bC~3(+d!fLSsfJ!dNoBw44aJHzbekN96BT7((D z3aO;G%-*nif<<QbX#iHPEqtL4MY7j2t3=ys&nLj8fGJ<f`dNmdnDyovdbTcusTYJL z*uZ&Qjy8C52mal{_hK~Q?YJc5Qvg1l(EK%M;J!H<D6N6Bd`4&KTfup_Xn`;X$qkGW zA~Y@N_p24CkTsIyw6U8#mu9}6)WhG8kaj6`9dHt%8!R9ym6NgTt&C8w(iH;}@QKqe zXMd)luiad`h?$U(9TNQ<@Hdy#9_6P@0;&90{sVOy`!?<VAvVO@zIi0jthkX^WZ)@* zF0iqS2Mn6maLV*~Jk|^w`^AvHhrMg4SF!@OXDkZ$pS0kZv|*j$vBn|Prubw14%UI; znHMVE;8~Nsr(vfN8QUzMtx{0*h6BhX*JS5c-qwp#RRHS#8@9?Z*NGGo3cX)Jr~2{g zb*j8Q8EeN1Z#rz25yn~DtyI83ZM_uk0}jpOI1kDyCodB?Te@K^%xPqzehKDJ(K9H; zOgOteLtSsdz1F-f-_l+49CV`Qkf~x%^mI6$hkNLAA_^St-%~C^<9(h1K78IC!Yk}X z>C2e3J&y(rbc05eO1+%^cS=8+TA`gRw)r%FF%oUl^1PAJo`c89O4JN4d9OO*b;3)J z9uu34cYEW?Xwzqz&||pI(V4W5E$M(I+khTz(8VbhzCStL84agI;U$vhL?R5_{OP6* zFtCWSMo}Rl;{2=qkveVP<CE8}b@d^FX9I%))J1_65+(MGV2_>>JLqOHFl<hAicKee zX0BN4J9?yy`(8!)H$D`zZT+2enO8(^hkJ<?6lnRlOZ`OCap!D2uOFp#9W-xc5Ai`B zi3PyTSF~V5#1DifX}@~skLTmiyVbW%fWEcJD1>4*k#$bTf`)@P^?E088q`5Ay%-Bl zgLWhFreJK7!5vr;Wba-ci(|$=q4Z++hSP^oZp>3#Ws3ZKHr*DzA|!#`+RTVOuFAY- z9pRBynkI6Rc0_C4MCOJ~5K(_x10>)WoMiXv=SZcCH~!c_ZTs6sUrOZ1M<Sc+e$EM! zcG`C{^=9he79696`N~YH8O2OpoPpLSS@7<)Gni!X3iXRN60OW62+6Gvi4IVl58Z2D zCSQRiqVR%2IROGX$DOE39IgxS%tVJ-?jWn^Ae~IPyLS*E^blwfYXvj-H4Gfwi#k0r z#acsy+tk9UwqF0XF2XHO-=#4IqC6|Ypn>MSplw;6)%nt-!KIwMV<;cm`zK53%!Iu4 z!0=|{u)wrO2b{6#c@+fvs1@k5A#$yEHB;}x8|CHamLda;v`53J-WA9I@2-2c%6C}| zbv%IEP*7#GcLrOvbgVF^4E)=0iWBN0NZoDF7}gx2$MhNg|Bgq<wanXQj~0eRievP8 z8&GahBkRG-r|pM2|4S9C8V_>j(@95AYRVK|2l0#@y0|jI6DRu(o(;YCNR|Fn?6HI@ zT`@^JvDpyCp0rhV$ZF4`IL=WQ$@&A4O1;W$OE^`ErT!j~vlm}}kph<YX!h)D8)k^r z*ipFIzD2egq>k3C#EpdwLG`|<H{@liY8E(#-%Wi!;uoHQi(Iq!R}=#-Ni<XXDh}g? zwAkbR;pe*0;pE79qCCWcy<fGx!@=j~(b(@;Vuw{Osk1@vl5TrtR7FSiZYM)tr@ykw z1ziJlg*m5KBw2+}C0YyB0>V)~>bt)H$}-(MF>y0W-EDf8e1U-e8k4*i%V%S<v?S80 zH@a;NJj40Z3+2>+s&`^a-aDYeN}$awkLp>b4Q(#Q2Tj4*YZm`b*~Elj3e;#`M8iV4 zEs77o;Rn@qf(akGvdl%U+8c#aTIqg1))c2SDEP7<IOLB}QQ)ZG6}SiT!LQ#+NB`Nv zh~At72Y1Y+YLg;k&<6NhZMV9od1XA!OU~yuE-lH3B*I<B(L7e9%`Z=i)J9dz7rLi| zpXKnf7?|;9f@ozv-B|u*uCaQMr+=JVtR6CSd1V(i0`tC`VoA9%3I$rd-?#mVN)2<0 zU}l5)?ab9lTwx`u%Dq+0_UzE!M}gU&!vy>^FZ<73cYMCPQ|euJ2AW+Az?e%3foAhg zLGv2(;?6049Ms{gq=Y>8pa(r?J%^F_EMo!+rRX$Ne4d3qrp}VP$FZW5;wx3}&92Zp z)>Jl<kWnj#1t{3AC^-~q*c-yFNN!exJ @`{<hZH4y&$}gB+4(2|;`%&gRXPUld z0pu+kCj}7g5K3anMUz4QTx?jaDy27AVL@sUD`{+Fixr_YZYzCmU_D-wrATv5qrQ-D zqh>v)Zpjb(w-y)|QO^8q5h#h6P<L~(EEODY2iecI7u)G7)p#FBjC>^U4^X#OkVrrd zEMjBCgVFTkmk#uri~?Xy)Uj?|>|6x-FL&x3H%-Mv&Fb*%21B8g=xxd2rgzfAtz=iF zK0(x-11@1KbbLKaETK>BE<^ls2Bo}ZN~Vj<se<&U`3B>r9M^ko!5ZHl?c*W+PY)(( z2~wOUQp@0__1ptYN=Vm0o${nPh>f$*QzZw0mne0Lcx0&krXXh}!<w}{H`5wo|EMV4 z<sb^wQQNXT&A4#%6g;FILAFLMzzu$VO3!nC`o8E9k<0NBCX1rJ5(TPZM`D>jYDf24 ztj?Du^{s{2ex6=yLN||tzS$I{qO%-LggHn?7O{6gI|lrKF)&Fsn%e%oFf1?)7vOL2 zLer@b-)pHklK)`b8@ok~^mB7aQOq8E6zD;v+=mm})p0iHWV`tqAB^Cqh^|3~4q!TA zOruVK*gH$~i-GlHKJXrh6=oyDspzBKegk{ZNe@)3J*%T%dv-<K9QmJ8J?^DuOTO~) zEN)2+EXleZw*NsVgQ-Ew*e3aV)Y&F=_o$l?!Zr+@W!a!??VOU4AXld-R0j=F&tSU@ zrZ+(NSBsb164Vr<+pQRX2OMwy8!x15=i$=TZ%`v>yR~ez#~ma?q;}P>0LlahHP#HX zttYhiO8rr!P*6#8%0ckGk<fYd6Zb%k0vLUmuiVa}qCni@)`~G`I!mJ+n|&)Lz_N}H z3lsgOVbzGIs1(W~$)QJCp){q}enuKm`-+0Bql;uiq!*@lIK7Aw3Uh7&U(BF)?aB~Q zCPd^)51@l;j0Ns(lq6slrYfu8xj5(|60CK{lyK7vF+XEgPzAY_Yi<tY0QyFxto_vh zA{yq!Ym30}W|VqRQoAkzz&Zv_^=-d@1QU^Uc*RT<6Ee@mj1x*xIS&`Le}xoy5Is?Y z7;Af7&KY`8g0;<bDMnkki<=X^oZJA$<!r}!&TUn^y#=>_y2r1(fr86mv-S<*>L<tt zPX^X>Lc5V{o%Gj0U&<l;xLZnKh(KQs5Q~OUXSjj~`ITW<!y57g<&PfH3vF)K`fg*& za)i%n-zLe5L3Z<aU?WO1AZ&j<G_w5^=UVz^kY;ji_e!<YZ!*?D$_`5R?O-57FxeeV z<`0Vrz}ThP9n2OS(VmDh0-!X<1D#bhgPgNs()t`_7q|t(4RABHo3p?`K1L|*WLho< z5UC6Kh+BB3sU}pFzk05fzVG$NLc6~ZT>2~-nk)(g(N4AxlZO(d(=61yv@iPn$VvN8 zxO?fgFQya9bFl;9eLY1Q()YX(`k8=0zb93{KCSQf5L|ku2TBS|VUMPPXw=c}flkTk ztTqu<sheRIXvf)3pB{a+?-%1~#G?4wPUcnd+1BX56F5xm)9+6CklcV|%xCqercy|I zI0ubnrQR5IE^1#5|7Wvx@yVJHs_VkW!)eX3FKZf0vb+I-_F-!dBmE!E7Zr}9Te<39 zDf{^1LaRPZ^pTazw1t-+(>{L_+)4k<g<j>`7|Lng7RIUnD`cPUQ7v_GLAbrBrHH4z zt+*W1d#rfZ(tcFB6W18`UH|D{wZ97p5BM$s|A@dp{}~EjQt1#jDS!0j!v`1Mveb3G zv1HVeMHa0GDvv$TU1PrR$Qp|!BPTC0d1iLtt<U+PAr})qf8TYZbp9WiyxC2WKdZgn zx9?qn?t005>wo9j{+=K)5^9uh5{t(1kryl<<(Gclso_nFiL24wiM-k4e@V)g@TUtO z>7Oqg&le_*6^rbYADD?m!U~>&CtmNm6|;06yqep!;+VgC!`5YDku6W4zsUI8lym$C zW<v9%qdfKUDI%RMR_7ANiFDp}=WQ|*YbQM|A14ylOH$@8I?dM{^?sMfg{Y~8`xo)H z?5djhL`mrf&2(3>`0R6m@@CWY1){~{a|L1IH_Drp-U?ebE<Y`JE9n^DG~bV>Q9d&v z$xy^o`eA`sRIDu$+9|!&6$+C=Rkw(*Xg^2_y*po|^ZVQ5)9c6a)73VGsXQwBV-<e= zDWRQ;aua3>y&T{4kHn?PJK1ow-IS1h*ua?8ql%R^{q|iSn;;bO_I<L&!bZZRdS17> zx_WuBV0P9gFY)3H`}T@W)(%E}kjxY5OqY)izcWUpRvsY)JG}{m({PT)y#3|vE$%gE zCUzuY*{@dXMkvv;HVJ{37xR>xW=g~%K4MX+27RfhW8?!#6NbBYHIR7jnipD<cs?1~ zE;3CQG<B5s<4knKi{n(fVzAw)mW_-0r#78kt1MpZKlTT32*xFKRtOLMsV)(}G!#C# zqQfNA&)*^r0$2W}&PU1o9UU_@xBtj{K=9{%R|+(-tfE{;NvVq`UL2y;CB7r-5|qaz z4~=p@z@35~Pq0=rbF(@hQa4pqZgP^$SHh2*Ni;1ssy#?j^EXLXJHEotypF%uZ=IUM zIjsl$E@9I-iTLweQFdAU9(}*`PqS`DJS{o3XD&{^X6>R4D*AVny3{pm)*gaVc<_{N zY!Z`h?wKSK{N${*rQBl^tr??#D@*%<;F0#dnLqR2N<IZO)nHLgwy36G`kpZjiLFk{ zKTqctoj=Z6qSExvdZGUuC6O>qpxktr+vrDZ^p)m=B;PyG0Vse&HZ2-{{|Q@kQPT+s zcPl&6N%;1KH{)`LBAOm@N7{v^eA9&<cex{-KUVl)EPvmpZ`p1?%@M0=Izzd9E3j~M z?1+JEA#d;BAMd?Q`+Xj4smx`AmU?-p((JHpNdZc#$pJzS^Rw>rH4gF~SX`cTR*TNS zTmNe@bfB2OZ^H;-(mPJKh9Rz2_>@mR{}XiIxnXWoprs_zIUS`csP|KcxcZac36#2& zn7D?Z1$)bPCchhHRSqj1IW2h`iEHD53ChZ1k%A|DaEPbWHHzDCnM5;h4&R_F5xV6P zSem@mVFIsP^p#z}3wM1b9I#o_R*4wXI}dv=xUFREwqe=rj6IfJ6?7DvCU-3*?OL;m z>47s}IV>B(uoR_9@Smu{r+{<U^cyMd=S3}9nhyj6%WCe<g)O{Kw!;oEmJT}m4!hmm z^IgkhVd(FiPq7oI7Km@FEcj}*8GcGnBEG7`R5ny)i?|ED7;bApWeeVFZ~qZaWBOj6 z_v^<v+E8<asu1I&XAe$GDpS!M6}O`EanGyepN%<HxILOw!RL#oiojv5<_<tjtU4xC zu%+C2GwD}GfdP!GV2jczsGFcEGpV1i)P+ZeVq}vX8!hQ6I#;aKW%Wf9`1JWYo0IN; zkAnql>+OH>{m$IF>d@OX!R#h~`u>;G@7E~dulxEH<Wh&1Mk@)EJd|cPjp?$4g8dLV zt@X+zg%;GH<?N#9@w8PN{fQwGm_h)5-%EXwFjK&s*Dw!ffai8JpU$9X@`ohyF^>w& zBC)LR#8OT5@-;kiF>Cny%k^5@SA>Q(ITxx+QtnS}DkCi@O_(tL7hR}#s93ZH24Tn< zgfk36VUm7W=^t3^fX!TWxWwaS3HuA2L~1!kmOAP=eDVDcX>R)^E~^%Q{q_2I=C%$W zSAwY4Crp?!ypN}zGXYDg(Kg*4@eMrJYdmm!PewcQRpLGyGJ__6{kejt5~K8yIk-~B z@(pgx$-$K(bqeySy_ussT}hN(aOA6zFoH`dOjp0<?UNcYA9#ol-wOQRA?8`+!l;S7 z?mO3f{XD+m+l%J7jh2Z2qr9bO=A=3N{W<e4E&lJ}iv{@Tj&(Uf5Ruyq%^L56*r2T_ z{S=RF&{(XRF;itr&P;so-)cKGPalBybAEnk-_dH-w~{*ZMPkupGNvGHYqz$;-!1ty zXD_!sDIYgWj@GOLkWG#qU+I#7bTb{!MAGCXN$D9!_Icz)ff60y)x@!R{4lZX0J_Ge z+X`rd!P@$hN{&fq1^L&GBQ5HmUV9w=B)#Z&VaacB(UFGOqq1_XEt5-YJLKY|^}ubK zwGPauqGmA<#-tJG?<m%G>l6h;E)hzzBlV7;CDxH5vtSJ9keZRgJ06sdrj4GNwJ>ys zk|;eK#<@Z)D&cOFcHAhnld;<n+U+8H)zk%8_uzelx!V2^Pq;ZcEV)JT;HBJ<m^t!l zf$WgfFv#P}3Fy(n{r+V1ZczfI8<t{GJ9i3wWTETCiwAk6Mla^AO~$D{s$84gj8&c7 z=JAAXYHYhs%KDSqJ!EN8(;z;l^y}$f?b-B2Q2TlNUVbG?cN)nKS6g+<oqwN7Qg*46 zv<^%hc%ut#Ixu_S%`Dyne(Rf@d%S&MZ<!S*NJm<`&CBSSRa@ca>Zi*y`0*G}VNlV* zN$kSjHA$4djjz<|#T~ANSoNEYP=ahkDFUkxcaA)wb${$HcpG+EfQ9vIVSKxQrz=kA zvA-i~waM`~>a=2>Hs`I4b)cxNADOqV9f2ELm1*i3^#_vg4bu;o9zYbZr<r&B<$fGO ze)NIK{eG4(rlW4=0;Sez6kgO?N?dABEb*Nbb9mQONNwfXCG)gk`?YI0^~%Rq7cy_s zvm0AA=NYsbk~dlN*VczV>XP~YR;n-S_#H8eTjQiNBh+(#<&j1G3KRG<$Hpxnz4S+R zTOTQ<d!^NZ+o?VLsX!e7?Y+uSvj`Y!$r$W55xYIG!+M#WEAOp~x@*C%ozoFX?c%o% zNyImW<o?$S^e3J1n3uN8q;(>v2|WTbE<H{YPT<O3Cl>i|BEz%^$@lOlvFsFnNS*vl z1cN<z>esj9FRS5b!)nJ$#Df6%7ANeQ?3cdFAZSL=jMk}N)*VgjuHn`NQ7;U!keo$* zz!E)>R33yXIFc$1EL$69iF<H<+~j35|G)*C9x_R8By_Xmlj`CgSz?!d)y&8BrV1xA zxe}eO6?eMcxL{40$Qo3!S$_KA5W2K)_t+hO`4dED<ydvZyKqdppw>92QAw#af{E-5 zL>B5+HbwrQU3#3zZW_{}o!p|%Jf&7!PGqhSSt3N{K_WY$?e}m6iLBOjZP@7JN@3zb zo7Px-mI@thB+uIIX4)>R){V9Mxtp~WOwDm_I@}QqtEs;o#ctKeG3oq2r(w|_V7Q_! z`!0B!;kMrLXzk58{~u9r0uSZ-{*TWyGtAi67}=)C5@l<%41+=?Yp0wRYql0jBKuHC zO;MC7VI<r4RCLm6EwZJIOdH{oB}*}xERALSuV<Y9>-TbAFDJ(J-1oKL*K^(P>jpTQ z=%i`T@I361p^q7SVUp(3^=Z<W(fh9dXeyyJI|NyqLjnQ+zHr41)FB0^Mp_&Uyun#} zdc3`+ColufzeBt92}87L*)f0c9N3nY`)g`M!FxGG=S^FT<^M+%U_|sYJR793;AQv{ zLpgUxc9v(VX;AE6WKR>vKHZIR_;fXX$AR51Qb2JMb_QnfQ>S;l9|KdWrOuPa>UJ(% zwY2HJJneP@1P%9&5J-d@|X_Jf1T-7`D?l6`O1IFeQ_r@oG#dsb@7XQ2E-oi?BT zXZJu+<3DJEqm}2hC^>~62yR@~QgdF^naYGbl^_$T!B;$WaA4O!Q{(-#%>y|Xe#rha zy=<#;XIW}~xp<kl6gak=$hVyiElSwBbZm4viC;8bUc7vGKAW{{wsd+lA(6j-Y<cU_ z;8;QyUwe$Rdzn06^}auce`7kbz{)wnWPZ>5)%dIN=lGZ<Y|Z|$k<q2geQ&1uHXM=Y zx}r$+8s)KGCyr?Jr!MpCKL2T_k~!6AZI=1vzN^z$3L~X!G)5sZC-&`%j_K@4>3aux z3g-mmKqnl^Z9;GL8Aoe`FYgHz>8<NpQbZHlJ#!&1q*yh{lN>a;A3;w<LU0U=Po9(s zx7r<<IO!8UcQ*#pn>#vb9jW)ioYhO4{2pF1s`X;s3%PK@r~)G{Ii%^QdwXr;PcKGX z(wBatg}{Jv|JkkWeA<_M<AoCe21jN$1#|?=o(*vJnU!l-d-HScqg0bB#YY6=D!E57 zhE=N`nHyHEd34aQ$2TJHi}54(_KkiK=ieP`mwX?U^6sB-wL8DJnniX+(cWzckGCIL zZ)R0;S|(R-H1^WFxv|?eBaUXx?@iO*tv73C{mnd0%B75{Py_AFI7N?yawo>xJHKST z)6BKZ^)WkR7G?Is<hN)p{igQKsH-=vOnz@Osq^{n8So(R`<8%>e%}uSZ1fJ;VA6l& zyG1~4d%IG*%G>s|c9mD{B@7<MMBV>;Z~M=NcF};o*X;*yMwy!&4(MxlynItB;D(8N z!IFAFN|2TF!kO;$_k8n-m>?_1g?-(TU3}e%gF#+43+uZ_ShW{s_Xj6jnjHzw-?iY- z-P~z;VOFMcuDFI8qu3quK3;i3HMlNq_C{q-<yPk5V#<yf-3i(I37u-pnN9ci2T>e; z=pKmdoEN*F6TEce<MqmryN^>o4p+V_K7-tM56a#7<8-A{@#BjhjVdj-{}`;Cy?X|8 zKR0N?`q=IRoV%NykC_~Buavznb0D+w25Y_XlEDF-qgBK7KaN(7)BTLvoG}8UHhHYy zC~j4kRp9*E;&Si#9!Ia1=`zRs*V7{m%bR0=IjT2IW4ls~Y81N&#x-(XF@`nkx;RIo z^TysXLX1^!ZZz)oHtU$4cdU68-o$oytZ8H~IWAo<GiS`pxo<U`Iaa*UZ~Q><M(?8= z+yjkgD8){`<I0Y4DP;!6XLQ`{jUHRM6HFePyDu0$rnqwgj=C4`GyD)ZzQJ)bW8C;J z@1wWN=w)YuYaM=DZCc*-JL=Y4XDt1DXi(qo-@>=*ila>5{s@ZC|DGO{x8ry0E$yyD zdw!E{dEMcix|LrXg?VcmG--W0;TGpEPwiGt=hrI%cX|76N!&Vo>qpnU&i0@OZ-#bn zGUDlE8W=`NWlEVu>19Tm@Z_`-UJfNY@)@m<ZdTJ;-5C~{t$sHx(pnpCE{bOojH7fi zV+@~LWpa*iDUKmcLlKU4TK#QDa$D6he;ArJH6C-^=r^FG<>^0Qt(E6H;G{JfIH099 z={>MZE9A()TCE4(13iw1S{j2LZ@+DHb&O3;J9l&2wX`_KAEF$`cVNl!cthhm$K&;l z`y3sOt};kIzSq*$Im!nwUP?n&cd*W0Z#uUr@6IXh%L~C@w+~9L-v9pD)utVr9x$gi ztnM$F@>w1D{&P+f_VV>jvv(d!T+R(XzTHo6^@-KWtWp-E*l4?7)_h&CegR*9Cc!P? z!hC;e|BfXCc4#sG;P`&Gx~uc6$Ftp+6=pc@R@dkMW?R`W9U1?pRNZ#TmTlz}=R8hu z>AyN3Qx<9+7c+jEy|u`4XN_}>LydJDEjn>(e`z4ys`bw$`tr@GW~I7wQ&vXt8iP+w z;)Mo3nZ&Od%r=Tw7#uN)Hy->)$tq>az$l(HXm1iv9!xc^B@8Z@@SU3a8~y~)Lz~YY zS-el*O5zyTS`Kb*!VSJDt8tjHV)J)RL>*1IGpqC1%e^YK`!PGD=*f-l&FotH3AdyD z%!rMT4cyFgyIt8%%m~BB8f6o^C&)*u+<J1ld&(wkPWT-?=$=#bMCP&4(Gy31xagV8 znmsO}Q<^^?qf=geenekNe3(jFSmzgM6k+7I)nrz+>7!x9YQI2}2o*mspPzDc)i<A0 zl@7lCyr8ro?PqQ>YuWU~_c)<x%aP+^O$X?2O}4aoWIkL+UscA>d}#5Z&Nc4lxL2vA z^~}?~oWf|04?Sg`)-%!nr|4EgAOCxAM@h8thYx$x3+C7TU#ctR);O|Ps$kyp|J0c; zC?6u-d+Fm5d*wdtDRpw0aNXddT5`+vaCY0Jy?oo@Q|p(RVFuUjOTXF;OYMD76sC9m z@!kgoVQSZxi?|Z&N8EU_*UQ~aZ?&cEl`P=uu2*MrUDtDp|CPEvQTlV2zufvBw=vf* zjJ9ppC9j_;+pD{r+9o|bqFd+0GivK_<!cNt7cT0x)s-xowAI~Rv~9C0T0GpA?ZP)3 zKCP4Qz&n<eekT??d`?$=XOveP!L{!?w=e5yQS~+ML7h<Bs2g0ntY)33y5+Xd+ed$7 zWgE|{6qXy!6AL?x=LvS|zSG(CzSpB4=zZ;@uj!txqmiV;4Wq&IytkvL=<&&&8l)xt z!b^0uccTkROVWkrCi9kdPkqA)c3Zu}PurD{_Rw#Ta=pW8omh=>!&^pNO$`I%Fe{A& z|DmV!ugxu&>0jTq4AT!aw+PYi`48pLZ#TDOE190VTBu}t^{Oq2C$F(!5@x7TW%94O z2FLrzR=X`eM-JLW(~q~eY|~J+`{;Lsezm?OY?pyc^ttzbS$T!w21WKdy>??#S@Cvb z*;yNl!Z#Oj++Li2uaw14eV>+9Yd7YeW$8Rtm}R*$Sbdjzaiv~SjqZuv!Rv}H>z>$C zW;?bgt4Zgv?uL@R1%HH#{N?it|Ezv(uVcA;O6{VQd-e5!UipMwQwkTeU8*k*n93X6 z`Lp?AfALxD>n<I&om1N`5?n2k2KwX`btY|2IbS^J{xxkNM<-<G!yTvc26SF)y<R7O zNqqUz0H&eCaB+Qul<}f`L!@!6Rl`=JSVDtUz#rihgCl>8Q#ySA#HKh0oZZ}z@BinG zxca+afh5a}Z?;PD>J92g&Jt201J9mL>3@0B_pDrsHfhr6koVc!-&7iYxSmb(k>~IF z79~38dY1lbLZ0vN%}liJPQ{H^o8{jX{;HF&wf=TU)a&j^v*bdZ+thE~qV{f|vtL=u zM-=|@f+LNjG`zS$56lc)3vD<QWBlD1Ju!YC$&m2TB}*NTm-!;Si>6|j9T<X<W+>w! z@icsjH*7g{<$3)f?+{(ckV`67B~R%7eD&3s+VoW#OE!d3tlvh%rz$O?n>q@kgu8;Q z92S&B(^<7?v&og5yGwp;{l&3j!&Rjgj%U2qM|OR=@-8#HukB}R`=N&RqIUZGGNa#G zq8&ySZ`yyi|9sniwf*O-cI)=zZ`+@TfAs&3H)8}F)rfQr0xm@Rn4^49EkFj|>M(jN z`N5<#&tYaqfOeVZ?lT)61scbQKT;3eVpOFx@rlu!*Y;s=XrWv2;_B<=Mvt=E3`+Hj zeu;=Gx(B;|$!W9O`=lUD^ST@(n%=gnsoa13tz%6tcWRSVF;DlF_WSD8so$mP?#+Lm zZ1N0<h51J9(*;>`zC(Dsjw92Wc2*6e(R7QBmJ{^bZ$?+q`L{bcCR`c2S)b|1tjK`M zcP&@x)t6?~Ha%fJ)OcL(W}ezT#pb7XpQif`dD?CCA2PCwZ)zEH{MouazbSFk#==>% z!N2U5ck)hTl_<sC;*txcT>3J&#K*fyRu`w_l=>M{l0~VyCw2tmc6HIkUFWwJ1?q-4 zO&z}|<@WX5w;tVz9l?q3--(y|LnqP}-@ALph`S=)zb1bBL9#q|^#rNRaMAd+$Wdjt zo=e?pA3M9`6hB$@IG6M$sJ%1xyzdp?JK?K~4(cq?-k-{PaA#U|HK*{DrJT8Or}v^C zy}#*Cp}3Xx!}?!w^b;MQ=ahQr4^5O}4E)mOv&#}P=Xvb(fS;O8l72r)O_JUNj#_n% zjoOUGO*P2CQ-}P!bK5TaF+L}yNtv9okb7$K{Culu<|9YXcSCO-`&x(kZ^oVrNd4Zl zDYEdn`Xei!*-!2Kv@gv!;}SQ%_|bX4Yoe<+Z*nqxt|UgMw}VAVo0JIWlq|^i?jL*G z)pM=y%CzL&xefC<to}5<=d|LOy7ltuY3HKl6Z2NgWz15^*z%=kSsq%Gj<Hr##DS3u zCDNg5`gX)-{8=5W)GIbtlipV}{h%mPxh5v<@PNr(g<!q+@9s)<)n)SEPM6=QQ(K}Y z-0r)?E-&G0G{rZ6#?hCrPK^XVuzj@akfvwjXtPqBL1E7OuQ%T7yf)X(v*E}`H@`Q( zJstl(<*xC}Olf3U-Szp1a1L|-k!g(G)4<A(p>tU`T&mLs9%eZg@@$Laclgm-KV&7? zEKs`xJ1uX_?hjITGPnOB(%n$$d55nlXWI11Eg^Y+jpHFIPp*>h7<IL2f}a2Kk94MT z*R#BCPT5jY_XqZ}<I-08Djhy!z)Dxm=|n~=$7qXn%<1mT%Kn`(SIVs3kG+4MKOW^+ z{)WBT@o*bkkugtn_X-@Z3*HobKVZOLD~~y&aiFeX#{YnQmrm!4yfJe|ZQ|G)MlJ2* zoQ~z5DVvLg_h*xu;vD1YA15nk1LysV^Y30S{L6bj{y+{ZnsDFpzVKM~-HuF8jl+AQ z8PhSh>ej}kPXBews&G+1d50*c<dIdE`jL6naLJ<I#<_EbHRkRHE)>SN8J%PIO+vKU zWKltMzsYlDch&c=?*`s{>rT1&?x>lU?||h|4fb3KUu}ItiP`RN|A-p|o)t>=dT)Ab zob(~)W^e90RmVyDk=QmXCfDr4Vcpvezeo+7;$MV-OW`5Rc?Gj<*WMfBgYEGd4-<@M zH~DQfT-sBd9`L72Jc0hrz<3ckkZO2}BA0FSSkHaIaI@Xli<8BT|BP-Pn`6|_#?(xn z8voWYnR79{P{wgHO>KKkQyG)^{?R{#LKSYmFGU>sjwJr~TI2aza+U9o4G*G6KR>FI zB(`7182@beszXji^sSXL?7c&>o{M6zr{S#vk>F!`1Iwju!@qgBkxXd+z)&XO`6xmQ zyHZ6N6WiOd5bt%m_GYNr|Lv5^{8w|&-0$iGC5|!I6uS4?e`B6**=Bo#o^j%yf5(mC zr|0i(Zw$?8`J655I(y=0Ohi2YN1HiL^!cqrI8vs$$)n3RPYi{?!TzPt(0Zy=BhMod zFr~<x?UM+R`+#QiK#L<x%B(Vp(R^foEwXtnM)bj2m=j9Qqpw0|`g}cCtxayJ!6YBr z(<f%AHFxCSYpuJ&LiMps4USFUxkF+)v<);Ht>>aFocSRXeOl`0FYNKxyQqXC8U$*_ zoUnY@=H~|9j{Cyw;mZpCE@L5M0Vn#%-uR1z3*(EO8r--8xGR!Rq6J1SOcg6V^peIr zMk8}@qi~7;xta^(dB!23(L~kB<P9kG3iT)6rc507B$r6du*bqaEkq5qq%@?~Nv%t7 zg5^J#iW2ueFT|~7Wc-k(k<*2yhYE(4T2xzYG0D{iQ5dF1x^2cBL0+DFmYRBhNIwee z-S~6owsZ@bM2viZvkfLG1QSl=CK9RXRB}ruBE-PTbNBA0N~!#t{I5z6&W!A=?M%S) z$kc+WBnzCOmb2g6(|`Xy#C6n1DHWjS?6e{AC<P{_3#&79&J`IR%a58WvWoBcKWIU= zAQ;*(8$v!SMIzHg?wTQ6M0%(JLou8u%)(4>+@R5H(-oRzOKK*xt@9kHYerg6BYxcD zTzRg181f0n%s4%?mOMHsLF9JQ)Do%Ac?2TY2*z84P%&oSZv(gwX);eT(q9>eiHsmg zUx_-z7fTolHj%6@rBh!?kmrh&(*9je&7bbyS{wX9ghtNCkIIZzC`S-rHLXKJ7$R4f zrbd#~vP_X=;TqyDr29Y*YV3MIWo{`!M4zX{nt3}~UbV=#$d45+#L{fuT~$Pq;8I~m z2X-28@ng4;TShQuc-|rCmHT*HYb(wS|2#2PP4JI=DN^ThW~`RhkQhkDiwJ8&tZPXu zXaeH;Qjuam>Kd;k(b2@48~}=85(47S7oretwRGq&Z`qF#HFd6)HPrIj<!R$-Q!9E2 zqcfxurHjh?10w_AZ!*ded@llo38?DuqJiQYOmZx4J(I9cEm4S$P%5}*^}V5Am8Zo8 zWS(V$OG^}Y1I>tM$2Eqp`(Pz&Mp~XCb{^#ZEB-^W(Cr-sujwW6c%G9$8zK==!(5Cg zF9g)#UbRBH9VXg)So>EH+Wl0US}iI)^Z_`M$q}Nr>b2^%=(QY!R~En`4vt<ytyPx` z69J=x+cfuaD6baDpdzbiHu1z&EZiHbi<w5k)Ad~Clg2Q&k-VI}NF=wUV`0<r;)YsQ zj-K|O_VFr$P4h(RiDE+v`Wo3M)yOSB5bJnp0T112SjYWXQC={*yw7qaYcQ0GJKr`z zpHxQv8Y?j(_cXe?d;-1o{*VNGDXy%}3lFP{7g<@|8=6fd5i20~4^1*rQpklirq9gq zvK5-H|1<enU6n~_it$Tm2Om=B6pxc00`m`9TPLD3w1-&FsJ>Z!wfgF!LX;jz%f3&T z_`()r9ZwLDtD=$X2&2+Nb22nZk4HBb5vFa1a*H}vptkljnJ78l3C9+7xc$P%ONY-P z@?pv=>j;ExMPlT+ilsC)58e(7bPfCB)jOJ15!H;kwFzJgThT=%h><euG!vZow-;4c z7Qk1N4hQz|WdsXwrjbJlVA>r-0NnrL$QNR*X^|N6LR2dPOn6)CFI1;3;<SePWk53! zSU8frsJOC#eslrg1*~8JjIV@1@6AktR`a0vPjHYISJRjZ)HPdO&FdM922mPO8jrO@ zg+Px=HK;6A07;EN&D13u0F-gW7Y&H@VM$aw&j^uA-KnWeK)yit53S&m*10Z$g0;*@ zqV~+3gj>FwbIQ!Ic0+sEMq2D#=wW++m4BfjBcRA3>ePn^QVSl~fn>POzJNMIsaVwW zXv~AAOXw0;0w`wS%AHAJl6&O3P|eW0eaY3UFv+(ICnbq5G@yG@k7*ArzE#i`9SJMx zJ7*>)cwh_`HQ8vGy$($36*F8LdS^0GWX2PkDFK$O*A+U#1-^MRxW-AW4m$-bb|Qh^ ztO44#6#DBdI!vTy>lrWjI8Ht+?}e4OqiX5(17!qwdjn5bBZ9Awmev}Qg+9_VG^RPs zP_ji`6tj;2x&|SMA`kX3n7y_WiS~}wl{JC=q`HGUoI%}TnLX0PZYcdx_M^<+Z-pQZ zCFn)Y2L7mt5nuRWtSOSHz{i(hg?&mcrP@qvwD&@@j|c)QHFzZ(0PWk(cPuf|YyxUu z6eYvrJO@f%w8e}(8lX<I-4P9a7w3lahaeh2SMcZ+A=F-@g~48XG0D*is7HyF3PkN| zVE;BTJ)o!VNKuw%pN1WlSmF(f-??_&&D#lh&4^2j>39PM=#3byL6>N$BF#1t?AKBj z_B%<lIZOo!w!QV(P8dzRZt!|-F>DyR%YiSy1>}Sci#NJ;6y@#irR~7Thh3TX0P|<B z6KoiOta?ug^Z|A7)4;trEUfOvvT_L%H(h^Gtptsh<C&$-&{>N``R9z#%K-n_nUeoP zMBq(k%4jkQrHa_+=Yhq+us}5d^Gq<nJXM#{j_t6zvuX<{xjz;^&!OG<9mSReY4!-p zyk2pVd;X4^3(+z&Q3C%7s6XIN;t<p$yP!TE?nKV}qVBU)5hrTcE9X*bV2wHxzaWU` zbgQ%!K=vGe11<0P8E}U0+`U9d?q0iU&PR*P{Ix<c8*AqQFg*eXH4ZmrA}vrr9#kYS zxC|+dK}2UHz%)PeypO*oLL#Dm<vx0^WnOC>?mjK4qt@Av9c2~TW1^==W!T{0iQ_#; zqMC~|Mu#|V4iYt;$pvS}tQB*S4%deE^AM^vOpbLiD5)i|)=0De!3bGLt_36>A$j~P zP989S=MXb^dVKwbXrY;1(9aB*y;#wQSLcy>#=PCQKl@E-Xi>a0s_&iz;)_g};bkbW zwq*c<)&TQf(t&AI$F{gz34=!z9&cVTd>C#IG}btl9t1`4lp?+$V58wa)sj3YUJa5w zk2?0BzrzO3jDqHM9J}zz2kuGFn6NvbVc!(prj-=Q{i$0hvyZxjJ%yk?cnOxBF9^-( zEV4T6s5HBLEe3UYGj`Ne@>Z8nY<T^xx9oDnuF-Bh+pwSy?X3R+b_osy5az)y<K;k4 zEonA2tH3UgaqJHzL27VJ-QiFY+`=pEFOCOY;6NspHcz2ivvL8dNeZJ1Fv(J7D?HJt zXFw<QL9|KTdC5qyC=u`N$M5px<;Kzbm`{<U$-QXfIqDc^6v5f*F<lap!L2;bNy{z( zz3bk6A=?Ml<IvJNU|=Nb*Z3g5@De0x-X47@Oy`=Q$-qPNOZU4kJJu+;T)51x!!)tI z+!J;#t(%cz*9k1UPO!~>qm^v}lzt3eAeRJ)?5ddL`2;k^R4GpRhQoJrX@e$ocG>cK zOp*fn%vAz;JG1{O`{YtCB=4(npFl5LF5u_PhaRw3M!<Dc8lgKokt#9+H->}WfUXLY zgXB={4|;?8OEH~FU9itE`++?m!;C7~&{ufq6OHNC2kZqfZn+Z#bUOge>V$5md;R~@ zttTeljt6%hv_5Fvb2Y*qbp(Kz3=g~CbUqqxp;p3l+(-qWB8(cw6Jd}UpJr1fhWe_* ziOv8snD`gxwRVLHJWi?`K3t{D%*Y9fB*^1?RmedxSUa{}(A-_mJvAxO8<^6{!7)fv zWp8j*C-jM*i!cy8Z!39J4Pa(0QHR5lm97*Wm99*|p!!pZpuw~<28Tv7n4ATA{6s;n zJ03u6W(oLSQ^8+I;^n=g$xPPA0`SLS>#o>B`o#5ANriz9-D@gg2ICVo$jG`2Xhsv> z-PV=of>M119$r-vJbVOR5Hczua+wR#o$gN1luwB0L@mgS#JThs7;5O5`j`RT)>Y~g zFzRelN@vJo;%q%5LVtqPZB!{I$agxO(0s>Sq-7mQND71Rg%rhH3H5p1x<FA1k?#&h z1HD@6BZ#z&yu?QjVYYntxgfwXkI-Uf{H}b2@#`L|bPq0h|9yp`nmD?^Eh`H=jj^6{ z0qh=8RGk6o2NsDFo%n_-*1l7QSV;9|U%y15DJj<?HUex4_`EI!#IIW;=yzpQcj12& zB~XfLD-@eCNs}mIkb+>5vfvCFFv+vqz!@BWYBdX9&AaN%XLZ1Rto$A93>~#;D9JoL zsY}hd$Dzj|##j)Ax`n{c2RXkKM71X((&18&^W&2u)&M;XpcfAG+P@;H5Y-GHiKe1? ze}v4$O?a*tDwRjDL=%{dM!7tzB8ZZeBdA+G(IlWZYf&##S<;yY4~$uVBJW9J#QKRf z&^`{x)Z5K3K2e^)bc=T9mM#~+bpd<S25*q>8p@167p*jt$=+RBUgWtwP6n>Tbd`<3 zOixPYJaK%AO>q6>&auDGK37)d^6A<#jQU(O<vDZhxFwLh?r-_g;(Tp?hC10mqu&}1 z1UKMFA`>Ob+Xf-Q9@HYa6!d7vkS=^-9}wypUu<tuZhVnzRpRP*UqX^2jGp6(-RnR+ z=>+-i-L9~`H!bMbQ(qa3Ux1YVtVe*eW?;*sAOrZlUduWv%TokGECNO)zsj(<=2rEU z>MM)T>u4A->k^ReKf66F#o`h`1{KuYI0X1TXJ|%IsuMFq{+0W$masgrpL1j7d5+u< zLp`l?nFm`8uJ&s!Q&rvPm}=W=4B*m{1>8VdGb6;_?+#NfWoJs*bzM&i?-$=MQnfZu zi`X(@(d{0TgQ!jjRlW%d#mF;jWLQ3W6MLS?8?MzJ!UcuekuelgdnX1(1c7CeU>U2> zWnO-f_aZNP6R1FtVppHQl1#}EUTSdQAs*6fViebVJ60t@{9Zf$8&$q4#%u^A7?pYm z+<XL7Az%<rzWW2p%^Jal0H=C5W>5HbZ$|Zrg98wX$~jyU{OXgN0if?3c8<|z79Mu0 zx;AS`wFj%7wWcE7XK0cKgVkRZ(PZE+4+g2*F$v%d)-qX^Nus1K%UzW0=#6J^=qk%m ze4tm@i9s6NXqgqE7Ub>j!J=c_R41bGfQgQ+bVkgM7xmJ1BN_nL9a<2#8bUd3td(Z7 zAP#$3&<8!@vG+jk0FszBMfcKyL#MYXLZr8xxC9OKYOK!M^UOFU7r?BBDnMoVfLr~A z*#oe!t!r6uWSy`PgoVm-2=mc+=GPZ|R>>mQC?O79Mn6XWwr!OV%SYJ#Co(Eic`&FS zOJ#)!$#<VPs|SzQ+TyD*(F0F^DFof+$q~wZ5~Qx$OOXS)l^pPALi;;I;M$Luz$3;J zM-jz44G(CNzczNK19r9ulND09ByU)}M4Q#<P7mq<iClqEZF9j^IQaoJAV~S&Cy$f~ zgz~>gj}%q70TaJ-0l=T@gyHd({UDbga15Ij>wa6A6I@rcBngInx|@So&pA+8e4@Z5 zIn$WEk~;0j3zw8!f}WbJ)L|C%06-fnU!6G{NquPt;6AD>ei|HW7nCQwXp%t*1xw0? zcj9tq2{1+t1P*s9CJ#|l;noxCL)KuDx8<N!nu#aPz%Ig*9^^+Sgo05eJDN8MEk^mk z@u<L0M?Bd_=>dVdX3mF!;2%Nx#-xV?quIPYtRKJ;IoT&kq(L4brd)@Hdnm%B{4J1z zcqAt4{=o?T$d9MBKC8&22)z95h6*`m15vNSglTX1v$4Vill1KhP3>1;ySEaPu!u8Z zz_HTfZG|40G>O*{uQL3qWEkM_x_wK)6Qg_3=PtU@8!r6gWEd~kQwic##(<%aLct{a zD5g|lo5UjvdY!6rzX_^le~FtzHO!8IjsX>~G64YTlo+gN!$xrkkK%LQXgMq;2E?k` z5F#EK_*vOywEr!vjnGsw-~pUyl6v)U6j3$Qv=zeJ;^b2flZ5LTKI$o|qzLb4ThSD) z;Xq0p0?dXu3TR6Wd^B|WSUHj;LHoZ^sXhV9nIq_CnuoxDn<NBsRlK*6fJe5Nm!rF= zKGR!LW@t$qV-?X++Bd_dPguLpul^j>_uHld^`CEW(iKlNf~v_Yg{FgsK5$)&m{N#X zKP8Szi_h)AKyvx=q&}0?arGZ?`tf=n0l7bMYS<%H3I1*_aGRR6xHeSOZhhNc&-k`2 zWx?oc`Zvu(QZ$>%hF_R^;z5$y1%Gk(+XP}tCX)O`LupwON_tHv;R;nsmo#{X&%<tx zBqu;H`DP9O@~Bn5VWdL;jd^)e3qEniaTY_bsORDmeI(KP-*76f;U$skbJnoo3w)Zi zOpr?{1TgZmL@4nTXZyY~OvT_E(4$nvg7AN2zd-^gPbR5}fMTCt)+Rb(HsAj)Ok3uU zm~y(lnn&L1G2Y`YS9oiCvFV6~)`q)xl4abUoZrjoInl}5^}aI5Rc5ouj*rLv)Bnwu z9(Z$2Gp6ar`SaYknWgwP8QS(2S=RxUW>N>eppvD5DE+cNPK%<6KU=NrjZP8QY3rmh zRd5dXsx!m%nI(=or?X7k^o$T{!0_RFi|h0foJ}P&98~o}0p-|1Kx047IK!#I?YL+Q z)l6>pL&*Gw@i3j9c^mJL^;TCK7v1)cPP6-!lG{llfJq`TyqU&qheYA_MaZ!LFrgd= z-xA!N&~F65eV`i_S1Y82nE?dJW5Gc>%R~U&_ztVdJ-V+Cj`?sFHWf@p96m%9elO&4 z@jjvi-F>=haMTa890m}ULvXqCtu9PPX;ao2aeY^K=={o~_t0dTvbJitqS6a9@;+}A zf@^b^fnjH>q{Ev5NkLZtR$xY*A1GhY%g+77ItRa#$T&$H4x^^(>emayrM*K=Q3Qaq zaH0wyqSJwK9;tL8fJWHw)(7Y!RY+Qx(C~Vg+Gc*G;Z)X~H<orjQ#TuD{J&jC^r7k( zr8v9<coCQko&7GB#8F*N!-r?L!ZvGR9G<0wnUZl2Y?{6UF-}P=UOFosD?2RO){LPH zH?&W-{>;ldW{6N<DP=*^@~^|~wo)xLj-OgFT(7pY=4k?wG{6h20-%WKw&t6?nO`Wn zfG&r*$@K&te{!=Fb9i<IeFC2z91hY}>;eOJv2=c}CY_}W)8B=L#b|jj8O=ZIjgD6% z9)6+<2-4^O(fy5xG`rSf0nYm`D1Pp<NaJn>Pzt5C%P1=iaKl;vZQ=z&3pl#2ST_cy z=kkQ1v>>>y57-Zn20HI~yt}BX(;=)mKXcCm-$GqJ`17J<Mq)fAVh4<3;T~?g4-ApY z@OTI+N_FN*4lA^6$XbS_r*TiToj{+Wj`2r88`QCOB$LOhyRvE^&@YkzlzIh08lQ1e z9podth;<*uDE(4)^ufUj&2_C-aGnoY^Gul~i&<fWVaYZF;NYD@_Pli<`{Di6=TN$A zD~C8Dm`$J6kCo*)?9;~{Vb-ZtF>wHdwl)Ht5&3T(7cdC;7>9cvHara=b|q}<a7bG| zN-Q>q8(2^DYok#ArU8D30quPLqi=4ofI8XJG#MBIr_`0TfG$`<8t9`_iWhM#*S_66 z>0v-(U(DqGjs7)`VgFpj9za11;9oZqRJyXxp?^a#e8re}Qom`K3|y9i6Xgf^;R-1A zG9oJp?pt`^6yuczRoK_w(FdUE)*r<r-NK}GUSUhJ+|_)JQrv%(Dam0pZ`0u~&DEsY z5`tU9BZ#pyuDAkf6h~c35bIni2YSTacbMp-&bfD-ToplUx2w8t#cSS9`~Y|)+1}_~ z_DMJW9iYLz7hb^i^Bjz9p=wzUj2_Y-_GxQi?_gkHnifn*!=Dk&U>tDng#+0c5iBhK z1*wa@0rnCiLVO{IvF;T`r>xAl($4>=Se=nrjqL&C!xr3=*XXV4^nwrx;`kO?W03v$ z6I$L2`h3uMVHv=zQ}1lN8$uxYIL7H0mmP$=pS@}oCNrJvW>>Z1js8`ab(Z~k1YIe& z$&$sKMe;Bwt4epyZIZc|QdXn$)IFqN@j}^zPXFCQcF&di?;Nr*&EYDvF~Wr3NEWIz zC5h>#zF)q?HU3`CRC&dY%Tj#EOU_aR+<q16vT0k_DT=>yu~gAdyDto`=B_*XM~0>T z<`K!??AqJ`I~tG{`iu1e20@5y6t^CRdb;cft2JxeLu2l%w`f~pH?iZ`Hea$|Xgnyf zL@HJdBjdhp<&0&W`D40X#g;dC&aY2Qdlguujy)pzw^0CfzpuTm&NZHTvp9iaZ20(E z*#~`(=EYnO{;4NDNpC+r{7ItP$61_VTFPWdv;P92!az4|fgwzaJ8V-@CF6cezLJEi zs^4pf<|qvR3+0(ssDBA%u-qdF)9k}=?1(I|Tm?$e?mCA7Z#$V54B0Nb29EV`58PP* zcGIhzLCQT3W}0mEx?mI#x&A1tTucTLEx6_b1i1+!LOHXUSu(l1H}!?o3;^hneXc-F z^t{o`7e;D5R8ygp7G5F%=TXBrH11|0#)ivPpTz(KfRUg$`wE;~DZw=#;6lVOB9xhA zT3$4HCW<}c>V<mh@1+L9urM{cF8&H246K6^r~sR0-i?a2P|7GZuqU+Cw<rA42iyj+ zXbUkhaAjWB6me;G11t={SDYzX^Ttqq^K-;x7~1Dz>`3nEjhj&LbQSWbCqVpmqPHGP zn=!KgOQmg(TW1F7gi>Eo==24lXNqFrQa`zP0aV97D7mvh?n<=&jK{c$V4u-z(S2$< zYJgVfi4$LZf;!b|G!n8uVP*e^PC=6dDdg0umB93m&(pHEue?PeFr3jjWRGd*^Pu8d zl4nP`9=(3XQLsLW7}!blmSPzw8}(y|FF>y-g1J*OfZhu^D;am`!7Eb`sX`on#M=@o zMnY`3PF+I;k@mxg!%#utb3Q;otcC={<`o311K0~FCZRsZQ`DhxHTeDGVOS>=&!+Xu zHX&sQ!nK#9O#$AUpg9ew7|~GTI|E;lz7Vxc$5(;xW$PJThLml7Bh|$~a96$t%sWyZ z<=25kXSl@s`Dkb?b~o&=-#*TYT&j)%N>bn!V#+uLwW-l)WMdS@6VWyK2-Xx2uR%fq zy$%Y1B(f2w7hLw|pFW3R_MrDv$}=LmsQBCSA^@u);xJkUMLza1v@TT;zgnQNR!|pr zT+$ctsvL*!9iuVh1sAJl;4Zco8d&_xoAv;&ZsaEB7(zlYzhgWb3YEs>f|7V4mU|tw zCNUZ}^N&E1^P9ojj$dZLAroL1q{EK^Qw6h3W|V1-7gwQ6n~nI24o>K)I;E`u7lHe1 z{)4EPhT>rK_|ww<yYZRPdMmq>LU$P_*d>5GjJ|`CBSD~NR)9h@1JoFNTm%Z2w<Ps5 z1(H#APZ``rn~|j2I4|ikOZbd=jq!8<trl><1EWDNi}YCy$JLn;FmrS=oFIJgMKf9q z?7#rnV&K1zcps%BDAV{iqPLPu>u|vT8Cu@F@_2b!T`^xxz_euy$~2AdsK)YPpMmgp zL<&uH#=S5ipHL!wHKpkK=0f-CL&jc{J`cDi%aR~~Wtb;f=f)}O+g+o+LQ4d7!<gTw zvH}dni>(49=Myq645Jh*@#cj<Z~p90`Z}=j<s!ZTxpbKTl%!w`oGqJBCglcLx)_nu zn?SuP>toS_HsFXSpmB+1O#bFD&ZjjBH=~Q1P2aW}l>gD(w;)HH8S`!79wcdtgu`g? zZ-jCJ*L`8jowsPD`VlM!yb4yvqNiaSDk~SA?S}F_NEs(Pwgt4lJ>-%G>VX#UDv3Ne z_g#XLL(&sc2CHie40PHPoWt5uumR0DK6X7pJ>m%wbn?<}C_O?cD_x+h0LWluMnYHP zzD*qcCU%MhsTb=>F*-`_37GX06;_v;?LoVPc0H27sRPb2`rs!3KpTZB0Jj70T~81M zx5BkCvflftcP1`?w(^!=byY~icd5Z(55Paz7X)**gkg_3X371#_nG&Z6d9E7o%;|w zQda~PWtTRRv>`z-XB40@5GNtc?hs(T$G^mEa5x^^vxiWJ!^#tkkt@N<Mt7^he61q_ zAdDD+a&>#6ye|<b*92Hlj0QPqK_JHoF^3^F*xtb_QwTWMgLxOE&{xY7H!C53pojs- zW`SdYtl7c|`<N{PwWpv-%IqP5+N*;15`)5{V20xand<v34p#O#TJkAG`Oou?UoWB< z9i9#`GAML|=5*ybgaW0Xk*+-F7!d(ARV=7ij4o>5EKlLjS7&3p-vlpo=GpgLf3g)I z)dRbYy9YsFPwrd}dc-CKtIl|i*P)tz6O6|I%(2g|CUO4E_-1y8YrpN1&KcF2UiR;@ zY!~OEJX^|*Xweyec428pp_)x^>inVcjqDOc_CA~jW<{y5uCf(jZq-Q;Jzx>eL{%D{ zQSbJly=GI&q+FIA!~hBYT(G%o`d!PMClt8BmcrI63xpAwOZpSRl_PFejx|?j-Nxt2 zv=OuJf=ZDI2P(E(g;_Yc-CPuK08+iBAX)pH@e-1?0Zh9$iX;56az8dHLWO*_-{~N3 zG`QLDabDR><$fP<0GSU7WZWs~0Mv!e?&=l+Bn6_w(9<p_7@9RGNbiHS{XqF+Hc!Y? zU#&Tdkd5X={lSaW+$fuIykC-g5#p#*lNd&c$an?s-gFMK;MGg_7_D(suktZ1FJiSG zCvbSF(ge+6Ypz6SdE2&mO)0ib(k)xt2tk3a0YT*jKik<j_BwWQ7Vc)Ry5TF8D|b}) zm@6X=3|lCatQF|;X(TOH7{WooY;!ybOzgDuY=by#O*vQ*3d?K2>mJ6#B(kn|g}L@I zav<W$hTu>%xF4pY*e<gQfFvOZELHf7(TMG6T#x;ES%!==QT7Xk_wmfbc7Vzel#K&5 z?g6t`op9?es)Vr;07r0p)(QM~_bLa}T7i>*uY`gwMB-${JW<=asqssk4nFDcbAbXS zuBn^@0F-!ew+k3xRBFVkiS^wCG`m!Fs#=X|@fuCkzNFxI4B$d>fC~V%Za>=gZUjui z#N5XVrS=84;JXHP*c_Y>mBdW2{8?QXjdmkKVq$f7$qKavXx1^BcqCZZCrI0UVbYxx zfbu0pP{>#F=;+$+Ka{cq&-9ZPSAkMf3<fMEfanHltEkQk*a_ioiOO=YR&<+anCiYP zptcn$4&KVDb*f64HSPYZGM%-tsOD0H+QfY*L5jld$By?Ki{pr_&Ja65@OGdsXfYe! zR)+ehdIOyZ$q5BWsHQ9Z0Y!gH!HFHpo$%5aUenyDlz=*<IgGADM1r1T&AvJ=sC`82 zD)Znk7KlfvzR~V@GU5sVE4X4cKz;2hA8-d@SB!m{(shK`vvgtB4~m5Z&0@9UB|NJu zzU9?x{X>pbIPhEnTMMA__lgmrEwJf$VF&a#fKg!%j~pR(8(>6&>hR&uZqDl6!_gVN z=OdISK7l|HFo4eZH*8w~oH7&p<cDvmzjOK&1@*b{_6I?gIY?SrYH$J!e5?BV86){p zZ~}U4{c<8G0SoH2q|}9K#T<YMDR%Lp(wYyC^Tbb-)9~<4Mm3~(N#X}aBj9O4uib(< z79kncTpT+ns13&hv%hoj=wC}=jDPQE@!_sY@G@TWJEw8_KGZEk@e@`+g@6DR%K0-i znJ%vqpkjbkg3Ja92WS?+B7tD0CJY4&0P1B{$-L=Aj9I<(VYFz5x5eUZ2@W0&83+Z} ztx;)DvUBtjrk#H*PVsJ>&LM577l%NW5(RD<3Vp?B?ZdF0pgEB5A=E;hr#t~0p}rs- z^o8!D6bU5-Y?6;cMnwt$Qgl?%d{W>GO?oSC_{hONQuY|BLjuvTmdXOC-IUj8R3vR7 zXukh|5;K#&Q~?ko;Y0=WHDt;Lr5iv@kNRqLiaipYvVnBNP~Ye5Ca_tLHx*F4W_&YZ zSd%d?D>joyKcZQs_h*HSlM?&jENSo(t1onf?0H4VLl`P{55;?MN#9P<&Tke^Crpd? z80iOkysN+oDE-2t>sqx!NrOy!f?EGe7fjmN;6u@j#Ai5VXxPS0oj!iP8Jb8)q&y-q z<u8;HR4h}PT$I)~WTI7&11Y(P23;xQrGQ!JTS6r=DNle6^G%m_exU-Tr62~^A`}0} zEBGqG{wO4v+ScHS(t@$h;?Dk`bA2LHO2CT(?_+RKT=;oT3G82d0h=^=0!nnkXo{2) zh?O0412K{*k~wfO78h)PJPBRe=lpM9;6svRABl$32&Lf>X`jixfNdDi`Fs9n)xZHX zBnozUQX2mMU9(`J;1nmrsD{-7+_HvCh{z|X{o+KnG`UfQ-I|&N+>=hp#wLC9N7rcp z+%xdm#`xx)Vhf%HcD*0YGA1f1f_0{XpAtY7k(y}R`34{eR*ibk9jIF@NDB_sr0*m+ zKQ(lsq?$3&DH=*z;&Lf||4~BADJrP~O2;rS4N)Yem%oluUu{5jBU1N?H7MF##PM@c zG)a=vk!>PSHpR()G`x)+RdVv)miO&!$|tnE^;o=I2+BX(R+)gwkUR;W>F}E_l9qV6 zFD#s>((ubB<%4)fAaL6%rL97_AuQ#(_-@h@ABN9o`Yf?Q-rKT4^ma#u6qKJ+jZ|8G zky7?3Ir(q90%e@$U(xu6F&bX{7#Lf?HXNAjz1ObdfiV21@uD^v(mq`)3@WlvG;RBe zLstOdq^(zc#5Kfc8zDny&Lti*N_uB+SPkQfzhSAVIZ7KF&FU|3Poz9j+Gw+AWqlIs zz5g#Y%%6ph!wzbQ?Els{lSRlu{*yyvl#%fN6FL|<b|;AX!|iFn`t`)53kTS%zg|kB zktd)1FNOH8o=H;LMUMi}JFj2Rb{8FZ?A9~zkz*_7K$C-+h8qAdm{l#uS}RR#GC?*} zObXwoZ>A}DL~1xSsnLW<{Y_`LkVzegW{T3_b6VcieR`l#-s?XR`Z%Wl=f;!71|K?9 zyQdtfm%wQI5r@XjKv3*n71{(-S^vUoK>i+XQ~jTfO_8{yDGII8&@a}285pZQC`4B@ ztd)8LL)Ax2YSmdj%7e*h>>ydePZSN~rKRbL*wsx^Ye-fAT7FiB`BJhe{eLRd=ChWA z5l!i9Pr_8cjWC@$9p)`s|K3P7^b`+*KscxkdsujSQsE=!fcIS0NrhQFJYBbGH}dkV z64|JoWF;}GWcBi_$lyA9wx1GaLxWa|ElJX^8Y{htBa9#Xi>3gc%u~YQ-j{cVLY4Ju z8{XLwf)Xha&0jv_SeU^;5;XHQ{lKL8G(7_qEBCKXnOfcSmp)*Yd$#^DrKX;sf6`b3 zrSoc-e^TzVXh9~G6E@M!rRnA{kFO3BRB>+V(l!DMKK(3rYBv$)5>la#P4^05V3WR6 zXu>c5tim&QkfrI%;z5a*b1vK;g$+%4`|4?dFhT>Wh2S?>2XT|8awNJbhg8@0=dY%6 zzg7QlBdcH|LioY0m<{0js#P;+XU)h)@t1r=ft<axeIZg=-=vXN_>jOvn#@Uc62zti z94g#M3}vv+I6iYKhu;6%7-rX;p!eg#KR5k}gR{ZQv)bgKJ_4~x9diH-@Q*DOJMcHH z@sD4ME>mdBM!9AtN_OyqpCJP5mTRg)Uj9)>g}tcD3|x$CQa9pXm<sfOxtrCOr>qdw zCq{e~bc84$%QA-FWYhbPREuSh>N@_MrZ)WXA~tyed{iF}<yPG(Nv2=*<x}Q^<uXA? z+G9P?uXD`j(uhFZ%gBNM)u+prTT2Bc;taGr&xydc`VILC(sVdw!nQaFWpLTJ@6B0* zLG=nz8bq9iHR7pPk)-c7)C@V4(?h};v}fWOEvgOjEG(gKcQqox6p|cI+X}fbZs>3e zbUk^B0mr#`R%T4t?vU0hUnNYM$`HI(?P;k(APjORX-pmlZ3hLOlS46@<c$PPt|w{; zD&bxOmxC}HPO=gkm5q>1)VynR2%cvNywTu@jkd;Q&S`RX=J^OH!qCIx8XN*xLCXT} z05@+4*Rml>in^9#*C)lw*t)g_#@w{zAQBN7QjMmxD=PG!R>B$<R#w^^lNM{0qDv04 zBoes=s0wFrfFbAdskc+rU{hAf+oY;+tI5kUyhIvVT^QXBmZ`076OZ8`L&zIUkkS?- zl~*T8Q8Ln-)Uo&|NtC%fYg*%jOM+c2N)D)^yZ9AzA<|odiG*n>QZw|xo>~&bM_V;` zFNX-3*%jin?R_?UIWU5uR1C@&Tg3XbjDW`tNLs5Tsw|SVH=bco2G{02!^i=Htu&UQ z#aM~}!s5x$@)ixbGD8|I>5ZU*QDQt-l*#GB;5~<$F-fu81jL#V+01{GifIbl{v<c@ z1nUxM<b4)mVR#0u2|`4!7f!Z;z6d2Tz8LrFkVdO>{R%Z5bZOe?5uBTeq3|5I``}^C z*VG;P?o`|OwY)V$_PEvDPPDBXKctOY{gZ4Z1TEEEg{Fz2I%}qo=Da!?lh&KRr%sNA zU7(12y<QDF2+V+LBNxi@OP7ZNvC-TvQ*U#^acwKHH<;0}(Bitwe98GLkE+bY<5XX+ z=6o(}@=%JL5=5)Z8xiu=d|SRb)jq}`N*SWj4op(DxC1p7o6*uy*e?>4xWJe9dWk+| zBohjeTrrpgV#gIjC_E3z&+m|))@D#0%F^5Er=9)Q35v%eucZ?-$(AAxLvo{$u$DCO zttMaFPntJuMW7(FwvNOy`f6h-BA3E}mmro(8MHt0<g3mF=#!RaUt0NrY`JAaP}_@F zB6s#X#G5j;>#}f9PS;0p^+3LP{DAqMQ4X{2(qr@1WsJknm3VFV#I@JT9@UAjmDAiY z87(DOy@W<D47Qdxoqr9@bmV0X^3}OJ{UW2SJ%3l{zV7*ZQSwqG7M`k+FDdxk>y^*d z!I-jVvKcL*O!Tv%N2av(`R8AIRs1=I(c}gXZAHum&T?~%e5Mq@K9=zg)I$<*)?{(; z(?d`H{c}(u;mwVd4XPMf-X`u%+Ldj2u?kxHeM!*De}dLHX&_j@NQJ-Z!H*C94W4`@ zfXS@p9irN+ILDj+_d$#L$>V?9Tjsy+zt#~jr9EWgQJuROtq@ACOZgVS-A!`RT&8#{ za(52x&-B_IZ^F~(_O$k3BzUSst`I0pBo-8=P7D<0V|?P`CT<m!2YlnDun^mR|G+?o z%ZmeThRQN{z6=9=WqYW00*&{4Xbva)n`2xj$;z<0Ku#cL&=!5D=TwI-odm}hFT|@E zN+*MnzYX{vgF8SzOyA>y3rd{cdHmG@(2Dd-E6}~yw8mI3t~O|eUweM@V${%Xkkt*I z-S;^^K9viT#(6$O>K#fFuaBnS(Elpnrvy+?J>Db^78Ax@1%3M>4()U6KyI!10es)@ z23lU1{ekO6L&q>_8~OubJjkom)_g=$#hI_qw9Y8E#ek>sy0qOBaCoal9q!_8a!`_f z$D6F-30AKIzN!OhVeSZG-G`v0Q7QenS(nHA(dva7BhQVi2w~uA-FX3xSPiCsZSJD~ z_nww~oF@12kO8uqyNAf#4-O9kVa2({zW@@xCHopQ9MS6^(#|t^Is#XafNUbh%8F!o zROzAadm$nFtUea~o_|&!wVIxme&Blq)t7O#4veB4I;DC0xSv|QG!iZ?mSY^m#fHo= zs=OaVp?<F_pbM7=AcntPsprTEih=;ci=hc8=3*d7!(UgyICq85%j+7UL<!&b4nRPX z$VhPD%aB{Fgi!|uqts)yv8>DlD&LZ~FqEIcfsp3`*BzJ|{X6-(TH@$ki3DNdt#o8H zo;N~7v!<3hjOKv!_Yvd(c_I~x!eVzXe?G4O?At#S0}c71RA0Umw;W_!*HWjsd~MNc zs0t8P`)&DNRQuY@#RPcq0+SnB<{*C<qzPN|H(=3Zj?2NMg>U)&8?!;$xjql5HfC06 zgT}~bVesn^+24_FtHC|ucymMj?gz9+uw<jU)c1&)R-jAtJ->JgyvaNaoDhLfdv6CC z^Q~{fYlz*n?SD6;Y0-50FB3dds2CRycU3U{k!@-;<~`Jf`$I(#UP4{}SI^&BKvf^c zPlr3cWB$A>(g1DeF#aNFqwppen*l6@=KNV9&5s-b$#}iSZ?Mdtk(5*0V4Yb@I?%lw zhJhHPql{37YqOO6?Y?ZUh1bA7iT^6Gu5#|GTv#D_4x_q{;cV#R0WIS&WFvLe$9%fT z@B2{tl@!`UytRAfN$`f}XV@B)A*{p6LMoV`#7Ps6`b3(rPO7T{#u@^!kKNz+=`4b= zbGw|s*ggS|Z{p1Exu~5#(_RkMH?B98bqE_Iokhn~LY2e~BXYqXUXB`}il@B)LdjYT zWNN|#Wuq7g7yfy!G}T2-6e4X~KD^<VEpl5weDWT!S>NRfCfuXI0?s(l1*wxDJX-N; z{*(8sSUunl9G1SjgAk)M1ww2qglO}?n}I#x{rgm-2(ah4<<)J7Didk*(**!N=xy+2 zxIz-;u4WsxX}B}1u5mp^wgEd_4lL10_>bkyu!#!3qaz~}qDz=O2msj9mbt=~Cy+j< zBbseSf;j(d*X=8q%KF57Z6|v|5REI})(`vMi#{b!cvRQ9SY?mtEpGVAKJt*!3LD-* zN74<b*jvc4e3EB4FVFM#!LB-l@o)XneQaFWM?Zj|-bPc<&CBONFk<(fD<15`u4XP$ zhM|N2NieVp>gE17|GoV0WiEDwz9$jrn@X<mX=M|G48#75@B$l;rhqB}3mtw<!#BQx zjs{~?4mnP1pR55nj1@IPO<h>0#KjY4Y?nd01#>@G?w&TLX<@R%!yZClmdt8!&B<IT zhM_(LLsJQK1OLAu_Ctag1`k=(SlIyh!Oek2j*yG<x*By+LDXK{CWH#&m@d5G8J-b5 z0b#j2l^ln!YMFmf-vR^4L4W$_D|(uM4fpU_QZ(|ZT?H@n4`L*ezV8y``p`R))|!br zmu;gT8gBx1m2k6DT4rDck90>sW*bo}2qjFzNRW4TZtpOaTA}_M0V#9p&%^BsVjHw~ z=-sQ6ko$fRZ50ud-n1XP8b3@CxELLdkY+@bCd!~kmo5N2IzZUax5;mS<tJwWKs?tP zm9Qn=;h87`?H(o7>0HaUR9!`=*Xp}iVn<J|8ZKPn5|2a<dVmG>(OyxXjl!DUbw2TG z@Ke;POaA*QxQLX&QzDXZDoBw#C8CA_T?%d+hMpG$74>ymUYN+RG4yf#!?-O*Yg|-o z`#`(Ws+4uC;tCJmjJiA!LRoeAKWRUrF`o;w$U1r}oB0O{%6Kqj>(d*U)-G`S4s}mF z3#QPE$J^wUkqxYg6x&nMus5eRMd<Wci!#p(JDv4I*IdV3XRGfjh=&yD3IWlCyfXX# zNdrD4%4`t*-RvjPAbUj!KMdEe^zJPSX|=-}0qf^N(wp)j%(Q)4*<7up!`(+{_Bk5z zN&%eUi=2l7McAqX9ZTFjC?{)$Ms=UMh7^6>#oNE4%af?h$11PrGE8?B7o$T7^PWrX zrOo2Iw<$<xaX?cl_;dg5*a_8d$rb%3fPRm`>~R@^eoLW~y}=#i%QY9^Hc@$zZO*Vn zco80u1#iK!)ZOD9cx#4BRup_u998iA|3qgBdsP1Dv*KyIu;3?@;3vcTJsK*``ytlP z3GiVFlNC3?JqEsx0KRT{#ryMMc41Gdz=Hc<xuO=#+j7N548(e{jRX9npiG7`XI8(< zWYO2%$$w!qQ-qFUV25&opy#3tetFAR^%eDupfUdjj*f2>2(<JVc+yNH&CCaNbX{_3 zBV^QqNdu;+C;j9a3H=(8C|l>?HQ(W~WAOFkp(w4fg4&A#578?p;A&|M%E5vaN6LmU zvk#F-idGjG7&{J0I~?LC@S{B!P|x>yF=9Xy90arqz^siK3`%}0yeo|>Ga0A9fN=#N zA@*BP12)b`5Llw6QSPZFt^^i1q}}R(5I*CwLVp1#%R?(EVk*;)1LB|*lF=Ja!5ar0 zo8Tq|%ZFk>02{o+@`ZWKIIi^4ig-d0P>Ode^ykSagH)1N=-Z?6M+EYJjFFX!7D8oz zOf2srIOpzl7OTK=$7N?g@^M69e5sVsG~23p-_atzF2pU8LTM_h-03TxVY~q>1qYV0 zyy7rtVDxr^ynvur5?rL(P%*pLLJXx?SPCbt86(t{St_vH)py3>XJpvTK436#%}Wy) zFk`S1X6LiPxj(`Q7C2Lmq&=gds`1#}LJ5Iy3j?i?7@=c21Z^L17JIRZ<(^NCvzVD; z7p{noA`IF#e`F<iSb~X}gFl?#yTVR*-St0HXn7D8tTMgH1SICMAa8*~S7`#%(eX21 z1gj(fOK-p^^{%px*Mi75!-2?G+Ox9C7vjWQxv<bUHGvZ!e@Mgs(IvZj)uSnlWu&wS zEH<vzR<jcHpxHGY^k4@{H4igdW^mUse?E3R4b8QHp9aF=L*`}LK7t%^<_MJ)i?GP0 zC7`F!woogd;qf#NWx&GX6$Yxiq0(}wZowTUXjw@BQ&5`OM0Ji+n%O&E#@>HTkQ6d& zMbRV%?pC3BK<P5kA^aHXRp9c27YN0Wg4Ch*6>T>27zTYvq6+j0UQS1-EfV0rab<zy zgxg*kOusPUnLR5u9%n)hswD=&HMbRC)sJ=96bb_;4tVFA{89$JoHZ4zvxdkXtuS)} z8bc(&8;_epUyIhvefCqg4*#%?X#OxWfqG;1?-e#iTLMvd`z<8M*!o2re$JjikMgsN zTS9EW7j@(hYQc$uHm^1w1EXlvQH|t46Hlaj{|=Z_VuKdP^j)mLRxi(6uLwMzD80!` z7|o)hWynYfjEc8M_x}0Ej%TYtC^epG2Op}jFf)T}7(*Owgi_4MPqWiqt!$QvGqV3t z6atFTD@nDB0S-U27EPy1;VoQDbv7yv3G-8{;4q+E3$r`SL393M{{y`f!YC_`Ar!bg zhE6xne+kDC*zzMdo0WcAk?<MN$E|R`$y7k_6A-kxEFh>o=mt#OQZ1Y9yB~~Om7NQn z9RlMv0Ox_iiPN6kI6Xki?_Y>JHgS4i?mNpf5xIvuYvckG9>4n%6USN5+Me~bC?_W6 zapn)kk&N9}#9d{EuNfD=JYw9|f4+&WBT^tQ5muJv5UIO9bNP4dC3OD5Xx$|^j=d0? z=fGsg_g-l-hvD4CUx2C;LLN;ZqQ#mec>Pc{_8qcYOEGi5bxfh|5%0j`NDz&`dLBAP zmi?%Aa!Qs(yd9P`&=G>gb&^ne4HuhU79)A@T6W(QE`$*0^MU-oy`prIIwV<3Gc=_5 zc%+OST?+C>el|R94LnUk@H7bsO2=TZtJVus$k{io^@PT8h0nnfQ^xI(*45pFfq|@L zg^bbfW1&LgV+pni8cUa_>l%R7y;yW`q2pMnt?hou8<%0a#sh&S@Fk(B2t%v*8wE7k zGWI^Ojfewlue{P{BJ(k$P}W>hhnX39d`lR488UA^#<{yJcLRYryrUc#3tg1(|B>}A z@KC33-@iF9<2VjEB}|b+lv5=+%qXEK<iDj=Q3%^XM=4>b$Wo@*#a4(8dTXPyo#arI z!zz@mDA`s@O3JAW!~4Bw_IW<f`@Gxd|KE1zH}~(pulv5Q^Yy(bB?)1g1o@Q=Vg!j; zG&v7LenEWpHsxk|%<m=@l4)-CI-nH5n3_My6S<>vq8&(M!chD$xolEJy!b6KPdn)L zAjJt?30Zb3uz%Z3ZaSAkODRnvLtVOZFNFxfJNhOYnc;%FCmi84Qy}9+%4wrACali( z?GJEw76fo3(cp90`lFI9b3`Gyio`m5CSWem6V`1eJ67_S=XM*M(|~e$lVDF6ty}9~ z(h~0~gL?IYg6aE?-BsGBJ1`<0uAjniuT>525YL7$cIa}Mn&HD_=qaz!l10=1u8LJv z=3RwVPc)`62+D-(&8mNDLGSvq;F;_0Sp(gHhHba#QzElU*Vrusx7uATI8B?vCTyxG zzju`>s-st7rP`tAgiSol?QJc_{!ElDcqJvOmOevG1$SbfddD>EPuVXE6mVrbJ{vv0 zFOSL39AK)v)hBb;u`;qeVTV5ol0B-^ATzUcBPVU3i(+`?#H&bW9_JV&F!LOvlg%4< z#Uy@JW=ht#6zj9lqf354!}+u3)_Kbiu5p0NrJ2#xMk99b!m%6l`RBl8W@%OKpYQ_G zU^#zy8-|g*$bRq?tc(o2oR1t(8EunNkjKJ2#@9-=b^u!BpgLce$B{TZefS%18Io14 z$`hahP18Si*5ijiVNr&4?)1cQIvKE11p}^dPCKAIx<oo)r8XH*Ev-i$Udk-GnmKQ= z9byK5IXNfD;;4M9H7SwDq`Kc$GMJ+p0#`flxS2s7&Obj$^}l<>k@(s>==;z%BlEyq zDRC}=+Irl&V#&1_pUN-j_=Oz0>%j-05+<~n{JqN`>~?(r02-iA#7HKkYGqV;Z2ZT8 z&GnN439A4a1?~J>yA7_mWLyKND=2>*HIO){D`kY?kI$b(c&0b+%{1_{Rg<*$LAqFG zS3MHuNA)CAGzs#E(;fZ&?adRNhovLRkrDR>OGb=@f!8Ft6M1Zp%M~Xj^E3H-Ek@hq zeed=Qvfx;Fwh|Ahaa$f`@F?vQlPTlt2V&|Ch-bHhJ{p339x&5x(SYv)`9IaGpFX1f z81&bQ1^Vy~uj)(bvZl}FJ0Bfa<qLW89=iwMKBX-lNM7TZ@xZn|EC}ZMwuEtOGDfX+ z0Sp5nkER{e@c4aD&37M+7)DP#Fhk-Yx(?+MM4y;T_Vt-ls{cXgP<P&|Z?<B9hB^b4 za4Rd5IZkR=8BN-bZW(BM+eAG47~h_Z|6V%&q>{>99;l-raj#G62`LDg(NuB({m0~9 zB@)f-`gR3bfoxsi4C1PlA!M9D%c%*@z`nPGb(196wOiqBF)r4LNp<NBMkd%=Jo{Wi z%6L7Yhi(rzg4N(0M;&w8Y^&=!-VcAph@IR?J8qp8dJ;R1s#OWc1K+0BX&s+rNk@!C zPqq`5O9(&Z2Ew@%h$Y!e79WY+GIz-ouhP8TA*&x~d!xHGHCpzzpO$!FCRukSNBczO z;)=urGNUZ%fu1Bh<ba0sFe0UDR6{ZtY+#L-PslF4k4fpm^>{JbRP!|(L5&Hy(S>~B zBs%K8tLMom*7J-5Ju7}4w}qT;nJtV*`<7)D`qBk!;Si-+%0A2jSDDb@);*W7lnS)e z+k;|QoSq5+qeQKlZX0-?COJiRZ&68JAwz}+5PiMf@kpaij|vUx$~y^{j!#dzg7mct zw1?09I*(7N(5cM+jVB!$ot;Js$zgBQ0}gw5LQc#Ml_Ze#n-KD3G@Wcx1;@&$b9qJT z0j*JQDFr$?m~tE?oifD9Z9v7aB$tN=$MbJ8MgxI_bAQR!J3`z`(T#IzW)ecg9U4bY zBF?wtOltAK<Z%O_hb)y0nLif`&(xHR7)r5!)1|rO&<HELMi+WGv=FMkV>`-`O&?t3 zC>H=Jy2Ygp%|p>w&Ca_GO@`7;&CSfsE?q^G<t9G|oxRUpc9DqCE#T!1-(>;V01wh1 zguan1!ahUM_R#O)w_uqr&mD?f(~saP>OuTA{d;fe%+MI3+L)pNq6(VgcKebul3jV! zAI5ah#GH+^#qd3xHg#YmUQ@ss@r#mCG)q7K$)oM~HC+iaH!VZ=)RiMuM+oKH)K%tI z(iBZ!8fLE*pnomlrQHaj01v~pjZb)?-n$ezNH~X2u3qHYJ@FlpWriAnb6rS7dtZ9@ z)y&$J^QD*WnjE(A+7XGoyvAhz79i2FhafMJ#eP^KeAgdSCkDarPp1T`F!L)iYiWts z49Gn7Ioc~{&GfVBmXi|QQI!nFlbqy+49evL*~9#4u<qhq$-I8QQM;5lbN*nhK=)l) zQ`SmlI<q^1klY^46diwo2h7jZyH>)ox^>7FhHWGaFyB!DBX<})Jo5Dn{1v$jY=OSP zy?_4;?6wpAaSupnoD2_JF}f1JAvBI2(y2M_k7$c~-jZo4Yd~c~GfOzM4wF~zQ$;J? z#GTsRRt>4rZQdsb$=H1^oQRgpETWx7t2cvVnp#5N3H`b)ADqu;Jq=XCaiSAH$QoHA zt`aR}U=~rLF7vX#%ddj=FP>hr8od~44VTP-)YJWk<ZXI*TM}B=hBtit_iYUFw*D5$ z8zPy^N3iVeHHp5`aFR+n_f5i~pPQ%~)GUM2x8Lfk3sLsQ4ZCN<<D`KbY8J^*G>bex z&E5XkPDklEi6n~IU4}$48(4t@`Rghgk@Gw0MCLnC<nPyr@Ec}rl(2476F9Qx^B3>Z zO&su+SRAie9{ffeQX{gY#>yivh}||?c;vJd6Aln=cL?8S1<mNrKDha&cQZ1-`MvfY zRgu2kKo)O5l_WV#1FHP$zx@fG5IOl57;2YRNQB^>K(&wLIPF!xs!Wc3x}P)LX(mL` z+-V%-F{<6I*r1c?O+vKW5_W155^a&ldQ>rOJ9o-uG^}CF%pg?UW!4`WvAH3}|0lD4 zkN`zg;SIc!X!Gu0EGE(h_GVU_x?inbnz{IeI6d?I%G^c!=FjxY_1b3<x(K%__iJqC zzK3>xH-g9g8m5KL4AmW|&YV6vXCP2JKI;GrFi8g5c`g(4Mt>ZztY0mRcda-wJ{w}C z_(YR&-k*>2jN*TWT+B17%o?9Q5K!Mgz4F)bO`4W73>G`5joS>kyj|?Sf8nT^g;T<K z$glH{e+!$?A5gE;tsk?Fx7`={uF24{q*QCv+9KEOb;g6i`m(E^R}MUWyR|OexJmfD zndhqC+_`j8OQicDGOSFYC;q{W_`*pU@n1u{3zcaTtRbx@lN|AFl!)+BCkA(?UvN0G zP4`ah2cGNX()r?}6Ct9T&5`HPb#2Fx<$2z36ZS)v`8>hInjt6g#N~Zc#cJa9Lnr;O zTZ;9E28YV7Zc2y?5dAW=dCg$BXnIC+Pr$QAn^D_Tj`i)lc}f{Po6*@BG1=vL^^Yvp zKX_|w%-b-qc~gAt+o=y(*0+DFx62t9wtdW5+xXVgtK!~UTd&8#BP_4Sn|4mk@V1VT z%ivf8=GZU0Tu?9Pbu4sX#iphW18&BC#*slIipChm%Glj{%-J~E%3fH%D<{@GgONSt zm0R=H+DkJhRw-jRO+CXYyV17oXuYP_jnEIO#^(Navd$&1xs!6P%>T}>=Ma4E>Tl6N zDVB`T|L(CXaIY2%YF0Mm6K9r(McgmslSo0$KPEl$s7j-ZejPbujzUD1N1A2G0d#Vn zq`mkV$HDSysaAI?c4Ty~X48dJfBv&Lv*~)YM)&nJPQfY0f?vOihj^g}7+tA3<0tBc zUw37%t@$4HtA5A<r7qQse!pi{Z5|bw<1qt_x0ia<GoH5qpOGmt6Wi96W$;$C<=y>l zz-!<b@*CVU0vE;EjV^!d*He79GW+wz-Rh#IC-FJsd0x++#0w_6hH9VW@h2R+w=_g} zRq)#mUaJey`t@_rhb*scL0^qE`_JFatXKaKSr{*f+iV;ewo@x3>}mh~w+9jMdsX}P zoW<`OBUlX%J+)rWOV>9&nQqbd=aYoDvGq;o#+`1TDgAUn`MRBL3}^GR>$UDL#2cO& z)J17s{k-f(z`ai=vz^l0+_O6xf0MheWj(gx*_Qg>7};Xw>v@|VId0Ca{j{f%S2wQL z7_Z*!<Gk2D4{f2R`&KB1Z?|_|HX1LWoqU}bY#NAkz?9kghbcw>Q7yNVBz+PNUYNFD zT}jsa=)(Na)MGnU3kyvH8AdmvXB0jXBn<?z8|)}K(@7<a)xGhQvbNQJT{M!O$u1?o ze5FiUt7a0t>)H+{HzaJBW+7k@#rK)RA~JnoD}Tnsp)l&ld?(qDP?W;cAM7W#IO+($ z!CR{))c&-5)ET@*8KsjNnFd6iJ$lKE#{{=pdsr3u3Gjzc$9L&3ZmD+%q#2x9D<A&i zmY(i&KRDkrH>+fBrpcM|War*GHK2~kmF$>z?9AF^KBacmfev2!E|#QHnerGI&q-g7 ztav})GsdD@Yi|G9@oB>F*X}Wn`Dez}Tt?AAwEMY0Hp_2Y)dvCXQJE?4+^Zv1)=+sE zC<T<LP`7oULOpyVK<N142&Cm%^Q%tKpiFPBjAOAIkl}i7X68%u{Ejrc>4PdZq?34T zY4^IVhl%WDNr<x-EPVyv=d;A1*~vl5EG?PwWB$7PZ~!Zmj&0Oxq^9P-avxRanto%b z9rQHiP)V7*mBf=gY70*OqBLg!t>*k14odprIWj%A?4PF0Dy{lE=rYl2zpTAn3)>T2 z$Q|iygsd>5|A2nQm*cYk;`lM76Iza)Yn3G%*Q%*9cP~)Em}IXOGCaPUlCc_Xj4_+W zyYlLD;kz^J)g<MB9*T3$T%jvoG{0z0T1*omu>9o3OE9P}UYE?}3XS=w15YxD(z9na z*E*U*2K%C%hJDw5onoy9$5RwR+j>k^R`imQ)VlTJu2%bAWvD0{pKj_sM&{+d;Si}{ z3g<V!Qjm-k#^0#DU>fvbP7se6Rr>ehpf;dR4z6$?v?G<x^7+OzVm>MdCYX!|`>u8~ ze(uj-^!GVI{v7JZ-ryqLBE6VA1rJw1hbdkxTe<hl=SAzhJCOH)1q4?K*u8F&F0FR7 z8)xy&*~h+n3$EmEd(@kJk<n0|nT*Z+Ojk79>!W{meyh80*hO8M(jGaN)0CP|OqC!u zuF?4$H8ze&)=!{-0etOuETvx7E-EFO(e(#bAE(GNgfZi&&I(2VNW|<_M=hcO^5P05 zeV}vi80_KoX0*L0d3?G>EK2s%U~sRW)yBI74(lS0AO}uX7M!8b*_@S-$!3Cv(2t{M z!;br$o=Rl)*nid@n*Kk1!qIxe_Gq%K1K<*tU!Lb(?AKS>-a(4-&v5JFe@7mLMWYUX zI0u0gXiQG%sTi%m^-T}CtJzF{`RW)znr4?ga{4|ChrG!0Yz13|lw5S}Y0aXNTD5%A z^YvX~&Ax7uOxlv6c;vZX`u!tc38q>WV=D6zLmJd5rslGzF$6W5d7)(7fX*Y9WZdD% zuS*n(bNqaTyh>{@rne9Ud?^c;VC>Ld^NYo2E0RP%O2#f`C>Htkezca-LrFmUy8mIS zw)J|1xkYaLW=G1SH5i8sQS?q+2*=#7KjQeFN~;E8vPV^v7sGojN^iZ1r&l5G?wDvk znt5+Rr|7Pt+OBv?s2!PKfjc_0b#G8jG$M$-4m8F?l~057nHj64gn?1!_={xJake`q zZ?#MBSPiVlVW(E?WUM7}9^J(zRl++tqHVop7yGNBCU+^dej}!V_rD}Tbn%^~WRl*{ zOi@gn+dww(#yDdIu{HjT*-qbtC=bueNlB$O9lzhxnvSl@9AED<xr?0iaCOt{e3~dQ zvN%P<he)6B;Q?@}UNRz2&V&C+-lNsqdc6X4`PT^?;(Y&$^_6y6b}rxPq$mQ_UpnNx z#mQQo&0XS@0Y7&jQe$eO8N<`SPkJjiN|zpn98xaOx+L6N1mXO4)x*sQ_Ky5{OB>4O zAAz+G{K3CP{lJBD0j?KWI0xjE>Al^$>tf)2>B0vIThRKKnpqf9g<BLTG4e!tf|^D@ z?PA0o{?E6+U5U@-^dWU?m_nj?z0W08l9=U$7Me8A8eY_UxHXc91f1M>LQ6|n8RU3* z!dYd7-qRAkwcm_;^}|%jEca^xrU2fY*zL1c!L7M|1u6eqEum#Uj#Z?uW#M`^@f<Kx z?0Fmqe0l#&x3nx+F>wYebK|?=w>A;{wrY?8zxBSRZ_9UG*!0_-P&cF5HGesDzKS{b z?Ue0}Pjl;XhYh~{R$q}eQ?2cfHH%e06k41)9^3SJS=-FTa%)b^jv0stXhjyO&<grK zpL?~J<+X{XlU+^=M5Egh?_|p$J3w1>ZumZjt6jnJ@S1ORz`$zLt+Y2QlRS0eh}X(^ zmm&`UQ8ix-8?Dy30=lABfU*Qq-%{;OrHt<|zAjL_yiX$v6`lUOW<>>|5^aN4brW^( z%PzE@O?a~mP#G5#S6i*Uw_O8^&2F?FM{&1W^<UW=t{p*vfhFzU7j;s7qgW|xH1aCF z-h#{c*i_qwLd<9JDq+EBADz@`Gp>y9Sw_aX+E--KbmFSFD^<@^j3-{<4ErSHc3W7F ztJ->LstSU>t8;rMfWT2HT{H3_D`sb`;mfWHOzl#^R43GrRkNFE)tx#MsFjGPt7RxD zPSVqKCLV5gsvcIHd?+1rl5+Fkc#HP$n3mi+EQm*@IYm%&r`RS3U0I7`{&qyATTF-6 zd*Fn?ezxJ78j>~J$flX4sAc9~CPy&CzG0g{^;gqN*3}`Z&><ALJ8D$PCcNllI!;Ab z>ur_~3y-4=c_ja`3^P=ZYu(0gQ@1+*>9M}n_1XuS!$cz4xVNFwS1mMBx(7B2?47XC zq~-j}sg%-OLbXeKtKlP-UDbaqn83u|HVftPeYVMBQ+>86Yx~V#Ip1f#&j>M_F{jy} znXht{V#X5aBsZDN+YBV}yV#@gZEk&trZC|z8C+}GRh_Dy+9=a`_4CfnBy$HFV0x35 zW|r)ue14kO?b(tI^?gUi+Lo+ioC;2$jKm8rAynpaLCrQhDW%qUp$+h-KUDqFu+VpA z3W9tKLio&=jE9+zGUY-S*-ryg@p@;v7Z50}_H+r^^A$n2Y)6BXdO3_cE9lTn;|tTR z?a^hrL6cwdljIC`uMpI{d2UU%qrd5mA*MW1p>1#v8{B;b5G4r06MEiu2=cnE(ET{m z!S8c|<V<G*-_5_^m4fs|Xd3!f6>5&8ob0v%_iD^W8-=>6&&hvBV}r{&m1~1>O)|&) z^eJ`&19>5`fI6`bH3#I0ZzhZAANA9^?~z^R+>=t2Bg5RM$Zn{X;E(7d+io&Za<Gf5 z?8(6nJi4o7D(2{$>5lXSZE|P^{HwADJQ}PsgH6*3&&D220(rM!TB?+;`W%nloM1Ao z#tUEPVp_I&)lY>W=dWhoOlGKLoo(3WU;X#$eY{dQEm^uw6EcRQj&z-lGAeI>fSW{# zT;Fa~N7Mk-WuGy*9?y7c2s`WePx_xa4E(u<>t+s=%RXg7y?t|ny!L5!>tu%c<o-MA zdj^?Vx^!mip<3QaswW_!zALJJn%^PlgX@8<^0osbp#K>>6`lZ=uqKDTj`v^&5Fy4r zEjG@>TmFR4gxRejs0m?o{~|cVEtX|=|0PgX8S1AQG55);WcpFT4Vz6({8pUf<-M^& znyxFedoBpIvSDdvhG}uZDP{cj*ualfN7xM?t)w^XBh6(!gWaGjq5W&7$L0N6WQBKF zT`C}3$Kvg5HM|<_K#)}8MMw|-f!97;iAUk)znj@dHZv?ZseT4bioJ-&IB-qhOX^&a z67P{EIdDTB|G{{LwS|(Ul$rIziQ_04&zU>DE(QjrGynYEybZ@N@j+S&yMcwpWbT~P zy|y97ObbbLIfVyQtY#z*Oqbs3fen1&b9&-6b;<4+3TpHv`_12Lqr8N&X+Hh);uc)R z18D<0V6@W6wRQ$HsPnShM|P@q`$^fv=v-WrhOmDr5k}ivxYl)a8s2y~)7G^A>?zlM zjEt?B>Sj&(cZGYN`{pjQPa0t25GVV7_PWi<$XfZiE+plNsrs@r)4V>I>6)^xfd$2y zx~!=NcRZ^2+-f}MjeJs|_1cE*_SFv*J2!qet$wfbIP69JcH`<Y6yHQER3j&M(~EoC zXJnLX=hlB=CS9;~Onc*>6l3d{ea8CL!R4Ky+EmH|g&Yfb<B;-&+_EKA37P8Ij(PPF zTSV$GX~Rs53lAKrr&o7quLz2s7aG|>%Y145ioegM`mfAnt5f&BtV!CXQ?b-O_f2|c zu$jBhDqF{2c3LLYT2FiQWkZqyZDGTEGcCUgriTYfDLRnzMme*Xa%;>x)R$R|3ZOEZ zc(v}$yG?f1{IF$xttv0%UG%Z<j)P5IM-crHKKL{HK<0yQ*F={rav%L;teL*S+TZcW zVS|z{98Kqk_wD?n6%_h|2b45dJV>!PRnwT2U76s-**IbO>9*bAYHQx6iIwjRn#S!) zj*ezs{qShZowVh}!B4Z=cH5kEzPG`4%YIgj{(}wf#|<9mCSMmtezLFqZhyJr=a%F( z;;n1{?3w+2MqFCeNt)r7!UK8<Fdw<VaWL&<?(TVUK#?vqH3xc)+YrhsO|y`VuZK@M z&#y&3U8ul%FnCMY#MjLJAl(}bHLEx-*S8{BS>>%0%>^-Oat}GQ;D&Ezyn=DI0?y0* z`3>nby@m6Uf~XdHctE}CNF2%Ildod;G|4b+xEtnmpM<k`$*2y@R@UyG>~STUN9MXr zgu+H;xU{Yq`2Vtc5ad))$qZx=un1u6`^7C8RqD3$C$jm5Z=8b1kRN=3ahx71yOJ#g zp3nr<)9f4PejB0G3z<Be0TVv`;Wt^MhQMY7Snar`&m}{RD_NZ%|7BorL+4eZhHNPW z_zle}UZ3lG<=nVJ^O-^b;6(BZ@>WjZt}6+=ZlBmWue&L8lD=?vnq~egwnwH$!z7h* zH&?!SgX^f-JSyedqUB&QO?0)y-)Js+G`Gxkt=1AECf?<6rxyYE`OZ|Uf=>mH;Ruz( z4IaajcO1AyF_fA(b6_^#=Q&NSds#2=r1dtBT)4tAYLMS!S#{)E>K#pBpwQ&LsPPLv zS2tcoq=Tu_dQ!qIwoz#p$5z^{wA)B5@XNRDhtCUKdFMW}JOl_z4!>)hhUhc81Ooyj z0O+jZi!2{b6&+&Lf7h<w*|e0nWlk!>BD+F;W?kc{Lwp~OtdYp<1|xC5IltibZ#!pm z4gEMUnotB`HauUD{~9wvz0<Kr0T*vFFu!3d>7M?<!UMTLqQ%6JK&5vw&*QXY{xqwC zJOq8#F`3;QfwJ&oG6QSf&*J-};~)U&uGtr%2Jed+Dy_10ndWAz{aC%482_tGD<Tx3 zhxeOE#~*$yP!^0@(i8Wu=CkTr4avmIb-6`*S-Sz->Or=tUyj{dPNnSU^p@ky8GpO* z8Z~z2tz&e_SuXm940AUo48DSG2;{f{t2$-BRxbk^cwEHCS^?R}J#M$sfa%e?yp@GO zcVO>gDrKQ(tSJf^&NwRyi#8W=P~c-R43V~OiIc3*w9K>;G?>%CLy=dfmar(@1-${5 z!$z7W0kD_^aCYs=^mF>DG6;!n&>j9t&U!2z2fL6;xs!&ABL-ZM1TEo?p4EBh<NK!# zFR*|OAosG!n@%WRFJPwib39rv6d;Ja3D7lqgmn+=2&{)ATEkZ6rTjbc8nWlky5zwB ziOKNT!lI@nn9KyJswDu?7Xd5a?W8ZglmNB?`C_LBX$<?ey8C4nd7t^5@0`4!&b6)7 zL}n9ov<sQ$NHYTrnGDtIiSiCzYuVibnO;xWtmi8fvGgKK(84W^(hEuuE#ntF1||(| zZtH3g&d?^G0W@z}quw8T&B&6SG%x@}OBH>1H-C{&Ai1)=A3=Flxn`PX5lW|+!$-*q zDkT3MK$CMpN)t{+Bn|JN$cuB3{GTgBe)p2PR8*W_k+gs6@NlawP@gM6lUDVMCcaIw z7zbf&i~~fHU6`5|to)t}t-AmtMMq)fkCJ6`xJB=2Hr>k_{z9%&6S$6cARkt(@2aAH z2BEna1`<^w3+KZL(VS`S<y1mi8u8M&;E>U|zEQ{RQ1g>{SGDA4pobEx?X>idz*Wo! zS6t<5NjG};d?iobU#2yEx(Yjl+U?QG@goMT2`13+?srUaW-;>4p5jfiR*QV*T!@o* z5nF&o<S{NU@c|mC5j~+Qs`tH)So0<Jn_i+N&fIDL$6_3-7wYiq(W#fmFT#f`GFu4n z8)WZ!O0OKCWU!1tT3XKqD72OGJ-7}F2q6RUHeIMat#m>{d03P5kcM9}FJ0_ByB<K@ z))fNLR(`={0O?47YRT9QU?Vy5@LnK#qn_KS_c);}I-j>1fUk_TLI&>k`V4$&^2#iR z^pf#tM!i`@y98-oPD~fE8F_UXOlKW#5n=By?*6iWCV6}{Z6P=EEHc<PxRxjGpGQ8b zWbWp__7H32S054S^&+|O+M;q`ZSON2wi75D_C5y=HDNb<1<G~07b{u2wSDb;izOd% zP=A)SUk^WA=qWZNW5olG&$*SySeS-88;>-1`Qc)LYwo%AoF1!!D+LougfCPwRzXON znheXen)u}{zkN!y2z~=Oe{|$v(rGXBbg;iOPTrhR!WU}|ce1If^+SHx*2x4uSH=G@ zpQ2MGdO0Ek!VCAh5b-7o<g?!HY1O6axn)NB0aqqdV=8}9^J%&V`$szScLpe1s>tp= zij(Z1CsZl5=7wJ)GUK*9Smcu^K(Nf?PVZU4Fk)_1;)a(IiUJ)a|25kKc;>hg0W5ZV zu0U0sc>`E!y7*x4EbjF17)IiJ1??8^2;jc%H{`aQ{+TjruP6lwT1S1UlEC$4>J3$4 z&x*p1DZ=p05o0v9jt`|LX=+=OqIU5M91h>$aa-H~a*7i@dYzokaV@achVdoX?9hqb zIPx~}`a!N`GyC`iAUYx$umj1#4vA5y#n-0cd^F4ixT&mFf!@zl`&MwqKU@9CEi2RG z0>`=nxnLJc^@ZUvKAmhl+|Cq+f4mq;M9fxvI_A?vxZGA%{M_xO>r1bI{0;ynl?FV0 z7Y;d*7{$W-pB$r@Tgg!ee0uY&ASJqsgJ9K{DL>*e5X_*hw6r3Ji~s(ir9_$R793LT zdyMSzH>G9NDq7zaY^M&bM_6QVUJoeW9n`&d@W0_2lK)?}9z+GYSeYHx$QsmY`J9s9 z!NTUC&b~#t>PKeZvC=Z=HCrxpq(@D;fxj_>EJbE`kglVXP{qbwq=Fo5wwgnl14zXs zh5|AJGtIOt9A+0}yy<gj3tL_<i?PDF4CeqS`?W`mDURx$lo2ajxMqE2kGHWA=@ia$ zSOhS?LJyC4@MjbkIuWHJMy0Ps;JW4vvSz1GpH5q7#H2T$2_96B;6Ry?@AH?6#=fuQ zVx!bDZdm}GxmAZeqJhG%yeBt+B%~)aTfOqid8gx1)AASBey6u*kTVx!6XeQ6e&5Ix z)S0H^#P~&cLq1NFB0$<FR~kn2J|Iz(8LSbgHs>LBiZ8-@b!jO#4vtLYZ+dhZq~IDp z%|-ZO;cowYCj-dGh+2l)R#u-3gw|_ik5YB0TL&m23fof%GIXid)Aeamfz>yY$OQ*E zp`Zm-tRptyHt6{czx;?iz;t)En@1Ga5GAGL=*TRx5|n0F^-SXzoF+mu0Khb}B*R8v z*eO5;fBLTmS1vhk&jE7d8Bu(uTwj>e$w|mBN3P(7()WzS5o3NqTuKzM6QI*Hzg8qC zG5`ojK;jLP?vsN8SVQSe8m_}g6yRTE)54SK&97--uYei?1=^9I{L&>xi~GhJP_LzC zd$9B2{p3IcxLG;ePa>#*S%{Hifa<ZfeII~bY!XM@+8rYLSw_r02H~RXsGjhqbLBz+ zHE){m0lTTIE9H>WypCO0SI54qZ?&k|nGST5=EszTZI1Yo!{RLKQJNZVz+(2P3+4g- zrLC0H67|G6h3^KGu$28%`jEo#I&UNb+RWh*9qg?4m<882(t8XrKAf|$`of}<;gA{w z6{VJRwLQN^Srg4BRt`HC9w=)xu&{kK_JZ5K8rNv<a41>m*g~QKFzJ|&^b915uphrS zS0MO6WP8AwSY2UwR`VP<HilK&?IN;hYI|n)IY@px2kif=3%|L7<58-S2CU`j+?G{e zwE4|(iXNo~Epu2QA-&W2V{&*bVl~d9AHP5>xS=3CelZ!)$V3+Ry$4BfX@4N%?dFp> z7zY~de-R3b0frw47?{O0ATjfkl@-3j7O4duGdX}a$#z*m|4{7vMrKl`*?gv&uFvpj z{i76W2Trg%f5u~1=tnNVOwnR~9H9~4)R;N2vshg+FF3-4Hv7923!rM^qkv~_o}hVX zKb?@n|Gdg_qn!EzVcRS!5M$|+VEv=~-Pn4&<_QH4cDqE65p#g~zbux)o&G|NFj7zY z<pS}s5mtNNi^SP_02UvkbX;&Hl>%(=hz((a0S_@1rS@+xuF0dbfMmVbLWyL2VZ*?# zzqlr=vZYOrs3oySXP3R%x226C(M#0&{I;wi3Y8I1ySS0ev=`TMXMb$F_FCfJ{dM&o znxL}N{{8FsRtv4F8~;9Xtih^^JAnpp`zODh+NJS8=E2cxC$^`!y=E?)Qxxi!`0J|k z=ev`Bd^~gSBBN!?e)O;nZ0|iWpm=u4iN?kq%Ri2#J-1yCg>-8pQljcwgJ~(F-Kcl1 zSkWfqq3GElCnzimkZ|LGi{wwO_4=l0r><dvYq`lBhLKV49R5E=<`=)wFH^(OOnO2h z6=Kw@424||iaI4fcGG9?^~vbos497%e()S0N_f1e)(3SLP}k*N#5;0*zx{2)r;k7} z{JY2;8oZh_I{wrP=&gFDMDX0ny2B9`nXL4dnUb%6FrRdsc}fMSv=&}8@i92U4^q;y zxvjogBH$mEgbe@m&O2;E@m-TEw?s+8%9Rc;q}0@)AAti7>T1Wc*}ZX9n=sM*D=V8^ z^YaOgyE7Z;mbU15qfjE_5^FWQW$mH=Z{XB4o-0>{Jq2M$N-%E<ED0Tz1?s|*`&;yc zMYg?c_NJRC|J5zp#rM$-vZW)KWB!Msj}J<OeY6CJf}&s)a#&kjd<@*2SX$Fn{HqUp zeXf~`=cpUj7HeX)_Z)*3B9JH{NsD5^h;1uRjLM)m=4`G4ySEUk`&_Q!b<vhy2)VX} z!&kv=mpO5S)*GVChj$2SwwxA4qri09LwJp@rHYs+ShgNc5xCm_CNje3Q+iRC(<VEp zMY9>EKfGyGq35I0*jX1nbJ!g+PV9>+x|MnI?A~;mTJ?#qT>$Q5Vi7K}^24<|grd1j zd3nULdk+g5zfXVoQKy(?)Vsu=VHDE)l$Ii{^~<CAqnrb3tvqEz0zY9$x%U&uC5~Rw ziCknAHwDzqlf~NOPwcmb8a#EF%5b=vMAo}i(eqvL`P_8A3%7)Xx9QYo)7CHsN{xUZ zQmN>{2DI~`e0?T!IE+Gl#qdhL)!39EI-pyY<Wp+H_Kl>TVHZtDAdW=c0PvALTHgaR z;Dm@y8t2qe|Nhqu{FkC>>*B`G2&%-&fw<*4qU#c>?Gm7<4{P(J=?Hb8yx=9FSU#Q0 zuzHMgnWf?-_^tB4zjY~B6CTfb?L}}YfA<Td)$w8d&Ey6z67TD^q^Vus(`(7F@rm*c zl&Q<Ms>~Ms#NT*hg;;}MaLPmc10PLBbcUZxwr`$55vEljE5xg4GXe6%FTuvHWuO_) z!$Uh9hV3LzLD_3_i8bT~-@lb3Iw_fM06tm~t4=;z1dGZ#;6jDWb~_`<KYaQtB{`zZ z3aAkN8D|7&e{Dr6n@E`tdt#Ivr3u87DjCwnRAc~W({0m!+hHp*%p4l%i`dj=fHAAl zg+-N~v36YFk(Heck4DS$6<R`5la-w;4|L^$&}-1|Mf~}Wn;0xQY9*BRhkdP`Fr1Zp z_Xuh**XjI%Ef>I4(#%^qxD|iu1fnk8GVH)SR41SMI!{oO)9)_J9E_yoaS!8|K7x+o z+M~C$@~N^;f$ok-zHG07@1!>3h;+5{ttV-Iy#^~I3tqd5)u&P5(i5otnN)=2aH<n1 zw(>5#hCU8#NwV$M{DN0WaH7`>yzFIN#7taod_|Kl89RDQ75&hHz6I<s9HDG{o?k@I zv)&7ooL}z+j32m?w^VF29J!sM<WfX8VfZ<~v9Q7wcv=%q!zgChZ#ImQ1sn%OuC)@J zQzHO?+lNz&vLHG}eYxFm1v;~QbHKl4L-ON}?>FZb=?V^+!HzQ<-qmX9Q^^;HNoN0A z)}!?$)ao_>z~;y0>}SD9BuFrxy_Iyey=C$`;2F$ebpicToXNHtv{U&MXSuB=4m3I4 zI*EAtB{J%qs!YT%i+w5=iM=JG#wiJlCJ$3a4NIvyzIfqqC9DatEb4VGPBQ$@OgIm_ z^&0r_R){Tz(IK5_v`P0C+&+;wvPx^gZujE2$)1Pp$`NL7{W}=vA2DGIYxD_9hT3i0 zI?0$a<zYE{Dvq;@0@~K#%N+n!lo4L!<%hxqIKHO*xA`)-J6mN`O3T1wUbgl>olmo| zxwN|F7`anhQO~p&Tw_;uhcdfelP~L~6NdAAFiMdd9zfG6xgV{`wSJker@(HH!b|Qg zpeNoymn&X=+(}Wa<OQ*MeYxZa_W`(Wy+CeQ1)7_3`~3E=g0#l-Pb1(AOq$u}EXePZ zz3jtjId*<)VPC$Udzt0+GD;_B`TG{Q&lEk|#prxPLkRk$XkELLcH{=foEt3R@wQ1< z=(Dc&hi?D~7f&sP?y6WPs9C@uQ#as-U!cnmt6X=`E<WP~`2_~xpacf+RG;n97@4nx zr<Gr%@qHrm_0ODM<iJQsOoVsg25QDV8}tchUBSdSp@gJnCbU%$r}z^Q_C((`#;VVv zim9#4aF696nV)kHC1OrftY}Hqf-La2tu<gpk7zP7>Rwmr>>n9kDvHSo&wMHMpaqJz zG|3ODr>K-BGmNXYYk%1>1;`L|0|vA2xIoTon(R(Dn&kPyie$P+se2cCCh9Y?xp>)? zB{ZY<eZVOiS%9DXIugxy+3+egA3AQSdrL0GZda$NGpjdL%`=6$U^52<Hxqa{WQ`cL z%YfLuf0d%nyghNG5OdWKM%cF7ml`laLRJ+Qh~b&gEX@N@VFbAU8>_!<FekB*4L4ki zX3WN-4w4aNpHVI`#zFFRrrAM1Ye09Dw{pRJ+aH*cc|V-884o#Bee%TcP{&x36El_D zqQA&NR;4r#uQm(^dVUmeP23);s-T6AJe2Y@fCtjS&W87tzK1RB5`$VdOUet1tn!oG z#buczi}2;s52b&?+{rIztCVUnBr{3{7Tm~nQ$fCXTF~JrxjrrsA9sZ(hg+U0pMrz2 zTfgr;EoJ0sDbm756a;2l-7b|XM0j6tN{y0%IG#)?4Maw;S!ou-CUL{kGBsg%*wZT@ zn1*0ByFISm{Ob}Xee*40{b~Qj;6zpPlV)B5Qt>aM`MrO!nV(bs758H^`DgSuMzOj~ zatxRuA?NL$u9xv>d=^m*R`RJpEDr&*u24LSuxasBaNH;B8IUjAE{K5$6>tG2lBq*@ zF#|Q2_#ub<uoBF%uhfpRF-|itg>&orp7uz(u?P7_Ag4x@4*BaZvpZbM6GvS@7m*8$ zCTJDFb(9Fh_Lq=F9GXYzrHDwv?Au|{-EoUS)=t!uIFu;gd>RKwtV4J|-RR+?(rO}- z0_bX8@_keJ|LEVG0XumX*Y6nYF@`BI0N3+rBhCuCFXT<DUWWxU9B4Y>{TE^>l~NAa zd+BDHUwc_Ze_?t0NX!CCaLBbxSr~4x2!^wRJeSe8zdf0jvSS&6%GlZ*XCPegvAj;g z-0hVJqwM3*WM0Im=)RD_;O{pw$}qTH%LF!<Cu`&tVIbob5zilE_ZgO++(CHzh>5^; z41o~eC@#aMc_MA(J(5VK^l1{#Hm1Pe`0D<5H|S3GeO-c@HUqxwl%W2>l7ulxkZ&?C zs^kQUrIiJ-uvGT&?`MuY7nt>_x1KDrYjHlzpgcTEy<t*Es(>B`!P~{N`}(0=jm3Zk zJedG&kM|OiEYA(kLkm3ECe!k%n7|`SpCFa`Q?DjuP&(n0u48m=IF6K8#7+A7(xzr- z?=i=%3>GC6h6mD(+Ifa#WANb3mTkoJ{F89R4qA5ewzIKmBUz>KXT=IwmsHq`fn++V zTqkG*YM`}61akLA3<Qc*6Nnc^Y2m3p9ua6`Q}L9HX=!nD#j^6mkv-%NvHL2Cw%DhO zYr!FWITu*RsvuDUI+xufUlFEI22tM-t&pE@^?3=g2pi{zPlA7QsGJd}c8-p>Ea3^~ z#V8}S_&~XE_~+xYM)da6C7c3TasE#fxIOAshxQ+dn%lRa-R^^*<c1Dukgxh>*&<^i z2mo5wbCPy=#Hn|0MCS&;u^(R(t0&ec>yD%`#o^N5i$9lk_0&=}HYZnHh$M0-R+rmT zOJ7x#9Pa+~SYhYsk=dfCO?x4bPChQ|M;-6XSQ9(|UQ8aSSn6;L5BLtn(4`L$i6vey z$T#3*H&-kgn}sHF&p&*_dPZ1ch?st|vh>9iu9Zj(3<9MgAe_X}iNZX#uD$X}zh~z` zz6-KdU3jWEwLWkmUsq~SCX!%;W;~S9W8!xcEz<Mo5^lUck$7+;Cb41${8zL3R`8n* zh->YlC4(s@{<>w@qn=9wIj&IXi!s$D>{4{lhdLZLH8wxbrIS-vcUk{To{UH7@i@?n zL-P{6u?kW`t&MH3jvzl>kjP);Mq`%K3CrV0O8Bho+9bSJ#s-fw$m4%c1+UIKm`b_C zGr7~JV?lF>o}DmZN<RPjPkgS(Z87kSSt#A`CRp=@AJQ?!uL7@lBk_Uu>kzbmUHafn z8STHoQ&kx<q1pWY#9&3d4H)Hq4KaKR{X+^U(^RcGBw@5F;xj`AG9e-JPp^od#cgpA z*W$|^cv&^sbZ^EKkzvFy6WgWZK5vOv6ZV{rQ^hwdDO<NZt5?8d5GK7)498ag5X_DV zKY>R3%+iCUi-#0G>=HdF&5~~E<Xp_%i$}3V*J<XZp5d$yoO5~wO&@qJmC#<m=3|wF zJ)H9fOr{uyhA3hvAx<`yCtiso8Q*M)r8<jtfdX@a6cQ`4;@Hn;{BHwJJh$&=cw|92 zT!>%cMx$WI3_%cuns;3lq{&?m_+SZe+K)q8D3Ckx)t}!S$b`mV<$Uz(x?4#T8>RM@ z$P=kz)wkHfl0}pD{Ed^EB2St*e{vs;Rk|RpCR$Gnbzl<EGcSXN4^%PlMO%xtRMQM( zjF{qX$QvYk7eo2?FAS`IyW%p!6wDx9ylkVWn^>*ikRzQyRZW!_&w}}zI4aI=+O9fi zPnVH_*(Lr(P;<ACQZpWc*JvbfNxVjoOiwW)jHYb98&VD@kk>vrE9n0C+L6kWqlGGs z25^jsHp31d3M`m}6<ak`{5*%J|KE1_WLXXiHe-Su%tpye|A6N3$)OQcz||VYXmHut zu5UysCT@Etn=Oh&2u5kBlP*y-MJI{r6>rP!iNq!UooF4SL-66@!ogS6wk+NO2GFNf zBFLq-g<@oQh<WWyA0O>+GkPi+P`gA3QScPmB3Nr{obuw7h?txN_=?4cep#m2N$>~& zlTWof4`(%Zrb=&WJlIs6Hhf*HdJHI|k;pUgDkb!x&HruF4t3(BFC;`a>sro!2*V^X z&IzYsB2JVHN<Zr{*tFB+8U98uIl7@f+Dv>?6|(P!#H0^WIvmFcx?d7<ik{eZfzlCt zwIHjWq*R7@%qg30T9+N1P;F=6`g~x0u#>FPfF?14d*T)AsI>RQw21>T;^`h~d4-ec z8?@`8dQZG8ME!?&WgrHq3i5#vkap85SJBioEk<=nV#iPHK<&1gFuP;PBvX26C*@%i z4tdiPE0ozHgx_s0y?@Fhs>X*riK;QF1XW{FLDWf922m#swrtC!zR>zieC7m)*q;;W z#Bd#If@T4$FkX{r!92u8`i7bqHpWA`A10Y(v_OczPX+O+a8Ed%jzd0Oz=cQ%bQ7^B zE<;z8hzE|iN3<1N=~AhOhhBM2+K`7jY5p#Y=tdNEnZiF@hJ2SInQ%OdDK;0}|1^Ko zUaR=|;uR8=MHDY#z^kQp>rff)j&)id6w$}gbn#}%WCNw@jOaNn#Rl%wiOwO+qB~(| zl4c}|e^M#uk@>X89Sr=z5G<ik2k}dRYvu6?J$}KbGx1tNZSM>59H==TI$5x9hb9r4 zQWkNjlpVHWSVhrOy3Fy@ZVRMmYCc>IB(6W#z|(T=XGPPh3RV~U1|?t8zhxRKwG}Ol zEzS#quF^fSuierOYAEo5EnBluU76xezq&wP=)c7cwBbP6NiDm63+ABz7NE4wOf5uu zz#9I>s~S2!f>_Qy)h#kpf&w3?eqx3M?XXDg!-x!_Rp?c9{O{<9bNiV8v}B@$SbUcA zu!x!tML|zTnDC(7e;Vv?KTU~&lsZPKeaj{a+ElLpD-KbhQpXtbCB<*p#~4$BI_d5u z3iyq_CTh1HtcQy8V5z_?b?n-Lga>jKD<0A)BO}cVlv-*mf)S^Z0FwRBe>6FW^`<_> zeXai|LEsuU8ySrZVs>kW5ejTrkci&GCH==2BzAxjBi9$Gt61nWL5gUrpoV{TYhJQg zwdRyWd$M>eYoy76_q96#`EHt=P-_IvQdMyPEZ$*Y3mdNgDH>5Im!R!o=j{cFV{{DM zf`2n?N}{z>fw}d-9>#I1J9XY-YN`zWxc@`J`IowjmhzwwYMBm&#TBk`loetQz^d}p zd{0~D&rEInM_aROjiSY@RH71e67w*&{m@s<hZ#8?GPqEy@n#NN)0;4_+0?y%mJ-e9 zpL#*!z$EjYLjKv%bqoG}|E;3`RM)jTw_eic%qrNFowcE9&F$FE##8;VwN^7*zx+I7 zJtNuX=9G+9t2=9c@;|;RWJcPLy;G+z`zbrCZs&c^Wm>^|ww9N$e=?qIl%H%`JgLI0 zMquabi8uAYeU=B@w0vD1lEdFbploy_!GgV=`(`cQ%8^l`mk{)kr5I-Gc>GNcdkXPT zDnj&h%?!0ee)CFwnnzkzraScWnW2jQVRte<tk^WB{WvUoe)TTxTm}EI`JvCOPuj-V zTo1F2(Z9Z$l9v~sV<!<e+MxD&ik*~4|Ag11zM+pJe_!Tt@NT98Bog&$<K2+1jW;Up zKU>`<f(RG`S}sLd_-M~zZgJ;kpQ<}sy~NA--}ykZzIIrKyP@XGQ$fwUi;6skKPw4t zP)H`OVN2kjU#?@3)<HFIz2;kn4QMknYH4bTZ{AZ5DvDaxpU`97-)h`3<N8&|H+I#W zhn{#!%@k*WRHaaq!jz;eD1RM&Ub%``tfCh1z(bMND+5N+wJ0eI<<>vHL8i^8knc<Q z5089a1z>CP(Jz9aoeI!Sc1%Y~G5KI6^6NvQ=W(^;bU@Yocw_Nj(W=#<|1P2u3RC|Q zS)<Fg<5Rg;Zy$W~oc6Kmi!OM+45gDY@nh0+q&nNt44Goj54<D!mFSrSv}zD<d{%WN z)@q@>zP5+0De+Z5h+aCmGHLa6T)}DRc|)fOWaO(q(o#Mvi)8ESc6SEfnkLk?)Oij1 zchW*a?!45Q11{v51Tj>N+{Z6vij~m%KE=|j5y0E_;0xklSRF&UIv&&Ljz+Zs%_it* zoH8Luydpcv4rw~UIuk_Q8Dl2hln`K+WW4}>zaoo0>&aUhKksqaEm^4pV*`=}`r#_L zYzy?^7q9MpQLb8OHEss$k;Zo`sQmQ}N!UhD1eeKWBF>DSNK~;@m(j3ek-^bgPze&> zP=$v6*f3_0R@Y3ixy}_=@P!w|=#{KChSD*0w{4TH({73V`3))x0<6c;*J`~@n=Cu& z8y%rpuN5Ir-gGNIm-^AI?!r5@<;@aAgR7M!rzl<yr-*|${MRXR;aeSGTW#EFe9|EF zzfX}{l`^w>OlLwqfgg15%Le1H-~=2a!Me1pEpeG|7Lp4?iis$)1CWV6BUYy=8@Jtg z$W!E6_h}<U_g2n^aI`xCl+3N?^$T>r$xiuueaw#eQFI%t@!8cTpu=9Eo+nkZE?bW) zRWDbxvp%4PY|ZJ`m#LID2H-sQQz;$oZJPknI$hgYiEK$8jaeNA$93Cj$ddXm_L#Z8 zPj<zr-_uKmX-NT429B0#Psou&FrlM@baT_)$vw)p<!NsMX*xN7Fcbrk2wzD9m3HTC z`0kz*kh5*NB-~&%jcIgaA&Phs-*^gYR=Ix*#n3v!>~;5t{u01(<k2{1gUAM-X^_vJ zB_o%94}W8nT!0-;FHoq;4!zlhyCa)#b>Qfm`AIKzK2y-qKX1IHPM-g=9KTfX(N?w| zcc@;iU7-ZiLoet%2Z*@-N&OgsM@~Rs$oDyO7I90e_N5&Q`fqBlt^2Yg#6pD)-Bijc zKGThusQO)GHLgWU3A9`ELcT}xn>&@@_g_yw=5Vzqw5$po>b}Gag03kERWg;x$CS!{ z1gae_=sq{G4cSNY;p==P$k(ho@|nAb$rhF2zuQ>w2`}Upq32kS^2(dA49JuEB*=;C zzJ58p>nw5*y>5Wd?39NK7Sa9vk&+Q$86y14?EVja5RaOn4)U&ERMJ(#^<lX~9i6;r z^25$v+ct<8n+yqynbQk5-`%tNXfPH_I~;cqrIscw9S8w$d9|dRN0_iT9JlCty!Xgw ztKm2l-H;t@WzoTNMBFMAg&+WNkaqa9O2>xY*d9A<KE6(dW<laJE>3bjJ<2V5n|3OQ zYo))0&i~Mp;DN<1$ZyswaA+K25p&CCGpi__Nxp*bJVQF(Cc`@p((#nY9JkRE?{Ivw zKZo7PyOSkCTh2}W&ItR~L@#wh!roKr@Q1kbV;9r&bCtneI^3NYSn@?ZZjn>IRY%iq z-td`LL{F~oHP{0v5A$9{VZycKJ2G|JzVX=zvvvIZBk&t%{Xw&FYYo1LY?J_B10|14 zzakzkx$j=zCc>A1OXj+7$}m=9h&Um_joX$WyUusSw}~E-R^;v*PZ##AYQ2sy*3wu< zZm{dRLT5z#SGAfTdGMepLy)FA&$kI>;u^zCp+9=Y+Hiwww%ZNEyu4Z9tbs6vUmo6O zeIQQ++d06uNk$^5=J(nX<($NrsGx+zL6o8FzjG3EFs@w*Qx>DGA<E0}0dUazIlXv2 zrlrh19)D7_mVBg)+B5QzGxs>qz$l^h#+_4!FOZMQkxN%xxSox7>i7l%{mZ<UzTJCD z)sgz8TO1h)v3M(0lKS<;BfpD8;1Q8kvQHh8SR_VR|I^tr9ieqW*)O6QWVZ2|M4F3? z&cFX>iI;wo)Hdnhcz89r%Oc#?#85l=?OX-3M7cE!)%;TsYZYz22yrK}sbb%IYjb1$ zQHj_=0b3LjwA{rqkkm56YHsc&|0Ow%vwQ>75V+XR0ncMYUKdreUmW9G*&kn;eA7CI zD-Bs9UgLFRyEAu#t9T}V<KTaK&k6SAS4>Ud=c8r^q7HGc#Qi!Fd0iaxUz}ESNuY-= zc14wmd~Qkp#MJ*OuZu(eA9>wd%mC*aYC2o_TVX}h+YM3)+xi??OEF*d`S%wN-=tC= z=^-@sEJ3tTY5=>&&mM;5I59uyP=49z-&Y$-M9#1hd1>TIOQIzB#3#3?g2TWH1I0f_ z^`=U!FyvwIB?Mo`mu=7&S%STLX|+Oq-?l^?%P*MbNn&iSl_%kH!X8YbUuvz-qzbmB zbA=eFQvtrkaPn^JHDK7d4Sj5dAjP#c5hNZubcksRmnxLuX{pF(>HsWR%1*^$0R42X zScl_k2W&Ri$`!E=+?L{&K&0Q#Igi*?i@}mf__lUh55sV+JUFpT^7jsbm)N$q4Aqk< zpSdVmov`P&=((a{n<bUg_O0moN;r5!=3lu`p{3FaJ)wTM)ea%3xT+{70(`wOMg4`Z z1=8ap_T+COTHbjFnqklZrN3vuffL78Kq{QP75(9J>ajag`jX*E5^v^jy5t*qHjd~7 zw{k_uNQo07goetz7YNG<M0U}-|6Ai5T+@2fr?{X(<G=hI3R8NQ3sk!hIvm=|KrtGc z0YIel+ZM;$fu>^SFzSjyW+N^ml(UC$3Z9D?NnioKk;LLW6DJQcPv{|W64y6QmhcnN z8eBM$TrW?;9_=OS-sgfieIiHxRrgS1KNNO-7?#>8zSx9cP}~}i{3*HCco=qmMU(O# znd(JIixFj91Vm?M@Wq>OE^O;{!9FcU(O*mCg*Z7Ryk1O(zwNXr6!cdJ@Z`IkiDxJ7 z*Cq@v5OlwUaF`MZ6YV(skxUM8yR@uUL!|F1Yg_L#rG|B>x~h<+@%9iNdx>^Ed=eal zW|L3V#!k)ZN)XVeeOt5k7z-N3<Ux@h$d7CRRPO~H6>UvQ*pUDFumaq=MU%*GK;RCc z!1Ol+4|_tQv{!Qc!*)7kIOH5SDVlvgr{1Ax625ei=f%A~K<80GK#p+v4#EAQSg+2e zA=zHorS-##kuuEL6!J@ZE1&*-;KxPz&{_ofp;f+Kdu=-EcbhytP5!bpUQfJd=()9{ zmN489?YtIlRf2klWW>viK7Cvri832px>b;`C*$E|ka!ZwbSC+fxK^b!2I*&v5{8R; zX<uofD%)ba{{PVMP}JVJsQhQqDki;_#Gt`j!eMoy_y#$XQHY-`CW5qm8^5`O1ol5* zu_WHNL-J`X`hmj3;SKlj{5fa!P`G5?9!o=^C~2$2m~v<F?&J@VoA&EK_(Tb!=-igy z!#Px%wk13tr`i9iM!$VPuVP3}G)uyE^xcE^+DalGy(ahFaHbPYC0`5;XZ@2Ea9^L1 zisKB0HW4z~Fe=jbARwKlsg$}?Hr(tEIqWw{=`>-;cTB%iU0SYq6Gp~nR=6$^moJ&j ztp-w|*mu0t9l|%}TMOVz9(uN;;`y|!sn<8#4(~DF_7Cu0y^bxlCAQo%TH-T?{8O>l zx2O7xknK-?gXI331&^Q$$Akw!xb-EJ5^-0G{?=`aH{*j0t!9|NHw&E5!1h}5NHBS3 zqSTh$su)GYs4pPu7VP?Raf}5aN;;zO5v6d&8n#;_5(mN7w*^U!xH{71iZk0#U`B`p zdD`*wNQU(8mhPfbo*=k@7+(f0<w;8}LWW~cb8(Su!@U_s?cxY;1cE-4?n3x1c|;NS z&Mn56+j8v8X_iN6$51iD$XM&?W0p~SY)>bO^5&d(VgbsXt)r5uEQ!*L;x@#aU%9^n z|DK$=_HY-%G`uBchyf|k<TjUT(VvAk6_yJ|=<Rf(H@D0|u)`bk*(s}r7*JTaB{app zFd9fWvXMU{UReimKL$x25If?&8QlNjmL>W3!yf=HlrXZg{ed9AToKo8Or~mT-+L0x zJ2bL$tt7gHgw@I;l+m(e+&Kh1mf>(B+J>!0fMo?w4^x)IwLz;ehWz&N)!Q}nDbuR2 zfwezGt6(Op!z<<4?Zq^;XW^B~{)OegMd|d}v`79;hi9Z6sS_h7Aq<8Pf*+gR4rm9L z7}RF2f3lvV-IBXJX3ka~qKlCbA1>!3Wtn|Q@7ODxXu&a^16S8|2UvgHK6wO04TL?) zHY6|;NjJK4u<;l&>CiZKukgbqTt&FS@ue}olqH8v11pt*;c$4uOJp0q8s|`VZ_8_F zXRL_TjKnA<NRg0_AMucq)y6w{Z7roYpZ+>zfJD6;r@dlWWOfsXA@dlQZ9AGig}1-V z(X0JbwtH`L9uw-TL+Z#Jaw*DnRZ1CxJ?F=A^p*!adHwiPS?S&<uNl%ONCa(ZnWjo< zCAFnctlanQ#T}Erqourjv1(s>MEoMe^B`W1A`IAkg(bPkITeT7w&Gf})7r<#o<;#X zGG0z{i=XDzf6|qH#gc6IbmZ#EOZ4CKyn`RkNxwo!5GUceAU~EV{c7EVPcNlk)slNQ z0?+4!tW?b%@ois@uW$t`iPqe1Lb;4N@HgDa`w)i?c9o#S$mhaXKzUe3n3KG~e}6@5 ze<b~tH}MQ019Rg}r#@S}k!Wc7i2vI6#g@jqN?B|yU|lWE!id6tj&SB>9BpgDaQigd zc`Xw<!WX~UB3UOKf@#a64+cu?_DY&sz%H_PddS6ax?T+K2tDhWtpj?;WzkaBG!6ZL z?_|k$k%-&<00T|=sU3sF=kPaiA<`J=9*vLB@RU?a;nTdnD;(*TQ6~A#+X_2YlpP#d zE-_2mp+x;|-}e$HjK#e=FN~Z6u}RP&Xa!asd|Fy|xv*_XY5$Xo>Q8N0?2U{FNTcw* zEfZ*%*H4q@l5o?Qv*h{a36@OB<4=bh=M=<dwZ+|d95f=osR=5{jnhlUfGI7u>;q$u zqx{207H!gMsaQNVCm^hO;-oz$@^VchL-=Ch;eB-GpyCTRv>RM?<8fi^@u(vs{Uq-A z@Gdx^?oDe0q~iu=gBGI;gII|p``&%4ih<YU{}D1LceQurl&7G#8od@(c|^!$M=~UW zAZT!z)RdaflLxxO45Wn*OS#*Y^$^*T4;(1NcQ11~Ge*m^`-9@<2C=~CPaO)1i!?Ub zNK|!A_9uqKUeNI}fhbHM5H3^|{V>3$i00TN(HtAm+oNYy<6eeZxBt6Zd^A*+gMi3x z^LOxCb!)%-Mhi*nxoL~cDe@5mdf?y~ov4nD{Vej)H?UH}oC6rde2@iqC-WV6vDmB; zQ9qP4qSB|wd(%V^l$AfnanR%J29xoS_B$TjPcL)wm>!uOjRx`_7%qN_@W@_6@TpI? zqm@P|{QLTssjm#s-_<air>4?a3#Y_86UMV_-l+`-Hyj*YPdH%B0fsO;5*<>r`-6KZ zxqIWj!~rorrD|cu%bZvBZ(p`XT0d&^2A~q(r8#g}`N7c$<S)j-;ke$0#{j2E|1g&G zy6P@i{`04m3M!2zf<88LFxk3)hw7l2G@dl5^S9}Fhn&wLQav>1^V!N{6AP2xHwUsj z8d(;U+{Z?9;z8TTxXeHdO6pOX)E<NpKAKL8;$}x$^D?65{%_^8OR3R2fBlkJb9zF@ zznb8?blNL*8A1IYTO9QhV&QF~B#s^D%z7oVe|!FQ><rc-4w&(0PX!U`jZHX+d6bkD z<vn_f-lCiOs6p&h$IqkcBn@vA7P{e`fSgGq$kgmShLknMpQv9VVKVQ$Wg%DxFOd5_ zc!8cc`jjYL=e;-XpYu3uVz)G!_wGGi?dhN|C6_Y_Vf|>#hQ}#UZp=F@&4KNi<hgvI z1+Do((hJTh?ANw+ZEkBqW-dt7NIGqO!!9x^tl#PV;i6R27rsa63t#xS{2a9Xf$7gO zdPmETwOnq=-*{B1EXTj3M)T;;ZYn8m$Ac2yj$H|fw%}^7Kzc!!x&L?V#KYbcyr_{B zYlgP|gY6_tXH(r}`aBp-9=TLN&brV8GSy4jNR<Al%El)pX^mf^`5U5u1)~YS$n{HU zacJX+$+x8I<e~!e>Kk?vPHN%2q)^nwRBJz`?>1EovvP!!>UV2L9UtvZ4ME3yRE=uq zHGTOHqR8602vJ?JIi<u-x`3{6$Yuo=%XRABcI6-jz$Kz(t8H~2FOgJz9O4?#JlIKt z_|iGws;pr<jjJ6iBUE{7O@_e$snp@NACgjO_wLW<Z;V!sfja(nxnexL_*yn==^lc3 z1no)t-vt9x4=5w0Y{7u;_h~n13xjTDz8Qr-%YlD?O0@#zr`QzyVbCHr_s!u2>c|Iq zV}sFD=>T^~)j<qcdrlA>ui5KDqmjd2t_twhWcakYNzI-16)NQKy2!3w3z)(L!^(qA z8yo;LvD)S`9sp*uS<-68dF}NNzsM)8wi<VO%;YcH7>&T=T}_(D+r<=K@KGDdIf`1g z=qN{CV2JNB-#|14y$MgcU~^Xo-r`7_(WccpujScEH&Kh#xbD&S5b*n7HfX3Q_ESch z6%W2}Z~*VOu#=F@@n)R>XS{ls&g;it(iPv&C?Q2{YUm6(f<hY*o2f{GasQXFH;;$1 zdjrO2?AaCBw-#E2C}bOLmM5V^v>39N(qi9cN~o++N%n^#6rw!YMk!0OM1(@3Y}xl^ zX5Q<}OwaduKcDya`TaBZ+~>Z|b*}xK``qU`=Xoq23`%QN%Y|b)J;k{CVAc?(y}uMe zLglaLPChu9!ZD^$*xaQb=7N2stJz_27?L5tEe+x%I4vo|9A*GZj5R5`Kf|rFaZ0L6 z({@P&92QpSzk7QhrX@K-qnXzi*uh&%z<N>GCrwpT@7$7Lbrx3C>Yw#*F2h{9F0-Mb zTPLW{8M}V{*dq3Z9Zp6y;Cg>JcdGpooU2#K!tv`l8%$<0@g7nxPwd~ooqAX~Ai<6& z%j|hAWC+6wX_c}M-`crN_lG&MFon@)4>zl>i9?D`IK-x<9e{*wTSQPD%w#IBaLN?< zsAXKR4iFC?Vup;iz}S8mvAaa}C&DMh9q@6akI7+Uvy;azF1+0pX-J{@EH6rtL^d42 z-hy3L4xMDu6byA&4dA1dNz=-tF`9xkS*mL=Vmf8N1ZU{Z*m2{){^z`HmrV7}71vDJ z>upj}Sz~EQS3XlcWsZt2?}GHx-Eek8y0Y%=ei$YJ4J3>?Z50;F!viqLgbh`~omER8 zoc&Vhqe!PD2qIL6$R&)R(@C^5iI9xy%I=H{U4c4W;gCXLK4}Ltq^P+IbVm*C1C>)- zQn?KW&vL^6CsbTCh~027L-Z1)AZRJCNyXIFDu08#7@!l6DEtOXO(7*8mLnyKP&f=D zOwYmYC%a*OKnW%ehg817sztnW|5d$n0gBaB?D~<7aH9EdAC3ikSXiWGQWNCH?4ft` zGo&YR8$SGdW0l}ut4DUOp_(D8Sa_KogN3pfIPwe*RQJ$Je)1gFD0#xB@nwJYa2RIp zLWgrtoBgick1(<<44CP$Lt-w>++M<NF+K(hp0e3rECp^Pnf5@?0c`Bp_4BqYGIyi( zVQUhdn||BiP5UdXKC&3KgDFHQ*s0bV@~@d``LnzT=avO~Ekxy;APL@K;)*aN74NUu zEMs>^llkG_QYJb#{8;8qXm0N(2x-Ic+0|DVTz|gla~T9m$U_ES^MhT)10>-DW-X~; zawqJy64h29O~4Kl{kZ3?^vs;G>&!n7V9Y=Jw@haA8T8zyQF@JdDA@pxkM515yhrcw zTM^F{!nSw%)pm@IXP0F4Ci}VTu~#)g08*gbK1*=0iwbF0SrMAA1f|Y0&j%6T{oVm{ z*P~1_+4g#n9I;6zy9^@^2g2gaClB^CPH4%-+=U=*77r^PEM^p|1jA2TeS6#b^mnQ7 z%PcA{EEd&t5QLnAG?|BiT4v%Zws+1lmx-56Ha|D-ns<X~X>fOTSP8w}^MDQ{8<Jo} zI0Yq{!T7^7O_v{%(TU6gpJ6{JE2;qtQ-emxRh<X4Jk|SM7Qaj#mRPGj;X4e!h9GSM zRWt#PxqvQU*c;qpV^oGN%$htqCd~wqE*lOc7u?AEH~JFS0S8)anURBuYDvz4IFxIx zf8nzMq}PblIKxI<ZBL^ity69^14i>kDA!}=5Ez?80Dnrj7L&^5DOn%I2^J9i4mfnd zjLH_H3=_q+cUG((I9WYP*OhW+f=vS9wRBXJLvNFGFa~TS4iX;@n5Dw9+~>#Rv62HN z=FR8IrDFslPKuC1AQ1wd!S|2n$5WYxMN(^*HjSD?(-sN=Sbmsr25FDlXPZ`Af6m%W z6ZR(Q>ZJ|DcW*1Qn_0HZ2tfYv?T~?N=+>9{M~mQpFk~ICU$z~ErS)&w;8cY}y@x9J zV?<I91**vi;L+*gZ-iErnc=eGcQqV<nBAOC*#+ksVAem$Ze}$;rsU)r`dD;SH)$QL zdJKeJ3X~3m4e6#n_6KWR@kS;w&r0-+Ak;7E*O2U*^Peg?teMoTVnyG)?Ngz!&Z_`J zcogH0<GF`L2no~bPvKkVK(19O<m|^Ysln>>@NA_w=GjX0)JiyrWBw=c?T?KcOA7u8 zd2DVC;r|%^x})9!gTSq}I$%=TH{TxwyP`o%1wAlv>?bRL0JHNCup1hpRPeFIV^paE z-t2r>`d-HYoxqPbD1V{j7@ocEqqJe5&TN?ZNz52=1&(^W&Q0XK{l1YlQ7Kn-odb?( zP%naG-L5+(9e6F(0OK=k=ZD`rt~F=q#!kCNP}#0msnN!3jt--zF$pm0INLB0t)dG% zH`tHrd@jMl&|~x~7RCW^&PFbrEq^is58s~@BqGx&v?v3M`m!hCOc7~h?IvjxG<)-A z)|8Rfhgrg~1A*-GIw%US*dXK0rg9IPiKLo9P6rq-<TeZ*g)Pq@%*!9c!ER+8D5tq* zRRmwN#0gWPEW?=V7`H=g9uX%WYGJSYAIBM!IbdcV`3GNC5I+p53V8oG$1%MFG4OhD zV?_xj5hctXo_!p|#INv;&5&beJp{2-ps@{gM)T+O3NE8=%pYapD8R;L%n<jVrn0ke z0k=$F(IS3`1b8M?if5yh7Ue2Tu&{XY#8_)5etP+PU+juI>KIZpgIEzCBu>Ll6S7#^ z!F0+_n79OP{dgcL(2Rr~Hp|AR1#O@Ku^J9;9WB91a?Y!MCeXw*3Bgc(zKqBlvSKv4 z2z&2{&94%l!RSr;voxA_g-LiATRXxUrcR<>!IaC!-q1yrN-dN$3*WZFzy;WqSq-fW z$j}>7^dON4Y+wV!210f|#X&B?DZTy5r6E1>GFD44IE4IhiCbxuTPvyn<v2<>_Mgsv zc;kXL&-DmB+s*L)#e#Wo_bCa3#+n5ldu0kW{<h)c_+<jEvQ@Ed(UW<Iw=dQw3;jD# zSoazKcQuR~JZY~84Rw%N&F3(|GD=V_4e^tl-B0F23X6C>bgb{g+U?I963QO+?~KT| z)kt10Q~8$qI$l+x`CEp_D;1n_p7X?xm0cgnDLQ>ll-<9VnKJM%hVFkPrRXTxmPhL} zJ5hRn`>I^~QaJSKqt5lmi3y}fk4=j%3ml<-e7jGf;vQ_~;=y}u=*8TjjVT`|_*AGS zUZkXuwywbK*YJ*ust^5z%WSHaR5qs8eUsWKehp_k9rUr{`ul;0q&w5W4{<ReIMo#? zu9M6v8SZ9;$SNHyYxLn5e-;d9(*bnR6fB2>vB;7#8SqtxA2M(_=jB*lXZ1#49OgJ| zXD>lY0v4v{Ij)+LnVHNAB4idhK#V?|u=Rnnp3~4+0h?PdWWih{9;gu7Kk<|r22VZk z?0PM5{sA_oYgu!R-$hc!Mbyk?-TnzzRxHO=RtK+2UJ^^Urd|s#M(`9+(Y*&Un7?o5 z&}&h{-1vq`urE=aw@mBdg$W~U$SVP;m&i}#Wt~JB5wMliI+<&~y*e&3b*ljeCc$n? zpjQ&;jp}LoY(NaqmVxYC&Rw!~Iv2Hohyg4OpE)Fv^A}fRLmu!Cro3Xt)^*;-B(cJh zWFQ<-M!1gHVuu+J1HBrPSMJ!TAT7u@0rcEAzO_Xiz^Hh=ARnnp*3xU4;NDg&3T^hV z0+xDwFOQ)=%&Cc|wKxH4FK-qZdwGLq=C4YSMj@h?-Y8S2Gj7EqGvLCa4Bv=c(1%|n z;qgNlm9gfko!3WT`}ah4!O>?R<1%@sn1ssvlFjTh@MX44IIU$pv_26@6RH6ahexOR z*5%=t@hsi2K%G@Dy@wB|w`Y-24cvJe&+={JoE}V+b6#YF(lP;Jgn<lkfWppSti#|F zPIDY&$EaxWIYO2#e4p5>mwplKp_>H`;$X%m3{-}T6TLu$1fFfht02R13Pk7|vsnIO zb|yF+B&N%E+sLN0WXSLc#z3wMv?321xxnXD^x&|bKL<>9L-R^NwuwGQWnmNXwI^g@ zaGPbTBW!T2V~tJeNsJa1mVs}^hNuP0EEXHU4q*QEwM^H%A?-tuAu`CMt{E55^My2Q zlsbb%mK=(4D97fr&D;izyk;)q0)r_FuvJ=}Jd;W<F&{h54XM#Abd!Qz;Xi(UoVO~e z90;czHV_Xa?guWdIrN-HkMnA*e`o{K_w|fYa~dpc*<_AmL3t;1*e*CoLYlkI`0QWa zk}YWUx+)kW*tqKucnTQgJP^UfU<)?ntK-8uoD&|A)T7Tx(?Od-1f1+i5w;Vs=KCze zZgkioj!r=W_4s<dn!sB}unmxM#4Ol-4Y325GA>uuSAB!88UDI-n_BP&vXQI(UW2pD zlvo(H!VPF(s0Dc56C>VY)i*sX12x5CILEJouqp(xb=!<_5^ugm1EzgxEs6#p0N4T` z?1m#$xerz5;^6fGX0Ept3!w_ID<D)5f>=Dhmw4bblGg>Wpto~vFlubE;bVYd5m;(4 zI3r}vr$0mH-3Uv3?N+oQycd@+ab9EvA-Nv}&fA1xel6`8qX_&L6`_pwiVbEiJHg)l zg_&ZGJ*;GbHZS7RS}ehV7B0~w5Lxw&2jJMvT;exF3l+2$jEF$nUyMd!0U<Q{6csTu z$hDvqji91~jG{0a2m4S_2BWB!A2H>NiUt@(aDpKeg`gsHMo|{6;(Ju2$|!QC6>UXD zvy38PRB^?e-y<HTSRn=HY!ukD(q@ck$ovpu^cna95(gDbSFi(r3rWp_!N3JXE%`8X zF)m*)bg(`oqA|K~cQL@e(!iEct6Gd!y%8ZRT(Bi!_aYeO4QLY%qlPwxv7iEJhP-hQ zv;Z0GK?}L{bmfs*u-i1G_fH65>bAk5hBsbcfZ175hLv3d!52`3tG?}pDPATs7gZFH zc#*9#L)-DtWawU&9$$FKY32e)km!uV&Pf>s{Z1^iivaM2c;E=sSml8T1Boahkuc#8 zCm$?m>7|dO2y+ziF~EIe%}lT55}RPR4cHu634p_xR5-;aBh$!Wy*Qv9#SRKOf7cG0 zgp|SVXJ|8#ZN(J`I}UjoyMv%wN|fz-qHQca2qk#XkniI7CB@Al$4aBU*?|)*)y;~z z6lR15F%T(<GK<Dfkmzbo6JN74->7G!$C|iEsNL^X3$p^YS7qpCSTBcooAyABya0s$ z;GFevoN318vH=X1Lyn!U*aQjacm&lhVY=P1ZwAfitja-dCNv@{gA?R8z65>)ls6Av zrXlRrTV<Ac_zlaz2%ZH(u}SOU<d8F<^RD`OGdM-(+*y`uLj1-PEZaVxMr|KM;RrF5 zqObf~6;J?C&~YVv6GDY6ipnQ}fC+jxobZhq!ue)w!o2hM%~gX|dddl`w6G{w3#Xd~ z2?A%!+gRdY!B_%2y?haRBhJJjAV>K%^g759>nji!zkI=D>tRU}vNGd|Fi^`UR6NFS zyb<%p0qp7~l*10FysQoAyFCvOQ;EiNaXQ4e4v4@x#Nuo!Lj9Ccv<Hb|D-!q2NxDhP z)gqs1B(SjnXVPX56`-?-V@Ec<YaM{S5T?b8=0O%zbJ+>eben-pTLcNfU+%`lCRl>9 zj;{Mb2Lw2-0wKv*3Ty`gobwD#1LN<(;T~$QXCMTz>f3#F1n-J72U|i+N<rq8;HwG| zWDPqQvT&t=sa{8I_)5?nrXJ3{0ZC*K>m}d|0G2L^CYm6#uGc4nI&QC_8CsdsXQerp zR^RNCJE!RF<?R&KmC)6t<?RfaE>{KI>fbwR4^dCP%iea(#!bYHU*JSuC+qv?-03^K z-ZTC6j%mRrBRhK=r&mjPyuzJ3wKJ7`_dc)6Zga|newX#HQ|e43(%Ai)>&%W%hNV4( zmFL)PX00{`GMN3EqGqwL_ytv4;{=FITgN>?W*$4zwp>JYvt*!dck|(b#RQ^^4v*RK zXaN8OCa=dX_=0>4n=$u3Ut@V6F=xz@rf0Ou1ai!DX8vn*-;Hzne!NU(X|Qn!wBe_i zmYK2{5Dv*Agu)mA6UH#WWDywojP0Z<cE8{5v`i@k7>`=kyom6rVz<sq?BBf^dMcpV zl^KePFpMGzRP+=bj&8GV+tVSAJE9{H+^@%)u%aSnM$t=D)byTXA!Qv$>lmOD0d&cc z7{P5LXF?IbybOM+kga&n<GCIZ3m-!)%e2ZCh+APMI?1IRsM*)taB5J600Znj4NM3O zWTcT8TO0yH?!7^40g=+n#vo;!PD)iTVpg8PtT&R%gXn`uQtAvzEg}Ith-QXnCz<I& z_Sg=Ra#+siSRe>NDf_+8_DJ+tB>ECq*?VIymC>^9E!6UJ_!2J@$iW~GdN`)7J$sQv zqyO-+9`SJzNi-V!SL;Sl+KL9!Mi~>(th;WT(P<=6b42rm7*Myu&7iIs8Px+S=Wpr` z0OZ^Dv=#w!(E!zvIx1Nh41>2rWp#*!ttT0Z4CF!0j>!x|T1qCBiNOXBVq+5av*FLo z{R7B_0JS6_OpIs|WBsB)sL%I<ispD(K`wm^xy%L9s$NDCVAeCzF);gRh&zP9H1(+L zL>V#%!*HXK2zGvffFjd_VbCX!=(`USGI}vVf0Q&tD@%orS%$pc?07RW`h0$<TN1df zV6hfz19$k_Kur)bBJ0<@2=b5SsNtw=c^{IOG=m&^85{`(b-5WWm!MWC_5u=*F#Ahc z86-|363J*J6+_`*09qpUs5F1$VaOk{Gg<#y$wHbVF_gI?jx<Nr1}R7|WVLt@2|=0* zwV((tgUlXQO<j+Dvm4C0j=)ZDV42ovusv#&kKxb(TF_|ycMVX-pd<qwACf6IERKS= zFVGtD`*I@iDk{MaHRBmRSpw^1d^FO&$0)8Q^A0-d#gq|aS14c}%Q7^SB2c}YiQe+- zXj*l!o0u?0BQO=sv^CfOBE;Zcy721Ckm#$mHg)=|$<&oG+&3RdAaoH$uwKl)%w0fC zfHLxrC0j)EQ9YyX17;uVN(>Oqrm)4)en=6fMTq4nB!Yf242x!xv0GiaKq}J6RYi3n zmdi+U(#?KZkp@-|9=&{-L5cTagbmjJ!O>w_^q4W|j~o>i{zIb#7CIqornIV(*DIXi zG&>f$mk;_QFSo?Qo>w;lw$SxBD2gbC{a7$J&M+MCIpRMT4Txo&8LciNlh3n3`v?XI zGAgdpkSVc)8W$lDAYH6AqgLC$z}Rk(onaI~{HTFqX7HfpHw+2WdF_!xWq5|ta45FE z@RJ3`bb09(KSQNA%MHqr=4FTuOyqj(CzN(r7tJt%H5a7T;RPh8-Haxv$ir2-v8XB+ zqv{oURRp$<0hWrwl(m-7n_fP2>K{W+1f;UOWX|7YWq4(S&SW>lr63i=$oP4rTXRNu zz_B8ty?9{aW~Pn7NNOlTStbVJgD=Q5*PT&CBE9ehnbe=YDG2pV4r?qdVW9KDA&?fW z5h+v}un{5T!q6>)#E)1BqBRoEuoe;wkIhtc*^vZ|8A^f)RnB`6@VOTb5@e_{XdkT+ zG0@xEUkrQ9KR^S+0$4tlp|?{=fJ^JD()rj}mghk>5fEA6l@O$W54N2g0-dna!@~># z8CwS-7@2|;!vJRydt+abPN0LLMe$11LA{``BWbBKOrcVQRwj<9abi&8PE)ZvvfYwD zL~9|U$$Usp$qYU9g&?7l*g9AD>BaQ%(4+mb4pL~nJ&K%t><phdj#wEygxXzXrqe>f zAcR+UAfXZ%v`}PF*~}gU$A_hNsuY^{&Y&U0;GYhv5_M22`-DK71a5{lYZ1F~;iy?0 zqla3i^-#=GyuYdIp%na_7MyV?5GN#YAV@w2fwV3PzO4gyM;M(YDFT6&V^AYPj7Gko zkQ0}MTnTkhH1|RsRQ(?Q-#P;PgQ$l>(UgFCD3x*4(Gu9{9qr6sBo|j{WE}ns0g_Oh zOYnkDaE6azS5%r^jU&-Z|B=9MBo#j;Bmwb1U^G*!5C+Gjh%ywOibCHLER7I6!*nS^ zsP61#R5y{~3sf<rabIs#MrC7!CIS*)BCKF0%24?j!Gwxz*G&r2Wo`thGRpc;53jZX zX>AvS3j*SzOIrd}6<`SJ8Dh3^GjMkGPk5z;p~lT9yy6+2N{a_C5&&vq10C6wR5E`R zHs5*nuAA<o)!)s(uDI#mShbzScfuj5N^ZIj8|~`}=VpW4qz?MQLUhO3Z8@Qh?|uc= z5o~7-)h^6Tc52^f+%=<?=<R$>dYM}2tC8rvnHt)-Ju&RAyTi5M#^9%7YG1m($>5uR zJ#6gxyxJ*I@H#v|GI6Kd!qdgVE~g>`3TxsKw^sn>-tcRaD_<g^_E(?zUSn#F(|!tT zTdFE~r)$+*mN7xU=bWEpTWU<c{fg@o+!mgPL9ruQ*<w$momO_wcp59&uE@5+@{5Nl zrh!atUsQGC*SMaTI<IK9-Wq9pn4;gVO!j_~6-#(CCtz&4Y20=t-&oqDT{&=DTiCI^ z#;R>$vAnjYRF9tEY4hLxtPbbs$I|8x52&|CbaO&SdJ=krbL{+1KJk{vH_s_Q2{WIx z_mh9(EsL+6i!z=()UMpq=)vd8*LK#p7Es8xS>%MerjEn}>}um}6Poc(c$MG<w1ykU z!!J9m$Zc2Vd7@R<_v?GJ^5K__Z<+_co0?y|HkGWgG;-pr#J3;cwT@pbnm)|iYxJ_Q zW~g&X+oHP0xob%RE?J#R3d@OZ0XHH|kMDI0xEE;}Lu#rKatpY*vPDg|u14qtk9hox z>D15iMbo-!ijU8C`tXTQqKlUJq%0vMQoU$eZ8WaaXM?!=lk=Vb5F^8%nmP>xs41oa z%r5c#47jk0=Wiw4jm#>TzS{aQ(!9%u6I$NT>9Ywgq{3;I8R0I-O$^<0Zgb&X(A`4e z6&85X)75D}*(82YZ9-qHV7INkVS7R0JN?8Qy8%752X%Odt6bt*hB|U2`>R|Q)q`cm zl*F}ta(zx$dD<J%fp3W?c;fO^m)Vv#80*#yUdY_=@XY%`dnhbXeI4t;7FVL;D-QKG z$+^oC9@TdlH^~xi)HggVPw`k6r<zcHFUM|%NzOSY|J+s0xFPeNkAn%uDTmvYeZ}IO zV{mT=Pso{W#lIP>H#SK>57~-?^@9xw+IQ=3Lj$#{zD$65x3QFIWkmh7vDCqK<$a;` z8TFHnu@ALxt;op<#XVZtBWIxeT7CJMagFi9U47$@hrj<m|B(A3_fs*soyKf~M;=PN z%QUEnuHO`=xGZUW%6Mb;;_iLKdXt&c3Ev*nV^8pm5$tCcGj;N)A@ii;d-M7}tfXd( zbx^BTo(-xORyM9o*1Ox1)$^!I%S;`F#uayX=<Yz4%L$!NEuwJ2<$@Jebooq2vjbV{ z#MivO`}3;Vue!)bI{31>ofZ%OrtE81D@(1O;ClM1D{xy`>d}{t6MMa+53Fn|JK3f4 z_xaAV1^$L9s+65&i}Iu!)6F8(kUh3*8)|TmFYTtt4v5#~JFRW6!O8H55N=G%KNYJv zsZx0-$8Jra=433scIjXQ&bT6adQ}v^KWOj35mCH8+?9!tqSrWUCLPukb3&)ZMM%NL zVbdWIN`IdhNo`othO&JT&09$k(-%aV>z$Q#!>4oWmCmiH)Hom88`12rb{rtWMMw{! z8I~HQEx6R_d(#6z|NMp;cJqrxW3|6aub)pmEBs>2;KbM8nh(wjn?b3bm(EQnWiJ?$ zuFEP~koW3&5WsnX;4p4E8kk2;d5Mcjec#PzT?YJ(pXvzGYcE*P>8LbqFYw~I(6<?v z{L;N>K@vdEHg=cPnL3Uq=AAtXFxxvS_v?ndj4oO*YL(49`@XvtRG{cQ9-Mj5%Af0k zv`p&T?nEnp)%gbNvY40QPx<V}13JKOto%hU4Bh3m9e<x!ac@Pr!=|n~9BP_4j<XKR z7g`p2aki^v;0o0;x}7|DGEP5zC0Y^t$vgX%X!85vpC@=S;*~YS12$y{@AlhmQ9b-R z!sbwrc82qA&E%CW`pph=3OS*l%tds9R~+=SRQ$jdhi^xi$fi~ePe-KawJU3;d>W{$ zRrM3hFfiX6A=NZ|_k_BtA0}g<F51BZr%yGeoYWV}$f)D9^$^MMk8`gcHq-Y`KQGdg z^2uhVT4elOabf^>hNJ$4%p)Q%MLn~{^tD8G#YQ`xP$}iHsCFUfZ(SC#bayJ29$~R8 zv93tSv6~OeqbQAVTjo2?v!CF}?K?NWI>=tIB4kPWtG&QaC>Jj`ZabfmXJZ6+`*Nja zam4l`)?`e3BOEy0zG${L&*tOk*1-(xikNow*vW>`q;{^l+wETp{nVCw^K|3dzr>bU z>D+I3j#(^R;kPuE-Bswvoy+$`H?lp$I*{L2DL1UZ57&CDU8~T~IVZGzU9Pvxr2V|K z)miy@D;;9{*JG(Q=iiK8tPg|i%-q#DKW_JPOBZrN`+2wtM!mID=R?25c`2{4Rl6VG z%T4+)b>Re0b*@OzeGg1^v`Tqwj-AiG>gbrs`YC2^rL6O4jTadp7L@~(EUKs0hfd1& zMfPinR)+gYSF5YS8K26b!m0rq=5<4dR0B5qDnS}`ifOxYz~1U>M*bp|vHiK(Vxda* zJ{zHS0ykb?lh|)AsvEqb46DEbK$}wzp;<AL@1~5nO`n_>NqINb_f#yD`eABIE0#Oj zeu)6J`Cw(%jHF0!YX9!R?bR8gy{SN?VD;*;z1$nBjoX#WH&;K>wu9q0bWVKbZnB%Q z&$-B!?|fO!qWZfhw+#^Evp4w5-CAy&rtcPm0>SYr8sXZumv;oK8hugK2zm4L1W)h= zQSF3p_EU%0<UuG>4d2%d@&*rxYKP|7O>wYA%PFa7M7^o5*AkUz_-;2yu)WM392Vo= zU|aKj|KL~Ue}scK*RQIIsA}9OerFr`Msc)FBu_=;?rsjYgRSeyq}_V$Ez}KUT-;(7 zb;C+YtlwOI7J+%*b&ha|TI(=}GoWf7kES>e6CUKGQ?IO~ZkaTt9$%5(Jf23iS((f@ zKTHazY;~Ax&Z3&M@~#ZE;-YiXVPlAW76+;C=bGgt^{63p_(Sb21Szs`4CMgTW2Jd3 z`8MU+aH<IuZX`RK@0~NvpnBwlQtU@?k#p=TxpSsz-B0?maB1g<^IIn2dD+}aLsACm z5Fw%XD1n`Pcr<Q=%1o9#!9!*q?u!Dtyyy62$hei&HWJUg(cEGN31?CLA$5d(zCLeH z-Ws^YPm=ZWh7t9=Z*?DFB(W3c{PSd{xb^Gn_fzi%z4_pocgA#O`%b*=&#=rf>t%(V zsq$xQKa^Q7bGv^0P-gY|<_e2tZ_@`Yqc1OJ{<7@V7b}{%YH71S+1hvaPUo2C_CKfd z%5UvC_p>lhI|@Lhsz0pDykKpVCt)g9c)akqxwGX$;Z>zp&%6lhW%*X({k@i6$5N|4 z#M~DcIc2$E@2DqMaN~Yj)7+)UUt*qTcr#bvKn+30lYXNchRHF#(S~1ISGg;sWBlfH z4^Ot8;Hk(R4ovcptWa+)i!DAhxyjExP4{~5@|@z9vqvU9{D$s67}iPby_PNJm-We? zZA8;AI%9E9E3cn?i@$=!!O3HDQStzIeNM`#d;etWhz>C)-CM8%Om=R<XmVRCc5cFC zQVN=>9_fjkJL?zLsueMJ(a(QG@~`%mGS!NMfY>{&BrGTWOS-pztNOm-Q+^k7F7o>A zVh>C`pBTKJz2XsEY9vmWUu>&*5bSjfE};{PhWM=2i=}<(=Z9^ME}lNY)8pP!@n}w{ z6pHb1m)}YjvpBfe;6Xaj-a_8fqjWdr(4x5q?r3{UMZ|=>5kAv@UC$At)U?&}rIVX@ zhAU$ya<fl+G-s}Iwq}h0EO#q$V#7#^VS9_tg9(z^UJq&GMXT1KVclC3yFafE_eM`J z8>K$+$Cnx$o6PWcFI|)kyfG0mQSnr))H}<6iz{DGS!;#3#X)alZ*};MoOtLV)FV{7 zxz%9!)yRV3$(FN!e|>U(Sh3}7=K0}?%&&i+x1#LqVK#~{TGM?%J@B$|W`lK*QoC9M zwPt3(D84X2yCFLLyj753L-Z54Y-xyw8oE(G4Pts5W)4~hiC9$66e>_*#5F^<4&oRw zQjl((AwLy+5bdxWlzGI8V(_5E(B0AZ{!jnd2Ub!A+L^l7e>y%`%~b!iyrDs|^;VPg z$Tv{Q^4{b=;Pz7H?1fgX;kDSN{OrX~udg>L8M;57wb%S&KAABasEONDal2`8I3P7A zeb!U6Z|ivatc4~X95^7UNj_Ul^W>*xjS(J?=V{lrjxcG)e_ECqIinfh`Z@yOZ8VXm zmnSy)1E6r*>zjMBu5l-;AHUeP?5xSRr6Q{7lgET<d&|B%O=N&re0)vN;-GxnYd(vE zS%1eF=08!*xaOV{+N3igVdS?y*}vtryv0Fh_2ja+rnDk`XJbFsWDCvB0I=V~?*vb> zVsbEGxv1HY8UN|EWb)S5fe{J4A+HlW<9(lA?;7FmR&S&18d)wT==&!Y2QQeqQ19iW zlec!`pPY9gTwjPcRQj8|uRArX*I-q(7<a@^cigm%!fkPIwYnI0xclHLXSej<=UwoJ zR(Sv=7yQvoUhpk$me&bicwO9_ao#2M;3_U#tlPOQ#e8pZ>XFs?%&a!u8w>CTP|TM_ zHqFdxt>hZ71_*ywX;6Uf%>|pB(1m1|zJ#3gim(OAr()e>#glrPAq&KXORaG(i>`bC z%-&)no2=(g=q}lsYzSNS8MGI8a2+9Kowr(%JW^t|SE1))t2?Ow(mB_}IiClY97bS( z`FOH$ZfEN{g}|`Kld)0GBRdrW6N|$hD;3T;!b83n9;_DCtv?k4V?4BvOgfJ+e_pjJ zz4y4mFxp`xP~j|e`g4{KdXIEhw<cQqfu*(PTKVmT%T6mlp%J{bX87ayIH{V}1}n|* z9J>+Y2R5k2`Ug!ga`M}8L`hDvq18@7C**wMds3l?aL&bz`L@$aC%!7ARR3Uq`pR1S z(R&!#hA-MqZ-Zx(AAa;?i*4kypMIXnUf{DW+&jIu_J`EQ#n>pP=>>(}>wePVV+xmI zqwKDzhmW1$3D?<Jo2jGVWvzWH$8K6pp*PBdn~P*uxP`~+pW7jL^Af8iQ7(0;+<!!2 z?cN^SelCT8yE%6KoEx8Sk!*bHs}K-V{O&EG;IiWUI)(BFA=eCR3#RsLv^gkNI3>6- z1z5SP%@w^_vij|A1?@-0!6ED~retzLL&^%wj$JHLcC?Gjv9k*f**=0*$SYC?THo&A zDl;-DQg*Pr4^M7~U|spRxD*tQ?^O^~xOakwYqiiUX5z?LvjWTZPb)glLOtA1m9(tr zSR5QHDGoAdZ|Q#GsR}pxr>F-tZ*({o2kn1(bgZYi_h5UApYSv9INtPm-cysAy#`|! zi+fE<<)4+b`a!@M3pnL2JD%>T_-r;??37a4Jm+Yli%xXM>{HV;Pt9jTu@gpPpNdls zihc6i^334a-eRdkV<N3Ou8l4_F~z|h8=q}{Dt5{{-BScmnVs^M?TPHLITdz%Z}Eij z7=KP^hoo!I?VR+NCKvcdluzXrcNm7XoNp70U1&L#`4YP3(p}GzGmoBfxqm{vXtw$H zn)v*Y&SiNowG%vA?&pS{WiodzU-RlU3>48Kd}vBNw)gj%t`_0Iqk`G5zt^PN8-Htu zC!d5Dg4zAQ14P@^w8DM|$h51m_r7hy!X@YdGr{0x<8|T7-<`8x?VGwH;iC)16mfRl z2j>%+OEji-!`s6WGkb6N`9w*QeeEu2HmD)0@9u-HX=bl)#$v0mvk%Wd4kP>dhf7u# zo|SqMoaAJ2ELV&?x;))|L%(9r-FUyJ{>0JsX>qbTSBOv8$1Ia8uf=Niy}s7;I=QkM zuWFmqw3X{rkC>i7j=SUgy*EsnA1-`)zxDfOk6a@4`LL#R^PMqWibuX;T~A0LUW=NP zOpSaaszs`4!r5M~Y07^wdqI0v&CAAq=%mlDot>{6nl5lRbuJ(I9gqTGR8=j)ZRg&a zrs#rMVc%a3UCVY}K{w$+s*`ig_lRWNo2D$^Ul)HHcP;Nv@nzRab#VRw#qQ3|Z>OBQ ze4~=1Rg(&4cY7($e*D_Ge9Egb#JTcjGS~0&M8JW6-z3$n{iaE)%Qx29t8@9PSLKa< zz8ABOpaeGO-yU7Q?9Fa3+U%Tz--v!CHJT(kYr0VtzLMU5m;GAnMs;N4>r(8hY`R>t zAp2Emw}!p{Tdy~&LtT|`7ra%H3dtAU2wK6auapYO63w>$RSRJvO`R0XS5s9D`&At` z0&_Jreq863lCrkls?jUb9cY6<9G9x@b6S;jqXvblzUpgQfFY~BfB3k%>Y9>UN#`Y1 zpLZS9>W)d3fVvKn{TNBv=bfd)n7XL;xiNg<JX^K2!`PRum3R=KUsVUmc1gN(<?yQs zRI9-)@a~&~uTt#C9>V2SXX(f9vZuw}JVhyNr!T(Jgc@s-uR262scsZuHR{f4(L}0x zA~oTSXd<cp*XB@|-O_TShy%xKMrwpR{q6<J8{8<qp$2XgYc;Awb=|M%!sP*Os<#@h z;a9ZEnx`6dPa+|-ALl%Kv+*l8;X&iRM4n((HC*Gb2*8<Nu<RPDs^<P;nT%StQG>(_ zoZW0$*P2u)HKsAXYfUkf^x+qeTC^xdM-BJ>7Z01Wn^d81GSn4H+zW_E!Hu!~foi0r zMtA$Em<f{OtR%NoE19R1%3e)6xX4p_xBun@>F?EiQ3_8f$z&0)vL@3*fUEM_;Zo|o z37ow~`fBtrWkcxVFon648q?1{LgDPe+b_*}P^EfE&6f`>>h_QtraVhYhZf^SgJLH7 zo~(L{P$YWrX{+Jf@*dQarLh0Yj-J$v)e@1UPh_SZQq$CAOJL*#APM)N?&vAwCZtR> zw~%*o6QU-z_mDnLIgcog^x8qoLn9QyYSRA2=2r4LM8UQmY3C))9^BEzB5wR+{{xF^ z)v1rQvc4*7WUO-bNZVf4UzF=PDH4=2vDiY^4JF*2fT3JsKNSt)AZw}h*@%tvlDP*S zZ7#H~^!kKswRG01QR%*FXn>{1R0Ys%>H*33T=3unVgpkyeYm`Tv0+%Z_PcD4Qu?Y_ zsA{SE->Xh+&K@Nrx^FJ)d&f;|3|1{2^C;tX_V_ZK;(S@Z<nW?I^#-+|Xhl%Q8m~C1 z;c`uld(kvrG^ucUvTJEqDDG{IxZ9c*TXg4=ggEuln}cd9YL?tmFRnyJGKb2$8Fv~Y z&}6p=(WcIS-qvt&6JjFW?FYW1QVH?C4^yk1OQ+NnRecPHsOJW9MYS_$8$2kI!C@Yk zT9>!7@r?xVB~LnD)*mVx4&bYvd@`$&JbCW2e%6y~AHNGE?@qqc&!gG5-#PWoX3gfd z*KeULx^3B>yJ<KeqN!F?>XYw-rV^0oY@<ht<I?(M_srR&@Q`~XU~{$JkgFy$a9?@7 zsf`<FG_+k)STyO=vZki#?tr^Z{v+klZ$y*xAJ5u^ZqURZ9J&AxePf$Ws`?ztI+|5` z`Oxa(a30*Odb>~+gAIxai3<cEbX8!y&uNKi95%d_4OZ=@-d!M^^9flXq>-fo0i~pv zssq5(t>b+TOV5g_>&EfUQ@0lKKT&GAr9M%Fx=9es76Q9TZ>I)|3AY!NR5dclT=1{I zq=n+!jdy?+hFqxH)$V_j)m^Dd+)XZ2&u&6oKf7yR2Kfs#5p;h+SVS9**a-Lg+lr~$ z<3fnKhy_z+AhxKOaPN)iIPUOj5jW}jLUaoy4p0z67owHd{JQ0f3Af<szVQvAgqsUl zuFYULK_LrSqCq_4lUZa=V5sl#svO|M|4mk@<|@YRUyXUQxqJI~1pL}}Z(wsSrfQ7u zu9kNt>8)xZhSj>2lz)E;N?9lYizF+I4~bG_yQR~~#{d*}Xmv9;DRzO~mDcQHTTm#V zSvT%(yJXT$l^n-`j>jQs_b;|+7Xs_JH^T6<m~^8bXmKniNCP?U@Q=H1|LU{iKIj}4 z5?x96`@gyXS)gW@$(EoS{Tqs5`%;i%#eV<rez|eAabBo4<WeH~$~B|s+~t~;$xk)> zW2A<aNr>m<5^E||MEe{W_kwboJ1A_P(6B<Vow_yO2lJH>bID(<3BhkRcTg3_s5joY z7g9>B302<(#*_*vTwwJSlMX&n?ZLG^XYx)FxOz5Gb=dg!Q5Bn{cl*?J<9@AaiE2BH z%Ui7)0Q7?w3qvAsdp7SVTrP&c^i}~NwhM{eQdcK*FGxbo#qCXGmEf?X+<o53Zw@wX zP0AI~&YZuHv|0_JZ$<YMSc~j4_b2le=vMie`?)E8YD7nvi1s^N4bE|jUv+IW8{c&B z(w^Y3r5$US;4qhdBQ`!Pj?JZCqT{mEm`g@5k4NF{U2gX5(fwDnudDYrF0C(Tb$|bL z$}3=d?{%f>$?Eul#R1Zrli!nlgee$`4CQpdRti6b{rBHtxCj>0AGrl#Tt4cRVv}#A z-}q1K21ru2tt|OzhlDS&l}k%kwTal?a-P~Bz#(0u{v&a!@t>Q~-!ptR_T7l?S4pX^ zovxkIt`zDEb*~i)V1NFDH9(%Cq2F06vLW-IQ0DSm+^JRayAMYZnZkV5ZQb!%<E>`^ zr~2iH`02Atmjd`Ih5N!1mx9V~NLTUwF#X>4&@PHien~cfjk-BV1a|{>8yCWQsb7t< zhr+EJF8w~bzjoSrdI&SL)S`*ey|wX_OM1wEekz%F6}KGV`v3VkMHSsuJaq{m!}(r* z`u_sQ;5jk>+z>8|+K69l;HGR~G&P+*>yrLHOq1R)b|{}oittL|d6u1E%}AS4)C-JN zSyBgab`tk60@LH%pwK+}+mCZN;QO}34r`1XH(pchXwiyD%bmvcmn)6)k7C88k7Q+N z?qo`vy_+%OaYBk&il8eQc-2qau(r#g7<Uw#O!m;$%sH;-9TQbl??QOQ)Py^GZ@cYy zc&h1-dOl_~-B_uJ(L4^~JbuYd#i*9p2-o9JzC0QE=Q#$D=O1mmau(Y$kR@EdMUq(! zu0{zqiBASfneuAmq)p^TL<&SrgEY3cU1;+-bCURSzb0F;KCzbV{U^#EvZ0@HPSGYz zN2=)2T)`IKCkiuL1c{N(A2yjrwQ&|K_rEaC8ks+qjr}Tpr1o%-&xz=TtnF!mucgLK z&8vQ0`_qs%<qU(MSdT|PiPxIXvLq3nYyX9fjt)1A&v1y1<h`T(h{>zwaCXC|E#hCS z1u*Vn!s&_zac*av4;#LC!J@6Ec;IZ5T-QHMclhzXtW%x;aCq+1RLIF|x`XLV${ZNF zb0lMcEI9RUi==ZL5Pmer4~eeFNHePFu&!pWDfU8KF<Vlqpew77cE#=#siUn%vCeVz zdssc5-#?skld0g!GliK!jc1tT3yydBT6ce+MpwM%hv(5xTt%{oO)p>6DM|k%K7(Rh zYP;;S9aS5?eAq<L%{_R^qB-J;1*eO1zuzfPy!HHDrh*iUz6X}X+2vv7-T<K~D?>hg zPfa5_O+yE;8Yj;%=Lvn8=tI9h7OA_cR_v?=zhHV)B!!pzk0mqpo_R8`38$UY=5Xd_ zj#P-))B@v)oVE3nwz-vUpN-V8#kn*d$2@l%))MDS{|UIqT<SR;U>R1NgNbf`KWi=7 zn&aoR=h^Y9?klITr|@gqwzFU}a*kIXa<}fL{r@o6y+`tk(D!e`ID){`FQp&~0JMq^ zjC|UXyv0{7!Ps1HH9VF2V<;}bgTpFXKGz{O3Y$3XaR1)OOIOY3JFEP(m(S2_&S&Oz zCRj5vwNulo@<hwEk9^o}>AJVs+HB?XT`x3fx>A0`%dr0f0bGS)PO;*Pb!oGn({16I zim@cBIdP%3b!^0i(<N}4i+IhN)n%Od)Je><Tw>%E0p8e7g`#z7aUrA1l3$3mMWg>i z%A{7Et-HgZEl-!N7}@vPEz|N;F<ULQhAnEi1c$gipREJAJEIzt|8wi|{#Oyv(rp5B zs*j^H;FSVz+l!`F&dWh(ZZcQZ{KLUsKRm_XA}ZW3Q!htcfP14P>NeZJGdF{MW-Y2# zpkVs`NzCbiO%m@;+8BWx=O|~ZI_Sl*Divyh1uS`{HHS4Y#qx9*tcP42F!5t&KBhFz zBVgaNa5tu5<D3j~PE>w5la*=1T(@9BN@VzDjEgg>k(kevQkSz|Q>xPA-dX9RZOSHC zE$h7jBiqV!A3d>4V$!tsboGdd#w0JA^0kt$viK<ErbyoRRC^RuoRqoMRy8+eSKHs{ zpN_QT;ASweV~_XR4ip%NUyXf0oOL`U$<)zMtXWPFF0gsh-XFZxw!%2Kdg;SXqSyV7 z!)n;PlLo66JU-^0#rSp*#pOql2A-BIRf)Mku!(+HSeu(R1O`N<RDRS#u4esP`PM7R zXHBuE+dl<a$fF--`je;V@zbblRRi=VS5i(O)WhgE`Q36fQ#UHFJ!yl@v#8x?o1xi+ z<b)aX+{Y<oo2zJ${+GUmMFchx?k!Xl-(tVtE|mz`@u$7XK345x@fl`Yt(OjIJ}(*@ z0ZwR@7%Fi(8+G9RiP~jy8L<)7dm1&%ex%CJZDLFN72eJ}jVjAlnmKo3LYS*w&uy(* zev>VQZ6aK=Y$sl;kovyHu5ZDq;k~<w1VMe}KAz92Ol%0h<X6>AE~}7A6B@!}lEP<| zzLB`BXqVMK=8hk0T#J)7EP3yjg(vzV)$y!u0U8eT>d`r3Eq6-XnWQG?V|Sj6yE^yu z$l+}-*4gHMIr8CCmSQCDjKKGA_j8H!pY~{R5GPz7N!}bPS~*<R?V`zAY|u8MYK_T@ ze9Kax##u|emT|DE33TG8y-imrWqY_pi=?EKNi;5K=?vDYPpR#bJa(L{A{jUn`Nf$i zUz72|B-c$aWmkCpr+?VW9kq9i6@Sq0?|Si2B61A1q35lgqZApn*v8o+C2XgigYCH@ zP?NFJI?P|N*GS%grfT!5g)PYvk|ssd*(QBHfoxW(nLm$#ZM*sk_QGxFL%~$e=6i>E zHE`R*59)az-x$_RdUzE2fJ_D6RKo7WC+_9C%*FYRZJN8V$rrkcO&Zi~Qn5|=hPNy& z6%Rt{!>YGbU^`*HiuP~1W=yrJ<E4VJu7q3{#N@AEUMMeLe&{ez>^CqkYD+$~Q4ZnD z!|!|&w;gVGC*LDZpa2bAm{%<ew|!zMxEqp3`E|^KxuYniUC4Gbm&K`$$$ac->+H6x z+K-|}1q<$nXzt3bF3vlyC!gDP^+XnNwhh3$+xAF&R$|g@Q^;&*kHU)6uZU$A)9@d0 z24TdP5DGi%48r2|Ygf11>X<eJ3trZhIGRPgc6i;eRMZE}T<9>pcnutr+pPr@rGE;2 zKN^<tq6vEF_~@C@<-8z)0$+_CW6*O=0h@Y51qyNo17Zcop7>tT2_*0gKN#uXBOW~} z*s{HC%+>lau~DTW<$=WtWm=@bCepx;(~6`m6~c-Br?VvG^BQ0aY$T)Jisv_Dl1DE; zY6m7Sv>)T`%r;mxXoIe#=}U$ir&Yc|S20skzCo%nw#RR-er{Q~$d$7*5AOffjz`p6 zusW8ij$<N6i-M;DEJ++L(jiv_cId5cJG1l22DqF^7h&#*e%sI`o8~WQ)5M$FJSvPG zisnt-i4ivli#wKFNn-QRwYV2I%3Mo~G!^=pxl%QaD7&EK97b$v=<1)fCU3GOSIg$^ ztXjF2;dTn!p{tUE9qN6_W#xa;ZQ?h!X<mt`V`chRt}ylj*=Kg*aMdYXZ@~Db@88U! z#Poe)KKPc7kb7?hhn4>h^dwCvF_k4Lym54s?Jl)kpV{_s;>*swAB|-IUuX>#V6YM1 zIHbtbac=fC<8gpuZ@Irf!N$a9+Vk>2^^vUuo0_9Xk3k4i%JQ{z>c>n)^<>67w6`5j zk#5h>EK9L5na%qM;ir&R{Buw92L%btEAKd)<B#d^9X*2wR1{~{vZdO)=I4$J+;L;_ zS-o^swmv5-TpO2T;OV7PapLBtL{y2($nxxt19y*N63gmCwkV;0$vqAyEZ&P0yiSZT zuuC$D)3v>zE*kw-Be|E9IEOu5R__A?%sEQW7PTW)9hX$3Lau055Y%rsau&R#*N|5D zUmaa5xk78ES6*?`+alUVf*sc<3Z+#fMucm#mtRyD74{7}5e>sqzkQ8G!7ns4S=`3u z;_`AD+;AgUFRwTw%XY_CsqdCu*AHb8pY=GzT6iXm%U=oNgO-e6<R})Ax1@Yy`nS$) z`4a5cvidE#qbgrX1oaWD)YqeO))w(y6mpqSQmWjQAkNO-|CIDGyG_!YqXIVw)BknL z_V4W*mH*%Z9)_|29TP~8Mj+KT;oY)M)xsr}`kc~_3&6e4Rjz^x28HYmfSJv0Lau8I zkEf+HGl;@1HdHe=8afQ1oL-8vVyg5||F3?Mes%4)mXKsBEdvFKp^mfZQmjRMo3_@r z@Qj0c&g%6B?8_Vlg^7cQx5et(+N;-xn{6qy#sa*}%^#GWOupd>IbVT0tKM;)d6=Fa zQ0=sB!W{O`zjscoo!v)s_B8W;_yO^?GVz3gy`I0i@4j7JfE+I2tKQh;p245R3Q`JK z_sOkg_5ak(X84_O{?gXqnsVP-15=(6|69&_-sA3&!F+bp^DH24XVGPhPaZ1v1dXb2 zu@lE_UL>Aiz`u<D{U75NO0Sb6mMq3rY{^aezm2tXVypkZW?A8&1gtr1(94RC$jK35 zR`aLX-=2|w|J`m}U(j`=CPUdI%gD`=_{`>yNschT^82r`S}Of9W5oUF)`4RpS#9hM znSB5Ik8GL&MzdUOcZI{5!F}H^qRa9zogC=c!Q<#cYd=D>uAud%m$1+h&F<h$=wg+Z zQi6lUS*O*-NxA94jkaa+bYlfw_*@k)=%H7mfFYrZ{`8830%KIwNvU4WH5=Z`3fhmn zFcKTh)2)P&(Vzc81uOxsO8-wg0Ws1#Z`&-06@N%<8QHo#3|j`N`GHkzLL#GG<-t(s zMaXNrj?8Ql1j!jmw~TM4n~ZMSS5ye}Ba^E5*LZAVh0}csjL#J@xy{G!VEAm_V(h33 z_wYij28p2}Kcl3+BQ=Lpm^&6%g#EAumBK@d2izX`6Xz>BzrAkZBF_5L0wSy4&;h0B z#$@fLN!F^B5Sv?#$vTWH#){O!Rk*wpdzo^;%@X$T@xqCsQCUT5#O%Q;eO&zq7~8VC zEym?3w7L2P3ZAd}x&k)yyus}ZPHV@Q+a9vZKBFU5r~fTH`8i|6puy-AQ#X0Az*wGm z3Aep%RJn5Iul0#r(|SyEG<8_YAE%!1Oo$wFhAuPmiS#ac*L6PEa8E~7#J@BP&AY68 zLL*ia+|!0**%)zhCX5J6K{uaFN3ROzz*x=+ce`<isIY9NI;JOReWJy27!6{*R$zc| zg#LPC(aN4QNt#|!XxbQ9<wu1<(Un)jwTa%?F)|E7-go^i?Ow^I(HE_bk1RSYUC<V- zKKbu!7I8kP+*iG6BZ9=1LWlHn%TC%z9iZ?>Ui^9e9YaqvPS2rn0$?gSj9MYilA&mz z@g#O=mC@n_D0%mb*Y=sTb2!9@<e;Qo=ZZ8H+Sbua))W|5i>hDdMsB+C5#ju;;>6@E z=*^5QAB^m&(B<4~JM@%JhQ}it=@7Akg%g0}8vE43&gD!0l=?sJ)F~HB{CPj11a7we zma<{h4%#{yl6xg+@~_^V&ug4Sf2`SHJ734vE>~RMoARx6J=_QFRcp}ETMThf7qzh1 z&8EF@;NPHX^Mb3>PnPZ-OYn}9Xk!Q_+O?%Vd}iu>-Un=taBMlpdx-SXErGHSGQN=# z!Oe#&V<19GL64NSdf;E{E&myk`Pc04G=(VMxbkg$_>VafExxiS>VY0>Tr|l`{ny?f zt~#xs-bD9}f1P{R|Eus~Rn0aD!rIj9AJFv&j%>2sT#*vm`E9%0hN_M?pR$uO<F+(S z!fPJ%TrG{>2f-7rysEP=`o{cwuuN91p+e*IV4?NI`(@LxPK8BGPk{P6Ctum7jlOOV zg#UL3fLz3Rc=oGFDlKsm@OOn59hOSe@+V%Z6hC*(S*dBGteW=%Cb`@@T^R29TI&6U z{h;Wop8Vw5;gHU6>xd@3Wi+?imzKEr{3bXoKC)2-qv@;kZn;8RtpPJUSiw*8RRn-s zmO-e;EI8*?S__~b^~sRVygM|ig^16}h0hi3_~CZ9*2w~}V1|EH@F!XbIETY8)~eZG zRt*=&+=1#hvs9V%f?EU5RoNzjW;8S{)vyXQ0<XAt%cetx)}ioL3;6NFAvE-?RertA z-psZ-lS364tLD}8c;;6#LBm3TVqnc?SQr@MyYi}Yj2-HPX`KE&e=w=90ozk94fHH3 z%`pKN2a|Tq1a<DhFbWT9zdnm8h32|8O$wJ`TN;6sjlQH?@+Hg+;Lbe7e=GW2%^hY3 z@d56eAaG`n*`QlRhXqV8{UigUWPa02h?G?tnv`!y3X5D-`n`mK`K}O~UM`~g5OP@H z<;?KUik-5rTN?7g@G(7WPx+n0A?$%tx|s8Zlx}7%Kw$`f;!hhO7g>SUG8fYLl~){e z(*W7nU=Sp#p?z04-SAw|U;pcWQtD{^|5p0w35|q$0R7o|kkUV5iH;bL84VILh6ah2 zMku2z)7Q-sD5u>fgAv{g;<aj~-R)ram<9NAMR#V~wAZ0=(+B$^?klfu(Qm=6rEKLP za93zhx%!_Xg?d+;3uBVC)w~twW^AA}TAZ(tC`hSwx<AQQ<tLSB0pgg~R?F2+`Nm7E z{Zao6#ziK)4L_83jVV0CjHyG(Kxf`1jGCfjf9r6K0UwqySC3nwjO;JPN2LUma#);| ziX4;2spLd|_%t^n9PXB+>+9+<kP%kzzo=L~)HCV3=-<0Kk>6D|mEujBtiZ3nB)&WX zudDbEZvopFaBHW|o(F;qvEdyCj|;Jp8d_ws-5K4uV-8vKl^~^eP8MJsZkFl7Ls`BM zkOW&A?Ta`oG;lmdQfLec^n>ESyw<X2I~BlUYMq3T=`VxL{{**W+&X!tgXhwL;kV#4 z4Jpl~Ez?i_=SDW=n<A!1ywF;*w(Ez!_UrkpZkDygOCT*UGx#kO{|}nK|BG`I65~W< zXI|8R+w@iK*Pffza_#lhffumsX!#OwUhrWCJ#nWGBUKY-_$mp0F%$6_@~MgkEaD9Y zR_dc8Ml-JN30!%<Nj%Hc21zQ83s0`}CM^a|k{`ioJ-JHdLxIA*6(x&%C$S#0op}eq zLVd1ky+d|PXAwfiOlKZA8?8A3&?LkXFf<Sy*zfe9mlU!<deu8VSCPOmUQh!t{a9mv zE_ia*MgSmwBb`9MwCtL4Qv(Nq3Brj)K^z_FP?jth^B%PZXzVJTPv8nU3pDHx%n$fn z9s8+q`rEC>Q4w0x(}R`}g@8h)mI{Hy_A5xbNxLq+bDCGN)j@KK`?}AT3_|ymNNnGY z%pm91tUpo`%?x^zYMnf-p2Y3!KcBQKCu92l+Ib|yCi|j0>i9<gKmQv1Q&|veKi2k` z9!N@EKvMX14PnuKEp)Dk8!59H1v3!Ctwt4Of6EDIzymyhmi`UAzkpnL4qY~=9pmBt zF79+~VM5=nqW4J0;LR!DI3YZuIFxr)Z;?*+@|C1rQT9cp5N8MKpFzL{yQbmU0r=%< zR$vKqNMe%3rw2Q(xLJaijoX7anqGnkj!2FBnvSQ1LZH+1Qdkg3tVId*3f5I<{ficT ze$pRxDxgS;YQ;e<S+!i)REGhFVxWD|j{n?3&=Mlfw?MEbnhfK?mLcHqwX8E8f8VtU zbzrtT1^ZvCl%+&I5ypzAbM&P=fbP~0i68a<(Eo|zRQW2VEGDB;T`@iUiH>U@jfO+= z{k+j^t|sc{ez*ywbTj`(;V^ub6oc~wkuP8J8D&b_aIUBudPdY@fz-L8QlJ4u(+)bg z34s+n7dl{J%ouP$3q_z6czxdsBhXXOMg7Kh9B>Pg!<2$&f~r3bxCxvqg24qudel2H zI;AGN1d~|wM(f(j{#)X~$QIg*1$xu<;%)GVrN1}BxGGu_p8c62#P(FHz+|=nXH^HP zUe0m2mEDa7K!0`X^+49<o7Bg35;eTmd!MV~vQUeTLR<a)%X`&wF`Z9WEMh}UGY4Q2 zfT`&u4E5SM3(O@W;}ilmNq*^?7c4jiH_!;C_l#s@{ve<M{{zjyzz$Hr^uOnx%BwV! z9&Lx{Wt(-UG3T!rQ@1_NIDdiZsvWdh3(m`wpF!oz=UFwV5R6~~1UqGw)K?x$>qucx z^SlOB2w_qxd^Yjq4C)SH4D!BR<FsI*HEWgqm7v+b{)8`D@y#Htk3y5XO+AB}Ygzw@ zIfwr(jS;^R^ay+%W-25KBx$eIf8LS3z?p8(wrxqwOm7+;fEQ6MI%w?CRx{`MX_#5n zYBG(}z#V-tdi+#;m&lY4?d|@NLwk{<-}Yw>ZrncR|0_8r9N${{v4&VuLz_r2vJ7na zl>I?5%h!tzy6F$DtB;)Nkd1&ENT_c^)spEq`xM%axSn;guZpM7N7NrN$k<T^uj!AA z6^nL%FQh{~ZPhNj*Wswf=2lB@?5tD9Wo7ubMrXk70sUIOyyN?{2CmUxw%-46QYs7y zAiQm_6i1Bws)!k^0GjNJB4*oS1flaC2EtR&z$~lKnytzbK#ZUxMz)LwU7{FwMlX#< z+tgH<rPab&4L)tE<LqMk7Oi!$8b$uo#y3I?zc$}J)NuU3u}2o$oK7(x8y)|p!6n|S z?^$8Gc)Ha;Cn_h0uj?xJ9xb-6p@$E{C)D<M_r8i!=2fn9FcOXJ*BUeVKb?IEJk;y{ z|5u%p(?Xl*klYlK(8g4XbBZjXvW+Y`Dau&0@7$ZCqLMvJIVoh%HkPqes1Rit!dQ}_ zF!pVP`F)=6=+rsqo_p{AKd+kke!tK1d7jVm{=7fWNXuf&o=rs#_3yX1W>O)PUv9$M zNiLTbRak-Dx4ssqet=wls%_q}fxGNFqyCK51pPa#fX1y5b=m4#TrShTfmdujR`8Bn z;Rb*EdogS8IbqDMP|*89*BPSPIQBH$dG=R#z7j5zqtuu`+{r3@k7eYROE*zITF(w7 zlbIW^0+4ZfBtHNBptkjdHr5>lW1Sx~Z6TMsr&wQEgB@w)E_*8#7G2&<Je4P@iTU=o zz2|HDRm<Vs+?-B(kA8d4cbIDi$Tr?7b1gj+^~sgI&EXr!7+rm@;NFepm=p+f3^84X zN#)}ga|6kp-+gm~uYv|rR0BQ@YL~3UjBKu%eB4nvImKS!+rF|UsRBW)TP-GzbjiD% zv`P>gRfp?Wtt(;*^*MYMVBuVJ$%%Xg8rOb;{^SZQ)^egz0ZxiN9Gm)*u&0Hx3f4il z+vX>O%G|!gZnk+Q6R1g>8S<dC+DLb8OXn;Mi){BymTEpdEs8xQKnM86Y{^CMOK4}o zAQlq|{;}3iq5?Nzv!Svkq@A`=WicrNoi+E&u<jV^(x%hRMj6kzM(B0>(88o;!t$)d zT_QL@Ri4QYXd2X|CNb1?IB~##sW}ktEq3W~_^w75WNvJ89#bo+Rcm<na*Pj#wvsu~ zrln=tr+=9#e~sb&w5(TQMx=>^2vXxX73`%IBb}!c6OnleR1VWbrwCxamte4rE}=M5 zu7<rExJffRS(vj1GqRdUXdW`<z@*4b`{;B_P>bhE>_~fG4d^S_ynXBImU!#Z@36W& z<~V4pX{%r!I0@P|c|pt|MMI`yrnY6zeq3(*;>PsW%4JB~X6-%=?5?il+DL0w@)v~_ zH1*YxSvpCF60EUqqZsM-#m4E6k82~br%W_u5_mUycnN!&V_j<XInLsXQ62ExJY{8k zHTleJhEFq+kj(%ehQ((eJsfR5h!&5;BSWJdtjFv!Y9p5>-N=^s@F6Y<xq%58%?DyQ zuuJWpt!O^hGOP|}7hDXIW$Ai2I&Q=ccJ<Xb&nmRdo3z7Fkaog)F}d`CeC6dC4&O`d z^Np~d<8`v{URSO{ii;yJygATob+y(g4SU)u%W6UUgHd&4W*IEs76Vc9U2QNGiw8T{ z-p9hXQML@@gXrndt0y1ifS9bzDoPI-XWvPY;Z7@r6C$TF%ck%h2Y5q|zE`w5=390- z@WPbEzSi^%|0wHsJBE65<xoUdX5k^5i{kSV^rf68aT`MPqq*-}1YTeqOFm@f<plu^ z*0K$}(D5S~^F3p@Wl&R=n686$A2Zxy`O#__I9A{VF8YY^5aw%}SDEE+FQFkyp;Gtc zaK-rDi*<%jqLU9cpB~2JdJ*DWCX7qqb<3U{Ih#v8AR2aby!DCdyDAHpRhZ{PygkP6 z8I`}71M}(Rec8q<cDTUcYJOg9Vua<BWtis&L89fTvsi`P@S+DebNX0U@u8u2dd?qQ zjMAUA&*}vv^ckDH94RV_>99*Z1y3}uk7!IKM4GrWE6Yk28)~G>#>c0Ki_?`O3mvIl z05+gIO<G+A2Nnr>5EB<gcmlh+_vmge@;bDBz?syfw!V{Cw>_xfcc9q}J5EeCBWkYz zaDrKKHr=BkoMK<*LKqzE;(cjgg<++w=yQ)r`D?Z&kO6qVXeh_}N?N7pdAqz7w#cck zq{(pdxj1tYxv+w_yf5d10m_Z60xvAQvY#6e$Lz8sZ%=Kn0pP>B3t?^KgJ<$}weB~g z5R=B?RS`EPMdtMo($xxVUKfmI7<+lAEdt@T^z-H;pi%<X-HH#dKizo+7ah*NWh=kG zt*4;&RC`YW{15EdVzH#=q|z5R>QLSLD>pV<Lq1r~ClhA4<Xr5hu8F887hN+z+!1^F z7~Oy(k8z?MY;-bIYVX+$34~W{fUH)%boUPK)Ks^k&59M#O`+@Mz>Oda#Kuc%HYHVE z%rD|MC!Y^}652S-gOA3uNdx&Hg4)21Y+s#%+_D)%@^e=w^gWq^Tu`|evwI+^dB9PY z148mj{DU0*;~mS9_uufZGv#s`Nd^CeUDv-ooAzMX%M0OT;Kr}+eB<A{bNyF$);#cM zcUJs^JEQh5?reLD)mMA|B{^H5>;s9b^F_YaZkds}!`>~_jv;?TuT`BdTxbh3iE|Bs zEhY<#5}C!b<Efd&^DbQ;%(*E(6VLA<e?v6fmDq&bO6>5E4TO1nkD?eJNhe`{p**1| zPv%wL(9}|2a@Sat7pah`MId)cYC>8?Sr|a@W?PS9V^emj_;Sc(J&H-K0iCfZso|FZ z^?2NGl$6D=E-;;>?A^@s<}P@g5MgAX#&b<g3b-KnP`zpBp5Gm+KKg2)9OY{23hYP= zuUH%JOJwOresgfcm(Z1bP%ln2v}1sDx7A<c826Qaq2r838!u^oY%~&f0M3uIf)0N3 z>D#F?uFJ7SWKaYPU;GR}taqKu;@&2CM)OQOpac~c!(VNPIEb_5-B!7b=58*`mAo_% z0fPWGnvP3vJ}qmt!T!1&Ycna%dC;RC^OazWy^2e#%d1q4i6d94w)-pt4Sgt|eLgkq zni0yg5#sKs{7q>i+n|2J_HS;iOK)ON6Q*bPl-iXXy^pev!VQRG#+WdAu*h`zGD&A& z(t60(F9Xq(v;Buv_agIsQ1X25?lC_RgV@<t>7Las?H}UunfsOj8hvTQGbyw;d~MJE zvKH`PS(AaH!VuHzav-*=Ir0F-!$nK|!E)?GSM$(t#}&Da5K(TFwD#^4Cl}a;!#rV7 zwt!!23(M=e)X1%t+HjuZv>#a|60)T+<&Jy2K2=7Ru@)lFc{AyJCSFWi)-aXGnywsd z;qcwh7%k(#wg=veDADim!n$c19je&Vj|}_I<;DcRB3uq_(XNH-!SI3@*4@x<_OP#8 z0<&W>vd=h-P63eg=Ua&>0%eqJUwNkHT-vZaHhRgByfJ=`$I{t&ROZnS*!(rr<`p1A zk(%Y#YRcwIrL4t{<eSNDOYtfg^glg=`6_f}-&;qOr#T&e>|ro%>I6yF*1A>mx`A_k zHdlr}Pu@`F!0>YHjTL%E;x1i=`H*2eDf!TBikY;Yix9TC^%|}N2$px(C^kER^9*^{ z*0r<?4Hs=|I*N6lrI_SQIN__`vx+_qeSUCSixu!xpQwFCge;(*)wBfyvMzGK9=Jsq z36an=npN4n3}%<OFm8hRo+K<SgHJ6FSSjFka_y<5N0IoWWMJ%=Z;rKhn>TZCy2b|> zRyYY!YHKvjlZL(|Mr7FvtS;ANksmA2Nr8VEe<}~3J20V5r+t2Mkj0I~T2C})vOGT1 zGl3MM#}L_Rxy<Z^@fL7Ppq{OP)DBLpE<8FBdkXMBx(T>9tlc>O8TAg-02#Gwoy5Q{ zZSy)V8*nOA6pH!o?}EMA`r%wWgWklLTZP3I0&ilDU%<M7uX4aoNyK*Aa;}B0KC|Cp z2Pww8Fgx$)#?Opy3_EBDARa48*$P)}o@X&%2cff?z;uGzJby$^bg2x3T~z}1GTTEf z0Fb5NhIITchV|<if8y}1NzItdcjptvMr-q8&3TO{$u{zIZp_uD+@E169&tJa(Wnm| zeDQ4yq%1%P+KiMUqt<Cad-nq!VBDLToxeUleq9KFYqDhZnbg%+>uBo^W8H?OF55^> zp~+?))qV0i5Btq>V51scly!N$Fu3Rf_<-lgv0?#Hy_6XB)z~QQwvxnyk2Y(BhWRl& zvzPJVr}4%0yA$2H#=ItEUa!G6Pi0IF_Gu3T%lc^isiiiFW}&X@vfE%Kf@+T`ab6(3 zBTjO?T40-!mdz_5BzmtpO^CZ-PviIqdl|-`@(ONza(?na0)&H9Ny|%A&pb%2sr>jV z0$=dt?T@PsU+04P2G+HNb%CTn`ASmOGK}aMKs7aXO5ns!Xh%vm>k{3;E4emGP&L5L zQHaIWm0j2m%)Vyi-D`gqfB72hV17g+@X4}>#>uXJ;I_B8Qo#c7&XNZQvDROhIfU8S z>_lc$@Vs~k(5Xvf>{<js#@@u)<25U>%_T8C9b3$wabZt)Cb|-?owZwz+!WY46m*G* z#NU8&V95LKkARQ(Z3Iah`y~Ng0y&97IQHhim(m1k55(S+Wg&n{pzhc<4EARLkYE^} zOLigYN9-U$UmP~W`T-nI0XZcEgC{4Z&(!?P^jVJ86=M){^B-)ta$ph!C@rvG=;rM{ zt56`~e7zIH^g1tph9C<kv@OC)=(or$SX&<pv;!q3ObMNHnV*DwCVs&5-4>ki#rn*{ zZ>3?c8lpB4F>}2(#Nofg^erY%wWFC5Y89+;8Nm6$rL7msU^44P;BQf~EEc+kLO3nA z@(Ri$@QLS>`Bs<7_UMihms&cIf-ODV*yyd1kNGkb4*}VghbjainT~u6gbs4ziElZu zha;DrXV`R6FT;luD{#$WDAY?5*DPJ(bzN!ebR4OzQwnOr8r#e2Y<tnUZiz3xwI{9h zp84DbWdjrf#<E!VDNYF`%aQoRwCVR5;2Hfr0@6@!?5(66di^hSZyMm@aQAycuCnsj ztw?~`TjHIy=Xq&ww+dGB_*@2uc=Lj1q;Jn^o6D#)CQ5^cBuZTa*pH1d9|R<$n}c2> zV-k+ec8vI}#*B7?Q5qYz&JWL&A*c+pi2zO^ZC$#X@odDKWe7!t1u6p#+OBpjEaoGC zuz`qCmzzh5hEog3Ay4VDtne3X=|qspuRZ^O-n29&s_8%@R_)92vK(s2L<Ylj6T}|! z$tdteFa%?48)i5C>mXxn3wH1`o_Y~4mDd$`u6Vo-0^ml>fRH?3uR2YKHHzCQEw|if zF$hfQJPU5CB{v_%&Ue2g2rdpOV^6DHkh68Cc@sH&bLQwT0US%13<Bo6X+a9Jdlnut zPlVdK-$=a5XG&+mrtTSlPrQT(GxI)F#MsmC;l#xxZW;EbP%DHr_lFMPTrhEQSFIZ) zn{F3}PD|vy4jmO5fVl9EH`Czr*>cFf=#v9rw<nR&{Trg`b>SX1z+q{(yb1k~#2RcS zDThY#kj%oM=eekd)WZ9Xk{($_pJ5k5@=Pdx69u;}!!EVGR{$Ca@0RR*p$C@`cVGq< zjhh?P_#>PqAw*k%BalQu?3r_30ZJy2KiL;h68G0$hTJ^iwx3H7dlIr5)|DgZU79St z4C^PfzweKkgY;s)acP<oZp>)7WR54$8#anh?Da*0*Z0^#2T0LV=kkwB>4Jz8B`u6Q zC#m3(XK*l`It=hK14qgnnEt1unB6wp#0MTI*UV3VWy``%^Ypu9ux=y%qJGdCB!VMr zu~>$!55UMR=e5|u7=La|UmOMyPzrdWwHkZUgB1Ep0{JQ@=p|S12mptXoGZY~tg{`w zH{tK<NZU_(+R%FlK0qY@D%inljL{8dAme>w8I0oqG~OJT-Z(6$RUVSZi{;oQNRQ^L zFbp$tZ=sGUBTyYPCcZRlfhIz`C(5lI#qe%KgOj_ySvY?<fpzzo$4C=i0(?R`8FyYY zB}g2YRfh9EsgjLqg0|%HHw5s%7!&VeG<tqm6ti1g8+SARBm^7ledv`Nf|#9iPVFh^ zbuz}<yXa6I`6wJM5ZdR7aK=|X!s-_VR{S-?+6-x@@k(H$Le2KkXC3R?i<n`QD_>Fy ziyYDTIzF(o8Is(sBq_r*$kEu-{BOEuO~}ym<*;ralNt;2dL>(w7+;`dl~L5N9P2W2 zN+_dJ*1crs!5+B>O=CgcYq0rGDH03moxdI&F0r|0v<n**V<-Qff*Hu?*wab{&wcm% zCcTk0VmW+A(jHg8XftHxmu@VsG@Wz4l;zL(h3TKxan?iGgum>u_+WPmJvfD?)>yK| zd2SbRL8frAXNuTc_UI#NB)(Y2+hI4Wa=3NRE2_6G$t~Ky<VGrVVJ;R3P{uwxzwv+1 zQr8Z=(EpMmRJ>P9FuB@5bAO<eI|VJ=Cdr{;`c$uw*8PH?6X9@5wS8YLzR@0FB7Uey zP2b~JBu4ZZ9mpF~aeB0CFrSslnwy?u$mMsHP#(B+(h590ypDXbdA0UjxK9XM@uEOV z)b6S~r?#&GHfx{flE4R)q)Z}|M_XgBYw~&1grcfkqfnoaTkZM1alN&IOtG^b7LN{V zh07)9JW%PB)b*;Ke*W5i=lS?XoGVQqt6KbZo7F9c!o6CX)`x#JuckZE=K85D1Naq8 zY<5v5!DbF~BLT>Cl#u|`3hLjkyZ+2sCww8cubRvPov7i9%J1c*7xd5Z#X)otjX(MD zc<;8S#*T=@oGJ|6R<SmqX~zTJDD%PBhT`HT0?7f_s}=l=QZ543X`l!%5fnFLXP`!A z)bMi|&7HlOCi8=~W#idAfij8$8MkzlQ^qxS>Sk-?Yzm}O@APn;H7qM}{@jtjgI<@N zyRb32q-uac=nBb|%p?4GEO|pfQ)h2SJImGMLk5!a0c_4V9(?2T(}3B|Uw1l}8SLJs zRM-ZH6=A7G>YQ>SP|m*#{WxIbT<(gL$39LH29-5;Y)%e%b@Bpi82P7Jbo%N}-B7ZR z@Ql@FM>+Zpr+T$uI3-IwRJ*(AqXsa<Ezo@`3#2@z{`&JJn;-6>)k&=hA)EF2zr8VH ze0@W}YudV^;FRm0w!0DKw9U|tGhLZrkOuDah|r?CQ_Jd|9{UMs+RY7h8d%rE3C$M= zvMY-Q)TAv8-5u_&+2(!vk?5GX9x}kW2C{mNQ(t;mm~x8n_(GNrA<Ce+E`>=#;>St? z<+csu#sz0|NyD$ekoHcNr+}`m#Odm>4vMVSreNX=<ga>Yc)Hc_)p)>1m3D>oVn7@t z(nGD}>6+y1La=A&#UI9&x7Jw*Kj0Oobb?ufbd-9jSxtMgv^p*t97ue8k7<?RSFqih zrAWG~pKP;FGCQ>K=#}FEa#|i6!G-B0{=*WwD~JPmpow{gKC-=%%Y|9HmPd>@m4Nd5 zO`VU5APx^W>J7+L0aiS6QTpN^M#0H7yn`ybii<W#IVl~PymA(w#rD<8=Dqv1V7q2~ z2fonZU)ushqXdqL!5*$dABu1uVXS2Jl>^a_Y+Ux7Wy25@67jS_$@a@(iM%4}iQ+`J z8eM{Si*Bprw&7RQ=^T?{QK(KbRa>Vrs}SCrgx1EyWLCkU@6AkA;3tK$YTAJtL>t3W zi|Hu>WrG8*l1a&Q{<K#%E3YVLkeOOJ)@w6W;Yi(_uf}qNZRzB_*=U*-IwW_q{)0%d zuw4|93c<oT&7D&+%=#b@<UGslv)rdoOJ2x36KxQ-Co0j1E=lYF2{tW6_5y3A;pd(; zWXmYbXmtnoGQlIR8i|KWy12PBZZ$Lk>>Qx)vMn5VkuTO*B!*Al#v4~EYCV?RwobI= zgazr$^NZ2;;Z6cZT;jJ`@`{l&FDwsfXs?uc^wFPP?3M(xcVa_ke(JLfvG#D)sP62Q zNp7tikDldmITu<h%cA!EzCN>#6~yBKA6#6nJ)<9GFR_dDrq=RdaLRvw4XEhKzHYe_ zU+j3ileSAn&LqJlAyRd~YIE{`ADeRfL`wF6nyr$N>=nZV^Zrz)o`BrA1=Ax2P6?3$ z+2@^dA#O5Her}aMv0c~AA+>49GrTkW9NFbz<~F5P&9kxU4k)X|AhdYH*UByHXm&(v z6M{2(O;CdWjqm|cZ76S}imXt=s)-Z|#PgZO+MtRgeN<r0)gK5ik}Du>5AF>2?(nR> zN4o=67qmZG7~*T#y0t1rq9pg4GOSUW@jf`Ms3AyEhpxx(fBINJNV=bv@uc(dA*gOt z+R__440fm|mfuH54|mojQSxpo3ztZg_|KLy`OQTMj|7KnPI=0?{v%KD$Q=v#%n&@X z4lMqNe06yg{Jp~H?oQgiK6rXm#ho~Z%1K0Dk{R9q1criI|C!57o@7_16~IuS<tJc> zKN-m#+<fCYED?3Ln9v5=@>A)&J_#gPd+*8<-0~_?s0PT-9a)8N3%P@Lq!oBuxVNgr z#06M#a&BL*`-xrRs^&(SRq44n?AG!K6%5on7Hpd3f4SUsNBpi-i;n?*I|nHV)e&y_ zhU&6<T}^{sK1vPb3LdH`_~$lFm(*M)OTuTN$HeDzc;w69Rrx~42^HAA{<-B|g#!uI zI<ZK(bC#f1)e*HaDE~E!b)o`7p%^GW$eL%uA+gG{A1qULivpprkHEBjt{Q|At@ zd*QoUJ*zt=loHt-<wHFx<G~s=T!8RG4E@{<)nKS%DQWPh|KYV=w{?%n@c{b;|1|o6 zoR9GY1XUM94h#HGqaU7;jk6cC#B+x=+h@2^?KAVOwy%#fO?GQ(2{RjihIm%+D>wxZ zxN&`!C$(~yurT2wHQalLuvU468XQ<$r`-XZx~ZtRw5Ko4J)Rm&hNhyT%`4zgDvC>f z>Rz}FpkFwCSLVqj)uE$;rJQfJlwk0~1QaOR*~~Dxl}Z6z?d;<d-RERKgD*3!RPwBD z^%R`0NUCzoHy~`ae)q&nqH=>8#LSX9@{#)?=l7@9`*XF!kv1H`=w=toZ&-rgZ&~2U z2w4Te8`qay^MtUlFtOz;)yzWdBP8>dqM~9(2Ci#-@(*mXB4PsCxo%|q4YypP6;&;N znF)_vc)nftzrU!#DOPs@BAA`XK@j<3QviYF+d6a`+kJjoB0S<UW`keY#$tK<pnwZ~ z;6QtmsTh1`n|Vta;gV0P%eHE2@sEuAx#e#}pdtx%o-ORP*AVS9WfL*f<Oc~V>XDX6 zOuqiLqgNV}E9(v;cRG<G93Mhxd<)FLgPnn`<jg{0KqsU?opJMAj}Isx;%*D&RmVFv zAh`#pLgF(*Y)h?E{Wu{5CyP35p5ku%+e((0yRfjmQ;2DbcpIeQ|Na`d+*CLKzPG41 zax6>$5wfoYUk;$|VtERlOLsJQ+&aH5Sm4XXgS?OM-<ZS;HX9W=T4Y>r+-<O<qT-{Y zM?$y+$hHak$C7zddxNi|7edLuRudr!wWq!uvnhjb2n(}86;%g2i8-KFy0Sv4EdPzi z)du9`N=>8}t<Zv<&I)1tJ+~3-D=Dly+}aje9OZ9wWWy&^UB2S7xb}G*sq%hr?E%M- znp1zBdBU9kFq}4=lxo3OtW!CBP3t&qJvhSGzF%D%N!`94y^uO2lcJRPO45M@3!E}0 zBVsIKdx8Q#Ua{+I%>Qn#su+up$W_@JV^<c2ggE;zh8sr<nMh#c%pZIo<|J%DG7;v( zDWvU`lwx5t;b8~vWU#RYZiPsrj2n*w5K|hgn31ZeSNZR*uM0(rK%b?`nRBJv1mLSE z_{#JjS2aX4ZtDCGW0kk-GM}ZQis)Xu%W$!{bjKY^iK^52J3yArffXZzJb3k*>w!R! z4t>CyIBlacLQT@YFP$xDpWHKEH(X&Fe!5~<&g9zBWHfFaDEwv12}GEk&c8*hj-3dz zd&e8uZUdg9;^Df3T_469)v`B@@PhqnE8o6E;zc5ZfrlFyZ*~MqUPbL3bbggkooID# zBd)`igc<(wJ>y@GhVeRq_tSzotX`uvtk!b8Xwo~UXwS_PDdSg(35>W`a`_`PH|>lr zv_;#D&~LY>l6BtDD?lWSlyzJA4{YE5c-#BZ`5h;v5a#{<`j0#KYrMcV+F}NqlJM{` zVxSs1g-1FDT(#Fp+`AW;{B+-8D01LTd>7FpQ@&C-MyuGWx=+T<I6&x@4xKYotz9A5 z#ipxpBGGfipT^K|cleqdP2uZ~2P%)15?c>X=&4n|;ae&;ypJI5TwAAbN1%awhAswp zauj**j>x63R{W}KI;r!f^QpG?{WyZj%^TSl%HqEI{}~`l#gKg-s{gaaWA^#Y{?qDQ zXp7koV^_hcy+>1^;Bo7pz9(l}T|?-#Vb>l#*{bSi<R)5;fRsb&)cmN?>Ykf~2z?9~ z=dnrArJ7NZaw%Z|e8HPYvTcYNe8gsbE0x;jwZRAX#+i!AzN;|xD!qwYQW7-p_!8$N zOjWc=6rfdfpk7O_{~V2fbj)w^mW+BDRBs5f;m3Qzz82^<M(lLo3kPX<94sw>gyQC+ zwyc>psZccYq~!k-)%Ilh7axrp(#zijU!6Ugj>y)i{-2QHfL6Fr@h&LthBG?7Z5O5m z-9Gu#N+KV<dy+1XGXPZiOV~>{og>+5(JH=h<lvM^8}ki^17%zZh|xfG=A8bDx~_|0 zQog(CsA_uyr6CgrQe?Mo`os4h_7`^Go>mC-6Am1Ezoi!tKPY3IFP~25Ik^+442^s1 zP2_>CPtJeV!$VKa1{f+b$Jca4!TAJ$1v;t&w)^k+$ls0nn<2QA^}*5#5&y<$&Z9rU zcWB?pUrN9**E1Ex-kKn@f0HcouX4iIeE%*P{<~3fCfeSVP}n_i@}-k0JbyTsCeAyH z6$);-84z=KI%Cp+R^cf3?(HV=QqU%l{eMnQKLADs{C(+Xu?h`x0+li%N^bCUac?yW zO_ibOKgwU500)?MjJC^<b}6{(wk_VE8J8?{!s!)oU)q2%4I+zV+XF|0saBi<htZLP z-lJ*Jr31boWpsX!mqRKw5F?xrX4si{hlnUh6TZoipFgSmnGGN?6p*#*Qp@twlGLku z6NKZ+H=%-NSBn7FkCGy*>>mYZZ|zongUweBZ`BpDqOBG%e_t6$sP^-AxZvTw?s9$l z`nrG4sb(vB>|LX(Qz52wvyR!bQ#kBw{LY;M+jR`<OECh&^6m5dZi}w+!r%}6doP{K zA8=SH5tp7U6Bf#SckBHp;%=FJuLt$Kq|VcGWb(+)GV#C6&2N+y8KK6yb)>uzN%JHp zHaBOS2n)?9Jm2*EHw~{w-+?m=E91rEJA2ZW$BXmAe?JhtxX|_FvCZid#O<j9VMNu> zr{NzjD*W8kPR|}xnIDx*Jf+ZM?=@!bF}8~38jyawWTGmgKHhXUf0Jc;;cr9sLUt;W z&mX}2@f>E6pY0lg_5B>35|q-PU9u!uuN)ap(aU{&ePCI2-K3q?-lBN4p4>GxA9IW? zpLItErPe7&-mK+%$XhBgV6*G<MZ!Ub@29bi@?rNrq3L|Jt<*YG4^0&5#+{tKzBZwU zM_SwZU05;Eq0cE*K5{f`dEzy(G~-+S^&;i*L+WeqZjC$RT!?=RWTUp-)lxI>zf1Z$ z<jIilr}ryu^U4mt(bKjjap7y`XoO!k_D8jyM$^c%{4}aOWWLqB@o>2+r~2Ac-^|Ew ze8$p}xE*%r`)T6ahVj;D149ZwT9MejAgLc?1WrhL5}P`7rmJk$9mQ8hjr+%aP!uB% z?BQ;G{eI!vgO!Gi%(@(^PK-f$#&YZ*?tSvzYvhYA-W*uwMUEzETk6k#49O~Nau}K8 zS+|f4JxbM`m&ZN>J$hveD0XpR>sm5iv;^!wK)7kJIXBssf89|Bq0<X1d)&Yq&RKHq zD3@kkm2a(G3F@iHOC8HItq;fvP%05xXCg{D8-L2#A$#a>$PU%N`BQbf!f;B`tlU=f zO4N$<TCVN>R)uCK<?}w@{@Wp{;7LZswv|l(2sGk?(!RS}b7jvZXbIS_G|&Uhw3_kf z$`WpFN^*@Nl_&O=MDNcuoGSbd)QlcX3b9(5A47~2f3|ce*B`s4unfNBnSUSj&A`Ba zeH!~QLF$VCS97J}vbL{Ee0d2Jwlw$*sU>*!y`g%XcLjZ}r*^t$)8Z$dNQ<{^v^{z| zu9kOBi9c12T$8bLx4_l8_zY3X3*D;Y%EB{0Tb9QcHc8j-Q(tQ!z<b76Kf~^dN91${ z(Q;vdFj_A{QNP;U=bt64TX5UY?7}{iChEnO$L=0UeVbM{UTyYI)qzcn`jrfh<UZe2 zIms;B%iEFUNsLUrJ)|<9YP#Dk^YWR!eHR&bQcN=w60MR7s=u!re*Zi!mT0-zc$bPl z?<TN+m*WltonHCNKc9&^c`*4hzGl<8`TYhupv>PMzY&Sm$vtq{+GaD+_$0oXhU~qM zr3Xmama)!HCudg*A^*&-@%kxQNbVPR*XM0i*R96;cH>?<_9KclT!nbs?<Vlm#pha& zH&49z#_37;0#MBTMI3U)^nEF-i<>jH|D6}9h773`>vn0$dU^!-oN0{SHN3l{Gua|t z@b3p){EO+BJho|N%7FoWT0u&~SDVR-pqif?H?@!qw2dkJ?{}RbHnp`;3kFm?FU>r$ zDppI7<(Vq<3k|5>guwEc?TC>{ef)ETQw)+~=`qr#3-!O9v#h1Z-1f`C0B7^^Qr9P& z?mk3w<8pNg`beSyrg~mIt}g%FG4p$+cH^u7OPTcK4C7m=!v-3bDaARC9!Y<~FD(U7 z{-_^>*^#1)`Uuq;KkSFdM;~g&c-uEF%U&1`WB@%&lXv7rH|w}v`ZTRM_5jo^1n>4j z^1|*Ffrhb8KV@=_nPm%n8~&#kVERkol^1mQKC*G-IRe>)`yMSxz8@n6@2sh(7673B zMt^}sIAho^S$|9={@^G2e|k=B_${ySykaX|ezzQkJ^98JzTltK%=uOh%(wEZ_k>SK zMs{(A0cHYDVS#$Fy1NG!d<R?#*Ph&Ma78cc@yzm`nu~;!6@hrR`5VLmmyp|FOU)1C zDmz-~=2-RHNQ{9->wki%jO}CP`FG_PuGdUe#xN=+=KFYJ;vX8QICF6}F6VBWDXunq z>l-g0iFtK39a9QiC1jkftK741SMi{S<YSi)@5ItCpQ#t4Jkz%?{a21J?<Pqu1HkT^ zBK@hl{SQx`fKkgdAyDT1)}p}plZ}msmjmX672$~{Fnj;eZl0Ks8B$%b4Yhwmgy5c| zN8YDT6Z>09DH`d>py<hY?eTN9Q7!3nM$bu(O6lpww|2@kE>C>J`t_@Kpv;dv5=6fJ zQ^GR9;d3X2#tFY1<;7S1e~I2p?jKRsvm|~1^*zC-h86z4P*2CI$Dq@nCQ(a@a@^30 z^qw{Kjh9ldZ>2sw<$P674yp(yCwhxvujkFflcI1qAWd$gL}WwB_-bsBQtCoK2i2|j zl-EAuuo6^#BeEhvg^3Z_@U`Z2-BfZ%AoSxabErlA`5=$O1mlsJ#!%@7+h~tJP7p~) zPbytno}nJZm=OP_hI445n!xi`Gtg#VEPjZ_M1o&-btROFppkrh`7Rl!7)WyI+`Ty- z8U-H`IANp%kIof0TefJ)!phyUsHx|?QekYvLMac!X(SB^F|nrhpD(r?HqJ!T&l~%s zj|8AurAu9Oamo3eXE(t;O2~S{o~WtJo6%-`Wx?%rRVq}<#jfE);T|k=9yG{A#AAN| z%sO_m#b6wcMukmE?MJ$9!o@gM&$u>X$N=fb>1@M0G<1G|Q6e*gFWfUGgHsH;`D}wX zVRV)=-t$ofI^Tvxl-}=myWh`@e29v<L%d)Fcl?Qh5GvMRWx?N;i6d_&j<81R(dok$ z(DVFXx^&f{KJzty?qYHqoqqM|!{0#=<944l^1I|kZ(NYcqL5tY{U^^(ufYBfj7xyl diff --git a/verisimdb/verification/proofs/idris2/ConnectorSafety.idr b/verisimdb/verification/proofs/idris2/ConnectorSafety.idr deleted file mode 100644 index 332041b0..00000000 --- a/verisimdb/verification/proofs/idris2/ConnectorSafety.idr +++ /dev/null @@ -1,200 +0,0 @@ --- SPDX-License-Identifier: MPL-2.0 --- Copyright (c) 2026 Jonathan D.A. Jewell (hyperpolymath) <j.d.a.jewell@open.ac.uk> --- --- ConnectorSafety.idr - V11: Connector type safety (eliminate unchecked --- JSON casts). --- --- V11 in standards/docs/proofs/spec-templates/T1-critical/verisimdb.md. --- --- Corresponds to: connectors/clients/*.res (ReScript SDKs that previously --- used Obj.magic to cast untyped Js.Json.t into typed values). --- --- Claim: every json -> typed conversion goes through a total validator --- that returns `Either ValidationError (ValidatedValue s)`. There is no --- public constructor for `ValidatedValue s` that bypasses the validator, --- so any code path producing a `ValidatedValue s` has a proof-of-shape --- at the type level. --- --- The proof consists of three parts: --- (1) `validate` is total (by Idris2 totality check + %default total). --- (2) Soundness: if `validate s j = Right v`, then the wrapped payload --- has the type `SchemaType s` (by construction -- the only way to --- build `MkValidated` is with a value of the right type). --- (3) Schema-injectivity: `validate` never produces a `ValidatedValue s'` --- with `s' /= s` (by the type of `validate` itself). --- --- Idempotence / round-trip with encoders is deferred to a separate proof; --- this file covers the "cannot lie about shape" guarantee, which is the --- Obj.magic-elimination claim V11 is really asking for. - -module ConnectorSafety - -import Data.List -import Data.Maybe - -%default total - ------------------------------------------------------------------------- --- JSON values (minimal untyped representation) ------------------------------------------------------------------------- - -public export -data JsonValue : Type where - JNum : Double -> JsonValue - JStr : String -> JsonValue - JBool : Bool -> JsonValue - JNull : JsonValue - JArr : List JsonValue -> JsonValue - ------------------------------------------------------------------------- --- Schema descriptions ------------------------------------------------------------------------- - -||| A Schema is a self-contained description of the expected JSON shape. -||| We cover the cases that occur in the verisimdb connector clients; -||| nested objects are represented as JNull placeholders at this proof -||| level because their proof would need heterogeneous records, which is -||| separate work (V11 bullets do not require it). -public export -data Schema : Type where - SNum : Schema - SStr : Schema - SBool : Schema - SArr : Schema -> Schema - SOpt : Schema -> Schema - -||| The typed value that a schema describes. -public export -SchemaType : Schema -> Type -SchemaType SNum = Double -SchemaType SStr = String -SchemaType SBool = Bool -SchemaType (SArr s) = List (SchemaType s) -SchemaType (SOpt s) = Maybe (SchemaType s) - ------------------------------------------------------------------------- --- Validation errors ------------------------------------------------------------------------- - -public export -data ValidationError : Type where - TypeMismatch : (expected : Schema) -> ValidationError - ArrayElementError : (idx : Nat) -> ValidationError -> ValidationError - ------------------------------------------------------------------------- --- ValidatedValue: the proof-carrying wrapper ------------------------------------------------------------------------- - -||| A `ValidatedValue s` is a value of type `SchemaType s`. The only -||| public way to produce one is via `validate`; external code cannot -||| construct `MkValidated` with a value of the wrong type because -||| Idris2's type system will reject it at the call site. -||| -||| This is the structural Obj.magic elimination: it is not just a -||| convention, it is impossible to pass an unvalidated `JsonValue` -||| through the `ValidatedValue` API without either calling `validate` -||| or proving conformance directly. -public export -data ValidatedValue : (s : Schema) -> Type where - MkValidated : (s : Schema) -> (v : SchemaType s) -> ValidatedValue s - -||| Extract the validated payload. The type of the output is determined -||| by the schema the wrapper was indexed by, so consumers statically know -||| what type to bind. -public export -unwrap : {s : Schema} -> ValidatedValue s -> SchemaType s -unwrap (MkValidated _ v) = v - ------------------------------------------------------------------------- --- The validator ------------------------------------------------------------------------- - -mutual - ||| Validate a JsonValue against a Schema, returning either an error or - ||| a proof-carrying ValidatedValue. Total by construction. - ||| JNull under SOpt yields Nothing; JNull under any other schema is an - ||| error via the catch-all. - public export - validate : (s : Schema) -> JsonValue -> Either ValidationError (ValidatedValue s) - validate SNum (JNum x) = Right (MkValidated SNum x) - validate SStr (JStr x) = Right (MkValidated SStr x) - validate SBool (JBool x) = Right (MkValidated SBool x) - validate (SArr s) (JArr xs) = - case validateAll s 0 xs of - Left err => Left err - Right vs => Right (MkValidated (SArr s) vs) - validate (SOpt s) JNull = Right (MkValidated (SOpt s) Nothing) - validate (SOpt s) j = - case validate s j of - Left _ => Left (TypeMismatch (SOpt s)) - Right (MkValidated _ v) => Right (MkValidated (SOpt s) (Just v)) - validate s _ = Left (TypeMismatch s) - - ||| Validate every element of a list against a common schema. Threads - ||| the index so array-element errors can be attributed. - validateAll : (s : Schema) -> (idx : Nat) -> List JsonValue -> - Either ValidationError (List (SchemaType s)) - validateAll _ _ [] = Right [] - validateAll s i (x :: xs) = - case validate s x of - Left err => Left (ArrayElementError i err) - Right (MkValidated _ v) => - case validateAll s (S i) xs of - Left err => Left err - Right vs => Right (v :: vs) - ------------------------------------------------------------------------- --- Type-level soundness (by construction) ------------------------------------------------------------------------- - --- Note: the "if validate returns Right vv, then vv is indexed by the --- schema passed in" claim is a tautology of the type of `validate` --- (the Idris2 type checker rejects any RHS that disagrees with the --- declared return type). It is therefore not stated as a separate --- operator -- the signature of `validate` IS the theorem. - ------------------------------------------------------------------------- --- Lightweight sanity properties (single-step) ------------------------------------------------------------------------- - -||| JNum validates under SNum with the same payload. -public export -validateNumRoundtrip : (x : Double) -> - validate SNum (JNum x) = Right (MkValidated SNum x) -validateNumRoundtrip _ = Refl - -||| JStr validates under SStr. -public export -validateStrRoundtrip : (x : String) -> - validate SStr (JStr x) = Right (MkValidated SStr x) -validateStrRoundtrip _ = Refl - -||| JBool validates under SBool. -public export -validateBoolRoundtrip : (x : Bool) -> - validate SBool (JBool x) = Right (MkValidated SBool x) -validateBoolRoundtrip _ = Refl - -||| JNull validates under SOpt as Nothing. -public export -validateOptNull : (s : Schema) -> - validate (SOpt s) JNull = Right (MkValidated (SOpt s) Nothing) -validateOptNull _ = Refl - -||| A wrong-tagged JsonValue produces a TypeMismatch error (sampled -||| at the SNum schema -- the full mismatch table is exhaustive on -||| JsonValue and Schema and is enforced by `validate`'s totality). -public export -validateNumStrMismatch : (x : String) -> - validate SNum (JStr x) = Left (TypeMismatch SNum) -validateNumStrMismatch _ = Refl - -||| And the symmetric: SStr rejects JNum. -public export -validateStrNumMismatch : (x : Double) -> - validate SStr (JNum x) = Left (TypeMismatch SStr) -validateStrNumMismatch _ = Refl - ------------------------------------------------------------------------- --- End of module ------------------------------------------------------------------------- diff --git a/verisimdb/verification/proofs/idris2/DriftMetric.idr b/verisimdb/verification/proofs/idris2/DriftMetric.idr deleted file mode 100644 index c9bea570..00000000 --- a/verisimdb/verification/proofs/idris2/DriftMetric.idr +++ /dev/null @@ -1,263 +0,0 @@ --- SPDX-License-Identifier: MPL-2.0 --- Copyright (c) 2026 Jonathan D.A. Jewell (hyperpolymath) <j.d.a.jewell@open.ac.uk> --- --- DriftMetric.idr - Formal proof that the VeriSimDB drift metric is a proper --- metric (reflexivity, symmetry, triangle inequality) and that threshold --- detection is sound. --- --- V8 in standards/docs/proofs/spec-templates/T1-critical/verisimdb.md. --- --- Corresponds to: rust-core/verisim-drift/src/calculator.rs. --- --- Model: each Octad modality emits a fixed-width feature vector over Bool --- (feature-present / feature-absent). Drift is Hamming distance: --- the count of feature positions where two snapshots disagree. --- This is a metric (non-negative, reflexive, symmetric, triangle-inequal) --- and the threshold-detection predicate is sound by construction. - -module DriftMetric - -import Data.Nat -import Data.Vect - -%default total - ------------------------------------------------------------------------- --- Boolean xor primitive ------------------------------------------------------------------------- - -||| Exclusive or on Bool: True iff arguments differ. -public export -bxor : Bool -> Bool -> Bool -bxor False False = False -bxor False True = True -bxor True False = True -bxor True True = False - -||| xor is commutative. -public export -bxorCommutative : (a, b : Bool) -> bxor a b = bxor b a -bxorCommutative False False = Refl -bxorCommutative False True = Refl -bxorCommutative True False = Refl -bxorCommutative True True = Refl - -||| Self-xor is always False (a bit never differs from itself). -public export -bxorSelfZero : (a : Bool) -> bxor a a = False -bxorSelfZero False = Refl -bxorSelfZero True = Refl - ------------------------------------------------------------------------- --- Counting True bits ------------------------------------------------------------------------- - -||| Convert a Bool to 0 or 1. -public export -ifB : Bool -> Nat -ifB False = 0 -ifB True = 1 - -||| Count True bits in a Bool vector. -public export -countTrue : {n : Nat} -> Vect n Bool -> Nat -countTrue [] = 0 -countTrue (True :: xs) = S (countTrue xs) -countTrue (False :: xs) = countTrue xs - -||| countTrue on a single element. -public export -countTrueSingle : (b : Bool) -> countTrue [b] = ifB b -countTrueSingle False = Refl -countTrueSingle True = Refl - ------------------------------------------------------------------------- --- Drift metric: Hamming distance ------------------------------------------------------------------------- - -||| A modality snapshot: presence/absence of n features. -public export -State : (n : Nat) -> Type -State n = Vect n Bool - -||| Drift is the count of positions where two states disagree. -||| This is Hamming distance on Bool vectors. -public export -drift : {n : Nat} -> State n -> State n -> Nat -drift xs ys = countTrue (zipWith bxor xs ys) - ------------------------------------------------------------------------- --- Metric axiom 1: identity of indiscernibles (d(x, x) = 0) ------------------------------------------------------------------------- - -||| Drift from a state to itself is zero. -||| Every position xor-against-itself is False, so no bits are counted. -public export -driftSelf : {n : Nat} -> (xs : State n) -> drift xs xs = 0 -driftSelf [] = Refl -driftSelf (False :: xs) = - -- zipWith head: bxor False False = False, so LHS = countTrue (False :: rest) - rewrite driftSelf xs in Refl -driftSelf (True :: xs) = - -- zipWith head: bxor True True = False, so LHS = countTrue (False :: rest) - rewrite driftSelf xs in Refl - ------------------------------------------------------------------------- --- Metric axiom 2: symmetry (d(x, y) = d(y, x)) ------------------------------------------------------------------------- - -||| Pointwise symmetry of zipWith bxor: xs xor ys = ys xor xs at each position. -||| Qualified `DriftMetric.bxor` prevents implicit binding of the lowercase name. -public export -zipWithBxorSym : {n : Nat} -> (xs, ys : Vect n Bool) - -> zipWith DriftMetric.bxor xs ys = zipWith DriftMetric.bxor ys xs -zipWithBxorSym [] [] = Refl -zipWithBxorSym (x :: xs) (y :: ys) = - rewrite bxorCommutative x y in - rewrite zipWithBxorSym xs ys in - Refl - -||| Drift is symmetric: swapping the two states gives the same count. -public export -driftSym : {n : Nat} -> (xs, ys : State n) -> drift xs ys = drift ys xs -driftSym xs ys = cong countTrue (zipWithBxorSym xs ys) - ------------------------------------------------------------------------- --- Metric axiom 3: triangle inequality ------------------------------------------------------------------------- - -||| Per-position triangle inequality on xor: -||| if x != z, then x != y or y != z. -||| Equivalently: ifB (bxor x z) <= ifB (bxor x y) + ifB (bxor y z). -||| Exhaustive over the 8 Bool combinations. -public export -xorTriangleBit : (x, y, z : Bool) - -> LTE (ifB (bxor x z)) (ifB (bxor x y) + ifB (bxor y z)) -xorTriangleBit False False False = LTEZero -xorTriangleBit False False True = LTESucc LTEZero -- (F,F,T): L=1, R=0+1 -xorTriangleBit False True False = LTEZero -- (F,T,F): L=0, R=1+1 -xorTriangleBit False True True = LTESucc LTEZero -- (F,T,T): L=1, R=1+0 -xorTriangleBit True False False = LTESucc LTEZero -- (T,F,F): L=1, R=1+0 -xorTriangleBit True False True = LTEZero -- (T,F,T): L=0, R=1+1 -xorTriangleBit True True False = LTESucc LTEZero -- (T,T,F): L=1, R=0+1 -xorTriangleBit True True True = LTEZero - -||| Monotonicity of addition on both sides. -plusLteMonoBoth : {c : Nat} -> LTE a c -> LTE b d -> LTE (a + b) (c + d) -plusLteMonoBoth LTEZero q = go q - where - go : {c' : Nat} -> LTE b d -> LTE b (c' + d) - go {c' = Z} prf = prf - go {c' = S c''} prf = lteSuccRight (go {c' = c''} prf) -plusLteMonoBoth (LTESucc p) q = LTESucc (plusLteMonoBoth p q) - -||| Helper: countTrue of a cons = ifB head + countTrue tail. -public export -countTrueCons : {n : Nat} -> (b : Bool) -> (rest : Vect n Bool) - -> countTrue (b :: rest) = ifB b + countTrue rest -countTrueCons False rest = Refl -countTrueCons True rest = Refl - -||| Rearrangement used in the triangle step: -||| (a + c) + (b + d) = (a + b) + (c + d) -||| via associativity and commutativity on Nat. -plusSwap : (a, b, c, d : Nat) -> (a + c) + (b + d) = (a + b) + (c + d) -plusSwap a b c d = - rewrite sym (plusAssociative a c (b + d)) in - rewrite plusAssociative c b d in - rewrite plusCommutative c b in - rewrite sym (plusAssociative b c d) in - rewrite plusAssociative a b (c + d) in - Refl - -||| Triangle inequality: -||| drift x z <= drift x y + drift y z -||| Proof: per-position case split on (x_i, y_i, z_i), combined with -||| monotonicity of Nat addition over vector zipWith. -public export -driftTriangle : {n : Nat} - -> (xs, ys, zs : State n) - -> LTE (drift xs zs) (drift xs ys + drift ys zs) -driftTriangle [] [] [] = LTEZero -driftTriangle (x :: xs) (y :: ys) (z :: zs) = - let - -- Per-position bit inequality. - bitPrf : LTE (ifB (bxor x z)) (ifB (bxor x y) + ifB (bxor y z)) - bitPrf = xorTriangleBit x y z - -- Recursive inequality on the tails. - tailPrf : LTE (drift xs zs) (drift xs ys + drift ys zs) - tailPrf = driftTriangle xs ys zs - -- Combine them: bit + tail on both sides. - combined : LTE (ifB (bxor x z) + drift xs zs) - ((ifB (bxor x y) + ifB (bxor y z)) + - (drift xs ys + drift ys zs)) - combined = plusLteMonoBoth bitPrf tailPrf - -- Re-associate the RHS: (xy + yz) + (txy + tyz) = (xy + txy) + (yz + tyz). - rearranged : LTE (ifB (bxor x z) + drift xs zs) - ((ifB (bxor x y) + drift xs ys) + - (ifB (bxor y z) + drift ys zs)) - rearranged = - -- plusSwap a b c d : (a+c)+(b+d) = (a+b)+(c+d). - -- With a=xy, b=yz, c=drift_xy, d=drift_yz: the LHS is the goal's - -- shape and the RHS is combined's shape. Rewriting turns the goal - -- into combined's type, then combined discharges it. - rewrite plusSwap (ifB (bxor x y)) (ifB (bxor y z)) - (drift xs ys) (drift ys zs) - in combined - -- LHS rewrites to drift (x::xs) (z::zs), and each RHS summand to - -- drift (_::_) (_::_), via countTrueCons. - lhsEq : drift (x :: xs) (z :: zs) = ifB (bxor x z) + drift xs zs - lhsEq = countTrueCons (bxor x z) (zipWith bxor xs zs) - rhsEq1 : drift (x :: xs) (y :: ys) = ifB (bxor x y) + drift xs ys - rhsEq1 = countTrueCons (bxor x y) (zipWith bxor xs ys) - rhsEq2 : drift (y :: ys) (z :: zs) = ifB (bxor y z) + drift ys zs - rhsEq2 = countTrueCons (bxor y z) (zipWith bxor ys zs) - in - rewrite lhsEq in - rewrite rhsEq1 in - rewrite rhsEq2 in - rearranged - ------------------------------------------------------------------------- --- Threshold soundness ------------------------------------------------------------------------- - -||| Drift-detection predicate: fires when drift strictly exceeds threshold. -public export -driftExceeds : {n : Nat} -> State n -> State n -> (threshold : Nat) -> Bool -driftExceeds xs ys threshold = - case isLTE (S threshold) (drift xs ys) of - Yes _ => True - No _ => False - -||| Threshold soundness: if drift strictly exceeds the threshold, the -||| detector fires. This is the intended direction — false positives -||| are impossible by construction (the `isLTE` branch decides exactly). -public export -driftExceedsSound : {n : Nat} -> (xs, ys : State n) -> (t : Nat) - -> LTE (S t) (drift xs ys) - -> driftExceeds xs ys t = True -driftExceedsSound xs ys t prf with (isLTE (S t) (drift xs ys)) - _ | Yes _ = Refl - _ | No contra = absurd (contra prf) - -||| Classical contrapositive on LTE over Nat: if S b does not fit under a, -||| then a fits under b. The `No` branch of `isLTE` supplies exactly this -||| negation, so the converse soundness proof below reduces to an application. -public export -notSuccLTEImpliesLTE : (a, b : Nat) -> Not (LTE (S b) a) -> LTE a b -notSuccLTEImpliesLTE Z _ _ = LTEZero -notSuccLTEImpliesLTE (S a') Z contra = absurd (contra (LTESucc LTEZero)) -notSuccLTEImpliesLTE (S a') (S b') contra = - LTESucc (notSuccLTEImpliesLTE a' b' (\lt => contra (LTESucc lt))) - -||| Converse soundness: if the detector doesn't fire, drift does not exceed -||| the threshold. Together with driftExceedsSound this characterises the -||| predicate completely. -public export -driftExceedsComplete : {n : Nat} -> (xs, ys : State n) -> (t : Nat) - -> driftExceeds xs ys t = False - -> LTE (drift xs ys) t -driftExceedsComplete xs ys t prf with (isLTE (S t) (drift xs ys)) - _ | Yes _ = absurd prf - _ | No contra = notSuccLTEImpliesLTE (drift xs ys) t contra diff --git a/verisimdb/verification/proofs/idris2/FFIOwnership.idr b/verisimdb/verification/proofs/idris2/FFIOwnership.idr deleted file mode 100644 index c7552acc..00000000 --- a/verisimdb/verification/proofs/idris2/FFIOwnership.idr +++ /dev/null @@ -1,113 +0,0 @@ --- SPDX-License-Identifier: MPL-2.0 --- Copyright (c) 2026 Jonathan D.A. Jewell (hyperpolymath) --- --- FFIOwnership.idr - V12: FFI pointer validity + ownership discipline. --- --- Scope (spec V12): --- (1) Non-null before dereference. --- (2) Ownership state is explicit in the type. --- (3) Double-free is blocked by type shape (`free` only accepts Alive). - -module FFIOwnership - -%default total - ------------------------------------------------------------------------- --- Ownership state ------------------------------------------------------------------------- - -public export -data PtrState : Type where - Alive : PtrState - Freed : PtrState - ------------------------------------------------------------------------- --- Raw pointer model + non-null witness ------------------------------------------------------------------------- - -public export -record RawPtr where - constructor MkRawPtr - addr : Nat - -public export -data NonNull : Nat -> Type where - IsNonNull : (k : Nat) -> NonNull (S k) - ------------------------------------------------------------------------- --- Owned pointer token ------------------------------------------------------------------------- - -||| Ownership token indexed by pointer state. -||| -||| - `Owned Alive` carries a non-null proof. -||| - `Owned Freed` is a tombstone token that cannot be dereferenced or freed. -public export -record Owned (st : PtrState) where - constructor MkOwned - raw : RawPtr - nonNull : case st of - Alive => NonNull raw.addr - Freed => () - ------------------------------------------------------------------------- --- Core API ------------------------------------------------------------------------- - -||| Allocate an owned pointer token from a non-zero address witness. -public export -alloc : (addr : Nat) -> NonNull addr -> Owned Alive -alloc addr nn = MkOwned (MkRawPtr addr) nn - -||| Free consumes an `Alive` token and returns a `Freed` tombstone token. -||| No function in this module turns `Freed` back into `Alive`. -public export -free : Owned Alive -> Owned Freed -free (MkOwned rp _) = MkOwned rp () - -||| Dereference is only available for `Alive` tokens. -public export -derefAddr : Owned Alive -> Nat -derefAddr (MkOwned rp _) = rp.addr - ------------------------------------------------------------------------- --- Proof obligations ------------------------------------------------------------------------- - -||| Non-null invariant for dereference: every dereference target is provably non-zero. -public export -derefNonNull : (o : Owned Alive) -> NonNull (derefAddr o) -derefNonNull (MkOwned _ nn) = nn - -||| Capability witness: only `Alive` pointers are freeable. -public export -data CanFree : PtrState -> Type where - FreeAlive : CanFree Alive - -||| There is no capability to free a `Freed` token. -public export -noCanFreeFreed : CanFree Freed -> Void -noCanFreeFreed FreeAlive impossible - -||| Freeing requires the `Alive` capability at the type level. -public export -freeWithCap : (st : PtrState) -> CanFree st -> Owned st -> Owned Freed -freeWithCap Alive FreeAlive o = free o - -||| "Double free is impossible" witness: a second free would need `CanFree Freed`, -||| but that type is uninhabited. -public export -doubleFreeImpossible : (o : Owned Freed) -> CanFree Freed -> Void -doubleFreeImpossible _ cap = noCanFreeFreed cap - ------------------------------------------------------------------------- --- Sanity checks ------------------------------------------------------------------------- - -public export -allocOneNonNull : derefNonNull (alloc 1 (IsNonNull 0)) = IsNonNull 0 -allocOneNonNull = Refl - -public export -freePreservesAddress : (o : Owned Alive) -> (free o).raw.addr = o.raw.addr -freePreservesAddress (MkOwned _ _) = Refl diff --git a/verisimdb/verification/proofs/idris2/OctadCoherence.idr b/verisimdb/verification/proofs/idris2/OctadCoherence.idr deleted file mode 100644 index 805bf1ee..00000000 --- a/verisimdb/verification/proofs/idris2/OctadCoherence.idr +++ /dev/null @@ -1,327 +0,0 @@ --- SPDX-License-Identifier: MPL-2.0 --- Copyright (c) 2026 Jonathan D.A. Jewell (hyperpolymath) <j.d.a.jewell@open.ac.uk> --- --- OctadCoherence.idr - Formal proof that the VeriSimDB Octad coherence --- invariant (all 8 modalities remain mutually coherent) is preserved by --- every transaction-wrapped operation. --- --- V1 in standards/docs/proofs/spec-templates/T1-critical/verisimdb.md. --- --- Corresponds to rust-core/verisim-octad/src/store.rs and --- rust-core/verisim-octad/src/transaction.rs (the atomic-across-modalities --- write path). The Rust code uses a TransactionManager to guarantee that --- related modalities update together; we model only the *typed shape* --- of that guarantee and prove coherence preservation holds by construction. --- --- Model choices: --- - Each modality's data is a record whose only relevant coherence-field --- is an abstract Nat (hash / id / version marker). --- - Consistent m1 m2 o is computed pairwise: --- * (Graph, Document): graph.edgeDocRefs = document.id --- * (Vector, Document): vector.embeddingDocHash = document.contentHash --- * (Provenance, Temporal): provenance.temporalVersionRef = temporal.temporalId --- * every other pair: trivially consistent (unit). --- - Op is a closed enumeration of coherence-respecting updates. Each --- coherence-relevant Op updates BOTH sides of its relation in one step --- (the shape of the Rust transaction boundary). - -module OctadCoherence - -%default total - ------------------------------------------------------------------------- --- The eight modalities ------------------------------------------------------------------------- - -public export -data Modality : Type where - Graph : Modality - Vector : Modality - Tensor : Modality - Semantic : Modality - Document : Modality - Temporal : Modality - Provenance : Modality - Spatial : Modality - ------------------------------------------------------------------------- --- Per-modality data ------------------------------------------------------------------------- - -||| Graph modality carries the hash/marker of the document IDs its edges -||| reference. For coherence with Document, this must match document.id. -public export -record GraphModality where - constructor MkGraph - edgeDocRefs : Nat - -||| Vector modality carries the hash of the document content it embeds. -||| For coherence with Document, this must match document.contentHash. -public export -record VectorModality where - constructor MkVector - embeddingDocHash : Nat - -||| Document modality carries its primary ID and content hash. -public export -record DocumentModality where - constructor MkDoc - id : Nat - contentHash : Nat - -||| Provenance modality points at the temporal version it describes. -||| For coherence with Temporal, this must match temporal.temporalId. -public export -record ProvenanceModality where - constructor MkProv - temporalVersionRef : Nat - -||| Temporal modality carries its version hash and primary ID. -public export -record TemporalModality where - constructor MkTemporal - temporalId : Nat - versionHash : Nat - -||| Opaque payloads for the coherence-trivial modalities. -||| Their contents participate in no cross-modality invariant. -public export -data TensorData : Type where - MkTensor : Nat -> TensorData - -public export -data SemanticData : Type where - MkSem : Nat -> SemanticData - -public export -data SpatialData : Type where - MkSpatial : Nat -> SpatialData - ------------------------------------------------------------------------- --- Octad aggregate ------------------------------------------------------------------------- - -public export -record Octad where - constructor MkOctad - graphData : GraphModality - vectorData : VectorModality - tensorData : TensorData - semanticData : SemanticData - documentData : DocumentModality - temporalData : TemporalModality - provenanceData : ProvenanceModality - spatialData : SpatialData - ------------------------------------------------------------------------- --- Pairwise coherence predicate --- --- Computed dispatch. All non-listed pairs collapse to `Unit` which is --- trivially inhabited. The three real cross-modality invariants are: --- --- (Graph, Document): graph.edgeDocRefs = document.id --- (Vector, Document): vector.embedHash = document.contentHash --- (Provenance, Temporal): prov.temporalRef = temporal.temporalId --- --- Each non-trivial case is stated symmetrically (both (A,B) and (B,A)) --- so that Coherent does not depend on ordering. ------------------------------------------------------------------------- - -public export -Consistent : Modality -> Modality -> Octad -> Type -Consistent Graph Document o = o.graphData.edgeDocRefs = o.documentData.id -Consistent Document Graph o = o.graphData.edgeDocRefs = o.documentData.id -Consistent Vector Document o = o.vectorData.embeddingDocHash = o.documentData.contentHash -Consistent Document Vector o = o.vectorData.embeddingDocHash = o.documentData.contentHash -Consistent Provenance Temporal o = o.provenanceData.temporalVersionRef = o.temporalData.temporalId -Consistent Temporal Provenance o = o.provenanceData.temporalVersionRef = o.temporalData.temporalId -Consistent _ _ _ = () - -||| Internal Coherent representation: the three irredundant cross-modality -||| invariants stored as a single record. This is bi-directionally -||| equivalent to the pairwise spec form (see `pairwise` and `fromPairwise` -||| below) but admits much simpler proofs because there are no type-level -||| dispatch catch-alls to unfold. -public export -record Coherent (o : Octad) where - constructor MkCoherent - graphDocCoh : o.graphData.edgeDocRefs = o.documentData.id - vecDocCoh : o.vectorData.embeddingDocHash = o.documentData.contentHash - provTempCoh : o.provenanceData.temporalVersionRef = o.temporalData.temporalId - ------------------------------------------------------------------------- --- Transaction-wrapped operations --- --- Coherence-relevant ops update BOTH sides of their invariant in one step, --- matching the Rust TransactionManager's atomic write boundary. --- Coherence-irrelevant ops update only their own modality. ------------------------------------------------------------------------- - -public export -data Op : Type where - ||| Atomically set graph.edgeDocRefs and document.id to the same new Nat. - UpdateGraphDoc : (newId : Nat) -> Op - ||| Atomically set vector.embeddingDocHash and document.contentHash. - UpdateVecDoc : (newHash : Nat) -> Op - ||| Atomically set provenance.temporalVersionRef and temporal.temporalId. - UpdateProvTemp : (newVersion : Nat) -> Op - ||| Update tensor payload only (coherence-irrelevant). - UpdateTensor : TensorData -> Op - ||| Update semantic payload only (coherence-irrelevant). - UpdateSemantic : SemanticData -> Op - ||| Update spatial payload only (coherence-irrelevant). - UpdateSpatial : SpatialData -> Op - -||| Apply an op to an Octad. -public export -applyOp : Op -> Octad -> Octad -applyOp (UpdateGraphDoc n) o = - { graphData := MkGraph n - , documentData := MkDoc n o.documentData.contentHash - } o -applyOp (UpdateVecDoc h) o = - { vectorData := MkVector h - , documentData := MkDoc o.documentData.id h - } o -applyOp (UpdateProvTemp v) o = - { provenanceData := MkProv v - , temporalData := MkTemporal v o.temporalData.versionHash - } o -applyOp (UpdateTensor t) o = { tensorData := t } o -applyOp (UpdateSemantic s) o = { semanticData := s } o -applyOp (UpdateSpatial s) o = { spatialData := s } o - ------------------------------------------------------------------------- --- Main theorem: every Op preserves Coherence. --- --- The proof is a per-Op, per-pair case split. For ops that touch only --- coherence-irrelevant modalities (tensor/semantic/spatial), the three --- cross-modality invariants are unchanged and the old witness carries --- over. For the three coherence-relevant ops, the paired update sets --- both sides to the same fresh Nat so the invariant becomes Refl. ------------------------------------------------------------------------- - -||| **Main V1 theorem**: for every Octad and every Op, coherence is preserved. -||| -||| Proof is per-op: destructure `o` and the Coherent witness, construct -||| the post-state witness by reusing unchanged invariants and emitting Refl -||| where both sides are set to the same fresh Nat. -public export -opPreservesCoherence : (o : Octad) -> Coherent o -> (op : Op) - -> Coherent (applyOp op o) -opPreservesCoherence (MkOctad _ _ _ _ _ _ _ _) (MkCoherent gd vd pt) (UpdateTensor _) = - MkCoherent gd vd pt -opPreservesCoherence (MkOctad _ _ _ _ _ _ _ _) (MkCoherent gd vd pt) (UpdateSemantic _) = - MkCoherent gd vd pt -opPreservesCoherence (MkOctad _ _ _ _ _ _ _ _) (MkCoherent gd vd pt) (UpdateSpatial _) = - MkCoherent gd vd pt -opPreservesCoherence (MkOctad _ _ _ _ _ _ _ _) (MkCoherent _ vd pt) (UpdateGraphDoc _) = - -- graph.edgeDocRefs and document.id both become the same fresh Nat. - -- vd unchanged: vectorData untouched, and document.contentHash untouched. - -- pt unchanged: provenance and temporal untouched. - MkCoherent Refl vd pt -opPreservesCoherence (MkOctad _ _ _ _ _ _ _ _) (MkCoherent gd _ pt) (UpdateVecDoc _) = - -- vector.embeddingDocHash and document.contentHash both become the - -- same fresh Nat. gd unchanged: graphData and document.id untouched. - MkCoherent gd Refl pt -opPreservesCoherence (MkOctad _ _ _ _ _ _ _ _) (MkCoherent gd vd _) (UpdateProvTemp _) = - -- provenance.temporalVersionRef and temporal.temporalId both become - -- the same fresh Nat. gd, vd unchanged. - MkCoherent gd vd Refl - ------------------------------------------------------------------------- --- Bridge to the spec's pairwise form ------------------------------------------------------------------------- - -||| The spec-prescribed pairwise coherence relation. From an internal -||| `Coherent o` witness, derive `Consistent m1 m2 o` for every pair. -||| Irrelevant pairs reduce to `()` and are inhabited trivially; the three -||| non-trivial pairs dispatch to the matching Coherent field (both -||| orientations). -public export -pairwise : {o : Octad} -> Coherent o -> (m1, m2 : Modality) - -> Consistent m1 m2 o -pairwise (MkCoherent gd _ _) Graph Document = gd -pairwise (MkCoherent gd _ _) Document Graph = gd -pairwise (MkCoherent _ vd _) Vector Document = vd -pairwise (MkCoherent _ vd _) Document Vector = vd -pairwise (MkCoherent _ _ pt) Provenance Temporal = pt -pairwise (MkCoherent _ _ pt) Temporal Provenance = pt --- All remaining pairs reduce Consistent to (); enumerate the 58 pairs. -pairwise _ Graph Graph = () -pairwise _ Graph Vector = () -pairwise _ Graph Tensor = () -pairwise _ Graph Semantic = () -pairwise _ Graph Temporal = () -pairwise _ Graph Provenance = () -pairwise _ Graph Spatial = () -pairwise _ Vector Graph = () -pairwise _ Vector Vector = () -pairwise _ Vector Tensor = () -pairwise _ Vector Semantic = () -pairwise _ Vector Temporal = () -pairwise _ Vector Provenance = () -pairwise _ Vector Spatial = () -pairwise _ Tensor Graph = () -pairwise _ Tensor Vector = () -pairwise _ Tensor Tensor = () -pairwise _ Tensor Semantic = () -pairwise _ Tensor Document = () -pairwise _ Tensor Temporal = () -pairwise _ Tensor Provenance = () -pairwise _ Tensor Spatial = () -pairwise _ Semantic Graph = () -pairwise _ Semantic Vector = () -pairwise _ Semantic Tensor = () -pairwise _ Semantic Semantic = () -pairwise _ Semantic Document = () -pairwise _ Semantic Temporal = () -pairwise _ Semantic Provenance = () -pairwise _ Semantic Spatial = () -pairwise _ Document Tensor = () -pairwise _ Document Semantic = () -pairwise _ Document Document = () -pairwise _ Document Temporal = () -pairwise _ Document Provenance = () -pairwise _ Document Spatial = () -pairwise _ Temporal Graph = () -pairwise _ Temporal Vector = () -pairwise _ Temporal Tensor = () -pairwise _ Temporal Semantic = () -pairwise _ Temporal Document = () -pairwise _ Temporal Temporal = () -pairwise _ Temporal Spatial = () -pairwise _ Provenance Graph = () -pairwise _ Provenance Vector = () -pairwise _ Provenance Tensor = () -pairwise _ Provenance Semantic = () -pairwise _ Provenance Document = () -pairwise _ Provenance Provenance = () -pairwise _ Provenance Spatial = () -pairwise _ Spatial Graph = () -pairwise _ Spatial Vector = () -pairwise _ Spatial Tensor = () -pairwise _ Spatial Semantic = () -pairwise _ Spatial Document = () -pairwise _ Spatial Temporal = () -pairwise _ Spatial Provenance = () -pairwise _ Spatial Spatial = () - ------------------------------------------------------------------------- --- Corollary: repeated applications preserve coherence. ------------------------------------------------------------------------- - -||| Apply a list of ops left-to-right (earliest first). -public export -applyOps : List Op -> Octad -> Octad -applyOps [] o = o -applyOps (op :: rest) o = applyOps rest (applyOp op o) - -||| Coherence survives any finite sequence of Ops. -||| Useful because real transactions consist of multiple field updates. -public export -opsPreserveCoherence : (o : Octad) -> Coherent o -> (ops : List Op) - -> Coherent (applyOps ops o) -opsPreserveCoherence o coh [] = coh -opsPreserveCoherence o coh (op :: rest) = - opsPreserveCoherence (applyOp op o) (opPreservesCoherence o coh op) rest diff --git a/verisimdb/verification/proofs/lean4/RaftSafety.lean b/verisimdb/verification/proofs/lean4/RaftSafety.lean deleted file mode 100644 index 17554d2d..00000000 --- a/verisimdb/verification/proofs/lean4/RaftSafety.lean +++ /dev/null @@ -1,226 +0,0 @@ --- SPDX-License-Identifier: MPL-2.0 -/-! -# Raft Consensus Safety — Commit Invariants - -**Proof obligation V4** — companion to -`nextgen-databases/verisimdb/src/registry/MetadataLog.res` and -`KRaftCluster.res` - -Lean 4 only — no Mathlib. - -## Scope - -Single-node commit safety invariants of the KRaft-style Raft implementation -used in VeriSimDB's federation registry: - -1. **V4-A Commit monotonicity** — `commitIndex` never decreases -2. **V4-B Append isolation** — appending never changes entries at `< commitIndex` -3. **V4-C WF preservation** — well-formedness is maintained by valid appends -4. **V4-D Log Matching (single node)** — sequential index uniquely locates entries - -The distributed Log Matching invariant (no divergence across replicas) is proven -in the companion TLA+ model (`verification/tla+/RaftConsensus.tla`). --/ - --- ============================================================================ --- § 1. Data types --- ============================================================================ - -/-- A single Raft log entry. -/ -structure RaftEntry where - term : Nat -- Raft term (epoch) - idx : Nat -- 1-based sequential log index - cmdId : Nat -- abstract command identifier - deriving DecidableEq, Repr - -/-- Raft node log state (`commitIndex = 0` means nothing committed). -/ -structure RaftLog where - entries : List RaftEntry - commitIndex : Nat - deriving Repr - --- ============================================================================ --- § 2. Well-formedness --- ============================================================================ - -/-- A log is well-formed: sequential indices, bounded commit pointer, and - monotone terms. -/ -structure RaftLog.WF (rl : RaftLog) : Prop where - idxSeq : ∀ i (e : RaftEntry), rl.entries[i]? = some e → e.idx = i + 1 - commitBnd : rl.commitIndex ≤ rl.entries.length - termMono : ∀ i j (ei ej : RaftEntry), - i ≤ j → j < rl.entries.length → - rl.entries[i]? = some ei → - rl.entries[j]? = some ej → - ei.term ≤ ej.term - --- ============================================================================ --- § 3. Transitions --- ============================================================================ - -/-- Append one entry to the log. -/ -def RaftLog.appendEntry (rl : RaftLog) (e : RaftEntry) : RaftLog := - { rl with entries := rl.entries ++ [e] } - -/-- Advance `commitIndex` monotonically. -/ -def RaftLog.advanceCommit (rl : RaftLog) (n : Nat) : RaftLog := - { rl with commitIndex := Nat.max rl.commitIndex n } - --- ============================================================================ --- § 4. V4-A Commit monotonicity --- ============================================================================ - -theorem commitIndex_monotone (rl : RaftLog) (n : Nat) : - rl.commitIndex ≤ (rl.advanceCommit n).commitIndex := - Nat.le_max_left .. - --- ============================================================================ --- § 5. V4-B Append isolation --- ============================================================================ - -/-- Appending a new entry does not change entries at positions below - `commitIndex`. -/ -theorem append_preserves_committed (rl : RaftLog) (e : RaftEntry) (i : Nat) - (hwf : rl.WF) (hi : i < rl.commitIndex) : - (rl.appendEntry e).entries[i]? = rl.entries[i]? := - List.getElem?_append_left (Nat.lt_of_lt_of_le hi hwf.commitBnd) - -/-- Committed entry is still present after any valid append. -/ -theorem committed_entry_stable (rl : RaftLog) (e_new : RaftEntry) - (hwf : rl.WF) (i : Nat) (hi : i < rl.commitIndex) - (x : RaftEntry) (hx : rl.entries[i]? = some x) : - (rl.appendEntry e_new).entries[i]? = some x := by - rw [append_preserves_committed rl e_new i hwf hi]; exact hx - --- ============================================================================ --- § 6. V4-C WF preservation under valid append --- ============================================================================ - -/-! -A **valid** append carries: -- `e.idx = rl.entries.length + 1` (next sequential index) -- `∀ last, getLast? = some last → last.term ≤ e.term` (term monotonicity) --/ - -/-- `getElem?` of the last element in a singleton-appended list. -/ -private theorem getElem?_snoc_last {α} (l : List α) (x : α) : - (l ++ [x])[l.length]? = some x := by - exact List.getElem?_concat_length l x - -/-- In bounds `getElem?` of appended list equals original. -/ -private theorem getElem?_snoc_left {α} (l : List α) (x : α) (i : Nat) - (h : i < l.length) : (l ++ [x])[i]? = l[i]? := - List.getElem?_append_left h - -/-- The only in-bounds position past `l.length` in `l ++ [x]` is `l.length`. -/ -private theorem snoc_out_implies_eq {α} (l : List α) (x : α) (i : Nat) - (hge : ¬ i < l.length) - (hlt : i < (l ++ [x]).length) : i = l.length := by - simp [List.length_append] at hlt; omega - -theorem appendEntry_WF (rl : RaftLog) (e : RaftEntry) - (hwf : rl.WF) - (hidx : e.idx = rl.entries.length + 1) - (hterm : ∀ last, rl.entries.getLast? = some last → last.term ≤ e.term) : - (rl.appendEntry e).WF := by - unfold RaftLog.appendEntry - constructor - · -- idxSeq - intro i x hget - by_cases hi : i < rl.entries.length - · rw [getElem?_snoc_left _ _ _ hi] at hget - exact hwf.idxSeq i x hget - · -- new element - have hlt : i < (rl.entries ++ [e]).length := by - cases Nat.lt_or_ge i (rl.entries ++ [e]).length with - | inl h => exact h - | inr h => - exfalso - have : (rl.entries ++ [e])[i]? = none := - List.getElem?_eq_none_iff.mpr h - simp [this] at hget - have heq : i = rl.entries.length := snoc_out_implies_eq _ _ _ hi hlt - subst heq - rw [getElem?_snoc_last] at hget - exact Option.some.inj hget ▸ hidx - · -- commitBnd - simp only [List.length_append, List.length_singleton] - exact Nat.le_add_right_of_le hwf.commitBnd - · -- termMono - intro i j ei ej hij hjlt hgi hgj - simp only [List.length_append, List.length_singleton] at hjlt - by_cases hj : j < rl.entries.length - · -- j in old log → i also in old log - have hi_lt : i < rl.entries.length := Nat.lt_of_le_of_lt hij hj - rw [getElem?_snoc_left _ _ _ hi_lt] at hgi - rw [getElem?_snoc_left _ _ _ hj] at hgj - exact hwf.termMono i j ei ej hij hj hgi hgj - · -- j = length (the new element) - have hjeq : j = rl.entries.length := by omega - subst hjeq - rw [getElem?_snoc_last] at hgj - -- hgj : some e = some ej - obtain rfl : ej = e := (Option.some.inj hgj).symm - -- goal: ei.term ≤ e.term - by_cases hi_lt : i < rl.entries.length - · -- i in old log: chain ei.term ≤ last.term ≤ e.term - rw [getElem?_snoc_left _ _ _ hi_lt] at hgi - have hlen_pos : 0 < rl.entries.length := - Nat.lt_of_le_of_lt (Nat.zero_le i) hi_lt - have hne : rl.entries ≠ [] := (List.length_pos.mp hlen_pos) - have hbound : rl.entries.length - 1 < rl.entries.length := by omega - have hlast_get : rl.entries.getLast? = some (rl.entries.getLast hne) := - List.getLast?_eq_getLast rl.entries hne - have hlast_idx : rl.entries[rl.entries.length - 1]? = some (rl.entries.getLast hne) := by - rw [List.getElem?_eq_getElem hbound] - exact (congrArg some (List.getLast_eq_getElem rl.entries hne)).symm - have hterm_e := hterm (rl.entries.getLast hne) hlast_get - have hei_last : ei.term ≤ (rl.entries.getLast hne).term := - hwf.termMono i (rl.entries.length - 1) ei (rl.entries.getLast hne) - (by omega) hbound hgi hlast_idx - exact Nat.le_trans hei_last hterm_e - · -- i = rl.entries.length: ei occupies same slot as ej (= e) - have heq : i = rl.entries.length := by omega - subst heq - rw [getElem?_snoc_last] at hgi - -- hgi : some e = some ei → e = ei - exact Nat.le_of_eq (congrArg RaftEntry.term (Option.some.inj hgi).symm) - --- ============================================================================ --- § 7. V4-D Log Matching (single-node) --- ============================================================================ - -/-! -In a well-formed log, the `idx` field uniquely determines the list position: -two entries with the same `idx` are at the same list position. --/ -theorem log_matching_single (rl : RaftLog) (hwf : rl.WF) - (i j : Nat) (ei ej : RaftEntry) - (hgi : rl.entries[i]? = some ei) - (hgj : rl.entries[j]? = some ej) - (hidx_eq : ei.idx = ej.idx) : - i = j := by - have hi := hwf.idxSeq i ei hgi -- ei.idx = i + 1 - have hj := hwf.idxSeq j ej hgj -- ej.idx = j + 1 - -- ei.idx = ej.idx implies i + 1 = j + 1 - rw [hi, hj] at hidx_eq - omega - --- ============================================================================ --- § 8. Summary --- ============================================================================ - -/-! -## Proof summary (V4) - -| Label | Property | Theorem | -|--------|-----------------------------------|-----------------------------| -| V4-A | Commit monotonicity | `commitIndex_monotone` | -| V4-B | Append isolation | `append_preserves_committed` | -| V4-B′ | Committed entry stability | `committed_entry_stable` | -| V4-C | WF preservation under append | `appendEntry_WF` | -| V4-D | Log Matching (single-node) | `log_matching_single` | - -The distributed Log Matching invariant (no divergence across replicas after -commit) is established in the companion TLA+ model. --/ diff --git a/verisimdb/verification/proofs/lean4/VCLSubtyping.lean b/verisimdb/verification/proofs/lean4/VCLSubtyping.lean deleted file mode 100644 index bdb4a1c7..00000000 --- a/verisimdb/verification/proofs/lean4/VCLSubtyping.lean +++ /dev/null @@ -1,203 +0,0 @@ --- SPDX-License-Identifier: MPL-2.0 -/-! -# VCL Subtyping: Transitivity and Decidability - -**Proof obligation V3** — companion to -`nextgen-databases/verisimdb/src/vcl/VCLSubtyping.res` - -Lean 4 only — no Mathlib. All arithmetic discharged by `omega`. - -## Model note - -`QueryResultType`, `ProvedResultType`, and `SigmaType` from the ReScript source -are omitted: they reduce structurally to `Hexad`/`Array` subtyping and add no -new proof complexity. The core structural rules are the subject here. --/ - --- ============================================================================ --- § 1. Type universe --- ============================================================================ - -/-- Primitive scalar types. `VecN n` carries the vector dimension. -/ -inductive PrimType : Type where - | Int | Float | Bool | Str | Uuid | Timestamp - | VecN : Nat → PrimType - deriving DecidableEq, Repr - -/-- The eight octad modalities. -/ -inductive ModalType : Type where - | Graph | Vec | Tensor | Semantic | Doc | Temporal | Provenance | Spatial - deriving DecidableEq, Repr - -/-- Core VCL type universe (structural core, no dependent query-result types). -/ -inductive VclType : Type where - | Prim : PrimType → VclType -- scalar primitive - | Arr : VclType → VclType -- covariant array - | Mod : ModalType → VclType -- single modality designator - | Hexad : List ModalType → VclType -- hexad carrying ≥1 modalities - | Unit : VclType -- top type - | Never : VclType -- bottom type - | Pi : VclType → VclType → VclType -- (domain, codomain) function type - deriving DecidableEq, Repr - --- ============================================================================ --- § 2. Termination measure --- ============================================================================ - -/-- Structural size of a type, used as well-founded measure. -/ -def VclType.size : VclType → Nat - | .Prim _ => 1 - | .Arr t => 1 + t.size - | .Mod _ => 1 - | .Hexad _ => 1 - | .Unit => 1 - | .Never => 1 - | .Pi d c => 1 + d.size + c.size - -/-- All types have strictly positive size. -/ -theorem VclType.size_pos : ∀ t : VclType, 0 < t.size := by - intro t; induction t <;> simp [size] <;> omega - --- ============================================================================ --- § 3. Primitive subtyping --- ============================================================================ - -/-- Primitive widening: only `Int ↪ Float` (safe numeric promotion). -/ -inductive SubPrim : PrimType → PrimType → Prop where - | refl : SubPrim p p - | intFloat : SubPrim .Int .Float - -/-- `SubPrim` is transitive. - The only non-trivial path is `Int <: Float`; there is no `Float <: X` - rule beyond reflexivity, so the intFloat branch after intFloat is - vacuous (Float ≠ Int). -/ -theorem subPrimTrans {p q r : PrimType} (h1 : SubPrim p q) (h2 : SubPrim q r) : - SubPrim p r := by - cases h1 with - | refl => exact h2 - | intFloat => - -- q = Float; only SubPrim Float r via refl fires here - cases h2 with - | refl => exact .intFloat - --- ============================================================================ --- § 4. Core subtype relation --- ============================================================================ - -/-- Structural subtype relation for VCL, formalising the 7 rules - of `VCLSubtyping.checkStructuralSubtype`. -/ -inductive VclSub : VclType → VclType → Prop where - /-- S-Refl: `t <: t`. -/ - | refl : VclSub t t - /-- S-Bot: `Never <: t` (Never is the bottom type). -/ - | neverBot : VclSub .Never t - /-- S-Top: `t <: Unit` (Unit is the top type). -/ - | unitTop : VclSub t .Unit - /-- S-PrimWid: primitive widening (`Int <: Float`). -/ - | primWid : SubPrim p q → VclSub (.Prim p) (.Prim q) - /-- S-ArrCov: array covariance (`Arr a <: Arr b` iff `a <: b`). -/ - | arrCov : VclSub a b → VclSub (.Arr a) (.Arr b) - /-- S-PiSub: function subtyping (contravariant domain, covariant codomain). -/ - | piSub : VclSub d2 d1 → VclSub c1 c2 → VclSub (.Pi d1 c1) (.Pi d2 c2) - /-- S-Hexad: a hexad with MORE modalities subtypes one requiring FEWER - (having more data satisfies a lesser requirement). Implements the - contravariant modality rule from the VCL formal spec. -/ - | hexadSub : (∀ m, m ∈ ms2 → m ∈ ms1) → VclSub (.Hexad ms1) (.Hexad ms2) - --- ============================================================================ --- § 5. V3-A Transitivity --- ============================================================================ - -/-! -### Theorem: `VclSub` is transitive - -If `a <: b` and `b <: c` then `a <: c`. - -**Proof.** Well-founded recursion on `a.size + b.size + c.size`. - -In the only arms with recursive calls: - -* **`arrCov`**: both arguments shrink by 1 (the `Arr` wrapper is removed), - so `a'.size + b'.size + c'.size < (1+a'.size) + (1+b'.size) + (1+c'.size)`. -* **`piSub`**: each recursive call lands on strictly smaller sub-expressions; - the combined measure drops by at least 2 (the two `Pi` wrapper costs). - -All other arms terminate without recursion. --/ -theorem vclSub_trans {a b c : VclType} (hab : VclSub a b) (hbc : VclSub b c) : - VclSub a c := - match hab, hbc with - -- S-Refl on left: a = b, return hbc directly - | .refl, hbc => hbc - -- S-Refl on right: b = c, return hab directly - | hab, .refl => hab - -- S-Bot: Never <: anything - | .neverBot, _ => .neverBot - -- S-Top: anything <: Unit - | _, .unitTop => .unitTop - -- S-PrimWid composed (Int <: Float <: Float = Int <: Float) - | .primWid h, .primWid h' => .primWid (subPrimTrans h h') - -- S-ArrCov composed - | .arrCov h, .arrCov h' => .arrCov (vclSub_trans h h') - -- S-PiSub composed: - -- hab : Pi d1 c1 <: Pi d2 c2 (hd1 : d2 <: d1, hc1 : c1 <: c2) - -- hbc : Pi d2 c2 <: Pi d3 c3 (hd2 : d3 <: d2, hc2 : c2 <: c3) - -- goal: Pi d1 c1 <: Pi d3 c3 need (d3 <: d1) and (c1 <: c3) - | .piSub hd1 hc1, .piSub hd2 hc2 => - .piSub (vclSub_trans hd2 hd1) (vclSub_trans hc1 hc2) - -- S-Hexad composed: ms3 ⊆ ms2 ⊆ ms1 ⟹ ms3 ⊆ ms1 - | .hexadSub hs, .hexadSub hs' => - .hexadSub (fun m hm => hs m (hs' m hm)) -termination_by a.size + b.size + c.size -decreasing_by - all_goals (simp only [VclType.size]; omega) - --- ============================================================================ --- § 6. V3-B Decidability --- ============================================================================ - -/-! -### Theorem: `VclSub` is decidable - -For any `a b : VclType`, either `VclSub a b` or `¬VclSub a b`. - -**Proof.** Directly from the law of excluded middle. - -The constructive witness is the decision algorithm -`VCLSubtyping.isSubtype` in the companion ReScript module, which is a -total function returning `Result<unit, subtypeError>` — it always terminates -with a definite answer. The algorithm is structurally recursive on the -`VclType` constructors with the same termination argument as `vclSub_trans` -above. --/ -theorem vclSub_decidable (a b : VclType) : VclSub a b ∨ ¬VclSub a b := - Classical.em (VclSub a b) - --- ============================================================================ --- § 7. Corollaries and summary --- ============================================================================ - -/-- Reflexivity (direct from `VclSub.refl`). -/ -theorem vclSub_refl (t : VclType) : VclSub t t := .refl - -/-- `Never` is the bottom type. -/ -theorem vclSub_never (t : VclType) : VclSub .Never t := .neverBot - -/-- `Unit` is the top type. -/ -theorem vclSub_unit (t : VclType) : VclSub t .Unit := .unitTop - -/-- `VclSub` is a **preorder**: reflexive and transitive. -/ -theorem vclSub_preorder : - (∀ t, VclSub t t) ∧ - (∀ a b c, VclSub a b → VclSub b c → VclSub a c) := - ⟨vclSub_refl, fun _ _ _ h1 h2 => vclSub_trans h1 h2⟩ - -/-! -## Summary - -| Property | Statement | Proof | -|--------------|-----------|-------| -| Reflexivity | `∀ t, VclSub t t` | `vclSub_refl` (constructor) | -| Transitivity | `VclSub a b → VclSub b c → VclSub a c` | `vclSub_trans` (WF recursion) | -| Decidability | `VclSub a b ∨ ¬VclSub a b` | `vclSub_decidable` (LEM) | --/ diff --git a/verisimdb/verification/proofs/lean4/VCLTypeSoundness.lean b/verisimdb/verification/proofs/lean4/VCLTypeSoundness.lean deleted file mode 100644 index c7033068..00000000 --- a/verisimdb/verification/proofs/lean4/VCLTypeSoundness.lean +++ /dev/null @@ -1,289 +0,0 @@ --- SPDX-License-Identifier: MPL-2.0 -/-! -# VCL Type Inference Soundness (V2) - -Formal model for the query-core fragment implemented by `src/vcl/VCLBidir.res`. - -The model focuses on the pieces that drive result typing in production: - -- query synthesis (`select`, `selectP`, projections) -- checking against expected type via a small subtyping layer -- progress and preservation for the core reduction rules (`fst`, `snd`) - -This is intentionally query-core, not lambda-calculus. The source checker in -`VCLBidir.res` walks a query AST and returns `Result<vclType, typeError>`; it -does not evaluate lambda/application terms. --/ - -namespace VCLTypeSoundness - --- ============================================================================ --- 1. Types --- ============================================================================ - -inductive Modality : Type where - | graph | vector | tensor | semantic | document | temporal | provenance | spatial - deriving DecidableEq, Repr - -inductive PrimTy : Type where - | int | float | string | bool | uuid | timestamp - deriving DecidableEq, Repr - -inductive ProofKind : Type where - | existence | integrity | consistency | provenanceProof | freshness | access | citation | custom - deriving DecidableEq, Repr - -inductive Ty : Type where - | prim : PrimTy → Ty - | array : Ty → Ty - | modality : Modality → Ty - | unit : Ty - | never : Ty - | queryRes : List Modality → Ty - | proof : ProofKind → String → Ty - | sigma : Ty → Ty → Ty - | piType : String → Ty → Ty → Ty - deriving DecidableEq, Repr - --- ============================================================================ --- 2. Expressions and values --- ============================================================================ - -structure FieldRef where - mod : Modality - field : String - deriving DecidableEq, Repr - -inductive AggFunc : Type where - | count | sum | avg | min | max - deriving DecidableEq, Repr - -inductive Lit : Type where - | intLit : Int → Lit - | strLit : String → Lit - | boolLit : Bool → Lit - deriving DecidableEq, Repr - -def Lit.ty : Lit → PrimTy - | .intLit _ => .int - | .strLit _ => .string - | .boolLit _ => .bool - -inductive Expr : Type where - | lit : Lit → Expr - | fieldGet : FieldRef → Expr - | agg : AggFunc → FieldRef → Expr - | select : List Modality → List FieldRef → Expr - | selectP : List Modality → List FieldRef → ProofKind → String → Expr - | proofWitness : ProofKind → String → Expr - | pair : Expr → Expr → Expr - | fst : Expr → Expr - | snd : Expr → Expr - | unitVal : Expr - deriving DecidableEq, Repr - -inductive IsValue : Expr → Prop where - | litV : IsValue (.lit l) - | fieldV : IsValue (.fieldGet f) - | aggV : IsValue (.agg g f) - | selectV : IsValue (.select ms fs) - | selectPV : IsValue (.selectP ms fs pk c) - | proofWV : IsValue (.proofWitness pk c) - | unitV : IsValue .unitVal - | pairV : IsValue e₁ → IsValue e₂ → IsValue (.pair e₁ e₂) - --- ============================================================================ --- 3. Typing --- ============================================================================ - -abbrev Ctx := FieldRef → Option PrimTy - -def Ctx.empty : Ctx := fun _ => none - -inductive HasType : Ctx → Expr → Ty → Prop where - | tLit : HasType Γ (.lit l) (.prim l.ty) - | tUnit : HasType Γ .unitVal .unit - | tField : Γ fr = some pt → HasType Γ (.fieldGet fr) (.prim pt) - | tAggCount : HasType Γ (.agg .count fr) (.prim .int) - | tAggSumI : Γ fr = some .int → HasType Γ (.agg .sum fr) (.prim .int) - | tAggSumF : Γ fr = some .float → HasType Γ (.agg .sum fr) (.prim .float) - | tAggAvg : Γ fr = some .int ∨ Γ fr = some .float → - HasType Γ (.agg .avg fr) (.prim .float) - | tAggMinMax : Γ fr = some pt → (g = .min ∨ g = .max) → - HasType Γ (.agg g fr) (.prim pt) - | tSelect : HasType Γ (.select ms fs) (.queryRes ms) - | tSelectP : HasType Γ (.selectP ms fs pk c) (.sigma (.queryRes ms) (.proof pk c)) - | tProofW : HasType Γ (.proofWitness pk c) (.proof pk c) - | tPair : HasType Γ e₁ τ₁ → HasType Γ e₂ τ₂ → - HasType Γ (.pair e₁ e₂) (.sigma τ₁ τ₂) - | tFst : HasType Γ e (.sigma τ₁ τ₂) → HasType Γ (.fst e) τ₁ - | tSnd : HasType Γ e (.sigma τ₁ τ₂) → HasType Γ (.snd e) τ₂ - --- ============================================================================ --- 4. Operational semantics --- ============================================================================ - -inductive Step : Expr → Expr → Prop where - | fstPair : IsValue e₁ → IsValue e₂ → Step (.fst (.pair e₁ e₂)) e₁ - | sndPair : IsValue e₁ → IsValue e₂ → Step (.snd (.pair e₁ e₂)) e₂ - | fstSelectP : Step (.fst (.selectP ms fs pk c)) (.select ms fs) - | sndSelectP : Step (.snd (.selectP ms fs pk c)) (.proofWitness pk c) - | fstStep : Step e e' → Step (.fst e) (.fst e') - | sndStep : Step e e' → Step (.snd e) (.snd e') - | pairStepL : Step e₁ e₁' → Step (.pair e₁ e₂) (.pair e₁' e₂) - | pairStepR : IsValue e₁ → Step e₂ e₂' → Step (.pair e₁ e₂) (.pair e₁ e₂') - -inductive Progress (e : Expr) : Prop where - | done : IsValue e → Progress e - | step : Step e e' → Progress e - --- ============================================================================ --- 5. Progress --- ============================================================================ - -theorem progress {Γ : Ctx} {e : Expr} {τ : Ty} (ht : HasType Γ e τ) : Progress e := by - induction ht with - | tLit => exact .done .litV - | tUnit => exact .done .unitV - | tField _ => exact .done .fieldV - | tAggCount => exact .done .aggV - | tAggSumI _ => exact .done .aggV - | tAggSumF _ => exact .done .aggV - | tAggAvg _ => exact .done .aggV - | tAggMinMax _ _ => exact .done .aggV - | tSelect => exact .done .selectV - | tSelectP => exact .done .selectPV - | tProofW => exact .done .proofWV - | tPair _ _ ih₁ ih₂ => - cases ih₁ with - | done hv₁ => - cases ih₂ with - | done hv₂ => exact .done (.pairV hv₁ hv₂) - | step hs₂ => exact .step (.pairStepR hv₁ hs₂) - | step hs₁ => exact .step (.pairStepL hs₁) - | tFst hSigma ih => - cases ih with - | done hv => - cases hSigma with - | tPair h1 h2 => - cases hv with - | pairV hv1 hv2 => exact .step (.fstPair hv1 hv2) - | tSelectP => - exact .step .fstSelectP - | tFst _ => cases hv - | tSnd _ => cases hv - | step hs => exact .step (.fstStep hs) - | tSnd hSigma ih => - cases ih with - | done hv => - cases hSigma with - | tPair h1 h2 => - cases hv with - | pairV hv1 hv2 => exact .step (.sndPair hv1 hv2) - | tSelectP => - exact .step .sndSelectP - | tFst _ => cases hv - | tSnd _ => cases hv - | step hs => exact .step (.sndStep hs) - --- ============================================================================ --- 6. Preservation --- ============================================================================ - -theorem preservation {Γ : Ctx} {e e' : Expr} {τ : Ty} - (ht : HasType Γ e τ) (hs : Step e e') : HasType Γ e' τ := by - induction hs generalizing τ with - | fstPair hv₁ hv₂ => - cases ht with - | tFst hSigma => - cases hSigma with - | tPair h1 _ => exact h1 - | sndPair hv₁ hv₂ => - cases ht with - | tSnd hSigma => - cases hSigma with - | tPair _ h2 => exact h2 - | fstSelectP => - cases ht with - | tFst hSigma => - cases hSigma with - | tSelectP => exact .tSelect - | sndSelectP => - cases ht with - | tSnd hSigma => - cases hSigma with - | tSelectP => exact .tProofW - | fstStep hsInner ih => - cases ht with - | tFst hSigma => exact .tFst (ih hSigma) - | sndStep hsInner ih => - cases ht with - | tSnd hSigma => exact .tSnd (ih hSigma) - | pairStepL hsL ih => - cases ht with - | tPair h1 h2 => exact .tPair (ih h1) h2 - | pairStepR hv hsR ih => - cases ht with - | tPair h1 h2 => exact .tPair h1 (ih h2) - --- ============================================================================ --- 7. Multi-step soundness corollary --- ============================================================================ - -inductive Steps : Expr → Expr → Prop where - | refl : Steps e e - | trans : Step e e' → Steps e' e'' → Steps e e'' - -theorem preservationSteps {Γ : Ctx} {e e' : Expr} {τ : Ty} - (ht : HasType Γ e τ) (hs : Steps e e') : HasType Γ e' τ := by - induction hs with - | refl => exact ht - | trans hstep hrest ih => - exact ih (preservation ht hstep) - -def HasTypeVal (v : Expr) (τ : Ty) : Prop := HasType Ctx.empty v τ ∧ IsValue v - -theorem type_soundness {e v : Expr} {τ : Ty} - (ht : HasType Ctx.empty e τ) - (hs : Steps e v) - (hv : IsValue v) : - HasTypeVal v τ := by - exact ⟨preservationSteps ht hs, hv⟩ - --- ============================================================================ --- 8. Bidirectional synthesis/checking soundness --- ============================================================================ - -inductive SubTy : Ty → Ty → Prop where - | refl : SubTy τ τ - | provedForget : - SubTy (.sigma (.queryRes ms) (.proof pk c)) (.queryRes ms) - -theorem subTy_trans {a b c : Ty} (h₁ : SubTy a b) (h₂ : SubTy b c) : SubTy a c := by - cases h₁ with - | refl => exact h₂ - | provedForget => - cases h₂ with - | refl => exact .provedForget - -inductive Synth : Ctx → Expr → Ty → Prop where - | fromTyping : HasType Γ e τ → Synth Γ e τ - -inductive Check : Ctx → Expr → Ty → Prop where - | exact : Synth Γ e τ → Check Γ e τ - | subsume : Synth Γ e τ → SubTy τ τ' → Check Γ e τ' - -theorem synth_sound {Γ : Ctx} {e : Expr} {τ : Ty} - (hs : Synth Γ e τ) : HasType Γ e τ := by - cases hs with - | fromTyping ht => exact ht - -theorem check_sound {Γ : Ctx} {e : Expr} {τ : Ty} - (hc : Check Γ e τ) : ∃ τ', HasType Γ e τ' ∧ SubTy τ' τ := by - cases hc with - | exact hs => - exact ⟨τ, synth_sound hs, .refl⟩ - | subsume hs hsub => - exact ⟨_, synth_sound hs, hsub⟩ - -end VCLTypeSoundness diff --git a/verisimdb/verification/proofs/lean4/WALIntegrity.lean b/verisimdb/verification/proofs/lean4/WALIntegrity.lean deleted file mode 100644 index e2283094..00000000 --- a/verisimdb/verification/proofs/lean4/WALIntegrity.lean +++ /dev/null @@ -1,202 +0,0 @@ --- SPDX-License-Identifier: MPL-2.0 -/-! -# WAL Integrity — Sequence Monotonicity, CRC Protection, Replay Idempotence - -**Proof obligation V6** — companion to -`nextgen-databases/verisimdb/rust-core/verisim-wal/` - -Lean 4 only — no Mathlib. - -## Scope - -Structural invariants of VeriSimDB's Write-Ahead Log: - -1. **V6-A Sequence monotonicity** — sequence numbers strictly increase -2. **V6-B CRC validity** — every entry in a well-formed log passes its checksum -3. **V6-C Replay compositionality** — `replay (xs ++ ys) s = replay ys (replay xs s)` -4. **V6-D Checkpoint idempotence** — replaying from checkpoint N yields the same - result whether entries 0..N were applied once or multiple times - -## On-disk entry format (modelled abstractly) - -``` -[4 bytes: entry_length] [4 bytes: crc32] -[8 bytes: sequence u64] [8 bytes: timestamp i64] -[1 byte: operation] [1 byte: modality] -[4+N bytes: entity_id] [4+M bytes: payload] -``` - -The proof models sequence and CRC structurally; timestamp, operation, and -modality are abstracted into `cmdId : Nat` since the ordering invariants are -independent of those fields. --/ - --- ============================================================================ --- § 1. Data types --- ============================================================================ - -/-- A single WAL entry (abstract model). -/ -structure WalEntry where - seq : Nat -- monotonically strictly increasing sequence number (1-based) - cmdId : Nat -- abstract command payload identifier - crcOk : Bool -- CRC32 validity (true iff checksum matches) - deriving DecidableEq, Repr - -/-- Abstract WAL state: tracks how many commands have been applied. -/ -abbrev WalState := Nat - --- ============================================================================ --- § 2. Transitions --- ============================================================================ - -/-- Apply one WAL entry to the state. -/ -def applyEntry (s : WalState) (_ : WalEntry) : WalState := s + 1 - -/-- Replay a list of WAL entries onto an initial state. -/ -def replay (entries : List WalEntry) (s : WalState) : WalState := - entries.foldl applyEntry s - --- ============================================================================ --- § 3. Well-formedness --- ============================================================================ - -namespace WalLog - -/-- A WAL log segment is well-formed relative to a base sequence number. - Both fields use `entries[i]?` so CRC can be checked at any valid index - without a separate membership proof. -/ -structure WF (entries : List WalEntry) (base : Nat) : Prop where - /-- Position `i` carries sequence `base + i + 1` (1-based, no gaps). -/ - seqMono : ∀ (i : Nat) (e : WalEntry), entries[i]? = some e → e.seq = base + i + 1 - /-- Every entry at a valid index has a passing CRC. -/ - crcValid : ∀ (i : Nat) (e : WalEntry), entries[i]? = some e → e.crcOk = true - -end WalLog - --- ============================================================================ --- § 4. V6-A Sequence monotonicity --- ============================================================================ - -/-- In a well-formed log, earlier entries have strictly smaller sequence numbers. -/ -theorem seq_strictly_increasing (entries : List WalEntry) (base : Nat) - (hwf : WalLog.WF entries base) - (i j : Nat) (ei ej : WalEntry) - (hgi : entries[i]? = some ei) - (hgj : entries[j]? = some ej) - (hij : i < j) : - ei.seq < ej.seq := by - have hi := hwf.seqMono i ei hgi -- ei.seq = base + i + 1 - have hj := hwf.seqMono j ej hgj -- ej.seq = base + j + 1 - rw [hi, hj]; omega - -/-- Sequence numbers are unique in a well-formed log (position determined by seq). -/ -theorem seq_injective (entries : List WalEntry) (base : Nat) - (hwf : WalLog.WF entries base) - (i j : Nat) (ei ej : WalEntry) - (hgi : entries[i]? = some ei) - (hgj : entries[j]? = some ej) - (hseq : ei.seq = ej.seq) : - i = j := by - have hi := hwf.seqMono i ei hgi - have hj := hwf.seqMono j ej hgj - rw [hi, hj] at hseq; omega - --- ============================================================================ --- § 5. V6-B CRC validity --- ============================================================================ - -/-- Every entry at a valid index in a well-formed log has `crcOk = true`. -/ -theorem crc_valid_at (entries : List WalEntry) (base : Nat) - (hwf : WalLog.WF entries base) - (i : Nat) (e : WalEntry) (hgi : entries[i]? = some e) : - e.crcOk = true := - hwf.crcValid i e hgi - --- ============================================================================ --- § 6. V6-C Replay compositionality --- ============================================================================ - -/-- Replaying an empty log is the identity. -/ -theorem replay_nil (s : WalState) : replay [] s = s := rfl - -/-- Replaying a singleton advances the state by exactly 1. -/ -theorem replay_singleton (e : WalEntry) (s : WalState) : - replay [e] s = s + 1 := rfl - -/-- Replay distributes over list concatenation. -/ -theorem replay_append (xs ys : List WalEntry) (s : WalState) : - replay (xs ++ ys) s = replay ys (replay xs s) := by - simp [replay, List.foldl_append] - -/-- Replaying `n` entries from state `s` yields `s + n`. -/ -theorem replay_length (entries : List WalEntry) (s : WalState) : - replay entries s = s + entries.length := by - induction entries generalizing s with - | nil => rfl - | cons hd tl ih => - -- replay (hd :: tl) s is definitionally replay tl (s + 1) - have key : replay (hd :: tl) s = replay tl (s + 1) := rfl - rw [key, ih (s + 1), List.length_cons, Nat.add_assoc, Nat.add_comm 1] - --- ============================================================================ --- § 7. V6-D Checkpoint idempotence --- ============================================================================ - -/-! -A **checkpoint** at position `n` records the state after replaying entries -`0..n-1`. On crash recovery, we seek to position `n` and replay entries -`n..end` onto the checkpointed state. - -The idempotence theorem: the final state is the same whether entries `0..n-1` -were applied once (normal path) or had been applied at checkpoint time. --/ - -/-- Splitting a log at position `n` and replaying the suffix from the - checkpointed state yields the same result as a single full replay. -/ -theorem checkpoint_idempotent (entries : List WalEntry) (s : WalState) (n : Nat) : - replay (entries.drop n) (replay (entries.take n) s) = replay entries s := by - rw [← replay_append, List.take_append_drop] - -/-- Prefix replay + suffix replay = full replay (the `no_double_apply` form). -/ -theorem no_double_apply (pfx sfx : List WalEntry) (s : WalState) : - replay sfx (replay pfx s) = replay (pfx ++ sfx) s := by - rw [replay_append] - -/-- The checkpointed state after `n` entries equals `s + n`. -/ -theorem checkpoint_state_value (entries : List WalEntry) (s : WalState) - (n : Nat) (hn : n ≤ entries.length) : - replay (entries.take n) s = s + n := by - rw [replay_length, List.length_take, Nat.min_eq_left hn] - --- ============================================================================ --- § 8. Corollary: full replay count --- ============================================================================ - -/-- A complete replay of a well-formed `k`-entry log from state `s` yields - `s + k`. -/ -theorem full_replay_count (entries : List WalEntry) (s : WalState) : - replay entries s = s + entries.length := - replay_length entries s - --- ============================================================================ --- § 9. Summary --- ============================================================================ - -/-! -## Proof summary (V6) - -| Label | Property | Theorem | -|--------|-----------------------------------|-------------------------------| -| V6-A | Sequence strictly increasing | `seq_strictly_increasing` | -| V6-A′ | Sequence numbers injective | `seq_injective` | -| V6-B | CRC validity at any index | `crc_valid_at` | -| V6-C | Replay over concatenation | `replay_append` | -| V6-C′ | Replay length equals entry count | `replay_length` | -| V6-D | Checkpoint idempotence | `checkpoint_idempotent` | -| V6-D′ | No double-apply of prefix entries | `no_double_apply` | -| V6-D″ | Checkpoint state value | `checkpoint_state_value` | - -The on-disk CRC32 computation and the concrete `applyEntry` state machine -(modality stores, entity IDs, payloads) are verified by the Rust test suite -in `verisim-wal/src/`. --/ diff --git a/verisimdb/verification/proofs/lean4/lake-manifest.json b/verisimdb/verification/proofs/lean4/lake-manifest.json deleted file mode 100644 index 1d2142d6..00000000 --- a/verisimdb/verification/proofs/lean4/lake-manifest.json +++ /dev/null @@ -1,5 +0,0 @@ -{"version": "1.1.0", - "packagesDir": ".lake/packages", - "packages": [], - "name": "verisimdb_proofs", - "lakeDir": ".lake"} diff --git a/verisimdb/verification/proofs/lean4/lakefile.lean b/verisimdb/verification/proofs/lean4/lakefile.lean deleted file mode 100644 index d8e49993..00000000 --- a/verisimdb/verification/proofs/lean4/lakefile.lean +++ /dev/null @@ -1,9 +0,0 @@ --- SPDX-License-Identifier: MPL-2.0 -import Lake -open Lake DSL - -package verisimdb_proofs where - name := `verisimdb_proofs - -lean_lib VeriSimDBProofs where - roots := #[`VCLSubtyping, `VCLTypeSoundness, `RaftSafety, `WALIntegrity] diff --git a/verisimdb/verification/proofs/tlaplus/.gitignore b/verisimdb/verification/proofs/tlaplus/.gitignore deleted file mode 100644 index 326d23b6..00000000 --- a/verisimdb/verification/proofs/tlaplus/.gitignore +++ /dev/null @@ -1,4 +0,0 @@ -# SPDX-License-Identifier: MPL-2.0 -# TLC model-checker scratch output. -states/ -tlc.log diff --git a/verisimdb/verification/proofs/tlaplus/Normalizer.cfg b/verisimdb/verification/proofs/tlaplus/Normalizer.cfg deleted file mode 100644 index d5bfe05f..00000000 --- a/verisimdb/verification/proofs/tlaplus/Normalizer.cfg +++ /dev/null @@ -1,25 +0,0 @@ -\* SPDX-License-Identifier: MPL-2.0 -\* TLC configuration for Normalizer.tla (V9). -\* -\* Bounded instance: 4 modalities (abstracted from 8), 3 distinct values, -\* MaxSteps=3. State space = |Values|^|Modalities| initial states (3^4=81) -\* times up to MaxSteps transitions. Convergence happens in 1 step, so -\* effective space ~= 81 * 2 = 162 reachable states. Finishes instantly. -\* -\* Run with: -\* java -cp tla2tools.jar tlc2.TLC -config Normalizer.cfg Normalizer.tla -\* or via `just verify-tlaplus`. - -SPECIFICATION Spec - -CONSTANTS - Values = {"v0", "v1", "v2"} - MaxSteps = 3 - -INVARIANTS - NormalizerSafe - -PROPERTIES - Convergence - -CHECK_DEADLOCK FALSE diff --git a/verisimdb/verification/proofs/tlaplus/Normalizer.tla b/verisimdb/verification/proofs/tlaplus/Normalizer.tla deleted file mode 100644 index de500d51..00000000 --- a/verisimdb/verification/proofs/tlaplus/Normalizer.tla +++ /dev/null @@ -1,139 +0,0 @@ -------------------------------- MODULE Normalizer ------------------------------- -\* SPDX-License-Identifier: MPL-2.0 -\* Copyright (c) 2026 Jonathan D.A. Jewell (hyperpolymath) <j.d.a.jewell@open.ac.uk> -\* -\* V9: Normalizer determinism + convergence. -\* Corresponds to rust-core/verisim-normalizer/src/lib.rs (StorageRegenerator). -\* -\* The verisim-normalizer resolves drift between the 8 modalities of an octad -\* by picking an authoritative source modality and regenerating the others -\* from it. The real system has a (source, target) strategy table -- Document -\* is the usual authoritative source for Vector/Semantic/Graph regeneration, -\* with cosine-similarity drift measured against Vector and Jaccard against -\* Semantic. This spec abstracts that machinery into its essential claim: -\* -\* - Determinism: normalisation is a *function* of state. Given the same -\* input octad, normalisation always produces the same output octad; -\* there is no schedule-dependent or source-rank-tie non-determinism. -\* - Convergence: starting from any drift-ed state, repeated normalisation -\* reaches a drift-free fixed point in bounded time. -\* -\* The spec also checks that the fixed point is *stable* (Normalize is -\* identity on a drift-free state). - -EXTENDS Naturals, FiniteSets, TLC - -CONSTANTS - Values, \* abstract set of possible modality payload hashes - MaxSteps \* bound on normalisation rounds for model-checking - -\* Modalities and their deterministic priority ordering are module-level (TLC -\* config files cannot represent record literals, and the real system has a -\* fixed strategy table anyway). Priority is strictly injective by -\* construction here -- that injectivity is exactly what makes the CHOOSE in -\* SourceOf deterministic, and it is the spec's central structural claim. -Modalities == {"graph", "vector", "semantic", "document"} - -Priority == [graph |-> 1, vector |-> 2, semantic |-> 3, document |-> 4] - -ASSUME Cardinality(Values) >= 1 -ASSUME MaxSteps \in Nat -ASSUME \A m1, m2 \in Modalities: (m1 /= m2) => (Priority[m1] /= Priority[m2]) - -VARIABLES - state, \* [Modalities -> Values] -- current octad snapshot - steps \* Nat -- normalisation rounds elapsed - -vars == <<state, steps>> - -TypeOK == - /\ state \in [Modalities -> Values] - /\ steps \in 0..MaxSteps - -\* Drift holds when any two modalities disagree on the payload. -HasDrift(s) == - \E m1, m2 \in Modalities: s[m1] /= s[m2] - -\* The deterministic authoritative-source function. Under the injectivity -\* ASSUME above, CHOOSE returns the unique highest-priority modality. This -\* is the single piece of the spec that, if wrong, would make the whole -\* normaliser non-deterministic -- so it is the thing determinism hinges on. -SourceOf(s) == - CHOOSE m \in Modalities: - \A other \in Modalities: Priority[m] >= Priority[other] - -\* The normaliser: rewrite every modality to the source's value. -Normalize(s) == - [m \in Modalities |-> s[SourceOf(s)]] - -\* Non-deterministic initial state; all octads are possible starting points -\* for the model-check. -Init == - /\ state \in [Modalities -> Values] - /\ steps = 0 - -\* One round of normalisation. Guard by HasDrift so converged states are -\* stuttering; guard by MaxSteps for finite model-check. -Step == - /\ steps < MaxSteps - /\ HasDrift(state) - /\ state' = Normalize(state) - /\ steps' = steps + 1 - -Next == Step - -\* Weak fairness forces Step to fire while drift remains, which is what makes -\* Convergence true. Without it, the system could stutter forever in a drifted -\* state (that would be a real bug in the implementation, not the spec). -Spec == Init /\ [][Next]_vars /\ WF_vars(Step) - --------------------------------------------------------------------------------- -\* Safety --------------------------------------------------------------------------------- - -\* I1. SourceOf is well-defined: the CHOOSE always returns an element of -\* Modalities, and that element is the unique priority-maximum. Trivial from -\* the ASSUME, but stated explicitly so TLC exercises it on every state. -SourceIsMaximal == - \A other \in Modalities: - Priority[SourceOf(state)] >= Priority[other] - -\* I2. Idempotence of Normalize: normalising a normalised state is a no-op. -\* This is a property of the *definition* of Normalize; TLC checks it across -\* all reachable states including the non-deterministic Init. -NormalizeIdempotent == - Normalize(Normalize(state)) = Normalize(state) - -\* I3. Post-Step drift-free: immediately after Step, no drift remains. -\* Implementation: Step replaces state with Normalize(state), which makes -\* every modality equal to state[SourceOf(state)] -- so any two modalities -\* agree, i.e. ~HasDrift. TLC verifies by exploring. -PostStepNoDrift == - (steps > 0) => ~HasDrift(state) - -\* I4. Stability of fixed point: in any drift-free state, Normalize is the -\* identity. This is the "once converged, stay converged" guarantee. -FixedPointStable == - ~HasDrift(state) => (Normalize(state) = state) - -NormalizerSafe == - /\ TypeOK - /\ SourceIsMaximal - /\ NormalizeIdempotent - /\ PostStepNoDrift - /\ FixedPointStable - --------------------------------------------------------------------------------- -\* Liveness / convergence --------------------------------------------------------------------------------- - -\* Eventually, drift is gone and stays gone. Stronger than "eventually no -\* drift" because the system could in principle re-drift if Step non- -\* deterministically reintroduced disagreement -- the <>[] form forbids that. -Convergence == - <>[]~HasDrift(state) - -THEOREM NormalizerSafety == Spec => []NormalizerSafe -THEOREM NormalizerConverges == Spec => Convergence - -================================================================================ diff --git a/verisimdb/verification/proofs/tlaplus/OctadAtomicity.cfg b/verisimdb/verification/proofs/tlaplus/OctadAtomicity.cfg deleted file mode 100644 index 96752568..00000000 --- a/verisimdb/verification/proofs/tlaplus/OctadAtomicity.cfg +++ /dev/null @@ -1,24 +0,0 @@ -\* SPDX-License-Identifier: MPL-2.0 -\* TLC configuration for OctadAtomicity.tla (V5). -\* -\* Bounded instance: 2 transactions, up to 2 crashes per run. With 8 modalities -\* and 3 status values, state space is bounded by -\* (3 * 2^8)^2 * 3 ~ 1.8M states upper bound; in practice < 100k reachable. -\* -\* Run with: -\* java -cp tla2tools.jar tlc2.TLC -config OctadAtomicity.cfg OctadAtomicity.tla -\* or via the repo `just verify-tlaplus` recipe (see verisimdb Justfile). - -SPECIFICATION Spec - -CONSTANTS - Txns = {"t1", "t2"} - MaxCrashes = 2 - -INVARIANTS - OctadSafe - -PROPERTIES - EveryTxnResolves - -CHECK_DEADLOCK FALSE diff --git a/verisimdb/verification/proofs/tlaplus/OctadAtomicity.tla b/verisimdb/verification/proofs/tlaplus/OctadAtomicity.tla deleted file mode 100644 index 6467e88b..00000000 --- a/verisimdb/verification/proofs/tlaplus/OctadAtomicity.tla +++ /dev/null @@ -1,160 +0,0 @@ ------------------------------- MODULE OctadAtomicity ------------------------------ -\* SPDX-License-Identifier: MPL-2.0 -\* Copyright (c) 2026 Jonathan D.A. Jewell (hyperpolymath) <j.d.a.jewell@open.ac.uk> -\* -\* V5: Transaction atomicity across the 8 modalities of a VeriSimDB octad. -\* Corresponds to rust-core/verisim-octad/src/transaction.rs. -\* -\* A VeriSimDB octad is a record with 8 per-modality projections. A transaction -\* touches a subset of the 8 modality projections and must be either fully -\* committed (all 8 projections updated) or fully aborted (none visible). No -\* PARTIAL state is ever observable outside a transaction's own PENDING scope. -\* -\* This spec models the transaction lifecycle (PENDING -> COMMITTED | ABORTED) -\* together with fault injection (crashes during a PENDING transaction) and -\* proves the Atomicity invariant under an adversary that can interleave -\* multiple concurrent transactions and crashes. - -EXTENDS Naturals, FiniteSets, TLC - -CONSTANTS - Txns, \* finite set of transaction identifiers - MaxCrashes \* bound on the number of crash events per run - -ASSUME Cardinality(Txns) >= 1 -ASSUME MaxCrashes \in Nat - -\* The 8 modalities of a VeriSimDB octad. -Modalities == {"graph", "vector", "tensor", "semantic", - "document", "temporal", "provenance", "spatial"} - -Status == {"PENDING", "COMMITTED", "ABORTED"} - -VARIABLES - txnStatus, \* [Txns -> Status] - each transaction's lifecycle stage - modalityUpdates, \* [Txns -> SUBSET Modalities] - modalities already applied - crashes \* Nat - bounded counter for fault-injection events - -vars == <<txnStatus, modalityUpdates, crashes>> - -TypeOK == - /\ txnStatus \in [Txns -> Status] - /\ modalityUpdates \in [Txns -> SUBSET Modalities] - /\ crashes \in 0..MaxCrashes - -Init == - /\ txnStatus = [t \in Txns |-> "PENDING"] - /\ modalityUpdates = [t \in Txns |-> {}] - /\ crashes = 0 - --------------------------------------------------------------------------------- -\* Incremental per-modality application during a PENDING transaction. -\* Each modality can be applied at most once per transaction (enforced by the -\* set-union: \cup already-applied is idempotent, but we also disallow reapply -\* via the `m \notin` guard so TLC bounds the state space by 2^|Modalities|). --------------------------------------------------------------------------------- -ApplyModality(t, m) == - /\ txnStatus[t] = "PENDING" - /\ m \notin modalityUpdates[t] - /\ modalityUpdates' = [modalityUpdates EXCEPT ![t] = @ \cup {m}] - /\ UNCHANGED <<txnStatus, crashes>> - --------------------------------------------------------------------------------- -\* Commit: only allowed once all 8 modalities have been applied. The commit -\* is itself atomic in the spec (single state transition); the Rust source is -\* expected to implement this via a 2-phase commit or equivalent mechanism. --------------------------------------------------------------------------------- -Commit(t) == - /\ txnStatus[t] = "PENDING" - /\ modalityUpdates[t] = Modalities - /\ txnStatus' = [txnStatus EXCEPT ![t] = "COMMITTED"] - /\ UNCHANGED <<modalityUpdates, crashes>> - --------------------------------------------------------------------------------- -\* Abort: unilateral rollback. Clears all applied updates and transitions to -\* ABORTED. Can fire at any time during PENDING. --------------------------------------------------------------------------------- -Abort(t) == - /\ txnStatus[t] = "PENDING" - /\ txnStatus' = [txnStatus EXCEPT ![t] = "ABORTED"] - /\ modalityUpdates' = [modalityUpdates EXCEPT ![t] = {}] - /\ UNCHANGED crashes - --------------------------------------------------------------------------------- -\* Crash during PENDING. Recovery must act as an abort: rollback applied -\* updates, transition to ABORTED. This is the adversarial event that the -\* Atomicity invariant is designed to defeat: without the rollback, a crash -\* mid-transaction would leave some modalities updated and the txnStatus -\* visible, i.e. a partial commit. --------------------------------------------------------------------------------- -Crash(t) == - /\ crashes < MaxCrashes - /\ txnStatus[t] = "PENDING" - /\ txnStatus' = [txnStatus EXCEPT ![t] = "ABORTED"] - /\ modalityUpdates' = [modalityUpdates EXCEPT ![t] = {}] - /\ crashes' = crashes + 1 - -Next == - \/ \E t \in Txns, m \in Modalities: ApplyModality(t, m) - \/ \E t \in Txns: Commit(t) - \/ \E t \in Txns: Abort(t) - \/ \E t \in Txns: Crash(t) - -\* Fairness: every transaction eventually resolves (either commits or aborts). -\* WF on (Commit \/ Abort) is sufficient because Abort is always enabled while -\* the transaction is PENDING, so the disjunction is continuously enabled. -Spec == Init - /\ [][Next]_vars - /\ \A t \in Txns: WF_vars(Commit(t) \/ Abort(t)) - --------------------------------------------------------------------------------- -\* Safety properties --------------------------------------------------------------------------------- - -\* I1. The central atomicity claim: -\* COMMITTED => all 8 modalities updated -\* ABORTED => no modalities updated -\* This is the "all-or-nothing" observation externally visible on an octad. -Atomicity == - \A t \in Txns: - /\ (txnStatus[t] = "COMMITTED" => modalityUpdates[t] = Modalities) - /\ (txnStatus[t] = "ABORTED" => modalityUpdates[t] = {}) - -\* I2. No observable partial state: once a transaction is no longer PENDING, -\* its modalityUpdates set is either empty or the full 8. This is derivable -\* from Atomicity but stated separately because it is the property consumers -\* of the octad actually rely on ("if I read a committed octad, I never see a -\* half-update"). -NoObservablePartial == - \A t \in Txns: - (txnStatus[t] \in {"COMMITTED", "ABORTED"}) => - (modalityUpdates[t] = Modalities \/ modalityUpdates[t] = {}) - -\* I3. Status monotonicity: PENDING can move to COMMITTED or ABORTED, but -\* never the reverse. (This is implicitly enforced by Next -- every action -\* that changes txnStatus requires it was PENDING -- but stating it as an -\* invariant on (txnStatus, txnStatus') is not trivial without action vars; -\* instead we check a weaker "never COMMITTED AND ABORTED" which is trivially -\* true by type but confirms no action flips between terminal statuses.) -StatusMonotone == - \A t \in Txns: - \neg (txnStatus[t] = "COMMITTED" /\ modalityUpdates[t] = {}) - -OctadSafe == - /\ TypeOK - /\ Atomicity - /\ NoObservablePartial - /\ StatusMonotone - --------------------------------------------------------------------------------- -\* Liveness --------------------------------------------------------------------------------- - -\* Every transaction eventually reaches a terminal state. -EveryTxnResolves == - \A t \in Txns: <>(txnStatus[t] \in {"COMMITTED", "ABORTED"}) - -THEOREM AtomicitySafety == Spec => []OctadSafe -THEOREM Resolution == Spec => EveryTxnResolves - -================================================================================ diff --git a/verisimdb/verification/proofs/tlaplus/README.adoc b/verisimdb/verification/proofs/tlaplus/README.adoc deleted file mode 100644 index b0ff3fd7..00000000 --- a/verisimdb/verification/proofs/tlaplus/README.adoc +++ /dev/null @@ -1,108 +0,0 @@ -// SPDX-License-Identifier: CC-BY-SA-4.0 -// Copyright (c) 2026 Jonathan D.A. Jewell (hyperpolymath) <j.d.a.jewell@open.ac.uk> - -= VeriSimDB TLA+ proofs -:toc: - -== V5: Octad transaction atomicity - -Specification: `OctadAtomicity.tla` + -Configuration: `OctadAtomicity.cfg` + -Source being verified: `rust-core/verisim-octad/src/transaction.rs` - -Models the VeriSimDB transaction lifecycle over an 8-modality octad -(`graph`, `vector`, `tensor`, `semantic`, `document`, `temporal`, -`provenance`, `spatial`) under an adversary that can crash a PENDING -transaction at any point. Safety properties: - -* `Atomicity` -- `COMMITTED` implies all 8 modalities updated; `ABORTED` - implies none. (The central all-or-nothing claim.) -* `NoObservablePartial` -- any terminal transaction's update set is - either the full 8 modalities or the empty set. -* `StatusMonotone` -- no terminal status ever reverts (paired with - Atomicity closes off "committed then rolled back" bugs). - -Liveness: `EveryTxnResolves` -- every transaction eventually reaches a -terminal status (COMMITTED or ABORTED), given weak fairness on the -`Commit \/ Abort` disjunction. - -== Running TLC - -Preferred: `just verify-tlaplus` from the verisimdb repo root. Works with -or without a host Java install (falls back to an ephemeral -`eclipse-temurin:21-jre` podman container). - -Manual invocation: - -[source,bash] ----- -cd verification/proofs/tlaplus/ -java -cp ~/.local/share/tla2tools.jar tlc2.TLC \ - -workers auto -config OctadAtomicity.cfg OctadAtomicity.tla ----- - -The default config bounds the model to 2 transactions and 2 crash -events. State space: 134,160 distinct states at depth 19, ~18s on an -8-core laptop. Verified 2026-04-17. - -== V9: Normalizer determinism + convergence - -Specification: `Normalizer.tla` + -Configuration: `Normalizer.cfg` + -Source being verified: `rust-core/verisim-normalizer/src/lib.rs` - -Models self-normalisation as a deterministic state transition: pick -the highest-priority modality as source, rewrite every modality to -the source's value. Safety properties: - -* `SourceIsMaximal` -- the `CHOOSE` in `SourceOf` always returns the - unique priority-maximum (determinism hinges on this). -* `NormalizeIdempotent` -- `Normalize(Normalize(s)) = Normalize(s)`. -* `PostStepNoDrift` -- after any normalisation round, drift is gone. -* `FixedPointStable` -- on a drift-free state, `Normalize` is identity. - -Liveness: `Convergence` -- `<>[]~HasDrift` (eventually-always drift-free). - -Default config: 4 modalities (`graph`, `vector`, `semantic`, -`document`), 3 distinct values, `MaxSteps=3`. 84 reachable states, -finishes in < 1s. Verified 2026-04-17. - -== V10: Transaction serializability - -Specification: `Serializability.tla` + -Configuration: `Serializability.cfg` + -Source being verified: `rust-core/verisim-octad/src/transaction.rs` (concurrent path) - -Extends V5's single-transaction atomicity to concurrent transactions -under atomic two-phase locking. Each transaction declares a fixed -read-set + write-set; `Begin(t)` atomically acquires the whole -access-set (or blocks), `Commit(t)` releases locks and appends to the -commit log -- which is, by construction, a valid serial schedule. - -Safety (composite `SerializabilitySafe`): - -* `NoSharedLocks` -- 2PL mutex invariant. -* `LocksOnlyWhileActive` -- IDLE/COMMITTED txns hold no locks. -* `ActiveHoldsFullAccessSet` -- no partial lock acquisitions. -* `CommitLogInjective` -- each txn commits at most once. -* `CommitLogSound` -- log entries are only COMMITTED txns. -* `NoConcurrentConflict` -- the central claim: two conflicting - transactions are never simultaneously ACTIVE, so every schedule is - conflict-equivalent to the commit-log order. - -Liveness: `EveryTxnCommits` -- every txn eventually commits (atomic -acquisition rules out deadlock and therefore starvation under WF). - -Default scenario: 3 transactions (t1, t2, t3) × 3 modalities (m1, m2, m3) -with every pair of txns conflicting on at least one modality. 31 -reachable states, depth 7, sub-second. Verified 2026-04-17. - -== Cross-references - -* `developer-ecosystem/standards/docs/proofs/spec-templates/T1-critical/verisimdb.md` -- V5, V9, V10 -* `/home/hyper/Desktop/proof-debt-plan.md` -- Dependability / VeriSimDB V5, V9, V10 - -== Remaining TLA+ work - -None at this tier. V5/V9/V10 close the TLA+-shaped VeriSimDB proof -obligations. V2/V3/V4/V6 stay Lean4-blocked; V11/V12 are Idris2-shaped. diff --git a/verisimdb/verification/proofs/tlaplus/Serializability.cfg b/verisimdb/verification/proofs/tlaplus/Serializability.cfg deleted file mode 100644 index 0d4c7198..00000000 --- a/verisimdb/verification/proofs/tlaplus/Serializability.cfg +++ /dev/null @@ -1,20 +0,0 @@ -\* SPDX-License-Identifier: MPL-2.0 -\* TLC configuration for Serializability.tla (V10). -\* -\* The scenario (Txns, Modalities, TxnReads, TxnWrites) is fixed at module -\* level because TLC config files cannot represent record literals. See the -\* comment block in Serializability.tla for the scenario choice rationale. -\* -\* Expected state space: 3 txns × 3 status values × 2^3 lock-sets × commit -\* log permutations. TLC prunes heavily via the Begin/Commit guards; actual -\* reachable ~hundreds of states. Sub-second run. - -SPECIFICATION Spec - -INVARIANTS - SerializabilitySafe - -PROPERTIES - EveryTxnCommits - -CHECK_DEADLOCK FALSE diff --git a/verisimdb/verification/proofs/tlaplus/Serializability.tla b/verisimdb/verification/proofs/tlaplus/Serializability.tla deleted file mode 100644 index 290b4884..00000000 --- a/verisimdb/verification/proofs/tlaplus/Serializability.tla +++ /dev/null @@ -1,202 +0,0 @@ ----------------------------- MODULE Serializability ---------------------------- -\* SPDX-License-Identifier: MPL-2.0 -\* Copyright (c) 2026 Jonathan D.A. Jewell (hyperpolymath) <j.d.a.jewell@open.ac.uk> -\* -\* V10: Conflict serializability for concurrent transactions on an octad. -\* Corresponds to rust-core/verisim-octad/src/transaction.rs (concurrent path). -\* -\* V5 proved atomicity for a single transaction under crashes. V10 extends -\* that to *multiple* concurrent transactions: the spec asserts that any -\* concurrent execution is conflict-equivalent to some serial execution of -\* the same transactions. -\* -\* Concurrency protocol modelled: atomic two-phase locking (2PL+atomic-acquire). -\* - Each transaction has a fixed access-set (reads \cup writes) declared at -\* design time (the TxnAccess constant). -\* - Begin(t) atomically acquires the full access-set, or blocks. -\* - Commit(t) releases locks and appends to the commit log. -\* This rules out deadlock by construction (atomic acquisition) and by the -\* same stroke makes the commit-log a valid serial schedule of the concurrent -\* execution. The spec's job is to verify both claims. - -EXTENDS Naturals, FiniteSets, Sequences, TLC - -\* Scenario is fixed at module level because TLC config files cannot express -\* record literals for TxnReads/TxnWrites, and this keeps a single-file spec -\* straightforwardly model-checkable. Alternative scenarios live as separate -\* .tla files. -\* -\* Scenario chosen: 3 transactions, 3 modalities, *partial* conflict graph. -\* This is deliberately richer than the all-pairs-conflict case, which would -\* trivially degenerate to a serial execution (at most 1 ACTIVE txn at any -\* time, so serializability is free). Here: -\* - t1 reads {m1}, writes {m1} -- disjoint from t2 (can run concurrently) -\* - t2 reads {m2}, writes {m2} -- disjoint from t1 (can run concurrently) -\* - t3 reads {m1, m2}, writes {m3} -- conflicts with BOTH t1 and t2 -\* -\* Consequence: the reachable state graph contains states where t1 and t2 -\* are simultaneously ACTIVE (real concurrency), but t3 must serialise -\* against each of them (real serialisation under 2PL). This scenario -\* genuinely exercises `NoConcurrentConflict` rather than trivially -\* satisfying it. - -Txns == {"t1", "t2", "t3"} -Modalities == {"m1", "m2", "m3"} - -TxnReads == [t \in Txns |-> - CASE t = "t1" -> {"m1"} - [] t = "t2" -> {"m2"} - [] t = "t3" -> {"m1", "m2"}] - -TxnWrites == [t \in Txns |-> - CASE t = "t1" -> {"m1"} - [] t = "t2" -> {"m2"} - [] t = "t3" -> {"m3"}] - -ASSUME Cardinality(Txns) >= 2 \* serializability is trivial for 1 txn -ASSUME Cardinality(Modalities) >= 1 -ASSUME TxnReads \in [Txns -> SUBSET Modalities] -ASSUME TxnWrites \in [Txns -> SUBSET Modalities] - -\* A transaction's access-set is everything it reads or writes. 2PL acquires -\* the whole set atomically at Begin. -AccessSet(t) == TxnReads[t] \cup TxnWrites[t] - -VARIABLES - txnStatus, \* [Txns -> {"IDLE", "ACTIVE", "COMMITTED"}] - holdsLocks, \* [Txns -> SUBSET Modalities] -- currently held - commitLog \* Seq(Txns) -- commit order (the serial schedule) - -vars == <<txnStatus, holdsLocks, commitLog>> - -Status == {"IDLE", "ACTIVE", "COMMITTED"} - -TypeOK == - /\ txnStatus \in [Txns -> Status] - /\ holdsLocks \in [Txns -> SUBSET Modalities] - /\ commitLog \in Seq(Txns) - -Init == - /\ txnStatus = [t \in Txns |-> "IDLE"] - /\ holdsLocks = [t \in Txns |-> {}] - /\ commitLog = << >> - --------------------------------------------------------------------------------- -\* Begin: atomically acquire all locks in AccessSet(t). Fires only if no other -\* transaction currently holds any lock in AccessSet(t). Atomic acquisition -\* is what rules out deadlock -- no partial lock sets ever exist. --------------------------------------------------------------------------------- -Begin(t) == - /\ txnStatus[t] = "IDLE" - /\ \A other \in Txns: - (other /= t) => (holdsLocks[other] \cap AccessSet(t) = {}) - /\ txnStatus' = [txnStatus EXCEPT ![t] = "ACTIVE"] - /\ holdsLocks' = [holdsLocks EXCEPT ![t] = AccessSet(t)] - /\ UNCHANGED commitLog - --------------------------------------------------------------------------------- -\* Commit: release all locks, mark COMMITTED, append to serial log. Only an -\* ACTIVE transaction (= one that has successfully acquired) can commit. --------------------------------------------------------------------------------- -Commit(t) == - /\ txnStatus[t] = "ACTIVE" - /\ holdsLocks[t] = AccessSet(t) - /\ txnStatus' = [txnStatus EXCEPT ![t] = "COMMITTED"] - /\ holdsLocks' = [holdsLocks EXCEPT ![t] = {}] - /\ commitLog' = Append(commitLog, t) - -Next == - \/ \E t \in Txns: Begin(t) - \/ \E t \in Txns: Commit(t) - -\* Fairness: every ACTIVE txn eventually commits, every IDLE txn eventually -\* begins (or has no opportunity because all other txns hold its locks -- -\* but atomic acquisition means some txn always makes progress, so an IDLE -\* txn cannot be permanently starved in this model). -Spec == Init /\ [][Next]_vars - /\ \A t \in Txns: WF_vars(Begin(t) \/ Commit(t)) - --------------------------------------------------------------------------------- -\* Safety properties --------------------------------------------------------------------------------- - -\* I1. Lock mutex: no two transactions hold overlapping locks simultaneously. -\* This is the 2PL invariant; its violation would be a direct serializability -\* failure. -NoSharedLocks == - \A t1, t2 \in Txns: - (t1 /= t2) => (holdsLocks[t1] \cap holdsLocks[t2] = {}) - -\* I2. Locks are held only by ACTIVE transactions. IDLE and COMMITTED txns -\* own no locks. -LocksOnlyWhileActive == - \A t \in Txns: - (txnStatus[t] \in {"IDLE", "COMMITTED"}) => (holdsLocks[t] = {}) - -\* I3. Lock-set matches the txn's AccessSet when ACTIVE. No partial acquisitions. -\* This is what atomic-Begin enforces structurally. -ActiveHoldsFullAccessSet == - \A t \in Txns: - (txnStatus[t] = "ACTIVE") => (holdsLocks[t] = AccessSet(t)) - -\* I4. No duplicate commit log entries -- each txn commits at most once. -CommitLogInjective == - \A i, j \in 1..Len(commitLog): - (commitLog[i] = commitLog[j]) => (i = j) - -\* I5. Commit log only contains COMMITTED txns. -CommitLogSound == - \A i \in 1..Len(commitLog): - txnStatus[commitLog[i]] = "COMMITTED" - -\* I6. CENTRAL serializability claim: for any two committed conflicting txns -\* (sharing a written modality), the one that appears earlier in the commit -\* log is the one that accessed first. Because 2PL prevents overlap, this is -\* always true -- the commit log IS a conflict-equivalent serial order. -\* -\* Formal statement: if t1 and t2 both appear in commitLog and they conflict -\* (t1 writes some m that t2 reads or writes, or vice versa), then their -\* relative order in commitLog is consistent with the (trivially unique) -\* sequential order in which they held the shared lock. Since only one txn -\* can hold a given lock at a time and locks are released at Commit, the txn -\* committed earlier necessarily ran earlier. So the log IS the schedule. -\* -\* Operationally this reduces to: two conflicting txns both committed implies -\* they don't appear "simultaneously" in the log -- which is trivially true -\* of a sequence. What we really want to check is that ACTIVE sets of -\* conflicting txns never co-exist. That is exactly NoSharedLocks when -\* combined with ActiveHoldsFullAccessSet -- two ACTIVE txns have disjoint -\* access-sets, so no WW / WR / RW conflict can be concurrent. -NoConcurrentConflict == - \A t1, t2 \in Txns: - (t1 /= t2 - /\ txnStatus[t1] = "ACTIVE" /\ txnStatus[t2] = "ACTIVE" - /\ ((TxnWrites[t1] \cap (TxnReads[t2] \cup TxnWrites[t2])) /= {} - \/ (TxnWrites[t2] \cap (TxnReads[t1] \cup TxnWrites[t1])) /= {})) - => FALSE - -SerializabilitySafe == - /\ TypeOK - /\ NoSharedLocks - /\ LocksOnlyWhileActive - /\ ActiveHoldsFullAccessSet - /\ CommitLogInjective - /\ CommitLogSound - /\ NoConcurrentConflict - --------------------------------------------------------------------------------- -\* Liveness --------------------------------------------------------------------------------- - -\* Every transaction eventually commits. Under atomic-acquire 2PL + WF, no -\* transaction can be starved forever: every state has at least one enabled -\* action (either some IDLE txn can Begin, since at most |Txns|-1 txns can -\* be ACTIVE at once -- and a fully ACTIVE state has at least one Commit -\* enabled; and after any Commit, freed locks re-enable some IDLE Begin). -EveryTxnCommits == - \A t \in Txns: <>(txnStatus[t] = "COMMITTED") - -THEOREM SerializabilitySafety == Spec => []SerializabilitySafe -THEOREM AllCommit == Spec => EveryTxnCommits - -================================================================================ diff --git a/verisimdb/verification/tests b/verisimdb/verification/tests deleted file mode 120000 index 6dd24e02..00000000 --- a/verisimdb/verification/tests +++ /dev/null @@ -1 +0,0 @@ -../tests \ No newline at end of file diff --git a/verisimdb/verisim-architecture-visualisation.html b/verisimdb/verisim-architecture-visualisation.html deleted file mode 100644 index f7d71065..00000000 --- a/verisimdb/verisim-architecture-visualisation.html +++ /dev/null @@ -1,230 +0,0 @@ -import React, { useState } from 'react'; -import { - Database, - Shield, - Activity, - Share2, - Cpu, - FileText, - Layers, - Zap, - Lock, - History, - CheckCircle2, - AlertCircle -} from 'lucide-react'; - -const App = () => { - const [activeStep, setActiveStep] = useState(0); - - const steps = [ - { - title: "1. The Heavy Handshake", - desc: "Client sends a sactify-php signed request. Controller Quorum (KRaft) verifies identity and registers the intent.", - icon: <Lock className="w-6 h-6" />, - color: "text-blue-500" - }, - { - title: "2. Consensus & Commitment", - desc: "The Leader proposes the Hexad update. Once a majority of the Quorum agrees, it's committed to the log.", - icon: <CheckCircle2 className="w-6 h-6" />, - color: "text-green-500" - }, - { - title: "3. Snapshotting & Truncation", - desc: "To keep the core 'tiny', the log is periodically collapsed into a snapshot. Old entries are discarded.", - icon: <Cpu className="w-6 h-6" />, - color: "text-purple-500" - }, - { - title: "4. Federated Routing", - desc: "WASM Proxy uses the 'Trust Window' (symmetric key) for high-speed routing to Rust modality stores.", - icon: <Zap className="w-6 h-6" />, - color: "text-yellow-500" - } - ]; - - return ( - <div className="min-h-screen bg-slate-50 p-8 font-sans"> - <div className="max-w-6xl mx-auto"> - <header className="mb-12 text-center"> - <h1 className="text-4xl font-bold text-slate-900 mb-4">VeriSimDB Architecture</h1> - <p className="text-xl text-slate-600">The "Tiny Core" Federated Knowledge Infrastructure</p> - </header> - - {/* High Level Structure */} - <div className="grid grid-cols-1 md:grid-cols-3 gap-8 mb-16"> - {/* THE CORE */} - <div className="bg-white p-6 rounded-2xl shadow-sm border border-slate-200"> - <div className="flex items-center gap-3 mb-4"> - <div className="p-2 bg-blue-100 rounded-lg"> - <Shield className="w-6 h-6 text-blue-600" /> - </div> - <h2 className="text-xl font-bold">The Registry Quorum</h2> - </div> - <ul className="space-y-3 text-slate-600 text-sm"> - <li className="flex items-center gap-2"><CheckCircle2 className="w-4 h-4 text-green-500"/> KRaft Consensus Layer</li> - <li className="flex items-center gap-2"><CheckCircle2 className="w-4 h-4 text-green-500"/> ReScript Global Namespace</li> - <li className="flex items-center gap-2"><CheckCircle2 className="w-4 h-4 text-green-500"/> Snapshot Management</li> - </ul> - </div> - - {/* THE ORCHESTRATION */} - <div className="bg-white p-6 rounded-2xl shadow-sm border border-slate-200"> - <div className="flex items-center gap-3 mb-4"> - <div className="p-2 bg-purple-100 rounded-lg"> - <Activity className="w-6 h-6 text-purple-600" /> - </div> - <h2 className="text-xl font-bold">Orchestrator</h2> - </div> - <ul className="space-y-3 text-slate-600 text-sm"> - <li className="flex items-center gap-2"><CheckCircle2 className="w-4 h-4 text-green-500"/> Elixir GenStage Pipelines</li> - <li className="flex items-center gap-2"><CheckCircle2 className="w-4 h-4 text-green-500"/> Drift Detection Engine</li> - <li className="flex items-center gap-2"><CheckCircle2 className="w-4 h-4 text-green-500"/> Trust Window Management</li> - </ul> - </div> - - {/* THE STORES */} - <div className="bg-white p-6 rounded-2xl shadow-sm border border-slate-200"> - <div className="flex items-center gap-3 mb-4"> - <div className="p-2 bg-orange-100 rounded-lg"> - <Share2 className="w-6 h-6 text-orange-600" /> - </div> - <h2 className="text-xl font-bold">Federated Stores</h2> - </div> - <ul className="space-y-3 text-slate-600 text-sm"> - <li className="flex items-center gap-2"><CheckCircle2 className="w-4 h-4 text-green-500"/> Rust Modality Crates</li> - <li className="flex items-center gap-2"><CheckCircle2 className="w-4 h-4 text-green-500"/> Hexad Modality Payloads</li> - <li className="flex items-center gap-2"><CheckCircle2 className="w-4 h-4 text-green-500"/> Local Sovereignty</li> - </ul> - </div> - </div> - - {/* Interactive Sequence Flow */} - <div className="bg-slate-900 rounded-3xl p-8 text-white shadow-xl overflow-hidden relative"> - <div className="absolute top-0 right-0 p-8 opacity-10"> - <Layers className="w-64 h-64" /> - </div> - - <h3 className="text-2xl font-bold mb-8">The Knowledge Lifecycle</h3> - - <div className="grid grid-cols-1 lg:grid-cols-2 gap-12 items-center"> - <div className="space-y-6"> - {steps.map((step, idx) => ( - <div - key={idx} - onClick={() => setActiveStep(idx)} - className={`p-4 rounded-xl cursor-pointer transition-all duration-300 border-l-4 ${ - activeStep === idx - ? 'bg-slate-800 border-white translate-x-2' - : 'bg-transparent border-transparent opacity-50 hover:opacity-100' - }`} - > - <div className="flex items-center gap-4"> - <div className={step.color}>{step.icon}</div> - <div> - <h4 className="font-bold">{step.title}</h4> - <p className="text-sm text-slate-400">{step.desc}</p> - </div> - </div> - </div> - ))} - </div> - - {/* Visual Logic Diagram */} - <div className="bg-slate-800 rounded-2xl p-6 border border-slate-700 min-h-[400px] flex flex-col justify-between"> - <div className="flex justify-between items-center border-b border-slate-700 pb-4"> - <span className="text-xs uppercase tracking-widest text-slate-500">System State</span> - <span className="text-xs px-2 py-1 bg-green-500/20 text-green-400 rounded">HEALTHY</span> - </div> - - <div className="flex-grow flex items-center justify-center relative py-12"> - {activeStep === 0 && ( - <div className="animate-pulse flex flex-col items-center"> - <Lock className="w-16 h-16 text-blue-400 mb-4" /> - <div className="h-1 w-32 bg-blue-400/30 rounded-full" /> - <p className="mt-4 text-blue-400 font-mono text-sm">VERIFYING_SIGNATURE...</p> - </div> - )} - {activeStep === 1 && ( - <div className="flex flex-col items-center"> - <div className="flex gap-4 mb-4"> - <Database className="w-10 h-10 text-green-400" /> - <div className="w-8 h-px bg-slate-600 mt-5 border-dashed border-t" /> - <Database className="w-10 h-10 text-green-400" /> - <div className="w-8 h-px bg-slate-600 mt-5 border-dashed border-t" /> - <Database className="w-10 h-10 text-green-400" /> - </div> - <p className="text-green-400 font-mono text-sm">MAJORITY_AGREED (2/3)</p> - </div> - )} - {activeStep === 2 && ( - <div className="flex flex-col items-center"> - <History className="w-16 h-16 text-purple-400 mb-4" /> - <div className="flex gap-1 mb-2"> - {[1,2,3,4,5].map(i => <div key={i} className={`w-2 h-4 bg-slate-600 ${i < 4 ? 'bg-purple-400/50' : ''}`} />)} - </div> - <p className="text-purple-400 font-mono text-sm">LOG_TRUNCATED_AT_0x4F2</p> - </div> - )} - {activeStep === 3 && ( - <div className="flex flex-col items-center w-full px-8"> - <div className="w-full flex justify-between mb-8"> - <div className="p-3 bg-slate-700 rounded-lg">WASM</div> - <div className="flex-grow flex items-center px-4"> - <div className="w-full h-px bg-yellow-400 relative"> - <Zap className="w-4 h-4 text-yellow-400 absolute left-1/2 -top-2 animate-bounce" /> - </div> - </div> - <div className="p-3 bg-slate-700 rounded-lg">RUST</div> - </div> - <p className="text-yellow-400 font-mono text-sm">FAST_PATH_ENABLED</p> - </div> - )} - </div> - - <div className="grid grid-cols-2 gap-4"> - <div className="p-3 bg-slate-900/50 rounded-lg border border-slate-700"> - <span className="block text-[10px] text-slate-500 uppercase">Registry Size</span> - <span className="text-sm font-mono text-slate-300">4.2 KB</span> - </div> - <div className="p-3 bg-slate-900/50 rounded-lg border border-slate-700"> - <span className="block text-[10px] text-slate-500 uppercase">Consensus Status</span> - <span className="text-sm font-mono text-slate-300">Quorum: OK</span> - </div> - </div> - </div> - </div> - </div> - - {/* Hexad Structure Diagram */} - <div className="mt-16 p-8 bg-white rounded-3xl border border-slate-200"> - <h3 className="text-2xl font-bold mb-8 text-slate-900">Logical Structure: The Hexad</h3> - <div className="grid grid-cols-2 md:grid-cols-3 lg:grid-cols-6 gap-4"> - {[ - { label: "GRAPH", type: "Oxigraph", icon: <Share2 className="w-5 h-5" /> }, - { label: "VECTOR", type: "HNSW", icon: <Zap className="w-5 h-5" /> }, - { label: "TENSOR", type: "ndarray", icon: <Cpu className="w-5 h-5" /> }, - { label: "SEMANTIC", type: "CBOR", icon: <FileText className="w-5 h-5" /> }, - { label: "DOCUMENT", type: "Tantivy", icon: <FileText className="w-5 h-5" /> }, - { label: "TEMPORAL", type: "Merkle", icon: <History className="w-5 h-5" /> }, - ].map((mod, i) => ( - <div key={i} className="flex flex-col items-center p-4 bg-slate-50 rounded-xl border border-slate-200 hover:bg-slate-100 transition-colors"> - <div className="p-2 bg-white rounded-lg shadow-sm mb-3 text-slate-700">{mod.icon}</div> - <span className="text-xs font-bold text-slate-900">{mod.label}</span> - <span className="text-[10px] text-slate-500 mt-1">{mod.type}</span> - </div> - ))} - </div> - <div className="mt-8 pt-8 border-t border-slate-100 flex justify-center"> - <div className="bg-blue-600 text-white px-6 py-3 rounded-full font-mono text-sm shadow-lg shadow-blue-200"> - UUID: 0x12AB...990F - </div> - </div> - </div> - </div> - </div> - ); -}; - -export default App;