Skip to content

Latest commit

 

History

761 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Areev

Build agents that improve from experience. Keep control of what changes.

Areev is an open-source runtime for governed, self-improving AI agents. Run workflows, assemble context with CAL, and turn execution history into proposed improvements you can evaluate, approve, and reverse.

Persistent memory supplies the evidence. Pseudonymization controls what sensitive data reaches models. Optional Jev decision models and integrations with your model-training pipeline extend the same improvement lifecycle.

CI License: MIT OR Apache-2.0

Quickstart · Examples · Documentation · Benchmarks · Discussions

Run → Record experience → Propose a change → Evaluate → Approve → Apply
 ↑                                                                │
 └──────────── Measure the next runs; keep or roll back ────────────┘

Try it

Install the CLI on macOS or Linux; no Rust toolchain required:

curl -fsSL https://raw.githubusercontent.com/AreevAI/areev/main/scripts/install.sh | sh
export PATH="$HOME/.local/bin:$PATH"

Windows binaries · Build from source or use Docker. Prebuilt Linux binaries require GLIBC 2.39+.

Seed a sample memory, generate improvement proposals, and open the review queue:

areev init --db areev-demo.db --template demo
areev loop run --db areev-demo.db
areev loop list --db areev-demo.db
areev ui --db areev-demo.db

Open localhost:7437 and select Suggestions. Inspect each recommendation and its evidence. To approve and apply changes from the console, enable authenticated review. This uses seeded data and deterministic analyzers: no account, model key, or external service needed.

Areev's review queue: improvement recommendations with evidence and Apply or Dismiss actions

The screenshot shows the richer committed demo memory. After cloning this repository, open it with areev ui --db data/demo.db --ns accounting.

One improvement cycle, connected capabilities

Capability What you can do
Run agents Execute workflows with branches, retries, subgraphs, budgets, and human approval steps. Resume interrupted runs and verify their journals. Triggers start workflows from schedules, events, or memory conditions.
Remember experience Keep facts, lessons, plans, and execution history in a content-addressed store. Trace changes, inspect past versions, and retrieve with structural, keyword, and vector search.
Engineer context with CAL Query memory and assemble model-ready context under a token budget, with priorities, summaries, and reusable queries.
Protect sensitive data Apply namespace policies for pseudonymization, redaction, masking, and generalization before storage or on model-facing output.
Improve under governance Detect patterns in history, propose changes with evidence, record review decisions, and measure outcomes after application. Roll back supported changes when they regress.
Add Jev decisions Use optional typed decision models for scoring, ranking, and supported workflow decisions, with deterministic fallbacks and authority enforced by code.
Govern model tuning Export training corpora, connect your trainer, evaluate the resulting adapter, and govern its promotion and rollback.

Use individual components with an existing agent, or use Areev Run for execution. Embed through Python, TypeScript/Node, or Rust; connect agent hosts through MCP. Local storage runs in-process in a Turso database file, with PostgreSQL available for server deployments.

See how the components fit together

Areev architecture: typed memory, context assembly, pseudonymization, model calls, governed execution, and an improvement loop with a model-tuning integration

Architecture · Why Areev

Pseudonymization and anonymization controls

Give models useful context while replacing detected sensitive values with typed placeholders such as [PERSON_1] and [EMAIL_1].

areev add --db private-demo.db --ns support \
  --subject caller:john --relation contact --object "j.doe@acme.io"

areev anonymize set --db private-demo.db --ns support \
  --policy '{"mode":"egress"}'

areev recall --db private-demo.db --ns support --subject caller:john
# Returned fields include subject: [PERSON_1], object: [EMAIL_1]
  • Egress policies transform model-facing reads, including recall, CAL, and MCP. Configured runtime and decision-model calls also use pseudonymization.
  • Ingress policies transform supported content before storage; audit mode measures detections without changing it.
  • Policy controls include category rules, known identities, custom terms, contextual rules, and optional detector integrations.
  • Local rehydration restores mapped values for applications and runtime tools. Mappings stay outside model payloads; a sealed vault supports reconstruction across processes.

Pseudonymization is reversible and does not by itself make data anonymous. Coverage depends on the configured detectors and policy. Ingress and sealed vaults require encrypted memory; runtime pseudonymization requires encrypted memory with scope: "memory" for replay-stable tokens.

Privacy recipes · Security boundaries · Runnable clinical-referral example

Does the improvement help?

In a synthetic support-desk benchmark, a fixed model performed 300 held-out tasks across three independent task streams. Applying lessons derived from its earlier failures improved task completion; removing those lessons removed the gain, and restoring them brought it back:

Memory state Tasks passed
Before lessons 46.3%
Lessons applied 68.7%
Lessons rolled back 47.0%
Lessons re-applied 67.0%

The lessons in this run came from deterministic analysis with zero model calls for lesson generation. These are results on a workload we designed, not a guarantee for other agents. The model still makes calls to perform tasks.

Results, controls, and caveats · Reproduce the experiment

The loop starts with deterministic analyzers. Optional LLM analysis adds proposals grounded in recorded evidence and checked by an independent verifier. Auto-apply is off unless explicitly granted by host policy; destructive and LLM-originated changes cannot auto-apply. Evaluation gates constrain what can be promoted. Hooks, cron, CI, or your host invoke the loop; Areev runs no daemon.

Context engineering with CAL

CAL is the Context Assembly Language. Combine memories from multiple sources, set priorities and budgets, and render the result for your model. Context can degrade from full content to a summary or omission as the budget fills. Saved queries and templates travel with the memory.

Try a budgeted query against the sample memory above:

areev cal --db areev-demo.db \
  'ASSEMBLE "deployment" FROM facts: (RECALL facts ABOUT "acme") BUDGET 800 tokens FORMAT sml'

This makes the agent's briefing inspectable and reusable across the CLI, libraries, MCP, and console. CAL reference · Saved queries and templates

Decision models, including Jev

Attach TypeSafe Jev or a compatible decision backend to supply typed scores, choices, and probabilities for supported judgments, including recall reranking. Backends are optional and provider-configurable; code retains control over permissions, approval, and application of changes.

In the recorded LoCoMo retrieval experiment, Jev reranking raised hit@1 from 18.6% to 51.8% over the same lexical-baseline candidate pool across 1,982 questions. This measures retrieval, not final-answer accuracy; hosted decisions add latency and cost. Egress policies also cover decision requests.

Configure decision backends · Measurements and costs

From experience to model tuning

Use recorded agent trajectories to build a training corpus, hand it to your trainer, and bring the resulting adapter back through evaluation and review:

Execution history → Governed corpus → Your trainer → Candidate adapter
                                                    ↓
                                   Evaluate → Approve → Promote / Roll back

areev corpus exports trajectories with step-level weights and lineage. areev tune --cmd connects an external trainer. Adapter promotion is checked against a pinned evaluation set and recorded gating run. Areev supplies the corpus and governance; training runs in your chosen training stack.

Tuning walkthrough · Adapter governance

Build a complete agent

Start with the invoice-to-accounting agent: process invoices, route approvals, accept corrections by email reply, and use recorded experience in later runs. The same agent ships in Python, TypeScript, and Rust.

Run its Python example on macOS or Linux:

git clone https://github.com/AreevAI/areev.git
cd areev
python3 -m venv .venv
source .venv/bin/activate
python -m pip install areev

examples/agents/invoice-to-accounting/python/smoke.sh
examples/agents/invoice-to-accounting/python/improve.sh

The first script runs the desk through invoices, approvals, and corrections. The second shows a correction helping the next invoice, analyzes run history, and records a person's decision on a proposed fix. Both use synthetic fixtures and mock integrations, with no model key or live mailbox required.

Example What to explore
Invoice processing Execution, approvals, corrections, and improvement from history
Clinical referrals Pseudonymized model requests and locally rehydrated results
Incident response Prior incident knowledge informing the next response
Revenue-cycle optimization A proposed CAL query revision, human approval, and rollback
Data-subject requests Disclosure and erasure using the same identity selector

All agent examples · How to build an Areev agent

Use Areev in your stack

Surface Install
Python pip install areev
TypeScript / Node npm install @areev/areev
Rust libraries cargo add areev-store areev-core
CLI, console, and MCP server Prebuilt installer above, or cargo install areev

The Python and Node packages embed the engine; install the CLI separately for areev ui and the MCP server. With the CLI installed, connect Claude Code:

claude mcp add areev -- areev serve --mcp --db ~/.areev/code.db --ns claude-code

Language quickstarts · MCP reference · Docker deployment · Migration

Documentation

Guide Covers
Quickstart · Cookbook Installation, bindings, and task recipes
Run · Triggers · Packs Execution, scheduling, and shipping agents
Loop · CAL Improvement, evaluation, and context assembly
Security model · Erasure Pseudonymization, encryption, authorization, and deletion
GDPR map · EU AI Act map Requirements mapped to capabilities and their limits
Architecture · Grain types Storage, data types, and design decisions
Benchmarks · Quality Performance, experiments, and verification
FAQ · Changelog Common questions and releases

Local recall runs in-process with no server in its path. Published benchmarks cover desktop and Raspberry Pi hardware; hosted model and decision calls have separate costs. Areev collects no telemetry. Optional encryption at rest, retention policies, legal holds, and subject-level erasure support the data lifecycle alongside pseudonymization.

Repository quality and coverage Repository quality metrics generated from source and test counts Source line coverage by crate, compared with each crate's enforced floor

CI checks metric drift, per-crate coverage floors, backend conformance, and documented CAL examples. How the numbers are produced. The memory format and CAL conform to OMS. Areev uses Turso for embedded storage; see third-party notices.

Community and contributing

Share what you're building in Discussions or r/Areev. Contributions use the DCO; start with CONTRIBUTING.md and the Code of Conduct. Report security issues through SECURITY.md.

If governed self-improvement would help your agents, star Areev to help other developers discover it.

License

Dual-licensed under Apache-2.0 or MIT, at your option. Contributions are dual-licensed under the same terms unless explicitly stated otherwise. The OMS specification is CC0.

Built and backed by MindGryd Software Private Limited.

About

Self-improving agents, governed. Areev is the substrate for adaptive agents — agents that get better from their own history, under human authority, in steps you can inspect, undo, and re-measure.

Topics

Resources

Code of conduct

Contributing

Security policy

Stars

24 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages