Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
196 changes: 146 additions & 50 deletions README.md
Original file line number Diff line number Diff line change
@@ -1,77 +1,173 @@
# OpenDecision

An auditable, typed decision layer for AI agents.
**A framework-agnostic decision layer for AI agents.**

OpenDecision places explicit contracts, deterministic rules, provider fallback, human review, and privacy-preserving audit traces between an agent and consequential actions. It is a control-layer library, not a safety guarantee.
OpenDecision places a typed, testable contract between an AI agent and the action it wants to take. It uses [Laya](https://github.com/NandhaKishorM/laya) for fast `choice`, `score`, and `noul` decisions without adding another text-generation step.

## Production-foundation preview
> Status: early MVP. Do not treat uncalibrated model confidence as a production safety guarantee.

Version 0.2 adds:
## Why OpenDecision?

- strict decision-contract and provider-output validation;
- deterministic rules plus ordered provider fallback chains;
- fail-safe `raise`, `review`, and `block` behavior;
- policy IDs, semantic versions, risk levels, defaults, and inheritance;
- context hashes, decision IDs, provider attempts, and JSONL audit sinks;
- an in-memory human-review workflow with loop protection;
- reproducible accuracy, confusion, and Brier-score evaluation primitives;
- coverage, lint, strict typing, and Python 3.10–3.13 CI gates.
Agents repeatedly need to decide whether to continue, stop, call a tool, ask a person, retrieve more context, or escalate. OpenDecision makes those branches explicit and reusable:

## Example
- typed decision contracts instead of free-form generated text;
- normalized results across decision providers;
- community-maintained YAML decision packs;
- adapters that do not couple the core SDK to one agent framework;
- optional Laya loading, so importing OpenDecision never downloads a model.

```python
from opendecision import (
DecisionContext,
DecisionGuard,
JsonlAuditSink,
ProviderChain,
Rule,
RuleProvider,
)
## Install

rules = RuleProvider(
[Rule(field="command", operator="contains", value="rm -rf", decision="block")],
)
providers = ProviderChain([rules], terminal_answer={"choice": "review"})
```bash
pip install -e ".[laya]"
```

guard = DecisionGuard(
providers,
audit_sink=JsonlAuditSink("audit/decisions.jsonl"),
failure_mode="review",
)
To enable the Rich-powered CLI output:

```bash
pip install -e ".[rich]"
```

Laya and OpenDecision require Python 3.10 or newer. For the examples:

```bash
pip install -e ".[laya,langgraph,fastapi]"
```

## Quick start

```python
from opendecision import DecisionGuard

guard = DecisionGuard()
result = guard.decide(
{"command": "rm -rf /tmp/cache"},
{
state={"tool": "send_email", "recipient": "customer@example.com"},
question={
"type": "choice",
"instructions": "Should this command execute?",
"instructions": "Should the agent execute this tool action?",
"options": {
"allow": "Authorized and low risk",
"review": "Needs human approval",
"block": "Unauthorized or destructive",
"allow": "Proceed automatically",
"review": "Require human approval",
"block": "Stop the action",
},
},
context=DecisionContext(
actor="agent:ops",
tool="shell",
action_id="run-123",
policy_id="shell-command-risk",
policy_version="1.0.0",
risk="critical",
),
)

print(result.decision)
print(result.probabilities)
print(result.confidence)
```

OpenDecision returns model decisions and evidence; it does **not** ask Laya to generate a reason. Applications may add an explanation separately without confusing generated prose with decision-model output.

## CLI

Evaluate a decision pack against JSON state:

```bash
opendecision evaluate decision_packs/security/tool_risk.yaml --state '{"tool": "send_email"}'
```

Raw state is hashed for correlation and is not written by the built-in audit sinks. Applications remain responsible for authentication, authorization, durable review storage, encryption, retention, monitoring, calibration, and incident response.
Or render JSON for scripting:

```bash
opendecision evaluate decision_packs/security/tool_risk.yaml --state state.json --json
```

Rich output is enabled automatically when `rich` is installed.

## Decision packs

```python
from opendecision import DecisionGuard, load_policy

policy = load_policy("decision_packs/security/tool_risk.yaml")
result = DecisionGuard().decide(agent_state, policy.question)
```

The first pack includes a tool-risk contract with `allow`, `review`, and `block` outcomes.

## Tool protection decorator

```python
from opendecision import DecisionGuard, load_policy
from opendecision.tool_guard import protect_tool

policy = load_policy("decision_packs/security/tool_risk.yaml")
guard = DecisionGuard()

@protect_tool(guard, policy.question)
def send_email(to: str, body: str) -> None:
...
```

The decorator fails closed:
- `allow` executes the tool
- `review` raises `ToolReviewRequiredError`
- anything else raises `ToolBlockedError`

## LangGraph

```python
from opendecision import DecisionGuard
from opendecision.integrations.langgraph import decision_node, route_by_decision

graph.add_node("tool_guard", decision_node(DecisionGuard(), question))
graph.add_conditional_edges(
"tool_guard",
route_by_decision(),
{"allow": "execute_tool", "review": "human_review", "block": "stop"},
)
```

## FastAPI

```bash
uvicorn app.main:app --reload
```

Then send `POST /v1/decide` with `state` and a typed `question`.

## Architecture

```text
Agent / application
|
DecisionGuard + contract
|
Provider adapter (Laya first)
|
Normalized DecisionResult
|
ALLOW | REVIEW | BLOCK
```

## Scope of v0.1

- Python SDK and Pydantic contracts
- lazy Laya provider
- choice, score, and noul normalization
- YAML decision packs
- LangGraph node and routing helpers
- FastAPI example
- unit tests that run without downloading model weights

A dashboard, TypeScript SDK, additional frameworks, model-generated explanations, and evaluation tooling are intentionally deferred.

## Calibration and safety

Laya's documentation notes meaningful limitations: base checkpoints can be weak zero-shot on typed workflows, confidence may require domain calibration, and high-cardinality choices need special handling. Benchmark decision packs on representative data before automating consequential actions. Prefer human review when evidence or authorization is insufficient.

## Development

```bash
pip install -e ".[dev]"
coverage run -m pytest
coverage report
pytest
ruff check .
mypy opendecision
```

See `docs/production-readiness.md` for operating guidance.
See [CONTRIBUTING.md](CONTRIBUTING.md) for contribution guidelines.

## License

Apache License 2.0. Laya is a separate Apache-2.0 project and remains subject to its own license and notices.
30 changes: 8 additions & 22 deletions opendecision/__init__.py
Original file line number Diff line number Diff line change
@@ -1,33 +1,19 @@
"""OpenDecision: an auditable typed decision layer for AI agents."""
"""OpenDecision: a typed decision layer for AI agents."""

from .audit import AuditEvent, JsonlAuditSink, MemoryAuditSink
from .evaluation import EvaluationCase, EvaluationReport, evaluate
from .guard import DecisionGuard
from .models import DecisionContext, DecisionQuestion, DecisionResult
from .models import DecisionQuestion, DecisionResult
from .policies import DecisionPolicy, load_policy
from .providers import LayaProvider, ProviderChain, Rule, RuleProvider
from .review import InMemoryReviewStore, ReviewRequest, ReviewStatus
from .tool_guard import ToolBlockedError, ToolReviewRequiredError, protect_tool

__all__ = [
"AuditEvent",
"DecisionContext",
"DecisionGuard",
"DecisionPolicy",
"DecisionQuestion",
"DecisionResult",
"EvaluationCase",
"EvaluationReport",
"InMemoryReviewStore",
"JsonlAuditSink",
"LayaProvider",
"MemoryAuditSink",
"ProviderChain",
"ReviewRequest",
"ReviewStatus",
"Rule",
"RuleProvider",
"evaluate",
"DecisionPolicy",
"ToolBlockedError",
"ToolReviewRequiredError",
"load_policy",
"protect_tool",
]

__version__ = "0.2.0"
__version__ = "0.1.0"
13 changes: 13 additions & 0 deletions opendecision/__main__.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,13 @@
"""Package entrypoint for `python -m opendecision`."""

from __future__ import annotations

from .cli import run


def main() -> None:
raise SystemExit(run())


if __name__ == "__main__":
main()
111 changes: 111 additions & 0 deletions opendecision/cli.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,111 @@
"""Command-line interface for evaluating decision packs.

Rich is an optional dependency. When available, output is rendered with tables.
"""

from __future__ import annotations

import argparse
import json
from pathlib import Path
from typing import Any, Callable, Mapping

from .guard import DecisionGuard
from .models import DecisionQuestion, DecisionResult
from .policies import load_policy


def build_parser() -> argparse.ArgumentParser:
parser = argparse.ArgumentParser(prog="opendecision", description="Evaluate OpenDecision policies")
subparsers = parser.add_subparsers(dest="command", required=True)

evaluate = subparsers.add_parser(
"evaluate",
help="Evaluate a decision-pack YAML policy against JSON state",
)
evaluate.add_argument("policy", help="Path to a policy YAML file")
evaluate.add_argument(
"--state",
required=True,
help="JSON string or path to a JSON file containing agent state",
)
evaluate.add_argument(
"--json",
action="store_true",
help="Render output as JSON (useful for scripting)",
)

return parser


def _load_state(value: str) -> Mapping[str, Any] | str:
path = Path(value)
if path.exists():
return json.loads(path.read_text(encoding="utf-8"))
return json.loads(value)


def _render_plain(result: DecisionResult) -> str:
lines = [
f"decision: {result.decision}",
f"provider: {result.provider}",
f"confidence: {result.confidence}",
]
if result.probabilities:
lines.append("probabilities:")
for key, value in sorted(result.probabilities.items()):
lines.append(f" - {key}: {value}")
lines.append("limitations: Model confidence is uncalibrated; prefer review when uncertain.")
return "\n".join(lines)


def _render_rich(result: DecisionResult) -> str | None:
try:
from rich.console import Console
from rich.table import Table
except ImportError:
return None

console = Console()
table = Table(title="OpenDecision")
table.add_column("Field")
table.add_column("Value")
table.add_row("decision", str(result.decision))
table.add_row("provider", result.provider)
table.add_row("confidence", str(result.confidence))
console.print(table)

if result.probabilities:
probs = Table(title="Probabilities")
probs.add_column("label")
probs.add_column("p")
for label, prob in sorted(result.probabilities.items(), key=lambda item: item[0]):
probs.add_row(str(label), str(prob))
console.print(probs)

console.print(
"[dim]limitations: Model confidence is uncalibrated; prefer review when uncertain.[/dim]"
)
return "" # indicate rich rendered


def run(argv: list[str] | None = None, *, guard_factory: Callable[[], DecisionGuard] = DecisionGuard) -> int:
parser = build_parser()
args = parser.parse_args(argv)

if args.command != "evaluate":
parser.error(f"Unknown command: {args.command}")

policy = load_policy(args.policy)
state = _load_state(args.state)
guard = guard_factory()
result = guard.decide(state, policy.question)

if args.json:
print(result.model_dump_json(indent=2))
return 0

rendered = _render_rich(result)
if rendered is None:
print(_render_plain(result))
return 0
Loading
Loading