Skip to content

Repository files navigation

PolicyFuzz

Open-source policy fuzzing and adversarial testing framework for AI agents.

PolicyFuzz tests whether an AI agent can reach outcomes that its authorization, safety, and data-flow policies are supposed to prevent.

The core idea is simple:

Do not only test whether individual tools are allowed. Test whether combinations of allowed actions can produce a forbidden outcome.

PolicyFuzz is a local-first, developer-friendly, reproducible testing system. It runs in CI, produces machine-readable findings (JSON, SARIF v2.1.0), and provides minimal reproducible counterexamples for every confirmed policy violation.

Project status

PolicyFuzz is a working implementation, not only a design project. The following are implemented, tested (895 tests), and verified against the documented behavior:

  • versioned policy syntax (v1) with six rule kinds and fail-closed evaluation;
  • a deterministic in-memory simulator with resource limits and replayable traces;
  • bounded deterministic breadth-first sequence search with state deduplication and an optional risk-guided priority hook;
  • finding lifecycle: stable IDs, evidence, greedy sequence minimization, deterministic reproduction, terminal and JSON renderers;
  • CLI: init, validate, test, reproduce, version with stable exit codes 0/1/2/3;
  • machine-readable CI output: JSON result envelope and SARIF v2.1.0;
  • two tool-protocol adapters (MCP 2026-07-28, OpenAI function tools) that translate external tool definitions into the domain model without any network or SDK surface;
  • synthetic policy packs (customer support, finance, DevOps, sensitive-data access) with vulnerable and fixed variants.

Features described as "future" or "conceptual" in the design documents are not implemented; see "Unsupported features" below. The implementation follows the roadmap in docs/ROADMAP.md.

At a glance

A policy declares a security property:

{
  "version": 1,
  "id": "refund-exposure",
  "description": "Cumulative refunds above 100 per customer are forbidden.",
  "rules": [
    {
      "id": "refund-exposure",
      "description": "Cumulative refunds above 100 per customer are forbidden.",
      "kind": "cumulative_limit",
      "action": "refund",
      "field": "amount",
      "limit": 100,
      "scope_argument": "customer_id"
    }
  ]
}

policyfuzz test searches bounded action sequences, executes them in the deterministic simulator, evaluates the policy after every step, and reports a minimized, reproducible counterexample when a sequence violates it — here, two 60-unit refunds totalling 120 against the 100 cap.

Core principles

  1. Policy-first — security requirements are expressed as executable properties.
  2. Outcome-oriented — test forbidden outcomes, not only individual tool permissions.
  3. Compositional — analyze sequences and combinations of actions.
  4. Reproducible — every finding must have a deterministic or replayable reproduction.
  5. Local-first — sensitive agent configurations and test data should not need to leave the developer environment.
  6. Framework-agnostic core — integrations belong at the boundary.
  7. Open source first — the complete core remains available under the repository's chosen open-source license.
  8. Evidence over scores — findings must explain what happened and why the policy was violated.
  9. Fail safely — unknown behavior must never be silently classified as safe.
  10. No false claims — passing a test suite does not prove an agent is universally secure.

Intended users

Primary

Developers building AI agents that can call tools or take actions.

Secondary

Security engineers, AppSec engineers, platform engineers, and teams responsible for deploying multiple AI agents.

Not the primary audience

Developers building ordinary applications without autonomous or semi-autonomous tool execution.

Initial supported problem

Given:

  • an agent or agent-compatible tool interface,
  • a set of tool definitions,
  • a policy,
  • a state model or sandbox,
  • and synthetic/test data,

generate adversarial action sequences and determine whether any sequence violates a declared security property.

Supported features

  • Policy syntax v1 (JSON, or YAML when PyYAML is installed): six rule kinds — forbidden_action, approval_required, sensitive_data_flow, cumulative_limit, sequence_restriction, capability_combination — with strict unknown-key rejection, two-stage validation, and deterministic composition DENY > ERROR > REQUIRE_APPROVAL > INCONCLUSIVE > UNKNOWN > ALLOW. See docs/POLICY_SPEC.md and the policyfuzz.policy module docstring.
  • Deterministic simulator: 16 synthetic in-memory tools (no shell, no network), explicit seed, resource limits (max_steps, max_tool_calls, timeout_seconds, max_outputs_bytes), and byte-replayable traces via Simulator.replay.
  • Bounded BFS search: explicit candidate templates, state deduplication, seeded ordering, optional priority hook for risk-guided discovery order, and graceful incomplete reporting.
  • Findings: stable PF-000001-style IDs, structured evidence, greedy 1-minimal sequence minimization, deterministic reproduction, terminal and JSON renderers.
  • CLI: init, validate, test, reproduce, version; JSON and SARIF v2.1.0 output; --output artifact writing; --fail-on-inconclusive; --no-minimize.
  • Adapters: MCP 2026-07-28 (full tools/list + tools/call, content annotations, x-mcp-header, outputSchema validation, structuredContent validation, ttlMs/cacheScope caching) and OpenAI function tools (flat + nested shapes, strict mode validation, output_schema validation) — translation only, no transport, no SDK.
  • Provider adapters (src/policyfuzz/providers/): Anthropic Claude, Google Gemini, OpenAI, and OpenAI-compatible providers (Ollama, vLLM, LiteLLM, Groq) — each translates native tool/function definitions into the internal domain model with contract tests and deterministic local fixtures.
  • Policy packs: four synthetic vulnerable+fixed packs under fixtures/packs/.
  • CI: exit-code contract, JSON envelope, SARIF, GitHub Actions starter (examples/ci/policyfuzz-ci.yml) and an active regression workflow.

Unsupported features

Not implemented (do not assume they exist):

  • policyfuzz explain or any LLM-assisted explanation, authoring, or intent generation;
  • natural-language policy authoring — policies are structured documents only;
  • live tool transports: no HTTP, stdio, SSE, OAuth, or SDK client anywhere; adapters translate vendored JSON snapshots only;
  • real tool execution: the simulator is synthetic and in-memory; running real tools/servers requires isolation outside PolicyFuzz;
  • pagination, elicitation, tasks, sampling, or stateful sessions from MCP;
  • non-function OpenAI tool types, async execution, deferred tool_search loading, allowed_callers enforcement, orchestration;
  • rule kinds beyond the four v1 kinds (no sequence-restriction, capability-combination, or trust-boundary rule kinds — see docs/POLICY_SPEC.md for the conceptual categories);
  • property-based search, symbolic execution, constraint solving, distributed search;
  • JUnit, Markdown, or HTML reporters (terminal, JSON, and SARIF only);
  • hosted dashboards, cloud services, billing, runtime enforcement, or production exploitation.

Limitations

  • Search is bounded and incomplete by default. BFS explores the space of explicitly supplied candidate templates up to max_steps/max_cases. A PASS means "no violation found within the bound", never "no violation exists". Truncated runs report INCOMPLETE (exit 3), never PASS.
  • Only the six v1 rule kinds are evaluated. Policies over behaviors the state model cannot express (e.g. cross-agent flows, timing) are out of scope.
  • History-sensitive policies depend on state encoding. Deduplication fingerprints cover resources, data assets, identities, approvals, balances, data locations, trust levels, and unknown aspects — not the event log. A policy whose verdict depends on history invisible in state would need a fingerprint change.
  • The simulator is synthetic. Findings demonstrate policy violations against the modeled environment, not against a live system.
  • Adapters are translation layers, not sandboxes. A passing adapter run only means the vendored definitions translated and the local fixture behaved; it says nothing about a live server or model.
  • Terminal findings include a fixed generic "Why it matters" sentence. It is not derived from the finding's evidence; treat it as a placeholder for a future evidence-derived explanation.
  • SARIF physicalLocation points at the tested policy file (line 1); precise per-step source ranges are not emitted.
  • YAML support (PyYAML is a required dependency); JSON also works everywhere.
  • Performance: BFS is exponential in depth by design; trace serialization is O(n^2) in trace length; minimization is O(n^2) executions. See docs/PERFORMANCE_AUDIT.md.

Deterministic behavior

Given the same policy, initial state, candidate pool, and seed:

  • candidate order is fixed (canonical sort + one seeded shuffle, or priority tiers + seeded shuffle when a priority hook is supplied);
  • the simulator reproduces traces byte-for-byte (to_json equality) from a serialized TestCase (seed + initial state + actions);
  • finding IDs are a pure function of an explicit counter;
  • JSON and SARIF output use sorted keys for byte-stable artifacts;
  • repeated policyfuzz test runs produce byte-identical finding files.

Wall-clock timeouts (timeout_seconds) and real-time sleeps are the only nondeterministic elements, and only when explicitly configured.

Incomplete search

A run is completed only when the bounded space was exhaustively explored within max_steps/max_cases. Hitting a bound stops the run with completed: false and RESOURCE_LIMIT; the CLI reports INCONCLUSIVE and exits 3 — never a silent pass. --fail-on-inconclusive additionally escalates a completed run that saw inconclusive evaluations to exit 3.

Security boundaries

  • No network or subprocess surface in the core. The domain, policy, simulator, search, finding, and SARIF layers import the standard library plus policyfuzz.domain only; architecture tests enforce this. The adapters contain no transport code.
  • The simulator never executes real tools. All tool behavior is synthetic and in-memory; shell/network-sounding tool ids map to UNKNOWN_TOOL/UNSUPPORTED_TOOL, never to execution.
  • Untrusted input is data. Tool metadata, descriptions, and fixture content are preserved as data, never executed or sent to an LLM.
  • Resource limits are enforced (steps, tool calls, wall-clock, output size) and oversized outputs are rejected before allocation.
  • Findings contain only modeled data — whatever you put in fixtures. Use synthetic data; never place real credentials or PII in arguments, definitions, or issues.
  • PolicyFuzz is a testing tool, not a runtime enforcer. A passing run does not certify that an agent is secure; it tests the modeled environment within the configured bounds. See SECURITY.md and docs/THREAT_MODEL.md.

Non-goals

PolicyFuzz is not:

  • a generic vulnerability scanner;
  • an LLM observability platform;
  • a production runtime firewall;
  • an agent hosting platform;
  • a compliance certification service;
  • a guarantee that an agent is secure;
  • a replacement for human security review;
  • a general-purpose autonomous penetration-testing platform.

Documentation map

  • docs/PRD.md — product requirements and boundaries.
  • docs/PRODUCT_SPEC.md — functional behavior and user-facing semantics.
  • docs/POLICY_SPEC.md — policy language, rule kinds, and evaluation model.
  • docs/ARCHITECTURE.md — system architecture.
  • docs/DOMAIN_MODEL.md — core entities and relationships.
  • docs/THREAT_MODEL.md — security assumptions, threats, and abuse boundaries.
  • docs/TESTING_STRATEGY.md — testing philosophy and required test layers.
  • docs/ROADMAP.md — development roadmap and status.
  • docs/DEVELOPMENT_GUIDE.md — implementation conventions.
  • docs/DECISIONS.md — architecture decision records (ADR-0001 … ADR-0012).
  • docs/RELEASE.md — release and versioning rules.
  • docs/CI.md — CI exit codes, JSON envelope, and SARIF contract.
  • docs/MCP_ADAPTER.md — MCP adapter capabilities and limitations.
  • docs/OPENAI_ADAPTER.md — OpenAI function-tools adapter capabilities and limitations.
  • docs/PERFORMANCE_AUDIT.md — measured performance characteristics and known complexity.
  • SECURITY.md — security reporting and secure development expectations.
  • CONTRIBUTING.md — contribution process.

Repository structure

policyfuzz/
├── README.md
├── LICENSE
├── CONTRIBUTING.md
├── SECURITY.md
├── pyproject.toml
├── docs/                     # design documents, ADRs, adapter docs
├── src/policyfuzz/
│   ├── domain.py             # immutable domain model (schema_version 1)
│   ├── policy.py             # policy syntax v1: parse, validate, evaluate
│   ├── simulator.py          # deterministic in-memory execution environment
│   ├── search.py             # bounded deterministic BFS over candidates
│   ├── finding.py            # finding lifecycle, minimization, reproduction
│   ├── sarif.py              # SARIF v2.1.0 reporter
│   ├── project.py            # project config/fixture loading (CLI boundary)
│   ├── cli.py                # init/validate/test/reproduce/version
│   ├── adapters/
│   │   ├── mcp.py            # MCP tool adapter (translation only)
│   │   └── openai_tools.py   # OpenAI function-tools adapter (translation only)
│   └── providers/            # provider-specific tool adapters
│       ├── anthropic.py      # Anthropic Claude tool-use adapter
│       ├── gemini.py         # Google Gemini function-calling adapter
│       ├── openai.py         # OpenAI function-tools adapter
│       ├── openai_compatible.py  # base for OpenAI-compatible providers
│       ├── ollama.py         # Ollama (OpenAI-compatible)
│       ├── vllm.py           # vLLM (OpenAI-compatible)
│       ├── litellm.py        # LiteLLM (OpenAI-compatible)
│       └── groq.py           # Groq (OpenAI-compatible)
├── tests/                    # 895 tests (unit, contract, integration, packs)
├── examples/
│   ├── quickstart/           # six-step vulnerable+fixed refund fixture
│   └── ci/                   # GitHub Actions starter + local walkthrough
├── fixtures/
│   ├── packs/                # four synthetic policy packs (vulnerable+fixed)
│   ├── policies/             # valid/invalid/boundary policy fixtures
│   ├── search/               # search fixtures + README
│   ├── mcp/                  # MCP adapter fixtures
│   └── openai/               # OpenAI adapter fixtures
└── benchmarks/bench.py       # reproducible performance benchmarks

Development philosophy

Prefer a small, composable core over a large framework-specific implementation.

Every major feature should answer:

  1. What user problem does it solve?
  2. What invariant does it preserve?
  3. What is the smallest useful implementation?
  4. How is it tested?
  5. How is it reproduced?
  6. What happens when information is unknown?
  7. Does it introduce a new security boundary?
  8. Does it belong in the core or an adapter?

Quality bar

A feature is not complete merely because its happy path works.

A feature is complete when:

  • behavior is specified;
  • unit tests exist;
  • integration tests exist where applicable;
  • error cases are covered;
  • security implications are reviewed;
  • public interfaces are documented;
  • deterministic reproduction is possible where relevant;
  • no unrelated behavior regresses.

Development quickstart

Requires Python 3.10+.

pip install -e ".[dev]"
policyfuzz --help
policyfuzz version
python -m policyfuzz version
pytest
ruff check src tests
ruff format --check src tests

The version output is deterministic: policyfuzz 0.1.0.

First workflow (local, no dashboard)

policyfuzz init ./my-project
cd ./my-project
policyfuzz validate   # exit 0 when the project is valid
policyfuzz test       # exit 1: seeded violation (two 60-unit refunds exceed 100)
ls findings/          # PF-000001.json
policyfuzz reproduce findings/PF-000001.json --policy policy.json  # exit 0

Fix the policy (raise the cumulative_limit cap) and rerun: policyfuzz test then prints Result: PASS and exits 0.

Exit codes: 0 no violations, 1 confirmed violation (or reproduce mismatch), 2 configuration/usage error, 3 execution error or incomplete search (never a silent pass).

Machine-readable CI output:

policyfuzz test --format json --output results.json   # stable envelope, exit 0/1/2/3
policyfuzz test --format sarif --output results.sarif # SARIF v2.1.0, pass/fail only

See examples/quickstart/README.md for the full six-step fixture, examples/ci/README.md + examples/ci/policyfuzz-ci.yml for the GitHub Actions starter (install, validate, fail on violation, pass when fixed -- no credentials or external services), and docs/CI.md for the JSON/SARIF schema and exit-code contract.

License

This project is licensed under the MIT License. See LICENSE.

About

Adversarial policy testing for AI agents: deterministic bounded search, policy evaluation, and reproducible findings.

Topics

Resources

Contributing

Security policy

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages