Open-source policy fuzzing and adversarial testing framework for AI agents.
PolicyFuzz tests whether an AI agent can reach outcomes that its authorization, safety, and data-flow policies are supposed to prevent.
The core idea is simple:
Do not only test whether individual tools are allowed. Test whether combinations of allowed actions can produce a forbidden outcome.
PolicyFuzz is a local-first, developer-friendly, reproducible testing system. It runs in CI, produces machine-readable findings (JSON, SARIF v2.1.0), and provides minimal reproducible counterexamples for every confirmed policy violation.
PolicyFuzz is a working implementation, not only a design project. The following are implemented, tested (895 tests), and verified against the documented behavior:
- versioned policy syntax (v1) with six rule kinds and fail-closed evaluation;
- a deterministic in-memory simulator with resource limits and replayable traces;
- bounded deterministic breadth-first sequence search with state deduplication and an optional risk-guided priority hook;
- finding lifecycle: stable IDs, evidence, greedy sequence minimization, deterministic reproduction, terminal and JSON renderers;
- CLI:
init,validate,test,reproduce,versionwith stable exit codes0/1/2/3; - machine-readable CI output: JSON result envelope and SARIF v2.1.0;
- two tool-protocol adapters (MCP 2026-07-28, OpenAI function tools) that translate external tool definitions into the domain model without any network or SDK surface;
- synthetic policy packs (customer support, finance, DevOps, sensitive-data access) with vulnerable and fixed variants.
Features described as "future" or "conceptual" in the design documents are
not implemented; see "Unsupported features" below. The implementation
follows the roadmap in docs/ROADMAP.md.
A policy declares a security property:
{
"version": 1,
"id": "refund-exposure",
"description": "Cumulative refunds above 100 per customer are forbidden.",
"rules": [
{
"id": "refund-exposure",
"description": "Cumulative refunds above 100 per customer are forbidden.",
"kind": "cumulative_limit",
"action": "refund",
"field": "amount",
"limit": 100,
"scope_argument": "customer_id"
}
]
}policyfuzz test searches bounded action sequences, executes them in the
deterministic simulator, evaluates the policy after every step, and
reports a minimized, reproducible counterexample when a sequence violates
it — here, two 60-unit refunds totalling 120 against the 100 cap.
- Policy-first — security requirements are expressed as executable properties.
- Outcome-oriented — test forbidden outcomes, not only individual tool permissions.
- Compositional — analyze sequences and combinations of actions.
- Reproducible — every finding must have a deterministic or replayable reproduction.
- Local-first — sensitive agent configurations and test data should not need to leave the developer environment.
- Framework-agnostic core — integrations belong at the boundary.
- Open source first — the complete core remains available under the repository's chosen open-source license.
- Evidence over scores — findings must explain what happened and why the policy was violated.
- Fail safely — unknown behavior must never be silently classified as safe.
- No false claims — passing a test suite does not prove an agent is universally secure.
Developers building AI agents that can call tools or take actions.
Security engineers, AppSec engineers, platform engineers, and teams responsible for deploying multiple AI agents.
Developers building ordinary applications without autonomous or semi-autonomous tool execution.
Given:
- an agent or agent-compatible tool interface,
- a set of tool definitions,
- a policy,
- a state model or sandbox,
- and synthetic/test data,
generate adversarial action sequences and determine whether any sequence violates a declared security property.
- Policy syntax v1 (JSON, or YAML when PyYAML is installed): six rule
kinds —
forbidden_action,approval_required,sensitive_data_flow,cumulative_limit,sequence_restriction,capability_combination— with strict unknown-key rejection, two-stage validation, and deterministic compositionDENY > ERROR > REQUIRE_APPROVAL > INCONCLUSIVE > UNKNOWN > ALLOW. Seedocs/POLICY_SPEC.mdand thepolicyfuzz.policymodule docstring. - Deterministic simulator: 16 synthetic in-memory tools (no shell, no
network), explicit seed, resource limits (
max_steps,max_tool_calls,timeout_seconds,max_outputs_bytes), and byte-replayable traces viaSimulator.replay. - Bounded BFS search: explicit candidate templates, state
deduplication, seeded ordering, optional
priorityhook for risk-guided discovery order, and graceful incomplete reporting. - Findings: stable
PF-000001-style IDs, structured evidence, greedy 1-minimal sequence minimization, deterministic reproduction, terminal and JSON renderers. - CLI:
init,validate,test,reproduce,version; JSON and SARIF v2.1.0 output;--outputartifact writing;--fail-on-inconclusive;--no-minimize. - Adapters: MCP 2026-07-28 (full tools/list + tools/call, content annotations, x-mcp-header, outputSchema validation, structuredContent validation, ttlMs/cacheScope caching) and OpenAI function tools (flat + nested shapes, strict mode validation, output_schema validation) — translation only, no transport, no SDK.
- Provider adapters (
src/policyfuzz/providers/): Anthropic Claude, Google Gemini, OpenAI, and OpenAI-compatible providers (Ollama, vLLM, LiteLLM, Groq) — each translates native tool/function definitions into the internal domain model with contract tests and deterministic local fixtures. - Policy packs: four synthetic vulnerable+fixed packs under
fixtures/packs/. - CI: exit-code contract, JSON envelope, SARIF, GitHub Actions starter
(
examples/ci/policyfuzz-ci.yml) and an active regression workflow.
Not implemented (do not assume they exist):
policyfuzz explainor any LLM-assisted explanation, authoring, or intent generation;- natural-language policy authoring — policies are structured documents only;
- live tool transports: no HTTP, stdio, SSE, OAuth, or SDK client anywhere; adapters translate vendored JSON snapshots only;
- real tool execution: the simulator is synthetic and in-memory; running real tools/servers requires isolation outside PolicyFuzz;
- pagination, elicitation, tasks, sampling, or stateful sessions from MCP;
- non-
functionOpenAI tool types,asyncexecution, deferredtool_searchloading,allowed_callersenforcement, orchestration; - rule kinds beyond the four v1 kinds (no sequence-restriction,
capability-combination, or trust-boundary rule kinds — see
docs/POLICY_SPEC.mdfor the conceptual categories); - property-based search, symbolic execution, constraint solving, distributed search;
- JUnit, Markdown, or HTML reporters (terminal, JSON, and SARIF only);
- hosted dashboards, cloud services, billing, runtime enforcement, or production exploitation.
- Search is bounded and incomplete by default. BFS explores the space
of explicitly supplied candidate templates up to
max_steps/max_cases. APASSmeans "no violation found within the bound", never "no violation exists". Truncated runs reportINCOMPLETE(exit3), neverPASS. - Only the six v1 rule kinds are evaluated. Policies over behaviors the state model cannot express (e.g. cross-agent flows, timing) are out of scope.
- History-sensitive policies depend on state encoding. Deduplication fingerprints cover resources, data assets, identities, approvals, balances, data locations, trust levels, and unknown aspects — not the event log. A policy whose verdict depends on history invisible in state would need a fingerprint change.
- The simulator is synthetic. Findings demonstrate policy violations against the modeled environment, not against a live system.
- Adapters are translation layers, not sandboxes. A passing adapter run only means the vendored definitions translated and the local fixture behaved; it says nothing about a live server or model.
- Terminal findings include a fixed generic "Why it matters" sentence. It is not derived from the finding's evidence; treat it as a placeholder for a future evidence-derived explanation.
- SARIF
physicalLocationpoints at the tested policy file (line 1); precise per-step source ranges are not emitted. - YAML support (PyYAML is a required dependency); JSON also works everywhere.
- Performance: BFS is exponential in depth by design; trace
serialization is O(n^2) in trace length; minimization is O(n^2)
executions. See
docs/PERFORMANCE_AUDIT.md.
Given the same policy, initial state, candidate pool, and seed:
- candidate order is fixed (canonical sort + one seeded shuffle, or
priority tiers + seeded shuffle when a
priorityhook is supplied); - the simulator reproduces traces byte-for-byte (
to_jsonequality) from a serializedTestCase(seed + initial state + actions); - finding IDs are a pure function of an explicit counter;
- JSON and SARIF output use sorted keys for byte-stable artifacts;
- repeated
policyfuzz testruns produce byte-identical finding files.
Wall-clock timeouts (timeout_seconds) and real-time sleeps are the only
nondeterministic elements, and only when explicitly configured.
A run is completed only when the bounded space was exhaustively explored
within max_steps/max_cases. Hitting a bound stops the run with
completed: false and RESOURCE_LIMIT; the CLI reports INCONCLUSIVE and
exits 3 — never a silent pass. --fail-on-inconclusive additionally
escalates a completed run that saw inconclusive evaluations to exit 3.
- No network or subprocess surface in the core. The domain, policy,
simulator, search, finding, and SARIF layers import the standard library
plus
policyfuzz.domainonly; architecture tests enforce this. The adapters contain no transport code. - The simulator never executes real tools. All tool behavior is
synthetic and in-memory; shell/network-sounding tool ids map to
UNKNOWN_TOOL/UNSUPPORTED_TOOL, never to execution. - Untrusted input is data. Tool metadata, descriptions, and fixture content are preserved as data, never executed or sent to an LLM.
- Resource limits are enforced (steps, tool calls, wall-clock, output size) and oversized outputs are rejected before allocation.
- Findings contain only modeled data — whatever you put in fixtures. Use synthetic data; never place real credentials or PII in arguments, definitions, or issues.
- PolicyFuzz is a testing tool, not a runtime enforcer. A passing run
does not certify that an agent is secure; it tests the modeled
environment within the configured bounds. See
SECURITY.mdanddocs/THREAT_MODEL.md.
PolicyFuzz is not:
- a generic vulnerability scanner;
- an LLM observability platform;
- a production runtime firewall;
- an agent hosting platform;
- a compliance certification service;
- a guarantee that an agent is secure;
- a replacement for human security review;
- a general-purpose autonomous penetration-testing platform.
docs/PRD.md— product requirements and boundaries.docs/PRODUCT_SPEC.md— functional behavior and user-facing semantics.docs/POLICY_SPEC.md— policy language, rule kinds, and evaluation model.docs/ARCHITECTURE.md— system architecture.docs/DOMAIN_MODEL.md— core entities and relationships.docs/THREAT_MODEL.md— security assumptions, threats, and abuse boundaries.docs/TESTING_STRATEGY.md— testing philosophy and required test layers.docs/ROADMAP.md— development roadmap and status.docs/DEVELOPMENT_GUIDE.md— implementation conventions.docs/DECISIONS.md— architecture decision records (ADR-0001 … ADR-0012).docs/RELEASE.md— release and versioning rules.docs/CI.md— CI exit codes, JSON envelope, and SARIF contract.docs/MCP_ADAPTER.md— MCP adapter capabilities and limitations.docs/OPENAI_ADAPTER.md— OpenAI function-tools adapter capabilities and limitations.docs/PERFORMANCE_AUDIT.md— measured performance characteristics and known complexity.SECURITY.md— security reporting and secure development expectations.CONTRIBUTING.md— contribution process.
policyfuzz/
├── README.md
├── LICENSE
├── CONTRIBUTING.md
├── SECURITY.md
├── pyproject.toml
├── docs/ # design documents, ADRs, adapter docs
├── src/policyfuzz/
│ ├── domain.py # immutable domain model (schema_version 1)
│ ├── policy.py # policy syntax v1: parse, validate, evaluate
│ ├── simulator.py # deterministic in-memory execution environment
│ ├── search.py # bounded deterministic BFS over candidates
│ ├── finding.py # finding lifecycle, minimization, reproduction
│ ├── sarif.py # SARIF v2.1.0 reporter
│ ├── project.py # project config/fixture loading (CLI boundary)
│ ├── cli.py # init/validate/test/reproduce/version
│ ├── adapters/
│ │ ├── mcp.py # MCP tool adapter (translation only)
│ │ └── openai_tools.py # OpenAI function-tools adapter (translation only)
│ └── providers/ # provider-specific tool adapters
│ ├── anthropic.py # Anthropic Claude tool-use adapter
│ ├── gemini.py # Google Gemini function-calling adapter
│ ├── openai.py # OpenAI function-tools adapter
│ ├── openai_compatible.py # base for OpenAI-compatible providers
│ ├── ollama.py # Ollama (OpenAI-compatible)
│ ├── vllm.py # vLLM (OpenAI-compatible)
│ ├── litellm.py # LiteLLM (OpenAI-compatible)
│ └── groq.py # Groq (OpenAI-compatible)
├── tests/ # 895 tests (unit, contract, integration, packs)
├── examples/
│ ├── quickstart/ # six-step vulnerable+fixed refund fixture
│ └── ci/ # GitHub Actions starter + local walkthrough
├── fixtures/
│ ├── packs/ # four synthetic policy packs (vulnerable+fixed)
│ ├── policies/ # valid/invalid/boundary policy fixtures
│ ├── search/ # search fixtures + README
│ ├── mcp/ # MCP adapter fixtures
│ └── openai/ # OpenAI adapter fixtures
└── benchmarks/bench.py # reproducible performance benchmarks
Prefer a small, composable core over a large framework-specific implementation.
Every major feature should answer:
- What user problem does it solve?
- What invariant does it preserve?
- What is the smallest useful implementation?
- How is it tested?
- How is it reproduced?
- What happens when information is unknown?
- Does it introduce a new security boundary?
- Does it belong in the core or an adapter?
A feature is not complete merely because its happy path works.
A feature is complete when:
- behavior is specified;
- unit tests exist;
- integration tests exist where applicable;
- error cases are covered;
- security implications are reviewed;
- public interfaces are documented;
- deterministic reproduction is possible where relevant;
- no unrelated behavior regresses.
Requires Python 3.10+.
pip install -e ".[dev]"
policyfuzz --help
policyfuzz version
python -m policyfuzz version
pytest
ruff check src tests
ruff format --check src testsThe version output is deterministic: policyfuzz 0.1.0.
policyfuzz init ./my-project
cd ./my-project
policyfuzz validate # exit 0 when the project is valid
policyfuzz test # exit 1: seeded violation (two 60-unit refunds exceed 100)
ls findings/ # PF-000001.json
policyfuzz reproduce findings/PF-000001.json --policy policy.json # exit 0Fix the policy (raise the cumulative_limit cap) and rerun:
policyfuzz test then prints Result: PASS and exits 0.
Exit codes: 0 no violations, 1 confirmed violation (or reproduce
mismatch), 2 configuration/usage error, 3 execution error or
incomplete search (never a silent pass).
Machine-readable CI output:
policyfuzz test --format json --output results.json # stable envelope, exit 0/1/2/3
policyfuzz test --format sarif --output results.sarif # SARIF v2.1.0, pass/fail onlySee examples/quickstart/README.md for the full six-step fixture,
examples/ci/README.md + examples/ci/policyfuzz-ci.yml for the
GitHub Actions starter (install, validate, fail on violation, pass when
fixed -- no credentials or external services), and docs/CI.md for the
JSON/SARIF schema and exit-code contract.
This project is licensed under the MIT License. See LICENSE.