OpenART is an open-source framework designed to evaluate the safety and robustness of autonomous AI agents in dynamic, long-horizon, and stateful environments. It stress-tests agent runtimes against multi-step state poisoning, privilege escalation, and tool-use vulnerabilities across 10,000+ benchmark scenarios.
agent docker benchmark jailbreak red-team red-teaming tool-use prompt-injection llm-security llm-evaluation llm-agents agent-skills agent-security agent-evaluation agent-safety agent-safety-frameworks agent-environment tool-use-safety agent-redteam tool-use-security
-
Updated
Oct 3, 2026 - Python