OpenART is a new framework for evaluating the safety of AI agents in long-horizon, stateful environments where early actions can affect later outcomes. The paper argues that existing safety benchmarks
The paper examines Claude Code as an example of an agentic coding system that can execute shell commands, edit files, and interact with external services on a user’s behalf. Rather than focusing only