Evaluation Metrics for Safe Reinforcement Learning
| Source: arXiv AI
Tags: safe reinforcement learning, RL benchmarks, CMDP, AI safety, evaluation metrics, open source
Researchers argue that safe reinforcement learning benchmarks are broken: average-case metrics hide how often and how severely safety bounds are violated. They propose a new evaluation framework with distributional metrics and a safety tier system, plus an open-source evaluation suite (SafeRLEval).
Details
Safe reinforcement learning is typically formalized as a Constrained Markov Decision Process (CMDP) where an agent maximizes reward while keeping cumulative cost below a safety bound. Most benchmarks report average safety — whether the bound is violated on average — and little else. This paper argues this is insufficient. Average-based reporting fails to capture violation frequency, violation severity, consistency across tasks and safety thresholds, and whether training-time behavior represents the final converged policy. The authors propose new evaluation metrics that address each gap, plus a safety tier system for categorizing algorithms by safety and reliability at both training time and for the final policy. They demonstrate empirically that aggregate metrics, distributional reporting, and task-specific results each surface information the others miss. Their recommendation: report all three jointly rather than collapsing to a single value. They release SafeRLEval, an open-source evaluation suite supporting the framework. This is methodological work rather than a novel algorithm, but it addresses a real gap in how safe RL is evaluated in practice.