RoboHarm: Do Frontier Robot Policies Refuse Unsafe Instructions?
Summary
RoboHarm evaluates frontier robot policies (Claude Fable 5.1, GPT-6 Astra, MolmoAct2) on five risky tasks designed to elicit unsafe actions. The study measures how often each policy refuses unsafe instructions versus completes or fails to act, presenting per-task results, overall refusals, and limitations of the evaluation approach.