GRP-Obliteration: Unaligning LLMs With a Single Unlabeled Prompt
Summary
arXiv CS ML paper introduces GRP-Obliteration (GRP-Oblit), a method using Group Relative Policy Optimization to remove safety constraints from target models with a single unlabeled prompt. The work claims to achieve stronger unalignment while preserving utility and generalizes to diffusion-based image generation. It evaluates across 15 models (7-20B params) and multiple benchmarks, discussing policy containment and the role of formal methods in auditing AI safety.