DigiNews

Tech Watch by Johan Denoyer

← Back to articles

GRP-Obliteration: Unaligning LLMs With a Single Unlabeled Prompt

Quality: 8/10 Relevance: 9/10

Summary

arXiv CS ML paper introduces GRP-Obliteration (GRP-Oblit), a method using Group Relative Policy Optimization to remove safety constraints from target models with a single unlabeled prompt. The work claims to achieve stronger unalignment while preserving utility and generalizes to diffusion-based image generation. It evaluates across 15 models (7-20B params) and multiple benchmarks, discussing policy containment and the role of formal methods in auditing AI safety.

🚀 Service construit par Johan Denoyer