Ring-Zero: Scaling Zero RL to a Trillion Parameters for Emergent Reasoning
Summary
Ring-Zero investigates scaling zero-shot reinforcement learning to a trillion parameters to elicit emergent reasoning. It introduces a stable training pipeline with algorithmic and system optimizations, demonstrates that scaling improves sample efficiency and unlocks advanced cognitive behaviors, and proposes a structured evaluation framework for reasoning traces beyond final answers.