Ask a model if code is malicious and it reaches for its morals
Summary
Manifold's article analyzes how mixture-of-experts models evaluate code for malicious intent, revealing that the morality path heavily influences judgments and that routing decisions can alter answers. It details experiments with OLMoE and DeepSeek models, showing how the route and workspace cues shape outcomes, and demonstrates that the authenticity of the code evaluation can be transferred by swapping routing. Pruning experiments indicate the path, not the stored judgment, drives decisions, and the morality route sharpens malice-vs-benign separation.