"As a Language Model...": Chat Template Switches LLM Self-Referential Voice and Activation Steering Reproduces It
Summary
The paper demonstrates that chat prompts act as a switch that modulates self-referential voice in LLMs, amplifying disclaimers like 'I'm just an AI' when a template is present and reducing them when it is not. It further shows a directional activation in the models that can steer this voice, implying that self-descriptions are shaped by deployment prompts and are not literal facts about the model.