Emergent Introspective Awareness in Large Language Models
Summary
The paper explores whether large language models can introspect their own internal states. By injecting known concepts into activations and testing self-reports, it finds limited but real evidence of introspective awareness in some models, especially larger ones like Claude Opus 4/4.1. The authors caution that the effect is unreliable and context-dependent but may develop with future advances.