GPT-6 Astra, Looped Transformers, and Hidden Reasoning
Summary
A detailed discussion by Sebastian Raschka on GPT-6 Astra, looped transformers, and the question of hidden reasoning traces. The article explains looped transformer concepts, compares them to Universal Transformers and other variants, reviews recent papers, and discusses implications for performance, compute cost, interpretability, and the rumor that Astra hides its chain-of-thought.