Raschka Says OpenAI Astra's Looped Transformer Is Not a Big Deal
Original titleA lot of hype around OpenAI's Astra model here on my timeline today. Apparently, this goes back to a new article from The Information, wh...
AISummary
Sebastian Raschka argues that the looped transformer approach attributed to OpenAI's Astra is a minor architectural tweak, not a major breakthrough.
He explains that Nanbeige4.2-3B reuses its 22-layer stack twice, effectively doubling depth without adding weights but roughly doubling compute, and that the idea traces back to the Mixture-of-recursions NeurIPS paper.
He adds that layer reuse does not inherently hide chain-of-thought, though it could shift more computation into latent activations.
Source: Sebastian Raschka · x.comPublished · added here