More computation before the next token
A standard autoregressive transformer traverses its fixed layer stack for each new token. A depth-recurrent model can apply a shared core several times before decoding a token. It must be trained for this recurrence; adding a loop to an arbitrary trained model is not enough.
A small model, twice. A larger model, once.
The moving dot is a token representation. Two loops and six distinct layers are illustrative; animation time and block size are not benchmark measurements.
Smaller loop model
2 passes · the same weights θ
Core passes completed: 0 / 2
Larger standard model
1 pass · distinct layer weights
Stack passes completed: 0 / 1
What changes at similar quality?
Relative to the larger direct-answer model, at similar task quality. Outcomes depend on model, workload, and hardware.
- Model memoryGB of model weights ↓
Less is possible
A smaller shared core can reach the target quality through recurrence. Compare at the same numerical precision.
- SpeedSeconds per complete answer ↓
No fixed advantage
A smaller core may be faster per pass, but serial loops add latency. Measure time to the completed answer.
- ComputeTFLOP per complete answer ↓
Can be similar or higher
Fewer stored parameters do not imply fewer operations. Include every core iteration and attention operation.
- Memory bandwidthGB transferred from/to HBM per answer ↓
Fewer weights ≠ fewer transfers
Shared weights save capacity, but may still be read from GPU memory on each loop. Reuse in fast cache is not guaranteed.
- InterpretabilityMonitor recall at a fixed false-positive rate ↑
Both have opaque internal states
Direct answers provide little intermediate text in either model. There is no universal ranking of their interpretability.
Latent updates or a growing chain of thought
Follow the same illustrative timeline below: recurrent depth updates a bounded state before emitting a token. Explicit CoT generates additional tokens, each through another model pass, and retains them in the context.
Loop: update the workspace
Answer follows the latent passes
Until then: internal state updates without written intermediate steps.
Same shape, new values. No new text position for each internal loop.
CoT: append intermediate text
Each tile represents a phrase made of multiple tokens. The text remains in the context; hidden states also change during every pass.
Illustrative calculation: (17 × 6) + 8. The latent colors are not decoded arithmetic steps.
Repeated updates can transform and mix information within the latent state without leaving readable intermediate text. A text monitor therefore misses those updates. Explicit CoT supplies an additional trace, although it is not a complete or guaranteed faithful explanation. Activation analysis must track both layer and loop position.
What happens to the hidden state?
A trained recurrent core transforms its current hidden state and feeds the result back into the same weights. In this example it runs twice before output decoding. A standard transformer instead traverses its distinct layers once for each new token.
The state is a tensor associated with token positions, not a single stored sentence. Additional loops do not automatically enlarge its capacity or improve the answer. Input injection, caches, and stopping rules vary by architecture.
Scaling up Test-Time Compute with Latent ReasoningMemory and compute are different constraints
Recurrence lets a model spend more inference compute without adding a new set of weights for every depth step. Avoiding explicit reasoning tokens can also reduce sequence growth. This separates parameter storage from executed depth, but does not make the extra work free.
Weight capacity
Can the parameters fit in device memory? Reusing a core can reduce storage relative to a larger model that reaches comparable quality.
Bandwidth versus compute
How many operations run per byte moved? Weight reuse does not guarantee that a workload becomes compute-bound. Batch size, kernels, caching, and hardware determine the bottleneck.
Hidden-state capacity
A fixed-size latent workspace has finite representational capacity. More iterations add computation, not automatically more working memory. Saving a history of states introduces its own memory cost.
Harder to monitor does not mean impossible to interpret
Latent states remain available to activation probes and causal interventions when researchers have access to the internals. Huginn studies find probe-dependent results; LOTUS shows that explicit supervision can make latent steps more readable. Repeated transformation makes analysis more demanding, not inherently impossible.
In the spotlight through Astra
Astra coverage brought looped transformers into the spotlight. The Information
The Astra System Card reports reduced CoT monitorability, but does not document a looped-transformer architecture or establish recurrence as the cause. This diagram explains the general research idea.
GPT-6 Astra System Card