A Princeton Researcher Proposes Recurrent Looped Transformer (RLT) that Carries Decoder State across Every Token, Fixing 96 Blocks per Token with Unbounded Temporal Depth

| Source: MarkTechPost

Tags: Recurrent Looped Transformer, Princeton, Yifan Zhang, transformer architecture, sliding-window attention, decoder design, recurrent networks

Princeton researcher Yifan Zhang proposes the Recurrent Looped Transformer (RLT), which feeds each token's final decoder hidden state and its sliding-window attention cache into the next token — enabling structural depth that grows with sequence length while per-token compute stays fixed. This is a pure design spec; no trained model, benchmarks, or efficiency numbers are included.

Details

Standard decoder-only LLMs pass information between tokens only via KV-cache attention — nothing from the last layer of token t flows into the first layer of token t+1. Zhang's RLT closes this gap by pairing a causal encoder with a recurrent decoder whose full hidden state (final output plus per-layer sliding-window attention cache) carries into the next token without resetting at the prompt/response boundary. The reference design uses 48 encoder and 48 decoder layers with shared weights, so each token executes 96 logical blocks. After t tokens, the state path traverses t×48 decoder blocks — depth growing with sequence length while per-token compute stays fixed. Zhang calls this unbounded temporal depth, but the paper explicitly cautions that gating and contraction may suppress information from distant positions, meaning structural depth does not guarantee long-range reasoning. Three design principles guide the architecture: latent reasoning via accumulated state, model-hardware co-design (encoder runs token-parallel, decoder runs sequentially but can batch across independent sequences), and an RL fine-tuning replay contract. Zhang explicitly states no parallel scan is assumed for the nonlinear decoder and no prefill speedup is claimed over standard approaches. This is a technical report defining architecture and execution schedules — no training runs, benchmark results, or efficiency measurements are included. It is a design proposal worth tracking, not a deployed or empirically validated result.