ResearchResearch paperTraining & Scaling · Efficiency & Inference · Robotics & Embodied AI1 source · Oct 6, 2026

LayerRoPE: Dynamic Depth-wise Magnitude & Angular Superposition

Across 16 pre-trained LLMs from 9 families, spanning dense, mixture-of-experts and hybrid architectures and Pre-, Peri- and Post-Norm designs, we find that this growth reflects an emergent depth-positional encoding, carried by the only learned per-layer gain on the residual stream, the normalization weight $γ$: with depth, $γ$ grows in magnitude and rotates in direction, jointly encoding the layer index.

Key points

  • As data propagates through a Transformer, the norm of its hidden states grows by orders of magnitude with depth, a phenomenon framed as 'curse of depth' and nearly universally treated as a pathology to be suppressed.
  • We make this depth-conditioned encoding explicit with LayerRoPE, an implicit analog of RoPE along the depth axis, which replaces all layerwise $γ$ vectors with a single shared vector and depth-conditioned scalars, at a net reduction in parameters and $<0.02%$ change in FLOPs.
  • Across a model ladder scaled up to $100$B+ tokens, LayerRoPE consistently outperforms Pre-, Post- and Peri-Norm and Layer-Norm Scaling, reaching Pre-Norm's 1.3B loss with $3.4\times$ less compute; LayerRoPE is the only approach that shows strong convergence and improves near monotonically as depth scales to 512 layers.
  • Depth stability, our results suggest, calls not for suppressing the residual stream, but for depth-conditioned regulation of the computational blocks it feeds.

Sources (1)

  • [1]LayerRoPE: Dynamic Depth-wise Magnitude & Angular Superposition
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 6, 10:30 PM
    Across 16 pre-trained LLMs from 9 families, spanning dense, mixture-of-experts and hybrid architectures and Pre-, Peri- and Post-Norm designs, we find that this growth reflects an emergent depth-positional encoding, carried by the only learned per-layer gain on the residual stream, the normalization weight $γ$: with depth, $γ$ grows in magnitude and rotates in direction, jointly encoding the layer
    As data propagates through a Transformer, the norm of its hidden states grows by orders of magnitude with depth, a phenomenon framed as 'curse of depth' and nearly universally treated as a pathology to be suppressed.

Extractive summary: sentences quoted from the sources.

Before this

  1. Oct 6, 2026huggingface/transformers v5.19.0: Release v5.19.0
  2. Sep 9, 2026vllm-project/vllm v0.29.0
  3. Aug 26, 2026huggingface/transformers v5.16.0: Release: v5.16.0
  4. Aug 10, 2026vllm-project/vllm v0.27.0
  5. Jul 11, 2026vllm-project/vllm v0.25.0
  6. Jul 3, 2026huggingface/transformers v5.13.0: Release v5.13.0

Related