AION
Research paperLarge Language Models1 source · Oct 6, 2026

Hybrid Latent Attention for Looped Language Models

We propose Hybrid Latent Attention (HLA), which keeps exact keys and values within a sliding window of W recent tokens and stores each older token as a compact latent that the query of each loop reads directly, without reconstructing keys and values.

Key points

  • Looped language models apply the same stack of layers T times to each token, which deepens the model without adding parameters but multiplies its key-value (KV) cache by T.
  • The larger cache limits how many sequences a GPU can decode at once and slows each decoding step, which reads the whole cache.
  • We uptrain HLA on Ouro looped models (T=4) with 1.4B and 2.6B parameters, keeping the pretrained weights frozen and training only the added parameters to reproduce the original attention.
  • HLA retains over 97% of the original accuracy on math, knowledge and reasoning benchmarks, and 96-100% on long-context retrieval up to 16K tokens.

Sources (1)

  • [1]Hybrid Latent Attention for Looped Language Models
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 6, 08:14 AM
    We propose Hybrid Latent Attention (HLA), which keeps exact keys and values within a sliding window of W recent tokens and stores each older token as a compact latent that the query of each loop reads directly, without reconstructing keys and values.
    Looped language models apply the same stack of layers T times to each token, which deepens the model without adding parameters but multiplies its key-value (KV) cache by T.

Extractive summary: sentences quoted from the sources.