ResearchResearch paperLarge Language Models1 source · Oct 6, 2026

DeltaTTT: Layerwise Optimization for Nonlinear Recurrent Memory

To alleviate this optimization difficulty, we introduce DeltaTTT, which replaces joint inner-loop optimization of a two-layer memory network with layerwise learning.

Key points

  • Sequential test-time training adapts a memory network through successive updates, each computing an inner-loop gradient based on the network's previous state.
  • However, we find that this expected advantage does not consistently materialize in nonlinear memories: a fixed-base parallel TTT baseline outperforms its serial counterpart.
  • Our exploratory experiments point to a key underlying difficulty: nonlinear memories can be harder to optimize than linear ones within a single pass over the sequence.
  • Experiments on DeltaNet and LaCT backbones show improvements in language modeling and retrieval over their recurrent baselines.

Sources (1)

  • [1]DeltaTTT: Layerwise Optimization for Nonlinear Recurrent Memory
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 6, 03:40 PM
    To alleviate this optimization difficulty, we introduce DeltaTTT, which replaces joint inner-loop optimization of a two-layer memory network with layerwise learning.
    Sequential test-time training adapts a memory network through successive updates, each computing an inner-loop gradient based on the network's previous state.

Extractive summary: sentences quoted from the sources.

Related