DeltaTTT: Layerwise Optimization for Nonlinear Recurrent Memory
To alleviate this optimization difficulty, we introduce DeltaTTT, which replaces joint inner-loop optimization of a two-layer memory network with layerwise learning.
ProofPaper ↗
Key points
- Sequential test-time training adapts a memory network through successive updates, each computing an inner-loop gradient based on the network's previous state.
- However, we find that this expected advantage does not consistently materialize in nonlinear memories: a fixed-base parallel TTT baseline outperforms its serial counterpart.
- Our exploratory experiments point to a key underlying difficulty: nonlinear memories can be harder to optimize than linear ones within a single pass over the sequence.
- Experiments on DeltaNet and LaCT backbones show improvements in language modeling and retrieval over their recurrent baselines.
Sources (1)
- [1]DeltaTTT: Layerwise Optimization for Nonlinear Recurrent MemoryarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 6, 03:40 PM
To alleviate this optimization difficulty, we introduce DeltaTTT, which replaces joint inner-loop optimization of a two-layer memory network with layerwise learning.
Sequential test-time training adapts a memory network through successive updates, each computing an inner-loop gradient based on the network's previous state.
Extractive summary: sentences quoted from the sources.