ResearchResearch paperTraining & Scaling1 source · Oct 7, 2026

Lightweight and Versatile Learned Optimization by Recombination of Gradient History

This paper presents a lightweight and versatile learned optimizer that dynamically recombines gradient history, represented as averages over disjoint time spans.

Key points

  • The optimizer reduces the prediction space to one scalar coefficient per gradient average, shared by multiple parameters.
  • Progressively averaging older gradients minimizes memory cost of long history, while keeping their contributions independently accessible.
  • A 37k-parameter network trained in 0.87 GPU-hours generalizes zero-shot to unseen tasks, lowering validation loss by 9.1% and 0.4% on BERT-Tiny and GPT-Tiny, and improving test accuracy over Adam by 3.5 %p on a Vision Transformer and by 2.7 %p on average across nine graph models, with FLOPs overhead as low as 0.3%.

Sources (1)

  • [1]Lightweight and Versatile Learned Optimization by Recombination of Gradient History
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 7, 07:50 AM
    This paper presents a lightweight and versatile learned optimizer that dynamically recombines gradient history, represented as averages over disjoint time spans.
    The optimizer reduces the prediction space to one scalar coefficient per gradient average, shared by multiple parameters.

Extractive summary: sentences quoted from the sources.

Before this

  1. Oct 7, 2026GPT-6 and Intelligent UI for everyone
  2. Oct 7, 2026MIMESIS: Learning User Simulators as Training Environments for Interactive Agents
  3. Oct 6, 2026llm-openai-decisions 0.1a0
  4. Oct 4, 2026nerkyor/Qwen3.8-27B-Coder390-EfficientThink-Opus5.5-GPT6Astra-Grok4.7-DSV4Pro-K3-SFT-RLOO-MTP-DFlash2
  5. Oct 1, 2026Claude-shaped science
  6. Sep 29, 2026Introducing GPT-6.1 Sol

Related