AION
Research paperEfficiency & Inference · Training & Scaling · Large Language Models1 source · Oct 7, 2026

Phase-HDC: Replacing Optimizer History with Gradient Thresholds in Discrete Phase Learning

For a hyperdimensional classifier whose learned parameters are low-bit angles, which we call a phase memory, these records take several times more memory than the model itself.

Key points

  • Training a compact model often needs far more memory than storing it, because the optimizer keeps its own records of past gradients.
  • We show that this simple rule is the exact solution of a first-order loss model in which every changed parameter pays a fixed cost.
  • The price is an average loss of about five accuracy points against float32 Adam, while Phase-HDC is more accurate than 8-bit Adam on six of the eleven datasets, including byte-level text prediction, where 8-bit Adam collapses.
  • Once parameters must sit on a discrete grid, Adam's moments mainly decide whether a parameter moves at all, a decision that a threshold on the current gradient can make without memory, and coarse quantization of the moments breaks this decision for inputs that the data rarely contain.

Sources (1)

  • [1]Phase-HDC: Replacing Optimizer History with Gradient Thresholds in Discrete Phase Learning
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 7, 01:05 PM
    For a hyperdimensional classifier whose learned parameters are low-bit angles, which we call a phase memory, these records take several times more memory than the model itself.
    Training a compact model often needs far more memory than storing it, because the optimizer keeps its own records of past gradients.

Extractive summary: sentences quoted from the sources.

Before this

  1. Oct 6, 2026EmbeddingGemma 2: an open, lightweight multimodal embedding model
  2. Oct 6, 2026huggingface/transformers v5.19.0: Release v5.19.0
  3. Sep 22, 2026vllm-project/vllm v0.30.0
  4. Aug 10, 2026vllm-project/vllm v0.27.0
  5. Jul 11, 2026vllm-project/vllm v0.25.0
  6. Jun 10, 2026DiffusionGemma: 4x faster text generation

Related