Phase-HDC: Replacing Optimizer History with Gradient Thresholds in Discrete Phase Learning
For a hyperdimensional classifier whose learned parameters are low-bit angles, which we call a phase memory, these records take several times more memory than the model itself.
Key points
- Training a compact model often needs far more memory than storing it, because the optimizer keeps its own records of past gradients.
- We show that this simple rule is the exact solution of a first-order loss model in which every changed parameter pays a fixed cost.
- The price is an average loss of about five accuracy points against float32 Adam, while Phase-HDC is more accurate than 8-bit Adam on six of the eleven datasets, including byte-level text prediction, where 8-bit Adam collapses.
- Once parameters must sit on a discrete grid, Adam's moments mainly decide whether a parameter moves at all, a decision that a threshold on the current gradient can make without memory, and coarse quantization of the moments breaks this decision for inputs that the data rarely contain.
Sources (1)
- [1]Phase-HDC: Replacing Optimizer History with Gradient Thresholds in Discrete Phase LearningarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 7, 01:05 PM
For a hyperdimensional classifier whose learned parameters are low-bit angles, which we call a phase memory, these records take several times more memory than the model itself.
Training a compact model often needs far more memory than storing it, because the optimizer keeps its own records of past gradients.
Extractive summary: sentences quoted from the sources.
Before this
- Oct 6, 2026EmbeddingGemma 2: an open, lightweight multimodal embedding model
- Oct 6, 2026huggingface/transformers v5.19.0: Release v5.19.0
- Sep 22, 2026vllm-project/vllm v0.30.0
- Aug 10, 2026vllm-project/vllm v0.27.0
- Jul 11, 2026vllm-project/vllm v0.25.0
- Jun 10, 2026DiffusionGemma: 4x faster text generation