ApexQuant: Data-Free Elastic Quantization by Residual Re-Isotropization
We introduce ApexQuant, a calibration-free quantization method that recursively re-quantizes the residual error, serving as a refinement layer on top of existing quantizers.
ProofPaper ↗
Key points
- We establish that a fresh random rotation returns each residual to the uniform distribution on the hypersphere, which characterizes the rate of progressive error decay across successive passes.
- This result lets us determine, before any weight is read, how many passes a layer needs for a target weight-space error.
- We instantiate ApexQuant with three interchangeable stages, scalar, $E8$ and trellis, and validate it on four open-weight LLMs and on Earth-observation and medical domains where in-distribution data is often unattainable as imagery arrives under restrictive licences or due to patient material under privacy constraints.
- Progressive re-isotropization comes within a few percent of full precision at four bits and gives the best two-bit arm we measure, in a completely data-free setting.
Sources (1)
- [1]ApexQuant: Data-Free Elastic Quantization by Residual Re-IsotropizationarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 6, 07:49 AM
We introduce ApexQuant, a calibration-free quantization method that recursively re-quantizes the residual error, serving as a refinement layer on top of existing quantizers.
We establish that a fresh random rotation returns each residual to the uniform distribution on the hypersphere, which characterizes the rate of progressive error decay across successive passes.
Extractive summary: sentences quoted from the sources.
Before this
- Sep 22, 2026vllm-project/vllm v0.30.0
- Aug 10, 2026vllm-project/vllm v0.27.0
- Jul 11, 2026vllm-project/vllm v0.25.0
- Jun 29, 2026vllm-project/vllm v0.24.0
- Jun 15, 2026vllm-project/vllm v0.23.0
- Jun 10, 2026DiffusionGemma: 4x faster text generation