ResearchResearch paperEfficiency & Inference1 source · Oct 6, 2026

ApexQuant: Data-Free Elastic Quantization by Residual Re-Isotropization

We introduce ApexQuant, a calibration-free quantization method that recursively re-quantizes the residual error, serving as a refinement layer on top of existing quantizers.

Key points

  • We establish that a fresh random rotation returns each residual to the uniform distribution on the hypersphere, which characterizes the rate of progressive error decay across successive passes.
  • This result lets us determine, before any weight is read, how many passes a layer needs for a target weight-space error.
  • We instantiate ApexQuant with three interchangeable stages, scalar, $E8$ and trellis, and validate it on four open-weight LLMs and on Earth-observation and medical domains where in-distribution data is often unattainable as imagery arrives under restrictive licences or due to patient material under privacy constraints.
  • Progressive re-isotropization comes within a few percent of full precision at four bits and gives the best two-bit arm we measure, in a completely data-free setting.

Sources (1)

  • [1]ApexQuant: Data-Free Elastic Quantization by Residual Re-Isotropization
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 6, 07:49 AM
    We introduce ApexQuant, a calibration-free quantization method that recursively re-quantizes the residual error, serving as a refinement layer on top of existing quantizers.
    We establish that a fresh random rotation returns each residual to the uniform distribution on the hypersphere, which characterizes the rate of progressive error decay across successive passes.

Extractive summary: sentences quoted from the sources.

Before this

  1. Sep 22, 2026vllm-project/vllm v0.30.0
  2. Aug 10, 2026vllm-project/vllm v0.27.0
  3. Jul 11, 2026vllm-project/vllm v0.25.0
  4. Jun 29, 2026vllm-project/vllm v0.24.0
  5. Jun 15, 2026vllm-project/vllm v0.23.0
  6. Jun 10, 2026DiffusionGemma: 4x faster text generation

Related