ResearchResearch paperLarge Language Models · Retrieval, RAG & Search · Interpretability1 source · Oct 6, 2026

Quantize by Drift: Label-Free Mixed-Precision Post-Training Quantization for Text Embedders

Mixed-precision post-training quantization needs a per-module sensitivity signal; for a text embedder the obvious one -- the retrieval quality a module costs when quantized -- needs relevance labels that deployments rarely have.

Key points

  • We measure a label-free substitute: quantization-induced representation drift, obtained by quantizing one module, re-encoding the corpus, and recording how far the output embeddings moved from their full-precision positions.
  • What is specific is the observable: the deployed output representation a dense retriever ranks with.
  • Across five development embedders, configuration-level drift orders sampled mixed-precision plans against held-out retrieval quality at a macro Spearman of 0.911, the sensitivity transports across calibration corpora and retrieval domains in the usable regime, module drifts compose rank-consistently but not numerically, and relevance-derived sensitivity adds no consistent value.
  • Output drift is thus a robust coarse sensitivity signal, not a universally optimal allocation objective: it avoids the catastrophic failures of the transferred signed-geometry adaptation and can remain usable at stressed budgets where uniform collapses, but fine-grained redistribution around a strong uniform operating point remains unresolved.

Sources (1)

  • [1]Quantize by Drift: Label-Free Mixed-Precision Post-Training Quantization for Text Embedders
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 6, 11:44 PM
    Mixed-precision post-training quantization needs a per-module sensitivity signal; for a text embedder the obvious one -- the retrieval quality a module costs when quantized -- needs relevance labels that deployments rarely have.
    We measure a label-free substitute: quantization-induced representation drift, obtained by quantizing one module, re-encoding the corpus, and recording how far the output embeddings moved from their full-precision positions.

Extractive summary: sentences quoted from the sources.

Before this

  1. Oct 6, 2026huggingface/transformers v5.19.0: Release v5.19.0
  2. Oct 6, 2026EmbeddingGemma 2: an open, lightweight multimodal embedding model
  3. Sep 22, 2026vllm-project/vllm v0.30.0
  4. Aug 10, 2026vllm-project/vllm v0.27.0
  5. Jul 11, 2026vllm-project/vllm v0.25.0
  6. Jun 10, 2026DiffusionGemma: 4x faster text generation

Related