ResearchResearch paperEfficiency & Inference1 source · Oct 6, 2026

CurveTQ: Rotation-Free Trellis Quantization of LLM Weights via Curvature-Weighted Search

The best two-bit weight quantizers for large language models, such as QTIP and Proteus, rotate each weight matrix by a random orthogonal transform, which must be undone at every decoding step, then encode it with a trellis or lattice code under a Euclidean search; the layer Hessian enters only through error feedback between coding blocks.

Key points

  • We show that this leaves part of the Hessian unused.
  • Error feedback turns the loss into a weighted sum of per-coordinate rounding errors whose weights, the diagonal of the Hessian's LDL factorization, existing quantizers compute but never read.
  • We put these weights into the Viterbi branch metric, so the search follows the curvature within each coding block.
  • Around this search we build CurveTQ, a trellis codec with no rotation, which handles the weights' amplitude and marginal shape with a factored scale field and a closed-form quantile table, and stores a start state per coding block so the trellis can adapt to the residual that error feedback carries into it.

Sources (1)

  • [1]CurveTQ: Rotation-Free Trellis Quantization of LLM Weights via Curvature-Weighted Search
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 6, 11:14 PM
    The best two-bit weight quantizers for large language models, such as QTIP and Proteus, rotate each weight matrix by a random orthogonal transform, which must be undone at every decoding step, then encode it with a trellis or lattice code under a Euclidean search; the layer Hessian enters only through error feedback between coding blocks.
    We show that this leaves part of the Hessian unused.

Extractive summary: sentences quoted from the sources.

Before this

  1. Oct 6, 2026huggingface/transformers v5.19.0: Release v5.19.0
  2. Aug 10, 2026vllm-project/vllm v0.27.0
  3. Jul 11, 2026vllm-project/vllm v0.25.0
  4. Jun 29, 2026vllm-project/vllm v0.24.0
  5. Jun 15, 2026vllm-project/vllm v0.23.0
  6. Jun 10, 2026DiffusionGemma: 4x faster text generation

Related