CurveTQ: Rotation-Free Trellis Quantization of LLM Weights via Curvature-Weighted Search
The best two-bit weight quantizers for large language models, such as QTIP and Proteus, rotate each weight matrix by a random orthogonal transform, which must be undone at every decoding step, then encode it with a trellis or lattice code under a Euclidean search; the layer Hessian enters only through error feedback between coding blocks.
ProofPaper ↗
Key points
- We show that this leaves part of the Hessian unused.
- Error feedback turns the loss into a weighted sum of per-coordinate rounding errors whose weights, the diagonal of the Hessian's LDL factorization, existing quantizers compute but never read.
- We put these weights into the Viterbi branch metric, so the search follows the curvature within each coding block.
- Around this search we build CurveTQ, a trellis codec with no rotation, which handles the weights' amplitude and marginal shape with a factored scale field and a closed-form quantile table, and stores a start state per coding block so the trellis can adapt to the residual that error feedback carries into it.
Sources (1)
- [1]CurveTQ: Rotation-Free Trellis Quantization of LLM Weights via Curvature-Weighted SearcharXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 6, 11:14 PM
The best two-bit weight quantizers for large language models, such as QTIP and Proteus, rotate each weight matrix by a random orthogonal transform, which must be undone at every decoding step, then encode it with a trellis or lattice code under a Euclidean search; the layer Hessian enters only through error feedback between coding blocks.
We show that this leaves part of the Hessian unused.
Extractive summary: sentences quoted from the sources.
Before this
- Oct 6, 2026huggingface/transformers v5.19.0: Release v5.19.0
- Aug 10, 2026vllm-project/vllm v0.27.0
- Jul 11, 2026vllm-project/vllm v0.25.0
- Jun 29, 2026vllm-project/vllm v0.24.0
- Jun 15, 2026vllm-project/vllm v0.23.0
- Jun 10, 2026DiffusionGemma: 4x faster text generation