Quantization
Also known as: quantisation, quantized
54stories this week
59last 30 days
80all time
Timeline
- Oct 11, 2026 · Opinion / analysis · 1 source[P] Pecision models that score every allowed label from the logits: Jebadiah v2.1 (27B, 9B), open weights and self-run benchmark results [P]I've been building open models that treat a decision as a closed-set scoring problem rather than text generation.
- Oct 11, 2026 · Opinion / analysis · 1 sourceQwen 3.8 27B Q5 vs Qwen 3.8 Next Q3_S for document analysisSo far I've been using Qwen 3.8 27B Q5 with a 150K context window, but I'm wondering whether I should switch to Qwen 3.8 Next Q3S, since it has much more knowledge and could extract data much better than the 27B.
- Oct 11, 2026 · Tutorial / explainer · 1 sourceOMG! If you have a Mac with 64GB, try Qwen3.8-Flash-Next-oQ4e-mtp with oMLX!I was able to run Qwen3.8-Flash-Next-oQ4e-mtp on M3Max 64GB with oMLX!
- Oct 10, 2026 · Opinion / analysis · 1 sourceEngineer / developer observations of Gemma4-31B, Qwen3.8-27B, and 6.1-Sol for software engineering workModels: Gemma4-31B vs Qwen3.8-27B at the same quantization (an Unsloth flavor of Q4).
- Oct 8, 2026 · Research paper · 1 sourceRounding in Preconditioner Space: Redesigning 4-bit AdamW Optimizer-State QuantizationQuantizing AdamW's optimizer states reduces persistent storage, but quantization errors propagate through the moment recurrences and perturb subsequent adaptive updates.
- Oct 8, 2026 · Research paper · 1 sourceVFold: Symmetry-Aware Cross-Layer Value Cache CompressionIn this work, we propose a symmetry-aware value cache merging strategy that reduces cache memory while avoiding both harmful performance degradation and architectural overhead during decoding.
- Oct 8, 2026 · Research paper · 1 sourceDVD: Dynamic Vector Decoding for Efficient MLLM-based PerceptionTo address these challenges, we propose a dynamic vector decoding method named DVD, which unifies the representation of 2D and 3D perception tasks.
- Oct 8, 2026 · Research paper · 1 sourceAI-Based On-Board Maritime Object Detection for Earth Observation Payload Data Reduction on Versal Embedded HardwareWe study onboard vessel detection as a way to select what is downlinked, which reduces the data according to its content rather than coding every pixel; it is complementary to conventional onboard compression.
- Oct 8, 2026 · Research paper · 1 sourceThe Polytopal Neural NetworkWe propose Polytopal Neural Networks (PNNs), a framework that extracts distinct layer-wise aspects by enforcing a polytope-based structure that is used directly in subsequent information processing.
- Oct 8, 2026 · Research paper · 1 sourceOptimal random quantisers for spherically symmetric distributionsZador's celebrated theorem is a cornerstone of optimal quantisation: it establishes both the weak limit of the empirical distribution of an optimal $n$-point quantiser in $R^d$ and the decay rate of the associated $Ls$-mean quantisation error.
- Oct 8, 2026 · Research paper · 1 sourceDEX: Digit-Level Early Exit for Energy-Efficient MSDF Neural Network InferenceU-Net inference for brain-tumor segmentation requires billions of multiply-accumulate operations, motivating hardware that can reduce computation dynamically rather than relying only on fixed precision or static model compression.
- Oct 8, 2026 · Research paper · 1 sourceSDPAD: A Fully Spike-Driven Pipeline for End-to-End Autonomous DrivingWe present SDPAD, a fully spike-driven end-to-end planning pipeline that closes this gap.
- Oct 8, 2026 · Research paper · 1 sourceEvaluating Local Language Model Agents for Reproducible Data Engineering: An Empirical Software Engineering Study of Mobility WorkflowsContext: Large language model (LLM) agents are increasingly used as software and data-engineering assistants, yet evidence about locally deployable open-weight agents remains limited.
- Oct 8, 2026 · Research paper · 1 sourceFlyMark: Training-Free Invisible Watermarking of 3D Gaussian Splatting via a Fruit Fly ConnectomeA trained 3D Gaussian Splatting (3DGS) scene ships as a portable parameter array that can be copied, pruned, requantized, or repackaged outside its training pipeline, so ownership evidence is most useful when it lives in the released parameters and remains checkable long after the embedding tooling is gone.
- Oct 8, 2026 · Research paper · 1 sourceDeflating the Hessian: Rank-4 W4A4 Quantization for Multimodal Diffusion TransformersIn diffusion transformers, low-rank branches can mitigate 4-bit weight--activation (W4A4) post-training quantization (PTQ) loss by decomposing each weight into a low-bit residual and a high-precision low-rank component.
- Oct 8, 2026 · Research paper · 1 sourceRead What Matters: Query-Adaptive Quantization for KV CachesKV-cache entries are stored before their future queries are known, but each decoding query needs precision in different places.
- Oct 8, 2026 · Research paper · 1 sourceWhen Lower Reconstruction Loss Hurts: Distributionally Robust Refinement for Low-Bit LLM QuantizationMotivated by these observations and our analysis, we propose Distributionally Robust Quantization (DRQ), a post-hoc refinement process that minimizes worst-case reconstruction loss over a constrained set of input activation distributions.
- Oct 8, 2026 · Research paper · 1 sourceBridging KV-Cache Quantization and Linear Attention: From Theory to Pretrained Weight MigrationKV-cache quantization compresses individual KV entries into discrete codes but retains all entries, whereas linear attention recurrently aggregates multiple historical KV contributions into a fixed-size continuous state but can introduce interference.
- Oct 8, 2026 · Research paper · 1 sourceDynamics as Code: On Model Compression via Dynamic SystemWith the irrational winding as an example, earlier work introduced a dynamic system (DS) paradigm that reconceptualizes compression as compact weight representation: high-dimensional parameters are encoded by the index of a trajectory produced by a dynamic system, from which the vector is recovered during decompression.
- Oct 7, 2026 · Research paper · 1 sourceRethinking the Tradeoff Between Temporal Encoding and Nonlinear Computation in Spiking Language ModelsSpiking language models face a tradeoff between representing continuous semantic features over short temporal windows and retaining costly nonlinear attention operations.
- Oct 7, 2026 · Research paper · 1 sourceLossy Compressive Text AutoencodersOur work explores learning a compressed latent representation of text, at the intersection of data compression and representation learning.
- Oct 7, 2026 · Research paper · 1 sourceOrBIT: Structure-Guided Embedding CompressionWe introduce OrBIT, a structure-guided embedding compression framework that learns reusable local geometry from orbit dynamics and uses it to constrain a small set of shared codewords.
- Oct 7, 2026 · Research paper · 1 sourceResidualQuant: KV Cache Quantization for Looped Transformers with 2-Bit ResidualsLooped Transformers improve parameter efficiency by repeatedly applying shared Transformer blocks over multiple recurrent loops, increasing computational depth without increasing the parameter count.
- Oct 7, 2026 · Research paper · 1 sourceOpen-MMUnlearning: Unifying Methods and Evaluation for MLLM UnlearningWe introduce Open-MMUnlearning, an open-source, extensible framework that integrates target-model preparation, multimodal data processing, unlearning, and evaluation through shared interfaces and structured configurations.
- Oct 7, 2026 · Research paper · 1 sourceReal-Time Joint Audio-Video Generation by Parallel Adapter CompositionDeploying a joint audio-video diffusion transformer for real-time, interactive generation normally requires two essential modifications: block-autoregressive attention, so frames can be emitted before the whole clip is finished, and few-step sampling, so each block is cheap.
- Oct 7, 2026 · Research paper · 1 sourceCoverage-Aware Reasoning with Medical Tokens for Diagnosis PredictionLarge language models (LLMs) offer promising potential for next-visit diagnosis prediction, owing to their ability to integrate longitudinal clinical evidence and reason over it in natural language.
- Oct 7, 2026 · Research paper · 1 sourceCache the Encoder Within:Compact, Reusable Memory across LLM QueriesRepeated queries over shared documents incur redundant encoding, while caching model states introduces persistent storage costs.
- Oct 7, 2026 · Research paper · 1 sourcePhase-HDC: Replacing Optimizer History with Gradient Thresholds in Discrete Phase LearningFor a hyperdimensional classifier whose learned parameters are low-bit angles, which we call a phase memory, these records take several times more memory than the model itself.
- Oct 7, 2026 · Research paper · 1 sourceTR-PTQ: High-Accuracy Integer-Only Transformer Post Training Quantization via Taylor Region ReformulationPost-training quantization (PTQ) enables efficient deployment, yet transformer architectures remain challenging to quantize due to nonlinear layers.
- Oct 7, 2026 · Research paper · 1 sourceLayerwise Error Attribution for Fast and Robust Mixed-Precision Post-Training QuantizationMixed-precision post-training quantization is a network compression method that assigns bits layer by layer, under a global memory budget using a small calibration set.