AION
Technique

Quantization

Also known as: quantisation, quantized

54stories this week
59last 30 days
80all time

Timeline

  1. Oct 11, 2026 · Opinion / analysis · 1 source
    [P] Pecision models that score every allowed label from the logits: Jebadiah v2.1 (27B, 9B), open weights and self-run benchmark results [P]
    I've been building open models that treat a decision as a closed-set scoring problem rather than text generation.
  2. Oct 11, 2026 · Opinion / analysis · 1 source
    Qwen 3.8 27B Q5 vs Qwen 3.8 Next Q3_S for document analysis
    So far I've been using Qwen 3.8 27B Q5 with a 150K context window, but I'm wondering whether I should switch to Qwen 3.8 Next Q3S, since it has much more knowledge and could extract data much better than the 27B.
  3. Oct 11, 2026 · Tutorial / explainer · 1 source
    OMG! If you have a Mac with 64GB, try Qwen3.8-Flash-Next-oQ4e-mtp with oMLX!
    I was able to run Qwen3.8-Flash-Next-oQ4e-mtp on M3Max 64GB with oMLX!
  4. Oct 10, 2026 · Opinion / analysis · 1 source
    Engineer / developer observations of Gemma4-31B, Qwen3.8-27B, and 6.1-Sol for software engineering work
    Models: Gemma4-31B vs Qwen3.8-27B at the same quantization (an Unsloth flavor of Q4).
  5. Oct 8, 2026 · Research paper · 1 source
    Rounding in Preconditioner Space: Redesigning 4-bit AdamW Optimizer-State Quantization
    Quantizing AdamW's optimizer states reduces persistent storage, but quantization errors propagate through the moment recurrences and perturb subsequent adaptive updates.
  6. Oct 8, 2026 · Research paper · 1 source
    VFold: Symmetry-Aware Cross-Layer Value Cache Compression
    In this work, we propose a symmetry-aware value cache merging strategy that reduces cache memory while avoiding both harmful performance degradation and architectural overhead during decoding.
  7. Oct 8, 2026 · Research paper · 1 source
    DVD: Dynamic Vector Decoding for Efficient MLLM-based Perception
    To address these challenges, we propose a dynamic vector decoding method named DVD, which unifies the representation of 2D and 3D perception tasks.
  8. Oct 8, 2026 · Research paper · 1 source
    AI-Based On-Board Maritime Object Detection for Earth Observation Payload Data Reduction on Versal Embedded Hardware
    We study onboard vessel detection as a way to select what is downlinked, which reduces the data according to its content rather than coding every pixel; it is complementary to conventional onboard compression.
  9. Oct 8, 2026 · Research paper · 1 source
    The Polytopal Neural Network
    We propose Polytopal Neural Networks (PNNs), a framework that extracts distinct layer-wise aspects by enforcing a polytope-based structure that is used directly in subsequent information processing.
  10. Oct 8, 2026 · Research paper · 1 source
    Optimal random quantisers for spherically symmetric distributions
    Zador's celebrated theorem is a cornerstone of optimal quantisation: it establishes both the weak limit of the empirical distribution of an optimal $n$-point quantiser in $R^d$ and the decay rate of the associated $Ls$-mean quantisation error.
  11. Oct 8, 2026 · Research paper · 1 source
    DEX: Digit-Level Early Exit for Energy-Efficient MSDF Neural Network Inference
    U-Net inference for brain-tumor segmentation requires billions of multiply-accumulate operations, motivating hardware that can reduce computation dynamically rather than relying only on fixed precision or static model compression.
  12. Oct 8, 2026 · Research paper · 1 source
    SDPAD: A Fully Spike-Driven Pipeline for End-to-End Autonomous Driving
    We present SDPAD, a fully spike-driven end-to-end planning pipeline that closes this gap.
  13. Oct 8, 2026 · Research paper · 1 source
    Evaluating Local Language Model Agents for Reproducible Data Engineering: An Empirical Software Engineering Study of Mobility Workflows
    Context: Large language model (LLM) agents are increasingly used as software and data-engineering assistants, yet evidence about locally deployable open-weight agents remains limited.
  14. Oct 8, 2026 · Research paper · 1 source
    FlyMark: Training-Free Invisible Watermarking of 3D Gaussian Splatting via a Fruit Fly Connectome
    A trained 3D Gaussian Splatting (3DGS) scene ships as a portable parameter array that can be copied, pruned, requantized, or repackaged outside its training pipeline, so ownership evidence is most useful when it lives in the released parameters and remains checkable long after the embedding tooling is gone.
  15. Oct 8, 2026 · Research paper · 1 source
    Deflating the Hessian: Rank-4 W4A4 Quantization for Multimodal Diffusion Transformers
    In diffusion transformers, low-rank branches can mitigate 4-bit weight--activation (W4A4) post-training quantization (PTQ) loss by decomposing each weight into a low-bit residual and a high-precision low-rank component.
  16. Oct 8, 2026 · Research paper · 1 source
    Read What Matters: Query-Adaptive Quantization for KV Caches
    KV-cache entries are stored before their future queries are known, but each decoding query needs precision in different places.
  17. Oct 8, 2026 · Research paper · 1 source
    When Lower Reconstruction Loss Hurts: Distributionally Robust Refinement for Low-Bit LLM Quantization
    Motivated by these observations and our analysis, we propose Distributionally Robust Quantization (DRQ), a post-hoc refinement process that minimizes worst-case reconstruction loss over a constrained set of input activation distributions.
  18. Oct 8, 2026 · Research paper · 1 source
    Bridging KV-Cache Quantization and Linear Attention: From Theory to Pretrained Weight Migration
    KV-cache quantization compresses individual KV entries into discrete codes but retains all entries, whereas linear attention recurrently aggregates multiple historical KV contributions into a fixed-size continuous state but can introduce interference.
  19. Oct 8, 2026 · Research paper · 1 source
    Dynamics as Code: On Model Compression via Dynamic System
    With the irrational winding as an example, earlier work introduced a dynamic system (DS) paradigm that reconceptualizes compression as compact weight representation: high-dimensional parameters are encoded by the index of a trajectory produced by a dynamic system, from which the vector is recovered during decompression.
  20. Oct 7, 2026 · Research paper · 1 source
    Rethinking the Tradeoff Between Temporal Encoding and Nonlinear Computation in Spiking Language Models
    Spiking language models face a tradeoff between representing continuous semantic features over short temporal windows and retaining costly nonlinear attention operations.
  21. Oct 7, 2026 · Research paper · 1 source
    Lossy Compressive Text Autoencoders
    Our work explores learning a compressed latent representation of text, at the intersection of data compression and representation learning.
  22. Oct 7, 2026 · Research paper · 1 source
    OrBIT: Structure-Guided Embedding Compression
    We introduce OrBIT, a structure-guided embedding compression framework that learns reusable local geometry from orbit dynamics and uses it to constrain a small set of shared codewords.
  23. Oct 7, 2026 · Research paper · 1 source
    ResidualQuant: KV Cache Quantization for Looped Transformers with 2-Bit Residuals
    Looped Transformers improve parameter efficiency by repeatedly applying shared Transformer blocks over multiple recurrent loops, increasing computational depth without increasing the parameter count.
  24. Oct 7, 2026 · Research paper · 1 source
    Open-MMUnlearning: Unifying Methods and Evaluation for MLLM Unlearning
    We introduce Open-MMUnlearning, an open-source, extensible framework that integrates target-model preparation, multimodal data processing, unlearning, and evaluation through shared interfaces and structured configurations.
  25. Oct 7, 2026 · Research paper · 1 source
    Real-Time Joint Audio-Video Generation by Parallel Adapter Composition
    Deploying a joint audio-video diffusion transformer for real-time, interactive generation normally requires two essential modifications: block-autoregressive attention, so frames can be emitted before the whole clip is finished, and few-step sampling, so each block is cheap.
  26. Oct 7, 2026 · Research paper · 1 source
    Coverage-Aware Reasoning with Medical Tokens for Diagnosis Prediction
    Large language models (LLMs) offer promising potential for next-visit diagnosis prediction, owing to their ability to integrate longitudinal clinical evidence and reason over it in natural language.
  27. Oct 7, 2026 · Research paper · 1 source
    Cache the Encoder Within:Compact, Reusable Memory across LLM Queries
    Repeated queries over shared documents incur redundant encoding, while caching model states introduces persistent storage costs.
  28. Oct 7, 2026 · Research paper · 1 source
    Phase-HDC: Replacing Optimizer History with Gradient Thresholds in Discrete Phase Learning
    For a hyperdimensional classifier whose learned parameters are low-bit angles, which we call a phase memory, these records take several times more memory than the model itself.
  29. Oct 7, 2026 · Research paper · 1 source
    TR-PTQ: High-Accuracy Integer-Only Transformer Post Training Quantization via Taylor Region Reformulation
    Post-training quantization (PTQ) enables efficient deployment, yet transformer architectures remain challenging to quantize due to nonlinear layers.
  30. Oct 7, 2026 · Research paper · 1 source
    Layerwise Error Attribution for Fast and Robust Mixed-Precision Post-Training Quantization
    Mixed-precision post-training quantization is a network compression method that assigns bits layer by layer, under a global memory budget using a small calibration set.

Often appears with