AION
Research paperLarge Language Models · Efficiency & Inference1 source · Oct 7, 2026

Rethinking the Tradeoff Between Temporal Encoding and Nonlinear Computation in Spiking Language Models

Spiking language models face a tradeoff between representing continuous semantic features over short temporal windows and retaining costly nonlinear attention operations.

Key points

  • We introduce Spora, which jointly designs spike encodings and attention operators.
  • Binary temporal weights let $T$ spikes represent compositional values with up to $T$ bits of capacity, compared with $O(\log2 T)$ bits for spike-count readout.
  • Unipolar Binary Spiking (UBS) uses thresholds and spike-triggered residual decay to produce non-negative integer codes; Bipolar Binary Spiking (BBS) separates sign and magnitude and learns a scale for signed activations.
  • Conditional-decay analysis, matched-budget activation-quantization comparisons, event-workload statistics, and fixed-point evaluation further characterize the connection between encoding fidelity and computational cost.

Sources (1)

Extractive summary: sentences quoted from the sources.