Rethinking the Tradeoff Between Temporal Encoding and Nonlinear Computation in Spiking Language Models
Spiking language models face a tradeoff between representing continuous semantic features over short temporal windows and retaining costly nonlinear attention operations.
Key points
- We introduce Spora, which jointly designs spike encodings and attention operators.
- Binary temporal weights let $T$ spikes represent compositional values with up to $T$ bits of capacity, compared with $O(\log2 T)$ bits for spike-count readout.
- Unipolar Binary Spiking (UBS) uses thresholds and spike-triggered residual decay to produce non-negative integer codes; Bipolar Binary Spiking (BBS) separates sign and magnitude and learns a scale for signed activations.
- Conditional-decay analysis, matched-budget activation-quantization comparisons, event-workload statistics, and fixed-point evaluation further characterize the connection between encoding fidelity and computational cost.
Sources (1)
- [1]Rethinking the Tradeoff Between Temporal Encoding and Nonlinear Computation in Spiking Language ModelsarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 7, 09:41 PM
Spiking language models face a tradeoff between representing continuous semantic features over short temporal windows and retaining costly nonlinear attention operations.
We introduce Spora, which jointly designs spike encodings and attention operators.
Extractive summary: sentences quoted from the sources.