Common-Mode Errors Limit Low-Timestep Deep Spiking Q-Networks
Based on this finding, we propose Common-Mode Compensation Deep Spiking Q-Network (CMC-DSQN), which uses an auxiliary ANN to compensate for common-mode errors in the SNN outputs.
ProofPaper ↗
Key points
- Spiking neural networks (SNNs) offer sparse and event-driven computation, making them attractive for energy-constrained reinforcement learning (RL) on edge devices.
- In value-based RL, deep spiking Q-networks (DSQNs) combine such efficiency with action-value estimation for decision making.
- We investigate this degradation from the perspective of Q-value estimation errors.
- By decomposing errors across actions into common-mode and differential-mode components, we find that low-timestep DSQNs suffer disproportionately from common-mode errors shared across action values, which are particularly detrimental to temporal-difference learning through bootstrapped targets.
Sources (1)
- [1]Common-Mode Errors Limit Low-Timestep Deep Spiking Q-NetworksarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 6, 06:01 AM
Based on this finding, we propose Common-Mode Compensation Deep Spiking Q-Network (CMC-DSQN), which uses an auxiliary ANN to compensate for common-mode errors in the SNN outputs.
Spiking neural networks (SNNs) offer sparse and event-driven computation, making them attractive for energy-constrained reinforcement learning (RL) on edge devices.
Extractive summary: sentences quoted from the sources.