ResearchResearch paperTraining & Scaling · Efficiency & Inference · Reinforcement Learning1 source · Oct 6, 2026

Extending Pathwise Gradients to Discrete Random Variables via Finite-Order Relaxation

We propose a general framework to construct finite-order exact pathwise gradient estimators for a range of common discrete variables such as Poisson.

Key points

  • Pathwise gradients are preferred for continuous random variables because they are unbiased, low variance, and work with a single sample.
  • For discrete variables, however, the pathwise identity cannot generally be exact for every differentiable function.
  • Against other admissible solutions, our estimator is unique and minimizes weight variance; in contrast, prior works use categorical variables or augmented representations to approximate non-categorical variables that induces excess variance and computations.
  • To understand approximation bias for functions beyond the prescribed class, we also derive a non-asymptotic bias bound.

Sources (1)

  • [1]Extending Pathwise Gradients to Discrete Random Variables via Finite-Order Relaxation
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 6, 05:30 AM
    We propose a general framework to construct finite-order exact pathwise gradient estimators for a range of common discrete variables such as Poisson.
    Pathwise gradients are preferred for continuous random variables because they are unbiased, low variance, and work with a single sample.

Extractive summary: sentences quoted from the sources.

Related