ResearchResearch paperTraining & Scaling1 source · Oct 7, 2026

Derivative Gaussian Processes on a Two-Direction Budget

We propose a derivative GP with a budget of just two directions per observed gradient.

Key points

  • Gradient observations promise more accurate Gaussian process (GP) surrogates, but the cost of incorporating them has long stood in the way of realizing that promise.
  • One direction focuses on each gradient's direct contribution to target prediction, while the other aggregates its indirect contributions through correlations with the conditioning function values.
  • Within a Vecchia approximation, where each prediction conditions on $m$ nearby inputs in $d$ dimensions, this construction represents their $md$ gradient coordinates using at most $2m$ directional derivatives, giving $\mathcal{O}(m^3)$ dense factorization cost per prediction target.
  • For general conditioning sets, we bound the posterior approximation error relative to using full gradients and characterize when the error is small or the approximation is exact.

Sources (1)

  • [1]Derivative Gaussian Processes on a Two-Direction Budget
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 7, 05:08 PM
    We propose a derivative GP with a budget of just two directions per observed gradient.
    Gradient observations promise more accurate Gaussian process (GP) surrogates, but the cost of incorporating them has long stood in the way of realizing that promise.

Extractive summary: sentences quoted from the sources.

Related