Safe Meta-Policy Design with Risk Control
We study how to plan policy updates (i.e., meta-policy) before future candidates are trained, balancing the benefits of improvement against the risk of performance regression.
Key points
- Models can be retrained as new data arrive, but deploying every new version risks replacing a good policy with a worse one.
- We estimate the value and risk of possible switches from historical learning trajectories, represent an update schedule as a path in a directed acyclic graph, and select a schedule using dynamic programming.
- A leading-order analysis identifies the signal-to-noise ratio of policy improvement as a key driver of update frequency, waiting times, and risk allocation: clearer improvements support earlier, more frequent updates, while noisier improvements call for longer waits or greater risk expenditure.
- Experiments on synthetic and clinical trial data illustrate the performance--risk tradeoff and compare our method with alternative baselines.
Sources (1)
- [1]Safe Meta-Policy Design with Risk ControlarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 7, 04:48 PM
We study how to plan policy updates (i.e., meta-policy) before future candidates are trained, balancing the benefits of improvement against the risk of performance regression.
Models can be retrained as new data arrive, but deploying every new version risks replacing a good policy with a worse one.
Extractive summary: sentences quoted from the sources.