Policy Alignment: New Signals for Membership Auditing in On-Policy Distillation
In this paper, we propose Policy Alignment Membership Auditing (PAMA), a new auditing framework tailored for OPD.
Key points
- On-policy distillation (OPD) trains a student model by aligning its policy with a teacher model on trajectories generated by the student model itself.
- Our key observation is that a member prompt directly contributes to the teacher-guided policy update, while a non-member prompt only experiences indirect effects through cross-prompt generalization.
- Specifically, we introduce Teacher Alignment Gain (TAG) to estimate the teacher-aligned update direction from model outputs, and further combine it with student drift and uncertainty alignment signals for reliable membership auditing.
- We evaluate PAMA on six datasets and three teacher-student model families.
Sources (1)
- [1]Policy Alignment: New Signals for Membership Auditing in On-Policy DistillationarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 8, 07:48 AM
In this paper, we propose Policy Alignment Membership Auditing (PAMA), a new auditing framework tailored for OPD.
On-policy distillation (OPD) trains a student model by aligning its policy with a teacher model on trajectories generated by the student model itself.
Extractive summary: sentences quoted from the sources.