Collaboratively Guided Adversarial Robust Distillation with Teacher-Favorable Examples
Adversarial distillation transfers robustness from high-capacity teachers to compact students.
ProofPaper ↗
Key points
- Existing adversarial distillation methods mainly use teacher predictions on clean or adversarial examples to supervise student learning.
- However, teacher-favorable supervision within the perturbation neighborhood remains underexplored in adversarial distillation.
- We therefore propose Collaboratively Guided Adversarial Robust Distillation (CGARD), which jointly optimizes distinct student-adversarial and teacher-collaborative examples within the same perturbation neighborhood.
- CGARD combines collaborative teacher guidance with adversarial teacher supervision to improve robust knowledge transfer.
Sources (1)
- [1]Collaboratively Guided Adversarial Robust Distillation with Teacher-Favorable ExamplesarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 8, 06:12 AM
Adversarial distillation transfers robustness from high-capacity teachers to compact students.
Existing adversarial distillation methods mainly use teacher predictions on clean or adversarial examples to supervise student learning.
Extractive summary: sentences quoted from the sources.
Before this
- Oct 8, 2026One Block, Multiple Depths: Recurrent Vision Transformers with Depth-Programmed Experts
- Oct 8, 2026SpatialOPSD: Self-Distilling Spatial Intelligence from Verified Coding Agent Traces
- Oct 8, 2026Distilling Routed 3D Privilege for Spatial Reasoning in Vision-Language Models
- Oct 7, 2026MIMESIS: Learning User Simulators as Training Environments for Interactive Agents
- Oct 7, 2026Composing What Each Teacher Learned: Multi-Teacher On-Policy Distillation through Teacher-Relative Shifts
- Oct 7, 2026On-Policy Distillation Teaches New Skills but Not New Knowledge