ResearchResearch paperLarge Language Models1 source · Oct 6, 2026

WASD: Wasserstein-based Knowledge Distillation for Large Language Models

We propose Wasserstein-based knowledge distillation (WASD) for LLMs, which incorporates token-level semantic information via the Wasserstein-based distance with a cost matrix derived from token embeddings.

Key points

  • Autoregressive large language models (LLMs) have rapidly advanced in capability, but their increasing scale comes with substantial computational and memory costs at inference time.
  • Knowledge distillation (KD) offers a practical solution by transferring knowledge from a large teacher model to a smaller student model via alignment of discrete probability distributions.
  • To ensure computational tractability, we adopt the Sinkhorn divergence and derive a gradient-equivalent objective that can be efficiently optimized without introducing additional networks.
  • Our results highlight the importance of semantic information encoded in the token space for effective distribution alignment in LLM distillation.

Sources (1)

  • [1]WASD: Wasserstein-based Knowledge Distillation for Large Language Models
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 6, 03:58 AM
    We propose Wasserstein-based knowledge distillation (WASD) for LLMs, which incorporates token-level semantic information via the Wasserstein-based distance with a cost matrix derived from token embeddings.
    Autoregressive large language models (LLMs) have rapidly advanced in capability, but their increasing scale comes with substantial computational and memory costs at inference time.

Extractive summary: sentences quoted from the sources.

Before this

  1. Oct 6, 2026Self-Retrospection Distillation: Turning Post-hoc Experiences into Prior Foresight
  2. Oct 6, 2026RISED: Rubrics for Agentic Multi-Environment Selection and Self-Distillation
  3. Sep 29, 2026[AINews] AMD buys World Labs for $8.2B, as Atlas solves sparse reconstruction problem for robotics, design and more
  4. Sep 29, 2026Salt++: Context-Aligned Post-Training for Few-Step Streaming Multimodal Generation
  5. Sep 28, 2026Notes on NVIDIA Nemotron
  6. Aug 10, 2026huggingface/transformers v5.15.0: Release: v5.15.0

Related