AION
Research paperComputer Vision · Interpretability1 source · Oct 7, 2026

DisParQ: Self-Supervised Part Concepts for Interpretable Vision Foundation Models

We introduce DisParQ (Discrete Parts with Quantized attributes), a method that learns spatially grounded, discrete concept representations from a powerful frozen vision-only self-supervised backbone.

Key points

  • Concept-based vision models represent images through an intermediate layer of human-inspectable concepts, so what a model relies on can be traced to those concepts.
  • To capture how each concept varies across images (e.g., the type of a "wheel"), we learn continuous residuals alongside the concepts and then quantize them into discrete attributes.
  • We evaluate DisParQ across seven datasets, from general recognition (ImageNet, PartImageNet, Places) to fine-grained benchmarks (CUB, Cars, Dogs, Flowers).
  • We show that DisParQ closely matches its frozen DINOv2 teacher on ImageNet linear probing (83.2% top-1), achieves higher concept consistency than language-aligned models, remains competitive on fine-grained recognition, and enables cross-category part-based retrieval.

Sources (1)

  • [1]DisParQ: Self-Supervised Part Concepts for Interpretable Vision Foundation Models
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 7, 10:19 AM
    We introduce DisParQ (Discrete Parts with Quantized attributes), a method that learns spatially grounded, discrete concept representations from a powerful frozen vision-only self-supervised backbone.
    Concept-based vision models represent images through an intermediate layer of human-inspectable concepts, so what a model relies on can be traced to those concepts.

Extractive summary: sentences quoted from the sources.