ResearchResearch paperComputer Vision · Large Language Models · Interpretability1 source · Oct 8, 2026

ProtoSemImage: Image-Valued Prototypes with Deformable Row Alignment for Interpretable Document Classification

A Skip-Gram objective learns that color space end to end through a four-dimensional bottleneck, discourse boundary rows become differentiable typed difference rows, and classification reduces to 2D visual template matching: a deformable row alignment between a document image and the archetype bank, in the spirit of dynamic time warping.

Key points

  • Prototypes in classification models are almost always vectors, and a vector has no readable form.
  • This paper asks what happens when a prototype is an image.
  • Documents give the question a natural form, because a document can be rendered as a multi-channel image in which every token becomes a pixel, so a class representative can take the same shape and the same channel semantics as the inputs it stands for.
  • ProtoSemImage represents each class by one or more visual archetypes: prototype images in a four-channel HSV space whose channels carry named linguistic factors.

Sources (1)

  • [1]ProtoSemImage: Image-Valued Prototypes with Deformable Row Alignment for Interpretable Document Classification
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 8, 08:12 AM
    A Skip-Gram objective learns that color space end to end through a four-dimensional bottleneck, discourse boundary rows become differentiable typed difference rows, and classification reduces to 2D visual template matching: a deformable row alignment between a document image and the archetype bank, in the spirit of dynamic time warping.
    Prototypes in classification models are almost always vectors, and a vector has no readable form.

Extractive summary: sentences quoted from the sources.

Related