ProtoSemImage: Image-Valued Prototypes with Deformable Row Alignment for Interpretable Document Classification
A Skip-Gram objective learns that color space end to end through a four-dimensional bottleneck, discourse boundary rows become differentiable typed difference rows, and classification reduces to 2D visual template matching: a deformable row alignment between a document image and the archetype bank, in the spirit of dynamic time warping.
ProofPaper ↗
Key points
- Prototypes in classification models are almost always vectors, and a vector has no readable form.
- This paper asks what happens when a prototype is an image.
- Documents give the question a natural form, because a document can be rendered as a multi-channel image in which every token becomes a pixel, so a class representative can take the same shape and the same channel semantics as the inputs it stands for.
- ProtoSemImage represents each class by one or more visual archetypes: prototype images in a four-channel HSV space whose channels carry named linguistic factors.
Sources (1)
- [1]ProtoSemImage: Image-Valued Prototypes with Deformable Row Alignment for Interpretable Document ClassificationarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 8, 08:12 AM
A Skip-Gram objective learns that color space end to end through a four-dimensional bottleneck, discourse boundary rows become differentiable typed difference rows, and classification reduces to 2D visual template matching: a deformable row alignment between a document image and the archetype bank, in the spirit of dynamic time warping.
Prototypes in classification models are almost always vectors, and a vector has no readable form.
Extractive summary: sentences quoted from the sources.