AION
Research paperRobotics & Embodied AI · Interpretability · Large Language Models1 source · Oct 8, 2026

Is In-Domain Training Enough for Fine-Grained Industrial Anomaly Understanding?

A single multimodal large language model (MLLM) struggles to excel simultaneously at detection, localization, description, and reasoning in multimodal industrial anomaly understanding (MM-IAU).

Key points

  • We show that in-domain training does not close this gap.
  • On MMAD, a widely adopted MM-IAU benchmark, trained specialists reach at most 75.5% accuracy in defect localization, against 92.3% for human experts, and even detect anomalies less accurately than their untrained base model.
  • We therefore propose SiGMA, a spatially grounded multi-agent framework that divides labor between heterogeneous MLLM agents and a dedicated visual defect expert.
  • A multimodal searcher supplies industrial knowledge and normal references, the defect expert turns query-reference comparison into calibrated anomaly evidence, and a label-free reliability controller weighs each source by task-wise competence and query-level evidence quality.

Sources (1)

  • [1]Is In-Domain Training Enough for Fine-Grained Industrial Anomaly Understanding?
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 8, 04:59 PM
    A single multimodal large language model (MLLM) struggles to excel simultaneously at detection, localization, description, and reasoning in multimodal industrial anomaly understanding (MM-IAU).
    We show that in-domain training does not close this gap.

Extractive summary: sentences quoted from the sources.