Is In-Domain Training Enough for Fine-Grained Industrial Anomaly Understanding?
A single multimodal large language model (MLLM) struggles to excel simultaneously at detection, localization, description, and reasoning in multimodal industrial anomaly understanding (MM-IAU).
Key points
- We show that in-domain training does not close this gap.
- On MMAD, a widely adopted MM-IAU benchmark, trained specialists reach at most 75.5% accuracy in defect localization, against 92.3% for human experts, and even detect anomalies less accurately than their untrained base model.
- We therefore propose SiGMA, a spatially grounded multi-agent framework that divides labor between heterogeneous MLLM agents and a dedicated visual defect expert.
- A multimodal searcher supplies industrial knowledge and normal references, the defect expert turns query-reference comparison into calibrated anomaly evidence, and a label-free reliability controller weighs each source by task-wise competence and query-level evidence quality.
Sources (1)
- [1]Is In-Domain Training Enough for Fine-Grained Industrial Anomaly Understanding?arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 8, 04:59 PM
A single multimodal large language model (MLLM) struggles to excel simultaneously at detection, localization, description, and reasoning in multimodal industrial anomaly understanding (MM-IAU).
We show that in-domain training does not close this gap.
Extractive summary: sentences quoted from the sources.