iDETEX: Empowering MLLMs for Intelligent DETailed EXplainable IQA
Zhaoran Zhao, Xinli Yue, Jianhui Sun, Yuhao Xie, Tao Shao, Liangchao Yao, Fan Xia, Yuetang Deng
TL;DR
iDETEX introduces a unified multimodal large language model for detailed, explainable IQA, jointly addressing quality grounding, perception, and description. It leverages three task-specific offline augmentations plus a data mixing strategy and online high-resolution enhancements, with a foundation-model replacement to InternVL3, achieving state-of-the-art results on ViDA-UGC and rank-1 in the MIPI 2025 Detailed IQA Challenge. Key contributions include targeted augmentation modules, a task-aware mixing scheme, and online strategies that improve localization, perceptual sensitivity, and causal description of image quality. The results demonstrate robust, interpretable IQA performance with potential for scalable transfer learning to low-data regimes in diverse distortion settings.
Abstract
Image Quality Assessment (IQA) has progressed from scalar quality prediction to more interpretable, human-aligned evaluation paradigms. In this work, we address the emerging challenge of detailed and explainable IQA by proposing iDETEX-a unified multimodal large language model (MLLM) capable of simultaneously performing three key tasks: quality grounding, perception, and description. To facilitate efficient and generalizable training across these heterogeneous subtasks, we design a suite of task-specific offline augmentation modules and a data mixing strategy. These are further complemented by online enhancement strategies to fully exploit multi-sourced supervision. We validate our approach on the large-scale ViDA-UGC benchmark, where iDETEX achieves state-of-the-art performance across all subtasks. Our model ranks first in the ICCV MIPI 2025 Detailed Image Quality Assessment Challenge, demonstrating its effectiveness and robustness in delivering accurate and interpretable quality assessments.
