Controllable-LPMoE: Adapting to Challenging Object Segmentation via Dynamic Local Priors from Mixture-of-Experts
Yanguang Sun, Jiawei Lian, Jian Yang, Lei Luo
TL;DR
This work tackles the inefficiency of full-parameter fine-tuning of large foundation models for binary object segmentation by introducing Controllable-LPMoE, a dynamic priors-based fine-tuning framework. It combines a lightweight Dynamic Mixed Local Priors (DMLP) extractor, which generates task-specific local priors via heterogeneous convolutions and a Mixture-of-Experts gating, with a Bi-directional Interaction (BDI) adapter that uses Cosine-aligned Deformable Attention and Channel-oriented Adaptive Scale Enhancement to transfer information between frozen and trainable features. The approach achieves high segmentation accuracy with only 23.4M trainable parameters, outperforming 31 state-of-the-art methods across 18 datasets spanning COD, SOD, PS, SLS, SD, and GD. This demonstrates a practical and scalable pathway for adapting large-scale foundation models to diverse binary segmentation tasks while significantly reducing training resources and preserving strong universal representations.
Abstract
Large-scale foundation models provide powerful feature representations for downstream object segmentation tasks. However, when adapted to specific tasks through the full-parameter fine-tuning, the enormous parameters being updated often results in significant computational overhead, creating a bottleneck in training efficiency. Although existing methods attempt to fine-tune frozen models by directly embedding trainable prompts, these prompts lack inherent semantic priors, limiting the adaptability of large-scale models. In this paper, we propose a novel dynamic priors-based fine-tuning paradigm with fewer trainable parameters, dubbed Controllable-LPMoE, which adaptively modulates frozen foundation models by dynamically controlling local priors to enhance fine-grained perception for specific segmentation tasks. More specifically, we construct a lightweight dynamic mixed local priors extractor that captures diverse local priors from input images through heterogeneous convolutions while employing a gating network to dynamically output expert priors required for the subsequent fine-tuning. Furthermore, we design a bi-directional interaction adapter that employs cosine-aligned deformable attention and channel-oriented adaptive scale enhancement to interact and restructure between frozen and trainable features, achieving efficient fine-tuning. Extensive experiments validate the superiority of our \href{https://github.com/CSYSI/Controllable-LPMoE} {Controllable-LPMoE} approach, demonstrating excellent segmentation performance compared to 31 state-of-the-art (SOTA) methods and adaptability to multiple binary object segmentation tasks.
