WISE: Weighted Iterative Society-of-Experts for Robust Multimodal Multi-Agent Debate

Anoop Cherian; River Doyle; Eyal Ben-Dov; Suhas Lohit; Kuan-Chuan Peng

WISE: Weighted Iterative Society-of-Experts for Robust Multimodal Multi-Agent Debate

Anoop Cherian, River Doyle, Eyal Ben-Dov, Suhas Lohit, Kuan-Chuan Peng

TL;DR

This paper investigates robust multimodal reasoning through multi-agent debate by introducing WISE, a modular framework that assigns agents to Solver and Reflector roles and uses an orchestrator to coordinate iterative reasoning. It extends the Dawid–Skene aggregation to jointly estimate solver and reflector error models, enabling principled consensus across debate rounds. Across Vision-Language reasoning benchmarks (SMART-840, VisualPuzzles, EvoChart-QA, and SMART-840++), WISE achieves 2–7% accuracy gains over state-of-the-art MAD approaches, demonstrating the value of heterogeneous agent roles and two-stage feedback. The work contributes a scalable MAD architecture for multimodal tasks and a probabilistic aggregation method that improves robustness, with implications for zero-shot reasoning and ensemble design in multimodal AI systems.

Abstract

Recent large language models (LLMs) are trained on diverse corpora and tasks, leading them to develop complementary strengths. Multi-agent debate (MAD) has emerged as a popular way to leverage these strengths for robust reasoning, though it has mostly been applied to language-only tasks, leaving its efficacy on multimodal problems underexplored. In this paper, we study MAD for solving vision-and-language reasoning problems. Our setup enables generalizing the debate protocol with heterogeneous experts that possess single- and multi-modal capabilities. To this end, we present Weighted Iterative Society-of-Experts (WISE), a generalized and modular MAD framework that partitions the agents into Solvers, that generate solutions, and Reflectors, that verify correctness, assign weights, and provide natural language feedback. To aggregate the agents' solutions across debate rounds, while accounting for variance in their responses and the feedback weights, we present a modified Dawid-Skene algorithm for post-processing that integrates our two-stage debate model. We evaluate WISE on SMART-840, VisualPuzzles, EvoChart-QA, and a new SMART-840++ dataset with programmatically generated problem instances of controlled difficulty. Our results show that WISE consistently improves accuracy by 2-7% over the state-of-the-art MAD setups and aggregation methods across diverse multimodal tasks and LLM configurations.

WISE: Weighted Iterative Society-of-Experts for Robust Multimodal Multi-Agent Debate

TL;DR

Abstract

WISE: Weighted Iterative Society-of-Experts for Robust Multimodal Multi-Agent Debate

TL;DR

Abstract

Paper Structure

Table of Contents

Figures (17)