Stop Reducing Responsibility in LLM-Powered Multi-Agent Systems to Local Alignment
Jinwei Hu, Yi Dong, Shuang Ao, Zhuoyun Li, Boxuan Wang, Lokesh Singh, Guangliang Cheng, Sarvapali D. Ramchurn, Xiaowei Huang
TL;DR
The paper addresses the gap between local agent alignment and system-wide responsibility in LLM-MAS, arguing for a lifecycle-wide, global notion of agreement that integrates subjective human values with objective verifiability. It outlines a dual-perspective framework and a four-stage human-centered lifecycle, plus a human-AI meta-governance layer, to manage agreement, uncertainty, and safety threats across design, development, deployment, and maintenance. Concrete mechanisms discussed include agent-to-human and agent-to-agent alignment techniques (e.g., RLHF, SFT, self-improvement, cross-model distillation, debate), uncertainty quantification methods (memory/RAG, conformal approaches, runtime provenance), and runtime safety measures (provenance, unlearning, neural-symbolic guards). The proposed paradigm aims to deliver ethically aligned, verifiably coherent, and resilient LLM-MAS behavior in open, uncertain environments, with auditable oversight and interdisciplinary guidance guiding governance across the lifecycle.
Abstract
LLM-powered Multi-Agent Systems (LLM-MAS) unlock new potentials in distributed reasoning, collaboration, and task generalization but also introduce additional risks due to unguaranteed agreement, cascading uncertainty, and adversarial vulnerabilities. We argue that ensuring responsible behavior in such systems requires a paradigm shift: from local, superficial agent-level alignment to global, systemic agreement. We conceptualize responsibility not as a static constraint but as a lifecycle-wide property encompassing agreement, uncertainty, and security, each requiring the complementary integration of subjective human-centered values and objective verifiability. Furthermore, a dual-perspective governance framework that combines interdisciplinary design with human-AI collaborative oversight is essential for tracing and ensuring responsibility throughout the lifecycle of LLM-MAS. Our position views LLM-MAS not as loose collections of agents, but as unified, dynamic socio-technical systems that demand principled mechanisms to support each dimension of responsibility and enable ethically aligned, verifiably coherent, and resilient behavior for sustained, system-wide agreement.
