Aligning Compound AI Systems via System-level DPO

Xiangwen Wang; Yibo Jacky Zhang; Zhoujie Ding; Katherine Tsai; Haolun Wu; Sanmi Koyejo

Aligning Compound AI Systems via System-level DPO

Xiangwen Wang, Yibo Jacky Zhang, Zhoujie Ding, Katherine Tsai, Haolun Wu, Sanmi Koyejo

TL;DR

The paper tackles the challenge of aligning compound AI systems by modeling them as DAGs and extending Direct Preference Optimization to system-level preferences. It introduces SysDPO with two variants (SysDPO-Direct and SysDPO-Sampling) to handle cases with or without observable intermediate outputs, and provides theoretical guarantees for beta-perfect alignment. Empirically, it validates the approach on a joint LLM+diffusion task and a two-LLM collaboration system, showing that holistic, joint alignment significantly outperforms component-wise or prompting-only baselines. The work lays a foundation for scalable, reliable coordination of complex multi-component AI workflows and outlines avenues for future efficiency and scalability improvements.

Abstract

Compound AI systems, comprising multiple interacting components such as LLMs, foundation models, and external tools, have demonstrated remarkable improvements compared to single models in various tasks. To ensure their effective deployment in real-world applications, aligning these systems with human preferences is crucial. However, aligning the compound system via policy optimization, unlike the alignment of a single model, is challenging for two main reasons: (i) non-differentiable interactions between components make end-to-end gradient-based optimization method inapplicable, and (ii) system-level preferences cannot be directly transformed into component-level preferences. To address these challenges, we first formulate compound AI systems as Directed Acyclic Graphs (DAGs), explicitly modeling both component interactions and the associated data flows. Building on this formulation, we introduce $\textbf{SysDPO}$, a framework that extends Direct Preference Optimization (DPO) to enable joint system-level alignment. We propose two variants, SysDPO-Direct and SysDPO-Sampling, tailored for scenarios depending on whether we construct a system-specific preference dataset. We empirically demonstrate the effectiveness of our approach across two applications: the joint alignment of a language model and a diffusion model, and the joint alignment of an LLM collaboration system.

Aligning Compound AI Systems via System-level DPO

TL;DR

Abstract

Aligning Compound AI Systems via System-level DPO

TL;DR

Abstract

Paper Structure

Table of Contents

Key Result

Figures (8)

Theorems & Definitions (5)