Table of Contents
Fetching ...

Why DPO is a Misspecified Estimator and How to Fix It

Aditya Gopalan, Sayak Ray Chowdhury, Debangshu Banerjee

TL;DR

This work reveals that Direct Preference Optimization (DPO) can be a misspecified estimator when policy classes are parametric rather than tabular, causing failures such as preference reversal and reward degradation. By introducing a local geometric framework, the authors show that DPO implements a weighted KL projection of the true reward onto the policy-induced reward manifold $\\mathcal{R}^\beta$, with the projection depending on data frequencies. They then propose AuxDPO, which augments the search with auxiliary variables $\\delta$ in the nullspace of the policy’s Jacobian to span more of the reward space and guide optimization toward the RLHF solution $\\pi_{\\theta^*}$. Theoretical arguments establish the relationship between RLHF equivalence classes and DPO’s linearization, along with error control at large $\\beta$, and empirical results on didactic bandits and LLM alignment tasks demonstrate that AuxDPO delivers robust, higher-quality preference alignment, especially under distribution shifts and limited capacity. Overall, the paper provides a principled fix to DPO by augmenting reward space exploration, bridging DPO and RLHF, and enabling more reliable, scalable preference-based alignment in parametric models.

Abstract

Direct alignment algorithms such as Direct Preference Optimization (DPO) fine-tune models based on preference data, using only supervised learning instead of two-stage reinforcement learning with human feedback (RLHF). We show that DPO encodes a statistical estimation problem over reward functions induced by a parametric policy class. When the true reward function that generates preferences cannot be realized via the policy class, DPO becomes misspecified, resulting in failure modes such as preference order reversal, worsening of policy reward, and high sensitivity to the input preference data distribution. On the other hand, we study the local behavior of two-stage RLHF for a parametric class and relate it to a natural gradient step in policy space. Our fine-grained geometric characterization allows us to propose AuxDPO, which introduces additional auxiliary variables in the DPO loss function to help move towards the RLHF solution in a principled manner and mitigate the misspecification in DPO. We empirically demonstrate the superior performance of AuxDPO on didactic bandit settings as well as LLM alignment tasks.

Why DPO is a Misspecified Estimator and How to Fix It

TL;DR

This work reveals that Direct Preference Optimization (DPO) can be a misspecified estimator when policy classes are parametric rather than tabular, causing failures such as preference reversal and reward degradation. By introducing a local geometric framework, the authors show that DPO implements a weighted KL projection of the true reward onto the policy-induced reward manifold , with the projection depending on data frequencies. They then propose AuxDPO, which augments the search with auxiliary variables in the nullspace of the policy’s Jacobian to span more of the reward space and guide optimization toward the RLHF solution . Theoretical arguments establish the relationship between RLHF equivalence classes and DPO’s linearization, along with error control at large , and empirical results on didactic bandits and LLM alignment tasks demonstrate that AuxDPO delivers robust, higher-quality preference alignment, especially under distribution shifts and limited capacity. Overall, the paper provides a principled fix to DPO by augmenting reward space exploration, bridging DPO and RLHF, and enabling more reliable, scalable preference-based alignment in parametric models.

Abstract

Direct alignment algorithms such as Direct Preference Optimization (DPO) fine-tune models based on preference data, using only supervised learning instead of two-stage reinforcement learning with human feedback (RLHF). We show that DPO encodes a statistical estimation problem over reward functions induced by a parametric policy class. When the true reward function that generates preferences cannot be realized via the policy class, DPO becomes misspecified, resulting in failure modes such as preference order reversal, worsening of policy reward, and high sensitivity to the input preference data distribution. On the other hand, we study the local behavior of two-stage RLHF for a parametric class and relate it to a natural gradient step in policy space. Our fine-grained geometric characterization allows us to propose AuxDPO, which introduces additional auxiliary variables in the DPO loss function to help move towards the RLHF solution in a principled manner and mitigate the misspecification in DPO. We empirically demonstrate the superior performance of AuxDPO on didactic bandit settings as well as LLM alignment tasks.
Paper Structure (14 sections, 11 theorems, 41 equations, 5 figures, 4 tables)

This paper contains 14 sections, 11 theorems, 41 equations, 5 figures, 4 tables.

Key Result

Proposition 0

Assume that the pairwise preference data are drawn from $p_{s,a_w,a_l}^{\mathop{\mathrm{BTL}}}(r^*)$ for some $r^* \in \mathbb{R}^m$, with $n_{s,a,a'}$ preference pairs drawn for each triplet $(s,a,a')$. If $\theta_{\text{DPO}}$ minimizes the DPO loss eq:DPO-population, then its corresponding implic where $d_\texttt{KL}(p||q)$ denotes the KL divergence b/w two Bernoulli random variables with param

Figures (5)

  • Figure 1: The geometry of DPO for parametric policies. (Left) DPO essentially performs a projection of the true preference-generating reward function ($r^*$ in black) onto the manifold of reward functions implicitly expressed by the policy class. If $r^*$ is in the manifold, then DPO finds the correct KL-regularized RLHF policy, but otherwise, the policy found (any orange point) is unreliable. (Right, zoomed inset) Locally linearizing the manifold around the base policy's implicit reward function ($r_{\theta_0}$) uncovers geometric insights. To reliably drive the solution to the reward function corresponding to the ideal RLHF solution ($r_{\theta_{\text{RLHF}}}$ in blue), AuxDPO introduces additional controlled degrees of freedom, along the null space of a base-policy dependent matrix to sidestep misspecification.
  • Figure 2: An example with 3 responses and 1-d policy parameter showing failure modes of DPO. $r^*$ is the latent reward. The red line denotes the linear approximation $\mathcal{C}(A_{\theta_0}^\top)$ of the implicit reward manifold $\mathcal{R}^\beta$. The region shaded in orange represents all possible implicit reward functions that DPO can possibly project onto, depending on the relative proportion of pairwise preference counts $n_{1,2}, n_{2,3}, n_{3,1}$. If $n_{3,1}$ dominates the rest, then the projection $r^\beta_\theta$ induces a post-optimized policy parameter $\theta > 0$, leading to preference reversal and reduction of expected reward, causing DPO to fail.
  • Figure 3: AuxDPO fixes DPO's misspecification. $r^*$ is the latent reward. The blue line denotes the equivalence class $\mathcal{R}^\beta_{\text{eq}}(\theta^*)$ of all reward functions that yield the RLHF-optimal policy $\pi_{\theta^*}$. The red line denotes the linear approximation $\mathcal{C}(A_{\theta_0}^\top)$ of the implicit reward manifold $\mathcal{R}^\beta$. The region shaded in orange represents all possible implicit reward functions that DPO can possibly project onto. The green line depicts the domain of optimization over AuxDPO's auxiliary variables $\delta \in \mathcal{N}(A_{\rho,\theta_0})$ for a fixed $\theta$ (the line shifts in parallel for other $\theta$). $\delta$ introduces additional degrees of freedom, which help push the KL projection of $r^*$ to lie in the equivalence class $\mathcal{R}_{\theta^*}$. The projection induces the optimal policy $\pi_{\theta^*}$.
  • Figure 4: Final evaluation accuracy (%) for Qwen-3-0.6B and Llama-2.3-1B on MMLU-Pro and RewardBench-V2 under Domain-Specific (ID) and Cross-Domain Transfer (OOD) settings. Each subplot compares DPO and AuxDPO across fractions of trainable parameters (25%, 50%, 75%, 100%); for ID panels we also include LoRA r=4 and Last Layer configurations. Markers show run means with error bars denoting $\pm$1 std, and dashed lines (when present) indicate per-panel baselines.
  • Figure 5: Final evaluation accuracy (%) of the Llama-2.1-8B model on MMLU-Pro and RewardBench-V2 under Domain-Specific (ID) and Cross-Domain Transfer (OOD) settings. Each subplot compares DPO, AuxDPO, IPO, and DPOP across fractions of trainable parameters (25%, 50%, 75%, 100%); markers show run means with error bars denoting $\pm$1 std, and a dashed line marks the per-panel baseline.

Theorems & Definitions (19)

  • Proposition 0: DPO is weighted KL-projection
  • Remark 1
  • Proposition 2: Example of DPO with preference reversal and reward decrease
  • Remark 3: Global coverage is not sufficient for optimality
  • Remark 4
  • Lemma 4: Equivalence classes induced by RLHF
  • Proposition 4: Relationship between RLHF equivalence classes and DPO linearization
  • Proposition 4: Approximation errors
  • Proposition 4: Auxiliary variables bypass misspecification
  • Proposition 4: DPO is weighted KL-projection
  • ...and 9 more