CausalGDP: Causality-Guided Diffusion Policies for Reinforcement Learning

Xiaofeng Xiao; Xiao Hu; Yang Ye; Xubo Yue

CausalGDP: Causality-Guided Diffusion Policies for Reinforcement Learning

Xiaofeng Xiao, Xiao Hu, Yang Ye, Xubo Yue

TL;DR

CausalGDP introduces causal reasoning into diffusion-based reinforcement learning by learning an offline causal dynamical model and a base diffusion policy, then continuously updating causal masks to guide action generation in real time. The framework intervenes on action components via the do-operator and integrates these causal signals into the diffusion scoring and sampling process, yielding a causality-guided diffusion policy with theoretical stability and performance guarantees. Empirically, CausalGDP demonstrates competitive or superior performance across diverse, high-dimensional tasks (e.g., Maze2D, AntMaze, Humanoid) and shows robustness to different causal-discovery methods. This work highlights the practical impact of incorporating explicit causal structure into diffusion RL, enabling more efficient learning in complex control settings and paving the way for broader causal-guided policy design.

Abstract

Reinforcement learning (RL) has achieved remarkable success in a wide range of sequential decision-making problems. Recent diffusion-based policies further improve RL by modeling complex, high-dimensional action distributions. However, existing diffusion policies primarily rely on statistical associations and fail to explicitly account for causal relationships among states, actions, and rewards, limiting their ability to identify which action components truly cause high returns. In this paper, we propose Causality-guided Diffusion Policy (CausalGDP), a unified framework that integrates causal reasoning into diffusion-based RL. CausalGDP first learns a base diffusion policy and an initial causal dynamical model from offline data, capturing causal dependencies among states, actions, and rewards. During real-time interaction, the causal information is continuously updated and incorporated as a guidance signal to steer the diffusion process toward actions that causally influence future states and rewards. By explicitly considering causality beyond association, CausalGDP focuses policy optimization on action components that genuinely drive performance improvements. Experimental results demonstrate that CausalGDP consistently achieves competitive or superior performance over state-of-the-art diffusion-based and offline RL methods, especially in complex, high-dimensional control tasks.

CausalGDP: Causality-Guided Diffusion Policies for Reinforcement Learning

TL;DR

Abstract

Paper Structure (24 sections, 4 theorems, 72 equations, 2 figures, 4 tables, 1 algorithm)

This paper contains 24 sections, 4 theorems, 72 equations, 2 figures, 4 tables, 1 algorithm.

Introduction
Literature Review
Causal Reinforcement Learning
Diffusion Policy
Preliminary Setting
Diffusion Model
Diffusion Policy with Reward Guidance
Causal RL
Methodology
Causal dynamical models
Real-time Causality-Guided Diffusion Policy
Theoretical analysis
Step Size Stability
Performance Difference
Gradient of Guidance
...and 9 more sections

Key Result

Proposition 1

Consider the causal-guided reverse diffusion SDE eq:diffSDE under Assumption assump:lipschitz, the explicit Euler discretizations $a_{n+1}=a_n+b_{\text{guided }}\left(a_n, t_n\right) \Delta t+g\left(t_n\right) \Delta \bar{w}_n$ is stable as the time step $\Delta t$ satisfies $\Delta t \leq \frac{\de Under this step-size condition, the Euler solution remains stable and the discretization error stay

Figures (2)

Figure 1: Causality and Association illustration
Figure 2: Performance comparison of CausalGDP on Gym MuJoCo tasks. Learning curves of average episodic return on HalfCheetah-v4 and Humanoid-v4, comparing CausalGDP with Diffusion-QL over training steps.

Theorems & Definitions (8)

Proposition 1
proof : Proof of Proposition 1
Theorem 1
proof : Proof of Theorem 1
Proposition 2
proof
Lemma 1
proof

CausalGDP: Causality-Guided Diffusion Policies for Reinforcement Learning

TL;DR

Abstract

CausalGDP: Causality-Guided Diffusion Policies for Reinforcement Learning

Authors

TL;DR

Abstract

Table of Contents

Key Result

Figures (2)

Theorems & Definitions (8)