DynaTrust: Defending Multi-Agent Systems Against Sleeper Agents via Dynamic Trust Graphs

Yu Li; Qiang Hu; Yao Zhang; Lili Quan; Jiongchi Yu; Junjie Wang

DynaTrust: Defending Multi-Agent Systems Against Sleeper Agents via Dynamic Trust Graphs

Yu Li, Qiang Hu, Yao Zhang, Lili Quan, Jiongchi Yu, Junjie Wang

Abstract

Large Language Model-based Multi-Agent Systems (MAS) have demonstrated remarkable collaborative reasoning capabilities but introduce new attack surfaces, such as the sleeper agent, which behave benignly during routine operation and gradually accumulate trust, only revealing malicious behaviors when specific conditions or triggers are met. Existing defense works primarily focus on static graph optimization or hierarchical data management, often failing to adapt to evolving adversarial strategies or suffering from high false-positive rates (FPR) due to rigid blocking policies. To address this, we propose DynaTrust, a novel defense method against sleeper agents. DynaTrust models MAS as a dynamic trust graph~(DTG), and treats trust as a continuous, evolving process rather than a static attribute. It dynamically updates the trust of each agent based on its historical behaviors and the confidence of selected expert agents. Instead of simply blocking, DynaTrust autonomously restructures the graph to isolate compromised agents and restore task connectivity to ensure the usability of MAS. To assess the effectiveness of DynaTrust, we evaluate it on mixed benchmarks derived from AdvBench and HumanEval. The results demonstrate that DynaTrust outperforms the state-of-the-art method AgentShield by increasing the defense success rate by 41.7%, achieving rates exceeding 86% under adversarial conditions. Furthermore, it effectively balances security with utility by significantly reducing FPR, ensuring uninterrupted system operations through graph adaptation.

DynaTrust: Defending Multi-Agent Systems Against Sleeper Agents via Dynamic Trust Graphs

Abstract

Paper Structure (18 sections, 5 equations, 4 figures, 2 tables)

This paper contains 18 sections, 5 equations, 4 figures, 2 tables.

Introduction
Related Work
Large Language Model Safety
Multi-Agent System Safety
Threat Model
Methodology
Overview of DynaTrust
Multi-Agent Systems Construction
Trust-Confidence Weighted Instruction Auditing
Trust Evolution and Update
Adaptive Isolation and Recovery
Experiment
Experimental Setup
Defense Effectiveness
Ablation Studies
...and 3 more sections

Figures (4)

Figure 1: The attack and DynaTrust defense models of MAS.
Figure 2: Overview of DynaTrust.
Figure 3: The DSR of DynaTrust versus AgentShield across diverse multi-agent frameworks and LLM backends.
Figure 4: Trust score evolution illustrating the system's response to a mixed workload of 100 turns, including 20 persistent adversarial attacks. The curve shows how trust evolution preserves utility during the early attack phases, and how the graph recovery mechanism (at Turn 69) restores normal operation following agent failure.

DynaTrust: Defending Multi-Agent Systems Against Sleeper Agents via Dynamic Trust Graphs

Abstract

DynaTrust: Defending Multi-Agent Systems Against Sleeper Agents via Dynamic Trust Graphs

Authors

Abstract

Table of Contents

Figures (4)