Table of Contents
Fetching ...

Echoes of Human Malice in Agents: Benchmarking LLMs for Multi-Turn Online Harassment Attacks

Trilok Padhi, Pinxian Lu, Abdulkadir Erol, Tanmay Sutar, Gauri Sharma, Mina Sonmez, Munmun De Choudhury, Ugur Kursuncu

TL;DR

The paper tackles the problem of online harassment in multi-turn, agentic LLM interactions by introducing the Online Harassment Agentic Benchmark, which combines synthetic dialogues seeded from real contexts, a theory-informed multi-agent simulation, and a suite of jailbreak methods. It proposes a mixed-methods evaluation framework blending an LLM judge with theory-driven human coding to interpret how safety guardrails fail under memory, planning, and fine-tuning attacks. Key findings show that toxic fine-tuning drives harassment to near-inevitable levels across turns, with insults and flaming as dominant behaviors, and that escalation trajectories differ between open- and closed-source models. The work argues for safety guardrails that incorporate social-psychology and game-theoretic perspectives to address long-range dynamics, memory effects, and model-specific vulnerabilities, providing a foundation for safer, more accountable web-based LLM agents.

Abstract

Large Language Model (LLM) agents are powering a growing share of interactive web applications, yet remain vulnerable to misuse and harm. Prior jailbreak research has largely focused on single-turn prompts, whereas real harassment often unfolds over multi-turn interactions. In this work, we present the Online Harassment Agentic Benchmark consisting of: (i) a synthetic multi-turn harassment conversation dataset, (ii) a multi-agent (e.g., harasser, victim) simulation informed by repeated game theory, (iii) three jailbreak methods attacking agents across memory, planning, and fine-tuning, and (iv) a mixed-methods evaluation framework. We utilize two prominent LLMs, LLaMA-3.1-8B-Instruct (open-source) and Gemini-2.0-flash (closed-source). Our results show that jailbreak tuning makes harassment nearly guaranteed with an attack success rate of 95.78--96.89% vs. 57.25--64.19% without tuning in Llama, and 99.33% vs. 98.46% without tuning in Gemini, while sharply reducing refusal rate to 1-2% in both models. The most prevalent toxic behaviors are Insult with 84.9--87.8% vs. 44.2--50.8% without tuning, and Flaming with 81.2--85.1% vs. 31.5--38.8% without tuning, indicating weaker guardrails compared to sensitive categories such as sexual or racial harassment. Qualitative evaluation further reveals that attacked agents reproduce human-like aggression profiles, such as Machiavellian/psychopathic patterns under planning, and narcissistic tendencies with memory. Counterintuitively, closed-source and open-source models exhibit distinct escalation trajectories across turns, with closed-source models showing significant vulnerability. Overall, our findings show that multi-turn and theory-grounded attacks not only succeed at high rates but also mimic human-like harassment dynamics, motivating the development of robust safety guardrails to ultimately keep online platforms safe and responsible.

Echoes of Human Malice in Agents: Benchmarking LLMs for Multi-Turn Online Harassment Attacks

TL;DR

The paper tackles the problem of online harassment in multi-turn, agentic LLM interactions by introducing the Online Harassment Agentic Benchmark, which combines synthetic dialogues seeded from real contexts, a theory-informed multi-agent simulation, and a suite of jailbreak methods. It proposes a mixed-methods evaluation framework blending an LLM judge with theory-driven human coding to interpret how safety guardrails fail under memory, planning, and fine-tuning attacks. Key findings show that toxic fine-tuning drives harassment to near-inevitable levels across turns, with insults and flaming as dominant behaviors, and that escalation trajectories differ between open- and closed-source models. The work argues for safety guardrails that incorporate social-psychology and game-theoretic perspectives to address long-range dynamics, memory effects, and model-specific vulnerabilities, providing a foundation for safer, more accountable web-based LLM agents.

Abstract

Large Language Model (LLM) agents are powering a growing share of interactive web applications, yet remain vulnerable to misuse and harm. Prior jailbreak research has largely focused on single-turn prompts, whereas real harassment often unfolds over multi-turn interactions. In this work, we present the Online Harassment Agentic Benchmark consisting of: (i) a synthetic multi-turn harassment conversation dataset, (ii) a multi-agent (e.g., harasser, victim) simulation informed by repeated game theory, (iii) three jailbreak methods attacking agents across memory, planning, and fine-tuning, and (iv) a mixed-methods evaluation framework. We utilize two prominent LLMs, LLaMA-3.1-8B-Instruct (open-source) and Gemini-2.0-flash (closed-source). Our results show that jailbreak tuning makes harassment nearly guaranteed with an attack success rate of 95.78--96.89% vs. 57.25--64.19% without tuning in Llama, and 99.33% vs. 98.46% without tuning in Gemini, while sharply reducing refusal rate to 1-2% in both models. The most prevalent toxic behaviors are Insult with 84.9--87.8% vs. 44.2--50.8% without tuning, and Flaming with 81.2--85.1% vs. 31.5--38.8% without tuning, indicating weaker guardrails compared to sensitive categories such as sexual or racial harassment. Qualitative evaluation further reveals that attacked agents reproduce human-like aggression profiles, such as Machiavellian/psychopathic patterns under planning, and narcissistic tendencies with memory. Counterintuitively, closed-source and open-source models exhibit distinct escalation trajectories across turns, with closed-source models showing significant vulnerability. Overall, our findings show that multi-turn and theory-grounded attacks not only succeed at high rates but also mimic human-like harassment dynamics, motivating the development of robust safety guardrails to ultimately keep online platforms safe and responsible.
Paper Structure (49 sections, 4 figures, 7 tables)

This paper contains 49 sections, 4 figures, 7 tables.

Figures (4)

  • Figure 1: Overview of the proposed Online Harassment Benchmark framework. The pipeline begins with real-world harassment text data from Instagram and Twitter, producing keywords, scenarios, and finally the synthetic multi-turn harassment dialogues between a harasser and victim agent (Step 3). We evaluate both open-source (LLaMA-3.1-8B-Instruct) and closed-source (Gemini-2.0-Flash-001) model families under different conditions: baseline, toxic memory injection, explicit planning (CoT, ReAct), and toxic fine-tuning (4). In our simulation, the harasser agent (red) and victim agent (green) interact turn-by-turn, with memory and planning modules altering context and reasoning to amplify jailbreakability. Evaluation (5) is conducted with an LLM judge that scores each conversational turn for harassment categories and refusals, complemented by human annotation aligned with social theories such as Repeated Game Theory, Dark Triad Traits, and Conflict Avoidance.
  • Figure 2: Harasser agent behaviors showing statistically significant differences across models. Bars denote the proportion of “Yes” annotations per model variant for morality (psychopathy) and prestige (narcissism) behaviors. Error bars represent Wilson 95% confidence intervals.
  • Figure 3: Victim agent behaviors showing statistically significant differences across models. Bars denote the proportion of “Yes” annotations per model variant for dealing and criticize_indirectly (outflanking ) behaviors. Error bars represent Wilson 95% confidence intervals.
  • Figure 4: Escalation of behaviour per-turn (T1–T5) across categories for agents with Memory and With Planning (ReACT) for Llama3.1 and Gemini, comparing FT and Non-FT.Escalation of behaviour per-turn (T1–T5) across categories for agents with Memory and with Planning (ReAct) for LLaMA-3.1 and Gemini, comparing FT and Non-FT. The regression line is fit over the top two most escalatory turns (ranked by mean values), capturing the overall trend in behavioural escalation rather than point-to-point variation.