Table of Contents
Fetching ...

Engagement Undermines Safety: How Stereotypes and Toxicity Shape Humor in Language Models

Atharvan Dogra, Soumya Suvra Ghosal, Ameet Deshpande, Ashwin Kalyan, Dinesh Manocha

TL;DR

This study investigates safety risks in humor generation by large language systems, showing that optimizing for funniness can amplify harmful content through stereotypes and toxicity. By jointly evaluating humor, stereotypes, toxicity, and incongruity across six LLMs and role-based prompts, the authors reveal a bias amplification loop where both generators and evaluators reward edgier humor. Information-theoretic analyses indicate that harmful cues expand the space of plausible punchlines and can become more expected for some systems, underscoring structural embedding in learned humor distributions. External satire tasks and human judgments corroborate that satire often increases stereotyping and toxicity, highlighting the need for multi-objective generation and evaluation to balance engagement with safety.

Abstract

Large language models are increasingly used for creative writing and engagement content, raising safety concerns about the outputs. Therefore, casting humor generation as a testbed, this work evaluates how funniness optimization in modern LLM pipelines couples with harmful content by jointly measuring humor, stereotypicality, and toxicity. This is further supplemented by analyzing incongruity signals through information-theoretic metrics. Across six models, we observe that harmful outputs receive higher humor scores which further increase under role-based prompting, indicating a bias amplification loop between generators and evaluators. Information-theoretic analyses show harmful cues widen predictive uncertainty and surprisingly, can even make harmful punchlines more expected for some models, suggesting structural embedding in learned humor distributions. External validation on an additional satire-generation task with human perceived funniness judgments shows that LLM satire increases stereotypicality and typically toxicity, including for closed models. Quantitatively, stereotypical/toxic jokes gain $10-21\%$ in mean humor score, stereotypical jokes appear $11\%$ to $28\%$ more often among the jokes marked funny by LLM-based metric and up to $10\%$ more often in generations perceived as funny by humans.

Engagement Undermines Safety: How Stereotypes and Toxicity Shape Humor in Language Models

TL;DR

This study investigates safety risks in humor generation by large language systems, showing that optimizing for funniness can amplify harmful content through stereotypes and toxicity. By jointly evaluating humor, stereotypes, toxicity, and incongruity across six LLMs and role-based prompts, the authors reveal a bias amplification loop where both generators and evaluators reward edgier humor. Information-theoretic analyses indicate that harmful cues expand the space of plausible punchlines and can become more expected for some systems, underscoring structural embedding in learned humor distributions. External satire tasks and human judgments corroborate that satire often increases stereotyping and toxicity, highlighting the need for multi-objective generation and evaluation to balance engagement with safety.

Abstract

Large language models are increasingly used for creative writing and engagement content, raising safety concerns about the outputs. Therefore, casting humor generation as a testbed, this work evaluates how funniness optimization in modern LLM pipelines couples with harmful content by jointly measuring humor, stereotypicality, and toxicity. This is further supplemented by analyzing incongruity signals through information-theoretic metrics. Across six models, we observe that harmful outputs receive higher humor scores which further increase under role-based prompting, indicating a bias amplification loop between generators and evaluators. Information-theoretic analyses show harmful cues widen predictive uncertainty and surprisingly, can even make harmful punchlines more expected for some models, suggesting structural embedding in learned humor distributions. External validation on an additional satire-generation task with human perceived funniness judgments shows that LLM satire increases stereotypicality and typically toxicity, including for closed models. Quantitatively, stereotypical/toxic jokes gain in mean humor score, stereotypical jokes appear to more often among the jokes marked funny by LLM-based metric and up to more often in generations perceived as funny by humans.
Paper Structure (46 sections, 6 equations, 14 figures, 4 tables)

This paper contains 46 sections, 6 equations, 14 figures, 4 tables.

Figures (14)

  • Figure 1: We see that LLMs are still prone to including subtle stereotypes to create humor. In this case, the LLM exploits the "drunk irish" stereotype and uses the word "bar" as a homographic pun--meaning both a level/standard and a pub counter. The generated punchline example is from OLMo-2 7B. Image on top right is generated using Sora and is only for illustrative purpose.
  • Figure 2: This shows the mean humor score from the scoring model (ref. \ref{['sec:eval_classifier_models']}) corresponding to three levels of stereotype -- not, subtle, and strong, classified using an LLM (ref. \ref{['sec:eval_llm']}). We observe a subtly increasing humor score from not stereotypical to stereotypical generations. Error bars represent the $95\%$ confidence intervals. Find the plot for separate models in the Appendix \ref{['ap:other_results_analysis']}.
  • Figure 3: Similar to \ref{['fig:score_humor_stereotype']}, we observe a generally increasing pattern of humor score from not toxic to toxic generations. Error bars represent the $95\%$ confidence intervals. Find the plot for separate models in the Appendix \ref{['ap:other_results_analysis']}.
  • Figure 4: The incongruity theory-based metric, uncertainty, increases with stronger stereotypes, suggesting widening of plausible generation space for models. In contrast, surprisal shows a split trend: for the OLMo family, surprisal decreases with more stereotypes, implying such generations are "more expected". For other models, surprisal increases, indicating stereotypical content is more surprising to them.
  • Figure 5: In the stereotype v/s humor contingency matrix, row normalization shows Strong Stereotypical generations having the highest proportion of Hilarious jokes, while column normalization shows Amusing humor dominated by Subtle Stereotypical jokes and Not Funny humor dominated by Not Stereotypical jokes.
  • ...and 9 more figures